reasoning
19 stories and discussions about reasoning, aggregated from every source we track.
Jeeves – Reasoning improves Jev-like decision models - PostHog/jeeves
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked A while...
Per-step reasoning effort for Claude Code, chosen by Jev, without breaking the prompt cache. Unofficial. - ifoster01/jev-effort
Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood. Across math, coding, and science QA, we find that aggressive PTQ…
Junnan Dong, Linhao Luo, Senlei Zhang and colleagues at Tencent Youtu Lab and Monash University propose WFM, a Wiki Foundation Model for encoding and retrieving
Solving the game with reasoning, not reinforcement learning. An interactive snapshot of how models approached The Crux benchmark.
On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level…
Compare GPT-6 Sol and Luna pricing, context limits, reasoning, and use cases to choose a model for coding or high-volume applications.
Why think in words? A deep dive into Latent-GRPO, continuous thought recurrence (Coconut, SofT-GRPO, CoLaR, SLPO), policy gradients over continuous embeddings, and empirical benchmarks on Qwen3.6-27B.
Gensler's informal fallacies drill is back: new passages drawn from real arguments, more than one right answer, and a guide beside every question.
We just released Ternary Bonsai 2 27B: 5.9 GB of weights for a 27B model. How much reasoning survived compression? We test it on a few evals that I found interesting and discuss where it retains the performance of the…
Paradigma releases Limite 1B - Violetto, a 1-billion parameter transformer for high-throughput mathematical reasoning, alongside its evaluations, a training-time value model, and a custom vLLM inference plugin.
Explore Space Bunny’s SVG art and benchmarks, plus its new DeepSeek reasoning match and the Harmony/GPT-OSS identity clue.
Span-01 is the first hyper-parallel reasoning classifier built for unseen challenges, with true single-forward classification across AI traces.
Picks the reasoning effort for each turn of Claude Code with a System One classifier - totally-tim/effort-router
The safety of a tool-using language model agent is usually treated as a property of the model alone. We give controlled, full-precision evidence that it is instead a joint property of the model and the software that…
A framework for understanding where LLM inference goes during agentic tasks: Initialization, Reasoning, Orchestration, and Synthesis. Consider "Reasoning Yield" as a way to measure how much inference is spent resolving…