reasoning

19 stories and discussions about reasoning, aggregated from every source we track.

1.

Jeeves – Reasoning improves Jev-like decision models - PostHog/jeeves

241 points•nicowaltz•1 day ago•94 comments•
2.

This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked A while...

26 points•dj29•4 days ago•14 comments
3.

Per-step reasoning effort for Claude Code, chosen by Jev, without breaking the prompt cache. Unofficial. - ifoster01/jev-effort

4 points•ifoster41901•7 days ago•0 comments•
5.

Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood. Across math, coding, and science QA, we find that aggressive PTQ…

3 points•theanonymousone•3 days ago•0 comments•
6.

Junnan Dong, Linhao Luo, Senlei Zhang and colleagues at Tencent Youtu Lab and Monash University propose WFM, a Wiki Foundation Model for encoding and retrieving

3 points•omarsar•8 days ago•0 comments•
7.

Solving the game with reasoning, not reinforcement learning. An interactive snapshot of how models approached The Crux benchmark.

3 points•dustinlakin•10 days ago•1 comment•
8.

On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level…

2 points•simonpure•2 days ago•0 comments•
9.

Compare GPT-6 Sol and Luna pricing, context limits, reasoning, and use cases to choose a model for coding or high-volume applications.

2 points•flashbrew•7 days ago•0 comments•
10.

Why think in words? A deep dive into Latent-GRPO, continuous thought recurrence (Coconut, SofT-GRPO, CoLaR, SLPO), policy gradients over continuous embeddings, and empirical benchmarks on Qwen3.6-27B.

2 points•gfactor_ai•7 days ago•0 comments•
11.

Gensler's informal fallacies drill is back: new passages drawn from real arguments, more than one right answer, and a guide beside every question.

2 points•kotk•8 days ago•0 comments•
12.

We just released Ternary Bonsai 2 27B: 5.9 GB of weights for a 27B model. How much reasoning survived compression? We test it on a few evals that I found interesting and discuss where it retains the performance of the…

2 points•tosh•8 days ago•0 comments•
13.

Paradigma releases Limite 1B - Violetto, a 1-billion parameter transformer for high-throughput mathematical reasoning, alongside its evaluations, a training-time value model, and a custom vLLM inference plugin.

2 points•panecondito•8 days ago•0 comments•
14.

Explore Space Bunny’s SVG art and benchmarks, plus its new DeepSeek reasoning match and the Harmony/GPT-OSS identity clue.

1 points•devy•1 day ago•0 comments•
15.

Span-01 is the first hyper-parallel reasoning classifier built for unseen challenges, with true single-forward classification across AI traces.

1 points•nateb2022•1 day ago•0 comments•
16.

Picks the reasoning effort for each turn of Claude Code with a System One classifier - totally-tim/effort-router

1 points•totally-tim•2 days ago•1 comment•
17.
1 points•wslh•4 days ago•0 comments•
18.

The safety of a tool-using language model agent is usually treated as a property of the model alone. We give controlled, full-precision evidence that it is instead a joint property of the model and the software that…

1 points•sbulaev•6 days ago•0 comments•
19.

A framework for understanding where LLM inference goes during agentic tasks: Initialization, Reasoning, Orchestration, and Synthesis. Consider "Reasoning Yield" as a way to measure how much inference is spent resolving…

1 points•jdauriemma•7 days ago•0 comments•

Related topics