tokens

46 stories and discussions about tokens, aggregated from every source we track.

1.

tokens are going to be as cheap as electricity within the decade

346 points•teoruiz•8 days ago•225 comments•
2.

Analysis of Inception's Mercury 2.5 and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.

138 points•Retro_Dev•7 days ago•82 comments•
3.

tokens are going to be as cheap as electricity within the decade

32 points•op•8 days ago•14 comments
4.

A Kaggle benchmark of the step agents rarely test: counting what a tool returns. Ten models, 68 questions, one tool that returns the count and one that returns the rows. With the count, every model is right at a flat cost. With the rows, models that reason through the list count 330 ids right and spend 6 to 26 times the tokens doing it; models that answer straight away get 0 to 10 of 21.

15 points•xbill•2 days ago•5 comments
5.

WavexAI gives you unlimited AI usage for one flat price. No caps, no surprises.

6 points•wilsprouse•8 days ago•11 comments•
6.

Every Claude Code session starts the same way. You say “run the tests.” The agent says “I’ll look for how to run tests in this project.”

5 points•runtooldev•5 days ago•0 comments•
7.
5 points•mtokarski•13 days ago•8 comments•
8.

I spent a weekend benchmarking llama.cpp on my Xiaomi Book Pro 14. The machine has Intel's Core Ultra X7 358H and Arc B390 integrated graphics, with 30 GiB of unified memory. That last part matters. It means the GPU…

4 points•grigio•3 days ago•0 comments•
9.

Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

4 points•charles_irl•4 days ago•1 comment•
10.

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

4 points•birdculture•4 days ago•0 comments•
12.

Open-source AI agent decision layer for Codex, Claude Code, Hermes, Antigravity and VS Code. TypeSafe Jev MCP routing, 38 recipes, local gates and receipts; optional Laya-MLX on Apple Silicon. - qu...

3 points•vpbhardwaj•2 days ago•0 comments•
13.
3 points•yesitcan•8 days ago•4 comments•
14.

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

2 points•birdculture•about 3 hours ago•0 comments•
15.

Why Playwright MCP burns tokens and loses context on multi-step tasks, how to tune it, and benchmark results versus Stagehand: 32k vs 97k tokens per task.

2 points•smpandya•about 10 hours ago•0 comments•
16.

VRAM for LLMs is a bandwidth problem: every token streams the whole model from memory. Bandwidth per tier, the 20x offload cliff, and what fits in 16, 24 or 48 GB.

2 points•axrisi•about 20 hours ago•3 comments
17.
2 points•mindracer•6 days ago•0 comments•
18.

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

2 points•birdculture•6 days ago•0 comments•
19.
2 points•gmays•6 days ago•0 comments•
20.

Eight tools on a machine with 19,195 sessions already on disk, asked 100 questions only that history answers: build time, first answer, hit@1.

2 points•vshulcz•6 days ago•0 comments•
21.

A lognormal estimate of OpenAI researcher token usage at API prices: 1,125 assumed researchers and $4.24 million per day. Assumptions, equations, and three charts.

2 points•abtinf•7 days ago•1 comment•
22.

A PR review last Thursday morning triggered this thought. A developer had written a 40-line utility...

2 points•alexandersstudi•7 days ago•1 comment
23.

Tool definitions are billed on every turn, and past roughly thirty tools they start costing you accuracy rather than just money.

2 points•charrington•8 days ago•0 comments•
24.

tokens are going to be as cheap as electricity within the decade

2 points•benjaminclauss•8 days ago•0 comments•
25.
1 points•wincy•about 20 hours ago•2 comments•
26.

A visual walkthrough of what happens inside AI coding tools, built up layer by layer from a basic prompt to a full agent loop.

1 points•fagnerbrack•1 day ago•0 comments•
27.

TensorFold's new inference engine delivers decode speeds over 62 tokens per second on a single stream and 119 across five, with prefill at 2,500 tokens per second and a 256k context window. It outperforms prior vLLM…

1 points•soltanov•1 day ago•0 comments•
28.

The first frontier LLM with ten million tokens of context.

1 points•ClintEhrlich•1 day ago•0 comments•
29.
1 points•sastra•2 days ago•1 comment•
30.

How fast can your fingers generate tokens? A tiny tokens-per-second benchmark for humans.

1 points•cmrdporcupine•2 days ago•0 comments•
31.

What each step really saved, measured on a 50,000-row report, while I was building an MCP server for Google Sheets.

1 points•andrewmatic•3 days ago•0 comments•
32.

Telemetry audit, post-mortem, and failure-mode analysis of a 474K LOC codebase built with Antigravity + Gemini Flash (102.8B tokens). - Janson79jc/antigravity-100b-telemetry-audit

1 points•janson79jc•3 days ago•0 comments•
33.

Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

1 points•birdculture•4 days ago•0 comments•
34.

Claude apologizes, then hits you. Codex never runs the tests. DeepSeek distills your ult. 20 fighters, online PvP, free in your browser.

1 points•lexdoudkin•4 days ago•0 comments•
35.

An AI agent is typically given a mission: a task to pursue on a user's behalf. OAuth 2.0 issues access tokens for individual resource requests, but it has no durable, approved artifact that ties those tokens to the one…

1 points•mooreds•5 days ago•0 comments•
36.
1 points•ibobev•6 days ago•0 comments•
37.

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

1 points•charles_irl•7 days ago•0 comments•
38.

Nori LLM: the fastest large language model on the market. 1,000,000+ tokens per second. Optimized for humans and robot crawlers alike.

1 points•theahura•7 days ago•0 comments•
39.

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

1 points•kkm•7 days ago•0 comments•
40.

A framework for understanding where LLM inference goes during agentic tasks: Initialization, Reasoning, Orchestration, and Synthesis. Consider "Reasoning Yield" as a way to measure how much inference is spent resolving…

1 points•jdauriemma•8 days ago•0 comments•
41.

Why the price of an AI token in Russia can include more than model inference.

1 points•whitef0x•8 days ago•0 comments•
42.

tokens are going to be as cheap as electricity within the decade

1 points•williamcotton•8 days ago•0 comments•
43.

Know when Codex and Claude Code limits reset, so you can use your remaining tokens on something you want to build.

1 points•codeclimber•9 days ago•0 comments•
44.

Capture, review, and share a supported web-app state, then reopen the State Card against a running local app and current code.

1 points•redvoron•9 days ago•0 comments•
45.

521 decision units, zero generated tokens. Watch a reply assemble one word at a time, with the probability behind every choice.

1 points•skillseeddev•10 days ago•1 comment•

Related topics