Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens
DEV Community·15 points·xbill·2 days ago·dev.to
A Kaggle benchmark of the step agents rarely test: counting what a tool returns. Ten models, 68 questions, one tool that returns the count and one that returns the rows. With the count, every model is right at a flat cost. With the rows, models that reason through the list count 330 ids right and spend 6 to 26 times the tokens doing it; models that answer straight away get 0 to 10 of 21.
Read the full article at dev.to →
Related stories
- The Verge · 0 points · 12 days ago
- Hacker News · 4 points · 5 days ago
- Time to Spend Tokens or Meditate?inmve.github.ioHacker News · 1 points · 9 days ago
- Hacker News · 5 points · 9 days ago
- Hacker News · 1 points · 9 days ago
- Hacker News · 3 points · 5 days ago