benchmarked

6 stories and discussions about benchmarked, aggregated from every source we track.

1.

Many agent memory systems accumulate duplicate and contradictory memories over time. To address this,...

2 points•haoning_kan_20d7ddb19e07c•5 days ago•2 comments
2.

I benchmarked the Jev model against an LLM agent on real browser navigation: 13x faster per task, and one setup mistake that broke every single test.

1 points•speckx•2 days ago•0 comments•
3.

I work on AI evals and reliability for agentic systems. Skeptical of tooling hype, mostly because I keep measuring it.

1 points•max-t-dev•3 days ago•0 comments•
4.

91.5% recall at sub-millisecond latency from regex alone, 98.1% with a ML layer added — and the false positives that come with it. Full methodology, mapped to the OWASP Top 10 for LLM Applications.

1 points•Dikesa_27•4 days ago•0 comments•
5.

Open source runtime to migrate DRF to FastAPI (and more to come) 🚚 - sankaHQ/sanka

1 points•haegwan•8 days ago•0 comments•
6.

Benchmark harness (Metal/CUDA/CPU quantum-simulation comparisons): github.com/waratahlabs/quantum-metal-bench

1 points•sec-oops•9 days ago•0 comments•

Related topics