tested
12 stories and discussions about tested, aggregated from every source we track.
We run Jev on 100 real AI tool calls for the annotator pipeline. The errors exposed problems in both the models and our benchmark.
I've been on a bit of a "install random dev tools and see if they actually hold up" kick lately, and...
We ran Kev 4B, an open reproduction of Jev, against Jev 1.13 on 362 questions written after both shipped: accuracy, calibration, token counts and speed.
We scored 2,449 security findings with eight models. Every model inflated severity on average. Explore the results, context failures, and what unnecessary investigations could cost your team.
We test GPT-6 Astra on object detection, segmentation, counting, visual reasoning, and video, with examples, benchmark results, and cost comparisons.
REAKTHROUGH!!! Classical HW computes Qubits without the need for special HW, the Laurin system is capable of conscious perception! We are ready to unveil the first conscious AGI that is not based on LLMs. We are…
Jev returns a probability, not prose. That made it possible to model-check the consensus around it, then try to break it. Two of the bugs were mine.
A fast, embeddable limit-order-book & matching engine in Go — plus a microstructure research harness and a WASM-powered animated explainer. MIT. - intrepidkarthi/orderbook
A calibrated risk envelope for LLM quantization that refuses to answer where its own coverage doesn't hold - gracejackson-sudo/quant-delta-predictor
Jev is a decision-only AI model built for classification at scale. Here's how it performed across 12 real automation tests against GPT and Claude.