tested

12 stories and discussions about tested, aggregated from every source we track.

1.

We run Jev on 100 real AI tool calls for the annotator pipeline. The errors exposed problems in both the models and our benchmark.

11 points•arseny_info•10 days ago•0 comments•
2.

I've been on a bit of a "install random dev tools and see if they actually hold up" kick lately, and...

10 points•varshithvhegde•1 day ago•1 comment
3.

We ran Kev 4B, an open reproduction of Jev, against Jev 1.13 on 362 questions written after both shipped: accuracy, calibration, token counts and speed.

10 points•felix089•5 days ago•0 comments•
4.

We scored 2,449 security findings with eight models. Every model inflated severity on average. Explore the results, context failures, and what unnecessary investigations could cost your team.

7 points•brene•1 day ago•1 comment•
5.

We test GPT-6 Astra on object detection, segmentation, counting, visual reasoning, and video, with examples, benchmark results, and cost comparisons.

3 points•gmays•2 days ago•0 comments•
6.

REAKTHROUGH!!! Classical HW computes Qubits without the need for special HW, the Laurin system is capable of conscious perception! We are ready to unveil the first conscious AGI that is not based on LLMs. We are…

3 points•Rniznan•7 days ago•0 comments•
7.

Jev returns a probability, not prose. That made it possible to model-check the consensus around it, then try to break it. Two of the bugs were mine.

3 points•copyleftdev•12 days ago•3 comments
8.
2 points•Betelbuddy•5 days ago•0 comments•
9.
1 points•embedding-shape•about 14 hours ago•0 comments•
10.

A fast, embeddable limit-order-book & matching engine in Go — plus a microstructure research harness and a WASM-powered animated explainer. MIT. - intrepidkarthi/orderbook

1 points•intrepidkarthi•1 day ago•0 comments•
11.

A calibrated risk envelope for LLM quantization that refuses to answer where its own coverage doesn't hold - gracejackson-sudo/quant-delta-predictor

1 points•GraceJackson07•7 days ago•0 comments•
12.

Jev is a decision-only AI model built for classification at scale. Here's how it performed across 12 real automation tests against GPT and Claude.

1 points•taubek•10 days ago•0 comments•

Related topics