measuring
13 stories and discussions about measuring, aggregated from every source we track.
Remember when using autocomplete meant you weren't "really" programming? A real Stack Overflow...
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked AI models...
Sequential tests on a Taiwan Mobile 4G hotspot in Taipei. Download measured 55 kbit/s at 15:08 and 30 Mbit/s at 17:47; Tor and Snowflake connected throughout.
Optimizing code starts with measuring it, and a measurement is only useful if it is repeatable: a 2% improvement is invisible under 5% of noise. Yet on an …
We built an evaluation suite to assess model trustworthiness. Our results indicate that models developed from open-source models can be trusted, provided…
The weight of a television set has nothing at all to do with the clarity of its picture. Even if you measure to a tenth of a gram, this precise data is useless. Some people measure stereo equipment…
Optimizing code starts with measuring it, and a measurement is only useful if it is repeatable: a 2% improvement is invisible under 5% of noise. Yet on an …
Large language models (LLMs) are increasingly used as primary knowledge sources, yet their epistemic diversity - defined as the diversity of real-world claims in their outputs - has never been measured. Low epistemic…
Sequential tests on a Taiwan Mobile 4G hotspot in Taipei. Download measured 55 kbit/s at 15:08 and 30 Mbit/s at 17:47; Tor and Snowflake connected throughout.
CheatBench measures whether AI agents attempt to cheat when honest work is difficult. A benchmark from the Center for AI Safety.
Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characterize systematically. We develop a cross-linguistic psychometric…
Testing semantic search, filtered queries, and multi-table joins in AlloyDB to measure the true cost...