measuring

13 stories and discussions about measuring, aggregated from every source we track.

1.

Remember when using autocomplete meant you weren't "really" programming? A real Stack Overflow...

26 points•mikachu•2 days ago•7 comments
2.

This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked AI models...

7 points•unit_500_c36d1b1011fdf39c•6 days ago•4 comments
3.

Sequential tests on a Taiwan Mobile 4G hotspot in Taipei. Download measured 55 kbit/s at 15:08 and 30 Mbit/s at 17:47; Tor and Snowflake connected throughout.

6 points•7tehdt3cnw6kir6o•8 days ago•0 comments
4.

Optimizing code starts with measuring it, and a measurement is only useful if it is repeatable: a 2% improvement is invisible under 5% of noise. Yet on an …

4 points•dalvrosa•5 days ago•2 comments•
6.

We built an evaluation suite to assess model trustworthiness. Our results indicate that models developed from open-source models can be trusted, provided…

2 points•ronfriedhaber•7 days ago•0 comments•
7.

The weight of a television set has nothing at all to do with the clarity of its picture. Even if you measure to a tenth of a gram, this precise data is useless. Some people measure stereo equipment…

2 points•compiler-guy•9 days ago•0 comments•
8.

Optimizing code starts with measuring it, and a measurement is only useful if it is repeatable: a 2% improvement is invisible under 5% of noise. Yet on an …

1 points•ibobev•1 day ago•0 comments•
9.

Large language models (LLMs) are increasingly used as primary knowledge sources, yet their epistemic diversity - defined as the diversity of real-world claims in their outputs - has never been measured. Low epistemic…

1 points•daniel_iversen•5 days ago•0 comments•
10.

Sequential tests on a Taiwan Mobile 4G hotspot in Taipei. Download measured 55 kbit/s at 15:08 and 30 Mbit/s at 17:47; Tor and Snowflake connected throughout.

1 points•DemiGuru•5 days ago•0 comments•
11.

CheatBench measures whether AI agents attempt to cheat when honest work is difficult. A benchmark from the Center for AI Safety.

1 points•gumby•7 days ago•0 comments•
12.

Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characterize systematically. We develop a cross-linguistic psychometric…

1 points•anigbrowl•8 days ago•0 comments•
13.

Testing semantic search, filtered queries, and multi-table joins in AlloyDB to measure the true cost...

0 points•gleb_otochkin•6 days ago•0 comments

Related topics