tests

51 stories and discussions about tests, aggregated from every source we track.

1.
135 points•theanonymousone•4 days ago•69 comments•
2.

TL;DR GPT-6 Astra has started another familiar AI conversation. The model is more capable, Jensen...

129 points•hemapriya_kanagala•16 days ago•44 comments
4.

The perfect app for your band practice recordings. Built by musicians, for musicians!

37 points•motet-a•2 days ago•12 comments
5.

Run this in a terminal, then run it again under script(1): python3 -c 'import sys;...

25 points•kenielzep97•2 days ago•9 comments
6.

Last month a teammate pasted a Playwright test into our PR channel and wrote "AI generated this in 4...

14 points•speaklouder•12 days ago•12 comments
7.

AI coding agents are getting very good at finishing tasks. They modify files. Fix errors. Write...

13 points•robertadam987_•4 days ago•18 comments
8.

A test named test_all_adapters_importable asserted nothing. It would pass forever, even if every...

13 points•debashish_ghosal•11 days ago•0 comments
9.

Peter Bourgon has a web site, and this is that web site.

11 points•stchris•over 5 years ago•1 comment
10.

What the independent measurements of TypeSafe's Jev found in its first eight days: arXiv preprints, GitHub evaluations and blog benchmarks, each traced to its primary source. Accuracy, calibration, speed, cost, failure modes, the prior art, the open alternatives, and what is still unmeasured.

9 points•xbill•7 days ago•1 comment
11.

Defunctionalise your continuations and your tests can run any computation a step at a time.

9 points•crowdhailer•2 months ago•10 comments
12.

Last year, a critical payment processing service went completely silent during Black Friday traffic....

6 points•tarikmostafa•6 days ago•1 comment
13.

Are generative (randomized) tests significantly more effective than example-based unit-tests at discovering bugs? There's an interesting discussion about this on lobste.rs. One argument in favor of unit tests is,…

5 points•hwayne•about 5 hours ago•2 comments
14.

Every backend developer knows the classic trade-off when writing tests: Unit tests with mocks are...

5 points•mindinu•6 days ago•0 comments
15.

My last post about Filament Studio was about v1.2.0 and multilingual content, back in April. Since...

5 points•serhii_fedorenko•8 days ago•1 comment
17.

authentik is now OpenID Certified. Here's what the conformance tests found, and why the certificate matters less than the tests behind it.

3 points•sdko•about 11 hours ago•0 comments•
18.
3 points•ngruhn•2 days ago•0 comments•
19.

Evaluating Jev's calibration on two hard datasets.

3 points•cannedbread•4 days ago•0 comments•
20.

Link to: https://pogueman.substack.com/p/125-tests-of-the-new-ai-siri

3 points•frizlab•9 days ago•0 comments•
21.

“Fornell et al. didn’t run standard multifactor asset pricing tests; didn’t use standard t-tests; and didn’t conduct out of sample tests. When you do, there is nothing there.”

3 points•ashgreat•10 days ago•1 comment•
22.

Plain-English browser e2e tests cheap enough to run on every PR. Built on Jev and Playwright, open source, bring your own key. - sedum-dev/sedum

2 points•jgnatch•about 8 hours ago•1 comment•
23.

TOKI, ‌Gifu Prefecture--Japanese startup Helical Fusion expects to ⁠begin preliminary power-on tests at its ⁠demonstration fusion reactor in 2027, it said on Tuesday, a step towards its goal of becoming a global…

2 points•BaudouinVH•about 15 hours ago•0 comments•
24.

Record AI agent runs locally and find which part of the context caused a decision. - RehanMohammed985/runtape

2 points•rehanmoin91•about 19 hours ago•0 comments•
25.

OpenAI is scrapping the release of GPT-6.1 Astra, a next-generation AI model planned for an October debut, over safety concerns raised by researchers during internal testing, the Wall Street Journal reported on Monday…

2 points•doppp•2 days ago•1 comment•
26.

Learn Raku by fixing small failing tests, right in your browser.

2 points•hankache•3 days ago•0 comments•
27.

"This has potential for so much negative PR. It could portray us as ‘their AI is not good enough so they still need humans’ kind of coverage for this launch."

2 points•pluc•7 days ago•0 comments•
28.

"This has potential for so much negative PR. It could portray us as ‘their AI is not good enough so they still need humans’ kind of coverage for this launch."

2 points•toomanyrichies•8 days ago•0 comments•
29.

Pushgate is a checkpoint between your coding agent and GitHub. Choose the tests and security checks you require before a push goes through.

2 points•handfuloflight•8 days ago•0 comments•
30.
2 points•luu•11 days ago•0 comments•
31.
1 points•ngruhn•about 17 hours ago•0 comments•
32.

The perfect app for your band practice recordings. Built by musicians, for musicians!

1 points•dimonomid•1 day ago•0 comments•
33.

Drives your web page in a real browser and tells you what broke. No tests to write, no LLM. - awss1i/assay

1 points•awss1i•2 days ago•0 comments•
34.

A green test suite can preserve the same mistake as the code. Ask an agent to show a bug its tests catch, then check where the expected answer came from.

1 points•joshcsimmons•3 days ago•0 comments•
35.
1 points•hellerve•3 days ago•0 comments•
36.

It's crazy interesting to contemplate what all these ChatGPT A/B tests area about. This is just a sample...

1 points•tosh•4 days ago•0 comments•
37.
1 points•luu•4 days ago•0 comments•
38.

Daily tests and community votes on whether AI models got nerfed. Each model is only ever compared with its own first week.

1 points•csergiu•5 days ago•0 comments•
39.

Are generative (randomized) tests significantly more effective than example-based unit-tests at discovering bugs? There's an interesting discussion about this on lobste.rs. One argument in favor of unit tests is,…

1 points•jamilbk•6 days ago•0 comments•
40.

What the independent measurements of TypeSafe's Jev found in its first eight days: arXiv preprints, GitHub evaluations and blog benchmarks, each traced to its primary source. Accuracy, calibration, speed, cost, failure modes, the prior art, the open alternatives, and what is still unmeasured.

1 points•xbill•6 days ago•2 comments
41.

The venture capital firm is starting a for-profit alternative to traditional college in San Francisco.

1 points•jnord•7 days ago•0 comments•
42.
1 points•rudx•7 days ago•0 comments•
43.

"This has potential for so much negative PR. It could portray us as ‘their AI is not good enough so they still need humans’ kind of coverage for this launch."

1 points•cdrnsf•8 days ago•0 comments•
44.

I added a dependency to a test helper and forgot to declare it in the dev extra. Locally it was...

1 points•juanauriti•8 days ago•3 comments
45.

Humanity faces four defining “tests of power” – around war or peace, inequality, climate change and AI – that will determine whether the world moves towards a new era of cooperation or deeper division, the UN chief…

1 points•geox•8 days ago•0 comments•
46.

Are generative (randomized) tests significantly more effective than example-based unit-tests at discovering bugs? There's an interesting discussion about this on lobste.rs. One argument in favor of unit tests is,…

1 points•r4um•10 days ago•0 comments•
47.

TL;DR I moved a 90-spec Cypress suite to Playwright in 4 working days using Claude Code....

1 points•yureki_lab•11 days ago•1 comment
48.

Flaky tests may seem like a minor inconvenience—we often learn to identify which tests occasionally (or frequently) fail for no good reason and pay them less attention. However, it’s crucial to understand their impact…

1 points•mpweiher•over 3 years ago•6 comments
49.

Daniel Balcarek's article API Performance Testing: How to Design Realistic Tests makes a...

0 points•cherware•7 days ago•3 comments
50.

Google has launched a pilot program that pays publishers for their contributions to its AI-powered search features, according to a report from The Information . The pilot program reportedly includes around 100 publishers and comes as Google faces scrutiny over the impact of its AI features on web traffic. Digiday first reported on the pilot program , which began less than a year ago. As part of the test, Google is reportedly paying participating publishers for how much their content contributed to AI Overviews and AI Mode in Search, as well as its Gemini chatbot. One publisher that joined the program when it first started earned over $1 milli … Read the full story at The Verge.

0 points•Emma Roth•about 10 hours ago•0 comments
51.

What’s in the box?! Cores. So many cores. | Photo: Amelia Holowaty Krales / The Verge The Mac Studio review unit that Apple sent us to test this year is, put simply, kind of outrageous. It has an M5 Ultra chip with a 36-core CPU and 80-core GPU, 256GB of RAM, and 4TB of storage and costs $12,299. This thing is not for your typical content creation workloads. It's for AI developers and some of the most demanding 3D visual effects houses out there. Frankly, our usual benchmarks aren't cutting it. The latest version of Apple's most powerful computer begins shipping today with new chips and prices that extend even further into the stratosphere. The new Mac Studio uses the same design that's been with us since 2022 , replete wit … Read the full story at The Verge.

0 points•Antonio G. Di Benedetto•9 days ago•0 comments

Related topics