tests
51 stories and discussions about tests, aggregated from every source we track.
TL;DR GPT-6 Astra has started another familiar AI conversation. The model is more capable, Jensen...
The perfect app for your band practice recordings. Built by musicians, for musicians!
Run this in a terminal, then run it again under script(1): python3 -c 'import sys;...
Last month a teammate pasted a Playwright test into our PR channel and wrote "AI generated this in 4...
AI coding agents are getting very good at finishing tasks. They modify files. Fix errors. Write...
A test named test_all_adapters_importable asserted nothing. It would pass forever, even if every...
Peter Bourgon has a web site, and this is that web site.
What the independent measurements of TypeSafe's Jev found in its first eight days: arXiv preprints, GitHub evaluations and blog benchmarks, each traced to its primary source. Accuracy, calibration, speed, cost, failure modes, the prior art, the open alternatives, and what is still unmeasured.
Defunctionalise your continuations and your tests can run any computation a step at a time.
Last year, a critical payment processing service went completely silent during Black Friday traffic....
Are generative (randomized) tests significantly more effective than example-based unit-tests at discovering bugs? There's an interesting discussion about this on lobste.rs. One argument in favor of unit tests is,…
Every backend developer knows the classic trade-off when writing tests: Unit tests with mocks are...
My last post about Filament Studio was about v1.2.0 and multilingual content, back in April. Since...
authentik is now OpenID Certified. Here's what the conformance tests found, and why the certificate matters less than the tests behind it.
Evaluating Jev's calibration on two hard datasets.
Link to: https://pogueman.substack.com/p/125-tests-of-the-new-ai-siri
“Fornell et al. didn’t run standard multifactor asset pricing tests; didn’t use standard t-tests; and didn’t conduct out of sample tests. When you do, there is nothing there.”
Plain-English browser e2e tests cheap enough to run on every PR. Built on Jev and Playwright, open source, bring your own key. - sedum-dev/sedum
TOKI, Gifu Prefecture--Japanese startup Helical Fusion expects to begin preliminary power-on tests at its demonstration fusion reactor in 2027, it said on Tuesday, a step towards its goal of becoming a global…
Record AI agent runs locally and find which part of the context caused a decision. - RehanMohammed985/runtape
OpenAI is scrapping the release of GPT-6.1 Astra, a next-generation AI model planned for an October debut, over safety concerns raised by researchers during internal testing, the Wall Street Journal reported on Monday…
Learn Raku by fixing small failing tests, right in your browser.
"This has potential for so much negative PR. It could portray us as ‘their AI is not good enough so they still need humans’ kind of coverage for this launch."
"This has potential for so much negative PR. It could portray us as ‘their AI is not good enough so they still need humans’ kind of coverage for this launch."
Pushgate is a checkpoint between your coding agent and GitHub. Choose the tests and security checks you require before a push goes through.
The perfect app for your band practice recordings. Built by musicians, for musicians!
Drives your web page in a real browser and tells you what broke. No tests to write, no LLM. - awss1i/assay
A green test suite can preserve the same mistake as the code. Ask an agent to show a bug its tests catch, then check where the expected answer came from.
It's crazy interesting to contemplate what all these ChatGPT A/B tests area about. This is just a sample...
Daily tests and community votes on whether AI models got nerfed. Each model is only ever compared with its own first week.
Are generative (randomized) tests significantly more effective than example-based unit-tests at discovering bugs? There's an interesting discussion about this on lobste.rs. One argument in favor of unit tests is,…
What the independent measurements of TypeSafe's Jev found in its first eight days: arXiv preprints, GitHub evaluations and blog benchmarks, each traced to its primary source. Accuracy, calibration, speed, cost, failure modes, the prior art, the open alternatives, and what is still unmeasured.
The venture capital firm is starting a for-profit alternative to traditional college in San Francisco.
"This has potential for so much negative PR. It could portray us as ‘their AI is not good enough so they still need humans’ kind of coverage for this launch."
I added a dependency to a test helper and forgot to declare it in the dev extra. Locally it was...
Humanity faces four defining “tests of power” – around war or peace, inequality, climate change and AI – that will determine whether the world moves towards a new era of cooperation or deeper division, the UN chief…
Are generative (randomized) tests significantly more effective than example-based unit-tests at discovering bugs? There's an interesting discussion about this on lobste.rs. One argument in favor of unit tests is,…
TL;DR I moved a 90-spec Cypress suite to Playwright in 4 working days using Claude Code....
Flaky tests may seem like a minor inconvenience—we often learn to identify which tests occasionally (or frequently) fail for no good reason and pay them less attention. However, it’s crucial to understand their impact…
Daniel Balcarek's article API Performance Testing: How to Design Realistic Tests makes a...
Google has launched a pilot program that pays publishers for their contributions to its AI-powered search features, according to a report from The Information . The pilot program reportedly includes around 100 publishers and comes as Google faces scrutiny over the impact of its AI features on web traffic. Digiday first reported on the pilot program , which began less than a year ago. As part of the test, Google is reportedly paying participating publishers for how much their content contributed to AI Overviews and AI Mode in Search, as well as its Gemini chatbot. One publisher that joined the program when it first started earned over $1 milli … Read the full story at The Verge.
What’s in the box?! Cores. So many cores. | Photo: Amelia Holowaty Krales / The Verge The Mac Studio review unit that Apple sent us to test this year is, put simply, kind of outrageous. It has an M5 Ultra chip with a 36-core CPU and 80-core GPU, 256GB of RAM, and 4TB of storage and costs $12,299. This thing is not for your typical content creation workloads. It's for AI developers and some of the most demanding 3D visual effects houses out there. Frankly, our usual benchmarks aren't cutting it. The latest version of Apple's most powerful computer begins shipping today with new chips and prices that extend even further into the stratosphere. The new Mac Studio uses the same design that's been with us since 2022 , replete wit … Read the full story at The Verge.