wrong
57 stories and discussions about wrong, aggregated from every source we track.
[I shared this note with my team earlier this week, and am posting it here as well. I hope it is interesting or helpful for others working on building product in the age of AI.]
OpenAI’s proof seems eligible for a $1-million prize—but only by using a controversial loophole
I have a bookmark folder called prompting. Forty-one tabs in it. "The 12 prompts that 10x your...
For 25 years I've been the person in the room asking to see the data. In incident reviews, in...
Most engineering teams working on long-context agents hit the same billing wall around turn twenty. A...
Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review...
OpenAI’s proof seems eligible for a $1-million prize—but only by using a controversial loophole
Schema-valid is not content-correct · Part 1/3 It starts with a storyboard a...
Yesterday I had a very frustrating start of my day for the most unexpected of reasons - a (supposedly) simple CPU cooler upgrade for my desktop computer turned into a nightmare. It was also a very educational…
Alert! Your state might already be infected! learn how and what todo about it in this article.
I built one engine where an LLM reviews another LLM's plan. I built another where two LLMs debate a...
TL;DR: I've shipped 26+ projects by directing AI agents, and I still couldn't write a Python program...
I published a memory benchmark six days ago. I spent most of the build trying to make it fair rather...
Last week we wrote about the classification problem hiding in your LLM bill, and about how to read a...
Despite admitting to a “component issue,” the company refuses to say what went wrong, how many units are affected in how many countries, and why it won’t recall its $500 AI toothbrush.
Software operations has always asked one question: Is it broken? AI agents change that. Here’s the shift toward measuring correctness at the level of the run, and the thinking behind Amazon CloudWatch Omni.
I’m afraid of spiders, so I made AI look at 2,000 of them. What it got right, what it got wrong, and how much it cost.
When we picture the end of the world, the first thing that springs to mind is typically the cinematic version of “things going wrong.” We picture the musical score rising to a crescendo…
The argument that repeated statements of concern about AI from multiple experts — some of whom have quit their jobs over it — is just marketing hype or regulatory capture appears increa…
The one-line version We ran five current frontier models over a set of documented-failure...
The essential guide for any startup founder or investor when things inevitably don’t go to plan.
I, for one, welcome our new UNIX primitive AI overlords
Despite admitting to a “component issue,” the company refuses to say what went wrong, how many units are affected in how many countries, and why it won’t recall its $500 AI toothbrush.
OpenAI’s proof seems eligible for a $1-million prize—but only by using a controversial loophole
A free AI canvas for following your curiosity. Try Wander and Deeper, then make Drift your own.
rm -rf on the wrong host: in 2017 a GitLab engineer wiped the production PostgreSQL database, and none of the five backups worked. The full postmortem.
Delegation goes wrong when managers don’t establish trust.
Software operations has always asked one question: Is it broken? AI agents change that. Here’s the shift toward measuring correctness at the level of the run, and the thinking behind Amazon CloudWatch Omni.
Learn how VAT-inclusive and VAT-exclusive pricing differ, why gross × VAT rate is wrong, and how to calculate net, VAT, and gross safely.
What is a developer to do when they need something more tangible than a chat box? Enter canvases.
See what keeps going wrong, get a debrief of how you drove each session, and carry approved lessons into future Claude Code and Codex sessions.
There's no /admin on this site. No login form to try a password against, no session cookie to steal, no plugin list to check against last month's CVEs — not bec…
View a Shapefile (.shp) on a map in your browser, free and with no install. Converts to GeoJSON (WGS84). Handles Tokyo Datum and JGD2000/2011.
See what keeps going wrong, get a debrief of how you drove each session, and carry approved lessons into future Claude Code and Codex sessions.
Despite admitting to a “component issue,” the company refuses to say what went wrong, how many units are affected in how many countries, and why it won’t recall its $500 AI toothbrush.
Why a persistent document with a threaded comment rail beats a chat thread for real work with an agent, and the pace it made possible on one site build.
What changes when implementation becomes cheaper than verification?
A no-BS series on what actually goes wrong in RL post-training - trajectory eyeballing, rubrics and verifiers, task design, environment quality - and how to fix it.
The dangerous AI answer is not the one that fails. It is the one that looks right, reads with total...