3 comments
It requires a minimal amount of steering and guidance and produces structurally sound, well documented, correct code, with less bugs and failures than we used to have with an all-human team.
The code has better test coverage, and is also easier to read and reason about in general (C++) than what we used to accomplish. The hardware testing is done automatically on the bench with minimal human intervention, and our bench test coverage is much, much higher than it ever was before, so that’s also good.
The failure mode seems to be just never finishing, with the local Overton window shifting to smaller and smaller issues until it’s writing bug reports and fixes about phrasing in the comments and formatting choices that are more aesthetic than functional, or fixing hypothetical bugs that would never be reachable without significant architectural changes.
Anyone else with similar experiences in your work?
Im hoping i can still operate this way with Qwen3.8-Flash-Next, which we can run on-prem with existing hardware, but im pretty sure ill still have to keep some of the expert agents from frontier models, at least for now.
Read the full thread on Hacker News →
Related stories
- Hacker News · 4 points · 6 days ago
- Horizon Create and Horizon Studiodevelopers.meta.comHacker News · 1 points · 4 days ago
- Hacker News · 1 points · 11 days ago
- Hacker News · 3 points · 3 days ago
- Hacker News · 1 points · 12 days ago
- The Verge · 0 points · 7 days ago