Earlier this year, I believed planning was going to become the most important part of building software with AI.

586 points•jmvldz•6 days ago•509 comments•

509 comments

bcherny5 days ago
[I work on Claude Code] I broadly agree with the author’s point: plan mode was useful, and is no longer useful.

In Claude Code, all plan mode does is add a little reminder to every user message along the lines of “you’re in plan mode, please don’t code yet”. It’s something I came up with late on a Sunday night many months ago, when I got tired of asking Claude to plan with me first before coding in each new session. Something people might not realize is plan mode has always been a prompt — it has never changed the toolset because doing so would break the prompt cache, and so would be expensive for users.

This worked well for a while, until a few months ago, using early versions of Fable, I realized that I wasn’t using plan mode anymore because the model just got it, and because for the increasingly complex work I asked the model to do, planning had become interactive and iterative. With Opus 5.5, I feel Opus has gotten to that point too.

For codebase understanding, I sometimes ask Claude to generate an artifact that explains some aspect of its changes. For complex diffs to core parts of the system, I will often ask it to make diagrams or even interactive demos so I can better understand the change and alternatives considered. I don’t do this very often, but it’s a useful way to explain code when you need it. I ask Claude to attach these artifacts to its PRs also, so others can understand and future Claudes have the context.

stingraycharles5 days ago
> This worked well for a while, until a few months ago, using early versions of Fable, I realized that I wasn’t using plan mode anymore because the model just got it, and because for the increasingly complex work I asked the model to do, planning had become interactive and iterative. With Opus 5.5, I feel Opus has gotten to that point too.

For me it’s actually the opposite, and Claude Code’s plan mode isn’t nearly sufficient. Personally I ask Claude to write down a markdown file with its plan, then review the plan using plannotator, and then go back and forth (most of the time it’s actually the comments that are the problem, not the code).

Then start a fresh session, seed it with the plan, tell Claude to find ambiguities / friction points / oversights, resolve those, and then implement it.

Review once again with plannotator, go back and forth, and then send PR.

Maybe not the “vibe coding” that was once imagined, but this does ensure I am fully aware of the code and architecture, the quality, and this also prevents long term degradation.

mikepurvis5 days ago
I've recently gotten religion on the workflow that is many (relatively) short-lived agent sessions passing planning/handoff docs between themselves. It's better for my own task tracking, better for handling "oh btw I noticed XXXX", and better as a clear review point. Overall it just feels like it takes a lot of the formerly implicit context that was whatever we happened to have talked about and turns it into a much more explicit "this is what you need to know, now go".

Currently looking for a framework for managing this in a more formal way, and I think it's probably beads, but interested to hear from others.

jcelerier5 days ago
> (most of the time it’s actually the comments that are the problem, not the code).

I wonder if it's just a consequence of a gigantic training set full of comments completely out-of-date with the code, leading to the model considering this "normal"

nerdyadventurer5 days ago
There is a popular skill for this kind of workflows: https://github.com/obra/superpowers
rickye264 days ago
My workflow is very similar. But I just ask agents to write design in html instead of markdown, due to its richer layout and better interactivity. When the design is about UI or anything related to graphics, this approach is extremely efficient.

I also found that having the design reviewed by multiple agents has very little marginal value. The review agent will always find something to improve, but mostly it’s just nit and not anything super important.

I used to let Claude just upload the html design doc to Claude artifacts for me to review. Recently I switched to codex and started to use my own tool https://github.com/hyperlogue/r3 to complete this workflow.

dmix5 days ago
I do the same, I don't use Claude Code or Codex planning because it is mostly pointless, even with Fable/Astra. I just have multiple agents work on a markdown file which I manually perfect, often breaking into multiple different files for large features or PRs. I also create design 'handoff' documents which I feed into Claude Design or Astra along with screenshots and wireframes. By the time an agent does something I'm well prepared.

I've tried doing the incremental, iterative approach with just Code and it's just not as effective unless you're working on something simple or experimental. Or you're shipping to something non-serious or perpetually beta.

akersten5 days ago
It's useful because it let's me see the decisions the model will make before it wastes a ton of time implementing them. The model is smarter now but that doesn't solve for underspecification if it guesses my intent wrong
bcherny5 days ago
Interesting, I don’t see this very often with the latest models. Are you using Opus 5.5/Fable 5.1?

Either way, plan mode isn’t going away. You can always /plan or ask Claude to enter plan mode. We might re-map the shift+tab keyboard shortcut to something else by default for people that don’t use plan mode.

brokencode5 days ago
You can just tell it to write out a plan.md file.

I greatly prefer this, since it lets me iterate on the plan with Claude for a while without it repeatedly asking if I’m ready to implement the plan.

Once I’m satisfied, I usually start a fresh session and tell it to implement the plan.

For smaller plans, you don’t need the file. Just ask it to come up with a plan. I don’t recall the last time it just started implementing if I only asked for a plan.

solarkraft5 days ago
But that’s just the value of planning, not having it be a special mode.
bityard5 days ago
> wastes a ton of time

It wastes a ton of tokens as well and those are not cheap.

zarzavat5 days ago
It’s not that planning is dead, but rather that planning has outgrown the simple “Plan Mode” feature as models have become capable of taking longer turns.
aymandfire5 days ago
hi I’m the author of the post. I think that’s basically the distinction I’m trying to make.

Historically, plan mode served two different roles:

1. making the agent’s instructions precise enough to execute 2. helping the human understand what was about to happen

I think #1 is less necessary as agents get better. #2 is going the other direction, it becomes more important as the model is able to do more on its own because larger chunks of work are happening with increasing complexity.

Where I’ve changed my mind is the interface for #2. I increasingly think an interactive, iterative workflow is closer to how people actually build understanding than being handed a long generated document, especially one they didn’t author themselves.

The human-understanding problem is very real though

digitaltrees5 days ago
Awesome blog. You're a good writer. I enjoyed seeing your article on AI in 2018. Thank you for sharing your expertise.

Also I hope your delivery goes well. My wife (and co-founder) had a challenging delivery and it really put life in to perspective for both of us on a range of issues (how much women's pain is minimized in the health system requiring stronger personal advocacy than I would ever have expected).

As far as plan mode, I still find it essential in keeping agents on track. I build propelcode.app and have a variation on plan mode I still find useful, happy to trade notes on agentic coding if youre interested.

jeffreygoesto5 days ago
Did you just invent "agile" maybe. ;)

Can't help but think "Doesn't matter if a machine or a human with (even slightly) different background wrote it", maximizing information flow is maximizing common assumptions and "culture" to only have to communicate a small set of current information for the task at hand. Being a team means having built a joint context so to say. This has always been the purpose of design documents and they always were too big or too small. Because you did not write them, but the others. If you only produce code you think they are the past and useless. If you iterate and your team grows, you start seeing the value in always current docs that are containing just what is not in your everyday culture.

All the best for you and your growing family. I had a similar experience recalibrating my values...

thot_experiment5 days ago
Fascinating! I literally never use Claude without plan mode and I find it's basically useless without it, constantly wasting tokens going in circles on irrelevant things. Fable or Opus. I feel like neither has a good sense for how to architect things and if I don't use plan mode it usually wastes hours of time chasing it's tail or implementing kludges on kludges to get something working that would be a much simple fix elsewhere, especially when working on a larger codebase.
physicallyIllfr5 days ago
Hence why the Anthropic employee is telling you to not use it and just aimlessly throw tokens at a wall. It'll eventually get you there, sure, and consume more tokens. This sounds absurd, but trust that Anthropic (and any large company) is hyper aware of how customer behavior impacts their revenue, and they certainly will try to steer you into behavior that increases revenue.

Also this is a way less removed process that I want nothing to do with. The more removed I am from the process the more I hate my job, get burned out and genuinely wish that Anthropic never existed.

Even if it could "just know" or infer my intent. It wouldnt be desirable.

Edit: Oh yeah its Boris, hes one or the most disengenous shovel sellers on earth right now.

burningion4 days ago
I realize there are already a ton of comments, but I think you're missing the idea of _precision_.

Most people aren't precise when initially describing their problem.

Jumping straight into implementation skips the part where we refine and better define what it is we're trying to do, and think through the implications of those changes.

I suspect "trusting the model" doesn't really work at scale with finite resources.

taurath5 days ago
I'm actively watching understanding slip away from developers, code review getting paired down to no comment checkmarks, and codebases go to bloated messes that nobody can read. Axioms like engineers must understand and take responsibility for the code they ship are getting torn down, and the products coming out are reflecting conway's law, becoming impenetrably obtuse and always "so complex there are no obvious deficiencies" (as opposed to "so simple there are no obvious deficiencies" which used to be the aim).

The one thing plan mode helped is for the humans to get an understanding of the strategy, and be able to poke around and look at the design and architecture. You can achieve this with some self discipline and keeping shorter leashes on agents, but it feels like a losing battle. The best devs still put out good code, but the poor devs are learning nothing while their metrics look great. I can't help but think we are racking up immense amounts of debt that will very soon become due.

gonzalohm5 days ago
100% this. Some people are basically adding AI as a dependency for their projects. They no longer understand the code
jmb995 days ago
I’m a big proponent of writing simple, understandable code, so I’m playing devil’s advocate here a bit, but: who cares?

A significant (majority?) portion of developers have been shipping JavaScript/node applications for the last decade that contain hundreds of MB to GB of code from god knows where doing god knows what with dependency trees the size of redwoods. It’s not like your average mediocre dev really knew what was going on behind their gluing of frameworks together - at least from what I’ve seen.

If you have remotely competent tech leadership that enforces relatively intelligent patterns (a good one I’ve found is “write everything backend in rust”) you can make AI churn out monstrous amounts of code that… isn’t all that bad? And if you enforce it writing and updating a docs/API.md on every commit/PR you’re probably doing better than 80+% of devs I’ve ever met. Up until a few years ago it wasn’t uncommon to roll up to a new job that was a “legacy” pile of garbage concocted over 20+ years with no comments or API docs and a readme that tells you to ask for help from someone who has been dead for 5 years. At least AI code is full of comments (some of which might even be accurate) and there’s a finite (relatively low!) cost to figuring out “wtf is this doing and how is it doing it”

pianopatrick5 days ago
Seems to me that if the AI writes the code, then AI can easily copy the code.

I.e. how hard is it to point an AI at a piece of software and say "AI, copy this"?

Seems like sooner or later copying just becomes a matter of spending enough on tokens.

Seems in that world, all significant software projects get copied. That turns software into a commodity loss leader for other business models or an open source project. Similar to the way Chrome works for Google and the way Firefox works.

jstummbillig5 days ago
This has been the experience of everyone making decisions in any company without being the one doing the technical work. It's not a novel concept. It's actually the opposite, compared to technical people running companies.
swat5355 days ago
In many places it's demanded by upper management that devs use AI.. so even if engineers wanted to avoid using it, they would have to meet their quotas.

This is what happens when executives suffer from AI psychosis. They were already impatient, now with AI all they care about is feature velocity.

The faster they can hit that refresh button to see the features, the quicker sales can close the deals for them.

AI has basically sold them to wet dream.

fireflash385 days ago
Do you understand every part of your dependencies now, pre-AI?

I see it kind of like baking/cooking. Do you bake your bread from scratch? Do you grow your own wheat and grist your own flour?

I think over reliance on it or not even trying to understand what is happening is a big problem to be sure, but it's certainly not a new problem.

digitaltrees5 days ago
I feel the same way but I have successfully refactored some of the early experiments. Our team has settled on targeting a double output from the before times but more ambitious product vision because AI can teach us things we don't know. We actually target 2 days of coding and 3 days of learning with Ai so the increased efficiency allows upskilling rather than just pushing more code
Daishiman5 days ago
> The best devs still put out good code, but the poor devs are learning nothing while their metrics look great. I can't help but think we are racking up immense amounts of debt that will very soon become due.

I agree with this, but the reality is that it's only the result of models empowering devs, and power in good hands amplifies positive results while power in mediocre hands amplifies technical debt.

It's a good time to choose wisely who you work with.

Ronsenshi5 days ago
> It's a good time to choose wisely who you work with.

Very true, but this also makes me think what kind of ridiculous obstacle course future hiring process would look like.

In a land where anyone with a pulse can prompt AI to make an app for them - how would future hiring managers and team leads figure out who will drag codebase down with tech debt and who wouldn't?

rglover5 days ago
Just like rushing made messes in the before times so too does rushing via LLM. The exact same outcome will happen, but at a far greater velocity and scale than anything we've seen in this industry before. Old Testament, Mr. Mayor, real wrath-of-God type stuff!
Ronsenshi5 days ago
It's truly bizarre to read about all these people who just give up on any understanding about what they are working on.

I very regularly use plan mode not to even make a plan of action itself, but to better understand what possible issues might come up when implementing some feature or fixing some bug. And it is quite common for me to fix or rewrite certain findings that AI comes up because its assumptions are not quite right or don't align with overall goal.

And yet so many seem to be perfectly fine leaving all the decisions to AI - even if it's going in the wrong direction. I suppose that's all the people who got into software purely for money or status - never really caring about the actual thing they are working on.

taurath5 days ago
They believe they know, because they have LLMs to tell them, but they don’t often get to the point of being able to have a conversation about it
bityard5 days ago
When I draft my idea for the implementation of a feature or bug fix, I don't even trust a _human_ to understand what I mean the first time. There are _always_ either errors on my part, or erroneous assumptions on theirs. Everything from "this accounts for X and Y, but not Z which breaks the whole thing" to "this part of the idea directly contradicts what with you said earlier, what do you want to do about it?"

I can't bring myself to trust that an LLM understands what I mean better than any human would, no matter how "good" people claim they are getting.

TFA seems to be advocating for regular old vibecoding. Code now and ask questions later. Which is their choice, and is perhaps even a valid choice in many cases. But at least call it what it is.

wilbo5 days ago
Plan mode, for me, is my opportunity to develop and understand my own plan. Sometimes, rarely, Claude demonstrates it misunderstood my intentions by writing a plan to address the wrong problem. But usually it's about me fleshing out the scope and boundary of the intended changes.

I find that faster than code first, ask questions later. But it takes more time up front.

Skunkleton4 days ago
In the before times, and sometimes now even, the way I always coded was by thinking/discussing the work, then hacking up various interesting bits, then throwing it away and doing it the "right" way. Plan mode makes it harder to go back and forth between "hacking" and "thinking" phases.
Kiro4 days ago
Coding with plan mode is still vibe coding. No-one outside HN calls it things like "agentic engineering" or whatever. Everyone just says they are vibe coding regardless of how hands on they are.
isityettime4 days ago
Whether you use plan mode or not is orthogonal to whether or not you're vibe coding. "Vibe coding" means that you don't read or edit the code; you are running the software and iterating on vibes alone. That's the meaning since the original (and very recent!) coinage and hasn't changed.

If you're reading and editing the code, you're not vibe coding. If you're not reading and editing the code, you're probably vibe coding even if you feel "hands on".

tcdent5 days ago
The real reason why plan mode is dead is because you can just conversationally instruct the agent to not make changes to the repository or to make changes to selected documents only, and it will listen. There was a time when we needed to enforce this via selected tool use, but we have surpassed that.
esafak5 days ago
I agree. Plan mode came about because earlier models were loose cannons, doing what they pleased. Today's models follow instructions better enough to not need a separate mode. However, there is still value in using a separate, smarter model for planning than execution, and persisting it for auditing on completion.

If you do use plan mode you might like https://plannotator.ai/

I typically converse with the default model to point the plan in the right direction, then have it iterate with a smarter reviewer to find flaws until the plan file is converged.

Feathercrown5 days ago
It seems unwise not to implement sandbox measures just because the chance of misuse has gotten lower.
monster_truck5 days ago
Plan was never a sandbox or a permissions system
diegof795 days ago
Yes, but plan mode wasn’t that.

You can use Docker’s sbx or similar VM/containers for that.

brianwawok5 days ago
What? Nothing to do with that
throwaway274485 days ago
It also doesn't allow leaving an audit trail of plans and decisions (by default, anyway). Most of my mutating prompts look like "Propose a plan for change X and write to file Y" and "Execute steps M-N from file Y".

RE the article: I don't think it's obvious why this process is worth following until you find your time and attention wasted. Conversationally-building is the express train to waste. I'm not sure why you would even be talking to claude if you don't understand what you want to build.

almostroot4 days ago
Do you really trust it with a production code base or database? Telling it not to make changes feels an awful lot like "Make no mistakes". I also like how Plan mode on Codex asks clarifying questions. I'm in management now so I've used Work more than Code lately.
z3ratul1630715 days ago
true dat
cronin1015 days ago
Anecdotally, in the Opus 4.6 days, it felt like there was something special about using plan mode to discover the approach then clearing the context to execute on it.

A mixture of defending against a disastrous mid-implementation compaction (where suddenly things would veer off the rails) and also allowing the fresh execution to double-check the assumptions and notice any subtle mistakes before context was poisoned.

I’ve found that for large enough changes I still prefer having a parent theorizing about the root cause of issues based on evidence and then dispatching targeted child sessions to fixed based on theories and concrete telemetry examples.

There’s something clean about having sandboxed context and a session you can quiz about architecture while one is heads-down working against a spec.

noisy_boy5 days ago
For me, it was the agent that decided about plan mode, when that happened. Most of the time I didn't invoke it but kept discussing and tweaking and if the context was close to full, make it dump it out to disk and start a new session. I felt like that worked more flexibly for me because I could keep steering and adjusting until it was what I wanted. At that point I made it write a detailed spec. Like my homegrown plan mode.

The "standard" plan mode felt too stifling.

Read the full thread on Hacker News →

Related stories