DHH's opening keynote had shockingly little to say about Rails.
230 comments
Rails World 2026 Opening Keynote [video]
I've seen this one a few times lately. People should stop building UIs, nobody wants to interact with a UI. Everyone's app should just be an API that you can use with a chatbot. Except for my app — my app is a handcrafted miracle of artisanal UX and its UI will change the way you see the world.
It's exactly the old argument, just now with LLMs in the place of shell pipelines: in terms of functionality and value to users, software ought to be malleable and composable. We've known it since the eighties. But the model of selling a piece of software as a product as if it were a pair of shoes is incompatible with that. You need a big monolithic application to justify users paying a bunch of money for it, and you need it to have a fancy interface that makes an impression. And the whole software industry is built on top of that model. Where monolithic software is completely unfit for a purpose, like when it needs to be a component of a larger system, we rely on (mostly unpaid) OSS.
Except now LLMs let you, with very little technical know-how, plaster a ‘programmable’ interface on top of unstructured data/interface meant for humans, and because that's what we actually want, of course everybody does that. So the end result is a wildly expensive pipeline from API to UI and back to API again. I wonder how long the legacy ‘human-oriented’ layer in the middle, and the industry that's been built on top of it, will last.
(Separately, chatbots are not great as a UI for most things, and the problem of building the universal UI still also stands. But it turns out for a lot of things people would rather have a bad universal UI than a good special-purpose UI for each task.)
A chatbot using an LLM of course not, it suffers from hallucinations, it may do what you want but there is a change it won't and you have to fight it to get the desired result.
Chatbots are far WORSE than traditional UI for everything. If some product has a chatbot functions it's the first thing I disable, if it's not possible to disable it, I avoid the product.
And GUI applications are typically preferred, at least for the normal people and not us nerds, to CLI applications, since you know people like moving a mouse and clicking on buttons (or tapping them on a touchscreen) that learning commands: a chatbot doesn't make the CLI experience less awful for the average user, and for the nerd user, he prefers to use the CLI directly (replace asking the chatbot with man or --help and you don't need to emit tons of CO2 and transmit your personal data to a datacenter on the other side of the world to do stuff you did with MS-DOS)
This is a very software-engineery point of view and totally false for the ordinary user. Heck, it is even false for me as a software engineer.
> since you know people like moving a mouse and clicking on buttons
No, they like well-designed user interfaces, and the UI design of a typical CLI is abysmal.
*a small percent of apps behave like this
> it may do what you want but there is a change it won't and you have to fight it to get the desired result
The vast majority of all software made in the last ten years falls into this category. This includes the goddamn operating system itself.
I don't think you're making the argument you think you're making.
The chatbot we added to our B2B product is by far the most popular addition we've made this year. Our users are not tech-savvy, they use a lot of apps everyday and don't want to have to learn and keep up with just another UI. So they like being able to type their wants and needs in plain language (or speak it into their phone, if they are in the field) and get a plain language response back with embedded images and charts.
YMMV of course.
Even though most of my projects have a UI, I build a CLI/API version so that the LLM can interact with it directly. I have been using this "CLI driven development" approach for more than a year now and have had fantastic results. The CLI arguments make it easy for LLM to interact with software it wrote.
I usually ask LLM to build a lib, then expose as a CLI and a RESTful API.
Words are a great UI. They can be easily converted into audio and haptics. We already have many systems in place that do exactly this with words.
Sure, pretty graphics are nice, but their sole intent should be to convey information. A User Interface that does not allow a person to easily have information conveyed to them is a bad interface.
The modern AI world isn’t perfect, but for a large portion of people who have accessibility needs, it’s absolutely a positive impact on their life in a way that no single technology has been until now.
A question becomes, how do you evolve a chat experience to task specific actions?
I’m building an app to explore scripture. Chat is an amazing starting point. But it’s terrible once you get into reading the actual scripture.
I think we will see more of this in the future. Here is how I’ve envisioned evolving an experience out of chat. Curious if others have their own ideas.
Just wait a few months/years and then some smartasses will come with a novel approach nobody ever has thought of before. Let's call it "context menus". They will give rapid access to the most used features without the tedious typing.
Until MacOS 27 the text context menus had items "Writing Tools" and "Proofread". Quick and efficient. MacOS has dropped these items and replaced with "Ask Siri". Now I have to type "proofread" every time and pray Siri agrees that it should proofread the selected text. Once it has done that, do i have to type "replace"? I don't even know.
There is a reason why UIs became popular. Command lines are fine but hard to discover and tedious to use. With AI it's even worse. You can't know for sure the AI interprets your commands in the way you intended.
At that point, why even bother with "driving a CLI"? I have a hard time seeing how this all doesn't go away soon. At the current trajectory, I am not seeing a future where software like Basecamp or Fizzy or the cheapest alternative are competitive with "Claude, build a basecamp-style project management tool for my team."
I don't mean to say that thoughtfully built, opinionated software doesn't have intrinsic value, I believe it will always be "better" in certain aspects...I just can't fathom that a market for it will exist in very short order.
I would LOVE to be talked out of this perspective.
I’ve spent the last three years tweaking away at this algorithm. There are so many little details that have had to come together to make it an enjoyable experience. For example, take the one problem of balancing multiple language: What is the best way to balance multiple languages? How should you switch between them, and how often? What happens if you are over performing in one language and underperforming in another? What happens if you have more reviews in one language than another that isn’t in line with your priorities? How do you deal with competing learning speeds and language priorities? Spiky review burdens? What happens when you change your language priorities?
This is probably the most trivial part of the algorithm, and it took years of trial and error and dozens of iterations to get it to feel right. It requires knowing about SRS, the FSRS algorithm, language learning principles, and even then tons of experimentation, trial and error, and hundreds of days of daily usage.
I’ve asked LLMs every step of the way what to do. I’ve probably taken <1% of the advice I was given. Just asking an LLM “build an SRS app that balances multiple languages” would produce a result, but one that would definitely be terrible to use, and probably would end up slowing you down, not speeding you up. Explaining to it any of the issues that you would come across when using it will likely suggest to changes that will only make things worse, based on the thousands of terrible ideas I’ve seen confidently produced.
The hard part of a product is not the idea, but the experience of using it. The valueable work is normally not in the big feature specs, but the hundreds of little paper cuts that have all been smoothed over. The time saved is not because you don’t have to code it, but is from not having to become an expert in several different fields.
solid
It's all opportunity cost.
Edit: throw in mandated security audits and maintenance for SOC2 compliance or whatever else you're on too, depending on the industry you're in.
Claude can 100% build you your own basecamp quickly, but there is a lot of good thinking that has gone into a product. Edge cases. Integrations. Clarity of thinking and ways to extend it. Agents love standing on the shoulders of good crystalized cognition (grep, curl etc), but I think that extends even beyond base CLI tools.
Which is not to say that "cheapest alternative" won't be much more important than it is now, but the tools that will succeed are the ones that solve clear problems with accessible CLI/tool-call like interfaces and take significant load off the agent running, while doing it efficiently and cost effectively so they are clearly better than building it yourself.
To be fair, there will be a ton of margin collapse and that will be unachievable by a lot of current orgs / structures.
Also: the coded apps that AIs build are unmaintainable without a tech background. They may work for exactly their proposed happy-path and that may be great. But over time I can’t see how it’s possible these apps get maintained.
Tomorrow's software will need to be _perfect_, or else it won't differentiate from "Claude, build a basecamp-style project management tool for my team".
We're in the very early stages of LLM driven programming, so it's hard to say how it will all shake out, but my anecdotal experience is that LLM written software is very reliable and easy to extend and develop. I have a side project that is about 90% LLM generated code (about 25k lines of production code and a similar amount of test code). This is a revenue generating product and I've had no issues with reliability, security, or performance.
For what it's worth this app is a rails app and I have no plans to switch to anything else. Rails works nicely, the LLMs extend it easily, and almost everything is I/O bound so I don't need C++/Rust level performance.
there's an important caveat here - LLMs absolutely can write better code than me, but they can also write far worse code, and often don't seem to be able to tell the difference. I find myself spending a lot of time prompting the bots to refine their code in specific ways that I only know about because I read the generated output.
but my anecdotal experience is that LLM written
software is very reliable and easy to extend and
develop.
Extremely refreshing to read. The whole thing, not just this statement.Too much discourse creates a false dichotomy of "awesome, wonderful, hand-crafted code" vs "shitty AI slop."
A lot of hand-crafted code is pretty bad. This is true even with talented engineers: there are often edge cases they did not think of... and in many cases, could not have thought of.
And AI code is pretty good if it's steered and vetted by a knowledgeable engineer. It is clear to me, and has been clear for a long time, that "talented engineer plus AI" is the winning combination and will remain that way for at least the near future.
Talented programmers will retire and Father Time takes care of the rest. Where will the new wave of talented programmers come from? I use AI, but I am a hobbyist.
Also eventually you'll also have to trust it to write the deployment code or even run the deployment itself, otherwise SRE is going to be the bottleneck. And only then should I feel anxiety about the rest of my career (that, or my employer decide LLM are good enough to get rid of me, even if they are imperfect).
The actual code and architecture is where it still lacking IMHO. Especially in rails... Like it will just build the least scalable features if you let it do it's thing. Ten queries for what could be one. No separation of concerns. Huge files, lots of duplicated code and then tens of thousands of units tests that just grow like a fatberg.
If your app does anything serious, if you have serious traffic... you are going to need to review each session finely (and your DB schema with each deploy). It could be that this is maybe an indictment of rails more than LLMs, I guess maybe time will tell.
If you don't look at your code and need to debug a production issue, do you just panic chat with your agent to solve the problem?
I guess it all depends on context, where I work ops stuff is clearly the bottleneck for a variety of reasons (technical debt that we are constantly working around, secret management for compliance reasons, etc). We self-host everything from bare metal. Some people would need to rethink the infra from the bottom up before it is "LLM ready".
That doesn't mean my job is not threatened mid/long term, in fact, thanks to LLM it is possible to rebuild that in a reasonnable amount of time I think. It's actually one of my side projects to offer this as a service. But if that doesn't work maybe I should have a plan C.
My experience is that the more structure and constraints you can place on what the LLM/Agents actually do, the better they perform. If you just let them run wild a la "make me an app that will make a million dollars MRR, no mistakes", you get a clusterfuck. And it gets more clusterfucky the longer it goes on.
This is lacking a lot of nuance. There are many types of code. There are many situations where I'm analysing something one-off and if I get 33% success ratio, but can easily verify the result, I'm happy - still saved me time and money. They're are situations where I'm generating graphs from some dataset and I don't have to trust anything - I know what the result should look like, I just need the agent to drive matplotlib. There are low stakes dashboards which I'm happy to generate and develop entirely via agents - they'll embed the updated screenshots in PRs that I can yolo-merge - worst case is that someone complains about something not working next time they visit. Then there's lots of experimenting which was never stopped to hit production anyway.
Finally after all of that you get code that's actually part of deployable features. Of course the trust is nowhere near 100%, but if you have a healthy testing process (e2e, validating different browsers, or whatever is appropriate for your environment), then what's your trust in human developer+review? Because mine is nowhere near 100% either.
In practice there are places where I extremely don't care about the code and never wanted it anyway, places where I'll read the code to check the design or just in case, and places which agents are not allowed to touch (medical billing rules for example).
That is very not true for many cases. Agents generating code are usually painfully slow, finding workflows that replace that reasoning with running code usually speed things up in my experience.
My version of "code review" is "test failure investigation" and I have a hard rule in my repos that agents never modify existing tests while they're implementing features. This means that when I run the tests after they do a bunch of stuff, I see all the tests break. Mostly they're stale assertions and we patch them up. Sometimes they're regressions and we patch those up, and sometimes I notice something dumb and dive deep into a facet of the architecture that can be improved, spend some time exploring it then get the agent to implement.
I think it's a better approach than trying to read everything and catch bugs or improve quality because you wind up focusing the things that actually matter in the real world rather than the things you think might matter.
This is how I've always approached working with offshore devs too. Focus on testing for quality control, not "code quality". After all, you're going to look at the code you wrote 5 years ago and think it's shit anyway right? So all your code is shit.
Regarding "lack of understanding", here's a recent anecdote: I had a bug in a production (but relatively new) system. The customer was texting me saying that they couldn't scan a QR code because it kept "skipping and glitching". They sent me a short video. I described the problem to the agent and it figured out WAY faster than I would have been able to that the customer's clock was set incorrectly. They turned on network time and bingo bango, the thing worked straight away.
I don't thing "comprehension debt" matters at all, because if you want to know something about the code you ask the agent. I can't remember how anything works after 12 months anyway, so I would frequently have to spend ages grepping my own code when a customer came back and asked me to change something in a system we hadn't touched since last year. Asking an agent the same thing takes minutes and is way more accurate (and fun!)
That only describes a happy path though -- I also had many instances where there's an issue and just describing it to an agent immediately identifies a fix and it all goes faster compared to me having to "load up" the flow of codebase into my brain first. But I also still have instances where it thinks it identifies an issue correctly, spits a fix which doesn't work and looks wrong. You point out that it does not make sense because of X, it agrees and spits out a new fix, which is also wrong and you start out on these back and forth wild goose chases, at this point I usually give up and do it the old fashioned way by understanding what is actually happening. If it is within an agent loop, there may be no back-and-forth to waste your time but then you pay with wasted tokens when it will eventually gives up or you stop it.
>I don't thing "comprehension debt" matters at all, because if you want to know something about the code you ask the agent. I can't remember how anything works after 12 months anyway, so I would frequently have to spend ages grepping my own code when a customer came back and asked me to change something in a system we hadn't touched since last year. Asking an agent the same thing takes minutes and is way more accurate (and fun!)
I would question the last part. In my mind the more you let go of control over your codebase the more likely that it will drift way from a place where it is still comprehensible to you, and also your comprehension skills atrophy, and with that, your ability to ask good questions and to prod your agent in a correct directions weakens, leading to more wild goose chases and burned tokens. This is all keeping in mind that for throw away or small applications, maintainability is not that of important of a value so this doesn't affect all codebases.
:blinks: The architecture of your service relied on customer clocks being correct? That is, you built a distributed service without a clock synchronization primitive that relied on the clocks being sychronized?
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 3 days ago
- Aaron Patterson’s closing keynote (and antidote to the DHH funeral) at Rails World 2026: “Get in loser, we’re going programming”globalnerdy.comLobsters · 2 points · 3 days ago
- What About Rails?jardo.devLobsters · 62 points · 6 days ago
- Rails World 2026 Opening Keynote [video]youtube.comHacker News · 418 points · 7 days ago
- Lobsters · -2 points · 7 days ago
- Rails World 2026 Opening Keynoteyoutube.comHacker News · 2 points · 7 days ago