630 comments
1. Collaboration between always-on agents is a really, really powerful thing. It allows for domain-specific expertise that doesn't overload the context window, while still allowing for access to knowledge if they need it.
2. Domain-specific always on agents creates a good barrier of trust. One of the things I dislike about Claude is sometimes it's memory is all-encompassing. It's weird that it brings up things about my personal life when I'm talking about something related to my business. I've never had that happen with Grok Bot bots because I have one for my biz admin and one for my personal admin. They don't intertwine, which is quite nice.
3. Combined with cloud agents / cloud builds, things become really powerful for development. It was the first time that I felt there was a solution to the git worktrees / multiple streams at once issue. Each bot has its own computer and can spin up additional cloud agents. It comes at the cost of end to end speed - doing something via a grok bot often takes an hour end to end, whereas with a synchronous local prompt it'll take like 10min. The difference is I have to babysit one whereas the other "just works".
On the flip side, since using Grok Bots my inference spend has 2-3x'd. It's worth knowing that tradeoff. Nonetheless I think Luna is a fantastic driver for these, and OAI has very good pricing overall. I'd give these a shot - I think a lot of people would be surprised how helpful they are.
Both are just Codexes running in a permanently rolling session in dedicated UNIX user accounts. They're wired up to Maildir so receiving a mail activates Codex and makes it read the new message, there are autonomy wakeup timers, they have accounts in my bug tracker and CI systems. They're currently useful for:
• Triaging and working on customer support tickets. Sometimes I wake up and the fix/response for a ticket filed by a customer is already there waiting for my approval. Recently I started letting them directly interact with customers in specific scenarios.
• Triaging the bug backlog. One of them decided to spend its "free time" finding old bugs that were fixed without being properly closed, or are dupes, so it's cleaning up detritus in the tracker.
• They obviously do all the coding and debugging by just assigning tickets.
• They keep an eye on a "pet" server the company has, and have proven able to fix it in the past when it ran out of disk space.
• They handle non-business projects I have for them.
• They help out with the release processes.
The dedicated home dir is very useful and they use it all the time as part of coding and investigating tricky issues.
My setup relies heavily on email, as everything bottoms out in email anyway. Watching them mail each other out of the blue to coordinate stuff is pretty cool.
I only have one though. What do you have the separate employees for?
For free time, do you send it mail with cron?
The way I distributed cloud agents for this https://news.ycombinator.com/item?id=49687032 was via grok bot setting up Fable cloud instances.
The catch with offerings like Grok Bot and Dots is that it is a slippery slope towards letting the labs keep the agent's learning and SOPs. That is where we should draw the line. Agentic coordination _has_ to be built with open protocols and OSS implementations, we as users need to push for the intelligence to be a commodity, and most importantly we need to make sure that we own the agentic team's learnings.
So why don't I just have Claude write notes and summaries on specific things, and then it will always have a knowledge base? Am I missing something? I mean, if Dots and Grok Bot are not an extra charge / extra compute, then I guess that's fine, but in Claude Code you can make a AGENTS.md or CLAUDE.md file, and if you put it into any directory, Claude will read it when accessing that specific directory, so if its in your root where you launch Claude, your new Claude instance will read it and have all that in a dedicated smaller context window. But as it edits code, it can read the smaller ones too.
IMO best to keep work and personal data segmented on the hardware level. It's better for opsec in every single way and helps if you were ever to be subpoena'd or raided, your work laptop would be the only in-scope device for search/seizure.
If police raid your house looking for electronics, they're going to take everything down to the Roku stick.
It seems like a dumbed down reskin of Codex/ChatGPT Work but with the power-user features e.g. visibility/mentions removed. As a serious engineer why would I want that? Then the word agent becomes dot.
It also seems to be running a VM so the agent has its own computer.
I guess the intent was to pull together the Codex, claw and ChatGPT Work paradigms and simplify them?
Take off your "engineering" hat and put on your "normie" hat...
OpenAI's Dots, Meta's Muse, xAI's Grok Bot, etc are trying to target non-technical people and give them "24/7 AI personal assistants". Apple is also pursuing this angle by adding more features to the new Siri (partnership with Google-Gemini).
Perspective of product and market : we need our AI to be more deeply intertwined with customers lives instead of just being on-demand chatbot Q&A sessions. The "sticky" product to create this customer relationship is "AI personal assistants". Instead of just providing "answers" (chatbots) -- provide "completed tasks" (AI assistant).
Perspective of technical predecessors and concepts overlap : OpenClaw and Hermes self-hosted software on Mac minis and Linux boxes as "personal assistants" is geared for techies instead of normies. Instead, repackage those types of tools for normal people in cloud VMs to be hosts for "long-lived agents". Zero install required.
“ Turn feedback into tested fixes You’re a developer working on an app. Your dot watches customer feedback for recurring requests, scopes smaller improvements and bugfixes, builds and tests them, and brings you complete PRs to review with attached videos showing the changes.”
Probably this.
> I guess the intent was to pull together the Codex, claw and ChatGPT Work paradigms and simplify them?
Probably this in part, too. They're trying to figure out how to decouple agent lifetime from conversation lifetime, without accidentally making it too useful for end users.
I noticed in the past that tooling - both for AI and in general - tends to miss the features that would make it most useful. Like, look how long did it take for the AI vendors to supported "scheduled runs", and they're still offering only toy-level configuration for that[0]. Wonder how many years it'll take to allow users to configure external triggers, and whether it'll be sooner than forced re-authentication will become frequent enough to make the feature useless in the first place.
--
[0] - I understand they don't want users to run this too often, or to accidentally end up spawning a job every 10 minutes - but with limits on frequency in place, there is otherwise no good reason for this to be a limited dropdown, instead of the "repeats every" configuration that every calendar app and phone app already uses.
That's how I understand it.
Cloud compute means its work VM is isolated from your devices, so it can't destroy your local data as easily (unless it has access to your devices) and it's on 24/7. A big negative is that you are giving access to even more of your services and data to a third party and the lock-in into a walled garden happens once you start depending on it.
So everyone is trying to give the aunts, cousins, neighbor's dads some of the agent capabilities that developers have been closer to.
That said, there's the other 10% here that is not necessarily new capability unlocks vs codex, but general product experience that can have some benefit to people already used to what the latest models can do
Why do they want it? What are some use cases that make sense?
I get how i prefer an already researched-version of a bug versus a raw bug notice, but I can do this with webhooks in the correct environment.
I really have no idea what to do with my agents over night. I can not build more. I can not think of more problems. My RAM is full.
This is exactly where I'm at with AI. (Mostly via Claude Code, but I'm not sure the harness, or my workflow in particular, is the important part.)
I am increasingly wondering though, is there really a valid reason to keep blocking on my approval? Most of the time, I wind up saying yes anyway, because the model has a valid, efficient solution.
What if it's faster at this point to just fix the mistakes?
Scary thought, but it seems like we're close, or already there.
I only leave it truly unattended if it's working on a very tight improvement loop, for everything else I'm still checking in on it between working on other things.
I still find agents need a lot of guidance and steering to produce the kind of work I want, but auto mode in a strong sandbox is very useful to me to take a bite out of that. It's particularly good for exploring problems experimentally - where most of the exploration might be thrown away after settling on a solution.
Having used it sandboxed, I wouldn't dream of letting auto mode run outside it. It's very creative at trying to work around the constraints of the sandbox (legitimately, not to escape it) and the classifier for auto mode seems very permissive with the right context.
For context, I'm using per-project VMs with very limited egress and restrictive mounts. Self-built tool to glue it all together, currently unreleased. There has been an explosion of sandboxing tools recently, none of which was quite what I wanted. Heavily inspired by Gondolin <https://earendil-works.github.io/gondolin/>, but fat long-lived VMs.
For example I was optimising a checkout experience and I at least did 30 iterations till I was happy.
But there is also other stuff like namings. They pick good names but if you work with exchangeable vendors I need even more explizit names
However of course I’m trying to get as much in linting and agents.md
But even if I would say Yes to everything I couldn’t come up with things to do fast enough.
But maybe skill issue ?!
In the morning I come and look at the work and decide what to do next. It's great in that sense.
But not great for peace of mind. Because now there's always something to be done overnight...
Same thing in the development world. I need to do something when version x of some thing gets released. If I can instruct a bot to do it and it's safe enough (e.g. read only), then why not.
Are you sure? Based on their marketing video, usually you just have to tell it to swap one of the later slides or photos to put it first. :)
Oh, also you have to tell it those times you want to prioritize your children on your calendar.
Today the models are better, but they don't seem trustworthy or reliable enough that I want to give them them a lot of access or freedom. My coding agents sometimes still go off in the wrong direction, or say they did something other than what they did, or say they will do something and then immediately stop without doing anything.
The always-running (so almost never supervised) agent that is meant to do the same work as a person (and therefore needs _access_ like a person) seems like a notorious footgun from earlier this year was just made more powerful, and the companies that are supposed to know the most are telling you to connect it to everything.
Maybe people are ok with that, but I feel like it would be nicer to just sell a cheap SBC like a raspberry pi that you can plug in (and unplug!) and just pay for the tokens used instead of having your data stored offsite and paying cloud prices.
However, it is the case that other industries like 3d graphics and so forth have experienced a frontier shift so perhaps there’s still some advancement
Read the full thread on Hacker News →
Related stories
- Show HN: Groundtrack – Continual learning for coding agentsgroundtrack.devHacker News · 1 points · about 13 hours ago
- Hacker News · 45 points · 9 days ago
- Hacker News · 124 points · about 8 hours ago
- Hacker News · 1 points · 3 days ago
- Hacker News · 23 points · 3 days ago
- Hacker News · 37 points · 13 days ago