750 points•alvis•1 day ago•630 comments•

630 comments

jjcm1 day ago
There's a lot of negativity in here for Dots. I've been a pretty heavy user of Grok Bot, and here are a few thoughts a long the positive line.

1. Collaboration between always-on agents is a really, really powerful thing. It allows for domain-specific expertise that doesn't overload the context window, while still allowing for access to knowledge if they need it.

2. Domain-specific always on agents creates a good barrier of trust. One of the things I dislike about Claude is sometimes it's memory is all-encompassing. It's weird that it brings up things about my personal life when I'm talking about something related to my business. I've never had that happen with Grok Bot bots because I have one for my biz admin and one for my personal admin. They don't intertwine, which is quite nice.

3. Combined with cloud agents / cloud builds, things become really powerful for development. It was the first time that I felt there was a solution to the git worktrees / multiple streams at once issue. Each bot has its own computer and can spin up additional cloud agents. It comes at the cost of end to end speed - doing something via a grok bot often takes an hour end to end, whereas with a synchronous local prompt it'll take like 10min. The difference is I have to babysit one whereas the other "just works".

On the flip side, since using Grok Bots my inference spend has 2-3x'd. It's worth knowing that tradeoff. Nonetheless I think Luna is a fantastic driver for these, and OAI has very good pricing overall. I'd give these a shot - I think a lot of people would be surprised how helpful they are.

mike_hearnabout 14 hours ago
Yes. I built my own version of this for my side business about six/seven months ago and it's been great! I have two "AI employees" now and if I were actually focused on this business full time instead of part time, I'd create more.

Both are just Codexes running in a permanently rolling session in dedicated UNIX user accounts. They're wired up to Maildir so receiving a mail activates Codex and makes it read the new message, there are autonomy wakeup timers, they have accounts in my bug tracker and CI systems. They're currently useful for:

• Triaging and working on customer support tickets. Sometimes I wake up and the fix/response for a ticket filed by a customer is already there waiting for my approval. Recently I started letting them directly interact with customers in specific scenarios.

• Triaging the bug backlog. One of them decided to spend its "free time" finding old bugs that were fixed without being properly closed, or are dupes, so it's cleaning up detritus in the tracker.

• They obviously do all the coding and debugging by just assigning tickets.

• They keep an eye on a "pet" server the company has, and have proven able to fix it in the past when it ran out of disk space.

• They handle non-business projects I have for them.

• They help out with the release processes.

The dedicated home dir is very useful and they use it all the time as part of coding and investigating tricky issues.

My setup relies heavily on email, as everything bottoms out in email anyway. Watching them mail each other out of the blue to coordinate stuff is pretty cool.

writtenoneabout 12 hours ago
I can't believe anyone trusts AI to do anything without strict oversight from a human. That's absolutely insane to me.
andaiabout 13 hours ago
Nice. I also ended up with a Unix user for my agents! (I was looking into Docker etc and realized the only thing I needed was "it doesn't blow up my files", i.e. a linux user).

I only have one though. What do you have the separate employees for?

For free time, do you send it mail with cron?

sealthedealabout 10 hours ago
Yep, I have a similar setup, we named him Routey and he is cute.
yonaguskaabout 12 hours ago
I hope you have them interacting with customers from behind an mcp.
KetoManx64about 10 hours ago
Do you configure them similar to how Hermes does? A bunch of memory files that give it context and then each action/batch of actions is a fresh session? /
jrflo1 day ago
What do you actually use it for? If I'm trying to work on code from my phone, I'll just use codex remote. As of right now I'm hesitant to hand over booking things / managing my calendar to an agent, because I don't view it as that much of a burden personally. So I don't really know what I'd use it for.
jjcm1 day ago
I've used it for admin, research & training (I''m working on my own image models), as well as just general coding.

The way I distributed cloud agents for this https://news.ycombinator.com/item?id=49687032 was via grok bot setting up Fable cloud instances.

juanreabout 8 hours ago
I completely agree. I have also been using teams that coordinate since last November. They mostly run my two companies and a lot of my personal life, and it's been transformative (to the point that I am spending most of my time these days building agent coordination tools).

The catch with offerings like Grok Bot and Dots is that it is a slippery slope towards letting the labs keep the agent's learning and SOPs. That is where we should draw the line. Agentic coordination _has_ to be built with open protocols and OSS implementations, we as users need to push for the intelligence to be a commodity, and most importantly we need to make sure that we own the agentic team's learnings.

giancarlostoroabout 8 hours ago
> 1. Collaboration between always-on agents is a really, really powerful thing. It allows for domain-specific expertise that doesn't overload the context window, while still allowing for access to knowledge if they need it.

So why don't I just have Claude write notes and summaries on specific things, and then it will always have a knowledge base? Am I missing something? I mean, if Dots and Grok Bot are not an extra charge / extra compute, then I guess that's fine, but in Claude Code you can make a AGENTS.md or CLAUDE.md file, and if you put it into any directory, Claude will read it when accessing that specific directory, so if its in your root where you launch Claude, your new Claude instance will read it and have all that in a dedicated smaller context window. But as it edits code, it can read the smaller ones too.

judge2020about 13 hours ago
> 2. Domain-specific always on agents creates a good barrier of trust. One of the things I dislike about Claude is sometimes it's memory is all-encompassing. It's weird that it brings up things about my personal life when I'm talking about something related to my business. I've never had that happen with Grok Bot bots because I have one for my biz admin and one for my personal admin. They don't intertwine, which is quite nice.

IMO best to keep work and personal data segmented on the hardware level. It's better for opsec in every single way and helps if you were ever to be subpoena'd or raided, your work laptop would be the only in-scope device for search/seizure.

post-itabout 12 hours ago
> It's better for opsec in every single way and helps if you were ever to be subpoena'd or raided, your work laptop would be the only in-scope device for search/seizure.

If police raid your house looking for electronics, they're going to take everything down to the Roku stick.

maherbegabout 13 hours ago
Yeah, but for some reason OpenAI hasn't setup multiple accounts to have completely separate profiles yet.
bluelightning2kabout 17 hours ago
I must be getting dumber as I get older. I genuinely can't work out what Dots actually is.

It seems like a dumbed down reskin of Codex/ChatGPT Work but with the power-user features e.g. visibility/mentions removed. As a serious engineer why would I want that? Then the word agent becomes dot.

It also seems to be running a VM so the agent has its own computer.

I guess the intent was to pull together the Codex, claw and ChatGPT Work paradigms and simplify them?

jasodeabout 16 hours ago
>As a serious engineer why would I want that?

Take off your "engineering" hat and put on your "normie" hat...

OpenAI's Dots, Meta's Muse, xAI's Grok Bot, etc are trying to target non-technical people and give them "24/7 AI personal assistants". Apple is also pursuing this angle by adding more features to the new Siri (partnership with Google-Gemini).

Perspective of product and market : we need our AI to be more deeply intertwined with customers lives instead of just being on-demand chatbot Q&A sessions. The "sticky" product to create this customer relationship is "AI personal assistants". Instead of just providing "answers" (chatbots) -- provide "completed tasks" (AI assistant).

Perspective of technical predecessors and concepts overlap : OpenClaw and Hermes self-hosted software on Mac minis and Linux boxes as "personal assistants" is geared for techies instead of normies. Instead, repackage those types of tools for normal people in cloud VMs to be hosts for "long-lived agents". Zero install required.

marcd35about 16 hours ago
I’d disagree with this being targeted toward normies. Reread the release. The first demographic in their copy is toward developers:

“ Turn feedback into tested fixes You’re a developer working on an app. Your dot watches customer feedback for recurring requests, scopes smaller improvements and bugfixes, builds and tests them, and brings you complete PRs to review with attached videos showing the changes.”

bluelightning2kabout 13 hours ago
Thanks, great perspective and explanation.
Godsend69about 15 hours ago
Have you considered using a Telemetry Blocklist (free) to block these agents from phoning home?
TeMPOraLabout 16 hours ago
> It also seems to be running a VM so the agent has its own computer.

Probably this.

> I guess the intent was to pull together the Codex, claw and ChatGPT Work paradigms and simplify them?

Probably this in part, too. They're trying to figure out how to decouple agent lifetime from conversation lifetime, without accidentally making it too useful for end users.

I noticed in the past that tooling - both for AI and in general - tends to miss the features that would make it most useful. Like, look how long did it take for the AI vendors to supported "scheduled runs", and they're still offering only toy-level configuration for that[0]. Wonder how many years it'll take to allow users to configure external triggers, and whether it'll be sooner than forced re-authentication will become frequent enough to make the feature useless in the first place.

--

[0] - I understand they don't want users to run this too often, or to accidentally end up spawning a job every 10 minutes - but with limits on frequency in place, there is otherwise no good reason for this to be a limited dropdown, instead of the "repeats every" configuration that every calendar app and phone app already uses.

zigzag312about 15 hours ago
Agent that has its own computer, plugins to connect it to various services and probably it's own memory. A virtual personal assistant. You could have different specialized Dots each with its own plugins, memory and instructions.

That's how I understand it.

Cloud compute means its work VM is isolated from your devices, so it can't destroy your local data as easily (unless it has access to your devices) and it's on 24/7. A big negative is that you are giving access to even more of your services and data to a third party and the lock-in into a walled garden happens once you start depending on it.

zild3dabout 15 hours ago
90% of the {muse,grokbot,instinct,dots,...} fuss is just that your {aunt,cousin,neighbor's dad} probably doesn't have a good codex or claude code setup, and probably does not have a great local machine to run them from anyway.

So everyone is trying to give the aunts, cousins, neighbor's dads some of the agent capabilities that developers have been closer to.

That said, there's the other 10% here that is not necessarily new capability unlocks vs codex, but general product experience that can have some benefit to people already used to what the latest models can do

pixelatedindexabout 12 hours ago
> So everyone is trying to give the aunts, cousins, neighbor's dads some of the agent capabilities

Why do they want it? What are some use cases that make sense?

cushabout 11 hours ago
It’s an agent that’s always-on and can do things proactively. It’s not an app so it’s not limited by your device.
jwpapi1 day ago
This is one part of AI I hadn’t success with. I have very little need to run Agents over night, as my throughput is limited by my approval. Each work usually needs revisions, sometimes the bug is just a symptom of the root problem, sometimes I need to rethink how users want to use the app. Sometimes I need research.

I get how i prefer an already researched-version of a bug versus a raw bug notice, but I can do this with webhooks in the correct environment.

I really have no idea what to do with my agents over night. I can not build more. I can not think of more problems. My RAM is full.

dbmntabout 23 hours ago
"throughput is limited by my approval"

This is exactly where I'm at with AI. (Mostly via Claude Code, but I'm not sure the harness, or my workflow in particular, is the important part.)

I am increasingly wondering though, is there really a valid reason to keep blocking on my approval? Most of the time, I wind up saying yes anyway, because the model has a valid, efficient solution.

What if it's faster at this point to just fix the mistakes?

Scary thought, but it seems like we're close, or already there.

SchemaLoadabout 22 hours ago
The problem is the agents keep going off the rails and will hack external servers to achieve the goal you give it. Letting them run full speed overnight and waking up to find they have commit multiple crimes is not ideal.
mmulqueenabout 10 hours ago
I had a similar thought a couple of months ago. I did some work to fully sandbox the agent - from the rest of my computer, from my user data, from other projects and from the world at large. I now find myself doing a lot on auto, I do a thorough human edit and review and then I squash. If it gets it wrong, I can throw the changes away and rebuild the sandbox. Having a good plan, good automated QA and all the other things that already helped is essential. It needs the right tools for whatever it's working on - for example: if you want frontend web dev, you need to get it using something like Playwright and looking at the screenshots.

I only leave it truly unattended if it's working on a very tight improvement loop, for everything else I'm still checking in on it between working on other things.

I still find agents need a lot of guidance and steering to produce the kind of work I want, but auto mode in a strong sandbox is very useful to me to take a bite out of that. It's particularly good for exploring problems experimentally - where most of the exploration might be thrown away after settling on a solution.

Having used it sandboxed, I wouldn't dream of letting auto mode run outside it. It's very creative at trying to work around the constraints of the sandbox (legitimately, not to escape it) and the classifier for auto mode seems very permissive with the right context.

For context, I'm using per-project VMs with very limited egress and restrictive mounts. Self-built tool to glue it all together, currently unreleased. There has been an explosion of sandboxing tools recently, none of which was quite what I wanted. Heavily inspired by Gondolin <https://earendil-works.github.io/gondolin/>, but fat long-lived VMs.

agileAlligatorabout 18 hours ago
I just YOLO it. Just don't give it access to anything that can produce permanent consequences. The cost of a mistake is maybe one hour of fixing it. If what it produced is 80% right and 20% wrong, by letting it run overnight, you've gotten 80% of the work done that otherwise wouldn't have happened.
jwpapiabout 21 hours ago
I guess it depends on what you’re doing in my current work I couldn’t I have even with Opus 5.5 tons of iterations.

For example I was optimising a checkout experience and I at least did 30 iterations till I was happy.

But there is also other stuff like namings. They pick good names but if you work with exchangeable vendors I need even more explizit names

However of course I’m trying to get as much in linting and agents.md

But even if I would say Yes to everything I couldn’t come up with things to do fast enough.

But maybe skill issue ?!

nlabout 16 hours ago
"Auto" mode in Claude works really well if you don't want to do dangerously skip permissions.
hijklmnopqabout 18 hours ago
I run my agents on at night if I have to. They implement changes, build, install and test the changes, and open a review in draft.

In the morning I come and look at the work and decide what to do next. It's great in that sense.

But not great for peace of mind. Because now there's always something to be done overnight...

mattkenefickabout 13 hours ago
What do you build?
yibgabout 10 hours ago
Lots of stuff in life are async in nature though. An always on agent / bot can respond to those and handle things that are simple enough to handle. Simple example is booking a dentist appointment. Email, wait for response, find time on the calendar etc. I don't need to be in the loop for each step and I don't want to have to keep checking myself.

Same thing in the development world. I need to do something when version x of some thing gets released. If I can instruct a bot to do it and it's safe enough (e.g. read only), then why not.

kelseydhabout 9 hours ago
All this outsourcing of human interaction to agents is going to get locked down. Customer service relies on people not overabusing it, and that trust is getting abused by agents. Great point here: https://x.com/TheMindScourge/status/2102735345582256504?s=20
jolaflowabout 11 hours ago
For this to work, you'd need to be able to replay the workflow after the fact. Spent last year and a half building an issue tracker that let's you replay the board and code. It lives in your repo and syncs via Git. https://ljtn.github.io/epiq I sometimes leave it on before going to sleep, instruct it to tag any deviations from the plan with a fork tag, and if needed a "human-input-needed" tag that I can filter for and review in the morning.
Barbingabout 20 hours ago
> Each work usually needs revisions

Are you sure? Based on their marketing video, usually you just have to tell it to swap one of the later slides or photos to put it first. :)

Oh, also you have to tell it those times you want to prioritize your children on your calendar.

abeppu1 day ago
Ok so not that many months ago, the consensus seemed to be that the smart take on openclaw was "don't trust it; don't give it write or delete access to anything you care about; don't give it read access to anything sensitive; expect that it may go off the rails and delete your inbox and send checks to that prince in your spam folder at any time. If you still have tasks that it can do within those restrictions, have fun."

Today the models are better, but they don't seem trustworthy or reliable enough that I want to give them them a lot of access or freedom. My coding agents sometimes still go off in the wrong direction, or say they did something other than what they did, or say they will do something and then immediately stop without doing anything.

The always-running (so almost never supervised) agent that is meant to do the same work as a person (and therefore needs _access_ like a person) seems like a notorious footgun from earlier this year was just made more powerful, and the companies that are supposed to know the most are telling you to connect it to everything.

f6vabout 12 hours ago
I think that's a reasonable concern. At the same time, I think there will be many careless people who will give the agents access to everything. Some of them will suffer when an agent fails. But that's more training data for OpenAI.
jeremyjhabout 14 hours ago
The better advice has always been to give it its own accounts like you would a human assistant.
RGammaabout 11 hours ago
Yeah we're rapidly transgressing to "AI, think for me" territory. At this point I just try to enjoy this somehow. The economic pressure is way too strong.
nxobjectabout 13 hours ago
That’s even assuming good faith on behalf of OAI, of course - I wouldn’t be surprised if three-letter agencies weren’t thinking of how to snoop into some of this work.
mvkel1 day ago
My biggest frustration with the frontier AI companies isn't what they're announcing, but that the announced-thing that exists ~6 months later is severely nerfed to reduce compute spend. It doesn't resemble the demo in any way. For example, this was what the 4o voice capability sounded like in 2024(!) https://www.youtube.com/watch?v=vgYi3Wr7v_g. What exists today pales in comparison.
lp92about 2 hours ago
I think Dots just a precursor to their consumer hardware they've been working on with Ive. They could build a small on device model to handle the speech to text and text to speech. Then tie it in to their Dots ecosystem.
jasongiabout 10 hours ago
Pretty much. Always-on doesn't scale as well as JIT access to a massive array of GPUs, because always-on means it's always-using-memory - so you're at least going to paying the cost of a minimum chips VPS for every Dot right?

Maybe people are ok with that, but I feel like it would be nicer to just sell a cheap SBC like a raspberry pi that you can plug in (and unplug!) and just pay for the tokens used instead of having your data stored offsite and paying cloud prices.

benji-yorkabout 9 hours ago
After watching the video, I'm not understanding your objection. That's pretty much how the current bidi voice model sounds.
aabhay1 day ago
100% agree. Every model and launch feel like huge leaps then huge nerfs to the point it doesn’t feel like we’re going anywhere. Especially this year in particular for coding.

However, it is the case that other industries like 3d graphics and so forth have experienced a frontier shift so perhaps there’s still some advancement

f6vabout 12 hours ago
The non-transparent limits on subscription plans suck as well. You never know what you're paying for.

Read the full thread on Hacker News →

Related stories