A benchmark where frontier language models drive a real comma-equipped Toyota through a cone course, one command at a time, with a human supervisor ready to brake.

316 points•plurby•7 days ago•251 comments•

251 comments

jyoung86077 days ago
I'm not an expert in the LLM space, but I'm an external contributor to comma.ai's openpilot project and I'm and quite familiar with how its controls work, so I looked from that perspective. There's two questions here:

1) Could a cloud-delivered LLM figure out how to drive this route, based on those input data and given access to those output actuators? Looks like yes. Sure.

2) Could this work in the real world? Absolutely not. Three reasons: latency, latency, and latency.

openpilot's driving model updates the target curvature and acceleration at 20Hz. Every millisecond of the round trip time through every piece of its entirely-local driving stack is well-understood, extremely consistent, and tightly optimized. It has to be, otherwise you can't react to even minor bumps or wind gusts, much less rapidly-developing traffic situations.

Adding even a single speed of light RTT to a cloud service is meaningfully bad, and you'll need a whole lot more to encode and upload camera imagery to even start the time-to-LLM-response clock, and then send the response back down. By then the world around the car has moved on.

There's a reason Tesla and every other self-driving manufacturer need the compute hardware in the car.

aditya-ramabadr7 days ago
Great point! Yeah latency was one of the biggest issues here. To cope with that (and for safety reasons) the cars are driving at extremely low speeds. They also get timestamps with every tool call output etc so they can, in theory, "in context learn" about their own latency and choose motion durations and control how fast their iteration loop is to some extent. But yeah, this is just sort of a fun benchmark to see how good frontier LLMs are out-of-the-box at driving a real car, and probably not actually practical any time soon.

-Aditya, Tobias, Simon

jyoung86077 days ago
To clarify my parent comment, I think it was an interesting experiment and seems like it was done well, and it may well be informative about what various frontier LLMs could do with recorded or world model footage.

My only point is to say this sort of experiment is where it ends. Neither Anthropic nor OpenAI will be coming out with a "drive your car from the cloud" subscription until we have FTL communication, meaning never.

bayarearefugee7 days ago
> Could this work in the real world? Absolutely not. Three reasons: latency, latency, and latency.

That and also the fact that (in spite of their usefulness) LLMs still so often do incredibly dumb shit without thinking of the consequences that the idea of having them drive in public is absurd.

Recently was using claude code/opus 5 to diagnose an intermittent wi-fi connection problem and one of the first things it did was to bring the adapter down. The wi-fi adapter was the only way the system was communicating with the outside world so claude effectively disconnected its own brain as step 1 in figuring out what was going wrong. Things did not progress well from there. Easy enough to clean up its mess in this case, but luckily it wasn't driving a heavy killing machine at the time.

fragmede7 days ago
Let he whomst amongst us, that hath never committed such a sin, cast the first stone.
NewsaHackO7 days ago
>Recently was using claude code/opus 5 to diagnose an intermittent wi-fi connection problem and one of the first things it did was to bring the adapter down.

Do you mean restarting it? IDK, that would have been my first step too.

ivanjermakov7 days ago
> otherwise you can't react

I'm far from neuroscience, but humans don't need to operate at 20Hz to drive a car. And human reaction latency (event to measurable action) is often over 1s (under 1Hz).

chaos_emergent7 days ago
The reaction latency you’re referring to for humans includes perception, planning, and actuation, I’d separate that from the concerns of the hardware, which are mostly about actuation frequency.

From what I understand about AV (as a non-expert!), all three of those steps happen at different clock rates, ie you have a planner that’s updating continuously with observations from sensors at one rate, that planner then issues actions that get picked up by the actuators at another rate.

In that sense 20hz should really be compared to human reflexes without perception and planning; in scenarios where one is anticipating an action, response time can be as low as 150ms. in that context, I think 50ms/20hz is plenty reasonable for an automated driver.

mirrir7 days ago
I'm not an ornithologist but birds don't need to consume jet fuel to fly hundreds of miles either.
bonsai_spool7 days ago
> but humans don't need to operate at 20Hz to drive a ca

This is not a helpful statement unless you can claim what speed human sensors do work at. And it's going to be faster than the latency of $(sensor + server round trip) Hertz, not getting into LLM processing time.

jacquesm7 days ago
Humans have multiple layers of processing such inputs and your subconscious reacts a lot faster than your conscious train of thought in case something happens (and then you have to 'catch up'). For the same reason that you don't consciously think about what you do when you are walking or how to stop yourself from falling when you stumble. That's all out of the top level and pushed further down to stack, sometimes even multiple levels.
cozzyd7 days ago
Let's see how well you play counterstrike with a 100 ms ping...
blactuary7 days ago
Kind of funny to mention comma today of all days
aaroninsf7 days ago
Is that the same comma.ai project also in the news today?

https://arstechnica.com/cars/2026/09/aftermarket-driver-assi...

famouswaffles7 days ago
What did they do to Astra so cracked at vision (and computer use). That ARC 3 score turned out to be no joke/fluke. That huge gap between Astra and Fable (in this case) is basically every hard vison/spatial benchmark i've seen including non-benchmarks like playing games (Portal, Factorio, RimWorld).

SpatialBench - https://x.com/spicey_lemonade/status/2096365630190698516

ZeroBench - https://zerobench.github.io/

Robot Arms - https://openai.robocurve.org/gpt-6-astra/

smusamashah7 days ago
I think Opus 5.5 is at same level now. I have seen too many videos made by Opus 5.5 today on twitter.

https://x.com/victormustar/status/2102707412704919910 horse galloping pixel art

https://x.com/LexnLin/status/2102133072585965759 moving train pixel art animation

https://x.com/jkeatn/status/2102441348075057539 painting with code

https://x.com/LCSlates/status/2102503027340988559 video, very detailed prompt though

https://x.com/aj_dev_smith/status/2102504509637587339 generated song/music with code

https://x.com/aj_dev_smith/status/2102575577563570450 another song

prideout7 days ago
These are amazing but the parent comment is referring to vision comprehension, not generation.
famouswaffles7 days ago
I don't think so. These are cool but all of this is code to x. I'm talking about actual computer control.

Stuff like: - https://x.com/iam_zachi/status/2095992132620136677

Puzzles, games, painting software, robotic control and now driving. I haven't seen any other model fire on all cylinders like that.

bigwheels7 days ago
Do these examples demonstrate new levels of computer use capability?
dyauspitr7 days ago
Well, they have the best in class image generator so that probably has something to do with it
valine7 days ago
The bitter lesson is finally coming for the self-driving cars. The vision stack, 3D maps, lane selection grammar, occupancy networks, it’s maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment.

It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.

jvanderbot7 days ago
You might be interested to learn that the bitter lesson has already been grok'd by generations of autonomous car company engineers, and many or all have incorporated learned components (at minimum) in all their vehicle stacks.

There's also a very tangible limitation of the bitter lesson.

If, over time, compute climbs, and so compute-bound data-driven general architectures beat bespoke architectures (this is the bitter lesson), then it is not necessarily true that the most general architecture now beats all available bespoke architectures now (or even in the near/mid future - the crossover point is "eventually").

Bitter lesson is most tangible for long-running research directions. Sometimes you need something working as best as possible now.

AlphaSite7 days ago
Yeah. Every major self driving model that I’m aware of is fully e2e at this point. Going from fused sensor output to control+debug vectors.

This is more generalised.

But also since there’s a huge volume of data it’s too expensive to just keep scaling compute up (per car overhead) so there are necessary tricks involved.

I do think having a large model that can do this means that a small specialised model could be distilled form it though. Which is probably the most feasible path to production IMO.

seanmcdirmid7 days ago
A later entrant can potentially side step those investments if their now is later. Since self driving car ventures aren’t profitable yet and need to make up their investments over time, thats a real risk for them.
torginus6 days ago
But what might happen imo, is that these huge models might be much better at learning from data in an unsupervised manner, so a model based on Astra might become a much better driver in a short period of time (and given a smaller set of training data) - this is due to it understanding much more of the inputs, and being able to draw conclusions from it much more efficiently, thus the information available for training it is greater per sample.

Then the big model can teach a small model to become almost as good a driver. This might be substantially more efficient way to train stuff, and might be fairly quick and straightforward.

In practical terms, I feel this means we can see huge jumps in capability overnight. And this is a general indicator of AI progress, not only in this narrow scope.

VBprogrammer7 days ago
I'm not sure how you take that from the original article. My 4 year old would drive that course in an automatic car, if only he could reach the pedals. Heck, he's done harder things at Lego land.

I wouldn't let him loose on the road though.

I think, at the very least, the guardrails would have to deterministic, ideally with super human senses, for people to accept self driving cars on the road.

chaos_emergent7 days ago
Your four year old has billions of years of learning embedded in his weights/architecture for generalized motor control :)
Davidzheng7 days ago
Maybe 4 year olds already have the brain power to learn to drive if they so desired.
an0malous7 days ago
Interesting. Is your 4 year old raising capital?
ACCount397 days ago
Nope, no "deterministic guardrails" for you. The domain is simply far too broad and unstructured to allow for that.

Unless you mean "a typical AI with all the computation constrained sufficiently to always unfold the same exact way, given the same input". In practice, that just kicks the can to "given the same input" street.

The noise in the system is going to come from the input plane. Which is, I remind you, facing the real world. It's full of noise.

SoftTalker7 days ago
If I'm reading the chart correctly, it took over 5 minutes to drive 135m at a cost of nearly $8.00 in tokens. I don't think that's really in the realm of practical yet.
atonse7 days ago
Tesla's already solved this - their vision model does this phenomenally well.

And they've demonstrated adding a sidecar LLM to it as well, mostly for these kinds of "read these 3 street signs, what should i do next?" sort of situations.

matt_heimer7 days ago
The same Tesla that pulled radar to go vision only and a person was killed because the vision model didn't recognize a truck? https://www.bbc.com/news/technology-36680043

Not sure that counts as phenomenally well.

robots0only7 days ago
What do you think Tesla has been doing this for so long?
archagon7 days ago
Busting unions, discriminating against black employees, and funding Musk’s virulent white supremacy.
sschueller7 days ago
Lying and they still are.
prometheus19927 days ago
Wow! but WHY is this a benchmark?? for comparison tesla's model is approximately 10-15B parameter model (estimating from maxxing the hardware that comes with the car at 16gb ram).
N_A_T_E7 days ago
I would assume this is a proxy for general intelligence. A model that can drive a car and do a bunch of other real world stuff is closer to a generalized intelligence that can reason through any task.
jrflo7 days ago
Tesla isn't using a general purpose model, they're using many highly-specialized models for a more deterministic system than "hey chat drive this car for me"
therealdrag07 days ago
Why not? Benchmark all the things!
syntaxing7 days ago
Surprised they didn’t try Qwen’s recently open sourced driving model https://huggingface.co/Qwen/Qwen-Drive-1.0-4B
aditya-ramabadr7 days ago
Also heard about this! But the point of our benchmark was to evaluate frontier LLMs with vision out-of-the-box, which we wouldn't expect to have been specifically trained on driving real cars. The fact that they can do anything at all (even in an open lot cone course, at low speeds) is pretty impressive. I'm sure Qwen Drive and models specifically trained for driving would do even better.

- Aditya, Tobias, Simon

syntaxing7 days ago
Interesting, I think it would be interesting to gauge how a 4B model would run compared to a frontier one

Read the full thread on Hacker News →

Related stories