1777 points•OfficialTurkey•8 days ago•855 comments•

855 comments

simonw8 days ago
GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

Here's GPT-6 Luna pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

And GPT-6 Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Scroll to the bottom for the GPT-6 Sol max one: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

For comparison, here are the pelicans I got for GPT-6 Astra: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - I still like the Astra Max one best.

Here's a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: https://static.simonwillison.net/static/2026/gpt-6-and-5.6.h...

The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.

gizmodo598 days ago
6-luna is at the pareto for most of the tasks! I dont know how they make money here but its insane value from a closed source model. I'd go further and say it makes no sense (privacy, sovereignty etc aside) to use many other models as its not only expensive but also many providers don't have that much GPUs to serve at a significant volume. https://openrouter.ai/rankings?view=month#top-models 5.6 luna is already the most used model this month.
sieve8 days ago
My OpenCode Go stats for the last 30d:

Cached Read: ~6,500M

Input: ~150M

Output: ~20M

Approx $40 worth of usage across DeepSeek V4 Flash + MuseSpark Contributor 1.3. And a bit of both the GLM models. This is covered in a $10 subscription.

If I were to use Luna's API pricing:

$0.02 x 6,500 = $130

$0.20 x 150 = $30

$1.20 x 20 = $24

So $184. And this is assuming smaller coding sessions (<272K) beyond which Luna pricing doubles.

--

Cost wise, these models are nice for small stuff. Translations etc. Any model that does not provide multiple Mtoks of cached reads per cent is not very useful to me for coding workflows.

booty8 days ago

    I dont know how they make money here
Well, here's the neat thing: they don't!

Snark aside, Luna 5.6 was (is) an incredible game-changer.

krat0sprakhar8 days ago
Can't agree more. Between 5.6 Luna and Gemini 3.8 flash I'm so happy for the value I'm getting for my dollar (subscription pricing not API pricing) :)
user439288 days ago
6-luna is no improvement over 5.6, merely a price cut.

And info from the help page with message limits suggests the 50% price cut does not apply to the subscription, where they applied only a 1/3 price cut instead.

I'm not thrilled with this release.

Opus 5.5, which matches GPT-6 Astra performance at a cheaper price, is much more interesting.

InsideOutSanta8 days ago
> I dont know how they make money here

By raising it from investors.

matznerd8 days ago
Simon, love your work, one piece of minor feedback for the individual model pages is to make the font of the model name potentially bigger than (and above) the conversation id (which means nothing to the audience) "2026-09-22T18:28:00 conversation: 01m355zvyw8946qyraa8zpz6h9 id: 01m355zvyx47zxx5c6q6b3fg0m#".

I had all the tabs open individually and harder to scan which model is which... otherwise keep up the great work! I like the grid view a lot. (Also the pages have no OG images set, which impacts what the link looks like shared)...

simonw8 days ago
That's a good idea. It's the default output for my `llm logs` command, but that header could at least show the model ID.

OG images will require me to move away from publishing in a Gist and linking to from a JavaScript page that loads the Gist. Probably worthwhile though.

Cu3PO428 days ago
I find it very interesting that for both these models we such a clear progression of better images with higher thinking levels from 'hardly useful' to 'pretty nice'. I feel on many other models low and max are much closer.
saretup8 days ago
Not that this benchmark is super relevant anymore but these look worse than I expected.
simonw8 days ago
Yeah, it's interesting how much worse they are than the Astra pelicans. I think that reflects a tiny bit of genuine value still left in the benchmark, to be honest.
alansaber8 days ago
It would be extremely funny if the explosion in SVG generation capability in particular was a result of this benchmark
gtirloni8 days ago
What's the relevance of the pelican benchmark when models probably saw it during training? Didn't OpenAI stop testing against SWE-Something because it was tainted?
simonw8 days ago
If they train for the benchmark, how come many of the pelicans produced by their different models at different reasoning levels still suck?

That aside, the relevance these days is in comparing models and effort levels within the same model families - hence the comparison grids.

genidoi8 days ago
It's not a benchmark, it is a meme benchmark.
m_fayer8 days ago
I've been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It's the first model I've gotten attached to. I'm concerned that whatever model supercedes it, while technically better, just won't feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.
NorthSouthNorth8 days ago
Completely agree. I've been using 5.6 still even with Astra available to me for most tasks. It's funny how much of this is just "vibes" because I cannot quantify what it is. Astra is definitely better when I have an ambitious feature, but in like 9/10 tasks I prefer working with 5.6 Sol. A few weeks ago when the limits were seemingly higher, having 5.6 on fast mode was a good time.
bryanhogan7 days ago
I have also been using 5.6 Sol instead of 6. I found 6 to burn through my usage incredibly quick, making it somewhat unusable because I wouldn't be able to get anything done.

My results with 5.6 Sol were quite similar to 6, although I haven't tested it that much.

fnordpiglet8 days ago
I have issues with astra having a full task list in front of it and doing an Opus 5 move and announcing it’s about to begin then end the turn and wait. Typically I can get it to work one step at a time then stop. It’s maddening. 5.6 was a workhorse.
jauntywundrkind8 days ago
Astra is 100% conpletionist no chill alien.

It wants things beyond what the mortals (us) know to reach for. It's not good at explaining itself, it doesn't show it's thinking. It's often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.

capital_guy8 days ago
I tend to agree. it's by far the best coding model i've ever worked with, including astra and if i remember correctly fable, and it's unbelievably smooth at just getting the work done and communicating in simple terms.

if GPT 6 Sol is just 5.6 at half the price it will be everything i really ever wanted.

manojlds8 days ago
Does the price really matter when you are on subscription? Are we getting more usage or are we getting same usage and the cost for openai is lower?
apitman8 days ago
Similar for me. gpt-5.6-sol high has been my go-to for months. One of the reasons I'm pushing myself to try open models more is because it lends some level of guarantee I can continue to use the same tool as long as I want to. And I think we may just be getting to the point the open models are >= 5.6 Sol for coding.
redox998 days ago
Same. In fact I found 6 Astra to be a downgrade in situations where I didn't need the extra intelligence.
cmrdporcupine8 days ago
Yeah.

Astra was/is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague/lazy instructions ("I need to be able to test this on Windows, maybe a qemu VM or something? Shrug." ... 1 hour later "yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve & verify harness.").

And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.

But it also feels sloppier? Somehow. And too expensive to use.

We'll see how Sol 6 is.

Rapzid8 days ago
Yeah, I use Astra for destroying vaguely scoped asks and tasks, and then for high-level design and plan generations..

Otherwise I'm using 5.6 Sol for actual plan execution and review..

amluto8 days ago
I use Astra for rapidly consuming my token limit on a task that would not consume it on 5.6 Sol.

(I have not done anything quantitative here. For one thing, OpenAI’s billing pages and the codex-rs frontend make it pathetically difficult to get any real data. Some day I should wire up a proxy to extract actual stats.)

danabramov8 days ago
Same. The way I would describe it is that I can mostly leave 5.6 Sol overnight and trust that it makes good progress, maybe stumbling a bit and needing some correction for the remaining 20%.

If I leave Astra overnight, I'll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.

jijijijij8 days ago
The A in Astra stands for ADHD. It's featuring a neurodiversal net.
jeffnash8 days ago
At this point, the deciding factors for me between Claude Code 20x and Codex Pro 20x are:

1/ Usage limits: downstream of input/output cost, but resets and obscure windows and odd 20x plan / 5x plan != 4x usage math throw a wrench into it. Winner right now is Codex by a mile, especially when you factor in the fact that ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan. Always a bummer when asking if I should see a doctor about a rash means I can't code as much. It's also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems, giving better planning results or deeper code analysis without burning usage.

2/ Context window in the harness. Claude Code wins on this. There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing (ETA: noname120 pointed out this is no longer the case and it can be enabled again [1]). 252k is just not enough. Codex's compaction is very good, fwiw, but it happens so frequently that even a model as powerful as Astra sometimes loses the plot on long-running tasks.

3/ Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.

I've subscription hopped a bunch, and at times I've had both, but I keep coming back to Codex because it wins on 2/3.

ETA: apparently I haven't been Keeping Up With the Altmans and new 20x signups have been disabled for a few weeks. I am grandfathered in, which makes the comparison above pretty much moot.

[1]https://news.ycombinator.com/item?id=49806060

glub8 days ago
> Usage limits [...] Winner right now is Codex by a mile

This hasn't been the case since around July. If you measure usage in raw api costs, Anthropic is actually giving more on $200 than OpenAI now. This includes resets. Usage allocation difference would be humiliating for codex subs were it not for resets. But fixing usage limits with resets is ugly, and they're not good for your mental well-being.

> Context window in the harness

Codex now allows 1M for subs with config params. But generally speaking, you shouldn't really be using 1M context. If you accidentally send a request with say, ~700k context already accumulated in a session which is outside cache TTL, you're paying full cost of these 700k tokens.

> I've subscription hopped a bunch

OpenAI actually has a new strategy to prevent subscription hopping after their 2-3 month-long marketing push to get claude-folks to switch over:

you can't buy a $200 sub anymore. So if you cancel, you won't be able to get back in. Hostage situation, essentially.

EDIT: re: usage limits, oh-my-pi maintainer has been tracking this - https://nitter.xitter.cc/_can1357/status/2090075496948060372

rudedogg8 days ago
I’ve been a Claude user, switched to Codex expecting usage limits to be more loose but I can’t even get through a basic sysadmin task on the $20 plan using Sol medium before I hit the 5hr one.

I think I’m gonna move back to a Claude plan. I could barely hit the $200 limit if I went non-stop on programming tasks.

athrowaway3z8 days ago
I'm not sure the tokens can be compared like that between OpenAI/Anthropic.

When i swapped between a 200k Fable context into an Astra model (i was out of fable) the token usage in that context dropped to 150k or something.

Either there was a bug somewhere, or the same text got cut up very differently between providers.

cameronh908 days ago
To add my anecdote, while the Codex subscription appears to get you much fewer tokens as measured by cost, I find the amount of actual useful work that can be done by both subs to be about equal. Codex seems much less prone to burning millions of tokens just reading the codebase and doing nothing useful. That also makes it much quicker. Plus it actually does what I tell it with few mistakes first time, so less rework needed.

The Claude TUI is just so much better though so I'm hoping Opus 5.5 is actually good and not just benchmaxxed.

platinumrad8 days ago
Given that Anthropic models are very verbose and OpenAI models can be very concise, wouldn't a count of expected task completions be a better measurement than raw API costs?
matheusmoreira8 days ago
Anthropic has a separate meter for Fable. I used to get like five Fable sessions per week and that's it.

OpenAI has no such nonsense. No separate meter. No five hour limits. I get to use Astra at max effort on literally every task if I want to, and even this somehow lasts me several days.

Anthropic got caught playing stupid "20x refers to the 5h limit" word games with their customers. Meanwhile, I have statistically verified that OpenAI Pro 20x = 4 * Pro 5x = 20 * Plus, exactly as advertised.

I quantified cybersecurity lockouts on my code review benchmark and they were significantly lower on OpenAI:

https://www.matheusmoreira.com/articles/code-reviewing-lone-...

My benchmark also suggests even OpenAI's Sol models can match Fable performance at a fraction of the cost.

OpenAI also used to have a ton of very nice features: unlimited chat separate from codex, allowing turns to finish even at 0% usage remaining. Sadly these got removed after abuse.

As a former Anthropic customer, OpenAI is simply the better company. There is no way around it. Good place to be while the chinese open weights models catch up. Claude is good but it doesn't make up for Anthropic's shenanigans.

elxr8 days ago
Also, OpenAI is just a company I'd rather support than Anthropic.

While you're understandably not including the values of the $20 standard plans on both, I find the generosity of then token limits on ChatGPT plus vs Claude Pro (it's a huge difference) to be good representation of their respective attitudes towards the average user. You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious.

Also, Anthropic has zero models comparable to Luna.

InsideOutSanta8 days ago
> Also, OpenAI is just a company I'd rather support than Anthropic.

They're both pretty horrible, but I find it difficult to find arguments for why Anthropic is worse than OpenAI, other than their doomtrolling. Which, in the grand scheme of things, doesn't even register.

Edit: forgot about the SpaceX thing.

ketzu8 days ago
> You literally cannot use Claude pro to build real software

Interestingly I would have drawn the exact opposite conclusion looking at my Claude and codex usage.

I can't get anything sustained out of codex in chatgpt plus, while I have been using Claude pro extensively and put on a lot of experimental task and features.

I ran into codex exhausting a 5h window on code review in minutes (like 3minutes) multiple times, while I could get Claude to implement 2~3 medium sized features with the same usage consumption.

(I also really dislike the usage resets in codex, they always make me feel like I use them wrong because I often just want to reset the 5h window, but they can only do both at once...)

therein8 days ago
They are both companies I'd rather not support. Not that our support for them has any material impact. NVIDIA is bankrolling them directly and indirectly.
bix68 days ago
Reasons for this?

> Also, OpenAI is just a company I'd rather support than Anthropic.

felixgallo8 days ago
You'd rather literally support <i>Sam Altman>/i>? I mean, that's a position to take, for sure, but apparently several people still use Grok, so maybe it's not all that surprising.

"You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious" - that's way past ridiculous. Even just using Fable most of the time, working on several ambitious projects, I have a hard time hitting the limit with a Max plan.

noname1208 days ago
> It's also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems

As far as I know Codex (at least the GUI) can automatically call the ChatGPT Chat models (including Astra 6 Pro), you just need to @ a ChatGPT Chat conversation from within Codex and tell it when to use it.

> There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing

Not true, it works again[1]. I confirm that it works both on 5.6 Sol and Astra 6, possibly other models too.

[1] https://x.com/thsottiaux/status/2089082893804896524

jeffnash8 days ago
I actually haven't played with the GUI. I probably should now that the Linux version is in beta. My situation is kind of the reverse: I like using oracle to basically zip up my repo, ask GPT Pro to propose some sort of design or refactor based on the code, then provide a step by step implementation plan for a cheaper model to implement directly in a harness on my machine. It often takes upwards of 90 minutes to come up with something but I've never been disappointed by the results. I suppose I could do this and then save a step by referencing the oracle-created thread with the @ you mentioned

And re: the toml workaround, AWESOME! I appreciate you pointing these two things out, this is my highest-ROI HN comment thus far.

hintymad8 days ago
> Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.

I'm quite puzzled about why Anthropic is so hellbent on blocking other coding agents. It's not like Claude Code has any secret sauce, right? And doesn't Anthropic make monkey off API usage, and their magic is on the model side anyway?

glub8 days ago
It's for lock-in - same reason why it took them so long to finally support AGENTS.md.

But to be fair, they don't really enforce the harness rule that much anymore. I guess if your harness doesn't do a lot of weird things like a lot of cache misses, or triggers some distillation attacks, or some broader Chinese fingerprints, they're tongue-in-cheek okay with you using a third party harness.

spacebanana78 days ago
This feels like a horrible precedent. Billing based on data like commits feels like it opens the door to tech stack based billing in general - could we see different prices for people who use other devtools Anthropic doesn't like? Makes me feel grateful for open models
InsideOutSanta8 days ago
They want to lock people into using the Claude Code ecosystem to make switching to other providers more difficult.
nl8 days ago
Originally it was because Anthropic was so compute constrained they relied on the extra care the Claude harness took with caching (heavy use of cache breakpoints etc) that other harnesses didn't.

I think that is less of a factor now, and I think Anthropic have backed off some on being as strict (eg, AFAIK they never implemented the two-tier "claude -p" pricing model they were planning)

saralily6 days ago
Meridian and DirectSDK work well to use a Claude MAX subscription in alternative harnesses.
sodacanner8 days ago
In my personal experience I currently get a lot, lot more usage on the 5x Claude plan than the 5x Codex plan.

Having limitless webUI ChatGPT usage is much better user experience, though. I'll give them that.

(edit: Sol-6 is half the price, so maybe the usage limits are going to be way better.)

basisword8 days ago
I've been using Claude Pro and recently gave Codex a try again. Both on the $20 plans. I get so much more usage with Claude. It's night and day for me. Codex runs out constantly, whereas Claude I hit limits very rarely.
jeffnash8 days ago
I'm actually interested to see how the token discount maps to the usage limit consumption. The conspiracy theorist in me wonders if they're making up the discount and resultant load increase on the API end by reducing effective usage on the subscription end.
leokennis8 days ago
From the perspective of “an average person”, ChatGPT is delivering fantastic products.

- For general chat and web search, occasional image editing, small coding work, document review etc. ChatGPT Plus is basically limitless and “just works” since 5.6. I’ve yet to give it some task it cannot do.

- When given sensible instructions, it hardly annoys with weird phrasing, glazing, or annoying constructs.

- The apps are very good (ignoring the initially terrible Codex app)

It’s easily my best spent $23 a month.

jeremyjh8 days ago
You can get a lot of Codex usage out of that same sub on top of ChatGPT usage. Its a really good value and you can use that sub in any harness. In OMP I have Sol high as the orchestrator, Sol max as Planner & Reviewer, Luna max as task/coder. Very good setup. I'm on pro now and there are weekends when I use half a week's usage but I'll have 5 or 6 sessions going at once for many hours each day.
lionkor8 days ago
I find that subagents usually burn more tokens and take longer, and produce about the same quality. A real killer use-case is using a VERY cheap subagent to do a lot of work, or reviews. Don't be fooled into thinking that a "scout" subagent will gather enough info for a "coder" agent to just start working.
sergiotapia8 days ago
This is quite interesting, I wasn't aware omp has a way to set up different models for planner/orchestrator/task.
shepherdjerred8 days ago
What is OMP?
PestoDiRucola8 days ago
Not even for the average person. Luna is an amazing model for most coding tasks.
maxnevermind8 days ago
> From the perspective of “an average person”, ChatGPT is delivering fantastic products.

It is a honeymoon still, enshittification is coming, who knows how that will look like given how much more expensive to run LLMs backed user experience. Some back of the envelope calculations: 300 million US users * 20$ a month * 12 months = 72 billion $ a year. 72B$ is some spare change for AI labs. That assuming entire US will be subs which is unlikely and outside of the US there are not many rich countries consuming it, India is the next market, then Brazil and Philippines I think, not super rich counties to say the least. I believe total revenue to just pay for the capex build out by the end 2027 should be on the scale of hundreds of billions a year.

mikeg88 days ago
Analysis totally excludes enterprise customer demand and or paid API usage which will only increase as apps integrate this into future knowledge work workflows.
jr35928 days ago
> ignoring the initially terrible Codex app

Still needs a LOT of work IMO.

arnaudsm8 days ago
The current bugginess of Codex is the proof that OpenAI hasn't "achieved AGI internally" yet.
XCSme8 days ago
I use a lot of ChatGPT remotez and 70% of the times is unusable and buggy (prompts disappear, a lot of errors, buttons don't work, etc.)
pookieinc8 days ago
I don't see how anyone can be using Claude with prices like this, it's pretty incredible what the OpenAI team is doing, w.r.t model quality and pricing.

  Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
   Cache reads              $0.20              $0.50
   Input tokens             $4                 $5
   Output tokens            $20                $25
   Cache writes             $5                 $6.25


Model

Input

Output

Price reduction

GPT‑6 Sol vs. GPT‑5.6 Sol

$4 → $2

$20 → $10

50% cheaper

GPT‑6 Luna vs. GPT‑5.6 Luna

$0.20 → $0.10

$1.20 → $0.50

50% cheaper

hombre_fatal8 days ago
I mainly use Codex/Sol to review my plans drafted by Fable. But beyond that, Astra blows through usage limits too fast to be a daily driver and writes weird code despite what my "house style" is, and Codex is behind Claude Code in terms of critical features like seeing what's going on in subagents.

The parent + subagent workflow has become critical for keeping the reasoning agent (parent) context-lean while also letting me chat to the main agent while work is getting done.

My main process is to use Fable to reason and then spawn Opus subagents, and I get amazing results, and I'm always looking into what the subagents are doing.

jorl178 days ago
Astra is:

- Unbearably slow

- A token eating machine like no other

- Constantly compacting

- A model (like other GPT ones) that hides thinking traces and thinking summaries, which infuriates me

I've been in the Claude camp for a while, but the way it writes has left me with a a brick for a brain and wanted to see if Astra was as good as they say. Well, I can't know, because in the time it takes for it to actually build anything useful, I've moved to other ideas.

Unbearably, annoyingly slow. I keep thinking I must be doing something wrong.

hoangnnguyen8 days ago
If you want a mix between both codex/claude code/pi for leveraging different models and harnesses, you can give ai-devkit agent orchestration a try
mfiguiere8 days ago
Also, batch processing prices are still 50% off, which put GPT-6 Sol and GPT-6 Luna at $5 and $0.25 for output.

https://developers.openai.com/api/docs/pricing?latest-pricin...

joshstrange8 days ago
As someone who has used Claude Code and Codex the prices don't matter in the same way but I found that I burned through my usage way faster on Codex even though I regularly hear that the Codex plans go further. That was not my experience and the intelligence was comparable to what I was getting in Claude.

If these price changes mean that coding plans have effectively more usage then that's great, but Codex is surviving on resets from my own experience using it. I was glad to go back to Claude.

etothet8 days ago
For API usage, sure. But plenty of people have subscriptions where these differences effectively don’t matter.
esafak8 days ago
It should matter; if their costs go down you'll get more usage.
mchusma8 days ago
Opus 5.5 is incredible so far, its going to get used. Fable is much better than Astra for me in practice, and Sol is not marketed as better.

Its a great release, I will use both heavily.

Read the full thread on Hacker News →

Related stories