SpaceXAI's most powerful model for coding and knowledge work. Twice as fast, at half the price of comparable models.
539 comments
Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.
However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.
My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
Grok has its own feel too. It's not as bad as Claude, but one of the things that bugs me is that it is far too terse.
It regularly seems to come up with terms and descriptions for things in its chain of reasoning and then uses these terms in its output assuming you understand what it's talking about.
I find I often have to ask it to re-explain what it means.
GPT does this all the time, too (both Sol and Astra). I constantly have to tell it to not use terms that were not part of the initial prompt.
Separately have been using Grok 4.6 for a bit and it's also pretty concise.
I’ve noticed Astra doing this a lot as well.
First, after a while it's just as grating as Claudeish. Second, my hunch is that it constricts the actual thinking of the LLM, like the same way that Newspeak does in 1984. It shrinks the range of thought that can be expressed if used as an input.
I think the real way to do it is to have another Claude entirely deal with the user as a liaison, but to keep the thinking in whatever format it came in.
Latent space reasoning, if you think about it, is exactly this to a crazy degree: why even formulate a thought as words if you can just keep it as matmuls until the user needs it? And then, if the user needs it, have it always specifically formulated for the user by another LLM rather than constrict its range of thought? Anyway, that's my take.
It’s been really productive and I’ve been asking my agents to communicate using it more and more. I believe it’s relieved my cognitive load a bit while working with them.
That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?
The weird thing is, that's not what AI models seem to be doing. The prose is just weird.
I should try adding these tips to my system prompt. Is there a shorthand to describe such language use? I am not a native English speaker.
When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.
Part of intelligence is knowing your audience and communicating efficiently.
I am a huge fan of Gemini Pro for chat... gemini somehow just knows the most obscure stuff. I'll double check something Gemini said and find the source is deep inside a hard to access scientific paper. Google just has the best index of the internet.
Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
UPDATE: I tried again with the xAI API directly: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.
For comparison here's a fresh run against Grok 4.6: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
It used to be a mess in various interesting ways. Now, almost every big release can draw something perfectly functional.
So the question - without a correct answer - given the prompt "Generate an SVG of a pelican riding a bicycle":
Does the user want the least lines of code to make it functional, or the best looking version?
4.7 is definitely slower & more expensive. It feels kind of like they really had it burn tokens to claw up the benchmarks. But it's not super clear to me whether it's above the line or not. A part of that is that it is so slow that i haven't been making fast progress today with benchmarking it.
Overall, it it gets above my intelligence line its a good release...but you can read the tea leaves and tell the Grok team thinks this was a miss.
For example?
4.6 made more mistakes than SOL or Opus overall. Gave up a lot. And in my opinion, the rate of mistakes is kind of more important than how brilliant it is.
I think 4.7 may still be better, but I was hoping for clearly Sol/Opus level and so far it just isn't there for me.
This is such cope.
For now, I doubt anyone would notice your protest if you didn't announce it.
If Elon hadn’t worked with Orange Man Bad, then the Left would still be in love with him for his massive former donations to the Democrat political machine, and his work against climate change.
The whole “he’s a nazi” accusation is banal, and people are seeing through it now. That’s why we’ve moved on.
Read the full thread on Hacker News →
Related stories
- Grok 4.7 Intelligence, Performance and Price Analysisartificialanalysis.aiHacker News · 5 points · 9 days ago
- The Verge · 0 points · 1 day ago
- Hacker News · 2 points · 8 days ago
- Hacker News · 1 points · 10 days ago
- DEV Community · 22 points · 11 days ago
- The Verge · 0 points · 8 days ago