Use Ultrafast mode with GPT-6 Astra over WebSockets, with SDK examples and an HTTP alternative.

4 points•prodigycorp•1 day ago•2 comments•

2 comments

ademup1 day ago
Faster inference would solve nearly all of the issues I have with AI models. I really liked working with Opus 4.8 and I would greatly prefer a 100x faster version of it, with 100x the tokens, at the same price, to any of the openai or anthropic models that have come out since.
prpl1 day ago
4.8 at 10k tokens a second would be more groundbreaking than a Opus 6 IMO.

Read the full thread on Hacker News →

Related stories