sol
27 stories and discussions about sol, aggregated from every source we track.
GPT-6.1 Sol replaces GPT-6 Sol after just 7 days. It scores 1 point below GPT-6 Astra in the Intelligence Index at less than one quarter of the Cost per Task
I asked GPT-3 & GPT-4 to follow instructions to create drawings in p5js and compared the results
Analysis of OpenAI's GPT-6.1 Sol (max) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.
Raycaster's Biopharma Bench V0.1: 71 professional assignments across 12 private biopharma company environments. We evaluate whether frontier agents can navigate contradictory records, identify controlling…
Safety evaluations and safeguards for GPT-6.1 Sol, an addendum to the GPT-6 Astra system card.
Plus Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
Anthropic and OpenAI both shipped a new model yesterday. We ran them through our hard cases overnight. Here is what changed, and why partforge is still on Gemini.
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and…
GPT-6.1 Sol replaces GPT-6 Sol after just 7 days. It scores 1 point below GPT-6 Astra in the Intelligence Index at less than one quarter of the Cost per Task
The GPT-6.1 Sol release offers 5 models, each with different intelligence, performance, and pricing characteristics. Below is a comparison of the key metrics across the 5 models. For intelligence, the top model of…
Compare GPT-6 Sol and Luna pricing, context limits, reasoning, and use cases to choose a model for coding or high-volume applications.
Yesterday was Grok 4.7 (pelicans) and MiMo v2.6 Flash/Pro (more pelicans). Today Anthropic released Claude Opus 5.5, and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna. It’s …
Analysis of OpenAI's GPT-6 Sol (max) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.
Track Codex with GPT-6 Sol on SWE-Bench-Pro. A new high-reasoning baseline is being collected; degradation detection is paused.
Use GPT-6 Sol with AI SDK evaluation to review application data, interpret typed answers, and compare Astra and Luna with code examples.
Opus 5.5 and GPT-6 Sol cost about half as much per task.
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and…
Compare Opus 5.5 and GPT-6 Sol pricing, cache costs, benchmark claims, agent efficiency, and subscription support in OpenClaw and Hermes.
OpenAI and Anthropic both recently released new models aimed at lowering costs. Anthropic announced Opus 5.5, the latest version of its main mass-market workhorse model, used for tasks like coding and other complex knowledge work. And OpenAI announced GPT-6 Sol and Luna, the latest versions of its middle-of-the-road or smaller models focused on efficiency and speed. Read full article Comments