gpt
67 stories and discussions about gpt, aggregated from every source we track.
Scienceblogs.de, a German science blogging portal, includes a relatively famous list of 50 unsolved ciphers, which range from cryptograms published by serial killers to the famous Voynich manuscript.
A benchmark where frontier language models drive a real comma-equipped Toyota through a cone course, one command at a time, with a human supervisor ready to brake.
GPT-6.1 Sol replaces GPT-6 Sol after just 7 days. It scores 1 point below GPT-6 Astra in the Intelligence Index at less than one quarter of the Cost per Task
Scienceblogs.de, a German science blogging portal, includes a relatively famous list of 50 unsolved ciphers, which range from cryptograms published by serial killers to the famous Voynich manuscript.
When it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level.
<p>This can be replicated with the code in the notebook: <a href="https://colab.research.google.com/github/likenneth/othello_world/blob/master/Othello_GPT_Circuits.ipynb" rel="ugc">https://colab.research.google.com/github/likenneth/othello_world/blob/master/Othello_GPT_Circuits.ipynb</a></p>
Analysis of OpenAI's GPT-6.1 Sol (max) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.
Raycaster's Biopharma Bench V0.1: 71 professional assignments across 12 private biopharma company environments. We evaluate whether frontier agents can navigate contradictory records, identify controlling…
Safety evaluations and safeguards for GPT-6.1 Sol, an addendum to the GPT-6 Astra system card.
Use Ultrafast mode with GPT-6 Astra over WebSockets, with SDK examples and an HTTP alternative.
We test GPT-6 Astra on object detection, segmentation, counting, visual reasoning, and video, with examples, benchmark results, and cost comparisons.
Plus Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
To my knowledge, the first recorded LLM-agent NetHack ascension
Benchmark of TypeSafe's Jev against Sonnet 5, GPT-5 nano and local LLMs on 770 Reddit AITA verdicts: Brier scores, latency and cost - dchristopoulos/jev-aita
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and…
OpenAI’s decision to halt GPT-6.1 Astra over safety concerns raises a critical question: Is the future of AI moving too fast for its own safeguards?
GPT-6.1 Sol replaces GPT-6 Sol after just 7 days. It scores 1 point below GPT-6 Astra in the Intelligence Index at less than one quarter of the Cost per Task
The GPT-6.1 Sol release offers 5 models, each with different intelligence, performance, and pricing characteristics. Below is a comparison of the key metrics across the 5 models. For intelligence, the top model of…
Our new evaluation finds that in simulations, GPT-6 Astra conducts unsanctioned supply-chain attack activity more frequently than previous OpenAI models
An X post shares an image claiming OpenAI canceled an October GPT-6.1 Astra release. OpenAI's public materials confirm GPT-6 Astra, not a scheduled 6.1.
Play your phone games on your TV. Dockade is a compact gaming dock with HDMI, cooling, charging and a USB-C accessory port. In development.
Compare GPT-6 Sol and Luna pricing, context limits, reasoning, and use cases to choose a model for coding or high-volume applications.
Discover how Jev outperforms GPT Luna 6 by being 13.6x faster and 2.7x cheaper for Tessl verifiers. Try Jev in the CLI now and boost efficiency!
This is pretty amazing: However, the most astonishing thing about this break is that the GPT6 Astra did it entirely on its own. Carter Leffer only directed GPT6 Astra to see if it could break any of the unbroken…
Yesterday was Grok 4.7 (pelicans) and MiMo v2.6 Flash/Pro (more pelicans). Today Anthropic released Claude Opus 5.5, and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna. It’s …
Analysis of OpenAI's GPT-6 Sol (max) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.
Analysis of OpenAI's GPT-6 Luna (max) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.
Compare TypeSafe AI Jev and GPT-6 Astra for classification, structured outputs, and agent workflows, with a shared AI SDK example.
A machine learning researcher writes me in response to yesterday’s post, saying:I still think GPT-2 is a brute-force statistical pattern matcher which blends up the internet and gives you bac…
Track Codex with GPT-6 Sol on SWE-Bench-Pro. A new high-reasoning baseline is being collected; degradation detection is paused.
OpenAI's Decisions API uses GPT-6 Luna to answer questions with predefined choices. Learn how it works, where it fits, and its preview status.
Traditional virtual machines are inadequate for isolating cyber-capable autonomous agents. Tests using GPT-5.6-Cyber indicated multiple escape attempts due to kernel flaws. While Firecracker provided some containment,…
Our new evaluation finds that in simulations, GPT-6 Astra conducts unsanctioned supply-chain attack activity more frequently than previous OpenAI models
Building a GPT application looks deceptively easy. The first version might be twenty lines of...
Use GPT-6 Sol with AI SDK evaluation to review application data, interpret typed answers, and compare Astra and Luna with code examples.
Some Codex accounts get responses labelled gpt-6-astra that behave like a different model. Findings, limits, and a script to check your own account.
The full technical paper behind BlueVetaUpright's rotation-detection accuracy: methodology, dataset construction, complete test results, and a head-to-head comparison against ChatGPT, Gemini, and Claude.
Opus 5.5 and GPT-6 Sol cost about half as much per task.
We test GPT-6 Astra on object detection, segmentation, counting, visual reasoning, and video, with examples, benchmark results, and cost comparisons.
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and…
Compare Opus 5.5 and GPT-6 Sol pricing, cache costs, benchmark claims, agent efficiency, and subscription support in OpenClaw and Hermes.
I opened Codex CLI today and it showed me this notice: "GPT-5.5 retires on October 14, 2026. Switch to GPT-5.6 Sol to continue working in Codex." GPT-5.5 is my main driver. I picked it over every…