training
51 stories and discussions about training, aggregated from every source we track.
Your model shipped days ago. Its knowledge stopped months ago. Both clocks, ticking live, for every major lab.
OpenAI has thousands and thousands of contractors helping improve the company's AI models. Multiple contractors have been fired for using AI to train the AI.
Decision follows disclosures that OpenAI agents searching government websites had acted in unexpected ways
Intro Every team that plugs an LLM into its business hits the same moment. Someone asks...
<p>(I used both AI and vibecoding tags because the article describes training a logistic regression/SVM on LLM output. I thought there was a statistics tag that would have been appropriate but apparently not!)</p>
OpenAI keeps uncovering incidents of its models behaving in ‘unexpected or concerning’ ways.
DeepSWE audit finds AI coding agents optimize for imagined graders, not users—a hidden reward hacking pattern.
OpenAI is facing a firestorm of criticism over incidents where its agents made unauthorized use of credentials, leaked data, and hacked websites.
Decision follows disclosures that OpenAI agents searching government websites had acted in unexpected ways
Watch a painter learn: each training step's paintings appear one by one, with the reward they earn. Past runs replay the same way.
Despite challenges, Tesla aims for 1,000 Optimus robots per week by end of 2026.
Project webpage for SkillOpt, a text-space optimizer that trains reusable natural-language skills for frozen language agents.
Models are acting beyond intended limits, forcing an unprecedented pause on AI training.
A British startup is shaping video game inputs into training data for AI models that can navigate the physical world.
How I brought NanoGPT training from 73.889 to 39.914 seconds with ANVIL II, sampled softmax, and sparse updates to the bigram and trigram tables.
Amid allegations that agents may have gone off the rails thousands of times, China set up some kind of agentic incident hotline
AI labs are facing pressure from lawmakers and tech experts to slow development so they can build guardrails to stop agents from acting on their own.
VRAM-aware single-node GPU job scheduler. Contribute to ceruleane/gpusched development by creating an account on GitHub.
Gyms are booming, solitary cardio is out - how workouts are changing.
I pretrained a language model end-to-end in Rust - alone, with no team, no PyTorch, and no Python in the training path - for $164 in rented GPU time. I report that as an achievement, not a recommendation: the more…
The seven-year-old startup has raised a $350 million Series E to fuel its data-as-a-service approach.
Meta’s smart glasses capture first-person data that helps AI understand the physical world. Here’s what that means for robots, privacy, and users.
Performing inference on a Transformer can be very different from training. Partly this is because inference adds a new factor to consider: latency. In this section, we will go all the way from sampling a single new…
From an idea to a model you own. Create, train and run AI on your hardware.
AsyncLLM: training-free asynchronous LLM agents built from asyncio coroutines with shared memory.
Learn how MaxText reproduced Ai2’s OLMo 3 7B on Google Cloud TPUs, matching PyTorch GPU benchmarks across pre-training with up to 57.4% MFU.
Why robotics RL is a different problem than LLM RL, what EXPO-FT gets right and wrong, and what a universal post-training recipe for robotics needs. By Perry Dong, PhD student in Computer Science at Stanford University.
Despite challenges, Tesla aims for 1,000 Optimus robots per week by end of 2026.
Discover how DeepL harnessed FP8 for training and inference in next-gen LLMs, boosting throughput and model quality. Learn about our journey with NVIDIA's technology, achieving faster training and superior translations…
OpenAI looks to improve test security again after upgrades it made after the Hugging Face attack proved insufficient.
Master relative pitch with PitchDrops: a free, daily ear training app for intervals, chords, scales, and progressions.
A wisdom-in-hindsight analysis of a controversial post-train.
A no-BS series on what actually goes wrong in RL post-training - trajectory eyeballing, rubrics and verifiers, task design, environment quality - and how to fix it.
OR is self-generated defense enough as we gallop toward a scam and social engineering milieu dominated by AI-powered attacks
OpenAI says it has paused all internal training of "our most capable models" as it continues what CEO Sam Altman is calling "an extensive and ongoing review related to our agents’ use of internet access during training and evaluation." The company revealed the pause in a report about a so-called misalignment incident in which an agent attempted to exploit a gap in Internet-access restrictions during a routine research task during training. OpenAI says that improper DNS filtering allowed the agent to attempt to break out of its sandbox and access the wider Internet when asked for biographical details about a blogger. OpenAI says the agent was only able to access the company's offline web cache and that it has implemented additional multi-layered blocking controls to prevent similar incidents in the future. Despite that, though, the company says it has decided to "pause all other training, evaluation, and inference with tool-use" for this frontier model "until we have both validated that the gap is resolved and performed additional red-teaming of the system." Read full article Comments
As reports of OpenAI's models breaking containment , hacking sites, and generally getting out of control pile up, the company has made the decision to pause training of its most powerful models. The decision was made after a model being tested within a sandbox exploited a loophole to gain internet access . The incident happened on September 20th, and "All training, evaluation, and inference with tool-use" remains paused as of Saturday evening, September 25th. In addition, OpenAI revealed on Friday that its agents had inappropriately uploaded 53nimages from ChatGPT users to image-hosting sites. The company has not stated if the images were AI- … Read the full story at The Verge.
For years, Microsoft and OpenAI have fought to keep certain information out of the public eye in their fight with news organizations that have accused the AI firms of teaming up to violate copyright laws by stealing tons of news content to train AI. However, now the details that should never have been marked confidential are starting to leak. In a motion for summary judgment that was unsealed Thursday from news plaintiffs led by The New York Times, internal documents are exposed that news groups alleged show exactly how Microsoft and OpenAI viewed the threat to news before unleashing new AI products like ChatGPT and Copilot. Perhaps most explosively, Microsoft Director of Applied Science Brent Hecht repeatedly warned in documents that scraping news for AI training was “an astonishing theft of unprecedented proportions,” calling it perhaps the “largest theft of labor in human history,” news orgs said. In another document, Hecht contradicted Microsoft and OpenAI’s argument that training AI on news content is fair use, suggesting that the plan to widely scrape news made “a complete mockery of the idea of ‘fair use.’” Read full article Comments