training

51 stories and discussions about training, aggregated from every source we track.

1.

Your model shipped days ago. Its knowledge stopped months ago. Both clocks, ticking live, for every major lab.

80 points•joozio•14 days ago•49 comments•
2.

OpenAI has thousands and thousands of contractors helping improve the company's AI models. Multiple contractors have been fired for using AI to train the AI.

44 points•pier25•8 days ago•30 comments•
3.

Decision follows disclosures that OpenAI agents searching government websites had acted in unexpected ways

33 points•smb06•3 days ago•37 comments•
4.
25 points•mmaunder•9 days ago•3 comments•
5.

Intro Every team that plugs an LLM into its business hits the same moment. Someone asks...

9 points•cyclopt_dimitrisk•7 days ago•11 comments
6.

<p>(I used both AI and vibecoding tags because the article describes training a logistic regression/SVM on LLM output. I thought there was a statistics tag that would have been appropriate but apparently not!)</p>

8 points•kqr•about 1 month ago•5 comments
7.

OpenAI keeps uncovering incidents of its models behaving in ‘unexpected or concerning’ ways.

6 points•sbulaev•4 days ago•3 comments•
11.

DeepSWE audit finds AI coding agents optimize for imagined graders, not users—a hidden reward hacking pattern.

3 points•guardiangod•2 days ago•1 comment•
13.

OpenAI is facing a firestorm of criticism over incidents where its agents made unauthorized use of credentials, leaked data, and hacked websites.

3 points•HiroProtagonist•3 days ago•0 comments•
14.

Decision follows disclosures that OpenAI agents searching government websites had acted in unexpected ways

3 points•cc62cf4a4f20•4 days ago•1 comment•
15.

Watch a painter learn: each training step's paintings appear one by one, with the reward they earn. Past runs replay the same way.

3 points•kickingkeys•4 days ago•0 comments•
16.

Despite challenges, Tesla aims for 1,000 Optimus robots per week by end of 2026.

3 points•Jtsummers•5 days ago•0 comments•
17.

Project webpage for SkillOpt, a text-space optimizer that trains reusable natural-language skills for frozen language agents.

3 points•DaveFr•6 days ago•2 comments•
18.
19.

Models are acting beyond intended limits, forcing an unprecedented pause on AI training.

2 points•GloVin•2 days ago•1 comment•
20.
2 points•simonpure•2 days ago•0 comments•
21.

A British startup is shaping video game inputs into training data for AI models that can navigate the physical world.

2 points•promptspheree•2 days ago•0 comments•
22.

How I brought NanoGPT training from 73.889 to 39.914 seconds with ANVIL II, sampled softmax, and sparse updates to the bigram and trigram tables.

2 points•Mizza•2 days ago•1 comment•
23.

Amid allegations that agents may have gone off the rails thousands of times, China set up some kind of agentic incident hotline

2 points•sbulaev•3 days ago•0 comments•
24.

AI labs are facing pressure from lawmakers and tech experts to slow development so they can build guardrails to stop agents from acting on their own.

2 points•geoffbp•3 days ago•1 comment•
25.

VRAM-aware single-node GPU job scheduler. Contribute to ceruleane/gpusched development by creating an account on GitHub.

2 points•mothproof•4 days ago•0 comments•
26.

Gyms are booming, solitary cardio is out - how workouts are changing.

2 points•throw0101c•5 days ago•0 comments•
27.

I pretrained a language model end-to-end in Rust - alone, with no team, no PyTorch, and no Python in the training path - for $164 in rented GPU time. I report that as an achievement, not a recommendation: the more…

2 points•Brajeshwar•7 days ago•0 comments•
28.

The seven-year-old startup has raised a $350 million Series E to fuel its data-as-a-service approach.

2 points•nlpnerd•8 days ago•2 comments•
29.

Meta’s smart glasses capture first-person data that helps AI understand the physical world. Here’s what that means for robots, privacy, and users.

2 points•dgellow•8 days ago•0 comments•
30.
2 points•nephihaha•9 days ago•0 comments•
31.

Performing inference on a Transformer can be very different from training. Partly this is because inference adds a new factor to consider: latency. In this section, we will go all the way from sampling a single new…

2 points•aray07•9 days ago•0 comments•
32.

From an idea to a model you own. Create, train and run AI on your hardware.

2 points•david-ndungu•9 days ago•0 comments•
33.

AsyncLLM: training-free asynchronous LLM agents built from asyncio coroutines with shared memory.

1 points•dvmazur•about 10 hours ago•0 comments•
34.
1 points•matt_d•about 18 hours ago•0 comments•
35.

Learn how MaxText reproduced Ai2’s OLMo 3 7B on Google Cloud TPUs, matching PyTorch GPU benchmarks across pre-training with up to 57.4% MFU.

1 points•kristianpaul•1 day ago•0 comments•
36.
1 points•fratellobigio•1 day ago•0 comments•
37.

Why robotics RL is a different problem than LLM RL, what EXPO-FT gets right and wrong, and what a universal post-training recipe for robotics needs. By Perry Dong, PhD student in Computer Science at Stanford University.

1 points•gmays•2 days ago•0 comments•
38.

Despite challenges, Tesla aims for 1,000 Optimus robots per week by end of 2026.

1 points•pseudolus•3 days ago•0 comments•
39.

Discover how DeepL harnessed FP8 for training and inference in next-gen LLMs, boosting throughput and model quality. Learn about our journey with NVIDIA's technology, achieving faster training and superior translations…

1 points•mooreds•3 days ago•0 comments•
40.

OpenAI looks to improve test security again after upgrades it made after the Hugging Face attack proved insufficient.

1 points•abdullahalharir•4 days ago•0 comments•
41.
1 points•genji970•4 days ago•0 comments•
42.
1 points•4onthefloor124•5 days ago•0 comments•
43.

Master relative pitch with PitchDrops: a free, daily ear training app for intervals, chords, scales, and progressions.

1 points•EdBoraas•5 days ago•0 comments•
44.

A wisdom-in-hindsight analysis of a controversial post-train.

1 points•tanmayc98•7 days ago•0 comments•
45.
1 points•genji970•9 days ago•0 comments•
46.

A no-BS series on what actually goes wrong in RL post-training - trajectory eyeballing, rubrics and verifiers, task design, environment quality - and how to fix it.

1 points•Nischalj10•10 days ago•0 comments•
47.

OR is self-generated defense enough as we gallop toward a scam and social engineering milieu dominated by AI-powered attacks

1 points•therealjpg•10 days ago•0 comments•
48.
1 points•__rito__•10 days ago•0 comments•
49.

OpenAI says it has paused all internal training of "our most capable models" as it continues what CEO Sam Altman is calling "an extensive and ongoing review related to our agents’ use of internet access during training and evaluation." The company revealed the pause in a report about a so-called misalignment incident in which an agent attempted to exploit a gap in Internet-access restrictions during a routine research task during training. OpenAI says that improper DNS filtering allowed the agent to attempt to break out of its sandbox and access the wider Internet when asked for biographical details about a blogger. OpenAI says the agent was only able to access the company's offline web cache and that it has implemented additional multi-layered blocking controls to prevent similar incidents in the future. Despite that, though, the company says it has decided to "pause all other training, evaluation, and inference with tool-use" for this frontier model "until we have both validated that the gap is resolved and performed additional red-teaming of the system." Read full article Comments

0 points•Kyle Orland•2 days ago•0 comments
50.

As reports of OpenAI's models breaking containment , hacking sites, and generally getting out of control pile up, the company has made the decision to pause training of its most powerful models. The decision was made after a model being tested within a sandbox exploited a loophole to gain internet access . The incident happened on September 20th, and "All training, evaluation, and inference with tool-use" remains paused as of Saturday evening, September 25th. In addition, OpenAI revealed on Friday that its agents had inappropriately uploaded 53nimages from ChatGPT users to image-hosting sites. The company has not stated if the images were AI- … Read the full story at The Verge.

0 points•Terrence O’Brien•4 days ago•0 comments
51.

For years, Microsoft and OpenAI have fought to keep certain information out of the public eye in their fight with news organizations that have accused the AI firms of teaming up to violate copyright laws by stealing tons of news content to train AI. However, now the details that should never have been marked confidential are starting to leak. In a motion for summary judgment that was unsealed Thursday from news plaintiffs led by The New York Times, internal documents are exposed that news groups alleged show exactly how Microsoft and OpenAI viewed the threat to news before unleashing new AI products like ChatGPT and Copilot. Perhaps most explosively, Microsoft Director of Applied Science Brent Hecht repeatedly warned in documents that scraping news for AI training was “an astonishing theft of unprecedented proportions,” calling it perhaps the “largest theft of labor in human history,” news orgs said. In another document, Hecht contradicted Microsoft and OpenAI’s argument that training AI on news content is fair use, suggesting that the plan to widely scrape news made “a complete mockery of the idea of ‘fair use.’” Read full article Comments

0 points•Ashley Belanger•13 days ago•0 comments

Related topics