How I brought NanoGPT training from 73.889 to 39.914 seconds with ANVIL II, sampled softmax, and sparse updates to the bigram and trigram tables.

2 points•Mizza•2 days ago•1 comment•

1 comment

Mizza2 days ago
Less mathy discussion on Twitter, that also covers some of the speedup that comes from outside the optimiser https://x.com/classiclarryd/status/2104688738354028962

Massive result

Read the full thread on Hacker News →

Related stories