Train a tiny GPT in under a minute (CUDA only). Contribute to lostmsu/TurboGPT development by creating an account on GitHub.
Requires CUDA 13.4; build script is Windows only.
10 comments
So, why do people keep making these? Asking genuinely
* Edit, I did however just notice literally all of this is just another vibeslopped banger, so I am not quite sure how much fun there really is, feels more like coding for the sake of keeping the wheel turning.
I extensively used minGPT for home experiments on transformer architecture. It is great for learning!
However, if you want to scale the experiments up at home you need to go faster. Karpathy made optimized https://github.com/karpathy/nanoGPT, but it is tuned for "8XA100 40GB node in about 4 days of training".
13s is a bit overkill here (my machine builds that project in 30s). But it gives some space for experimentation with architectures that don't have optimized primitives.
But it is a byte predictor. You can train it on any file.
Read the full thread on Hacker News →
Related stories
- Nemotron-H: A Family of Accurate, Efficient Hybrid Mamba-Transformer Modelsresearch.nvidia.comHacker News · 1 points · 8 days ago
- Hacker News · 1 points · 2 days ago
- Ars Technica · 0 points · 8 days ago
- The Basics of Transformer Inferencejax-ml.github.ioHacker News · 2 points · 9 days ago
- Hacker News · 44 points · 9 days ago
- How To Train Interpretable Neural Networks That Accurately Extrapolate From Small Datastochasticlifestyle.comLobsters · 1 points · over 6 years ago