We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends…
0 comments
No comments yet.
Related stories
- Hacker News · 2 points · 4 days ago
- Ars Technica · 0 points · 2 days ago
- Hacker News · 88 points · 14 days ago
- Scaling to 100k Usersalexpareto.comLobsters · 18 points · over 6 years ago
- Lobsters · 6 points · about 14 years ago
- Hacker News · 6 points · 5 days ago