We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends…

2 points•thoughtpeddler•6 days ago•0 comments•

0 comments

No comments yet.

Related stories