TL;DR: We study the scaling laws of data weighting across in-house and open-weight LMs, finding non-monotonic behavior across scales. We vary the weight assi...

15 points•pranitha_m•9 days ago•1 comment•

1 comment

gwern8 days ago
Double descent?

Read the full thread on Hacker News →

Related stories