Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to…
3 comments
vatsachakabout 9 hours ago
How do frontier labs choose which ideas to include in a training run? It seems that there are too many to choose from
martianvoidabout 9 hours ago
I think maybe they have enough compute to try out all the promising directions at a smaller scale first and then port them to larger run
They also have incredibly smart people that can understand which improvements and directions are even worth pursuing for, I think these people are being paid in millions
simianwordsabout 9 hours ago
Why do you think they are buying compute and accelerating to RSI? It’s for this.
Read the full thread on Hacker News →
Related stories
- Show HN: Every Credit Cardevery-credit-card.numberplanet.comHacker News · 1 points · 10 days ago
- Hacker News · 1 points · 13 days ago
- Hacker News · 20 points · 6 days ago
- DEV Community · 9 points · about 12 hours ago
- Hacker News · 2 points · 9 days ago
- Hacker News · 1 points · 10 days ago