Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to…

20 points•mrkn1•about 13 hours ago•3 comments•

3 comments

vatsachakabout 9 hours ago
How do frontier labs choose which ideas to include in a training run? It seems that there are too many to choose from
martianvoidabout 9 hours ago
I think maybe they have enough compute to try out all the promising directions at a smaller scale first and then port them to larger run

They also have incredibly smart people that can understand which improvements and directions are even worth pursuing for, I think these people are being paid in millions

simianwordsabout 9 hours ago
Why do you think they are buying compute and accelerating to RSI? It’s for this.

Read the full thread on Hacker News →

Related stories