LLM Benchmark - One prompt, Multiple models, Multiple dates

168 points•fragmede•8 days ago•46 comments•

46 comments

drywater28 days ago
Finally, I was getting tired of seeing pelicans on a bike.
schainks8 days ago
Up next, Pelican ass-bench!
athrowaway3z8 days ago
6 months from now we'll be discussing if model training is assmaxxing.
killingtime748 days ago
Why is this flagged? Isnt the readership adults? (Do children want to read these dry articles on tech?)
ravenstine8 days ago
I'm impressed that Luna got as far as it did on max reasoning. Don't get me wrong, that's one weird looking ass, but even the top GPT model a year ago probably would have done worse based on my experience (with asking it to create meshes in general, not modeling asses... yeah, that's the ticket). Then again, I'm also not, because I've found it to be a great deal at its relatively low price point when the reasoning is set to high.
variety86758 days ago
The idea that they're going to benchmaxx this in the future amuses me

Read the full thread on Hacker News →

Related stories