The first end-to-end benchmark of computer use, continual learning, and long-horizon agentic capabilities, set in a real job.

2 points•moyicat•7 days ago•1 comment•

1 comment

drogokhal43737 days ago
that's terrifying to see the latest AI models have already beaten the best humans on knowledge works...

Read the full thread on Hacker News →

Related stories