Epoch AI
@epochai
We are a research institute investigating the trajectory of AI for the benefit of society. epoch.ai
Who pays for the AI that US workers use on the job? We found that most AI use at work happens on free plans, except in science and tech, where employers are more likely to pay for premium access.
What are the big questions about AI capabilities that govern the impact AI will have on the world? In this Gradient Update, Greg Burnham lays out his top questions, and how he thinks Epoch's work on benchmarking can help answer them.
Investors also committed $4.5B without Broadcom’s support. That junior tranche pays 8.5%, versus 5.75% for a supported senior tranche. Because the tranches also differ in seniority, the 2.75-point gap represents an upper bound on the backstop’s value.
On the compute side, a dedicated equipment company borrows as Google TPU systems arrive, buys the racks and leases them to Anthropic for five years. Broadcom conditionally supports the lease obligations backing $30B of debt, with a reported maximum exposure of $29B.
Continuing to scale AI compute at recent rates will require exponentially more capital. Is financing a bottleneck? Probably not yet. That’s the takeaway from a new Gradient Update by Campbell Hutcheson, which analyzes nearly $50B of debt associated with Anthropic’s buildout.
How much more AI compute does a dollar buy each year? About 49%, or a doubling every 21 months, based on the chips actually bought each quarter from 2023 through 2025.
Subpar scores aren't due to budget constraints. We give each model a 1M output token budget per puzzle, but top-performing models don’t consume anywhere close to all of it. For instance, Opus 5 uses more tokens on puzzles it ultimately gets wrong, but never more than 400K.
We’ve launched a new “game puzzles” benchmark, similar to our Chess Puzzles benchmark but using puzzles from a different, undisclosed game. This helps us test AI on reasoning-heavy tasks where they probably weren’t post-trained. The current record-holder is Opus 5, scoring 59%.
Workers use most AI outputs with minimal revision. Across AI-assisted tasks, 66% of outputs were used unchanged or with only minor edits, while 5% were majorly reworked or redone. More findings: epoch.ai/publication...
AI usually handles only part of a task. But when it’s used to complete most or all of a task, workers report saving time more often (53% of such tasks, compared with 37% of tasks where AI only partly helps). More findings: epoch.ai/publication...
AI is being used across all 10 common work tasks in our poll. Adoption rates range from 25% of workers who maintain records to 57% of those who design systems and software. More findings: epoch.ai/publication...
New Epoch AI/Ipsos survey: 1 in 5 US workers say AI now handles at least one task previously delegated to humans. More findings on how AI is changing everyday job tasks 🧵
DeepSeek-V4-Flash-0731 debuts with an ECI of 153, comparable to GLM 5.2 and roughly midway between Opus 4.5 and Opus 4.6. It's the second strongest open-weights model available today, behind only Kimi K3.
Serious cyber vulnerability disclosures keep climbing. In July, 21 major tech organizations published ~2,500 high- and critical-severity CVEs — about 5× the monthly record before Anthropic revealed Claude Mythos Preview could autonomously find software vulnerabilities.
In our leaderboard configuration, models must write in the obscure programming language Ada half the time. But frontier models do almost as well in Ada as in a popular language like Go.
MirrorCode tests whether AI agents can reimplement software projects from scratch. A solve requires passing 100% of visible and hidden tests. Claude Fable is the first model we evaluated to solve the C preprocessor and Pkl tasks in at least one run.
We've updated the MirrorCode leaderboard with results for Claude Fable 5 and GPT-5.6 Sol. Claude Fable 5 leads with a 64% solve rate, followed by GPT-5.6 Sol at 20%.
We're hiring a Head of People to help us double our headcount this year! You'll own recruiting, HR, and events and lead a growing team as we scale from ~30 to ~70 people.
We thank the many mathematicians who proposed problems. We’re also grateful to our editorial board, consisting of @thomasfbloom.bsky.social, @littmath.bsky.social, and Dan Romik, who helped us vet each of the proposed problems. Here are some of their thoughts on the project overall.
We’ve launched an expansion of FrontierMath: Open Problems! The benchmark now contains 50 significant, unsolved problems from research mathematics. AI has solved three so far, and solving all of them would be an incredible mathematical feat. Thread with more.
Parallelization constraints could delay or prevent a technological singularity, even after R&D is automated. Whether, and how quickly, an explosion proceeds will depend on the development of “parallelization technology”.
GPT-5.6 Sol has been climbing Slay the Spire's Ascension ladder on our Twitch channel for a week, no human in the loop. This Thursday, Claude Opus 5 takes over the climb — live, with commentary. Thursday, July 30 · 12:30 PT twitch.tv/epochaiplays
This problem was proposed by David Roe, who had this to say about the solution.
AI has found a presentation for the absolute Galois group of the field of 2-adic numbers. This is the second problem to be solved in FrontierMath: Open Problems, our benchmark of significant unsolved problems from research mathematics.
Claude Opus 5 gets an ECI of 159, slightly below Fable 5's value of 161 (while 5.6 sol holds the record with 162). However looking only at software engineering, we find it matches Fable 5s SWE-ECI of 161.
Join the EpochAIPlays launch stream later today, with live commentary by @AlephNuul! We will be benchmarking GPT 5.6 Sol against Slay the Spire 1.
One way of looking at this is the Cyber ECI we constructed to track models’ ability to develop software exploits to hack into realistic systems. We found that Mythos and 5.6 Sol were large leaps in cyber capabilities, and we’d expect internal models are even stronger:
For example, when Irregular evaluated GPT-5.6 Sol’s cyber capabilities on their challenging FrontierCyber benchmark they found that:
Moonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it between Opus 4.6, and GPT 5.4, which released in February and March 2026 respectively, and just ahead of GPT 5.6 Luna.
We stress-tested some AI detectors and found that they rarely flag human text as AI-generated. But asking LLMs to mimic a specific author causes detectors to misclassify text as human-generated ~13% of the time. For scientific writing, false negatives rose to ~26%.