Claude Opus 5 is a workhorse: Big Tech agrees: The world needs open weight AI models

Independent benchmarks for Claude Opus 5. Task crossover, or using AI to complement your main job. The expenditure horizon, where agents cost less than humans. U.S. threatens sanctions, China threatens back.

Share
Diverse team collaborates on laptops in modern office with whiteboard flowchart, enhancing productivity and teamwork.

In today’s edition of Data Points, you’ll learn about our top headlines, and more:

  • Independent benchmarks for Claude Opus 5
  • Task crossover, or using AI to complement your main job
  • The expenditure horizon, where agents cost less than humans
  • U.S. threatens sanctions, China threatens back

But first:

Claude Opus update gives near-parity with Fable at lower price

Anthropic released Claude Opus 5, a model that matches Claude Fable 5’s abilities on many tasks while costing half as much to run. On coding benchmarks like CursorBench and Frontier-Bench, it performs within 0.5 percent of Claude Fable 5 while reducing cost per task significantly. Claude Opus 5 now powers Claude Max by default and is the most powerful model available to Claude Pro users. It trails Claude Mythos 5 only on cybersecurity exploitation—a gap Anthropic engineered intentionally by skipping offensive cyber training—and matches Mythos on vulnerability detection, while triggering cyber safeguards 85 percent less often than Claude Fable 5 does. The model maintains a one-million-token context window and introduces five effort settings that allow users to trade computation for performance. (Anthropic)

Tech companies argue in support of open AI models

Microsoft published an open letter signed by 80 tech companies, including AI giants OpenAI, Meta, Google, and Nvidia, arguing that open weight AI models are essential for American technological leadership and economic competitiveness. The argument draws a direct parallel to open source software: Just as shared code built the internet’s foundation, openly available model weights will diffuse AI into factories, hospitals, and main-street businesses by letting organizations run specialized models locally instead of paying frontier prices for every task. The letter pushes back against regulatory proposals that would restrict open models in the name of safety, contending that transparency makes AI more secure by allowing broad scrutiny of vulnerabilities and that closed models create concentrated points of failure. The signatories urge policymakers to expand compute access for startups, invest in shared training infrastructure, and avoid conflating legitimate techniques like distillation with unlawful model extraction. (Microsoft)

Claude Opus 5 surprisingly outpaces Fable on overall intelligence

Anthropic released Claude Opus 5, which scores as the most intelligent model on the Artificial Analysis Intelligence Index with a maximum score of 61, just ahead of Claude Fable 5 with fallback (60) and GPT-5.6 Sol (59). Claude Opus 5 leads on agentic knowledge work benchmarks, scoring 1861 Elo on GDPval-AA v2 and 1720 Elo on AA-Briefcase, outperforming Claude Fable 5 by over 100 points on the former and 146 on the latter. On cost, Claude Opus 5 (max) averages $2.03 per Intelligence Index task, below Fable 5’s $2.75, though higher than cheaper models like Claude Sonnet 5. One notable weakness is recall: Its factual knowledge still lags Fable 5, and it hallucinates more frequently (50 percent hallucination rate on AA-Omniscience, up 14 points from Opus 4.8). (Artificial Analysis)

OpenAI study finds people use AI to work across roles

OpenAI’s analysis of over 800,000 work-related messages from U.S.-based ChatGPT users reveals that 43.5 percent of occupation-specific messages involve tasks traditionally associated with other jobs—a pattern the company calls “task crossover.” The data shows workers are increasingly performing work beyond their formal roles: a marketer might troubleshoot a website, a salesperson could analyze customer data, or a small-business owner might draft contracts. The shift is most pronounced in customer experience, design, and HR roles, where outside-occupation tasks account for 75 to 77 percent of use. Marketing and engineering tasks travel farthest across occupational boundaries, appearing frequently in messages from workers in unrelated fields. Smaller organizations show higher task crossover rates, suggesting AI fills gaps where specialized teams aren’t available. Job descriptions and organizational structures may reorganize around these emerging patterns before formal labor statistics catch the shift. (OpenAI)

Measuring the cost of agents vs. humans at the same task

METR introduced the “expenditure horizon,” a metric that calculates the cost threshold where AI agents become more expensive than humans at solving the same problems. The organization tested the metric on the NanoGPT speedrun—a community project optimizing language model training—where humans have achieved roughly one-percent speedups at an estimated cost of $2,500 each. When six AI models worked on the same task with up to $10,000 in compute budget, results were mixed: GPT-5 and Opus-4.1 produced no real progress, while GPT-5.5 and Opus-4.8 delivered one to 1.5 percent improvements, with expenditure horizons in the low four figures. Those gains barely dent the estimated $250,000 in total human effort on the project so far. However, METR only tested older models; newer releases like Opus 5, which shows dramatic improvements on problem-solving benchmarks, weren’t included and could shift the picture substantially. The study’s biggest blind spot is that it measures purely autonomous AI optimization, not the hybrid human-plus-AI workflows that actually drive most research. (The Decoder)

China and U.S. stand off over expected economic sanctions

China’s Ministry of Commerce warned it will take “all necessary measures” if the US sanctions Chinese AI companies for allegedly using American models to train their own systems through a technique called distillation. The statement came after Treasury Secretary Scott Bessent announced the US would scrutinize Chinese models for intellectual property theft, backed by complaints from OpenAI and Anthropic that competitors built rival systems at a fraction of the cost by distilling their outputs. China defended distillation as a widely used industry technique and accused the US of applying double standards, calling the scrutiny an act of “AI hegemony” lacking factual and legal basis. The ministry also alleged without evidence that American AI companies have distilled Chinese models, while calling for dialogue on AI governance. (Bloomberg)


Want to know more about what matters in AI right now?

Read the latest issue of The Batch for in-depth analysis of news and research.

Last week, Andrew talked about a significant cyberattack involving a closed model, the role of open models like GLM 5.2 in defending against such attacks, and the ongoing debate over AI safety and openness.

“I believe that open weight models, and more generally openness — despite some companies falsely saying it is dangerous — casts sunlight on technology and ultimately makes it safer.” 

Read Andrew’s letter here.

Other top AI news and research covered in depth:


Start today with DeepLearning.AI Pro!

DeepLearning.AI’s first-ever subscription plan for our entire course catalog includes foundational classics like the Machine Learning and Deep Learning specialization plus seminars on the latest tools and frameworks you need.

As a Pro Member, you’ll immediately enjoy access to:

  • Nearly 200 AI short and long courses from Andrew Ng and industry experts
  • Labs and quizzes to test your knowledge
  • Projects to share with employers
  • Certificates to testify to your new skills
  • A community to help you advance at the speed of AI

Enroll now to lock in a year of full access for $25 per month paid upfront, or opt for month-to-month payments at just $30 per month. Both payment options begin with a one-week free trial.

Explore Pro’s benefits and start building today!

Try Pro Now


Data Points is produced by human editors with AI assistance.