Opus Stalks the Frontier, Jev Classifies Everything, Running Two Models in One Agent

The Batch News & Insights: Moving forward on early stage, 0-to-1 projects and mature projects requires very different tactics.

Share
One engineer gathers feedback around a water cooler, while another calculates the result of a rigorous survey-driven A/B test

Dear friends,

Moving forward on early stage, 0-to-1 projects and mature projects requires very different tactics. For those who aspire to be skilled at AI Engineering, I've found that selecting the right tactic based on stage of project is one of the hardest but most important things to learn.

In the 5 letters on the AI Engineering Skills Map, you may have noticed calibration to the project stage was a recurring theme. This letter explains why. For many AI engineering tasks, like building evals, choosing software architecture, or getting product feedback, the right option usually depends on the project stage.

Take the task of evaluating an AI system — say, an automated customer-service email system. In an early-stage project, you might manually examine a dozen examples and check how sensible they are. A later-stage project might have hundreds of test examples and a written rubric for judging the quality of AI-written emails. A mature product might have tens of thousands or more test examples, a more detailed rubric, and rigorous processes for evaluating not just the quality of an email but its downstream effects (such as whether it causes a customer to be more likely to return).

Calibrating to the right stage of a project is important for many other AI Engineering skills. It is not helpful to over-design at the early stages, or under-design a mature product. For an early stage project, if your primary goal is to quickly test a product idea, a casual software architecture design might be okay, with minimal thought given to efficiency, data schemas, cost of third party services, and so on. But as a project matures, giving careful thought to tradeoffs like latency, availability, consistency, reliability, maintainability, simplicity, and cost will give you a better outcome.

Similarly, getting product feedback can range from pulling aside 2-3 people and asking what they think, to running large-scale user studies, A/B tests, and analyzing product usage data.

One engineer gathers feedback around a water cooler, while another calculates the result of a rigorous survey-driven A/B test

One way to gain experience with different approaches is to work on different projects spanning early stage and mature ones. Someone working in a startup might learn the quick ways to do evals, and someone in a large company the best practices for slower, more rigorous approaches. I have seen engineers from large companies jump into startups and ask for overly slow/rigorous approaches. Similarly, engineers from startups may move to large companies but see their applications hit a performance ceiling until they learn to go past quick, but less accurate, ways of doing evals. Project experience is valuable, but if you don't want to have to take years to gain experience in both small and large companies to master this breadth of skills, DeepLearning.AI is here to help!

Every week, I am in discussions about products that serve 100M+ users as well as products that do not yet have any users. The operating cadence is very different for these types of projects! This is why corporate policies that mandate a one-size-fits-all approach, like requiring certain types of testing before anything can be shipped, can be counterproductive.

Given that even large companies should have small, innovative projects, it's worthwhile for everyone to learn the fast, efficient tactics that let small teams move quickly. At the same time, to avoid hitting a ceiling and being unable to develop your projects beyond a certain point, it is also worth knowing how to do things in a slower, more rigorous way. 

Keep building, and I hope some of your early stage projects turn into large, successful mature ones!

Andrew 

A MESSAGE FROM DEEPLEARNING.AI

Enroll for free in Building AI Assistants with On-Device Memory

Give your AI application a memory of its experiences, stored entirely on the device. In “Building AI Assistants with On-Device Memory,” turn text and images into vectors, search them by meaning, and teach it to recognize new objects from a few photos. Enroll for free

News

Bar chart comparing AI models, showing Claude Opus 5.5 with low prompt injection risk, highlighting robustness.

Claude Opus 5.5 Leaps Forward 

A week and a half after CEO Dario Amodei proposed slowing down AI development, Anthropic released an AI model that promises to be first in a larger family.

What’s new: Anthropic introduced Claude Opus 5.5, a lower-cost successor to Claude Opus 5 that outshines Claude Fable 5.1 and all other current models in overall intelligence. Unlike Fable, it doesn’t retain users' data for 30 days, but similar to Fable, it falls back to Claude Opus 4.8 for what Anthropic deems sensitive cybersecurity and biology queries.

  • Input/output: Text and images in (up to 1 million tokens), text out (up to 128,000 tokens or 300,000 in Batch API)
  • Knowledge cutoff: June 2026
  • Features: Reasoning always on, five levels (low, medium, high, xhigh, and max, defaults to high), statistical watermarking of generated text, fast mode (2.5x speed at 2x cost)
  • Performance: Claude Opus 5.5 is first on Artificial Analysis’ Intelligence Index v4.3 (58) and leads Vals AI’s Vals Index (69.69 percent)
  • Availability/price: Via Claude.ai and external providers such as Amazon Web Services, Google Cloud, and Microsoft Azure; via API at $4/$0.25/$20 per million input/cached/output tokens; cache reads/writes $0.20/$5 per million tokens; batch processing $2/$10 per million input/output tokens; Zero Data Retention is available
  • Weights/license: Proprietary
  • Undisclosed: Parameter count, architecture, specific training data and methods

How it works: Anthropic trained the model on private and public datasets, including data from public websites gathered with their ClaudeBot web crawler, synthetic data generated by other models, and data gathered from Claude users who haven’t opted out from allowing training on their inputs and outputs. The knowledge cutoff date is identical to Claude Fable/Mythos 5.1’s, suggesting the models were trained on similar datasets. After training, the company fine-tuned the model to align with values it defined using a constitution. It was also safety-tested by evaluators selected by Anthropic, including METR and Frontier Design.

  • Anthropic’s internal alignment tests show Claude Opus 5.5 outscores every recent model and is more truthful and less likely to engage in motivated reasoning. However, the company reported that the model’s behavior appeared to change in response to tests.
  • According to Anthropic, Claude Opus 5.5 communicates more clearly and succinctly than Claude Opus 5 or Claude Fable 5, addressing a common complaint with those models. It also follows writing style instructions more closely. (Anthropic reported similar improvements for Claude Fable 5.1, which is only modestly less verbose than its predecessor.)
  • Anthropic also claims that on tests of knowledge work tasks like writing business reports, Claude Opus 5.5 passed Anthropic’s internal quality threshold on 16 of 18 attempts at various effort levels. Claude Fable 5.1 and Claude Opus 5 both failed every attempt.
  • Anthropic said Claude Sonnet 5.5 and Claude Haiku 5.5 would follow in a matter of weeks. This would be the first update for Anthropic’s faster, less-expensive Haiku-class models since version 4.5 in October 2025.

Performance: Both Artificial Analysis and Vals AI rank Claude Opus 5.5 first among all models in their weighted evaluations of overall intelligence.

  • On Artificial Analysis’ Intelligence Index v4.3, a composite of 10 evaluations of math, science, coding, and reasoning, Claude Opus 5.5 at max reasoning with default fallback scored a weighted average of 58, seven points higher than Claude Opus 5 and five points higher than Claude Fable 5.1 and GPT-6 Astra.
  • The model posts top scores on six of the ten Intelligence Index evaluations: Humanity's Last Exam (61.4 percent), SciCode (66.9 percent), GDPval-AA v2.1, AA-Briefcase v1.1, AA-Omniscience and AutomationBench-AA, and ties on a seventh, Terminal-Bench 4.0 (59.6 percent).
  • Artificial Analysis reports that while Claude Opus 5.5 costs less per token than its predecessor or Claude Fable 5.1, the model’s cost per benchmark task remains high because it uses more tokens than earlier Opus models. At max reasoning with fallback, Claude Opus 5.5 costs $5.98 per task, second only to Claude Fable 5.1 at $7.63 and well ahead of GPT-6 Astra at $3.26.
  • On Vals AI’s Index, Claude Opus 5.5 scores 69.69 percent, the top score by just over 3 percentage points, beating GPT-6 Astra. Counting fallbacks as failures did not meaningfully affect its score.
  • The model also scored first on Vals’ RSI Index (a measurement of a model’s knowledge of AI and machine learning), MedScribe (medical administrative work), ProofBench v1.1 (formally verified math proofs, where it achieved a perfect score), VibeCodeBench1-100 (extending a working web application), ProgramBench (rebuilding programs from a description), and Terminal-Bench 4.0 (terminal coding, science, and security tasks).

Behind the news: Claude Opus 5.5 arrived on the same day as OpenAI’s GPT-6 Sol and GPT-6 Luna, both less expensive models whose predecessors were rivals to Claude Opus 5, but both of which Claude Opus 5.5 now easily outperforms. These models were announced despite recent public calls from both Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman, among other leading AI figures, to slow AI development to allow for further safety and security testing. If these releases are any indication, we won’t be lacking for new, highly capable models anytime soon, even if they may come with restrictions.

Why it matters: It’s a big deal any time we have a new best model on the market, and Claude Opus 5.5 appears to be significantly better than the rest. Business customers working with sensitive data, or anyone that doesn’t want to share inputs and outputs with Anthropic, will be pleased that Fable’s data retention policies don’t extend to Opus. Claude models have long been great coders, but this model seems to be particularly good at knowledge work — creating documents and presentations, crunching data, and doing research, all areas where Anthropic had recently ceded ground to OpenAI.

We’re thinking: From a benchmarking standpoint, it’s impossible to know just how capable Claude Opus 5.5 would be, particularly at cybersecurity and biological tasks, if it didn’t fall back to Claude Opus 4.8. It’s also important that legitimate safety, biomedical, and AI engineering work may be refused out of fears that users will use the models in ways Anthropic doesn’t want them to.


Jev, marked in pink, stands out in the graph, offering high accuracy for lower cost compared to peers.

Models Built to Do One Thing Well

While most companies focus on generative and reasoning models, one company is betting on a class of models that isn’t either. All this new model does is analyze text and return answers to questions about it, but at higher speed and lower cost than a large language model. 

What’s new: TypeSafe, founded by OpenAI alumnus Diogo Almeida, released Jev, a general classification model that can answer any question with predefined outputs. It’s designed to be used to give other software tools enough data to make decisions, rather than as a general-purpose language model.

  • Input/output: Input: text, up to 64 thousand tokens. Output is defined by the user to be either a choice from a set of options, a score, or a yes/no
  • Performance: Similar performance to GPT-5.6 Terra and Claude Sonnet 5 at classifying an internal dataset
  • Availability: Currently in early access, $0.042/free per million input/output tokens
  • Undisclosed: Architecture, context window, training data, most training methods.

How it works: TypeSafe does not describe the architecture of Jev beyond saying it is not an LLM and not autoregressive, but it is transformer-based. The authors don’t describe their training data beyond saying they make it all themselves. They describe just one of their training methods, which they call “reinforcement learning for calibrated decisions” (RLCD).

  • Jev’s input is broken up into two parts: a piece of text (such as a description or a JSON object), and questions about that text. Both pieces are limited to 32 thousand tokens. Within these limits, users can add as many questions as they want. Each question is evaluated in parallel. Users are only charged for the text once as well as for all of the tokens comprising their questions. Output is free.
  • Jev’s output can be either a yes/no (true/false), a selection from a list of options, or a score between 0 and 10. The binary yes/no response is defined as the simple probability of whether the answer to the question is yes. The list selection option returns the model’s preferred choice from the list, its confidence in that choice, and a probability for each possible list item. The ten-point rubric score similarly comes with confidence and probabilities for each option, except instead of the final answer being a choice of one of the values, it computes a score based on the probability of each option.
  • In training, RLCD encouraged the model to assign probabilities to answers equal to how often they occur. For example, an answer with a 20 percent probability should be the correct answer 20 percent of the time.
  • Jev’s structure and pricing encourage users to ask many short, straightforward questions at once. For example, instead of asking Jev to classify an email message as spam or not spam, a developer should decompose the question into multiple criteria for email spam, like a domain mismatch, a request for passwords or other credentials, and so forth. The goal is to create a full classification picture that allows a software system to assess the situation and take action.

Results: The authors only computed results for their models on their internal datasets, citing a number of reasons, including benchmark saturation and companies overly focusing on improving benchmark performance rather than general intelligence.

  • Across four internal datasets (which used GPT 6 Astra and Claude Fable 5 to determine the correct answers), Jev achieved 67 percent accuracy, about the same performance as GPT-5.6 Terra and Claude Sonnet 5.
  • Across the same datasets, Jev cost about $0.0007 per example, Terra cost about $0.06 per example, and Sonnet cost about $0.12 per example.
  • Across the same datasets, the authors claim Jev is 193.6 times faster than unspecified LLMs.

Behind the news: Shortly after Jev's release, a host of similar classification models hit the market, some of them open and local rather than proprietary. Laya focuses on multilingual support, but also claims higher accuracy and faster speed than Jev. Bespoke Nimble fine-tunes Qwen-3.5-9B to act as a classification model. Kev likewise uses Qwen 3.5 as a base, but in three different sizes, and attempts to reconstruct Jev's architecture. None of the models claim to have distilled Jev. Without a public benchmark, it's difficult to assess their performance. Meanwhile, platforms like Vercel and Cloudflare quickly added Jev support, replacing costlier LLMs for use cases like tool selection.

Why it matters: Before LLMs became hugely popular, most researchers focused on training one model that can do one or a small set of tasks well. When LLMs became popular, people’s opinions flipped, and the AI community started focusing on building one model that can perform any task well. TypeSafe takes a middle road: let’s build one model that can perform any classification task well. The company wrote a catchphrase to describe its approach to development: “Build prod, not god.”

We're thinking: Jev won’t replace modern LLMs. It can’t generate code, it can’t talk to people, it can’t act as an agent. Instead, it can detect jailbreaks, flag missing details, judge user satisfaction, and more — all situations where turning unstructured input into structured responses can be tremendously valuable for software engineers.


Diagram showing main agent planning and reviewing tasks, while sidekick explores, writes code, and fixes bugs.

One Agent Works, Another Directs

Benchmarking a coding agent typically means scoring one model within one harness. Now a major independent evaluator has scored a harness that uses two models and found it matches top models’ performance at lower cost.

What’s new: Cognition introduced SWE-2, a model built for software engineering work, and made Devin Fusion, a harness that runs two models in one session, available beyond its cloud service. When using Devin Fusion, a more powerful model like Claude Fable 5.1 plans and reviews tasks, and a less costly one like SWE-2 carries out most of the work. All features below are for SWE-2, except where noted.

  • Input/output: Text in, text out
  • Architecture: Fine-tuned from Kimi K3, a 2.8 trillion-parameter mixture-of-experts model
  • Features: Three reasoning levels (medium, high, max);
  • Performance: On Artificial Analysis’ Coding Agent Index v1.5, Devin Fusion, with Claude Fable 5.1 as the lead model and SWE-2 as the sidekick model, tied Claude Code with Claude Fable 5.1 (62), at 36 percent lower cost per task
  • Availability/price: SWE-2 in Devin Desktop, Devin CLI, Devin Fusion, and Devin Web; SWE-2 free on self-serve plans through October 15, then $3/$0.30/$15 per million input/cached/output tokens. Devin Fusion on paid plans (Pro $20 per month, Teams $80 per month, Max $200 per month) in Devin CLI and Devin Desktop
  • Weights/license: Proprietary
  • Undisclosed: Context limit, knowledge cutoff, training data, and active parameter count

How it works: Instead of handing a task from one model to another in sequence, Devin Fusion runs two agents at once. A lead agent runs a planner model that oversees the session. A sidekick agent runs a cheaper model that completes lower-priority work. Each agent maintains its own tools and context.

  • The lead model resolves ambiguities in the user request, writes the plan, and reviews the sidekick’s work. For each task it delegates, it gives the sidekick a brief that sets the task’s constraints and success criteria. The sidekick agent reads, edits, and tests code and communicates with the lead agent. The lead agent reclaims a task when the returned work suggests the sidekick agent is stumbling.
  • The two agents pass each other briefs, results, and feedback, rather than entire conversations. This setup lets each agent keep its own context and prompt cache, preserving discounts on repeated inputs. Cognition argues that this step is where ordinary model routing fails. Moving a task to another model mid-session empties the cache, and refilling it at frontier prices reduces savings that routing would otherwise give.
  • In the cloud version, Fusion can change models mid-session. Lightweight classifiers run throughout tasks and flag when to give the sidekick’s work back to the lead or assign the sidekick role to a stronger model. These model swaps happen during compaction, when an agent summarizes earlier turns to shrink context. Compaction discards the cache, so the switch adds no cost.
  • Cognition trained SWE-2 for the sidekick job by using reinforcement learning to fine-tune Moonshot AI’s 2.8 trillion-parameter Kimi K3. The reward subtracts what an attempt costs, in money and time, from whether it succeeded. That let Cognition train every reasoning level in one run, where Kimi K3’s makers trained a separate expert for each reasoning level and merged them afterward.

Performance: Artificial Analysis independently ran two Devin Fusion pairings through its Coding Agent Index v1.5. Configured with Claude Fable 5.1 as the lead model, Devin Fusion matched Claude Code running the same model alone at a higher reasoning level and cost 36 percent less per task. Configured with GPT-6 Astra, Devin Fusion scored three points lower than Codex running Astra alone, again at a higher reasoning level, but cost 39 percent less.

  • On the Coding Agent Index v1.5 — an average of three evaluations of software engineering tasks, command-line tasks, and repository understanding — Devin Fusion with Claude Fable 5.1 set to xhigh reasoning as lead and SWE-2 set to medium reasoning as sidekick achieved a 62, and cost $7.90 and 35.8 minutes per task. This matched the performance of Claude Code running Claude Fable 5.1 set to max reasoning with fallback (62, $12.40 and 34.8 minutes per task).
  • Using Devin Fusion with GPT-6 Astra traded accuracy for a slightly steeper discount. Devin Fusion with GPT-6 Astra set to xhigh reasoning as lead and SWE-2 set to medium reasoning as sidekick (59, $4.54 and 24.7 minutes per task) trailed Codex, with GPT-6 Astra set to max reasoning (62, $7.47 and 29.4 minutes per task).
  • On Vals AI’s Code Migration test, which asks an agent to rewrite a program in another software language, Devin Fusion had similarly mixed results. With Claude Fable 5.1 as the lead and SWE-2 as the sidekick (57.3 percent at $42.00 per task), it outperformed Claude Fable 5.1 in Claude Code alone (54.6 percent at $70.97 per task). With GPT-6 Astra as the lead and SWE-2 as the sidekick (61.3 percent at $35.51 per task), it trailed GPT-6 Astra in Codex alone (67.7 percent at $44.36 per task).

Behind the news: Devin Fusion is not new, and Cognition is not the only company attempting to match frontier model performance at lower costs by blending multiple models.

  • Cognition described Fusion’s lead-and-sidekick design in June and ran it on Devin Cloud through the summer, reporting that the router drove 88 percent of the pull requests a set of its internal users merged. This month’s release added two things the earlier one lacked: a version of the harness that runs on a developer’s own machine rather than only inside Cognition’s cloud service, plus independent evaluation.
  • Sakana AI released Fugu Max and Fugu Ultra v2 the same day Fusion left Devin Cloud, extending the orchestrator models it launched this summer. Fugu selects a model from a pool of models for each step or subtask, sometimes several in parallel, rather than fixing a pair of models for a session.
  • The two designs disagree about a frontier model’s role. Fusion keeps one in charge of every session, planning, reviewing, and sometimes doing work. Fugu puts a dispatcher in charge instead, a low-cost model trained to split a task into subtasks and hand each to whichever model in its pool fits, so most of its capabilities comes from the collective rather than from a lead model.

Why it matters: With one model in a harness, token and dollar consumption grow together, so developers can monitor token use as a rough proxy for their bill. Fusion muddies these estimates because it burns lower-cost tokens at a higher rate. Artificial Analysis measured Devin Fusion with Claude Fable 5.1 as the lead, finding it consumed 70 percent more tokens and took nearly three times as many turns than Claude Code using Claude Fable 5.1 alone — but still cost less per task. You might think you could save money by using a sidekick model with a lower per-token cost, but that’s not necessarily true. SWE-2 appears to be the most efficient option for a sidekick model, both outperforming and costing less than models (for example, GPT-5.6 Luna) that are head-to-head more intelligent and cost less per token. In this case, developers really have to pick the right tool for the right job. 

We’re thinking: Between Cognition and Sakana, we’ve seen two very different versions of an architect/worker model architecture, but both have succeeded by training companion models that do a specific job well, whether that job is giving instructions or following orders. Model specialization remains a powerful way forward and (as the multi-model Fugu already shows) could be further decomposed to move beyond a simple two-model worker-planner structure.


A Sudoku puzzle shows a central green cell linked to blue cells, depicting note-passing between agents.

Agents Work Better When They Can Pass Each Other Notes

Large language models can work faster by dividing problems into sub-problems to be solved in parallel instead of in sequence. However, the speed-up is capped in systems that aggregate the sub-results at a single point. Researchers devised a way to remove that limit.

What’s new: A team at Carnegie Mellon University led by Xuecheng Liu and Daman Arora proposed an agentic harness called Message Passing Language Models (MPLMs). Under MPLM, an LLM distributes sub-problems among separate threads — each running its own copy of the LLM — that can communicate with one another. This approach solved two types of structured puzzles faster than alternatives.

Key insight: In earlier parallelization methods, an LLM breaks down tasks into sub-tasks while a coordinator thread spawns separate threads, assigns sub-tasks to them, and collects their output. The problem with this arrangement is that the coordinator can become bogged down in reasoning, tool calls, and the like, so the subtasks must wait. But problems in which the relationships between threads are known in advance don’t require a central coordinator. Instead, related threads can communicate with one another. Enabling threads to communicate directly makes the coordinator unnecessary and limits the total load on any one thread.

How it works: Under MPLM, a model and its iterations can write commands that start threads, send results to specific threads, wait for replies, or stop threads. The authors built programs that (i) used these commands to solve puzzles and (ii) produced text traces as they considered possible solutions. They trained Qwen3-0.6B-Base on those traces. The puzzles included examples of 3-SAT (deciding whether a boolean formula can evaluate to true) and Sudoku (filling a square grid with numbers, from 1 to the number of cells in a row, column, and box, so no number repeats in any row, column, or box). The description below applies to solving Sudoku. Solving 3-SAT involved a different process.

  • A parent thread tracked the parts of a puzzle that had been solved.
  • For each cell in a Sudoku grid, its thread determined the correct number by elimination (since a Sudoku grid comes pre-populated with some numbers and a cell can’t repeat a number already in its row, column, or box). While a thread was undecided, it waited to receive numbers from the threads that managed cells in its row, column, and box. As their numbers arrived, it ruled out numbers until one was left. Once it had settled on a number, it sent the number to the parent and related threads and stopped.
  • When every thread had stopped, the parent thread reported the solution.

Results: MPLM solved the puzzles faster and in fewer tokens per thread than two earlier approaches: a single thread and an agentic harness that runs parallel threads but routes their results through a coordinator.

  • Solving Sudoku grids from 4×4 to 25×25, MPLM was significantly faster on average. For instance, given 9x9 grids, MPLM solved 100 percent in roughly 15 seconds, whereas the parallel method solved 93 percent in roughly 60 seconds. Furthermore, MPLM’s tokens per thread grew more slowly as the grids scaled up. MPLM solved 72 percent of the 25×25 puzzles, where the other two methods reached limits of context or compute imposed by the authors before they learned to solve the problems.
  • Solving 3-SAT problems involving formulas that included 8 to 20 variables, MPLM’s accuracy was about even with the parallel method, roughly 92 versus 91 percent. It was marginally faster on average than the parallel method, but it processed some examples much faster — sometimes 2.5x faster — because once the model found a solution in one thread, it could stop the others. The single-thread method did not run beyond 12 variables, having filled the model’s context window.

Yes, but: MPLM's efficiency depends on knowing in advance which threads must communicate with which. It shines where that communication pattern is fixed and easy to work out, as it is in Sudoku and 3-SAT. For open-ended problems, the authors note that finding the pattern may take careful prompting or extra training. Moreover, the Sudoku examples were limited to puzzles that could be solved by elimination (known as naked singles), so the model never had to guess and check or reason more extensively across threads.

Why it matters: Models that perform tasks in a single thread can fill their context windows before they reach a solution. MPLM enables models to distribute the work among many threads that, collectively, can get the job done.

We’re thinking: The authors also prompted two larger models, Qwen3-30B-A3B and Qwen3.6-35B-A3B, to use their method to tackle problems on the LongBench-v2 long-context reasoning benchmark. Both models showed improved accuracy and faster responses (roughly a 2x reduction in average latency) in the MPLM harness. This suggests this method is generalizable to reasoning challenges beyond the relatively simple Sudoku and 3-SAT examples, although its effects do seem to be stronger on smaller models.