We Should Be Encouraged By AI’s Cybersecurity Abilities: Why GLM-5.3’s High ExploitBench score is good news
The Batch News & Insights: Anthropic recently released an encouraging analysis of the cyber capabilities of open weight model GLM-5.3.
Dear friends,
Anthropic recently released an encouraging analysis of the cyber capabilities of open weight model GLM-5.3. The study shows that GLM-5.3 approaches the cyber capabilities of Claude's closed weight Mythos. For example, the diagram below shows that, using a comparable number of tokens, GLM-5.3 succeeded in 12% of attempts to exploit vulnerabilities on a subset of ExploitBench tasks, compared to Mythos’ 14%. (Interestingly, the gap reported by Anthropic is smaller in GLM-5.3’s creator's report on the full benchmark: 54.4% success compared to Mythos' 78.0%.) This paves the way for cyberdefenders, including ones that do not have access to Mythos, to use a highly capable model to defend themselves. Given attackers’ access to similar capabilities, the urgency of ramping up defenses grows.
Further, the cost of finding vulnerabilities and using it to create an attack (that is, to create a payload that targets the identified vulnerability) is falling. Anthropic found that spending $20.40 on GLM-5.3-Flash tokens was sufficient to find a recently disclosed flaw in Google Chrome. The results reported used an unmodified version of GLM-5.3. It is also easy to obtain frontier open weight models that have had their guardrails weakened. Such models would be even more effective for cyberdefense (or cyberoffense).

A lot of AI risks, including cyber vulnerabilities and how to contain agents, are engineering problems to be solved. To be clear, these are hard engineering problems! But I find that a lot of the fear-mongering scenarios implicitly assume that we will make little or no progress on them.
For an analogy, consider aviation. According to the National Postal Museum, in 1919, one person died for about every 115,000 miles flown. Today, airlines fly about 6 trillion passenger miles. So, airplanes kill about 6 trillion /115,000 = 52 million persons a year, right? Because we have engineered airplanes to be much safer, we now see one death approximately every 30 billion passenger miles.
Similarly, even though mass media describes the OpenAI-Hugging Face hack as dangerous agents “going rogue,” this incident is leading teams everywhere to engineer better sandboxing and monitoring – important changes that are making AI safer. (On OpenWorker, our agent harness that supports security workflows, Rohit Prsad, Devika Verma and I are building on Nvidia's new safety framework to improve sandboxing.)
In aviation and in software, we improve safety by repeatedly finding problems, fixing them, and further identifying and addressing root causes that reduces the odds of future problems. Frontier models are speeding this process up.
The advanced cyber capabilities of open weight models do give attackers a window to do damage. Long term, defenders have the advantage, because they have more information with which to discover and mitigate vulnerabilities. The key is to carry out this important and difficult engineering work as quickly as we can, so we can all come out the other side with safer, more robust software.
Keep building!
Andrew
A MESSAGE FROM DEEPLEARNING.AI

The Data Engineering Professional Certificate is now available on DeepLearning.AI. Across four courses, design and build systems that generate, ingest, store, transform, and serve data, including batch and streaming pipelines on AWS and open-source tools. Taught by Joe Reis, co-author of Fundamentals of Data Engineering. Enroll for free
News

An Unexpected Open Weights Leader
Xiaomi, best known for its smartphones and electric vehicles, released the highest-scoring open weights model on Artificial Analysis’ Intelligence Index. It performs similarly to GPT-6 Sol, but costs less per task.
What’s new: Along with MiMo-V2.6-Pro-RL and its smaller sibling MiMo-V2.6-Flash, the company also released MiMo-V2.6-Distill-Qwen-9B (Alibaba’s Qwen3.5-9B fine-tuned on MiMo outputs). In addition to the models, Xiaomi also published more than 7,000 reinforcement learning (RL) task environments (software workspaces in which a model attempts tasks and receives feedback) and the code to train on them.
- Input/output: Text, images, video, and audio in (up to 1 million tokens), text out (up to 128,000 tokens, 44.4 tokens per second)
- Architecture: Mixture-of-experts transformer; 1.02 trillion parameters, 42 billion active per token for MiMo-V2.6-Pro-RL; 309 billion parameters, 15 billion active per token for MiMo-V2.6-Flash
- Features: Optional reasoning (on by default), tool calling, optional UltraSpeed mode that generates output up to 20 times faster at 10 times the cost per token
- Performance: MiMo-V2.6-Pro-RL ranks first among open weights models on Artificial Analysis’ Intelligence Index (46), and MiMo-V2.6-Flash (59.58 percent) ranks first among open weights models on Vals Index, with MiMo-V2.6-Pro-RL (59.47 percent) just behind it, within the margin of error
- Availability/price: Monthly subscriptions from $6 to $100; MiMo-V2.6-Pro API $0.435/$0.0036/$0.87 per million input/cached/output tokens, UltraSpeed mode $4.35/$0.036/$8.70 per million input/cached/output tokens; MiMo-V2.6-Flash API $0.14/$0.0028/$0.28 per million input/cached/output tokens; batch processing at half the standard price of each model
- Weights/license: All three models free to download for commercial and noncommercial use under the MIT license
- Undisclosed: Knowledge cutoff, specific training datasets
How it works: MiMo-V2.6-Pro processes images, video, and audio through separate encoders that feed its language model. According to Xiaomi’s technical report, the training recipe combines methods from earlier work, much of it Xiaomi’s own. The main innovation rewards coding attempts for quality rather than only for passing tasks.
- Xiaomi pretrained the model first on 27 trillion tokens of text from web pages, books, academic papers, code, and STEM material, then on 3 trillion tokens of text, images, video, and audio for multi-modality. In a mid-training stage, it continued training on records of agents performing coding, visual, and research tasks and extended the input context to 1 million tokens.
- After a brief round of supervised fine-tuning, the company fine-tuned the model in one RL run that mixed all task types rather than training separate specialists. Coding accounted for 68 percent of the tasks, while design of websites and other visual artifacts (games, 3D scenes, slides, SVGs) accounted for 13 percent; tool use accounted for 12 percent, cybersecurity 4 percent, and context following 3 percent. The RL run used Group Relative Policy Optimization, which generates several attempts at each prompt and scores each one against the others. Each training step drew 1,568 prompts and generated 16 attempts per prompt, for around 25,000 total attempts. The model attempted tasks in variations of one minimal agent harness, configured differently for general, coding, visual, and cybersecurity tasks, so success wouldn’t depend on using a particular harness. The RL run took 30 steps and just over 123 hours.
- Automated tests that scored coding attempts during training reveal whether code works, not whether it’s well made, so Xiaomi additionally used AI models to grade passing attempts by quality. For coding tasks the model often passed, an agent studied a set of earlier attempts and wrote task-specific checklists for the quality of its approach. During training, each attempt’s reward equaled its test result (1 or 0) multiplied by its two checklist scores. For coding tasks the model did not often pass, a grader model reviewed all 16 attempts at a task together, inspected the repository, and ran tests as needed. It ranked passing attempts by criteria such as the approach’s suitability, how few changes it made, and consistency with the codebase’s conventions. Training then reinforced higher-ranked attempts more strongly and lower-ranked ones less.
- To counter reward hacking, such as downloading a published solution instead of writing a fresh one, the team scrubbed leftover answers from task environments, blocked network access, and ran an agent that searched for exploits until it found none. Attempts confirmed to use a leaked answer received zero reward, the same as a failed attempt, and they stayed below 2 percent of attempts throughout training.
- The company trained separate teacher models on tasks whose results are hard to check automatically, such as game development, scientific research, and embodied intelligence (controlling robots and other physical systems). Then it trained MiMo-V2.6-Pro to imitate them through on-policy distillation, in which the student learns to align its output with the teacher’s prediction of which token should come next.
Performance: MiMo-V2.6-Pro topped open weights models on two independent composite evaluations, making a large gain over its predecessor. It matched some proprietary models at a small fraction of their cost per task, but it trailed the leaders and generated output slowly.
- On Artificial Analysis’ Intelligence Index v4.3.2, a weighted average of 10 evaluations across math, science, coding, and reasoning, MiMo-V2.6-Pro set to reasoning (46, $0.13 and 19.5 minutes per task) led open weights models, ahead of GLM-5.3 set to max reasoning (45, $2.01 and 10.7 minutes per task) and Kimi K3 set to max reasoning (44 and $2.00 per task). Its predecessor, MiMo-V2.5-Pro, scored 26. Among proprietary models, it tied Grok 4.7 set to high reasoning (46, $2.73 and 14.2 minutes per task) and trailed several models from Anthropic, Meta, and OpenAI.
- On the Vals Index, which weights seven agentic benchmarks of finance, coding, and legal work by each sector’s share of U.S. GDP, the smaller MiMo-V2.6-Flash set to reasoning (59.58 percent accuracy, $0.20 per test, 45.03 minutes) and MiMo-V2.6-Pro set to reasoning (59.47 percent accuracy, $0.39 per test, 52.33 minutes) outperformed all other open weights models tested. Both trailed 15 proprietary models, led by Claude Opus 5.5 set to max reasoning (69.69 percent accuracy, $22.30 per test, 72 minutes).
- On Vals’ CyberBench v1.1, which tests whether agents can reproduce and patch vulnerabilities in open-source software, MiMo-V2.6-Flash set to reasoning (75.36 percent accuracy, $0.05 per test, 32.48 minutes) ranked first of nine models and MiMo-V2.6-Pro set to reasoning (72.86 percent accuracy, $0.09 per test, 31.73 minutes) ranked third, ahead of Muse Spark 1.3 Max (72.74 percent accuracy, $3.34 per test, 19.75 minutes) and Claude Fable 5.1 (70.42 percent accuracy, $4.03 per test, 22.12 minutes).
Behind the news: Recently, Anthropic alleged that, over 20 days in March and April 2026, Xiaomi routed MiMo users’ conversations and coding sessions to Claude through the coding tools OpenClaw and OpenCode, producing more than 400,000 exchanges to use as training data. To date, Xiaomi has not publicly responded to Anthropic’s accusation.
Why it matters: Code that passes tests isn’t necessarily good code. Xiaomi found that a version of its model that was trained without the code quality grader picked up habits that make software harder to maintain: The model added code the task didn’t call for, let errors pass silently, and were lax in checking on incoming data until tests passed. Trained with the code quality grader, the model made smaller, more focused changes. Rewarding code that a human reviewer would accept is as important as rewarding code that runs.
We’re thinking: Open weights are great, but open recipes are even better. By releasing its reinforcement learning tasks, the software that runs them, and its training code, Xiaomi lets others repeat the steps that produced much of its models’ improvements. It also disclosed what those steps cost: $2.6 million for MiMo-V2.6-Pro and $0.9 million for MiMo-V2.6-Flash. Teams planning to train agents get a set of tasks they don’t have to build and a rough price for reinforcement learning at this scale. We hope other labs can reproduce and extend on what Xiaomi learned training these models.

Voice, Video, and Reasoning in One
Google released two new speech-to-speech voice models, joining a suddenly crowded field of models built to power voice agents.
What’s new: Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two speech-to-speech models that can serve as real-time voice agents. The Gemini 3.8 Live model is built for scale and cost-efficiency while 3.8 Live Extended Thinking is intended for high-complexity tasks. Both models are designed to reduce response latency during live dialogue.
How it works: Google released few details about the model’s architecture and training data, but did say that the models are based on Gemini 3 Pro. Gemini 3.8 Live can take in image and video input, supports 97 languages, and can execute tools and API calls in the background while continuing a conversation. Live Extended Thinking can also reason and think simultaneously. The models’ knowledge cutoff is January 2025.
- In speech-to-speech benchmarks from Artificial Analysis, a company that evaluates AI models, Gemini 3.8 Live Extended Thinking generally ranked above its less powerful sibling, 3.8 Live, in various tests, but users preferred 3.8 Live for some tasks. The Live Extended Thinking model is ranked first in the Speech to Speech Index benchmark, with a score of 82.6%. It was fourth in AA’s Speech Reasoning benchmark – which evaluates an audio model’s ability to answer reasoning-based questions – behind StepAudio 3 Realtime and two Qwen models. It ranked in the top spot in the Tau Voice benchmark, with a score of 68.6 percent. In comparison, Gemini 3.8 Live is somewhat lower in the Speech to Speech Index, at 76 percent – the fifth spot – and significantly lower in the Agentic Performance Benchmark, at 30.1 percent. However, 3.8 Live ranked second in the Speech Agent Arena Leaderboard, where people hold blind live voice conversations and pick the model they prefer for tasks like booking a dentist appointment. It also performed better than Live Extended Thinking in certain benchmarks that measured conversational dynamics and the percentage of correctly completed tasks.
- The models are available to developers through the Gemini API and app, Google Cloud Vertex / Vertex AI, and Google AI Studio. Gemini 3.8 Live is also used in Search Live, and Gemini 3.8 Live Extended Thinking powers voice experiences in the Gemini app and Google Workspace products including Docs and Gmail.
- Google says that all audio generated by its AI products is watermarked with SynthID, a technique for detecting AI-generated content.
- Both models are available for free, with some limitations, in certain products like the Gemini app, Google Search, and the Gemini API. Paid tier queries are not used to improve the model, while the free tier is. The paid tier is priced per million tokens; the input price is $0.75 for text, $3.00 ($0.005/min) for audio, and $1.00 (or $0.002/min) for images or video. For output, the models are priced $4.50 per million tokens of text and $12.00 or $0.018/min of audio. Artificial Analysis found that Gemini 3.8 Live costs $0.84 per hour of input audio — the cheapest model in the Index — and a significantly lower cost than Live Extended Thinking at $3.50 per hour.
Behind the news: Traditionally, voice agents have relied on a pipeline that converts speech to text, sends it to a reasoning model, and converts the response back to speech. This approach has historically been seen as more accurate and easier for developers to control because LLMs generally reason in text. However, each step can add latency, which detracts from the user experience. OpenAI recently released a speech-to-speech model called GPT-Live-1, a voice model developers can use to build voice-enabled apps. This model can listen and speak at the same time, allowing it to handle interruptions and respond more naturally, while a separate reasoning model runs in the background.
Why it matters: It’s correct to call Gemini 3.8 Live a voice model since its primary language input is voice rather than text, but its ability to understand and reason over image and video input makes it more versatile. Putting voice and video input together is particularly powerful for helping users use AI in real-time to deal with anything on screen: web interfaces, games, multi-application workflows, etc. This differentiates it from GPT-Live-1 and other models that work strictly with voice and audio.
We’re thinking: Voice-driven agents are particularly exciting for developers because they offer a new real-time interface for building applications. Speech is a natural input format that most people can use regardless of their experience using a computer. It’s particularly useful for mobile and automotive interfaces where text input is less convenient. It’s up to developers to create applications that democratize users’ access to AI.

DeepSeek’s Flash Leapfrogs Pro Again
When agents make a tool call, they send the same context back through the model every time. DeepSeek says storing and moving all that context has become a bigger obstacle to keeping serving costs down than the computation a model performs to generate output. Its new architecture sidesteps that bottleneck.
What’s new: DeepSeek released DeepSeek-V4.1-Flash, its first model with a redesigned architecture that uses far less memory per token of context. It also cut its API prices. This is the first model in DeepSeek’s V4 series to accept images as input, outside of an experimental preview. It’s also very fast, second only to Gemini 3.8 Flash in tokens/second.
- Input/output: Text and images in (up to 1 million tokens), text out (up to 384,000 tokens, 225.6 tokens per second)
- Architecture: Encoder-decoder mixture-of-experts transformer, transformer-based vision encoder, 552 billion backbone parameters plus 196 billion in a memory module, 8 billion active per token while reading input and 16 billion while generating output
- Features: Reasoning (none, low, high, or max via the API; any integer from 1 to 100 in the released weights), tool calls, context caching
- Performance: 39 points on Artificial Analysis’ Intelligence Index; top open weights model on Vals AI’s Vals Index at time of release (currently fourth, 51.32 percent)
- Availability/price: Via DeepSeek’s app and website; via DeepSeek’s API at $0.30/$0.006/$1.20 per million input/cached/output tokens during peak hours, half price at all other times
- Weights/license: Free for commercial and noncommercial use under the MIT license
- Undisclosed: Knowledge cutoff, training data
How it works: A paper details how DeepSeek redesigned its V4 architecture. The authors reduced the key-value cache (the keys and values a transformer stores for every token it has read so that it avoids recomputing them) and the computation spent reading input.
- Of the model’s 40 layers, half read and half write. Instead of each decoder layer computing its own keys and values over the full input, the decoder derives them from the encoder’s last layer, so most of a prompt runs through only the encoder. DeepSeek says this nearly halves prefill computation (reading prompts before generating replies).
- Most layers borrow another layer’s cache. In DeepSeek-V4.1-Flash, each layer uses one of three modes. Full layers compute keys and values and select the 512 most relevant entries to attend to; reindex layers borrow the previous full layer’s cache but choose for themselves which 512 entries to attend to; reuse layers borrow the cache and the selected 512 entries from an earlier layer. Only 4 of the 40 layers compute a full input cache, three in the encoder and one in the decoder. Every layer keeps its own uncompressed cache of the most recent 128 tokens, computed from its own view of the text, so it attends to recent tokens in full while the shared cache supplies the 512 most relevant entries from the entire input.
- DeepSeek-V4.1-Flash stores its full-input cache in 4-bit numbers rather than 8-bit, but every group of 16 numbers shares one 8-bit multiplier that restores the group’s scale and trains the model during fine-tuning to tolerate the rounding, which helps maintain precision. Sharing caches across layers plus 4-bit storage brings the entire input cache to 890 bytes per token, about four times smaller than DeepSeek-V4-Flash’s and 437 times smaller than DeepSeek-V1’s. With the combined changes, the computation needed to generate each output token rises by only 25 percent as input grows from 4,000 tokens to 1 million tokens.
- DeepSeek’s servers keep each prompt’s full-input cache for at least 72 hours, so repeated prefixes needn’t be recomputed. DeepSeek-V4.1-Flash only keeps them in server memory for minutes. If a request arrives after they’ve expired, the system re-reads the last 128 tokens of the prompt to rebuild an approximate version. An exact rebuild would require re-reading 128 tokens for each of 40 layers (5,120 tokens) because each layer’s short-range cache depends on its preceding layer. DeepSeek says the approximation barely affects output quality and reduces the cache to an eighth of DeepSeek-V4’s.
- DeepSeek pretrained the model on 45 trillion tokens of text and images. It was fine-tuned according to DeepSeek-V4’s recipe: supervised learning, reinforcement learning, then on-policy distillation, in which the model was fine-tuned to match what more than 40 specialist models would have written. The teachers were DeepSeek’s own checkpoints from different training stages and domains.
- DeepSeek says it did not alter its training algorithms; instead, the authors credit performance gains over DeepSeek-V4 to automated pipelines that generate training tasks and agent environments.
Performance: Independent evaluators found that DeepSeek-V4.1-Flash outperforms its predecessor and the larger DeepSeek-V4-Pro-0813. It costs less than half as much per task and generates output more than twice as fast. It briefly led open weights models on Vals.ai’s index (prior to the release of MiMo v2.6 Pro) and tops all models on one independent test of business-software agents.
- On Artificial Analysis’ Intelligence Index v4.3.2, a composite of 10 evaluations of agentic across mathematics, science, coding, and reasoning, DeepSeek-V4.1-Flash set to max reasoning (39 points, $0.27, and 4.9 minutes per task) outperformed DeepSeek-V4-Pro-0813 set to max reasoning (36 points, $0.67, and 8.7 minutes per task). Among open weights models, it trailed MiMo v2.6 Pro (46 points, $0.13 per task), GLM-5.3 (45, $2.01) and GLM-5.3-Flash (42, $0.25, and 11.1 minutes per task).
- On AutomationBench-AA, a test of how well agents complete business workflows across simulated app environments such as Gmail, Salesforce, and Jira, DeepSeek-V4.1-Flash set to max reasoning (68.9 percent) led all models tested, including GPT-6 Astra set to max reasoning (68.5 percent) and Grok 4.6 set to xhigh reasoning (67 percent).
Behind the news: Shrinking the key-value cache has been a DeepSeek theme since its second-generation model. Rivals have borrowed DeepSeek’s techniques and developed their own to likewise compete on cache size.
- In May 2024, DeepSeek-V2 introduced multi-head latent attention, which compressed keys and values into a small vector and reduced the key-value cache by 93.3 percent relative to DeepSeek 67B. Three months later, DeepSeek began storing caches on hard disks and charging a tenth as much for inputs the cache already held.
- DeepSeek-V4.1-Flash’s design builds on YOCO, a 2024 Microsoft and Tsinghua University design in which the upper half of a transformer’s layers reuse the cache produced by the lower half.
- The release arrived days before Anthropic accused DeepSeek and six other Chinese developers of distilling Claude models to improve their own models and to serve Claude’s output to their customers. DeepSeek’s paper says it distilled the model from more than 40 of its own checkpoints, taken from different stages of training, and doesn’t mention outside models.
Why it matters: Agents usually read more than they write. Because each tool call sends a growing transcript back through the model, storing and re-reading context can cost more than generating output. DeepSeek’s price cut is steepest where long-running agents use the most tokens: The cost for previously cached input fell 57 percent, while the cost of output fell 9 percent.
We’re thinking: A year ago, model makers strove to lengthen context windows. DeepSeek-V4.1-Flash keeps its predecessor’s 1 million-token context and instead makes it cheaper to keep: 890 bytes of cache per token, 437 times smaller than DeepSeek-V1’s and 4 times smaller than DeepSeek-V4-Flash’s. If other model makers follow DeepSeek’s lead, launches may soon tout cache size per token alongside context windows and cost per token.

Agents That Can Check Their Own Work Piece by Piece
When asked questions requiring extensive reasoning and tool use to answer, agents may search without making progress, repeat unsuccessful strategies, or settle for partially verified answers. Building answers iteratively can help to address these shortcomings.
What’s new: Researchers at the Beijing Academy of Artificial Intelligence developed AREX, a research agent that hones its work through repeated cycles of gathering evidence, reflecting on provisional answers, and launching targeted follow-up searches. The combination of an iterative harness and a fine-tuned model outperformed competing approaches.
Key insight: Some earlier agents check a candidate answer against a question’s requirements and accept or reject it. However, the check can do more work. When a tentative answer fails to satisfy every requirement, checking it requirement by requirement reveals which ones remain unsupported and where the available evidence conflicts. Instead of discarding the answer, the agent can keep the parts it has confirmed and use unmet requirements as the basis for a further question to research next. This turns a pass/fail judgment into a loop that drives improvement.
How it works: The authors built an agentic harness to answer difficult research questions. They used it to build a dataset and fine-tuned two models to work with it more effectively: Qwen3.5-4B, which is fairly small, and Qwen3.5-122B-A10B, a much larger mixture of experts.
- They designed the harness so the agent first took actions that are typical of research agents: It searched the web, read web pages, and built a provisional answer. At any point, it could call a tool to compress its traces into a brief research state that recorded verified facts, hypotheses it had considered and discarded, unresolved points, and its next plan.
- When the agent had decided it was unlikely to improve its answer through further research, it produced an answer with a confidence score from 0 to 100. If the score exceeded a threshold set by the authors, the agent accepted the answer. If not, the agent either refined its work (keeping useful evidence while focusing on remaining gaps) or restarted from scratch.
- To build the dataset, the authors created template prompts for research questions that require synthesizing multiple sources, planning or deducing over multiple steps, or integrating academic papers. They used the templates and an unspecified model to produce questions. Then they paired the harness with unspecified models, which answered the questions and recorded their traces.
- They used supervised fine-tuning to train the two Qwen3.5 models to reproduce the traces. Then they applied reinforcement learning to train the models to produce the answers. Because long traces contain many routine steps and only a few decisive ones, they used hand-written rules to flag the moments that mattered most, such as when the model found key evidence, abandoned a wrong hypothesis, or condensed its research state. They concentrated the training signal on those flagged decision points during both supervised fine-tuning and reinforcement learning.
Results: Across six agentic benchmarks, the AREX systems based on the fine-tuned Qwen3.5 models outperformed systems that used the same harness with similarly sized models as well as the Qwen3.5 models without fine-tuning.
- On BrowseComp, which tests multi-step web searches, the AREX system based on the fine-tuned Qwen3.5-122B-A10B achieved 82.5 percent accuracy. This result fell somewhat below the next-best competitor, the authors’ harness paired with Gemini Pro 3.1 — presumably a much larger model — which achieved 85.9 percent.
- On WideSearch-en, which evaluates synthesis of output from multiple retrieved sources, the AREX system based on the fine-tuned Qwen3.5-122B-A10B achieved 82.0 percent F1. This result outdistanced the next-best system, the authors’ harness paired with Kimi-K2.6, which achieved 80.8 percent F1.
- The authors’ harness with the fine-tuned Qwen3.5-122B-A10B boosted performance by as much as 10 percentage points (for BrowseComp) over harnesses paired with the same model that performed only the steps typical of research agents.
- Likewise, fine-tuning improved markedly the systems performance. On BrowseComp and WideSearch-en, performance of similar systems without fine-tuning fell by as much as 20 percentage points. On five out of six benchmarks, the AREX system based on the relatively tiny fine-tuned Qwen3.5-4B outperformed a system based on the much larger Qwen3.5‑35B without fine-tuning.
Why it matters: Research agents can perform better with both fine-tuning and repeated efforts to improve their output. The training strategy of identifying and amplifying the most informative decision points offers a practical lesson for improving agent behavior: Not all steps in a long trajectory are equally important. Focusing training on critical decision points yielded better results than treating every action the same way.
We're thinking: The authors used hand-written rules, not a learned model, to detect and target training examples. These let the model learn how to compress long traces into more valuable decision points. The combination shows that cheap but clear heuristics can outperform elaborate ones when the thing being detected is well-defined.