Agents That Can Check Their Own Work Piece by Piece: AREX condenses and reviews its own research

When asked questions requiring extensive reasoning and tool use to answer, agents may search without making progress, repeat unsuccessful strategies, or settle for partially verified answers.

Share
AREX diagram shows inner and outer loops for structured query analysis, involving observation and feedback.

When asked questions requiring extensive reasoning and tool use to answer, agents may search without making progress, repeat unsuccessful strategies, or settle for partially verified answers. Building answers iteratively can help to address these shortcomings.

What’s new: Researchers at the Beijing Academy of Artificial Intelligence developed AREX, a research agent that hones its work through repeated cycles of gathering evidence, reflecting on provisional answers, and launching targeted follow-up searches. The combination of an iterative harness and a fine-tuned model outperformed competing approaches.

Key insight: Some earlier agents check a candidate answer against a question’s requirements and accept or reject it. However, the check can do more work. When a tentative answer fails to satisfy every requirement, checking it requirement by requirement reveals which ones remain unsupported and where the available evidence conflicts. Instead of discarding the answer, the agent can keep the parts it has confirmed and use unmet requirements as the basis for a further question to research next. This turns a pass/fail judgment into a loop that drives improvement.

How it works: The authors built an agentic harness to answer difficult research questions. They used it to build a dataset and fine-tuned two models to work with it more effectively: Qwen3.5-4B, which is fairly small, and Qwen3.5-122B-A10B, a much larger mixture of experts.

  • They designed the harness so the agent first took actions that are typical of research agents: It searched the web, read web pages, and built a provisional answer. At any point, it could call a tool to compress its traces into a brief research state that recorded verified facts, hypotheses it had considered and discarded, unresolved points, and its next plan.
  • When the agent had decided it was unlikely to improve its answer through further research, it produced an answer with a confidence score from 0 to 100. If the score exceeded a threshold set by the authors, the agent accepted the answer. If not, the agent either refined its work (keeping useful evidence while focusing on remaining gaps) or restarted from scratch.
  • To build the dataset, the authors created template prompts for research questions that require synthesizing multiple sources, planning or deducing over multiple steps, or integrating academic papers. They used the templates and an unspecified model to produce questions. Then they paired the harness with unspecified models, which answered the questions and recorded their traces.
  • They used supervised fine-tuning to train the two Qwen3.5 models to reproduce the traces. Then they applied reinforcement learning to train the models to produce the answers. Because long traces contain many routine steps and only a few decisive ones, they used hand-written rules to flag the moments that mattered most, such as when the model found key evidence, abandoned a wrong hypothesis, or condensed its research state. They concentrated the training signal on those flagged decision points during both supervised fine-tuning and reinforcement learning.

Results: Across six agentic benchmarks, the AREX systems based on the fine-tuned Qwen3.5 models outperformed systems that used the same harness with similarly sized models as well as the Qwen3.5 models without fine-tuning.

  • On BrowseComp, which tests multi-step web searches, the AREX system based on the fine-tuned Qwen3.5-122B-A10B achieved 82.5 percent accuracy. This result fell somewhat below the next-best competitor, the authors’ harness paired with Gemini Pro 3.1 — presumably a much larger model — which achieved 85.9 percent.
  • On WideSearch-en, which evaluates synthesis of output from multiple retrieved sources, the AREX system based on the fine-tuned Qwen3.5-122B-A10B achieved 82.0 percent F1. This result outdistanced the next-best system, the authors’ harness paired with Kimi-K2.6, which achieved 80.8 percent F1.
  • The authors’ harness with the fine-tuned Qwen3.5-122B-A10B boosted performance by as much as 10 percentage points (for BrowseComp) over harnesses paired with the same model that performed only the steps typical of research agents.
  • Likewise, fine-tuning improved markedly the systems performance. On BrowseComp and WideSearch-en, performance of similar systems without fine-tuning fell by as much as 20 percentage points. On five out of six benchmarks, the AREX system based on the relatively tiny fine-tuned Qwen3.5-4B outperformed a system based on the much larger Qwen3.5‑35B without fine-tuning.

Why it matters: Research agents can perform better with both fine-tuning and repeated efforts to improve their output. The training strategy of identifying and amplifying the most informative decision points offers a practical lesson for improving agent behavior: Not all steps in a long trajectory are equally important. Focusing training on critical decision points yielded better results than treating every action the same way.

We're thinking: The authors used hand-written rules, not a learned model, to detect and target training examples. These let the model learn how to compress long traces into more valuable decision points. The combination shows that cheap but clear heuristics can outperform elaborate ones when the thing being detected is well-defined.