AREX-2

Learning to improve, round after round. Long-horizon reflection learned in coding tasks carries over to deep research.

AREX-2 (27B) benchmark results: BrowseComp 84.0, text-only HLE 52.6, Frontier-CS 70.7, DeepSearchQA 93.8, GAIA 92.2, and MLE-bench Lite 81.8, compared with selected open- and closed-weight models.
Six evaluations of coding, deep research, and agentic reasoning. AREX-2 · 27B MLE: Any Medal, averaged over three seeds, with skills. DeepSearchQA: F1. HLE: text-only, except GPT-5.5 (full set).

Scroll horizontally to explore the figure

What self-improvement means

Self-improvement at test time

Agents that improve
within the task.

Given more rounds on the same problem, an agent should turn them into a better solution by its own judgment of what to change.

Reflection guides the next revision.

Long-horizon execution keeps the loop productive.

Together, they make progress accumulate: useful findings carry forward, and the agent recovers from setbacks to keep improving.

Try, measure, reflect, revise One round of self-improvement: build a solution, read feedback, find what to change, and update the solution. Repeat the loop on the same task. 01TryBuild a solution 02MeasureRead the feedback 03ReflectFind what to change 04ReviseUpdate the solution LEARN FROMEACH ROUND Sustained iteration turns more rounds into a better solution An animated score curve rising over repeated rounds, with plateaus between improvements. Score Rounds on the same task

Teaching long-horizon reflection

Train on the path
to a better solution.

Machine learning and algorithmic programming offer verifiable feedback and room to improve. We select whole trajectories by their final outcome and process quality, keeping failed attempts and regressions in context. Supervision focuses on decisions that diagnose, repair, and improve.

AREX-2 framework. Left: construct and validate environments from GitHub repositories and online judges. Middle: an agent submits solutions, receives feedback, reflects, and improves over many rounds, acquiring operational knowledge on demand. Right: train on these trajectories so the model improves its solutions over multiple rounds at test time.

Scroll horizontally to explore the figure

These trajectories are mixed with the unchanged AREX deep-research data to fine-tune Qwen3.8-27B. The resulting AREX-2 improves over both previous AREX models on all four research benchmarks, supporting transfer beyond the domains where it learned to improve.

From the paper to the process Scroll down to explore the demo
Back to research AREX-2 Interactive demo

Interactive demo

Follow a solution
as it improves.

Unloading Boxes · Frontier-CS 164. Replay five hours of iteration. Explore the task, then inspect the code, edits, and feedback behind each score.