Learning to improve, round after round. Long-horizon reflection learned in coding tasks carries over to deep research.
Scroll horizontally to explore the figure
Self-improvement at test time
Agents that improve
within the task.
Given more rounds on the same problem, an agent should turn them into a better solution by its own judgment of what to change.
Reflection guides the next revision.
Long-horizon execution keeps the loop productive.
Together, they make progress accumulate: useful findings carry forward, and the agent recovers from setbacks to keep improving.
Teaching long-horizon reflection
Train on the path
to a better solution.
Machine learning and algorithmic programming offer verifiable feedback and room to improve. We select whole trajectories by their final outcome and process quality, keeping failed attempts and regressions in context. Supervision focuses on decisions that diagnose, repair, and improve.
Scroll horizontally to explore the figure
These trajectories are mixed with the unchanged AREX deep-research data to fine-tune Qwen3.8-27B. The resulting AREX-2 improves over both previous AREX models on all four research benchmarks, supporting transfer beyond the domains where it learned to improve.
Explore AREX-2
Interactive demo
Follow a solution
as it improves.
Unloading Boxes · Frontier-CS 164. Replay five hours of iteration. Explore the task, then inspect the code, edits, and feedback behind each score.