AgentSWE

Can Coding Agents Build the Agent You Actually Want?

Jiahao Wang1,2,* Hongjin Qian1,* Yuyang Hu1,3 Jiajun Zhang2 Zheng Liu1,†

1Beijing Academy of Artificial Intelligence

2Institute of Automation, Chinese Academy of Sciences   3Renmin University of China

*Equal contribution   †Corresponding author

Paper (arXiv, coming soon) Code Lite (8 tasks)
AgentSWE at a glance
(a) Existing benchmarks grade the patch a coding agent writes, or the performance of an agent someone else built. (b) AgentSWE commissions an agent from a natural-language requirement and grades what the delivered agent does on held-out cases. (c) Mean held-out score of five builders in the Codex harness.

Abstract

As the production of software passes into the hands of software itself, coding agents are increasingly commissioned to build agents fitted to their users' own workflows. This delegation carries a quiet assumption, that the delivered agent will faithfully serve the requirement it was built for, rather than merely appear to. No existing benchmark tests that assumption. We introduce AgentSWE, a benchmark for agent software engineering spanning three stages of the agent development lifecycle: Creation of a complete agent from a natural-language requirement, Editing of production agent codebases, and Optimization toward a measurable target. Each of its 25 tasks is framed as a commission rather than an exam: the builder iterates against a handful of illustrative cases, while acceptance exercises the frozen submission on held-out cases designed to span the full requirement, grading behavior against requirement-derived rubrics under hard pass/fail gates or, for Optimization, against the source benchmark's own objective. Evaluating a range of coding agents across models and harnesses, we find that none comes close to delivering what a requirement asks for: the gap between what a builder demonstrates and what it delivers is the norm, not the corner case. An accompanying empirical analysis of the failure trajectories identifies recurring failure modes, among them overfitting to the visible cases, hollow runtimes behind code that looks complete, and unsupported claims of completion, each a signal by which coding agents can learn to build the agents they are asked for.

The Benchmark

Creation · 10 tasks

Build an agent from scratch

The builder starts from an empty workspace and delivers a complete agent for a domain requirement.

  • Repository bug repair, database analytics, formal theorem proving
  • Desktop GUI automation, web research report, schema-guided web extraction
  • Scientific PDF translation, document-to-editable PPTX, document QA, vulnerability validation
Editing · 10 tasks

Extend a production agent

The builder adds new capabilities to a pinned open-source agent without breaking what already works.

  • AI-Scientist, Aider, Claude Code, Codex, DeepCode
  • DeepTutor, Dyad, OpenClaw, OpenHands, OpenWiki
Optimization · 5 tasks

Improve a working agent

The builder improves a runnable starter agent toward the source benchmark's own objective.

  • BrowseComp, τ³ (retail), Terminal-Bench
  • PinchBench, OSWorld
How AgentSWE tasks are constructed
How the tasks are built, and what each task gives the coding agent and keeps for evaluation.

Evaluation Protocol

The builder works in one persistent session against the public development cases and gets feedback on each submission. When development ends, the delivery is frozen and run on held-out cases it has never seen, in an isolated, pinned runtime where every model call goes through the benchmark's gateway. Creation and Editing deliveries must pass hard validity gates before a requirement-derived rubric scores them; Optimization deliveries are scored by gap closure under the source benchmark's evaluator.

Develop and freeze, then held-out evaluation

Results

AgentSWE-Lite

Start here. AgentSWE-Lite is 8 of the 25 tasks: three Creation, three Editing and two Optimization tasks. DeepSeek-V4.1-Flash is the runtime model and the judge in all three stages, so a single DeepSeek API key is all you need: no search key and no release assets. It is the lower-cost entry point for evaluating your own coding agent.

Mean over three builds per pairing in Codex (profile lite-v1.1). Creation and Editing: held-out Result (0–100); Optimization: S. DeepSeek-V4.1-Flash is the runtime model and the judge, so the Creation and Optimization levels are not comparable with the main results above.

# in a clone of the repository, with AGENTSWE_DEFAULT_API_KEY set in .env
echo AGENTSWE_PROFILE=lite-v1.1 >> .env
python3 -m agentswe doctor repo db pdf deeptutor aider openwiki tau3 pinch
python3 -m agentswe setup <task>
python3 -m agentswe run <task> --builder codex

What Goes Wrong

Overfitting to the visible cases

+13.1 → +1.8
Optimization gain, development vs. held-out

Eleven Creation runs finish development at 90 or above, yet their held-out Result averages 9.9. Builders fit their agents to the cases they can see.

Hollow runtimes

58.1%
of traced non-GUI held-out Creation executions fail a validity gate

Code that reads as complete often does not run. A static code review cannot tell zero-score deliveries from the rest; in Editing, 28 of 33 deliveries leave some added capability absent, never invoked or only partly executed.

Unsupported claims of completion

9 reports
claim completions their own records contradict

6 of 42 final Creation reports and 3 of 19 Optimization reports overclaim. Delivered Editing agents do it too: in 13 of 33 runs they report outcomes their own logs contradict.

Optimization gain on development and held-out cases
Optimization: gain over the starter on development (open) and held-out (filled) cases, averaged over five builders per benchmark.
Outcomes of held-out Creation executions
Outcomes of held-out Creation executions by builder. Only graded executions reach the Result rubric.

BibTeX

@misc{wang2026agentswe,
  title  = {AgentSWE: Can Coding Agents Build the Agent You Actually Want?},
  author = {Jiahao Wang and Hongjin Qian and Yuyang Hu and Jiajun Zhang and Zheng Liu},
  year   = {2026},
  note   = {arXiv preprint coming soon}
}