Abstract
As the production of software passes into the hands of software itself, coding agents are increasingly commissioned to build agents fitted to their users' own workflows. This delegation carries a quiet assumption, that the delivered agent will faithfully serve the requirement it was built for, rather than merely appear to. No existing benchmark tests that assumption. We introduce AgentSWE, a benchmark for agent software engineering spanning three stages of the agent development lifecycle: Creation of a complete agent from a natural-language requirement, Editing of production agent codebases, and Optimization toward a measurable target. Each of its 25 tasks is framed as a commission rather than an exam: the builder iterates against a handful of illustrative cases, while acceptance exercises the frozen submission on held-out cases designed to span the full requirement, grading behavior against requirement-derived rubrics under hard pass/fail gates or, for Optimization, against the source benchmark's own objective. Evaluating a range of coding agents across models and harnesses, we find that none comes close to delivering what a requirement asks for: the gap between what a builder demonstrates and what it delivers is the norm, not the corner case. An accompanying empirical analysis of the failure trajectories identifies recurring failure modes, among them overfitting to the visible cases, hollow runtimes behind code that looks complete, and unsupported claims of completion, each a signal by which coding agents can learn to build the agents they are asked for.
The Benchmark
Build an agent from scratch
The builder starts from an empty workspace and delivers a complete agent for a domain requirement.
- Repository bug repair, database analytics, formal theorem proving
- Desktop GUI automation, web research report, schema-guided web extraction
- Scientific PDF translation, document-to-editable PPTX, document QA, vulnerability validation
Extend a production agent
The builder adds new capabilities to a pinned open-source agent without breaking what already works.
- AI-Scientist, Aider, Claude Code, Codex, DeepCode
- DeepTutor, Dyad, OpenClaw, OpenHands, OpenWiki
Improve a working agent
The builder improves a runnable starter agent toward the source benchmark's own objective.
- BrowseComp, τ³ (retail), Terminal-Bench
- PinchBench, OSWorld
Evaluation Protocol
The builder works in one persistent session against the public development cases and gets feedback on each submission. When development ends, the delivery is frozen and run on held-out cases it has never seen, in an isolated, pinned runtime where every model call goes through the benchmark's gateway. Creation and Editing deliveries must pass hard validity gates before a requirement-derived rubric scores them; Optimization deliveries are scored by gap closure under the source benchmark's evaluator.
Results
AgentSWE-Lite
Start here. AgentSWE-Lite is 8 of the 25 tasks: three Creation, three Editing and two Optimization tasks. DeepSeek-V4.1-Flash is the runtime model and the judge in all three stages, so a single DeepSeek API key is all you need: no search key and no release assets. It is the lower-cost entry point for evaluating your own coding agent.
Mean over three builds per pairing in Codex (profile lite-v1.1). Creation and Editing:
held-out Result (0–100); Optimization: S. DeepSeek-V4.1-Flash is the runtime model and the judge, so the Creation
and Optimization levels are not comparable with the main results above.
# in a clone of the repository, with AGENTSWE_DEFAULT_API_KEY set in .env
echo AGENTSWE_PROFILE=lite-v1.1 >> .env
python3 -m agentswe doctor repo db pdf deeptutor aider openwiki tau3 pinch
python3 -m agentswe setup <task>
python3 -m agentswe run <task> --builder codex
What Goes Wrong
Overfitting to the visible cases
Eleven Creation runs finish development at 90 or above, yet their held-out Result averages 9.9. Builders fit their agents to the cases they can see.
Hollow runtimes
Code that reads as complete often does not run. A static code review cannot tell zero-score deliveries from the rest; in Editing, 28 of 33 deliveries leave some added capability absent, never invoked or only partly executed.
Unsupported claims of completion
6 of 42 final Creation reports and 3 of 19 Optimization reports overclaim. Delivered Editing agents do it too: in 13 of 33 runs they report outcomes their own logs contradict.
BibTeX
@misc{wang2026agentswe,
title = {AgentSWE: Can Coding Agents Build the Agent You Actually Want?},
author = {Jiahao Wang and Hongjin Qian and Yuyang Hu and Jiajun Zhang and Zheng Liu},
year = {2026},
note = {arXiv preprint coming soon}
}