Where can coding agents improve accepted delivery rather than simply produce more code?
AI coding agents can investigate a repository, modify files and run development tools. Their most useful role is to help a team deliver a verified change, such as a tested bug fix or a documented migration. Generating more code is an intermediate activity; the business outcome is software that works, can be maintained and does not create an expensive repair burden.
Start with bounded changes in a repository that already has a runnable baseline. Protect production access, agree on acceptance criteria and measure review effort as carefully as implementation speed.
Recent developments and the evidence gap
On 9 September 2026, Mistral described a project migrating 40,000 lines of a European energy operator's Fortran code to C++. The report says this was a first sprint covering part of a 300,000-line codebase. It describes difficulties with fully autonomous approaches and a move toward structured workflows with human oversight. This is a vendor-reported case study, not an independent estimate of average productivity. Mistral case study.
Anthropic introduced Claude Fable 5.1 and Mythos 5.1 in September 2026, with its newsroom dating the announcement to 1 September. It positions the models for coding and knowledge work; Fable is generally available while Mythos has restricted access. Reported capability gains should be tested on a team's own tasks rather than treated as a delivery guarantee. Anthropic release.
A useful counterpoint is METR's July 2025 randomized study of experienced open-source developers using early-2025 tools on familiar repositories. In that setting, AI access increased completion time by 19%. The study involved 16 developers and 246 tasks. Its age and population mean it cannot settle the productivity of September 2026 agents, but it demonstrates why perceived speed and measured results can diverge. METR study.
Our interpretation is that capability progress and integration quality must be evaluated separately. A model can improve substantially while a poorly designed development process still wastes time.
How a coding agent differs from autocomplete
Autocomplete suggests code at the point of typing. A coding agent can pursue a larger instruction: inspect an error, locate related functions, propose changes and execute checks. This broader reach creates both value and risk. A mistaken assumption can now affect several files rather than one line.
Think of the agent as a contributor working inside a controlled development environment. It needs a task brief, repository context, tools and feedback. It should not have permission to merge its own work into production merely because its tests pass.
Tests are necessary but incomplete. They can miss behavior nobody anticipated, and an agent may accidentally weaken a test to accommodate its implementation. Review the test changes as well as the application changes.
Practical example: repair a campaign reporting integration
An illustrative agency has a reporting tool that imports advertising-platform data. A supplier changes one response field, and the dashboard stops displaying conversions. The agent's task is to repair the adapter while preserving the existing definition of conversion rate.
The developer provides a redacted failing response, a previous valid response and the expected report output. The agent reproduces the failure in a local branch, updates the mapping and adds a regression test. It must also test missing fields, zero denominators and pagination.
An account analyst reviews the business meaning. A developer reviews the code and dependency changes. The corrected report is compared with a known reference before release. The agent has no advertising-account write permission and cannot change campaign budgets.
This example is a proposed workflow, not a claim about a real client or a promised reduction in delivery time.
Implementation sequence
- Choose the task class. Begin with bug fixes, small adapters, test additions or documentation grounded in code. Avoid an initial project that redesigns authentication or replaces an entire business system.
- Create a reproducible environment. Record runtime versions and package dependencies. Provide synthetic or approved test data. Confirm that the baseline builds before asking the agent to change it.
- Write observable acceptance criteria. Specify inputs, expected outputs, failure handling and behaviors that must remain unchanged. Include the business calculation, not just the appearance of the screen.
- Use an isolated branch. Keep changes reviewable and reversible. Supply read-only access to relevant documentation and scoped access to development tools.
- Require evidence with the change. Ask for the original failure, the fix, the checks executed, their actual results and remaining uncertainty. “Should work” is not a test result.
- Review independently. A qualified person should assess security-sensitive logic, new packages, data handling and altered tests. A second model can assist, but does not replace ownership.
- Release gradually. Use a staging environment and a defined rollback. Monitor the behavior that motivated the task, not just whether deployment completed.
For legacy modernization, add behavior comparison against the old system. Numerical output, rounding, timezone handling and obscure edge cases may be business requirements even when the original code looks untidy.
A practical task brief
Fix the conversion import failure using the supplied fixtures. Preserve the documented metric definitions. Reproduce the failing case first. Keep the change limited to the adapter and relevant tests. Do not modify credentials, production data or deployment settings. Report the commands actually run, results, changed assumptions and anything that still needs human review.
The brief reduces ambiguity. Enforce sensitive restrictions through repository and environment permissions, rather than relying only on the wording.
Measuring whether it helps
Compare similar tasks over a sufficient period; do not compare a simple AI-assisted fix with a difficult manual redesign. Capture planning, prompting, waiting, reviewing, correcting and debugging after release. Count abandoned attempts as cost.
| Measure | Why it matters |
|---|---|
| Time to accepted change | Includes the complete path to review approval |
| Reviewer minutes | Reveals whether effort moved from writing to inspection |
| Rework within an agreed window | Detects superficially successful changes |
| Escaped defects by severity | Prevents speed from hiding quality loss |
| Cost per accepted task | Includes model usage and staff effort |
An illustrative calculation: a task takes 120 minutes manually. The agent spends 15 minutes executing, but the developer spends 20 minutes specifying, 50 reviewing and 45 correcting. Human effort is 115 minutes, so the human-time saving is only five minutes. The agent's fast first draft alone would have given a misleading impression.
For a pilot, define a small set of repeated task types and compare outcomes against a manual baseline. Do not set a universal productivity target. The right threshold depends on defect consequences, task complexity and the cost of supervision.
Common failure modes
The agent may implement the wrong interpretation correctly. Prevent this by making acceptance criteria concrete and involving a business owner where calculations or workflows matter. It may also invent a package or use an outdated interface. Verify new dependencies and consult the actual version's documentation.
Large changes can become impossible to review. Split work by independently testable behavior. If a change requires hundreds of unrelated edits, first ask whether the scope is justified.
Do not give an agent production secrets to make debugging convenient. Redact records, reproduce problems locally and provide narrow diagnostic interfaces. Restrict command execution where possible and review any changes to build pipelines or permission configuration.
Finally, preserve the ability to work without the agent. Repository instructions, tests and human understanding should improve during adoption, not disappear behind an opaque conversation history.
First-month recommendation
Spend week one creating task fixtures and recording the baseline. Use week two for supervised low-risk changes. In week three, review accepted and rejected attempts together to identify recurring errors. At the end of the month, decide which task classes merit expansion and which should remain manual.
Adopt coding agents where they improve accepted outcomes at an acceptable total cost. Keep a narrow scope until the team can reliably explain what changed, why it is correct and how to undo it.
Sources and scope
Evidence cutoff: 10 September 2026. Recent vendor reports establish announcements and reported examples; the older METR study is explicitly historical and should not be generalized to all current tools. Implementation plans and numerical examples are original recommendations and illustrations.
- Mistral. “Modernizing complex legacy code with AI agents.” 9 September 2026. Source.
- Anthropic. “Claude Fable 5.1 and Mythos 5.1.” September 2026; newsroom announcement 1 September. Source.
- Joel Becker, Nate Rush, Beth Barnes and David Rein, METR. “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.” 10 July 2025. Source.
Turn the research into operating value.
Connect the use case, architecture, evidence, controls and operating model around a decision that matters.
