Skip to main content

folkfox

Skip to main content
Skip to content
AI ENGINEERING / EVIDENCE REVIEW

Hermes Agent refactored a million-line codebase. The audit matters more

Nous Research's 19-hour refactor is an extraordinary demonstration of parallel software work. It is also a better story when the savings boast is set aside and the engineering record is read closely.

Quick answerHermes Agent shows that agentic coding tools can coordinate a vast maintenance campaign. The defensible result is a 34.4% reduction in non-test Python source, while orchestration, telemetry, regression tests and independent review remain essential.
SECTION 01

Agentic coding tools need an audit trail#

At 10.11pm GMT on 15 September, Nous Research announced that 1,393 subagents had spent nineteen hours shrinking the Hermes Agent codebase by 34.4%. Its social post added a more combustible line: the work had saved nearly $2 million in engineering hours. Nous Research on X, 15 September 2026 The first claim is backed by a repository record. The second is a promotional estimate without a published calculation that readers can reproduce.

The merged pull request gives the firmer account. Non-test Python source fell from 1,063,826 to 698,363 lines. Files containing at least 5,000 lines fell from 37 to 6. Functions of at least 300 lines fell from 192 to 2. Across 2,655 changed files, the run produced 4,271 commits and the project reported 111,352 tool calls. Hermes Agent pull request #102117

Those numbers are remarkable without fictional precision. Removing code is not the same as completing a known backlog, and a line deleted is not a human hour recovered. A proper economic claim would need a counterfactual: which engineers, working with which tools, under which quality threshold, would have done the same work and how long would it have taken? The public record does not supply that experiment.

The cleaner headline is therefore smaller and stronger. Agentic coding tools coordinated a behaviour-preserving refactor across a million-line Python codebase in one long working day. The project then exposed enough telemetry and independent review to show where the machinery bent. That is useful evidence for any leader considering an agent governance model or comparing the cost claims around a coding agent. The trail leads to the review record, not the promotional number.

The repository record supports substantial structural reduction. Percentages for files and functions are calculated from the before-and-after counts in the pull request.Non-test Python source: 34.4Files at least 5,000 lines: 83.8Functions at least 300 lines: 99Longest function: 92.3100755025034.4Non-test Pythonsource83.8Files at least5,000 lines99Functions at least300 lines92.3Longest function
The repository record supports substantial structural reduction. Percentages for files and functions are calculated from the before-and-after counts in the pull request.
ItemValue
Non-test Python source34.4
Files at least 5,000 lines83.8
Functions at least 300 lines99
Longest function92.3
The repository record supports substantial structural reduction. Percentages for files and functions are calculated from the before-and-after counts in the pull request.
SECTION 02

This was not ordinary agentic refactoring#

What is agentic coding in this case? It is not an autocomplete system proposing a neat extraction while one developer watches. A parent process decomposed a repository-wide objective, dispatched bounded tasks, inspected results, repaired its own orchestration layer and continued. The striking unit of work was not the generated patch. It was the system that kept producing, checking and integrating patches.

That distinction matters because the academic record on agentic coding tools is still dominated by narrower tasks. An empirical study of agentic refactoring examined 15,451 refactoring instances across 12,256 pull requests and 14,988 commits. The authors found that agents often pursued maintainability and readability rather than only fixing defects. Agentic Refactoring, 2025 A second empirical study found that annotation-related changes dominated the refactorings it observed. How do Agents Refactor, 2026

Hermes sits at a different scale. Its project record describes repository topology, dependency boundaries, function length, code motion and test behaviour as one connected programme. That makes it closer to software evolution than a single-issue benchmark. SWE-EVO was created precisely because long-horizon evolution tasks touch many files and tests, exposing limitations that shorter issue-resolution benchmarks can hide.

The older benchmark tradition remains essential context. SWE-bench made real repository issues a measurable target. SWE-agent showed that the interface between an agent and its computer can materially affect navigation, editing and testing. OpenHands extended that idea into a platform with sandboxed execution, evaluation and multi-agent coordination. The Hermes run makes those research questions operational: interface design, delegation and validation become part of the code change itself.

The larger the work unit, the more verification and coordination become first-class engineering concerns.
ScaleTypical questionEvidence needed
EditCan the model make this local change?Diff, static check and focused test
IssueCan the agent resolve a repository task?Issue criteria, regression suite and review
EvolutionCan a system change structure without changing behaviour?Dependency map, broad tests, telemetry, replay and independent audit
  • EditCan the model make this local change?Diff, static check and focused test
  • IssueCan the agent resolve a repository task?Issue criteria, regression suite and review
  • EvolutionCan a system change structure without changing behaviour?Dependency map, broad tests, telemetry, replay and independent audit
SECTION 03

The hidden product was the harness#

The forensic post-mortem is more valuable than the victory lap. It records 1,393 delegated children and 93,284 calls in the replayable run, with a historical cost estimate of $19,302.59. Cache writes represented 58% of the estimated spend, output 24% and cache reads 19%, with rounding taking the shares just over 100%. Hermes Agent forensic post-mortem For agentic coding tools, this is the den where the real engineering lessons live.

That is not the same as the roughly $25,000 figure circulated in summaries, and it is nowhere near a verified $2 million labour saving. The post-mortem explicitly refuses to add its measured and modelled cost cohorts into one total because they overlap. It offers a bounded, modelled saving for particular optimisations, not an audited return on investment for the refactor. Careful readers should keep those categories apart.

More important, the run discovered failures in the machinery doing the work. The post-mortem tracks thirteen harness fixes, including a time-to-live repair. An independent review then reported three blockers and ten defects. The project says those findings were closed. This is not embarrassing residue. It is the central educational result: at high concurrency, orchestration code is production code, telemetry is product telemetry, and the repair loop must include the system that assigns the repairs.

Research on multi-agent software systems predicts the same pressure points. SEMAP names under-specification, coordination misalignment and inappropriate verification as recurring failure classes. Towards Engineering Multi-Agent LLMs, 2025 The Hermes record gives each one a concrete shape. A child can receive an underspecified boundary, two children can touch the same seam, or a passing test can verify the wrong behavioural contract.

Cache writes
58%
Model output
24%
Cache reads
19%
Historical estimated cost share in the forensic post-mortem. The published percentages total 101% because they are rounded.
At this scale, the harness is not scaffolding around the work. It is part of the software system being tested.
folkfox analysis of the Hermes post-mortem
SECTION 04

Verification is the architecture#

A million-line refactor can preserve tests while still moving risk around. Tests are executable claims about behaviour, but they are never the whole behaviour. They may couple to private implementation details, omit rare integrations or accept the same mistaken assumption as the code. Agentic coding tools therefore need layers of evidence rather than one green badge.

The first layer is mechanical: syntax, types, lint rules, dependency checks and focused unit tests. The second is behavioural: integration tests, end-to-end journeys and regression suites. The third is structural: limits on file size, function complexity, dependency direction and duplicated code. The fourth is forensic: durable logs that let a reviewer reconstruct who changed what, under which instruction, after which preceding change.

The fifth layer is adversarial human review. This is where a reviewer asks whether the decomposition was sensible, whether the tests guarded the right contract and whether the claimed economics follow from the evidence. Human review is not a ceremonial approval placed after autonomy. It is an intentionally different failure detector.

There is also a useful counterweight to orchestration maximalism. Agentless showed that a deliberately simpler localisation-and-repair pipeline could reach 32% on SWE-bench Lite at a reported experimental cost of $0.70 per issue. Agentless, 2024 Different tasks justify different machinery. The right agentic AI coding tool is not the one with the most agents. It is the smallest system that can make the change, prove the result and leave a legible record.

This is why an AI agent in a regulated workflow should be judged by its evidence surface, not its demo velocity. The same principle applies to background execution in Hermes Agent sessions and to the wider Hermes and Pantheon agent model. Autonomy changes who performs the keystrokes. It does not remove accountability for the outcome.

Constrain

Define repository boundaries, behavioural invariants, forbidden changes and stop conditions.

Observe

Record prompts, tool calls, patches, test results, retries, costs and dependency conflicts.

Verify

Run focused, integration, structural and regression checks against explicit acceptance criteria.

Replay

Preserve enough state to reproduce disputed changes and diagnose orchestration failures.

Review

Use an independent reviewer to challenge both the code and the economic story.

SECTION 05

Five honest questions for buyers#

The market for agentic coding tools is filling with claims about speed, labour replacement and unattended delivery. The Hermes case suggests a more discriminating set of questions. First, what is the unit of autonomy? A tool that completes a local edit is not equivalent to a system that plans, delegates and integrates a repository campaign. Follow the scent from the demo to the acceptance criteria.

Second, what survives after the run? Ask for the change graph, test record, task lineage, cost ledger and retry history. A polished pull request without the underlying trace makes it difficult to distinguish confident coordination from lucky convergence. An agentic coding assistant should improve the reviewer’s visibility, not simply increase the volume arriving at the review queue.

Third, how does the system handle collision and contradiction? Parallel children will eventually share a boundary. The platform should declare how it locks work, detects overlapping edits, rebases dependent tasks and escalates disagreement. Coordination is not a decorative feature once the work crosses files and teams.

Fourth, which claims are measured and which are modelled? Token spend can be read from a ledger. Engineering time saved usually depends on a counterfactual. Code removed can be counted. Maintainability gained requires a proxy and time. Vendors of agentic coding tools should label those categories plainly. Buyers should resist converting every technical metric into money before the causal bridge exists.

Fifth, who can stop the system? Every serious agentic coding workflow needs budget ceilings, test-failure thresholds, protected paths and a human halt. The foxiest implementation is not the one that never asks for help. It is the one that recognises when confidence has become theatre.

Hermes Agent deserves attention because the repository record is unusually rich. The codebase became materially smaller, the oversized structures almost disappeared, and the team published the machinery’s misses as well as its wins. That is a more educational contribution than the $2 million line. The achievement is not that software maintenance has become free. It is that agentic coding tools can organise large-scale maintenance differently, provided the evidence grows with the autonomy.

For teams building an evaluation plan, folkfox’s evidence-led strategy work starts in the same place: define the decision, preserve the trail and refuse the number that cannot yet carry its own weight.

@NousResearch
1,393 subagents and nineteen hours later, the codebase was 34.4% smaller, saving us nearly $2m in engineering hours.
15 September 2026, 22:11 GMTView on X
Questions

Frequently asked questions#

What are agentic coding tools?

Agentic coding tools can plan and execute multi-step software tasks using repository access, terminals, tests and other tools. More advanced systems may delegate work to child agents and integrate their results. Autonomy varies widely, so buyers should ask what the system can do without intervention and how its work is verified.

What did Hermes Agent actually achieve?

According to the merged pull request, the project reduced non-test Python source from 1,063,826 to 698,363 lines, cut files of at least 5,000 lines from 37 to 6 and cut functions of at least 300 lines from 192 to 2. The stated goal was no behaviour change.

Did the Hermes refactor really save $2 million?

Nous Research made that claim, but the public materials do not provide a reproducible calculation for it. The forensic post-mortem gives a historical estimated run cost of $19,302.59 and explicitly avoids one aggregate saving because some cost cohorts overlap. Treat the $2 million figure as an unverified vendor estimate.

Why were 1,393 subagents used?

The system decomposed a repository-wide refactor into many bounded tasks that could be run in parallel. The number demonstrates orchestration scale, not automatically higher quality. The post-mortem’s harness fixes and review defects show why coordination and verification must grow with the number of agents.

Can agentic coding tools safely refactor a large codebase?

They can contribute safely when the work has explicit boundaries, strong regression tests, protected paths, detailed telemetry, cost limits and independent review. A passing suite is necessary but not sufficient. Teams should also verify structural goals, integration behaviour and the assumptions encoded in the tests.

What should companies compare when choosing an agentic coding assistant?

Compare the unit of autonomy, repository and terminal controls, test integration, traceability, collision handling, cost reporting, stop conditions and review experience. Prefer a system that leaves a clear evidence trail over one that only advertises the largest agent count or fastest demo.

Keep reading

Read more on this topic#

Need an AI claim that can survive the audit?

folkfox turns fast-moving technical evidence into clear strategy, useful editorial and decisions that remain legible after the launch post fades.

Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.