The Hard Part of Agentic Coding Was Making Progress Hard to Fake
I built a 424k-line Rust/Python spreadsheet SDK with coding agents over five months. The turning point was the day my own verifier erased 24 functions it had previously called verified.
I spent five months building WolfXL, a Rust and Python spreadsheet SDK. Coding agents carried most of the implementation work. The system now contains 423,749 lines of tracked programming-language code across 16 Rust crates and a registered formula ledger of 416 functions under a three-oracle plus adversarial verification definition.
The part worth writing down is not the volume. It is that the most valuable thing I built was the machinery that made the agents' claims falsifiable, and the moment I started trusting the system was the moment it lowered its own score.
The problem with a big vague goal
The objective was underspecified on purpose: approach the capability of a commercial spreadsheet SDK. Formula calculation, rendering to PNG and PDF, pivot cache recomputation, four file formats, and the compatibility surface that real spreadsheet code depends on.
An underspecified goal plus a tireless implementer is a Goodhart machine. If the same process that writes the code also decides whether the code works, the tracker will go up and the product will not. Any single number an agent can both move and grade is worthless.
So the first design question was not "how do I get agents to write more code." It was "what would make it impossible for an agent to convince me it had finished something it had not."
The architecture: separate the writer from the judge
I split the work into lanes with different proof rules, and put the evidence system outside every implementation lane.
North Star charter: commercial spreadsheet capability
|
live tracker recomputed
from committed evidence
|
+-------------+-------+-------+-------------+
| | | |
calc lane render lane pivot lane adversarial lane
formulas XLSX -> PNG cache finds cases that
/PDF recompute should fail
| | | |
+------ three calc oracles ---+-------------+
| A frozen Python behavior |
| B LibreOffice recalculation |
| C Excel-authored cached values |
+-------------------+-----------------------+
|
out-of-loop verifier
recomputes every total
from raw entries
|
single-writer evidence ledgers
|
human integration boundary
(I merge, not the loop)
|
back to the tracker
The rules that made it work were mostly rules about who is allowed to say what:
- The tracker selects work. It never proves work.
- An implementation agent never owns the oracle that grades it.
- Calc has to agree with three independent sources, not one.
- Adversarial case discovery is a separate lane from implementation.
- A function counts only when its entire case set passes, not when a flag flips.
- Append-only logs can be union-merged; structured evidence has exactly one writer.
- Shared interfaces and integration stay sequential even when implementation runs in parallel.
Each loop iteration then looked like this:
- Recompute status from committed evidence.
- Pick the highest-value unblocked item.
- Make one implementation or evidence move.
- Gate it against an independent oracle.
- Count it only if the authoritative ledger moved.
- If the required evidence does not exist, stop and record a blocker.
Step 6 is the one that makes the rest honest.
The moment I trusted it: 200, then 176
On 2026-06-19 the calc snapshot looked great:
adversarial_resolved = 202
fully_verified = 200
Then two adversarial batches landed that added an out-of-loop verifier and a test which recomputed the ledger totals from entry-level evidence instead of trusting per-function flags. The stricter recount produced:
adversarial_resolved = 178
fully_verified = 176
implemented = unchanged
The reconciliation is recorded in the project charter, in the commit that documents it:
The 2026-06-19 snapshot over-reported
adversarial_resolved(202) andfully_verified(200); both were corrected to 178 / 176 when the verifier recomputed totals from entry-level evidence. That was a verifier re-score, not a work retraction: zero implementations were removed.
Twenty-four functions did not regress. Twenty-four functions had been counted as verified on the strength of a self-reported flag, and the new measurement refused to accept flags as evidence. Nothing was deleted; the scoreboard just stopped lying.
Development continued against the stricter number. Today the same ledger, on
canonical main, reports:
implemented_function_count = 416
fully_verified_function_count = 416
adversarial_resolved_count = 416
oracle_a (frozen Python) = 416
oracle_b (LibreOffice recalc) = 416
oracle_c (Excel cached values) = 416
headless = true
excel_gui_used = false
The progression that matters is the shape, not the endpoint:
200 self-reported, before the verifier existed
176 independently recomputed
416 current canonical ledger under the stricter definition
Any system that cannot produce a number going down does not have a measurement. It has a mood.
What one loop actually did for six hours
The clearest single run is a calc goal loop on 2026-06-20 that ran six hours and twenty-six minutes. Its own summaries record the frontier moving:
fully_verified: 176 -> 273 -> 321 -> 340 -> 367 -> 403
Its transcript contains 56 distinct test(calc): promote ... adversarial evidence commit references, repeated focused Rust test runs, LibreOffice
recalculation checks, and Excel-cached-value comparisons.
Three details from that run are the reason I keep it as the reference example, and none of them are about throughput:
- It stopped on a boundary it was not allowed to cross. A merge conflict would have required editing files outside the calc lane, so it halted and preserved the blocker instead of reaching across lanes.
- I changed the operating procedure when the process became the bottleneck. Rather than running full CI inside every hot iteration, I moved CI and review to the merge boundary and kept authoritative ledger movement as the in-loop gate.
- It closed a declared gate rather than a vibe. The commit that closed the
then-current 403-function Oracle C gate is in canonical
main.
"56 commits in six hours" is a throughput number and I do not care about it. "Selected the next ledger gap, ran the independent checks, refused to cross a lane boundary, and closed a declared gate" is an engineering number.
When the metric was wrong, I replaced the metric
The render lane started as evidence-neutral scaffolding in its own crate and grew into real PNG and PDF output. Its first honest measurement attempt failed.
I wanted a cross-engine absolute RMSE threshold against a reference renderer. It could not distinguish correct output from a catastrophically broken blank image, because two different renderers legitimately disagree by more than a blank page differs from a page.
The wrong move is to keep the impressive-sounding metric and tune the threshold. What the render lane did instead was drop to a narrower claim it could actually defend: a frozen-baseline perceptual drift gate, which detects regressions against a committed reference rather than pretending to prove absolute cross-engine correctness.
The canonical render ledger now reports 27 geometry fixtures and 27 perceptual comparisons with 27 verified references, zero missing, zero stale, and every readiness flag true. That is a smaller claim than the one I originally wanted, and it is worth more because it is true.
When the agent refused to invent evidence
The pivot lane was the highest-risk one, because pivot cache recomputation
touches the shipped writer path. Its gate compares regenerated
pivotCacheRecords against Excel-authored cached records after canonical
sorting.
Early on, three of nine tracked scenarios were blocked: the Excel-authored oracle evidence did not exist yet, and native preflight support was missing.
A headless loop optimizing for a green tracker would have found a way to call those scenarios passing. This one filed them as blockers, and the work to unblock them was the work of producing real oracle evidence: a differential recompute gate, a headless non-sum oracle authoring pipeline, and support for month-grouped date recomputation.
All nine tracked scenarios are now preflight_supported on canonical main,
each bound to an Excel-authored cached-records oracle with canonical_sort: true and a recorded workbook_sha256.
The story is not that the agents removed the external dependency. It is that they represented it honestly instead of routing around it.
Five months, four operating models
WolfXL did not start as an autonomy experiment. It started on 2026-02-14 as a prototype inside an Excel-library benchmark, and the operating model changed four times as the project outgrew each one.
February AI pair programming
prototype, extract, package, iterate
|
late April-May contract-driven parallel waves
publish contracts, dispatch named waves,
disjoint slices, separate reviewers
|
June persistent evidence-gated goal loops
bounded goals survive turns and compactions,
independent oracles define progress
|
July reviewer-heavy integration
scouts, implementers, reviewers, auditors
around immutable commits
| Date | What changed | Evidence |
|---|---|---|
| Feb 14 | Prototype adapter inside the benchmark repo | first hybrid-adapter commit |
| Feb 15 | Extracted to a standalone repository | first standalone release: 31 files, 9,630 insertions, agent co-author trailer |
| Feb 19 | Recursive formula evaluator | one commit adding 142 tests |
| Apr 19 | Parity harness, ratchet, and known-gaps ledger land on main | 303a16288 |
| Apr 24 | Contracts published before dispatching Wave 1B/1C; disjoint slices implemented in parallel; reviewer fixes applied | four branch commits, two merges seven seconds apart |
| Apr 27 | The native writer wave lands on main | 3bf845cef |
| Jun 17 | Autonomous goal-loop orchestration and a readiness gate | dc5c95d87 |
| Jun 19-20 | Adversarial evidence campaign; the verifier corrects 200 to 176 | 61ce2d718 |
| Jul 18-21 | Specialist review and controlled integration | converges through PR #373 |
The interesting artifact from the April period is not the code. Between April 24 and May 30, the preserved agent archive holds 133 subagent records across 25 parent sessions: 68 reconnaissance, 43 implementation, 15 code review, and 7 planning. The prompts show named waves, pod-owned slices, contracts published before dispatch, branch ownership, and a sequential merge order.
That is execution telemetry, not a productivity count. What it demonstrates is that the repository had an operating system months before it had a formal goal loop. The goal-loop layer made the program persistent across turns and context compactions. It did not invent the structure.
By July the mix had inverted: in one long integration session, most delegated work was review, auditing, or reconnaissance rather than implementation. That is what maturity looked like here, and it is the opposite of the usual autonomous-coding demo.
Scale, and what scale does not mean
Canonical main snapshot:
| Metric | Count |
|---|---|
| Tracked files | 3,677 |
| Programming-language files | 1,123 |
| Code LOC | 423,749 |
| Non-test code LOC | 305,437 |
| Test code LOC | 118,312 |
| Function bodies and declarations | 19,950 |
| Types | 1,758 |
| Rust workspace crates | 16 |
About 225,700 lines are the core product under python/, src/, and
crates/. Another 118,312 are tests, and roughly 69,735 more implement
benchmarks, oracles, ledgers, and claim-audit tooling.
That last number is the actual thesis of the project. Roughly seventy thousand lines are dedicated to benchmarks, oracles, ledgers, and claim-audit tooling whose job is to make the rest of the system difficult to misrepresent.
LOC also measures maintained surface, not originality. WolfXL deliberately mirrors parts of an existing Python library's compatibility surface, because matching a widely used API is the point of a compatible SDK.
What I did and what the agents did
The honest division of labor:
Mine. The objective and its decomposition into lanes. The rule that an implementer never owns its own oracle. The three-oracle calc definition. The decision to accept a narrower render claim after the RMSE metric failed. Lane charters, worktree layout, and reviewer contracts. Relaxing my own rules when they became the bottleneck. I retained authority over integration and every public claim.
Theirs. The overwhelming majority of implementation, test authoring, reconnaissance, and review passes. Sustained execution across sessions I was not watching.
The interesting part. The system was allowed to tell me no: to block on missing evidence, to reject a bounded commit in review, and to lower a number I would have preferred to keep.
What this does not prove
I would rather put the limits on the page than have someone find them later.
- No parity claim. Not Excel, not Aspose.Cells, not any commercial SDK. The
repository's own audit still reports
sota_claim_ready: false. - 416 of 416 is the registered calc ledger under its own phase-0-plus-three-oracles-plus-adversarial definition. It is not the whole Excel function universe.
- Session event counts, task-dispatch counts, and subagent counts are runtime telemetry. They are not counts of accepted patches, features, or simultaneous workers.
- A worktree is not an agent. Parallel branches and near-simultaneous merges show parallel work streams, not a precise concurrency number.
- The repository's capability registry resolves its 26 tracked capabilities to "verified" from committed progress events. Only calc, render, and pivot carry independent oracle evidence, so I do not count the other capabilities as verified anywhere.
- Co-author trailers prove agent assistance, not which lines a model wrote.
- Nothing here was human-free. I set the objective, changed the operating rules, owned integration, and kept publishing authority.
- The repository is private, so these numbers are attributable to named canonical artifacts and commits, but not independently checkable by a reader today. I will walk any of them line by line on request.
Questions I get asked
Did agents build this by themselves? No. Agents wrote most of the code. I designed the goal decomposition and proof rules, retained integration and publication authority, and intervened whenever my own process became the constraint.
What was the single highest-leverage thing you built? The out-of-loop verifier that recomputed ledger totals from raw entries instead of trusting per-function flags. It is the reason a number in this project can go down.
How do you stop an agent from gaming its own metric? Separate the writer from the judge, require agreement across independent sources, count only whole passing case sets, give evidence files a single writer, and make "stop and record a blocker" a first-class success outcome.
What happens when a metric turns out to be invalid? Replace it with a narrower one you can defend. The render lane traded an impressive cross-engine RMSE claim for a frozen-baseline drift gate, and the smaller claim is the one that held.
Why does the test and evidence code approach a third of the repository? Because the expensive failure mode of agentic development is not bad code. It is confident, well-tested-looking code that nobody can verify.
The transferable part
Strip out the spreadsheets and what is left is a method:
- Turn a vague objective into lanes with different proof rules.
- Give every lane an oracle it does not control.
- Recompute status from committed evidence, never from self-reports.
- Make blocking and abstaining first-class outcomes.
- Keep integration and public claims behind a human boundary.
- When a metric cannot distinguish correct from broken, narrow the claim instead of tuning the threshold.
That is the part I would bring to someone else's agent system on day one.
You might also like
Interested in working together?
Get in touch →


