What Agent Harnesses Actually Ship for Spreadsheets
In the previous post I measured what one bundled runtime, the LibreOffice copy Simon Willison found inside the ChatGPT desktop app, does to a workbook on save. Is that an OpenAI quirk, or the industry default across agent runtimes?
So I surveyed how agent harnesses handle spreadsheets: the major coding agents' document skills, the spreadsheet MCP servers, a popular agent sandbox image, and the benchmark used to grade spreadsheet agents. I read each one's source, then ran falsification probes against the strongest candidates. Every probe emits a receipt with source and output hashes, linked at the bottom.1
TL;DR: what the survey found
12
agent spreadsheet pipelines surveyed across eight lanes: coding-agent skills, MCP servers, a sandbox image, and one benchmark
0
pipelines preserve untouched OOXML package parts on save; every edit path rebuilds the file through a DOM rewrite or a LibreOffice filter pass
1
native recalculation engine (formualizer, inside agent-spreadsheet), which still drops charts, pivots, and external links on a one-cell edit
The matrix
| Pipeline | Edit path | Recalculation | Save behavior |
|---|---|---|---|
| Anthropic xlsx skill (Cowork lineage) | openpyxl + pandas | Headless LibreOffice, StarBasic macro, compiled C socket shim | DOM rewrite, then LibreOffice rewrite |
| Nous Research hermes-agent | openpyxl CLI scripts (forked from Anthropic's skill) | soffice convert spawn | Same double rewrite |
| Xiaomi MiMo-Code | openpyxl + pandas + xlsxwriter, manual raw-XML hatch | soffice convert spawn | Lossy rewrite; only the manual hatch preserves |
| hewliyang/headless-excel | openpyxl proxies | Persistent LibreOffice UNO daemon | openpyxl save, then daemon store(): double rewrite |
| Four LibreOffice-branded MCP servers | ExcelJS or openpyxl; two have no write tool at all | None | Lossy DOM rebuild, or reads flattened to CSV |
| PSU3D0/agent-spreadsheet (Rust MCP) | umya-spreadsheet, full re-serialization | Native in-process formualizer engine | Preserves VBA; drops charts, pivots, external links |
| sheetforge / xlsx-tools MCPs | openpyxl | None, or a soffice pass after every write | Lossy rewrite, optionally double |
| E2B default code-interpreter sandbox | openpyxl 3.1.5, pinned in the image | None; no office suite installed | Formulas written with empty cached values |
| SpreadsheetBench (NeurIPS 2024 eval) | Agent-authored openpyxl in Docker | Excel COM, or LibreOffice added later, just to grade | Value-only grading over rewritten files |
Two of these pipelines, Anthropic's skill and agent-spreadsheet, represent the most sophisticated engineering in the survey and still hit the same wall.
Anthropic's skill is the origin, and it tells you so
The xlsx skill in anthropics/skills is the lineage most other harness skills descend from. It edits with openpyxl and pandas, then recalculates by spawning headless LibreOffice with an injected StarBasic macro. When the sandbox blocks the UNIX sockets LibreOffice needs, it compiles a C socket shim (lo_socket_shim.so, via gcc -shared -fPIC) on the fly and injects it with LD_PRELOAD.
To recalculate a cell, that pipeline builds a temporary office-suite profile, injects a macro, and compiles a C shim inside the sandbox.
The skill's own documentation lists the failure modes directly:
"Never use
XLOOKUP,XMATCH,SORT,FILTER,UNIQUE, orSEQUENCE. ... they are spilling array functions and an openpyxl-written file has no spill metadata, so only the top-left cell of the range gets a value — andrecalc.pyreportstotal_errors: 0on the truncated result."
"A workbook that links to another file loses those links if you re-save it with openpyxl and then recalculate. ... openpyxl strips that value on save; LibreOffice then has to resolve the reference for real, fails, writes
#NAME?, and deletes every link."
"
data_only=Trueis destructive if you save."
A silent wrong answer that reports zero errors is catastrophic for an autonomous workflow. To measure actual execution beyond the documentation, I ran the exact recalc.py bundled in Claude Desktop 1.40609.0, hash-pinned and unmodified, against five controlled fixtures:2
- The openpyxl stage preserved the macro binary; the LibreOffice stage then rewrote its bytes while returning success.
- Chart part inventories survived, but their XML was broadly reserialized.
- Pivot part inventories survived, but their XML was broadly reserialized.
- Both external-link parts were deleted.
FILTERproduced only its top-left value (D1=10,D2:D4empty) while the script returnedstatus: successandtotal_errors: 0.
A second probe ran the same path over a six-function fixture: SUM evaluated correctly, while XLOOKUP, LET, FILTER, UNIQUE, and SEQUENCE cached #NAME?.
The warnings do not survive reimplementation
The skill has been cloned twice publicly, propagating the failure modes into downstream projects.
Nous Research's hermes-agent started as a verbatim copy, carrying Anthropic's proprietary license file and the StarBasic macro line for line, before being rewritten as an MIT fork that drops the macro and shim for a direct soffice --convert-to spawn. Xiaomi's MiMo-Code is a clean-room rewrite against the ECMA-376 spec, written specifically to avoid the Anthropic license. Its create.md encourages XLOOKUP and never mentions that the LibreOffice recalc path silently truncates spilling array functions. The documentation warnings that accompany Anthropic's skill were dropped in both descendants.
Downstream users inherit the failure modes without the warnings.
Even in-process calculation hits the preservation wall
The most ambitious architecture in the survey attempts in-process calculation directly: PSU3D0/agent-spreadsheet is a Rust MCP server that embeds an in-process formula engine (formualizer), avoiding both Python DOM serialization and LibreOffice subprocesses. I exercised its published v0.15.0 release binary directly on the shared test fixtures.
It computed correct scalar values across all six test functions, including modern functions like XLOOKUP that no LibreOffice-based harness handled.
But calculation is only half the engineering problem. When saving the file, its underlying serializer (umya-spreadsheet) still rebuilds the package from scratch:
- A one-cell edit stripped the chart definition (
xl/charts/chart1.xml). - The same edit deleted all five pivot table and pivot cache parts.
- It also dropped both external-link parts.
- And while scalar formula anchors evaluated correctly, dynamic array functions (
FILTER,UNIQUE,SEQUENCE) left their spill ranges completely empty even as the API returnedstatus: successandunsupported_formula_cells: 0.
This illustrates the core architectural trap: an agent runtime cannot treat a spreadsheet as just numbers and formulas. If the file layer rebuilds the package from scratch, updating a single number still destroys the surrounding corporate workbook. Charts, pivots, and cross-sheet models disappear on save.
On the shared fixture, WolfXL matched the scalar anchors, materialized the dynamic array spill ranges, and preserved the untouched package parts; the Cowork pipeline returned #NAME? for the modern functions. That cross-engine comparison is documented in its own receipt.
The defaults underneath
The sandbox layer makes the same choice. E2B's default code-interpreter image pins openpyxl==3.1.5, pandas, and xlrd, and installs no office suite anywhere in the build chain.3 openpyxl never evaluates formulas, so an agent writing =SUM(A1:A10) in that sandbox produces a file whose cached value is empty. Reopen it with data_only=True and you get None. The documented escape hatch is installing LibreOffice at runtime: 500 MB and tens of seconds, per sandbox run.
Even the benchmark inherits it
SpreadsheetBench, the NeurIPS 2024 benchmark for spreadsheet agents, grades by reading cached values. Because openpyxl never calculates, the authors had to spawn desktop Excel over COM, later a headless LibreOffice pass, just to produce values to grade against. Their own issue tracker calls the resulting misgrading "systematic false negatives."
At the benchmark's exact source revision, two probes:
- deleting all 14
XLOOKUPformulas in one task's declared answer range still passes the official grading function, because the recalculated ground-truth caches for those cells were blank; - comparing all 600 sample answer workbooks against themselves, which should pass trivially, fails three cases because the evaluator does not trim spaces in comma-separated answer ranges.
The evaluation layer measuring "how good agents are at spreadsheets" runs on the same file plumbing, and its failures are silent in both directions.
What to do with this
The pattern is structural: the Python ecosystem's fast spreadsheet tools are one-way (read-only readers, write-only writers), so anything that must edit the workbook a company already uses converges on a slow DOM rebuild or a 400 MB office suite in a container, and neither preserves what it did not touch.
The fix is not a better prompt. After any agent edit, verify two things against the saved file: the intended change landed, and the protected state around it provably did not move. The previous post shows the package-diff method and receipts for the Excel side, and WolfPPT applies the same verify-the-save discipline to decks (pip install wolfppt && wolfppt demo). If a workbook in your pipeline cannot afford silent drift, run the fit check before the agent does.
Evidence receipts
Every executed claim above maps to a compact JSON receipt with pinned source or binary hashes.1
Survey probes, comparisons, and benchmark audits (10 pinned artifacts)
- Cowork skill fixture probe: exact bundled script, five preservation/recalc cases. SHA-256
57c55de76e42b14428ed2f0633d311e8e5714517cc46a501d13a983dabbeeb68 - Cowork formula matrix probe: six-function fixture through the same path. SHA-256
24233a6f3e3c89139ae63a8a6eb70d78b2f053cd0be229b02ae77294c7a67c6e - agent-spreadsheet preservation probe: published v0.15.0 binary, four one-cell edits. SHA-256
1c12549944a76571e3e6499db9e087f1b4a1f8e6300fc201e223326b2d461675 - Cross-engine formula comparison: formualizer vs Cowork's path vs WolfXL on one fixture. SHA-256
7348949dd840a75a6cc214992211ef763add7df33ddb602faca80a083c1ad0cb - E2B cached-value mechanism check: the E2B-pinned openpyxl version writes formulas with empty caches. SHA-256
5879a9061f3103c7720e9dcbfba041b5af353c592d292b4a712e192a085cd361 - SpreadsheetBench grading drift probe: exact-source evaluator on the six-function fixture. SHA-256
b3c47b4ee82eb94ce632e9cc4bbf84c1b35a9f19f5e8d1344ddf2efd0f5e6acc - SpreadsheetBench false positive probe: all 14 answer formulas deleted, grading still passes. SHA-256
4fff4fd989cf6e3c8b85dd9baa026b242d8ea2ee43f891e56c3545307211770e - SpreadsheetBench self-consistency audit: identical-workbook grading fails 3 of 600 cases. SHA-256
48627174cc5c3f2915e019b2e419258f883df9fc70fe29dd6e7dbff5cb55c426 - SpreadsheetBench sample formula scan: 78,116 formula cells across the sample corpus. SHA-256
d8987d37c5a20648baf78da0399db0ee6538cbed2967861cff535c4700a46ce0 - SpreadsheetBench answer integrity scan: formula-error census in answer ranges. SHA-256
121cc5d84adf31fa2a0c00bfdbe2c290bb20cdcfe32e6addb1e0277a9920beed
Each receipt pins its source or binary identity and its inputs and outputs by hash, so every claim above can be regenerated from the referenced revisions.
Footnotes
- Probes measure exact-byte package preservation and formula cache correctness on pinned test fixtures. Benchmark timings for WolfXL are recorded in the companion study. ↩ ↩2
- The Cowork probe uses the exact skill code shipped in Claude Desktop 1.40609.0 with pinned script and skill hashes. Cowork's service-side LibreOffice binary was not available for inspection, so formula and export outcomes were produced with Homebrew LibreOffice 26.2.5.2 and are not attributed to Anthropic's service infrastructure. Which surfaces load this exact skill varies by product and account configuration. ↩
- E2B image composition is read from the project's own Dockerfile and pinned requirements file. The cached-value failure was reproduced in an equivalent dependency environment with the exact pinned openpyxl version, not executed inside E2B infrastructure. Other sandbox vendors were not surveyed. ↩
Building something where correctness matters?
Get the next post