Skip to main content
•10 min read

What Agent Harnesses Actually Ship for Spreadsheets

In the previous post I measured what one bundled runtime, the LibreOffice copy Simon Willison found inside the ChatGPT desktop app, does to a workbook on save. Is that an OpenAI quirk, or the industry default across agent runtimes?

So I surveyed how agent harnesses handle spreadsheets: the major coding agents' document skills, the spreadsheet MCP servers, a popular agent sandbox image, and the benchmark used to grade spreadsheet agents. I read each one's source, then ran falsification probes against the strongest candidates. Every probe emits a receipt with source and output hashes, linked at the bottom.1

TL;DR: what the survey found

12

agent spreadsheet pipelines surveyed across eight lanes: coding-agent skills, MCP servers, a sandbox image, and one benchmark

0

pipelines preserve untouched OOXML package parts on save; every edit path rebuilds the file through a DOM rewrite or a LibreOffice filter pass

1

native recalculation engine (formualizer, inside agent-spreadsheet), which still drops charts, pivots, and external links on a one-cell edit

The matrix

How twelve agent spreadsheet pipelines edit and save Excel files, observed September 3, 2026
PipelineEdit pathRecalculationSave behavior
Anthropic xlsx skill (Cowork lineage)openpyxl + pandasHeadless LibreOffice, StarBasic macro, compiled C socket shimDOM rewrite, then LibreOffice rewrite
Nous Research hermes-agentopenpyxl CLI scripts (forked from Anthropic's skill)soffice convert spawnSame double rewrite
Xiaomi MiMo-Codeopenpyxl + pandas + xlsxwriter, manual raw-XML hatchsoffice convert spawnLossy rewrite; only the manual hatch preserves
hewliyang/headless-excelopenpyxl proxiesPersistent LibreOffice UNO daemonopenpyxl save, then daemon store(): double rewrite
Four LibreOffice-branded MCP serversExcelJS or openpyxl; two have no write tool at allNoneLossy DOM rebuild, or reads flattened to CSV
PSU3D0/agent-spreadsheet (Rust MCP)umya-spreadsheet, full re-serializationNative in-process formualizer enginePreserves VBA; drops charts, pivots, external links
sheetforge / xlsx-tools MCPsopenpyxlNone, or a soffice pass after every writeLossy rewrite, optionally double
E2B default code-interpreter sandboxopenpyxl 3.1.5, pinned in the imageNone; no office suite installedFormulas written with empty cached values
SpreadsheetBench (NeurIPS 2024 eval)Agent-authored openpyxl in DockerExcel COM, or LibreOffice added later, just to gradeValue-only grading over rewritten files

Two of these pipelines, Anthropic's skill and agent-spreadsheet, represent the most sophisticated engineering in the survey and still hit the same wall.

Anthropic's skill is the origin, and it tells you so

The xlsx skill in anthropics/skills is the lineage most other harness skills descend from. It edits with openpyxl and pandas, then recalculates by spawning headless LibreOffice with an injected StarBasic macro. When the sandbox blocks the UNIX sockets LibreOffice needs, it compiles a C socket shim (lo_socket_shim.so, via gcc -shared -fPIC) on the fly and injects it with LD_PRELOAD.

To recalculate a cell, that pipeline builds a temporary office-suite profile, injects a macro, and compiles a C shim inside the sandbox.

The skill's own documentation lists the failure modes directly:

"Never use XLOOKUP, XMATCH, SORT, FILTER, UNIQUE, or SEQUENCE. ... they are spilling array functions and an openpyxl-written file has no spill metadata, so only the top-left cell of the range gets a value — and recalc.py reports total_errors: 0 on the truncated result."

"A workbook that links to another file loses those links if you re-save it with openpyxl and then recalculate. ... openpyxl strips that value on save; LibreOffice then has to resolve the reference for real, fails, writes #NAME?, and deletes every link."

"data_only=True is destructive if you save."

A silent wrong answer that reports zero errors is catastrophic for an autonomous workflow. To measure actual execution beyond the documentation, I ran the exact recalc.py bundled in Claude Desktop 1.40609.0, hash-pinned and unmodified, against five controlled fixtures:2

  • The openpyxl stage preserved the macro binary; the LibreOffice stage then rewrote its bytes while returning success.
  • Chart part inventories survived, but their XML was broadly reserialized.
  • Pivot part inventories survived, but their XML was broadly reserialized.
  • Both external-link parts were deleted.
  • FILTER produced only its top-left value (D1=10, D2:D4 empty) while the script returned status: success and total_errors: 0.

A second probe ran the same path over a six-function fixture: SUM evaluated correctly, while XLOOKUP, LET, FILTER, UNIQUE, and SEQUENCE cached #NAME?.

The warnings do not survive reimplementation

The skill has been cloned twice publicly, propagating the failure modes into downstream projects.

Nous Research's hermes-agent started as a verbatim copy, carrying Anthropic's proprietary license file and the StarBasic macro line for line, before being rewritten as an MIT fork that drops the macro and shim for a direct soffice --convert-to spawn. Xiaomi's MiMo-Code is a clean-room rewrite against the ECMA-376 spec, written specifically to avoid the Anthropic license. Its create.md encourages XLOOKUP and never mentions that the LibreOffice recalc path silently truncates spilling array functions. The documentation warnings that accompany Anthropic's skill were dropped in both descendants.

Downstream users inherit the failure modes without the warnings.

Even in-process calculation hits the preservation wall

The most ambitious architecture in the survey attempts in-process calculation directly: PSU3D0/agent-spreadsheet is a Rust MCP server that embeds an in-process formula engine (formualizer), avoiding both Python DOM serialization and LibreOffice subprocesses. I exercised its published v0.15.0 release binary directly on the shared test fixtures.

It computed correct scalar values across all six test functions, including modern functions like XLOOKUP that no LibreOffice-based harness handled.

But calculation is only half the engineering problem. When saving the file, its underlying serializer (umya-spreadsheet) still rebuilds the package from scratch:

  • A one-cell edit stripped the chart definition (xl/charts/chart1.xml).
  • The same edit deleted all five pivot table and pivot cache parts.
  • It also dropped both external-link parts.
  • And while scalar formula anchors evaluated correctly, dynamic array functions (FILTER, UNIQUE, SEQUENCE) left their spill ranges completely empty even as the API returned status: success and unsupported_formula_cells: 0.

This illustrates the core architectural trap: an agent runtime cannot treat a spreadsheet as just numbers and formulas. If the file layer rebuilds the package from scratch, updating a single number still destroys the surrounding corporate workbook. Charts, pivots, and cross-sheet models disappear on save.

On the shared fixture, WolfXL matched the scalar anchors, materialized the dynamic array spill ranges, and preserved the untouched package parts; the Cowork pipeline returned #NAME? for the modern functions. That cross-engine comparison is documented in its own receipt.

The defaults underneath

The sandbox layer makes the same choice. E2B's default code-interpreter image pins openpyxl==3.1.5, pandas, and xlrd, and installs no office suite anywhere in the build chain.3 openpyxl never evaluates formulas, so an agent writing =SUM(A1:A10) in that sandbox produces a file whose cached value is empty. Reopen it with data_only=True and you get None. The documented escape hatch is installing LibreOffice at runtime: 500 MB and tens of seconds, per sandbox run.

Even the benchmark inherits it

SpreadsheetBench, the NeurIPS 2024 benchmark for spreadsheet agents, grades by reading cached values. Because openpyxl never calculates, the authors had to spawn desktop Excel over COM, later a headless LibreOffice pass, just to produce values to grade against. Their own issue tracker calls the resulting misgrading "systematic false negatives."

At the benchmark's exact source revision, two probes:

  • deleting all 14 XLOOKUP formulas in one task's declared answer range still passes the official grading function, because the recalculated ground-truth caches for those cells were blank;
  • comparing all 600 sample answer workbooks against themselves, which should pass trivially, fails three cases because the evaluator does not trim spaces in comma-separated answer ranges.

The evaluation layer measuring "how good agents are at spreadsheets" runs on the same file plumbing, and its failures are silent in both directions.

What to do with this

The pattern is structural: the Python ecosystem's fast spreadsheet tools are one-way (read-only readers, write-only writers), so anything that must edit the workbook a company already uses converges on a slow DOM rebuild or a 400 MB office suite in a container, and neither preserves what it did not touch.

The fix is not a better prompt. After any agent edit, verify two things against the saved file: the intended change landed, and the protected state around it provably did not move. The previous post shows the package-diff method and receipts for the Excel side, and WolfPPT applies the same verify-the-save discipline to decks (pip install wolfppt && wolfppt demo). If a workbook in your pipeline cannot afford silent drift, run the fit check before the agent does.

Evidence receipts

Every executed claim above maps to a compact JSON receipt with pinned source or binary hashes.1

Survey probes, comparisons, and benchmark audits (10 pinned artifacts)

Each receipt pins its source or binary identity and its inputs and outputs by hash, so every claim above can be regenerated from the referenced revisions.


Footnotes

  1. Probes measure exact-byte package preservation and formula cache correctness on pinned test fixtures. Benchmark timings for WolfXL are recorded in the companion study. ↩ ↩2
  2. The Cowork probe uses the exact skill code shipped in Claude Desktop 1.40609.0 with pinned script and skill hashes. Cowork's service-side LibreOffice binary was not available for inspection, so formula and export outcomes were produced with Homebrew LibreOffice 26.2.5.2 and are not attributed to Anthropic's service infrastructure. Which surfaces load this exact skill varies by product and account configuration. ↩
  3. E2B image composition is read from the project's own Dockerfile and pinned requirements file. The cached-value failure was reproduced in an equivalent dependency environment with the exact pinned openpyxl version, not executed inside E2B infrastructure. Other sandbox vendors were not surveyed. ↩
XLinkedIn

Building something where correctness matters?

Get the next post