It's since then expanded to cover everything from editing tables, hyperlinks, footnotes, and a lot more. Now it's a pretty powerful tool that can trivially fill out a MNDA form, mark up a contract, author a poetry booklet, and fill out an invoice, which is now the eval suite where the numbers in the title come from.
You might be asking, "why did you do all of this?" Well, I'm building an agent harness for normies that are not gonna know what a token even is but just want their stuff not to take an epoch and a half to run. So I've got to make the tools be MUCH more optimal than they've even been.
I figure putting them out to the community and inviting all of you to help me might be a way to do that =).
Generally, they will have to use more tokens to reach the same outcome, per some research: https://arxiv.org/abs/2604.02460
It is one of the biggest facilitator of vendor lock in in the history of computing.
I say it’s as if “Claude Code & Microsoft Office had a baby...”
Code available: https://github.com/espressoplease/smalldocs
Discord: https://discord.gg/txjATTsDaq
Sample document: https://smalldocs.org/blogs/what-is-a-smalldoc
Invoked via Claude Code by saying stuff like: “sdoc me the plan for this feature”, or “dig into our logs and sdoc me a report on our latency"
And given that LibreOffice and Google Docs are pretty good nowadays and OOXML is an open standard now post the monopoly rulings of the 90s, it's not quite as bad as it used to be.
https://github.com/iOfficeAI/OfficeCLI
https://github.com/rcarmo/python-office-mcp-server
https://github.com/espressoplease/smalldocs
It would be great to see some benchmarks for token use, correctness, features etc.
I'm also working on letting agents read/edit word docs but exposing it as a simple MCP
www.vespper.com
That being said:
- every single command you write MUST have a help text. that help text should also tell agents what the easiest path forward it. for example, the locator system in this CLI relies on the positions on paragraphs. these can change after an edit. so I prompt in the help text to use batch edits.
- errors must be clearly marked. for example, if an agent is doing a replace, but the replace comes back with "0 edits made", that's an error. otherwise the agents carry on going on obliviously.
- a lot of the CLI tools decide to use JSON as the default output. I started out that way as well. for agents, that immediately means having to invoke jq or a similar other tool to get it to be understood. the weaker the model, the worse they do on that. I found that providing markdown with annotations beats JSON on most tasks.
- use image reading as a last resort. if it's possible to give the model the information it needs as text, it will do a better job working with it than it will thinking through an image. thus, for this CLI, i do have a render command, but explicitly prompt in the help texts to use it for things that are not obvious from the markdown like layout.