// HACKER NEWS — CYBERSECURITY
Launch HN: Vespper (YC F24) – SOTA Docx MCP
Today we're launching Vespper DOCX MCP: the first model fine-tuned specifically for editing Word documents, shipped as an MCP. On our internal benchmark, it allows agents to be 3× faster, 2× cheaper, and more accurate than the closest alternative.
Word documents are everywhere. In domains such as legal, finance, and healthcare, the Word document is the deliverable. Contracts, regulatory submissions, and audit reports get drafted, redlined, and signed in .docx, and companies usually have libraries of Word templates they work with on a regular basis.
That work is increasingly shifting to agents. Microsoft Copilot and Claude in Word have made in-document AI mainstream, while a growing number of vertical agents; particularly in legal tech—need to interact with .docx files.
However, AI agents are still struggling to do great work in Word documents.
We spoke with dozens of software engineers, mostly in the legal tech and health care spaces, who said they spend weeks and even months tuning their harnesses to edit Word documents reliably, and now are forced to maintain very complex in-house solutions.
In general, there exist several ways for agents to edit Word docs today, which mainly fall into three categories:
The current solutions work on simple cases, but they fall short when it comes to complex scenarios.
Before diving deep into the solutions and their drawbacks, let’s first understand what a .docx file is.
A .docx file is essentially a ZIP file with a hierarchy of XML files following the OOXML (Office Open XML) spec. Inside the ZIP, there are the following files: document.xml contains the main text, styles.xml defines reusable styles (kind of like a CSS stylesheet), numbering.xml defines list/numbering behavior, and separate XML files store headers, footers, footnotes, relationships, media, and document metadata.
These XML files are quite verbose. For example, in document.xml, even a short 4–5-sentence paragraph can turn into thousands of tokens once you add styles, metadata, formatting information, run splitting, and XML boilerplate. The text users see in Microsoft Word might be split across many XML nodes and might be persisted in a very verbose manner.