Word documents read, rendered and written, from scratch in Rust.
docboss reads .docx and legacy .doc files into one document model: text with the document’s own list labels, headers, footers, footnotes, comments and text boxes, as plain text, Markdown, HTML or JSON. It lays pages out the way Word does and renders them to PNG, writes DOCX, opens password-protected files and reads remote files by byte range. One Rust core behind a CLI, a terminal explorer and Python bindings, built from the ECMA-376 and Microsoft specifications: safe Rust, no C dependencies. MIT OR Apache-2.0.
abi3 wheels for CPython 3.12 and later on Linux x86_64 and macOS arm64, no Rust toolchain required.
The toolkit
Two formats, one model.
The CLI ships text, md, html, json, q, render, images, convert and create md, plus parts, hex, xml, fonts, diagnostics and tui for looking inside a file. DOCX, DOCM, DOTX and DOTM, DOC and DOT all go through the same commands; the format is detected from the bytes, never from the file name. Every read command takes --password, and most take an http(s):// URL.
$ docboss info report.docx $ docboss text report.docx --headers-footers --comments $ docboss md report.docx --images images -o report.md $ docboss html report.docx --standalone $ docboss q report.docx '.metadata.title' -r $ docboss render report.docx --page 1 -o page.png --scale 2 # the same commands on a Word 97-2003 file $ docboss text legacy.doc $ docboss convert legacy.doc -o legacy.docx $ docboss create md notes.md -o notes.docx $ docboss text locked.docx --password secret $ docboss diagnostics report.docx
import docboss doc = docboss.Document("report.docx") # or "legacy.doc" print(doc.format, doc.metadata.title) text = doc.extract_text(headers_footers=True) md = doc.extract_markdown() # headings, lists, tables, notes html = doc.extract_html(standalone=True) for block in doc.blocks(): print(block.kind, block.text) png = doc.render(0, scale=2.0) # PNG bytes pngs = doc.render_pages() # every page, all cores data = doc.to_docx() # a fresh DOCX, also from .doc texts = docboss.extract_texts(paths, threads=8) docx = docboss.md.to_docx("# Notes\n\n- one\n")
Everything in the file
Headers, notes, comments, tables and text boxes.
List paragraphs start with the label the document’s numbering gives them (1., a), •), computed the way Word numbers them. Table rows come out tab-separated, footnote and endnote references as [1] and [i] with the notes after the main text, and text boxes after the paragraph that holds them. Hidden text, deleted revisions and field instructions are left out; a field gives its cached result, so a table of contents reads as the file shows it. Both outputs below are from libreoffice-rich.docx, a test fixture in the docboss repository. Extracting text →
$ docboss text libreoffice-rich.docx --headers-footers --comments
Running header
Fixture Heading
Text with bold words and a note[1] here.
Commented[c1] word.
Picture:
• Item one
• Item two
A1 B1
Span
Endnote section
Ends here[i].
Page 1
[1] Footnote text.
[i] Endnote text.
[c1] Reviewer: A comment body.$ docboss md libreoffice-rich.docx
# Fixture Heading
Text with **bold words** and a note[^1] here.
Commented word.
Picture:
- Item one
- Item two
| A1 | B1 |
| --- | --- |
| Span | |
## Endnote section
Ends here[^i].
[^1]: Footnote text.
[^i]: Endnote text.Layout and rendering
Pages laid out the way Word does.
Word-like line breaking, tab stops, list labels, keep and widow rules, tables split across pages with repeated header rows, sections and columns, headers and footers with page numbers, footnotes, inline and floating images, and text that wraps around floating pictures, frames and tables. Right-to-left paragraphs go through the Unicode bidirectional algorithm, and Arabic and Hebrew are shaped with GSUB and GPOS. An anti-aliased rasterizer over docboss’s own TrueType and CFF parsers writes PNG, PPM, BMP or JPEG, with pages rendered on all cores. A missing Calibri is set in Carlito and Arial in Liberation Sans, metric-compatible substitutes, so line breaks stay close to Word’s.
DOCX pages against LibreOffice’s rendering of the same files, 60 files, up to 5 pages each. Page counts agree on 56.
DOC pages, 30 files. Page counts agree on 27. A tenth of the DOC files score 0.634 or lower.
The low scores come from fonts LibreOffice substitutes differently, tables and shapes it places differently, text in table cells that LibreOffice wraps around pictures and docboss does not, and tracked changes that LibreOffice shows with revision marks. Fidelity method → More renders →

docboss render word95-tables.doc --scale 2, page 1: a Word 95 file from Apache POI’s test data, read with its formatting, sections and tables.Writing and encryption
DOC to DOCX, Markdown to DOCX, and locked files.
docboss writes WordprocessingML packages from its model, so a Word 97-2003 file converts to DOCX with its styles, lists, tables, sections, headers and footers, notes, comments, fields, bookmarks, revisions and images. Output is deterministic: the same input gives identical bytes. CommonMark and GFM map onto Word’s built-in styles, and a builder composes documents from Rust. Password-protected DOCX (Agile and Standard encryption) and DOC (RC4, RC4 CryptoAPI and XOR obfuscation) open with their password; a tampered Agile package fails its integrity check. Writing DOCX → Encrypted documents →
$ docboss convert legacy.doc -o legacy.docx # DOC to DOCX
$ docboss convert report.docx -o fresh.docx # a fresh package from the model
$ docboss convert legacy.doc -o legacy.md # or .html, .txt, .json by extension
$ docboss create md notes.md -o notes.docx --size a4 --title "Notes"
$ docboss text locked.docx --password secret
$ docboss convert locked.docx --password secret -o unlocked.docxuse docboss_write::{DocumentBuilder, ListKind, Para, TableBuilder};
let mut doc = DocumentBuilder::new();
doc.title("Q3 Report").heading(1, "Summary");
doc.paragraph(Para::new().text("Revenue grew ").bold("12%").text("."));
doc.items(ListKind::Bullet, &["North", "South"]);
let width = doc.text_width();
doc.block(TableBuilder::new(2, width).header(&["Region", "Growth"]).row(&["North", "14%"]));
let bytes = doc.to_bytes()?;Remote documents
Read a document without downloading it.
Every read command takes an http(s):// URL where it takes a path. For a DOCX, docboss fetches the end of the file to find the ZIP central directory, then only the entries the reader needs; images are fetched only when a command uses them. For a DOC, it fetches the compound-file header, the FAT and directory sectors, then only the sectors of the Word streams. Requests are coalesced and cached, so no byte is fetched twice, and a server that ignores Range costs one full download. In Python, AsyncDocument does the same from asyncio. Remote documents →
To extract the text of Apache POI’s saut_page.docx, a 2.96 MB file that is mostly images. A DOC reads its Data stream too, so it saves less.
$ docboss text https://example.com/report.docx
$ docboss info https://example.com/legacy.doc
$ docboss md https://example.com/report.docx --images imagesimport asyncio
import docboss
async def main() -> None:
doc = await docboss.AsyncDocument.open_url("https://example.com/report.docx")
text = await doc.extract_text()
print(doc.format, doc.bytes_fetched, "bytes fetched in", len(doc.requests), "requests")
full = await doc.document() # the synchronous Document, for rendering
asyncio.run(main())The explorer
See inside any Word file.
docboss tui opens a tree of the document: sections, paragraphs, tables and runs, headers and footers, notes, comments, styles, numbering, media, fonts, the container and the diagnostics. Beside it sit an inspector with each node’s direct and resolved properties and its style chain, the XML of the selected part, a Markdown view, a page preview drawn in half-block characters and a hex pane. The same model answers parts, hex, xml and q, a jq program over the document’s JSON.
┌Tree 5/22──────────────────────────────────┐┌Inspector────────────────────────────────────────────────────────┐
│ Document (Docx) ││Paragraph │
│▾ Sections (1) ││ style: Normal │
│ ▾ section 1 · 11906×16838 tw · 9 blocks ││ inlines: 5 │
│ ▸ ¶ [Heading1] Fixture Heading ││Text │
│ ▸ ¶ [Normal] Text with bold words and a││ Text with bold words and a note here. │
│ ▸ ¶ [Normal] Commented word. ││Direct properties │
│ ▸ ¶ [Normal] Picture: ││ ParagraphProperties { │
│ ▸ ¶ [Normal] Item one ││ bidi: Some(false), │
│ ▸ ¶ [Normal] Item two ││ } │
│ ▸ table 2×2 ││Resolved properties │
│ ▸ ¶ [Heading2] Endnote section ││ ParagraphProperties { │
│ ▸ ¶ [Normal] Ends here. ││ widow_control: Some(false), │
│▸ Headers and footers (2) ││ bidi: Some(false), │
│▸ Footnotes (1) ││ } │
│▸ Endnotes (1) │└─────────────────────────────────────────────────────────────────┘
│▸ Comments (1) │┌Hex file (10436 bytes)───────────────────────────────────────────┐
│▸ Styles (19) ││00000000 50 4b 03 04 14 00 08 08 08 00 6f 6b 3c 5d 00 00 |PK..│
│▸ Numbering (4) ││00000010 00 00 00 00 00 00 00 00 00 00 11 00 00 00 64 6f |....│
│▸ Media (1) ││00000020 63 50 72 6f 70 73 2f 63 6f 72 65 2e 78 6d 6c 6d |cPro│
│▸ Fonts (6) ││00000030 92 51 4f c3 20 14 85 df fd 15 0d ef 2d ad 26 c6 |.QO.│
│▸ Container (17) ││00000040 34 2d 4b 8c d9 93 4b 4c e6 12 5f 19 dc 75 68 0b |4-K.│
│ Diagnostics (0) ││00000050 84 7b bb 6e ff 5e da 69 dd e2 de 38 9c 8f 03 1c |.{.n│
│ ││00000060 a8 16 c7 ae 4d 0e 10 d0 38 5b b3 22 cb 59 02 56 |....│
└───────────────────────────────────────────┘└─────────────────────────────────────────────────────────────────┘
libreoffice-rich.docx · Docx · 0 diagnostics ? help q quit$ docboss parts word95-tables.doc 106 \x01CompObj 25617 WordDocument 440 \x05SummaryInformation 216 \x05DocumentSummaryInformation $ docboss hex word95-tables.doc WordDocument --length 32 00000000 dc a5 68 00 63 e0 09 04 00 00 00 00 65 00 00 00 |..h.c.......e...| 00000010 00 00 00 00 00 00 00 00 00 03 00 00 89 11 00 00 |................| $ docboss fonts word95-tables.doc Arial -> Arial Courier New -> Courier New Symbol -> Symbol Times New Roman -> Times New Roman Times New Roman Cyr -> Times New Roman $ docboss diagnostics word95-tables.doc approximated WordDocument: text decoded as code page 1251, guessed from its bytes $ docboss q libreoffice-rich.docx --blocks -r '.[] | select(.kind == "heading") | .text' Fixture Heading Endnote section
Benchmarks
Measured on real documents.
Test documents from the LibreOffice, Apache POI and python-docx projects, against docx2txt, python-docx, docx2python, mammoth, pandoc, antiword and catdoc, with every text output scored against LibreOffice’s. catdoc is faster on one very large DOC and antiword and catdoc use less memory on DOC; the benchmarks page has those tables too.
DOCX text over 80 real-world test files: about 4× docx2txt and 28× mammoth.
DOC text over 80 legacy .doc files: about 19× antiword and catdoc.
Against LibreOffice’s text export on DOCX, the highest of the six engines measured; 0.985 on DOC.
Damaged DOCX files that still give text, where the other engines manage 1 to 8. No crash and no hang on 300 damaged files.
Under the hood
Eighteen crates, one implementation.
The ZIP container, the XML tokenizer, the compound-file reader, the font parsers, the layout engine and the rasterizer are docboss’s own, written from ECMA-376 for DOCX and from Microsoft’s [MS-DOC], [MS-CFB], [MS-OLEPS], [MS-ODRAW] and [MS-OFFCRYPTO] for the binary format. Both readers produce the same model, so text, Markdown, layout and DOCX writing behave the same for a .docx and a .doc. Eighteen crates make up the engine, the CLI and the explorer, all under MIT OR Apache-2.0; the Python extension is one more.
docboss-core
Format detection and one open/read over both readers; re-exports the model.
docboss-model
The document model both readers produce: sections, runs, tables, styles, numbering, notes, comments, media, diagnostics.
docboss-output
Text, Markdown, HTML, JSON and the per-block view.
docboss-layout
Pages from the model: line breaking, tabs, lists, tables, sections, headers, footers, footnotes, images.
docboss-render
Anti-aliased rasterizer and the PNG, PPM, BMP and JPEG encoders.
docboss-write
DOCX from the model, a builder API, Markdown to DOCX.
docboss-aio
Async reads over files or HTTP, fetching only the byte ranges needed.
docboss-docx
OPC packages and WordprocessingML, parts parsed on separate threads.
docboss-doc
Word binary documents, Word 6 and 95 formatting, RC4 and RC4 CryptoAPI.
docboss-zip
ZIP reader and deterministic writer, ZIP64, recovery from local headers.
docboss-xml
Zero-copy pull XML tokenizer with namespace resolution.
docboss-cfb
Compound files and OLE property sets.
docboss-crypt
Agile and Standard encrypted DOCX, XOR-obfuscated DOC.
docboss-font
TrueType, OpenType CFF, GSUB and GPOS; font discovery and substitution.
docboss-metafile
WMF and EMF pictures played into paths, text and bitmaps; DIB decoding.
docboss-mtef
Equation Editor 3 equations read into the math model.
docboss-cli
The docboss binary.
docboss-tui
The terminal explorer.
docboss-py
The PyO3 extension, shipped as the docboss wheel.
Reporting
“A lenient reader that says what it cost.”
Damaged files still open
A ZIP with a missing or broken central directory is read from its local headers, and a damaged deflate stream keeps what inflated before the damage. A compound file with a looping or truncated sector chain is followed as far as it goes. XML that is not well formed still yields balanced elements.
Nothing is dropped silently
Every approximated or dropped item is a diagnostic on the document: docboss diagnostics lists them, --layout adds what layout approximated, and Python and Rust expose them as a value. An empty list means the reader took the file as written.
A ledger checked by Lean
Every clause of the specifications docboss reads has a row with a status. A Lean 4 gate fails the build when an implemented row lacks a code citation and a test citation, when a required clause has no row, or when a citation names a clause the standard does not have. The ledger →
| Specification | Implemented | Incomplete | Not implemented | Out of scope |
|---|---|---|---|---|
| ECMA-376 Part 1 | 130 | 78 | 220 | 74 |
| ECMA-376 Part 2 | 15 | 6 | 3 | 40 |
| ECMA-376 Part 3 | 4 | 0 | 5 | 1 |
| [MS-DOC] | 96 | 31 | 332 | 11 |
| [MS-CFB] | 10 | 0 | 1 | 2 |
| [MS-OLEPS] | 9 | 5 | 26 | 0 |
| [MS-OSHARED] | 0 | 1 | 3 | 0 |
| [MS-ODRAW] | 0 | 3 | 2 | 0 |
| [MS-WMF] | 1 | 8 | 0 | 0 |
| [MS-EMF] | 2 | 11 | 2 | 0 |
| APPNOTE | 13 | 6 | 74 | 24 |
Compare
Switching from another library?
Side-by-side numbers from the same benchmark run, what each library covers, and migration snippets for the reading path.
docboss vs python-docx →
Text 4.6 times faster with headers, footers and notes included; python-docx keeps editing documents.
docboss vs mammoth →
DOCX to Markdown and HTML about 30 times faster, and DOC files through the same call.
docboss vs antiword and catdoc →
Legacy .doc text about 19 times faster; the C tools win on memory and on one very large file.
DOCX to Markdown →
Headings from styles, nested lists, GFM tables, footnotes, links and images.
Questions
- Does docboss need LibreOffice, Word or a C library?
- No. docboss is written from scratch in safe Rust with its own ZIP, XML, compound-file, font and raster code. The Python wheel needs only CPython 3.12 or later. LibreOffice appears only in the benchmarks, as the reference for word recall and rendering.
- Which file formats does it read?
- DOCX, DOCM, DOTX and DOTM through ECMA-376 WordprocessingML, transitional and strict, and DOC and DOT through [MS-DOC], including Word 97-2003 files and the formatting, sections and tables of Word 6 and 95 files. Word 2 files are read as text only. RTF is detected and refused.
- Can it render pages to images?
- Yes. docboss lays the document out and renders pages to PNG, PPM, BMP or JPEG, on all cores. Against LibreOffice's rendering of the same files, the median windowed SSIM is 0.985 on DOCX and 0.966 on DOC.
- Can it edit a document?
- No. docboss reads, lays out, renders and writes documents: it converts DOC and DOCX to a fresh DOCX, composes DOCX from Markdown or a Rust builder, and writes the model back out. It does not edit a file in place and does not run macros.
- What happens with a damaged file?
- The reader recovers what it can and lists every approximated or dropped item as a diagnostic. On 150 damaged DOCX files it returned text for 140, and on 150 damaged DOC files for 142, without a crash or a hang.
- What does it not do yet?
- Indic, Thai and other scripts that reorder glyphs are not shaped, right-to-left and vertical sections lay out left to right, there is no column balancing, and picture and pattern fills are not drawn. Embedded fonts and the theme part are not written, and encrypted output is not supported. The limitations page lists the rest.