docboss.dev

Word documents read, rendered and written, from scratch in Rust.

docboss reads .docx and legacy .doc files into one document model: text with the document’s own list labels, headers, footers, footnotes, comments and text boxes, as plain text, Markdown, HTML or JSON. It lays pages out the way Word does and renders them to PNG, writes DOCX, opens password-protected files and reads remote files by byte range. One Rust core behind a CLI, a terminal explorer and Python bindings, built from the ECMA-376 and Microsoft specifications: safe Rust, no C dependencies. MIT OR Apache-2.0.

$pip install docboss$cargo install docboss-cli

abi3 wheels for CPython 3.12 and later on Linux x86_64 and macOS arm64, no Rust toolchain required.

The toolkit

Two formats, one model.

The CLI ships text, md, html, json, q, render, images, convert and create md, plus parts, hex, xml, fonts, diagnostics and tui for looking inside a file. DOCX, DOCM, DOTX and DOTM, DOC and DOT all go through the same commands; the format is detected from the bytes, never from the file name. Every read command takes --password, and most take an http(s):// URL.

Python API →Command line →Rust crates →

$ docboss info report.docx
$ docboss text report.docx --headers-footers --comments
$ docboss md report.docx --images images -o report.md
$ docboss html report.docx --standalone
$ docboss q report.docx '.metadata.title' -r
$ docboss render report.docx --page 1 -o page.png --scale 2
# the same commands on a Word 97-2003 file
$ docboss text legacy.doc
$ docboss convert legacy.doc -o legacy.docx
$ docboss create md notes.md -o notes.docx
$ docboss text locked.docx --password secret
$ docboss diagnostics report.docx
report.pydocboss
import docboss

doc = docboss.Document("report.docx")  # or "legacy.doc"
print(doc.format, doc.metadata.title)
text = doc.extract_text(headers_footers=True)
md = doc.extract_markdown()  # headings, lists, tables, notes
html = doc.extract_html(standalone=True)
for block in doc.blocks():
    print(block.kind, block.text)

png = doc.render(0, scale=2.0)  # PNG bytes
pngs = doc.render_pages()  # every page, all cores
data = doc.to_docx()  # a fresh DOCX, also from .doc

texts = docboss.extract_texts(paths, threads=8)
docx = docboss.md.to_docx("# Notes\n\n- one\n")

Everything in the file

Headers, notes, comments, tables and text boxes.

List paragraphs start with the label the document’s numbering gives them (1., a), •), computed the way Word numbers them. Table rows come out tab-separated, footnote and endnote references as [1] and [i] with the notes after the main text, and text boxes after the paragraph that holds them. Hidden text, deleted revisions and field instructions are left out; a field gives its cached result, so a table of contents reads as the file shows it. Both outputs below are from libreoffice-rich.docx, a test fixture in the docboss repository. Extracting text →

plain text
$ docboss text libreoffice-rich.docx --headers-footers --comments
Running header
Fixture Heading
Text with bold words and a note[1] here.
Commented[c1] word.
Picture:
• Item one
• Item two
A1	B1
Span
Endnote section
Ends here[i].
Page 1

[1] Footnote text.
[i] Endnote text.

[c1] Reviewer: A comment body.
GitHub-flavored Markdown
$ docboss md libreoffice-rich.docx
# Fixture Heading

Text with **bold words** and a note[^1] here.

Commented word.

Picture:

- Item one
- Item two

| A1 | B1 |
| --- | --- |
| Span |  |

## Endnote section

Ends here[^i].

[^1]: Footnote text.

[^i]: Endnote text.

Layout and rendering

Pages laid out the way Word does.

Word-like line breaking, tab stops, list labels, keep and widow rules, tables split across pages with repeated header rows, sections and columns, headers and footers with page numbers, footnotes, inline and floating images, and text that wraps around floating pictures, frames and tables. Right-to-left paragraphs go through the Unicode bidirectional algorithm, and Arabic and Hebrew are shaped with GSUB and GPOS. An anti-aliased rasterizer over docboss’s own TrueType and CFF parsers writes PNG, PPM, BMP or JPEG, with pages rendered on all cores. A missing Calibri is set in Carlito and Arial in Liberation Sans, metric-compatible substitutes, so line breaks stay close to Word’s.

0.985 median SSIM

DOCX pages against LibreOffice’s rendering of the same files, 60 files, up to 5 pages each. Page counts agree on 56.

0.966 median SSIM

DOC pages, 30 files. Page counts agree on 27. A tenth of the DOC files score 0.634 or lower.

The low scores come from fonts LibreOffice substitutes differently, tables and shapes it places differently, text in table cells that LibreOffice wraps around pictures and docboss does not, and tracked changes that LibreOffice shows with revision marks. Fidelity method → More renders →

Page 1 of a Word 95 price list rendered by docboss: a purple title, a shaded section heading and a 19-row table of wines with bold italic names and prices
docboss render word95-tables.doc --scale 2, page 1: a Word 95 file from Apache POI’s test data, read with its formatting, sections and tables.

Writing and encryption

DOC to DOCX, Markdown to DOCX, and locked files.

docboss writes WordprocessingML packages from its model, so a Word 97-2003 file converts to DOCX with its styles, lists, tables, sections, headers and footers, notes, comments, fields, bookmarks, revisions and images. Output is deterministic: the same input gives identical bytes. CommonMark and GFM map onto Word’s built-in styles, and a builder composes documents from Rust. Password-protected DOCX (Agile and Standard encryption) and DOC (RC4, RC4 CryptoAPI and XOR obfuscation) open with their password; a tampered Agile package fails its integrity check. Writing DOCX → Encrypted documents →

~/reports
$ docboss convert legacy.doc -o legacy.docx      # DOC to DOCX
$ docboss convert report.docx -o fresh.docx      # a fresh package from the model
$ docboss convert legacy.doc -o legacy.md        # or .html, .txt, .json by extension
$ docboss create md notes.md -o notes.docx --size a4 --title "Notes"
$ docboss text locked.docx --password secret
$ docboss convert locked.docx --password secret -o unlocked.docx
compose.rs
use docboss_write::{DocumentBuilder, ListKind, Para, TableBuilder};

let mut doc = DocumentBuilder::new();
doc.title("Q3 Report").heading(1, "Summary");
doc.paragraph(Para::new().text("Revenue grew ").bold("12%").text("."));
doc.items(ListKind::Bullet, &["North", "South"]);
let width = doc.text_width();
doc.block(TableBuilder::new(2, width).header(&["Region", "Growth"]).row(&["North", "14%"]));
let bytes = doc.to_bytes()?;

Remote documents

Read a document without downloading it.

Every read command takes an http(s):// URL where it takes a path. For a DOCX, docboss fetches the end of the file to find the ZIP central directory, then only the entries the reader needs; images are fetched only when a command uses them. For a DOC, it fetches the compound-file header, the FAT and directory sectors, then only the sectors of the Word streams. Requests are coalesced and cached, so no byte is fetched twice, and a server that ignores Range costs one full download. In Python, AsyncDocument does the same from asyncio. Remote documents →

17.9 KB in 3 requests

To extract the text of Apache POI’s saut_page.docx, a 2.96 MB file that is mostly images. A DOC reads its Data stream too, so it saves less.

~/reports
$ docboss text https://example.com/report.docx
$ docboss info https://example.com/legacy.doc
$ docboss md https://example.com/report.docx --images images
remote.py
import asyncio

import docboss

async def main() -> None:
    doc = await docboss.AsyncDocument.open_url("https://example.com/report.docx")
    text = await doc.extract_text()
    print(doc.format, doc.bytes_fetched, "bytes fetched in", len(doc.requests), "requests")
    full = await doc.document()   # the synchronous Document, for rendering

asyncio.run(main())

The explorer

See inside any Word file.

docboss tui opens a tree of the document: sections, paragraphs, tables and runs, headers and footers, notes, comments, styles, numbering, media, fonts, the container and the diagnostics. Beside it sit an inspector with each node’s direct and resolved properties and its style chain, the XML of the selected part, a Markdown view, a page preview drawn in half-block characters and a hex pane. The same model answers parts, hex, xml and q, a jq program over the document’s JSON.

docboss tui libreoffice-rich.docx
┌Tree 5/22──────────────────────────────────┐┌Inspector────────────────────────────────────────────────────────┐
│  Document (Docx)                          ││Paragraph                                                        │
│▾ Sections (1)                             ││  style: Normal                                                  │
│  ▾ section 1 · 11906×16838 tw · 9 blocks  ││  inlines: 5                                                     │
│    ▸ ¶ [Heading1] Fixture Heading         ││Text                                                             │
│    ▸ ¶ [Normal] Text with bold words and a││  Text with bold words and a note here.                          │
│    ▸ ¶ [Normal] Commented word.           ││Direct properties                                                │
│    ▸ ¶ [Normal] Picture:                  ││  ParagraphProperties {                                          │
│    ▸ ¶ [Normal] Item one                  ││      bidi: Some(false),                                         │
│    ▸ ¶ [Normal] Item two                  ││  }                                                              │
│    ▸ table 2×2                            ││Resolved properties                                              │
│    ▸ ¶ [Heading2] Endnote section         ││  ParagraphProperties {                                          │
│    ▸ ¶ [Normal] Ends here.                ││      widow_control: Some(false),                                │
│▸ Headers and footers (2)                  ││      bidi: Some(false),                                         │
│▸ Footnotes (1)                            ││  }                                                              │
│▸ Endnotes (1)                             │└─────────────────────────────────────────────────────────────────┘
│▸ Comments (1)                             │┌Hex file (10436 bytes)───────────────────────────────────────────┐
│▸ Styles (19)                              ││00000000  50 4b 03 04 14 00 08 08  08 00 6f 6b 3c 5d 00 00  |PK..│
│▸ Numbering (4)                            ││00000010  00 00 00 00 00 00 00 00  00 00 11 00 00 00 64 6f  |....│
│▸ Media (1)                                ││00000020  63 50 72 6f 70 73 2f 63  6f 72 65 2e 78 6d 6c 6d  |cPro│
│▸ Fonts (6)                                ││00000030  92 51 4f c3 20 14 85 df  fd 15 0d ef 2d ad 26 c6  |.QO.│
│▸ Container (17)                           ││00000040  34 2d 4b 8c d9 93 4b 4c  e6 12 5f 19 dc 75 68 0b  |4-K.│
│  Diagnostics (0)                          ││00000050  84 7b bb 6e ff 5e da 69  dd e2 de 38 9c 8f 03 1c  |.{.n│
│                                           ││00000060  a8 16 c7 ae 4d 0e 10 d0  38 5b b3 22 cb 59 02 56  |....│
└───────────────────────────────────────────┘└─────────────────────────────────────────────────────────────────┘
libreoffice-rich.docx · Docx · 0 diagnostics                          ? help  q quit
~/reports · docboss
$ docboss parts word95-tables.doc
       106  \x01CompObj
     25617  WordDocument
       440  \x05SummaryInformation
       216  \x05DocumentSummaryInformation
$ docboss hex word95-tables.doc WordDocument --length 32
00000000  dc a5 68 00 63 e0 09 04  00 00 00 00 65 00 00 00  |..h.c.......e...|
00000010  00 00 00 00 00 00 00 00  00 03 00 00 89 11 00 00  |................|
$ docboss fonts word95-tables.doc
Arial -> Arial
Courier New -> Courier New
Symbol -> Symbol
Times New Roman -> Times New Roman
Times New Roman Cyr -> Times New Roman
$ docboss diagnostics word95-tables.doc
approximated WordDocument: text decoded as code page 1251, guessed from its bytes
$ docboss q libreoffice-rich.docx --blocks -r '.[] | select(.kind == "heading") | .text'
Fixture Heading
Endnote section

Benchmarks

Measured on real documents.

Test documents from the LibreOffice, Apache POI and python-docx projects, against docx2txt, python-docx, docx2python, mammoth, pandoc, antiword and catdoc, with every text output scored against LibreOffice’s. catdoc is faster on one very large DOC and antiword and catdoc use less memory on DOC; the benchmarks page has those tables too.

3,304 files/s

DOCX text over 80 real-world test files: about 4× docx2txt and 28× mammoth.

3,130 files/s

DOC text over 80 legacy .doc files: about 19× antiword and catdoc.

0.992 word recall

Against LibreOffice’s text export on DOCX, the highest of the six engines measured; 0.985 on DOC.

140/150 damaged files

Damaged DOCX files that still give text, where the other engines manage 1 to 8. No crash and no hang on 300 damaged files.

Full benchmarks, charts and method →

Under the hood

Eighteen crates, one implementation.

The ZIP container, the XML tokenizer, the compound-file reader, the font parsers, the layout engine and the rasterizer are docboss’s own, written from ECMA-376 for DOCX and from Microsoft’s [MS-DOC], [MS-CFB], [MS-OLEPS], [MS-ODRAW] and [MS-OFFCRYPTO] for the binary format. Both readers produce the same model, so text, Markdown, layout and DOCX writing behave the same for a .docx and a .doc. Eighteen crates make up the engine, the CLI and the explorer, all under MIT OR Apache-2.0; the Python extension is one more.

Browse the Rust crates →

Reporting

“A lenient reader that says what it cost.”

Damaged files still open

A ZIP with a missing or broken central directory is read from its local headers, and a damaged deflate stream keeps what inflated before the damage. A compound file with a looping or truncated sector chain is followed as far as it goes. XML that is not well formed still yields balanced elements.

Nothing is dropped silently

Every approximated or dropped item is a diagnostic on the document: docboss diagnostics lists them, --layout adds what layout approximated, and Python and Rust expose them as a value. An empty list means the reader took the file as written.

A ledger checked by Lean

Every clause of the specifications docboss reads has a row with a status. A Lean 4 gate fails the build when an implemented row lacks a code citation and a test citation, when a required clause has no row, or when a citation names a clause the standard does not have. The ledger →

Conformance ledger, clauses per status
SpecificationImplementedIncompleteNot implementedOut of scope
ECMA-376 Part 11307822074
ECMA-376 Part 2156340
ECMA-376 Part 34051
[MS-DOC]963133211
[MS-CFB]10012
[MS-OLEPS]95260
[MS-OSHARED]0130
[MS-ODRAW]0320
[MS-WMF]1800
[MS-EMF]21120
APPNOTE1367424

Questions

Does docboss need LibreOffice, Word or a C library?
No. docboss is written from scratch in safe Rust with its own ZIP, XML, compound-file, font and raster code. The Python wheel needs only CPython 3.12 or later. LibreOffice appears only in the benchmarks, as the reference for word recall and rendering.
Which file formats does it read?
DOCX, DOCM, DOTX and DOTM through ECMA-376 WordprocessingML, transitional and strict, and DOC and DOT through [MS-DOC], including Word 97-2003 files and the formatting, sections and tables of Word 6 and 95 files. Word 2 files are read as text only. RTF is detected and refused.
Can it render pages to images?
Yes. docboss lays the document out and renders pages to PNG, PPM, BMP or JPEG, on all cores. Against LibreOffice's rendering of the same files, the median windowed SSIM is 0.985 on DOCX and 0.966 on DOC.
Can it edit a document?
No. docboss reads, lays out, renders and writes documents: it converts DOC and DOCX to a fresh DOCX, composes DOCX from Markdown or a Rust builder, and writes the model back out. It does not edit a file in place and does not run macros.
What happens with a damaged file?
The reader recovers what it can and lists every approximated or dropped item as a diagnostic. On 150 damaged DOCX files it returned text for 140, and on 150 damaged DOC files for 142, without a crash or a hang.
What does it not do yet?
Indic, Thai and other scripts that reorder glyphs are not shaped, right-to-left and vertical sections lay out left to right, there is no column balancing, and picture and pattern fills are not drawn. Embedded fonts and the theme part are not written, and encrypted output is not supported. The limitations page lists the rest.