* feat(anydoc): support explicit hosted OCR Closes #419 Signed-off-by: Magnus Hedemark <magnus919@pm.me> * test(anydoc): update release contract expectations Signed-off-by: Magnus Hedemark <magnus919@pm.me> * test(anydoc): align hosted OCR hint contract Signed-off-by: Magnus Hedemark <magnus919@pm.me> * chore: refresh generated marketplace Signed-off-by: Magnus Hedemark <magnus919@pm.me> * chore: refresh generated llms catalog Signed-off-by: Magnus Hedemark <magnus919@pm.me> --------- Signed-off-by: Magnus Hedemark <magnus919@pm.me>
anydoc — office documents to GitHub-Flavored Markdown
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files into clean, LLM-friendly GitHub-Flavored Markdown. Local conversion stays on your machine; an explicitly authorized hosted OCR mode handles scanned PDFs through Firecrawl Parse when whole-document upload is acceptable.
Why Install This Skill
Office documents are opaque to agents. A .docx or .pptx is a binary zip; a .xls is an OLE container; a PDF can be anything. Reading them directly means parsing formats, handling encodings, and reconstructing structure by hand — exactly the work anydoc automates. This skill gives your agent a single, verified command that converts all 8 format families (21 extensions) into GitHub-Flavored Markdown with headings, GFM tables, slide structure, and footnotes preserved, plus the knowledge of exactly where fidelity is lost (Excel number formats, legacy PowerPoint tables, PDF tables).
The skill wraps the pinned @firecrawl/anydoc v0.2.4 CLI with a small helper script that adds input checks, friendly error hints for the known failure classes (scanned PDFs, encrypted files, malformed archives), batch conversion, dry-run planning, JSON output, and an explicit --allow-hosted-upload acknowledgement for hosted OCR.
What You Get
| Directory / file | What it provides |
|---|---|
SKILL.md + README.md |
The skill index (trigger, command map, verification steps) and this human-facing guide |
scripts/ |
anydoc — an executable Python 3 wrapper with convert (single file or stdin, -o output), batch (many files, per-file status, summary), and info (tool + pinned CLI version), plus global --json and --dry-run |
references/ |
Five focused guides: formats.md (what GFM each format produces, with fidelity caveats), cli-reference.md (verbatim --help, every flag, stdout/stderr conventions), errors.md (exit codes and the exact error messages), workflows.md (recipes: single conversion, batch, vault ingestion, piping, output verification), sources.md (upstream URLs, fixture provenance, verification procedure) |
tests/ |
Unit tests for the wrapper (argparse, pre-validation, hints, dry-run, JSON, batch) — runnable offline |
evals/ |
An eval manifest with fixture-backed cases covering docx→headings, xlsx→tables, pptx→slide structure, csv→table, legacy .doc, ODS preserved values, ODT, and the image-only-PDF OCR failure |
fixtures/ |
24 tiny sample documents (all < 5 MB): valid samples for every family plus error cases (image-only PDF, encrypted ODT, empty DOCX, unsupported extension) — used by the tests, evals, and recipes |
Quick Start
You need Node.js 20+ and npx (no other install — the CLI and its native binary are fetched on first use):
cd anydoc
npx -y @firecrawl/anydoc@0.2.4 fixtures/fixture-handmade-outline.docx
This converts the sample Word document and prints GitHub-Flavored Markdown to stdout (note the #/##/### heading lines). To write to a file instead:
npx -y @firecrawl/anydoc@0.2.4 fixtures/fixture-handmade-outline.docx -o outline.md
Or use the wrapper for the same job:
python3 scripts/anydoc convert fixtures/fixture-handmade-outline.docx -o outline.md
Triggers
Load this skill when the task involves any of these:
- "Convert this Word/Excel/PowerPoint/PDF/EPUB/CSV file to markdown"
- "Extract the headings, tables, or slide content from this document"
- "Summarize this report / spreadsheet / deck"
- "Turn this CSV into a markdown table"
- "Read this document into markdown for a knowledge base or vault"
- "Convert this PDF to markdown" — but only for text-based PDFs; scanned or image-only PDFs fail (anydoc does not OCR)
- "OCR this scanned PDF" — use local OCR by default, or explicitly authorize
--ocr hosted --allow-hosted-uploadwhen sending the whole document to Firecrawl Parse is acceptable
Do not load this skill for document generation or editing ("create a docx report", "build a PDF proposal", "validate this document") — that is the documents skill's job — or for EPUB authoring (epub skill).
Requirements
- Node.js >= 20 and
npx(the CLI is distributed via npm; the native binary ships as a platform-specific npmoptionalDependency, so there is no manual install or compilation). - Network once — the first
npxrun downloads the package and binary; later runs use the npm cache. For permanent or fully offline use, runnpm install -g @firecrawl/anydoconce. - Python 3 (standard library only) if you use the
scripts/anydocwrapper. - Local mode needs no API key or service. Hosted OCR uses Firecrawl Parse and may use
FIRECRAWL_API_KEY; it sends the whole OCR-required PDF and has no page selection.