Files
magnus919_agent-skills/anydoc/README.md
T
Magnus HedemarkandGitHub bb57268a68 feat(anydoc): support explicit hosted OCR (#420)
* feat(anydoc): support explicit hosted OCR

Closes #419

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* test(anydoc): update release contract expectations

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* test(anydoc): align hosted OCR hint contract

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* chore: refresh generated marketplace

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* chore: refresh generated llms catalog

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

---------

Signed-off-by: Magnus Hedemark <magnus919@pm.me>
2026-08-28 15:33:48 -04:00

4.7 KiB

anydoc — office documents to GitHub-Flavored Markdown

Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files into clean, LLM-friendly GitHub-Flavored Markdown. Local conversion stays on your machine; an explicitly authorized hosted OCR mode handles scanned PDFs through Firecrawl Parse when whole-document upload is acceptable.

Why Install This Skill

Office documents are opaque to agents. A .docx or .pptx is a binary zip; a .xls is an OLE container; a PDF can be anything. Reading them directly means parsing formats, handling encodings, and reconstructing structure by hand — exactly the work anydoc automates. This skill gives your agent a single, verified command that converts all 8 format families (21 extensions) into GitHub-Flavored Markdown with headings, GFM tables, slide structure, and footnotes preserved, plus the knowledge of exactly where fidelity is lost (Excel number formats, legacy PowerPoint tables, PDF tables).

The skill wraps the pinned @firecrawl/anydoc v0.2.4 CLI with a small helper script that adds input checks, friendly error hints for the known failure classes (scanned PDFs, encrypted files, malformed archives), batch conversion, dry-run planning, JSON output, and an explicit --allow-hosted-upload acknowledgement for hosted OCR.

What You Get

Directory / file What it provides
SKILL.md + README.md The skill index (trigger, command map, verification steps) and this human-facing guide
scripts/ anydoc — an executable Python 3 wrapper with convert (single file or stdin, -o output), batch (many files, per-file status, summary), and info (tool + pinned CLI version), plus global --json and --dry-run
references/ Five focused guides: formats.md (what GFM each format produces, with fidelity caveats), cli-reference.md (verbatim --help, every flag, stdout/stderr conventions), errors.md (exit codes and the exact error messages), workflows.md (recipes: single conversion, batch, vault ingestion, piping, output verification), sources.md (upstream URLs, fixture provenance, verification procedure)
tests/ Unit tests for the wrapper (argparse, pre-validation, hints, dry-run, JSON, batch) — runnable offline
evals/ An eval manifest with fixture-backed cases covering docx→headings, xlsx→tables, pptx→slide structure, csv→table, legacy .doc, ODS preserved values, ODT, and the image-only-PDF OCR failure
fixtures/ 24 tiny sample documents (all < 5 MB): valid samples for every family plus error cases (image-only PDF, encrypted ODT, empty DOCX, unsupported extension) — used by the tests, evals, and recipes

Quick Start

You need Node.js 20+ and npx (no other install — the CLI and its native binary are fetched on first use):

cd anydoc
npx -y @firecrawl/anydoc@0.2.4 fixtures/fixture-handmade-outline.docx

This converts the sample Word document and prints GitHub-Flavored Markdown to stdout (note the #/##/### heading lines). To write to a file instead:

npx -y @firecrawl/anydoc@0.2.4 fixtures/fixture-handmade-outline.docx -o outline.md

Or use the wrapper for the same job:

python3 scripts/anydoc convert fixtures/fixture-handmade-outline.docx -o outline.md

Triggers

Load this skill when the task involves any of these:

  • "Convert this Word/Excel/PowerPoint/PDF/EPUB/CSV file to markdown"
  • "Extract the headings, tables, or slide content from this document"
  • "Summarize this report / spreadsheet / deck"
  • "Turn this CSV into a markdown table"
  • "Read this document into markdown for a knowledge base or vault"
  • "Convert this PDF to markdown" — but only for text-based PDFs; scanned or image-only PDFs fail (anydoc does not OCR)
  • "OCR this scanned PDF" — use local OCR by default, or explicitly authorize --ocr hosted --allow-hosted-upload when sending the whole document to Firecrawl Parse is acceptable

Do not load this skill for document generation or editing ("create a docx report", "build a PDF proposal", "validate this document") — that is the documents skill's job — or for EPUB authoring (epub skill).

Requirements

  • Node.js >= 20 and npx (the CLI is distributed via npm; the native binary ships as a platform-specific npm optionalDependency, so there is no manual install or compilation).
  • Network once — the first npx run downloads the package and binary; later runs use the npm cache. For permanent or fully offline use, run npm install -g @firecrawl/anydoc once.
  • Python 3 (standard library only) if you use the scripts/anydoc wrapper.
  • Local mode needs no API key or service. Hosted OCR uses Firecrawl Parse and may use FIRECRAWL_API_KEY; it sends the whole OCR-required PDF and has no page selection.