feat(skill): add anydoc — office documents to GitHub-Flavored Markdown (#295)

* feat(skill): add anydoc core content and references

Add the anydoc skill content tree: SKILL.md (progressive-disclosure index
with frontmatter per ALLOWED_FIELDS), human-facing README, the five reference
files (formats, cli-reference, errors, workflows, sources), 24 committed
fixtures (valid + error cases), and a fixture-grounded eval manifest with 8
cases. Every documented behavior, exit code, and error message was verified
against the real pinned CLI (npx -y @firecrawl/anydoc@0.1.6); verbatim --help
and error transcripts are reproduced character-for-character.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* feat(skill): add anydoc wrapper script and unit tests

Implements scripts/anydoc, a stdlib-only Python wrapper around the pinned
@firecrawl/anydoc@0.1.6 CLI: convert/batch/info subcommands, global
--json/--dry-run, input and output pre-validation, friendly hints for the
no-OCR/encrypted/malformed/unsupported error classes, Node >= 20 and npx
availability checks, deterministic batch output naming with documented
duplicate/collision behavior, and exit codes 0/1/2. Adds offline unittest
suite (46 tests, real-CLI tests skip when npx is unavailable) and keeps the
wrapper contract documented in cli-reference.md and errors.md.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* feat(skill): ratchet anydoc evals to 14 grounded cases

Verify the pre-authored 8-case manifest and extend it with six
high-signal cases (PDF lower-fidelity pipeline, legacy .ppt table
flattening, ODP same-serializer, RTF, EPUB, CSV header promotion),
each grounded in real pinned-CLI runs against the committed fixtures.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* feat(skill): integrate anydoc into repo catalog and artifacts

Add the sorted anydoc catalog entry to README.md (between agent-skills
and api-design-and-evolution), regenerate the tracked catalog artifacts
(.claude-plugin/marketplace.json, .codex-plugin/plugin.json,
.agents/plugins/marketplace.json, llms.txt) with the ruby generators,
and add a routing note to documents/SKILL.md pointing office-document
to-markdown conversion at the anydoc skill.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* fix(skill): polish anydoc wrapper timeout, JSON shape, and docs

- run_cli raises CliTimeoutError on the 120s timeout; convert/batch with
  --json now emit one parseable JSON error envelope (error_class "timeout")
  on stdout before exiting, so --json always yields exactly one JSON doc
- batch JSON failure entries (pre-validation and CLI) now carry error_class
  ("io" for missing/dir inputs, mapped classes for CLI failures), so all
  batch failure entries share the same shape
- build_cli_command places -o/-f before the -- separator for dash-leading
  filenames, so `convert -f csv -- -weird` converts instead of misparsing
  ("unexpected second input"); absolute-path inputs unchanged
- workflows.md vault-ingestion recipe globs notes/* instead of docs/* and
  warns to run from a temp/vault dir, never touching repo-root docs/
- unit tests: +6 (timeout envelope x4, batch error_class shape,
  dash-leading filename); suite grows 46 -> 52

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

---------

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
This commit is contained in:
Magnus Hedemark
2026-08-06 20:11:19 -04:00
committed by GitHub
co-authored by factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
parent cff17c5974
commit f37dc73829
39 changed files with 3204 additions and 0 deletions
+9
View File
@@ -50,6 +50,15 @@
"strict": false,
"description": "Use this skill when creating, reviewing, or editing Agent Skills-format skills, or when implementing skill discovery and loading in an agent client. It covers directory structure, SKILL.md metadata, progressive disclosure, evals, and repository conventions. Do not use this skill for general software work that does not involve the Agent Skills format or lifecycle."
},
{
"name": "anydoc",
"source": "./",
"skills": [
"./anydoc"
],
"strict": false,
"description": "Convert Word (.doc/.docx/.docm), PowerPoint (.ppt/.pps/.pot/.pptx/.pptm/.ppsx/.ppsm), Excel (.xls/.xlsx/.xlsm/.xlsb), OpenDocument (.odt/.ods/.odp), RTF, EPUB, CSV, and PDF documents to clean GitHub-Flavored Markdown locally with the Any Doc CLI (npx -y @firecrawl/anydoc@0.1.6): headings, GFM tables, slide structure, and footnotes in one pass. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly. Do not use for generating, editing, or validating documents (use documents), for ebook packaging (use epub), or for OCR of scanned or image-only PDFs (anydoc does not OCR; route to OCR tooling)."
},
{
"name": "api-design-and-evolution",
"source": "./",
+1
View File
@@ -22,6 +22,7 @@
"./agent-council",
"./agent-evals-and-observability",
"./agent-skills",
"./anydoc",
"./api-design-and-evolution",
"./artifact-pyramids",
"./autogen",
+4
View File
@@ -28,6 +28,10 @@ Operate an evaluated agent with tools and authority in production through a runt
Reference for the Agent Skills open format itself — directory structure, frontmatter schema, naming conventions, and progressive disclosure model. Use this meta-skill when creating or reviewing any other skill in this repository.
### [anydoc](anydoc/SKILL.md)
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and text-based PDF documents to GitHub-Flavored Markdown locally with the Any Doc CLI. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly. Do not use for generating, editing, or validating documents, for ebook packaging, or for OCR of scanned or image-only PDFs.
### [api-design-and-evolution](api-design-and-evolution/SKILL.md)
Design, document, review, and evolve consumer-facing HTTP, GraphQL, RPC, event, webhook, and streaming interfaces. Covers consumer jobs, domain semantics, contracts, failure behavior, compatibility, deprecation, migration, and deployed-boundary verification.
+61
View File
@@ -0,0 +1,61 @@
# anydoc — office documents to GitHub-Flavored Markdown
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files into clean, LLM-friendly GitHub-Flavored Markdown — entirely on your own machine, with no API keys and no file uploads. One command turns a report, spreadsheet, or slide deck into markdown you (or an agent) can read, summarize, quote, and feed into a knowledge base.
## Why Install This Skill
Office documents are opaque to agents. A `.docx` or `.pptx` is a binary zip; a `.xls` is an OLE container; a PDF can be anything. Reading them directly means parsing formats, handling encodings, and reconstructing structure by hand — exactly the work anydoc automates. This skill gives your agent a single, verified command that converts all 8 format families (21 extensions) into GitHub-Flavored Markdown with headings, GFM tables, slide structure, and footnotes preserved, plus the knowledge of exactly where fidelity is lost (Excel number formats, legacy PowerPoint tables, PDF tables).
The skill wraps the pinned `@firecrawl/anydoc` v0.1.6 CLI with a small helper script that adds input checks, friendly error hints for the known failure classes (scanned PDFs, encrypted files, malformed archives), batch conversion, dry-run planning, and JSON output — so an agent gets predictable exit codes and messages instead of guessing. It also documents the exact error vocabulary of the real CLI, so failures like "PDF has no extractable text ... OCR is required" are recognized and routed correctly (to OCR tooling) rather than retried blindly.
## What You Get
| Directory / file | What it provides |
| --- | --- |
| `SKILL.md` + `README.md` | The skill index (trigger, command map, verification steps) and this human-facing guide |
| `scripts/` | `anydoc` — an executable Python 3 wrapper with `convert` (single file or stdin, `-o` output), `batch` (many files, per-file status, summary), and `info` (tool + pinned CLI version), plus global `--json` and `--dry-run` |
| `references/` | Five focused guides: `formats.md` (what GFM each format produces, with fidelity caveats), `cli-reference.md` (verbatim `--help`, every flag, stdout/stderr conventions), `errors.md` (exit codes and the exact error messages), `workflows.md` (recipes: single conversion, batch, vault ingestion, piping, output verification), `sources.md` (upstream URLs, fixture provenance, verification procedure) |
| `tests/` | Unit tests for the wrapper (argparse, pre-validation, hints, dry-run, JSON, batch) — runnable offline |
| `evals/` | An eval manifest with fixture-backed cases covering docx→headings, xlsx→tables, pptx→slide structure, csv→table, legacy `.doc`, ODS preserved values, ODT, and the image-only-PDF OCR failure |
| `fixtures/` | 24 tiny sample documents (all < 5 MB): valid samples for every family plus error cases (image-only PDF, encrypted ODT, empty DOCX, unsupported extension) — used by the tests, evals, and recipes |
## Quick Start
You need Node.js 20+ and `npx` (no other install — the CLI and its native binary are fetched on first use):
```bash
cd anydoc
npx -y @firecrawl/anydoc@0.1.6 fixtures/fixture-handmade-outline.docx
```
This converts the sample Word document and prints GitHub-Flavored Markdown to stdout (note the `#`/`##`/`###` heading lines). To write to a file instead:
```bash
npx -y @firecrawl/anydoc@0.1.6 fixtures/fixture-handmade-outline.docx -o outline.md
```
Or use the wrapper for the same job:
```bash
python3 scripts/anydoc convert fixtures/fixture-handmade-outline.docx -o outline.md
```
## Triggers
Load this skill when the task involves any of these:
- "Convert this Word/Excel/PowerPoint/PDF/EPUB/CSV file to markdown"
- "Extract the headings, tables, or slide content from this document"
- "Summarize this report / spreadsheet / deck"
- "Turn this CSV into a markdown table"
- "Read this document into markdown for a knowledge base or vault"
- "Convert this PDF to markdown" — but only for text-based PDFs; scanned or image-only PDFs fail (anydoc does not OCR)
Do **not** load this skill for document generation or editing ("create a docx report", "build a PDF proposal", "validate this document") — that is the `documents` skill's job — or for EPUB authoring (`epub` skill).
## Requirements
- **Node.js >= 20** and `npx` (the CLI is distributed via npm; the native binary ships as a platform-specific npm `optionalDependency`, so there is no manual install or compilation).
- **Network once** — the first `npx` run downloads the package and binary; later runs use the npm cache. For permanent or fully offline use, run `npm install -g @firecrawl/anydoc` once.
- **Python 3** (standard library only) if you use the `scripts/anydoc` wrapper.
- **No API keys, no services** — conversion happens locally; files never leave your machine.
+185
View File
@@ -0,0 +1,185 @@
---
name: anydoc
description: >-
Convert Word (.doc/.docx/.docm), PowerPoint (.ppt/.pps/.pot/.pptx/.pptm/.ppsx/.ppsm),
Excel (.xls/.xlsx/.xlsm/.xlsb), OpenDocument (.odt/.ods/.odp), RTF, EPUB, CSV, and
PDF documents to clean GitHub-Flavored Markdown locally with the Any Doc CLI
(npx -y @firecrawl/anydoc@0.1.6): headings, GFM tables, slide structure, and
footnotes in one pass. Use when a task needs the contents of an office document,
spreadsheet, presentation, ebook, or PDF you cannot read directly. Do not use for
generating, editing, or validating documents (use documents), for ebook packaging
(use epub), or for OCR of scanned or image-only PDFs (anydoc does not OCR; route
to OCR tooling).
license: MIT
compatibility: >-
Node.js >= 20 and npx. The pinned CLI is @firecrawl/anydoc@0.1.6; the native
binary ships via npm optionalDependencies (no install step, no postinstall, no
compilation). Conversion runs entirely on your machine — no services, no API
keys, no uploads. The first npx run downloads the package once (network
required); later runs use the npm cache.
metadata:
skills: anydoc, markdown, conversion, docx, xlsx, pptx, pdf, odt, ods, odp, rtf, epub, csv, office, documents, firecrawl
tags: conversion, markdown, office, documents
source: https://github.com/firecrawl/anydoc
allowed-tools: Bash Read
---
# Any Doc — office documents to GitHub-Flavored Markdown
The `anydoc` skill converts office documents, spreadsheets, presentations,
ebooks, CSV, and text-based PDFs into GitHub-Flavored Markdown using the pinned
Any Doc CLI (`@firecrawl/anydoc` v0.1.6). One shared document model and one GFM
serializer produce the same logical output across formats, and conversion runs
locally in milliseconds — no service, no API key, no file upload.
## Overview
Load this skill when a task needs the *contents* of a document the agent cannot
read directly: a Word report to summarize, a spreadsheet to turn into a table,
a slide deck to extract, a CSV to analyze, or an ebook or PDF to quote from.
The skill ships a small Python helper (`scripts/anydoc`) that wraps the pinned
CLI and adds input pre-validation, friendly error hints, batch conversion, and
`--dry-run`/`--json` output. Every recipe in [references/workflows.md](references/workflows.md)
also shows the raw `npx` invocation, so the skill works with or without the
helper.
## When to use
- **Convert a document to markdown** — Word, PowerPoint, Excel, OpenDocument,
RTF, EPUB, CSV, or text-based PDF.
- **Extract structure** — headings, GFM tables, slide titles, speaker notes
(as blockquotes), and footnotes.
- **Feed documents to an LLM** — one-pass conversion to clean markdown for
summarization, extraction, or retrieval ingestion.
- **Batch a folder** — convert a directory of mixed office files for a vault
or knowledge base.
- **Read a document from stdin** — pipe bytes into `anydoc -`.
## Format coverage (summary)
anydoc covers **8 format families / 21 extensions** through **12 canonical
parsers**. The canonical formats are `doc, docx, odt, pdf, ppt, pptx, rtf,
epub, xlsx, ods, odp, csv`; extension aliases map through them (`.docm`→docx,
`.xls`→xlsx, `.pptm`→pptx, and so on).
| Family | Extensions | Expected GFM output |
| --- | --- | --- |
| Word | `.doc` `.docx` `.docm` | `#``######` headings, GFM tables, `[^n]` footnotes |
| PowerPoint | `.ppt` `.pps` `.pot` `.pptx` `.pptm` `.ppsx` `.ppsm` | slide titles as plain paragraphs, bullet lists, speaker notes as `>` blockquotes, GFM tables (PPTX/ODP; legacy `.ppt` flattens tables to text lines) |
| Excel | `.xls` `.xlsx` `.xlsm` `.xlsb` | `## <sheet name>` heading + one GFM table per worksheet; number formats dropped (raw cell values) |
| OpenDocument | `.odt` `.ods` `.odp` | same document/slide shapes as DOCX/PPTX; ODS keeps formatted display values |
| Rich Text Format | `.rtf` | same document shape as DOCX/ODT |
| EPUB | `.epub` | `#` chapter headings, GFM tables, internal anchor links |
| CSV | `.csv` | one GFM table; label-like first row promoted to header; delimiter sniffing; UTF-16 with BOM |
| PDF | `.pdf` | headings + inline emphasis, but a lower-fidelity pipeline: tables flatten to text, footnotes and links degrade. **Scanned or image-only PDFs fail** — anydoc does not OCR |
See [references/formats.md](references/formats.md) for the full per-format
expectations and fidelity caveats, and [references/errors.md](references/errors.md)
for the exact failure messages (including the no-OCR error).
## Command Map
Commands are shown relative to the repository root. `<file>` is any document
path (for example `anydoc/fixtures/fixture-handmade-outline.docx`); `-` reads
the document from stdin.
| Need | Command |
| --- | --- |
| Convert one file (markdown to stdout) | `anydoc/scripts/anydoc convert <file>` |
| Convert one file to a markdown file | `anydoc/scripts/anydoc convert <file> -o out.md` |
| Convert many files to a directory | `anydoc/scripts/anydoc batch <file1> <file2> ... --out-dir out/` |
| Show the tool and pinned CLI version | `anydoc/scripts/anydoc info` |
| Raw pinned CLI, one document | `npx -y @firecrawl/anydoc@0.1.6 <file> [-o out.md]` |
| Raw pinned CLI, read stdin | `cat data.csv \| npx -y @firecrawl/anydoc@0.1.6 - --format csv` |
Notes:
- `scripts/anydoc` is an executable Python 3 script (shebang `#!/usr/bin/env
python3`); `python3 anydoc/scripts/anydoc ...` is equivalent when the
executable bit is unavailable.
- The raw `npx -y @firecrawl/anydoc@0.1.6` rows are the ground truth for
conversion behavior; the wrapper delegates to exactly that command.
- Always pin `@0.1.6` for reproducible conversions. `-y` answers npx's
"Ok to proceed?" prompt non-interactively — the CLI itself never prompts.
- Both forms share the same contract: one document per invocation, exit code
`0` success / `1` conversion or IO failure / `2` usage error, diagnostics as
exactly one `anydoc: <message>` line on stderr, and no prompts.
## Reference Routing
Load these on demand — one per topic:
- [references/formats.md](references/formats.md) — the 8 families / 21
extensions / 12 parsers, what GFM each format produces, and the fidelity
caveats (xlsx/xls number-format drop vs ODS preserved display values, legacy
`.ppt` table flattening, PDF lower-fidelity pipeline, merged-cell covered
spans, ODP same-serializer).
- [references/cli-reference.md](references/cli-reference.md) — verbatim
`--help`, every flag (`-o`, `-f`, `-h`, `-V`, `--format=x`, `--`), stdin via
`-`, stdout/stderr conventions including EPIPE, version pinning, Node >= 20,
and first-run/offline network behavior.
- [references/errors.md](references/errors.md) — exit codes 0/1/2, the verbatim
error-message vocabulary (io, unsupported, malformed, encrypted, EISDIR,
resource-limit, usage errors), the no-OCR caveat, and troubleshooting recipes.
- [references/workflows.md](references/workflows.md) — single conversion,
batch loops, vault ingestion, stdin/stdout piping, output verification,
large-file/resource-limit behavior, and startup cost.
- [references/sources.md](references/sources.md) — upstream URLs, access dates,
fixture provenance, and how every documented claim was verified against the
real CLI.
## When not to use
- **Generating, editing, or validating documents** — anydoc only converts
existing documents *to markdown*; it never creates, edits, or checks
documents. Use the `documents` skill for generation, inspection, and
validation of PDF/Word/Excel/PowerPoint artifacts.
- **Ebook packaging or EPUB authoring** — use the `epub` skill. anydoc reads
EPUBs to markdown but never writes or validates EPUB containers.
- **Scanned or image-only PDFs (OCR)** — anydoc does not perform OCR. Such
PDFs fail as `unsupported` with the OCR message; route the file to OCR
tooling or the hosted Firecrawl Parse API instead of retrying locally. Do
not claim local OCR support.
- **HTML and other web content** — HTML is not a supported input format; use a
web-scraping skill instead.
- **Binary media (images, video, audio)** — embedded images render as alt text
only; anydoc cannot transcribe media content.
- **Layout or rendering work** — output is GitHub-Flavored Markdown only; there
is no pagination, font, or template control.
- **Password-protected files** — encrypted documents fail with
`anydoc: document is encrypted`; there is no password or decryption option.
## Verification
Confirm a conversion before reporting it as done:
1. **Check the exit code.** `0` means the CLI produced markdown. `1` means the
document could not be read or converted — read the single `anydoc: <message>`
stderr line and match it against [references/errors.md](references/errors.md).
`2` means the command itself was a usage error (bad flag, missing input,
invalid `--format`).
2. **Check the output shape.** The markdown must contain the structural markers
your format actually produces:
- Word / ODT / RTF / text-based PDF: `#`/`##` headings. For PDF, do not
expect GFM tables or `[^1]:` footnote definitions — that pipeline
flattens them.
- Spreadsheets (xlsx/xls/ods) and CSV: `|`-delimited GFM tables. xlsx/xls
show raw cell values (`0.155`, `1234.5`); ODS shows formatted display
values (`15.5%`, `$1,234.50`).
- Presentations (pptx/odp): slide titles as plain paragraphs, `>`
blockquote speaker notes, GFM tables. Legacy `.ppt` flattens tables to
bare text lines.
- EPUB: `#` chapter headings and internal anchor links.
3. **Write large outputs to a file with `-o`.** `-o out.md` keeps stdout silent
and gives a reviewable file instead of streaming the whole document into
context.
4. **Verify tables survived.** If the source had tables and the output has no
`|` rows, consult the format caveats — PDF and legacy `.ppt` flatten tables
by design, not by error.
**Stop when** the conversion exits 0 and the structural markers match the
source format. Do not re-run or retry on a documented failure mode (encrypted,
malformed, scanned/image-only, unsupported) without changing the input; report
the documented message and route as [references/errors.md](references/errors.md)
instructs.
+159
View File
@@ -0,0 +1,159 @@
{
"schema_version": 1,
"skill_name": "anydoc",
"evals": [
{
"id": "docx-headings",
"prompt": "Convert this Word document to markdown and extract its headings.",
"expected_output": "The conversion exits 0 with empty stderr and emits GitHub-Flavored Markdown whose structure is expressed as ATX headings, including the lines `## Style heading stays a heading`, `### Direct level overrides the style`, and `# Direct outline without a style`.",
"assertions": [
"The output contains the heading line `## Style heading stays a heading`",
"The output contains the heading line `### Direct level overrides the style`",
"The output contains the heading line `# Direct outline without a style`"
],
"files": ["fixtures/fixture-handmade-outline.docx"]
},
{
"id": "xlsx-table-cell-values",
"prompt": "Convert this spreadsheet to markdown and show me the cell values as a table.",
"expected_output": "The conversion exits 0 and emits a `## Values` heading followed by a GFM table whose cells carry the raw cell values: `Percent | 0.155`, `Currency | 1234.5`, `Thousands | 9876543`, `Date | 2026-03-15` — the number formats are dropped, so the output must NOT contain `15.5%` or `$1,234.50`.",
"assertions": [
"The output contains a `## Values` heading",
"The output contains the table row `| Percent | 0.155 | fifteen and a half |`",
"The output contains the raw value `1234.5` for the Currency row",
"The output does not contain the formatted values `15.5%` or `$1,234.50`"
],
"files": ["fixtures/sheet.xlsx"]
},
{
"id": "pptx-slides-structure",
"prompt": "Convert this PowerPoint deck to markdown, keeping the slide structure.",
"expected_output": "The conversion exits 0 and preserves slide structure: slide titles like `Deck Title Slide` and `Numbers Slide` render as plain paragraphs (not headings), top-level bullets carry an indented nested detail, speaker notes render as blockquotes (`> Speaker note for the intro slide.`), and the slide table renders as a GFM table with the row `| North | 42 |`.",
"assertions": [
"The output contains `Deck Title Slide` as a plain paragraph, not a markdown heading",
"The output contains the blockquote line `> Speaker note for the intro slide.`",
"The output contains the GFM table row `| North | 42 |`"
],
"files": ["fixtures/pres.pptx"]
},
{
"id": "csv-table",
"prompt": "Convert this CSV file to a markdown table.",
"expected_output": "The conversion exits 0 and renders the file as a single GFM table with the first row promoted to the header: `| Kind | Value | Note |` with a separator row and body rows carrying the cell values `15.5%` and `fifteen and a half`.",
"assertions": [
"The output contains the header row `| Kind | Value | Note |`",
"The output contains a row carrying the cell values `15.5%` and `fifteen and a half`"
],
"files": ["fixtures/fixture-sheet.csv"]
},
{
"id": "legacy-doc-converts",
"prompt": "Convert this legacy .doc file to markdown.",
"expected_output": "The conversion exits 0 with empty stderr and emits the shared document serializer's GFM shape: a `# Fixture Document` title, `##` section headings (including `## Lists` and `## Table`), a GFM table with merged cells as empty covered spans, and `[^1]: ...` footnote definitions at the end of the document.",
"assertions": [
"The output contains the heading `# Fixture Document`",
"The output contains at least three `##` section headings including `## Lists` and `## Table`",
"The output contains a `[^1]:` footnote definition block"
],
"files": ["fixtures/text.doc"]
},
{
"id": "image-only-pdf-no-ocr",
"prompt": "Convert this scanned PDF to markdown.",
"expected_output": "The conversion FAILS by design: exit code 1, empty stdout, and exactly one stderr line `anydoc: unsupported input: PDF has no extractable text (Scanned, 1 pages): OCR is required`. anydoc does not perform OCR; the correct response is to route the file to OCR tooling or the hosted Firecrawl Parse API, not to retry locally.",
"assertions": [
"The conversion exits with code 1 and emits no markdown",
"Stderr contains the verbatim message `anydoc: unsupported input: PDF has no extractable text (Scanned, 1 pages): OCR is required`",
"The response states that OCR is required and routes to OCR tooling or Firecrawl Parse rather than retrying locally"
],
"files": ["fixtures/scanned-image-only.pdf"]
},
{
"id": "ods-preserved-values",
"prompt": "Convert this OpenDocument spreadsheet to markdown. I need the formatted display values.",
"expected_output": "The conversion exits 0 and emits a `## Values` heading plus a GFM table whose cells carry the FORMATTED display values — `Percent | 15.5%`, `Currency | $1,234.50`, `Thousands | 9,876,543` — explicitly contrasting with the xlsx/xls number-format drop.",
"assertions": [
"The output contains the table row `| Percent | 15.5% | fifteen and a half |`",
"The output contains the formatted values `15.5%` and `$1,234.50`",
"The output does not contain the raw values `0.155` or `1234.5`"
],
"files": ["fixtures/sheet.ods"]
},
{
"id": "odt-converts",
"prompt": "Convert this ODT document to markdown and show me the structure.",
"expected_output": "The conversion exits 0 with empty stderr and emits the document shape shared with DOCX/DOC/RTF: a `# Fixture Document` title, `##` section headings, a GFM table, and `[^1]:` / `[^2]:` footnote definition lines at the end.",
"assertions": [
"The output contains the heading `# Fixture Document`",
"The output contains `##` section headings such as `## Lists` and `## Table`",
"The output contains the footnote definition `[^1]: Footnote after an astral character.`"
],
"files": ["fixtures/text.odt"]
},
{
"id": "pdf-text-lower-fidelity",
"prompt": "Convert this text-based PDF to markdown. Will the table survive the conversion?",
"expected_output": "The conversion exits 0 with empty stderr and preserves top-level structure via `# Fixture Document` and `##` section headings (`## Lists`, `## Table`, `## Notes and special text`), but the PDF pipeline is lower-fidelity: the table flattens into the plain paragraph `Wide head End Tall B2 C2 B3 C3` with no GFM table, footnote markers degrade to inline superscript glyphs with no `[^1]:` definition block, and links are not emitted as markdown links.",
"assertions": [
"The output contains `# Fixture Document` and the `##` section headings including `## Table`",
"The table region flattens to the plain paragraph `Wide head End Tall B2 C2 B3 C3` with no GFM table row",
"The output contains no `[^1]:` footnote definition block"
],
"files": ["fixtures/fixture-text.pdf"]
},
{
"id": "legacy-ppt-flattens-tables",
"prompt": "Convert this legacy PowerPoint file to markdown and keep the slides' content.",
"expected_output": "The conversion exits 0 and preserves slide text: plain-paragraph titles `Deck Title Slide` and `Numbers Slide` and the blockquote speaker note `> Speaker note for the intro slide.`. The Numbers Slide table flattens to bare text lines (`Region`, `Total`, `North`, `42`) rather than a GFM table, unlike PPTX and ODP.",
"assertions": [
"The output contains the blockquote line `> Speaker note for the intro slide.`",
"The table content renders as bare text lines including `North` and `42`",
"The output does not contain a GFM table row `| North | 42 |`"
],
"files": ["fixtures/pres.ppt"]
},
{
"id": "odp-slides-structure",
"prompt": "Convert this OpenDocument presentation to markdown, preserving the slide structure.",
"expected_output": "The conversion exits 0 and emits the same slide shape as PPTX: `Deck Title Slide` renders as a plain paragraph (not a heading), the speaker note renders as the blockquote `> Speaker note for the intro slide.`, and the slide table renders as a GFM table containing the row `| North | 42 |`.",
"assertions": [
"The output contains `Deck Title Slide` as a plain paragraph, not a markdown heading",
"The output contains the blockquote line `> Speaker note for the intro slide.`",
"The output contains the GFM table row `| North | 42 |`"
],
"files": ["fixtures/pres.odp"]
},
{
"id": "rtf-converts",
"prompt": "Convert this RTF document to markdown and extract the structure.",
"expected_output": "The conversion exits 0 with empty stderr and emits the shared document serializer's shape: `# Fixture Document`, `##` section headings including `## Lists` and `## Table`, a GFM table with merged cells as empty covered spans, and `[^1]:` footnote definitions at the end.",
"assertions": [
"The output contains the heading `# Fixture Document`",
"The output contains `##` section headings such as `## Lists` and `## Table`",
"The output contains the footnote definition `[^1]: Footnote after an astral character.`"
],
"files": ["fixtures/text.rtf"]
},
{
"id": "epub-converts",
"prompt": "Convert this EPUB ebook to markdown, keeping the chapter structure.",
"expected_output": "The conversion exits 0 with empty stderr and emits `# Fixture Book`, `# Chapter One` and `# Chapter Two` headings, a GFM table containing `| Bolts | 12 |`, and internal anchor links that resolve to fragments such as `[Chapter Two](#epub-text-ch002-xhtml-chapter-two)`.",
"assertions": [
"The output contains the chapter heading `# Chapter One`",
"The output contains the GFM table row `| Bolts | 12 |`",
"The output contains the internal anchor link `[Chapter Two](#epub-text-ch002-xhtml-chapter-two)`"
],
"files": ["fixtures/book.epub"]
},
{
"id": "csv-quoted-cells",
"prompt": "Convert this CSV to a markdown table. Some cells contain commas and newlines.",
"expected_output": "The conversion exits 0 and renders the file as a single GFM table with the first row promoted to the header: `| name | desc | qty |` followed by a separator row, with quoted content intact — `| padded | comma, inside | 3 |` keeps the embedded comma and `| plain | multi line | 4 |` keeps the embedded newline.",
"assertions": [
"The output contains the promoted header row `| name | desc | qty |`",
"The output contains the row `| padded | comma, inside | 3 |` with the embedded comma preserved"
],
"files": ["fixtures/fixture-handmade-quoted.csv"]
}
]
}
Binary file not shown.
View File
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,4 @@
name,desc,qty
" padded ","comma, inside",3
plain,"multi
line",4
1 name desc qty
2 padded comma, inside 3
3 plain multi line 4
Binary file not shown.
@@ -0,0 +1,3 @@
a;b;c
"1,5";"2,5";x
"3,0";y;z
1 a b c
2 1,5 2,5 x
3 3,0 y z
Binary file not shown.
Binary file not shown.
1 col1 col2
2 naïve café
3 Αθήνα 数据
+8
View File
@@ -0,0 +1,8 @@
Kind,Value,Note
Percent,15.5%,fifteen and a half
Currency,1234.5,dollars
Thousands,9876543,grouped
Date,2026-03-15,ides of March
Duration,26:30:15,over a day
Tiny,0.0000004,four ten-millionths
Boolean,TRUE,yes
1 Kind Value Note
2 Percent 15.5% fifteen and a half
3 Currency 1234.5 dollars
4 Thousands 9876543 grouped
5 Date 2026-03-15 ides of March
6 Duration 26:30:15 over a day
7 Tiny 0.0000004 four ten-millionths
8 Boolean TRUE yes
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+177
View File
@@ -0,0 +1,177 @@
{\rtf1\ansi\deff3\adeflang1025
{\fonttbl{\f0\froman\fprq2\fcharset0 Times New Roman;}{\f1\froman\fprq2\fcharset2 Symbol;}{\f2\fswiss\fprq2\fcharset0 Arial;}{\f3\froman\fprq2\fcharset0 Liberation Serif{\*\falt Times New Roman};}{\f4\fswiss\fprq2\fcharset0 Liberation Sans{\*\falt Arial};}{\f5\fnil\fprq2\fcharset0 Microsoft YaHei;}{\f6\fnil\fprq2\fcharset0 Arial;}{\f7\fswiss\fprq0\fcharset128 Arial;}}
{\colortbl;\red0\green0\blue0;\red0\green0\blue255;\red0\green255\blue255;\red0\green255\blue0;\red255\green0\blue255;\red255\green0\blue0;\red255\green255\blue0;\red255\green255\blue255;\red0\green0\blue128;\red0\green128\blue128;\red0\green128\blue0;\red128\green0\blue128;\red128\green0\blue0;\red128\green128\blue0;\red128\green128\blue128;\red192\green192\blue192;}
{\stylesheet{\s0\snext0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052 Normal;}
{\s1\snext25\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\ilvl0\outlinelevel0\nowidctlpar\hyphpar0\ltrpar\fs36\b\cf0\f3\lang2057\kerning1\dbch\af8\langfe2052 heading 1;}
{\s2\snext25\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\ilvl1\outlinelevel1\nowidctlpar\hyphpar0\ltrpar\fs30\b\cf0\f3\lang2057\kerning1\dbch\af8\langfe2052 heading 2;}
{\s3\snext25\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\ilvl2\outlinelevel2\nowidctlpar\hyphpar0\ltrpar\fs26\b\cf0\f3\lang2057\kerning1\dbch\af8\langfe2052 heading 3;}
{\*\cs15\snext15\loch\b0 NotBold;}
{\*\cs16\snext16\loch\b Bold;}
{\*\cs17\snext17\loch\i Italic;}
{\*\cs18\snext18\loch\strike Struck;}
{\*\cs19\snext19\loch\cf9\ul\ulc0 Hyperlink;}
{\*\cs20\snext20 Footnote Characters;}
{\*\cs21\snext21\loch\super footnote reference;}
{\*\cs22\snext22 Endnote Characters;}
{\*\cs23\snext23\loch\super endnote reference;}
{\s24\sbasedon0\snext25\rtlch\af6\afs28 \ltrch\hich\af4\loch\sb240\sa120\keepn\f4\fs28\dbch\af5 Heading;}
{\s25\sbasedon0\snext25\loch\sl276\slmult1\sb0\sa140 Body Text;}
{\s26\sbasedon25\snext26\rtlch\af7 \ltrch List;}
{\s27\sbasedon0\snext27\rtlch\af7\afs24\ai \ltrch\loch\sb120\sa120\noline\fs24\i caption;}
{\s28\sbasedon0\snext28\rtlch\af7 \ltrch\loch\noline Index;}
{\s29\snext29\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\b\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052 StrongPara;}
{\s30\sbasedon0\snext30\loch\nowidctlpar\noline Table Contents;}
{\s31\sbasedon0\snext31\rtlch\afs20 \ltrch\loch\fi-340\li340\lin340\noline\fs20 footnote text;}
{\s32\sbasedon0\snext32\rtlch\afs20 \ltrch\loch\fi-340\li340\lin340\noline\fs20 endnote text;}
{\s33\sbasedon0\snext33 Frame Contents;}
}{\*\listtable{\list\listtemplateid1
{\listlevel\levelnfc255\leveljc0\levelstartat1\levelfollow2{\leveltext \'00;}{\levelnumbers;}\fi0\li0}
{\listlevel\levelnfc255\leveljc0\levelstartat1\levelfollow2{\leveltext \'00;}{\levelnumbers;}\fi0\li0}
{\listlevel\levelnfc255\leveljc0\levelstartat1\levelfollow2{\leveltext \'00;}{\levelnumbers;}\fi0\li0}
{\listlevel\levelnfc255\leveljc0\levelstartat1\levelfollow2{\leveltext \'00;}{\levelnumbers;}\fi0\li0}
{\listlevel\levelnfc255\leveljc0\levelstartat1\levelfollow2{\leveltext \'00;}{\levelnumbers;}\fi0\li0}
{\listlevel\levelnfc255\leveljc0\levelstartat1\levelfollow2{\leveltext \'00;}{\levelnumbers;}\fi0\li0}
{\listlevel\levelnfc255\leveljc0\levelstartat1\levelfollow2{\leveltext \'00;}{\levelnumbers;}\fi0\li0}
{\listlevel\levelnfc255\leveljc0\levelstartat1\levelfollow2{\leveltext \'00;}{\levelnumbers;}\fi0\li0}
{\listlevel\levelnfc255\leveljc0\levelstartat1\levelfollow2{\leveltext \'00;}{\levelnumbers;}\fi0\li0}\listid1}
{\list\listtemplateid2
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow2{\leveltext \'02\'00.;}{\levelnumbers\'01;}\fi0\li0}
{\listlevel\levelnfc4\leveljc0\levelstartat1\levelfollow2{\leveltext \'02\'01);}{\levelnumbers\'01;}\fi0\li0}
{\listlevel\levelnfc2\leveljc0\levelstartat1\levelfollow2{\leveltext \'02\'02.;}{\levelnumbers\'01;}\fi0\li0}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'03.;}{\levelnumbers\'01;}\fi-360\li1800}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'04.;}{\levelnumbers\'01;}\fi-360\li2160}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'05.;}{\levelnumbers\'01;}\fi-360\li2520}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'06.;}{\levelnumbers\'01;}\fi-360\li2880}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'07.;}{\levelnumbers\'01;}\fi-360\li3240}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'08.;}{\levelnumbers\'01;}\fi-360\li3600}\listid2}
{\list\listtemplateid3
{\listlevel\levelnfc23\leveljc0\levelstartat1\levelfollow2{\leveltext \'01\u8226 ?;}{\levelnumbers;}\f9\fi0\li0}
{\listlevel\levelnfc23\leveljc0\levelstartat1\levelfollow2{\leveltext \'01\u9702 ?;}{\levelnumbers;}\f9\fi0\li0}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'02.;}{\levelnumbers\'01;}\fi-360\li1440}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'03.;}{\levelnumbers\'01;}\fi-360\li1800}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'04.;}{\levelnumbers\'01;}\fi-360\li2160}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'05.;}{\levelnumbers\'01;}\fi-360\li2520}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'06.;}{\levelnumbers\'01;}\fi-360\li2880}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'07.;}{\levelnumbers\'01;}\fi-360\li3240}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'08.;}{\levelnumbers\'01;}\fi-360\li3600}\listid3}
{\list\listtemplateid4
{\listlevel\levelnfc1\leveljc0\levelstartat1\levelfollow2{\leveltext \'02\'00.;}{\levelnumbers\'01;}\fi0\li0}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'01.;}{\levelnumbers\'01;}\fi-360\li1080}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'02.;}{\levelnumbers\'01;}\fi-360\li1440}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'03.;}{\levelnumbers\'01;}\fi-360\li1800}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'04.;}{\levelnumbers\'01;}\fi-360\li2160}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'05.;}{\levelnumbers\'01;}\fi-360\li2520}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'06.;}{\levelnumbers\'01;}\fi-360\li2880}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'07.;}{\levelnumbers\'01;}\fi-360\li3240}
{\listlevel\levelnfc0\leveljc0\levelstartat1\levelfollow0{\leveltext \'02\'08.;}{\levelnumbers\'01;}\fi-360\li3600}\listid4}
}{\listoverridetable{\listoverride\listid1\listoverridecount0\ls1}{\listoverride\listid2\listoverridecount0\ls2}{\listoverride\listid3\listoverridecount0\ls3}{\listoverride\listid4\listoverridecount0\ls4}}{\*\generator LibreOffice/25.2.7.2$Windows_X86_64 LibreOffice_project/5cbfd1ab6520636bb5f7b99185aa69bd7456825d}{\info{\creatim\yr0\mo0\dy0\hr0\min0}{\revtim\yr0\mo0\dy0\hr0\min0}{\printim\yr0\mo0\dy0\hr0\min0}}{\*\userprops}\deftab709\deftab709\deftab709\deftab709\deftab709
\hyphauto1\viewscale100\formshade\nobrkwrptbl\paperh16838\paperw11906\margl1134\margr1134\margt1134\margb1134\sectd\sbknone\sftnnar\saftnnrlc\sectunlocked1\pgwsxn11906\pghsxn16838\marglsxn1134\margrsxn1134\margtsxn1134\margbsxn1134\ftnbj\ftnstart1\ftnrstcont\ftnnar\fet2\aftnrstcont\aftnstart1\aftnnrlc
{\*\ftnsep\chftnsep}\pgndec\pard\plain \s1\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\ilvl0\outlinelevel0\nowidctlpar\hyphpar0\ltrpar\fs36\b\cf0\f3\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain \tab}\ls1 \fi0\li0\lin0\ql\fi0\li0\lin0\ltrpar{\loch
Fixture Document}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Plain paragraph with }{\loch\cs16\loch\b\loch
bold}{\loch
, }{\loch\cs17\loch\i\loch
italic}{\loch
, and }{\loch\cs18\loch\strike\loch
struck}{\loch
runs.}
\par \pard\plain \s29\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\b\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Style-bold paragraph with a }{\loch\cs15\loch\b0\loch
NotBold-styled span}{\loch
inside.}
\par \pard\plain \s2\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\ilvl1\outlinelevel1\nowidctlpar\hyphpar0\ltrpar\fs30\b\cf0\f3\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain \tab}\ls1 \fi0\li0\lin0\ql\fi0\li0\lin0\ltrpar{\loch
Lists}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain 1.\tab}\ilvl0\ls2 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
First numbered}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain 2.\tab}\ilvl0\ls2 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Second numbered}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain a)\tab}\ilvl1\ls2 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Alpha sub one}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain b)\tab}\ilvl1\ls2 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Alpha sub two}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain i.\tab}\ilvl2\ls2 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Roman sub sub}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain 3.\tab}\ilvl0\ls2 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Third numbered}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Interrupting paragraph between lists.}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain 4.\tab}\ilvl0\ls2 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Fourth, continuing the count}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain IV.\tab}\ilvl0\ls4 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Roman starting at four}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain V.\tab}\ilvl0\ls4 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Roman five}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain \u8226\'95\tab}\ilvl0\ls3 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Bullet one}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain \u8226\'95\tab}\ilvl0\ls3 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Bullet two}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain \u9702\'3f\tab}\ilvl1\ls3 \fi0\li0\lin0\ql\tx0\fi0\li0\lin0\ltrpar{\loch
Nested bullet}
\par \pard\plain \s2\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\ilvl1\outlinelevel1\nowidctlpar\hyphpar0\ltrpar\fs30\b\cf0\f3\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain \tab}\ls1 \fi0\li0\lin0\ql\fi0\li0\lin0\ltrpar{\loch
Table}
\par \trowd\trql\ltrrow\trpaddft3\trpaddt0\trpaddfl3\trpaddl0\trpaddfb3\trpaddb0\trpaddfr3\trpaddr0\cellx6425\cellx9638\pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\intbl\ql\ltrpar{\loch
Wide head}\cell\pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\intbl\ql\ltrpar{\loch
End}\cell\row\pard \trowd\trql\ltrrow\trpaddft3\trpaddt0\trpaddfl3\trpaddl0\trpaddfb3\trpaddb0\trpaddfr3\trpaddr0\clvmgf\cellx3213\cellx6425\cellx9638\pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\intbl\ql\ltrpar{\loch
Tall}\cell\pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\intbl\ql\ltrpar{\loch
B2}\cell\pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\intbl\ql\ltrpar{\loch
C2}\cell\row\pard \trowd\trql\ltrrow\trpaddft3\trpaddt0\trpaddfl3\trpaddl0\trpaddfb3\trpaddb0\trpaddfr3\trpaddr0\clvmrg\cellx3213\cellx6425\cellx9638\pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\intbl\ql\ltrpar\loch
\cell\pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\intbl\ql\ltrpar{\loch
B3}\cell\pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\intbl\ql\ltrpar{\loch
C3}\cell\row\pard \pard\plain \s2\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\ilvl1\outlinelevel1\nowidctlpar\hyphpar0\ltrpar\fs30\b\cf0\f3\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain \tab}\ls1 \fi0\li0\lin0\ql\fi0\li0\lin0\ltrpar{\loch
Notes and special text}
\par \ql\fi-340\li340\lin340\noline\ltrpar\pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Music clef \uc0 \u55348\uc1 \u56606\'3f appears before this footnote}{{\super \loch \chftn{\*\footnote \chftn\pard\plain \s31\rtlch\afs20 \ltrch\loch\fi-340\li340\lin340\noline\fs20{\loch
\tab Footnote after an astral character.}
}}
}{\loch
reference.}
\par \ql\fi-340\li340\lin340\noline\ltrpar\pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
An endnote follows here}{{\super \loch \chftn{\*\footnote\ftnalt \chftn\pard\plain \s32\rtlch\afs20 \ltrch\loch\fi-340\li340\lin340\noline\fs20{\loch
\tab Endnote body text.}
}}
}{\loch
.}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Persian with ZWNJ: }{
\u1605\'3f\u1740\'3f\u8204\'3f\u1582\'3f\u1608\'3f\u1575\'3f\u1607\'3f\u1605\'3f}{\loch
. Family emoji: \uc0 \u55357\uc1 \u56424\'3f\u8205\'3f\uc0 \u55357\uc1 \u56425\'3f\u8205\'3f\uc0 \u55357\uc1 \u56423\'3f.}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Markdown specials: *stars* _under_ [bracket] `tick` #hash 1. dotted | pipe.}
\par \pard\plain \s2\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\ilvl1\outlinelevel1\nowidctlpar\hyphpar0\ltrpar\fs30\b\cf0\f3\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain \tab}\ls1 \fi0\li0\lin0\ql\fi0\li0\lin0\ltrpar{\loch
Links and anchors}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
External link to }{{\field{\*\fldinst HYPERLINK "https://example.com/page" }{\fldrslt {\loch\loch\cf9\ul\ulc0\loch
example}{}}}\loch
.}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Relative link to }{{\field{\*\fldinst HYPERLINK "file:///anydoc/tests/fixture-src/sibling.odt" }{\fldrslt {\loch\loch\cf9\ul\ulc0\loch
a sibling file}{}}}\loch
.}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
{\*\bkmkstart plainmark}{\*\bkmkend plainmark}This plain paragraph carries a bookmark.}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Jump to }{{\field{\*\fldinst HYPERLINK "#plainmark" }{\fldrslt {\loch\loch\cf9\ul\ulc0\loch
the bookmarked paragraph}{}}}\loch
.}
\par \pard\plain \s2\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\ilvl1\outlinelevel1\nowidctlpar\hyphpar0\ltrpar\fs30\b\cf0\f3\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain \tab}\ls1 \fi0\li0\lin0\ql\fi0\li0\lin0\ltrpar{\loch
Objects}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Inline image: }{\loch
{\pict{\*\picprop{\sp{\sn wzDescription}{\sv tiny red dot}}{\sp{\sn wzName}{\sv }}}\picscalex1700\picscaley1700\piccropl0\piccropr0\piccropt0\piccropb0\picw1\pich1\picwgoal10\pichgoal10\pngblip
89504e470d0a1a0a0000000d49484452000000010000000108060000001f15c4890000000d4944415478da63fccfc0500f000485018084a98c21000000004945
4e44ae426082}
}{\loch
done.}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Text box: }{
{\shp{\*\shpinst\shpwr2\shpwrk0\shpbypara\shpbyignore\shptop-1362\shpbottom-228\shpbxcolumn\shpbxignore\shpleft0\shpright2835\shpz1{\sp{\sn shapeType}{\sv 202}}{\sp{\sn dxWrapDistLeft}{\sv 72390}}{\sp{\sn dxWrapDistRight}{\sv 72390}}{\sp{\sn dyWrapDistTop}{\sv 72390}}{\sp{\sn dyWrapDistBottom}{\sv 72390}}{\sp{\sn posrelv}{\sv 2}}{\sp{\sn posv}{\sv 1}}{\sp{\sn posrelh}{\sv 2}}{\sp{\sn posh}{\sv 2}}{\sp{\sn dxTextLeft}{\sv 53975}}{\sp{\sn dyTextTop}{\sv 53975}}{\sp{\sn dxTextRight}{\sv 53975}}{\sp{\sn dyTextBottom}{\sv 53975}}{\sp{\sn lineColor}{\sv 0}}{\sp{\sn lineWidth}{\sv 635}}{\shptxt\s33\ql{\loch
Inside the text box.}
\par \pard}}}
}{\loch
after the box.}
\par \pard\plain \s2\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\ilvl1\outlinelevel1\nowidctlpar\hyphpar0\ltrpar\fs30\b\cf0\f3\lang2057\kerning1\dbch\af8\langfe2052{\listtext\pard\plain \tab}\ls1 \fi0\li0\lin0\ql\fi0\li0\lin0\ltrpar{\loch
Quote and code}
\par \pard\plain \s0\rtlch\af6\afs24\alang1081 \ltrch\lang2057\langfe2052\hich\af3\loch\nowidctlpar\hyphpar0\ltrpar\cf0\f3\fs24\lang2057\kerning1\dbch\af8\langfe2052\ql\ltrpar{\loch
Value below one millionth: 0.0000004 should survive.}
\par }
+1
View File
@@ -0,0 +1 @@
hello world
+194
View File
@@ -0,0 +1,194 @@
# CLI reference: the Any Doc CLI (pinned @firecrawl/anydoc@0.1.6)
Everything here was captured by running the pinned CLI on this machine
(`npx -y @firecrawl/anydoc@0.1.6`, version 0.1.6, Node v22). The CLI is a 4.7 KB
Node wrapper (`bin.anydoc = cli.js`) around a native NAPI binding that ships as
an npm `optionalDependency` per platform.
## Verbatim `--help` output
```
anydoc: convert documents to GitHub-Flavored Markdown
Usage:
anydoc <file> [options]
anydoc - [options] < file
Converts one document per invocation and writes the Markdown to stdout.
Pass - as the input to read the document from stdin. Never prompts; all
diagnostics go to stderr.
Options:
-o, --output <path> Write the Markdown to <path> instead of stdout
-f, --format <format> Name the input format instead of detecting it:
doc, docx, odt, pdf, ppt, pptx, rtf, epub, xlsx, ods, odp, csv
(extension aliases like xls, docm, ppsx resolve
to these)
-h, --help Print this help and exit
-V, --version Print the version and exit
The format is detected from the file content; the file extension is the
fallback for signature-less formats (CSV). stdin has no extension, so CSV
input from stdin needs --format csv. Scanned or image-only PDFs need OCR,
which anydoc does not do, and error as unsupported.
Exit codes:
0 success
1 the document could not be read or converted
2 usage error: unknown option, missing input, or invalid --format
Examples:
anydoc report.docx
anydoc slides.pptx -o slides.md
anydoc - --format csv < data.csv
curl -s https://example.com/paper.pdf | anydoc -
```
`anydoc --version` prints exactly `0.1.6` (verified; both `--help` and
`--version` exit 0 and write to stdout).
## Invocation forms
```text
anydoc <file> [options] # convert a path on disk
anydoc - [options] < file # read the document from stdin
```
- `-` as the input reads the document from **stdin**.
- The CLI accepts **exactly one document per invocation** — there is no batch
mode. Passing a second input exits 2:
`anydoc: one document per invocation: unexpected second input '<path>'`.
For multiple documents use a shell loop or `scripts/anydoc batch`
(see [workflows.md](workflows.md)).
## Flag reference
| Token | Behavior (verified) |
| --- | --- |
| `<file>` | Input path. Format detected from content; extension is the fallback for signature-less formats (CSV). |
| `-` | Read the document from stdin. If stdin is a TTY, exits 2 with `anydoc: stdin is a terminal; pipe or redirect a document into anydoc -`. |
| `-o <path>`, `--output <path>` | Write the Markdown to `<path>` instead of stdout. **Silently overwrites** an existing file (verified). Writing to a directory fails with exit 1: `anydoc: EISDIR: illegal operation on a directory, open '<path>'`. With `-o`, stdout stays silent. |
| `-f <fmt>`, `--format <fmt>` | Force the input format instead of detecting it. Values: `doc, docx, odt, pdf, ppt, pptx, rtf, epub, xlsx, ods, odp, csv`. Extension aliases resolve through the parser mapping (verified: `--format xls`, `--format docm` accepted). Invalid value → exit 2: `anydoc: invalid format 'bogus'; expected one of: doc, docx, odt, pdf, ppt, pptx, rtf, epub, xlsx, ods, odp, csv`. |
| `-h`, `--help` | Print help to stdout, exit 0. Works even when the native binding is unavailable. |
| `-V`, `--version` | Print the version (`0.1.6`) to stdout, exit 0. Binding-independent like `--help`. |
| `--format=x` | Inline `=` value syntax is supported for long options (verified: `--format=rtf` works). |
| `--` | End of options: everything after `--` is treated as a positional input (a filename starting with `-`). |
| Missing option value | `anydoc: <option> requires a value` → exit 2 (e.g. `anydoc: -o requires a value`). |
| Unknown option | `anydoc: unknown option '--bogus' (see anydoc --help)` → exit 2. |
| No input | `anydoc: missing input: pass a document path, or - for stdin (see anydoc --help)` → exit 2. |
## stdin / stdout / stderr conventions
- **stdin input** via `-`. Because stdin has no file extension, **CSV from
stdin requires `--format csv`** (CSV has no content signature). Without it,
CSV bytes fail with exit 1:
`anydoc: unsupported input: unrecognized file content: name the format explicitly`.
Verified success pattern:
```bash
printf 'name,role\nAlice,Engineer\n' | npx -y @firecrawl/anydoc@0.1.6 - --format csv
```
- **Markdown goes to stdout only.** With `-o`, stdout stays silent.
- **All diagnostics go to stderr** as exactly one `anydoc: <message>` line per
failure. Nothing is ever printed to stdout on failure.
- **The CLI never prompts** — no confirmation, no interaction. (`-y` on the
`npx` invocation exists only to answer *npx's* package-install prompt.)
- **EPIPE is handled**: if the downstream pipe closes early
(`anydoc big.xlsx | head -n 1`), the CLI exits **0** with no stderr noise
(verified). Piping into `head` is not treated as a conversion failure.
- **No environment variables** — the CLI uses only argv, stdin, and the
filesystem (verified by reading `cli.js`).
## Running it: npx invocation
```bash
npx -y @firecrawl/anydoc@0.1.6 report.docx # markdown to stdout
npx -y @firecrawl/anydoc@0.1.6 slides.pptx -o slides.md # to a file
npx -y @firecrawl/anydoc@0.1.6 - --format csv < data.csv # stdin (CSV needs --format)
curl -s https://example.com/paper.pdf | npx -y @firecrawl/anydoc@0.1.6 - # URL → stdin
```
### Version pinning
Always pin the version: `npx -y @firecrawl/anydoc@0.1.6`. An unpinned
`npx -y @firecrawl/anydoc` floats to the latest published tag, so conversions
are not reproducible across time. All behavior in this skill is documented
against **0.1.6**. The `-y` flag answers npx's "Ok to proceed?" install prompt
non-interactively; without it, bare `npx @firecrawl/anydoc` will prompt on a
cold cache.
### First run and offline behavior
- The **first** `npx` invocation downloads the npm package plus the native
platform binary (network required once). Verified with a fresh empty npm
cache: `env npm_config_cache=$(mktemp -d) npx -y @firecrawl/anydoc@0.1.6 --version`
prints `0.1.6` and exits 0.
- Later runs reuse the npm cache; measured warm startup is ~0.330.55 s per
invocation (see [workflows.md](workflows.md)).
- **Cold-cache offline**: if the package is not cached and there is no network,
npx itself fails with a clear fetch error before anydoc runs. The conversion
itself is fully local — only package retrieval needs network.
- **Permanent / offline-capable alternative**: `npm install -g @firecrawl/anydoc`
once, then invoke `anydoc` directly (still pinning is up to you). This
satisfies the skill's "no service dependency" claim: there is no server, no
API key, and no upload — the only network use is downloading the tool.
## Distribution and system requirements
- **Node.js >= 20** (package `engines`). Verified under Node v22.
- The native binary ships via npm **`optionalDependencies`** — one small
package per platform (`darwin-x64`, `darwin-arm64`, `linux-x64-gnu`,
`linux-arm64-gnu`, `linux-x64-musl`, `linux-arm64-musl`, `win32-x64-msvc`),
with **no postinstall script and no compilation**.
- The npm package `@firecrawl/anydoc` 0.1.6 is ~48 KB unpacked (the binding
package is a few MB per platform); published 2026-08-05T18:29:40Z.
- The Rust crate `anydoc` (crates.io) and Python wheels `firecrawl-anydoc`
(PyPI, imports as `anydoc`, Python >= 3.10) ship in the same release train.
There is no standalone Rust CLI binary (`cargo install anydoc` is an open
feature request) — the CLI exists only through the npm package.
## The wrapper: `scripts/anydoc`
The skill ships a Python 3 standard-library wrapper at `scripts/anydoc` that
delegates to the pinned CLI. It adds value beyond a thin npx alias:
- **`convert <file|-> [-o out.md] [-f <format>] [--json] [--dry-run]`** —
pre-validates the input path (missing file, directory input) and the `-o`
path (existing directory) before invoking the CLI, validates `-f` against
the 21 accepted format names (the 12 canonical parsers plus the 9 aliases,
exit 2 on an invalid name), maps known failure classes to friendly hints
(no-OCR, encrypted, malformed, unsupported), and forwards the CLI's exit
code. Stdin input via `-` is passed straight through. A dash-leading
filename is supported through the CLI's `--` marker with the options first:
`anydoc convert -f csv -- -weird` (the wrapper emits `-o`/`-f` before `--`,
since npx forwards `--` to the CLI and anything after it reads as an extra
input). Absolute paths never need this.
- **`batch <inputs...> [--out-dir DIR] [--json] [--dry-run]`** — converts many
documents one at a time, prints per-file status, continues past failures,
and exits 1 when any input failed. Output naming is deterministic: each
input becomes `<stem>.md` under `--out-dir`, which is created when missing
and defaults to the current working directory. Duplicate inputs convert per
occurrence (a later conversion overwrites the earlier output); same-basename
inputs from different directories collide on the same `<stem>.md` and the
last one wins.
- **`info [--version]`** — reports the tool name and the pinned CLI version
(`anydoc 0.1.6 (wraps @firecrawl/anydoc@0.1.6)`) without invoking the
converter; `info --version` prints exactly `0.1.6`.
- Global **`--json`** (exactly one JSON document on stdout; diagnostics stay
on stderr) and **`--dry-run`** (print what would run — the exact `npx`
command line and output paths — and execute nothing: no CLI spawn, no
output files, no directory creation). With `--json`, `convert` embeds the
converted markdown in the JSON document when `-o` is not given.
- Checks for Node >= 20 (missing `node`, or a version below 20, exits 1 with
a clear message naming Node.js and the required version) and for `npx`
(missing `npx` exits 1 naming `npx` and the pinned package
`@firecrawl/anydoc@0.1.6`); always invokes npx with `-y`; never prompts;
exit codes 0/1/2 mirror the CLI.
Run it as `anydoc/scripts/anydoc <subcommand> ...` from the repository root,
`scripts/anydoc <subcommand> ...` from the skill directory, or
`python3 anydoc/scripts/anydoc <subcommand> ...` anywhere (the executable bit
and `#!/usr/bin/env python3` shebang let it run directly). See
[workflows.md](workflows.md) for recipes and [errors.md](errors.md) for the
error vocabulary.
+222
View File
@@ -0,0 +1,222 @@
# Errors, exit codes, and troubleshooting
Every message below is a **verbatim real stderr capture** from the pinned CLI
(`@firecrawl/anydoc@0.1.6`) run against the committed fixtures in `fixtures/`
(and, for resource limits, generated oversized archives). The CLI prints
exactly one line to stderr, prefixed `anydoc: `, and never prompts.
## Exit codes
| Code | Meaning | Triggers |
| --- | --- | --- |
| `0` | Success | Normal conversion; `--help`/`--version`; also on **EPIPE** when the downstream pipe closes early (`anydoc big.xlsx \| head`). |
| `1` | The document could not be read or converted | Any conversion or IO failure below: missing file, unsupported input, scanned/image-only PDF, malformed archive, encrypted document, resource limit, `-o` pointing at a directory. |
| `2` | Usage error | Unknown option, missing input, invalid `--format`, more than one input, an option missing its value, stdin is a terminal. |
## Conversion / IO failures (exit code 1)
### io — the file could not be read
```
anydoc: io error: No such file or directory (os error 2)
```
This is the missing-file case (`to_markdown` path only; stdin and byte APIs
have no io error).
### unsupported — unknown format or unconvertible content
Unknown content **and** unknown extension (the extension is echoed as given):
```
anydoc: unsupported input: unrecognized file content and extension: unsupported.xyz
```
Verified against `fixtures/unsupported.xyz` (run from the fixture directory,
the tail is `unsupported.xyz`; when you pass a longer path, that path is echoed).
Recognized format but unconvertible content — a **scanned or image-only PDF**
(the CLI detects the page count and that it looks scanned):
```
anydoc: unsupported input: PDF has no extractable text (Scanned, 1 pages): OCR is required
```
Verified against `fixtures/scanned-image-only.pdf`: exit 1, empty stdout.
### unsupported — stdin without a format
CSV has no content signature and stdin has no extension, so CSV piped to `-`
without `--format csv` fails:
```
anydoc: unsupported input: unrecognized file content: name the format explicitly
```
Fix: add `--format csv` (e.g. `cat data.csv | npx -y @firecrawl/anydoc@0.1.6 - --format csv`).
### malformed — structurally unusable archive
An empty (0-byte) `.docx` and a truncated `.docx` both produce:
```
anydoc: malformed document: not a readable zip archive: invalid Zip archive: Could not find EOCD
```
Verified against `fixtures/empty--errors.docx`. Any other structurally broken
package surfaces the same class.
### encrypted — password-protected document
```
anydoc: document is encrypted
```
Verified against `fixtures/encrypted--errors.odt`. There is **no password or
decryption option** anywhere in the CLI or library — the only fix is an
unencrypted copy of the file.
### resourceLimit — fixed safety limits (decompression / nesting / node count)
Zip-bomb style DOCX (giant `word/document.xml`):
```
anydoc: resource limit exceeded (max_entry_bytes): word/document.xml declares 201326759 decompressed bytes
```
Image-bomb style DOCX (giant `word/media/image1.png`):
```
anydoc: resource limit exceeded (max_entry_bytes): word/media/image1.png declares 201326592 decompressed bytes
```
The **character-exact prefix** is:
```
anydoc: resource limit exceeded (max_entry_bytes):
```
with a tail naming the offending entry and the declared decompressed size —
the tail varies by entry, so match on the prefix. Verified also with a
generated 250 MB-entry zip (tail: `word/document.xml declares 250000000
decompressed bytes`). anydoc rejects zip/image bombs via `max_entry_bytes`;
conversion is **not streaming**, and the whole entry is checked before use.
### output-is-directory (EISDIR)
`-o` pointing at an existing directory fails with exit 1:
```
anydoc: EISDIR: illegal operation on a directory, open '<path>'
```
Verified: `npx -y @firecrawl/anydoc@0.1.6 report.rtf -o /tmp` prints
`anydoc: EISDIR: illegal operation on a directory, open '/tmp'` and exits 1.
Fix: pass a file path (or a path in a directory that exists); anydoc **does
not create directories**.
## Usage errors (exit code 2)
All verified verbatim:
```
anydoc: missing input: pass a document path, or - for stdin (see anydoc --help)
anydoc: unknown option '--bogus' (see anydoc --help)
anydoc: invalid format 'bogus'; expected one of: doc, docx, odt, pdf, ppt, pptx, rtf, epub, xlsx, ods, odp, csv
anydoc: one document per invocation: unexpected second input '<path>'
anydoc: stdin is a terminal; pipe or redirect a document into anydoc -
anydoc: -o requires a value (pattern: `<option> requires a value`)
```
Notes:
- `unknown option '--bogus'` echoes the offending token; `one document per
invocation` echoes the second input path as given; the `-o requires a value`
pattern applies to `-f` too (`-f requires a value`).
- Usage errors never touch the filesystem and produce no markdown.
## The no-OCR caveat (read before converting PDFs)
- anydoc converts **text-based PDFs locally** via `pdf-inspector`; there is no
OCR service anywhere in the pipeline.
- **Scanned / image-only PDFs fail as `unsupported`** with the exact message
above (`... OCR is required`). The library's stance: "Scanned and image-only
PDFs need OCR, which anydoc does not do."
- **Route, don't retry.** When this message fires: report the exact error,
state that OCR is required, and direct the user to OCR tooling or the hosted
Firecrawl Parse API. Do **not** retry the same file locally, do **not**
claim anydoc can OCR, and do **not** fabricate the document's content.
- There is no password option, no OCR option, and no retry-until-success
behavior to enable.
## Troubleshooting recipes
| Symptom | Message to match | Fix |
| --- | --- | --- |
| File not found | `io error: No such file or directory` | Check the path; anydoc does not glob or resolve relative to the skill. |
| Unknown file type | `unsupported input: unrecognized file content and extension: <path>` | Confirm the extension is one of the 21 supported; or force it with `--format <name>`. |
| Scanned PDF | `PDF has no extractable text (Scanned, N pages): OCR is required` | Route to OCR tooling / Firecrawl Parse. Never retry locally. |
| Encrypted file | `document is encrypted` | Ask for an unencrypted copy; there is no password option. |
| Empty/truncated archive | `malformed document: not a readable zip archive` | Re-download or re-export the file. Note: some damaged files still convert partially (see below). |
| Huge or malicious archive | `resource limit exceeded (max_entry_bytes):` | anydoc rejected the entry by design; do not bypass. For genuinely large real documents, use `-o out.md`. |
| `-o` "failed" | `EISDIR: illegal operation on a directory, open '<path>'` | Point `-o` at a file path inside an existing directory. |
| CSV from stdin failed | `unsupported input: unrecognized file content: name the format explicitly` | Add `--format csv`. |
| Command rejected | any `anydoc: ...` exit-2 message | Re-read the usage: one input only, valid `--format`, options before/after correctly placed. |
## Graceful recovery — exit 0 is not byte-perfect fidelity
The library skips broken parts rather than failing whenever some meaningful
Markdown is still producible. The upstream test suite ships `*--recovers.*`
and `*--skips.*` fixtures (e.g. `mismatched--recovers.docx`,
`unbalanced--recovers.rtf`, `corrupt-styles--skips.docx`): structurally damaged
documents often convert with exit 0, dropping only the broken part. So a
conversion that exits 0 can still be incomplete — run the output-verification
steps in [workflows.md](workflows.md) and [SKILL.md](../SKILL.md) when fidelity
matters.
## Wrapper (`scripts/anydoc`) error behavior
The wrapper mirrors the CLI's contract and adds pre-validation and hints:
- **Pre-validation errors (exit 1)**: a missing input path, a directory-as-
input, or an `-o` path that is an existing directory is caught before the
CLI runs — stderr names the path and the problem (e.g.
`anydoc: input file not found: <path>`,
`anydoc: input path is a directory, not a file: <path>`,
`anydoc: output path is a directory: <path> (pass a file path; -o does not
create directories)`), with no traceback and no prompt.
- **Usage errors (exit 2)**: an unknown option, a missing input, or an invalid
`-f` value exits 2 with a usage message on stderr before any CLI invocation.
The accepted `-f` names are the 12 canonical formats plus the 9 aliases
(`anydoc: invalid format 'bogus'; expected one of: ...`).
- **Friendly hints (exit 1)**: known failure classes get a hint plus a next
step on stderr — no-OCR (`scanned-image-only.pdf` → "anydoc does not
perform OCR. Route the file to OCR tooling or the hosted Firecrawl Parse
API; do not retry it locally."), encrypted ("the document is encrypted or
password-protected — supply an unencrypted copy"), malformed ("the document
is malformed or corrupt (not a readable zip archive) — re-export or
re-download the file and retry"), unsupported ("unsupported or unrecognized
file type — check that the extension is one of the supported formats, or
force it with `-f <format>`"). The raw CLI error line is always printed
first, verbatim.
- **Node check (exit 1)**: if `node` is missing or older than v20, stderr
states that Node.js >= 20 is required (`anydoc: Node.js >= 20 is required
but `node` was not found on PATH ...` / `anydoc: Node.js version v18.20.0 is
too old; anydoc requires Node.js >= 20 ...`), before any CLI invocation.
- **npx missing (exit 1)**: stderr names `npx` and the pinned package
(`@firecrawl/anydoc@0.1.6`): `anydoc: `npx` was not found on PATH —
conversion runs via `npx -y @firecrawl/anydoc@0.1.6`. Install Node.js >= 20
(which ships npx), or install the CLI permanently with `npm install -g
@firecrawl/anydoc`.`.
- **Batch exit policy**: `batch` exits 1 when any input failed; per-file
status lines (`ok <file> -> <out.md>` / `FAIL <file>`) and a summary
(`summary: N total, S succeeded, F failed`) print to stdout, failure detail
to stderr.
- **`--json`**: exactly one JSON document on stdout in success and failure
(result, exit code, output path, optional embedded markdown for `convert`;
per-file status plus summary for `batch`); human diagnostics stay on stderr.
- **`--dry-run`**: prints the plan (the exact `npx` command line and output
paths) and executes nothing — no CLI spawn, no output files, no directory
creation.
- The wrapper always passes `-y` to npx and never prompts, even on a cold
cache.
+402
View File
@@ -0,0 +1,402 @@
# Formats: what anydoc converts and what GFM you get
This reference documents every input format the pinned CLI (`@firecrawl/anydoc`
v0.1.6) accepts, the GitHub-Flavored Markdown each one produces, and the
fidelity caveats you must know before trusting the output. Every claim below
was verified by running the real CLI against the committed fixtures in
`fixtures/` (see [sources.md](sources.md) for provenance and the verification
procedure).
## Coverage: 8 families / 21 extensions / 12 parsers
| Family | Extensions | Canonical parser |
| --- | --- | --- |
| Word | `.doc`, `.docx`, `.docm` | `doc` (legacy OLE) / `docx` (`.docm` aliases to `docx`) |
| PowerPoint | `.ppt`, `.pps`, `.pot`, `.pptx`, `.pptm`, `.ppsx`, `.ppsm` | `ppt` (`.pps`, `.pot` alias to `ppt`) / `pptx` (`.pptm`, `.ppsx`, `.ppsm` alias to `pptx`) |
| Excel | `.xls`, `.xlsx`, `.xlsm`, `.xlsb` | `xlsx` (all four; calamine reads both OLE and ZIP) |
| OpenDocument | `.odt`, `.ods`, `.odp` | `odt`, `ods`, `odp` |
| Rich Text Format | `.rtf` | `rtf` |
| EPUB | `.epub` | `epub` |
| CSV | `.csv` | `csv` |
| PDF | `.pdf` | `pdf` |
That is **8 families, 21 extensions, 12 canonical parsers**: `doc, docx, odt,
pdf, ppt, pptx, rtf, epub, xlsx, ods, odp, csv`. These 12 names are also the
values accepted by `--format`; extension aliases resolve through the same
mapping (verified: `--format xls` and `--format docm` are accepted).
Format detection reads the file *bytes* first (PDF header, RTF open group, OLE
stream names, ZIP mimetype/content types). CSV has no content signature, so it
falls back to the extension or to an explicit `--format`.
## Shared output behavior
All document formats flow through one shared document model and one GFM
serializer, so identical logical structure yields near-identical Markdown
across formats. Behaviors you can rely on everywhere:
- Headings render as `#``######` with anchors.
- Inline runs preserve **bold**, *italic*, ~~strike~~, `` `code` ``, and lists
(bullet, numbered, nested, roman).
- GFM tables with header rows; merged cells render as **empty covered spans**.
- Footnotes/endnotes: `[^n]` reference inline, with `[^n]: ...` definition
lines at the end of the document.
- Markdown specials in source text are escaped (`\*stars*`, `\| pipe`).
- Embedded images render as their **alt text only** — raw image bytes never
survive into Markdown.
- Bookmarks/anchor targets render as raw `<a id="..."></a>` markers.
## Word (`.doc`, `.docx`, `.docm`)
Expected output: `#` title, `##`/`###` section headings, inline emphasis,
GFM tables, `[^n]` footnotes. DOCX, DOC, ODT, and RTF all share this document
shape; the same fixture converted as `.doc`, `.odt`, and `.rtf` produced
near-identical markdown.
Real conversion of `fixtures/fixture-handmade-outline.docx`:
```markdown
## Style heading stays a heading
### Direct level overrides the style
Direct nine turns the style heading off
# Direct outline without a style
Child style nine stops inheritance
```
Headings come from Word styles and direct formatting; `#``######` levels map
onto heading levels. Real conversion of `fixtures/text.doc` shows the full
document shape:
```markdown
# Fixture Document
Plain paragraph with **bold**, *italic*, and ~~struck~~ runs.
## Table
| | | |
| --- | --- | --- |
| Wide head | | End |
| Tall | B2 | C2 |
| | B3 | C3 |
## Notes and special text
Music clef 𝄞 appears before this footnote[^1] reference.
[^1]: Footnote after an astral character.
```
Caveats:
- **Merged cells** in Word tables render as empty covered spans (the covered
cells are blank, not repeated or filled).
- **Nested tables** flatten into a single cell (GFM cannot nest tables) — a
known limitation of the library.
- Legacy `.doc` (OLE) converts through the same document serializer with the
same shape; only the relative-link target rendering differs cosmetically
between sources.
- Fillable-form controls (DOCX content controls) lose their field layer;
labels and underline glyphs survive.
## PowerPoint (`.ppt`, `.pps`, `.pot`, `.pptx`, `.pptm`, `.ppsx`, `.ppsm`)
Expected output: **slide titles as plain paragraphs** (never markdown
headings), bullet lists, speaker notes as `>` blockquotes, and — for PPTX and
ODP — slide tables as proper GFM tables. Legacy `.ppt` flattens tables to bare
text lines (see caveat).
Real conversion of `fixtures/pres.pptx`:
```markdown
Deck Title Slide
- Top level point
- Nested detail
- Second point with emphasis
> Speaker note for the intro slide.
Numbers Slide
| Region | Total |
| --- | --- |
| North | 42 |
Grouped shapes below.
```
Caveat — **legacy `.ppt` flattens tables to bare text lines.** The same deck
converted from `fixtures/pres.ppt` renders the Numbers Slide table as plain
lines with no `|` table syntax:
```markdown
Numbers Slide
Region
Total
North
42
```
If the presentation's tables matter, use PPTX or ODP and verify the `|` rows
survived (see [workflows.md](workflows.md), "Output verification").
## Excel (`.xls`, `.xlsx`, `.xlsm`, `.xlsb`)
Expected output: each worksheet becomes a `## <sheet name>` heading followed by
a GFM table; the first row is used as the table header when it looks
label-like.
Real conversion of `fixtures/sheet.xlsx` (first table):
```markdown
## Values
| Kind | Value | Note |
| --- | --- | --- |
| Percent | 0.155 | fifteen and a half |
| Currency | 1234.5 | dollars |
| Thousands | 9876543 | grouped |
| Date | 2026-03-15 | ides of March |
| Duration | 26:30:15 | over a day |
| Tiny | 0.0000004 | four ten-millionths |
| Boolean | TRUE | yes |
```
Caveats:
- **XLS/XLSX drop number formats (issue #27).** Cells carry their *raw*
values, not the formatted display values: `Percent → 0.155` (not `15.5%`),
`Currency → 1234.5` (not `$1,234.50`), thousands `9876543`. A percentage
reading as a raw fraction is wrong by 100x in meaning — warn consumers and
sanity-check spreadsheets. Dates survive as ISO strings (`2026-03-15`).
- **ODS is the contrast case:** it keeps the formatted display values
(`15.5%`, `$1,234.50`, `9,876,543`) on the same logical content. If display
values matter, prefer ODS or a CSV export.
- **Merged cells render as empty covered spans** within the populated range
only. Real conversion of `fixtures/handmade-merged.xlsx`:
```markdown
| | | |
| --- | --- | --- |
| Merged across | | padded |
| tall | b2 | 3.5 |
| | b3 | |
```
- Hidden rows and columns are treated as visible and appear in the output
(known limitation) — check for hidden template or calculation content before
feeding output to an LLM.
## OpenDocument (`.odt`, `.ods`, `.odp`)
- `.odt`: same document shape as DOCX/DOC/RTF — `#`/`##` headings, GFM
tables, `[^n]` footnotes. Real conversion of `fixtures/text.odt` matches the
`text.doc` output structure line-for-line (only relative-link targets differ
in depth).
- `.ods`: same spreadsheet shape as XLSX (`## Values` + GFM table) but with
**formatted display values preserved** — the Excel number-format caveat does
not apply. Real conversion of `fixtures/sheet.ods`:
```markdown
## Values
| Kind | Value | Note |
| --- | --- | --- |
| Percent | 15.5% | fifteen and a half |
| Currency | $1,234.50 | dollars |
| Thousands | 9,876,543 | grouped |
```
- `.odp`: **same slide serializer as PPTX** — slide titles as plain
paragraphs, speaker notes as blockquotes, and GFM tables **kept** (unlike
legacy `.ppt`). Real conversion of `fixtures/pres.odp`:
```markdown
Deck Title Slide
- Top level point
- - Nested detail
- Second point with emphasis
> Speaker note for the intro slide.
Numbers Slide
| Region | Total |
| --- | --- |
| North | 42 |
```
One cosmetic difference vs PPTX: a nested bullet renders as `- - Nested
detail` on one line rather than as an indented sub-list. The table,
blockquote notes, and paragraph titles are identical in shape to PPTX.
## Rich Text Format (`.rtf`)
Expected output: the same document shape as DOCX/ODT (`# Fixture Document`,
`##` sections, GFM tables, `[^n]` footnote definitions). Real conversion of
`fixtures/text.rtf` matches `text.odt` structure; the one notable difference is
that relative link targets render with a `file:///` absolute path, e.g.
`[a sibling file](file:///anydoc/tests/fixture-src/sibling.odt)`, instead of a
relative path — a known cosmetic quirk.
## EPUB (`.epub`)
Expected output: `#` chapter headings (plus the book metadata title), GFM
tables, preserved inline emphasis/code, and **internal anchor links resolved to
fragments**. Real conversion of `fixtures/book.epub`:
```markdown
# Fixture Book
# Fixture Book
anydoc tests
<a id="epub-text-ch001-xhtml-chapter-one"></a>
# Chapter One
Opening paragraph with **bold**, *italic*, and `code` runs.
See [Chapter Two](#epub-text-ch002-xhtml-chapter-two) for the table, or jump straight to [the marked paragraph](#epub-text-ch002-xhtml-markpoint).
<a id="epub-text-ch002-xhtml-chapter-two"></a>
# Chapter Two
| Name | Qty |
| --- | --- |
| Bolts | 12 |
| Nuts | 30 |
```
Notes: the book title may appear twice (metadata title + injected title);
internal links keep working as `[text](#fragment)` links; external links stay
as normal markdown links.
## CSV (`.csv`)
Expected output: the file renders as **one GFM table**. The first row is
**promoted to the header row** when it looks like labels (≥ 2 columns,
non-empty, non-numeric, distinct fields) — this behavior ships in 0.1.6.
Quoted fields with embedded commas and newlines are preserved.
Real conversion of `fixtures/fixture-handmade-quoted.csv`:
```markdown
| name | desc | qty |
| --- | --- | --- |
| padded | comma, inside | 3 |
| plain | multi line | 4 |
```
Also verified:
- **Delimiter sniffing** — a semicolon-delimited file with decimal commas
splits on `;` and keeps `1,5` intact (real output of
`fixtures/fixture-handmade-semicolon.csv`):
```markdown
| a | b | c |
| --- | --- | --- |
| 1,5 | 2,5 | x |
| 3,0 | y | z |
```
- **UTF-16 (with BOM)** decodes to correct Unicode (real output of
`fixtures/fixture-handmade-utf16.csv`):
```markdown
| col1 | col2 |
| --- | --- |
| naïve | café |
| Αθήνα | 数据 |
```
CSV has no content signature, so **`--format csv` is required when reading CSV
from stdin** (see [cli-reference.md](cli-reference.md)).
## PDF (`.pdf`) — the lower-fidelity pipeline
Text-based PDFs convert **locally** through a separate pipeline (`pdf-inspector`)
that emits Markdown directly — PDF has no document model, so only Markdown
output exists. Real conversion of `fixtures/fixture-text.pdf`:
```markdown
# Fixture Document
Plain paragraph with **bold**, *italic*, and struck runs. **Style-bold paragraph with a** NotBold-styled span **inside.**
## Lists
1.First numbered
2.Second numbered
a)Alpha sub one
b)Alpha sub two
i.Roman sub sub
3.Third numbered Interrupting paragraph between lists.
## Table
Wide head End Tall B2 C2 B3 C3
```
**Fidelity caveats (verified on the real output):**
- **No GFM tables.** Table cell text flattens into a plain paragraph run
(`Wide head End Tall B2 C2 B3 C3`) — there is no `|` table.
- **No `[^n]` footnotes.** Footnote markers degrade to inline superscript
glyphs (`¹`) and the note bodies drop into the flow; there is no `[^1]:`
definition block.
- **Links are not emitted as markdown links.** They degrade to `<u>underlined
text</u>`.
- Numbered/bulleted list structure compresses (markers inline), and some
Unicode degrades (e.g. emoji without ZWJ).
### Scanned or image-only PDFs — no OCR
A PDF with **no extractable text layer** fails as `unsupported` with this exact
message (exit code 1):
```
anydoc: unsupported input: PDF has no extractable text (Scanned, 1 pages): OCR is required
```
anydoc **does not perform OCR** — the library's stance is explicit, and there
is no password, retry, or OCR option. When this message fires: report the exact
error, state that OCR is required, and route the file to OCR tooling or the
hosted Firecrawl Parse API. Do not retry the same file locally and do not claim
anydoc can OCR it. See [errors.md](errors.md) for the full routing guidance.
## Formats anydoc does NOT support
- HTML/SingleFile (open feature request only) — not an input format.
- Images (`.png`, `.jpg`, ...) — no image-to-text conversion.
- Password-protected/encrypted documents — fail with
`anydoc: document is encrypted` (see [errors.md](errors.md)).
- Anything without a recognized signature and extension — fails as
`unsupported input: unrecognized file content and extension: <path>`.
## Output-shape invariants to remember
1. One serializer: the same logical structure yields near-identical Markdown
across docx/odt/rtf — do not re-test each office format for the same
feature.
2. Spreadsheets: expect `## <sheet name>` + GFM tables; warn that xlsx/xls
drop number formats (issue #27) while ODS keeps display values.
3. Legacy `.ppt` and all PDFs lose tabular structure — add a
"verify the table survived" step or use PPTX/ODP and text PDFs.
4. Images never survive as bytes in Markdown — only alt text.
+97
View File
@@ -0,0 +1,97 @@
# Sources, provenance, and verification
## Upstream project
| Resource | URL / identifier |
| --- | --- |
| Repository | https://github.com/firecrawl/anydoc |
| npm package | `@firecrawl/anydoc` — https://www.npmjs.com/package/@firecrawl/anydoc |
| PyPI package | `firecrawl-anydoc` (imports as `anydoc`) — https://pypi.org/project/firecrawl-anydoc/ |
| crates.io crate | `anydoc` (same release train) |
| Browser demo (WASM) | https://firecrawl.github.io/anydoc/ |
| License | MIT |
## Access and verification dates
- Research and empirical verification performed **2026-08-05** and **2026-08-06**
on macOS (arm64) with Node v22.22.3, network access, and the pinned CLI
`npx -y @firecrawl/anydoc@0.1.6`.
- The pinned release **0.1.6** was published to npm at
**2026-08-05T18:29:40Z**; PyPI wheels for the same version were uploaded
**2026-08-05T18:29Z**. First release was 0.1.1 (2026-08-04).
## Fixture provenance
The committed fixtures under `fixtures/` come from two sources, both documented
here per the repository's attribution policy:
1. **The MIT-licensed upstream test suite.** Most fixtures were downloaded from
`https://github.com/firecrawl/anydoc/tree/main/tests/fixtures` (raw files
via `https://raw.githubusercontent.com/firecrawl/anydoc/main/tests/fixtures/...`).
They retain the upstream naming and structure:
- CSV: `fixture-handmade-quoted.csv`, `fixture-handmade-semicolon.csv`,
`fixture-handmade-utf16.csv`, `fixture-sheet.csv`
- DOCX: `fixture-handmade-numbering.docx`, `fixture-handmade-outline.docx`,
`fixture-handmade-rich.docx`, `fixture-handmade-tables.docx`
- Word legacy: `text.doc`; OpenDocument: `text.odt`, `sheet.ods`, `pres.odp`
- RTF: `text.rtf`; EPUB: `book.epub`
- PowerPoint: `pres.ppt`, `pres.pptx`; Excel: `sheet.xls`, `sheet.xlsx`,
`handmade-merged.xlsx`
- PDF: `fixture-text.pdf`
- Error cases from the upstream `*--errors.*` corpus:
`empty--errors.docx`, `encrypted--errors.odt`
2. **Generated samples** (created during research for cases the upstream suite
does not cover; deterministic, reproducible):
- `scanned-image-only.pdf` — a PDF with a single grayscale image and no
text layer, generated with Pillow, to exercise the no-OCR error path.
- `unsupported.xyz` — a small text file with an unsupported extension, to
exercise the unrecognized-content error path.
All fixtures are tiny (largest: `pres.ppt` at ~454 KB) and each is well under
the 5 MB repository limit. All committed copies are byte-identical to the
staged originals used during research (verified by sha256).
MIT license notice: the upstream anydoc project is MIT-licensed (Copyright
Firecrawl); the fixture files above are used under that license. The generated
samples carry no upstream copyright.
## Verification procedure
Every factual claim in this skill was confirmed against the **real pinned CLI**
(v0.1.6), not inferred from documentation:
1. **Environment warm-up**: `node --version` (v22.22.3 ≥ 20), then
`npx -y @firecrawl/anydoc@0.1.6 --version` → prints `0.1.6`; `--help`
the verbatim help block reproduced in
[cli-reference.md](cli-reference.md).
2. **Positive conversions**: the pinned CLI was run on every committed fixture
with stdout and stderr captured separately and the exit code recorded. All
20 positive fixtures converted with exit 0 and empty stderr; the captured
markdown was compared against the output expectations documented in
[formats.md](formats.md) (headings, table rows, slide structure, footnote
definitions, CSV header promotion, UTF-16/delimiter handling, merged-cell
covered spans).
3. **Error paths**: each error fixture and each usage error was run with
stderr captured verbatim and the exit code recorded (1 for conversion/IO
failures, 2 for usage errors). The exact messages appear in
[errors.md](errors.md) character-for-character, including
`anydoc: unsupported input: PDF has no extractable text (Scanned, 1 pages): OCR is required`.
4. **Special behaviors**: `-o` overwrite and EISDIR, stdin via `-` with and
without `--format csv`, `--format=x` inline syntax, `--` end-of-options,
extension aliases (`--format xls`, `--format docm`), EPIPE (`| head` exits
0 with empty stderr), the stdin-is-a-terminal usage error (via a
pseudo-TTY), and resource limits (run on the upstream `zipbomb`/`imagebomb`
fixtures and on a generated 250 MB-entry archive — all exit 1 with the
documented `max_entry_bytes` prefix).
5. **First-run/offline**: a fresh empty npm cache was used to verify the
first-run download path (`env npm_config_cache=$(mktemp -d) npx -y
@firecrawl/anydoc@0.1.6 --version``0.1.6`, exit 0).
6. **Startup timing**: repeated warm invocations were timed
(`/usr/bin/time -p npx -y @firecrawl/anydoc@0.1.6 ...`) — ~0.320.35 s each,
consistent with the documented ~0.330.55 s warm-cache startup range.
Repository checks applied after authoring: frontmatter and structure
(`ruby scripts/validate-skills.rb`), skill quality (`ruby
scripts/validate-skill-quality.rb --base origin/main`), reference caps and link
resolution, eval-manifest validation (`scripts/validate-evals.py`), and no
machine-specific paths or credentials in any committed file.
+190
View File
@@ -0,0 +1,190 @@
# Workflows: recipes for converting documents to markdown
All recipes use the pinned CLI `npx -y @firecrawl/anydoc@0.1.6` (ground truth)
and the skill's wrapper `scripts/anydoc` where it adds value. Commands are
shown relative to the repository root; `anydoc/fixtures/...` paths can be
replaced with any document path. The vault-ingestion recipe (section 5) is
written to be run from a temp or vault directory holding *your own*
documents. Each raw-CLI invocation converts **exactly one document** — there
is no batch mode.
## 1. Single conversion
```bash
# Markdown to stdout
npx -y @firecrawl/anydoc@0.1.6 anydoc/fixtures/fixture-handmade-outline.docx
# Markdown to a file (stdout stays silent; existing file is overwritten)
npx -y @firecrawl/anydoc@0.1.6 anydoc/fixtures/fixture-handmade-outline.docx -o outline.md
# Same jobs through the wrapper
python3 anydoc/scripts/anydoc convert anydoc/fixtures/fixture-handmade-outline.docx
python3 anydoc/scripts/anydoc convert anydoc/fixtures/fixture-handmade-outline.docx -o outline.md
```
Expected result: exit code 0, empty stderr, and GitHub-Flavored Markdown on
stdout (or written to the `-o` output file) containing `#`/`##`/`###` heading
lines.
## 2. Force the input format
```bash
# Extensionless or mislabeled file: name the format explicitly
npx -y @firecrawl/anydoc@0.1.6 ./data --format csv
npx -y @firecrawl/anydoc@0.1.6 ./report --format docx
```
Use `--format <name>` only when detection cannot work (CSV from stdin, or a
missing/wrong extension). Aliases resolve: `--format xls`, `--format docm`,
`--format ppsx` are accepted. An invalid name exits 2 with
`anydoc: invalid format 'bogus'; expected one of: ...`.
## 3. Read a document from stdin
```bash
# CSV from stdin requires --format csv (no signature, no extension)
printf 'name,role\nAlice,Engineer\n' | npx -y @firecrawl/anydoc@0.1.6 - --format csv
# Any document type can come from stdin; detection reads the bytes
curl -s https://example.com/paper.pdf | npx -y @firecrawl/anydoc@0.1.6 -
```
The wrapper supports the same: `cat data.csv | python3 anydoc/scripts/anydoc convert - -f csv`.
Piping notes:
- Markdown goes to **stdout only**; diagnostics are the single
`anydoc: <message>` stderr line.
- **EPIPE is handled**: if the downstream pipe closes early
(`... anydoc@0.1.6 big.xlsx | head -n 1`), the CLI exits 0 with no stderr
noise — piping into `head` is safe and is not a failure.
## 4. Batch conversion (raw CLI)
The raw CLI takes one document per invocation, so batch with a shell loop:
```bash
mkdir -p out
for f in anydoc/fixtures/*.docx; do
npx -y @firecrawl/anydoc@0.1.6 "$f" -o "out/$(basename "${f%.docx}").md"
done
```
Each failed document (error fixtures, scanned PDFs, encrypted files) exits 1
with its `anydoc: <message>` on stderr and produces no output file; the loop
continues with the next input. Handle or route those per
[errors.md](errors.md).
Or the wrapper, which is built for this (per-file status, continues past
failures, summary, and a non-zero exit when any input failed):
```bash
python3 anydoc/scripts/anydoc batch \
anydoc/fixtures/fixture-handmade-outline.docx \
anydoc/fixtures/fixture-sheet.csv \
--out-dir out/
```
`batch --dry-run --json` prints the plan (input → output, dry-run marker)
without converting or creating anything:
```bash
python3 anydoc/scripts/anydoc batch anydoc/fixtures/fixture-handmade-outline.docx \
anydoc/fixtures/fixture-sheet.csv --out-dir out/ --dry-run --json
```
## 5. Vault-ingestion pattern
Convert a folder of mixed office documents to markdown for ingestion into a
vault or knowledge base:
1. **Collect** the documents into a folder (mixed docx/xlsx/pptx/csv/odt/pdf
is fine — text-based PDFs only; see the no-OCR caveat in
[errors.md](errors.md)).
2. **Batch-convert** with the wrapper into a markdown folder:
```bash
python3 anydoc/scripts/anydoc batch notes/*.docx notes/*.xlsx notes/*.csv --out-dir vault/inbox/
```
(or the raw-CLI loop above if you are not using the wrapper).
> **Run this from a temp or vault directory — never from the agent-skills
> repo root.** The glob matches whatever directory you name, and the
> repository tracks a top-level `docs/` directory (distinct from the
> `documents/` skill): globbing `docs/*.docx` there, or deleting/cleaning
> those matches, would damage tracked repository files. Keep the source
> documents in their own folder (here `notes/`) and convert into a
> separate `vault/inbox/` folder.
3. **Verify each output** (step 6) — at minimum confirm exit 0 and that the
structural markers your formats produce are present (headings for Word/PDF,
`|` tables for spreadsheets/CSV).
4. **Failures are per-file**: the batch summary names what failed; route those
files per [errors.md](errors.md) (scanned PDF → OCR tooling, encrypted →
unencrypted copy, unsupported → check extension) and re-run only the
failures.
## 6. Output verification
Before treating a conversion as done:
1. **Exit code 0** — the CLI produced markdown. Exit 1: read the
`anydoc: <message>` stderr line and match it against
[errors.md](errors.md). Exit 2: fix the command (usage error).
2. **Structural markers** — check the markers your format actually produces:
- Word / ODT / RTF / text-based PDF: `#`/`##` heading lines
(`grep -E '^#{1,6} ' out.md`).
- Spreadsheets (xlsx/xls/ods) and CSV: `## <sheet>` headings and
`|`-delimited rows (`grep -E '^\|' out.md`).
- Presentations (pptx/odp): slide titles as plain paragraphs, `>`
blockquote speaker notes, `|` table rows (legacy `.ppt` has no `|` rows —
that is by design, not an error).
- EPUB: `#` chapter headings and `[text](#fragment)` internal links.
3. **Tables survived?** If the source had tables and the output has no `|`
rows, check the caveats: PDF and legacy `.ppt` flatten tables by design.
4. **Large outputs**: convert with `-o out.md` and inspect the file rather
than streaming everything into context.
Use the committed fixtures to sanity-check an environment once:
```bash
npx -y @firecrawl/anydoc@0.1.6 anydoc/fixtures/fixture-handmade-outline.docx # headings
npx -y @firecrawl/anydoc@0.1.6 anydoc/fixtures/sheet.xlsx # ## Values + table
npx -y @firecrawl/anydoc@0.1.6 anydoc/fixtures/fixture-text.pdf # headings, no table
```
## 7. Large files and resource limits
- **Conversion is not streaming** — the document is read and processed as a
whole, and safety limits protect against decompression and nesting bombs.
- **Zip/image bombs are rejected via `max_entry_bytes`** with exit 1 and the
prefix `anydoc: resource limit exceeded (max_entry_bytes):` (full examples
in [errors.md](errors.md)). This is by design — do not try to bypass it.
- **`-o out.md` is recommended for large documents** so the output is written
to a reviewable file instead of filling stdout/context; you can then read
the parts you need.
- Genuinely large real documents (as opposed to bombs) convert normally; the
per-document limit only rejects entries whose declared decompressed size
exceeds the cap.
- If a resource-limit error fires on a *real* file, the archive is malformed
or hostile — re-export the document rather than disabling the limit.
## 8. Startup cost and performance
Each `npx -y @firecrawl/anydoc@0.1.6` invocation costs roughly **0.330.55 s
of warm-cache startup** (npm/npx process startup) on top of the conversion
itself, which is a few milliseconds (measured ~5 ms for a PDF, <1 ms for a
DOCX once the process is warm). There is no progress output; conversions are
effectively instant. Plan for ~0.5 s per document in batch loops, and prefer a
single `npx` process per document (you cannot batch inside one invocation).
## 9. Offline / cold-cache behavior
- The first `npx` run downloads the package plus the native binary (network
required once); later runs use the npm cache. A cold-cache offline run fails
with a clear npx fetch error before anydoc executes.
- For permanent or fully offline use, install once:
`npm install -g @firecrawl/anydoc`, then call `anydoc <file>` directly.
- The wrapper always invokes npx with `-y` (non-interactive), so it never
hangs on npx's install prompt — even on a cold cache it fails fast if the
package cannot be fetched.
+672
View File
@@ -0,0 +1,672 @@
#!/usr/bin/env python3
"""anydoc — convert office documents to GitHub-Flavored Markdown locally.
Python 3 standard-library wrapper around the pinned Any Doc CLI
(`npx -y @firecrawl/anydoc@0.1.6`). Adds input pre-validation, friendly
hints for the known failure classes, batch conversion, and `--json` /
`--dry-run` output on top of the raw CLI.
Commands:
convert <file|-> [-o out.md] [-f <format>] [--json] [--dry-run]
batch <inputs...> [--out-dir DIR] [--json] [--dry-run]
info [--version]
Exit codes: 0 success / 1 conversion or pre-validation failure / 2 usage error.
With `--json`, exactly one JSON document goes to stdout; all diagnostics go to
stderr. The wrapper never prompts: npx is always invoked with `-y`.
"""
import argparse
import json
import os
import re
import shutil
import signal
import subprocess
import sys
from pathlib import Path
TOOL_NAME = "anydoc"
CLI_PACKAGE = "@firecrawl/anydoc"
CLI_VERSION = "0.1.6"
PINNED = "%s@%s" % (CLI_PACKAGE, CLI_VERSION)
NODE_MIN_MAJOR = 20
RUN_TIMEOUT = 120
# The 12 canonical parsers plus the 9 extension aliases the CLI accepts for -f.
CANONICAL_FORMATS = "doc docx odt pdf ppt pptx rtf epub xlsx ods odp csv".split()
FORMAT_ALIASES = "docm xls xlsm xlsb pps pot pptm ppsx ppsm".split()
VALID_FORMATS = frozenset(CANONICAL_FORMATS + FORMAT_ALIASES)
FORMATS_LIST_TEXT = ", ".join(CANONICAL_FORMATS)
def fail(message, code=1):
"""Print a diagnostics line to stderr and exit with `code`."""
print("anydoc: " + message, file=sys.stderr)
raise SystemExit(code)
def emit_json(obj):
print(json.dumps(obj, ensure_ascii=False))
def fail_command(args, command, message, code, error_class=None):
"""Record a wrapper-level failure: diagnostics to stderr, JSON when asked."""
print("anydoc: " + message, file=sys.stderr)
if getattr(args, "json", False):
emit_json(
{
"tool": TOOL_NAME,
"command": command,
"ok": False,
"exit_code": code,
"error": "anydoc: " + message,
"hint": None,
"error_class": error_class,
"output": getattr(args, "output", None),
}
)
return code
# --- pre-validation ---------------------------------------------------------
def validate_input(path):
"""Return an error message for an unusable input path, or None."""
if path == "-":
return None
target = Path(path)
if not target.exists():
return "input file not found: %s" % path
if target.is_dir():
return "input path is a directory, not a file: %s" % path
return None
def validate_output(path):
"""Return an error message for an unusable -o path, or None."""
if path is None:
return None
if Path(path).is_dir():
return (
"output path is a directory: %s (pass a file path; -o does not "
"create directories)" % path
)
return None
def validate_format(fmt):
"""Return an error message for an invalid -f value, or None."""
if fmt is not None and fmt not in VALID_FORMATS:
return (
"invalid format '%s'; expected one of: %s (extension aliases like "
"xls, docm, ppsx are accepted)" % (fmt, FORMATS_LIST_TEXT)
)
return None
# --- runtime checks (node >= 20, npx) ---------------------------------------
def node_version():
"""Return (major, full_version) for the node on PATH, or None."""
node = shutil.which("node")
if node is None:
return None
try:
proc = subprocess.run(
[node, "--version"], capture_output=True, text=True, timeout=10
)
except OSError:
return None
if proc.returncode != 0:
return None
version = (proc.stdout or "").strip()
match = re.match(r"^v?(\d+)\.", version)
if match is None:
return None
return int(match.group(1)), version
def runtime_errors():
"""Return a list of node/npx problems, empty when the environment is ready."""
errors = []
info = node_version()
if info is None:
errors.append(
"Node.js >= %d is required but `node` was not found on PATH "
"(conversion runs via npx -y %s). Install Node.js >= %d and retry."
% (NODE_MIN_MAJOR, PINNED, NODE_MIN_MAJOR)
)
else:
major, version = info
if major < NODE_MIN_MAJOR:
errors.append(
"Node.js version %s is too old; anydoc requires Node.js >= %d "
"(conversion runs via npx -y %s). Upgrade Node.js and retry."
% (version, NODE_MIN_MAJOR, PINNED)
)
if shutil.which("npx") is None:
errors.append(
"`npx` was not found on PATH — conversion runs via `npx -y %s`. "
"Install Node.js >= %d (which ships npx), or install the CLI "
"permanently with `npm install -g %s`."
% (PINNED, NODE_MIN_MAJOR, CLI_PACKAGE)
)
return errors
# --- CLI invocation ---------------------------------------------------------
def build_cli_command(file, out, fmt):
"""Build the exact argv the wrapper passes to the pinned CLI.
Dash-leading filenames (a path like `-weird`) need the CLI's `--`
end-of-options marker so the file is read as the positional input. npx
forwards `--` to the CLI verbatim, so `-o`/`-f` must be placed BEFORE the
separator — after it they read as extra positional inputs ("unexpected
second input"). Absolute paths never start with `-` and take the plain
form.
"""
options = []
if out:
options.extend(["-o", out])
if fmt:
options.extend(["-f", fmt])
argv = ["npx", "-y", PINNED]
if file != "-" and file.startswith("-"):
argv.extend(options)
argv.extend(["--", file])
else:
argv.append(file)
argv.extend(options)
return argv
class CliTimeoutError(Exception):
"""Raised when the pinned CLI does not complete within RUN_TIMEOUT seconds."""
def __init__(self, message):
super().__init__(message)
self.message = message
def run_cli(argv):
try:
return subprocess.run(
argv,
capture_output=True,
text=True,
encoding="utf-8",
errors="replace",
timeout=RUN_TIMEOUT,
start_new_session=True,
)
except subprocess.TimeoutExpired as exc:
try:
os.killpg(exc.pid, signal.SIGKILL)
except (ProcessLookupError, PermissionError):
pass
raise CliTimeoutError(
"conversion did not complete within %d seconds" % RUN_TIMEOUT
) from exc
def error_class_hint(stderr):
"""Map a failure message (CLI stderr or a wrapper pre-validation message)
to (error_class, hint-or-None)."""
if "document is encrypted" in stderr:
return (
"encrypted",
"the document is encrypted or password-protected — supply an "
"unencrypted copy; anydoc has no password or decryption option.",
)
if "OCR is required" in stderr or "no extractable text" in stderr:
return (
"no-ocr",
"scanned or image-only PDF — anydoc does not perform OCR. Route "
"the file to OCR tooling or the hosted Firecrawl Parse API; do "
"not retry it locally.",
)
if "malformed document" in stderr:
return (
"malformed",
"the document is malformed or corrupt (not a readable zip "
"archive) — re-export or re-download the file and retry.",
)
if "resource limit exceeded" in stderr:
return ("resource-limit", None)
if "EISDIR" in stderr:
return (
"output-dir",
"the output path is a directory — pass a file path (-o does not "
"create directories).",
)
if "unsupported input" in stderr:
return (
"unsupported",
"unsupported or unrecognized file type — check that the extension "
"is one of the supported formats, or force it with -f <format>.",
)
if "io error" in stderr:
return (
"io",
"the file could not be read — check that the path exists and is "
"readable.",
)
# Wrapper pre-validation messages (batch per-file entries).
if "input file not found" in stderr or "input path is a directory" in stderr:
return ("io", None)
return ("unknown", None)
def print_cli_error(stderr):
if not stderr:
stderr = "conversion failed"
line = stderr if stderr.startswith("anydoc: ") else "anydoc: " + stderr
print(line, file=sys.stderr)
# --- subcommands ------------------------------------------------------------
def cmd_convert(args):
error = validate_input(args.file)
if error is not None:
return fail_command(args, "convert", error, 1)
error = validate_output(args.output)
if error is not None:
return fail_command(args, "convert", error, 1)
error = validate_format(args.format)
if error is not None:
return fail_command(args, "convert", error, 2)
command = build_cli_command(args.file, args.output, args.format)
if args.dry_run:
plan = " ".join(command)
if args.json:
emit_json(
{
"tool": TOOL_NAME,
"command": "convert",
"ok": True,
"exit_code": 0,
"dry_run": True,
"input": args.file,
"output": args.output,
"command_line": plan,
}
)
else:
print("dry-run: would run: %s" % plan)
return 0
for error in runtime_errors():
return fail_command(args, "convert", error, 1)
try:
proc = run_cli(command)
except CliTimeoutError as err:
return fail_command(args, "convert", err.message, 1, error_class="timeout")
if proc.returncode == 0:
if args.json:
emit_json(
{
"tool": TOOL_NAME,
"command": "convert",
"ok": True,
"exit_code": 0,
"output": args.output,
"markdown": None if args.output else proc.stdout,
}
)
elif not args.output:
sys.stdout.write(proc.stdout)
return 0
stderr = (proc.stderr or "").strip()
error_class, hint = error_class_hint(stderr)
print_cli_error(stderr)
if hint is not None:
print("anydoc: hint: " + hint, file=sys.stderr)
if args.json:
emit_json(
{
"tool": TOOL_NAME,
"command": "convert",
"ok": False,
"exit_code": proc.returncode,
"error": stderr,
"hint": hint,
"error_class": error_class,
"output": args.output,
}
)
return proc.returncode
def cmd_batch(args):
out_dir = Path(args.out_dir) if args.out_dir else Path.cwd()
records = []
for input_path in args.inputs:
error = validate_input(input_path)
records.append(
{
"input": input_path,
"output": str(out_dir / (Path(input_path).stem + ".md")),
"error": error,
}
)
if args.dry_run:
plan = []
for record in records:
plan.append(
{
"input": record["input"],
"output": record["output"],
"command": " ".join(
build_cli_command(record["input"], record["output"], None)
),
"would_fail": record["error"] is not None,
"error": record["error"],
}
)
if args.json:
emit_json(
{
"tool": TOOL_NAME,
"command": "batch",
"ok": True,
"exit_code": 0,
"dry_run": True,
"out_dir": str(out_dir),
"plan": plan,
}
)
else:
for entry in plan:
marker = (
" [would fail: %s]" % entry["error"] if entry["error"] else ""
)
print(
"plan: convert %s -> %s%s"
% (entry["input"], entry["output"], marker)
)
print(" command: %s" % entry["command"])
return 0
if out_dir.exists() and not out_dir.is_dir():
return fail_command(
args,
"batch",
"output directory path is not a directory: %s" % out_dir,
1,
)
for error in runtime_errors():
return fail_command(args, "batch", error, 1)
out_dir.mkdir(parents=True, exist_ok=True)
results = []
for record in records:
if record["error"] is not None:
error_class, hint = error_class_hint(record["error"])
results.append(
{
"input": record["input"],
"output": record["output"],
"status": "failed",
"error": "anydoc: " + record["error"],
"hint": hint,
"error_class": error_class,
}
)
print("anydoc: " + record["error"], file=sys.stderr)
if not args.json:
print("FAIL %s" % record["input"])
continue
try:
proc = run_cli(build_cli_command(record["input"], record["output"], None))
except CliTimeoutError as err:
return fail_command(args, "batch", err.message, 1, error_class="timeout")
if proc.returncode == 0:
results.append(
{
"input": record["input"],
"output": record["output"],
"status": "ok",
"error": None,
"hint": None,
}
)
if not args.json:
print("ok %s -> %s" % (record["input"], record["output"]))
else:
stderr = (proc.stderr or "").strip()
error_class, hint = error_class_hint(stderr)
results.append(
{
"input": record["input"],
"output": record["output"],
"status": "failed",
"error": stderr,
"hint": hint,
"error_class": error_class,
}
)
print_cli_error(stderr)
if hint is not None:
print("anydoc: hint: " + hint, file=sys.stderr)
if not args.json:
print("FAIL %s" % record["input"])
succeeded = sum(1 for r in results if r["status"] == "ok")
failed = len(results) - succeeded
if not args.json:
print(
"summary: %d total, %d succeeded, %d failed"
% (len(results), succeeded, failed)
)
if args.json:
emit_json(
{
"tool": TOOL_NAME,
"command": "batch",
"ok": failed == 0,
"exit_code": 1 if failed else 0,
"dry_run": False,
"out_dir": str(out_dir),
"files": results,
"summary": {
"total": len(results),
"succeeded": succeeded,
"failed": failed,
},
}
)
return 1 if failed else 0
def cmd_info(args):
if args.version:
print(CLI_VERSION)
return 0
if args.json:
emit_json(
{
"tool": TOOL_NAME,
"command": "info",
"ok": True,
"exit_code": 0,
"name": TOOL_NAME,
"cli": PINNED,
"version": CLI_VERSION,
}
)
else:
print("%s %s (wraps %s)" % (TOOL_NAME, CLI_VERSION, PINNED))
return 0
# --- CLI plumbing -----------------------------------------------------------
def extract_globals(argv):
"""Hoist --json / --dry-run to the front so they work anywhere in argv."""
values, retained = [], []
for arg in argv:
if arg in ("--json", "--dry-run"):
values.append(arg)
else:
retained.append(arg)
return values + retained
def build_parser():
parser = argparse.ArgumentParser(
prog="anydoc",
description=(
"Convert office documents to GitHub-Flavored Markdown locally via "
"the pinned Any Doc CLI (%s)." % PINNED
),
epilog=(
"Examples:\n"
" anydoc convert report.docx\n"
" anydoc convert report.docx -o report.md\n"
" anydoc convert - -f csv < data.csv\n"
" anydoc batch a.docx b.csv --out-dir out/\n"
" anydoc info\n"
"\nExit codes: 0 success / 1 conversion or pre-validation failure "
"/ 2 usage error. JSON goes to stdout; diagnostics go to stderr. "
"The wrapper never prompts (npx runs with -y)."
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument(
"--json",
action="store_true",
help="Emit exactly one JSON document on stdout; diagnostics stay on stderr.",
)
parser.add_argument(
"--dry-run",
action="store_true",
help="Print what would run and execute nothing (no CLI spawn, no output files).",
)
subparsers = parser.add_subparsers(dest="command", required=True)
convert = subparsers.add_parser(
"convert",
help="Convert one document (path or - for stdin) to markdown.",
description=(
"Convert one document to GitHub-Flavored Markdown. Pass - as the "
"input to read the document from stdin (CSV from stdin needs -f csv)."
),
epilog=(
"Examples:\n"
" anydoc convert report.docx\n"
" anydoc convert report.docx -o report.md\n"
" anydoc convert - -f csv < data.csv\n"
"\nWith --json the converted markdown is embedded in the JSON "
"document when -o is not given. Exit codes: 0 success / 1 "
"conversion or pre-validation failure / 2 usage error."
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
convert.add_argument(
"file", metavar="<file|->", help="Document path, or - to read from stdin."
)
convert.add_argument(
"-o",
"--output",
metavar="out.md",
help="Write the markdown to this file instead of stdout (silently overwrites).",
)
convert.add_argument(
"-f",
"--format",
metavar="<format>",
help=(
"Force the input format instead of detecting it: one of %s "
"(extension aliases like xls, docm, ppsx are accepted)."
% FORMATS_LIST_TEXT
),
)
batch = subparsers.add_parser(
"batch",
help="Convert many documents, one at a time, to an output directory.",
description=(
"Convert many documents to markdown files. Each input is converted "
"independently; the batch continues past failures and prints a "
"summary. Output naming is deterministic: <stem>.md in --out-dir "
"(created when missing; default: the current working directory)."
),
epilog=(
"Examples:\n"
" anydoc batch a.docx b.xlsx c.csv --out-dir out/\n"
" anydoc batch notes/*.docx --out-dir vault/inbox/ --dry-run --json\n"
"\nDuplicate inputs convert per occurrence (a later conversion "
"overwrites the earlier output); same-basename inputs from "
"different directories collide on the same <stem>.md and the last "
"one wins. Exit codes: 0 when every input converted; 1 when any "
"input failed; 2 for a usage error."
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
batch.add_argument(
"inputs", nargs="+", metavar="<input>", help="Document paths to convert."
)
batch.add_argument(
"--out-dir",
metavar="DIR",
help="Directory for the converted .md files (created when missing; default: the current directory).",
)
info = subparsers.add_parser(
"info",
help="Show the tool name and the pinned CLI version.",
description="Report the wrapper tool name and the pinned Any Doc CLI version.",
epilog=(
"Examples:\n"
" anydoc info\n"
" anydoc info --version\n"
"\n--version prints exactly the pinned CLI version (0.1.6)."
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
info.add_argument(
"--version",
action="store_true",
help="Print only the pinned CLI version and exit.",
)
return parser
def main(argv=None):
parser = build_parser()
args = parser.parse_args(
extract_globals(list(argv if argv is not None else sys.argv[1:]))
)
if args.command == "convert":
return cmd_convert(args)
if args.command == "batch":
return cmd_batch(args)
if args.command == "info":
return cmd_info(args)
parser.print_help()
return 0
if __name__ == "__main__":
try:
code = main()
except BrokenPipeError:
code = 0
try:
sys.exit(code)
except BrokenPipeError:
# Downstream pipe closed early (e.g. `anydoc convert big.docx | head`)
# — exit 0, mirroring the CLI's EPIPE behavior, without stderr noise.
devnull = os.open(os.devnull, os.O_WRONLY)
os.dup2(devnull, sys.stdout.fileno())
sys.exit(0)
+810
View File
@@ -0,0 +1,810 @@
#!/usr/bin/env python3
"""Unit tests for the anydoc wrapper (`anydoc/scripts/anydoc`).
Offline by design: the core tests need no node, no npx, and no network.
Real-CLI tests (converting the committed fixtures through the pinned CLI)
are opt-in and skip gracefully when the toolchain is unavailable.
"""
import importlib.machinery
import io
import json
import os
import shutil
import signal
import stat
import subprocess
import sys
import tempfile
import unittest
from contextlib import redirect_stderr, redirect_stdout
from pathlib import Path
from unittest import mock
ROOT = Path(__file__).resolve().parents[1] # anydoc/
SCRIPT = ROOT / "scripts" / "anydoc"
FIXTURES = ROOT / "fixtures"
DOCX = FIXTURES / "fixture-handmade-outline.docx"
CSV = FIXTURES / "fixture-sheet.csv"
SCANNED = FIXTURES / "scanned-image-only.pdf"
ENCRYPTED = FIXTURES / "encrypted--errors.odt"
MALFORMED = FIXTURES / "empty--errors.docx"
UNSUPPORTED = FIXTURES / "unsupported.xyz"
TABLES = FIXTURES / "fixture-handmade-tables.docx"
PINNED = "@firecrawl/anydoc@0.1.6"
cli = importlib.machinery.SourceFileLoader("anydoc_wrapper", str(SCRIPT)).load_module()
def run_in_process(arguments):
"""Run cli.main() in-process; return (code, stdout, stderr)."""
stdout, stderr = io.StringIO(), io.StringIO()
with redirect_stdout(stdout), redirect_stderr(stderr):
try:
code = cli.main(arguments)
except SystemExit as exc:
code = exc.code if exc.code is not None else 0
return code, stdout.getvalue(), stderr.getvalue()
def run_script(arguments, env=None, cwd=None, input_bytes=None, timeout=120):
"""Run the wrapper as a subprocess; return CompletedProcess."""
return subprocess.run(
[sys.executable, str(SCRIPT)] + arguments,
capture_output=True,
text=True,
encoding="utf-8",
errors="replace",
env=env,
cwd=cwd,
input=input_bytes,
timeout=timeout,
)
def minimal_path_env():
"""A PATH containing only a python3 symlink (no node, no npx)."""
tmp = Path(tempfile.mkdtemp())
bindir = tmp / "bin"
bindir.mkdir()
os.symlink(sys.executable, bindir / "python3")
env = os.environ.copy()
env["PATH"] = str(bindir)
return tmp, env
def node_shim_env(version_line):
"""A PATH whose `node` is a shim printing `version_line`."""
tmp, env = minimal_path_env()
bindir = tmp / "bin2"
bindir.mkdir()
shim = bindir / "node"
shim.write_text("#!/bin/sh\n%s\n" % version_line)
shim.chmod(0o755)
env["PATH"] = str(bindir) + os.pathsep + env["PATH"]
return tmp, env
class WrapperCoreTests(unittest.TestCase):
"""Offline wrapper behavior: help, usage errors, pre-validation, plans."""
def test_script_is_executable_and_has_shebang(self):
mode = stat.S_IMODE(SCRIPT.stat().st_mode)
self.assertTrue(mode & stat.S_IXUSR, "scripts/anydoc must be executable")
with SCRIPT.open("rb") as handle:
first = handle.readline().decode("utf-8", "replace").strip()
self.assertEqual(first, "#!/usr/bin/env python3")
def test_direct_execution_via_shebang(self):
result = subprocess.run(
[str(SCRIPT), "info"], capture_output=True, text=True, timeout=60
)
self.assertEqual(result.returncode, 0, result.stderr)
self.assertIn("anydoc", result.stdout)
self.assertIn("0.1.6", result.stdout)
def test_help_exits_zero_with_usage_and_examples(self):
for arguments in (
["--help"],
["convert", "--help"],
["batch", "--help"],
["info", "--help"],
):
with self.subTest(arguments=arguments):
code, stdout, stderr = run_in_process(arguments)
self.assertEqual(code, 0, stderr)
self.assertIn("usage", stdout.lower())
self.assertIn("Example", stdout)
self.assertEqual(stderr, "")
def test_batch_help_documents_exit_semantics(self):
_, stdout, _ = run_in_process(["batch", "--help"])
self.assertIn("1 when any", stdout)
self.assertIn("input failed", stdout)
def test_help_works_without_node_on_path(self):
tmp, env = minimal_path_env()
try:
for arguments in (
["--help"],
["convert", "--help"],
["batch", "--help"],
["info", "--help"],
):
with self.subTest(arguments=arguments):
result = run_script(arguments, env=env)
self.assertEqual(result.returncode, 0, result.stderr)
self.assertIn("usage", result.stdout.lower())
self.assertIn("Example", result.stdout)
self.assertEqual(result.stderr, "")
finally:
shutil.rmtree(tmp)
def test_usage_errors_exit_2(self):
# Each usage error must exit 2 on stderr and name the offending token
# OR the missing input (argparse names the missing subcommand for an
# unknown root option).
cases = (
(["--bogus"], "command"),
(["convert", str(DOCX), "--bogus"], "--bogus"),
(["convert"], "required"),
(["batch"], "required"),
)
for arguments, needle in cases:
with self.subTest(arguments=arguments):
code, stdout, stderr = run_in_process(arguments)
self.assertEqual(code, 2)
self.assertEqual(stdout, "")
self.assertIn(needle, stderr)
self.assertIn("usage", stderr.lower())
self.assertNotIn("Traceback", stderr)
def test_info_reports_tool_and_version(self):
code, stdout, stderr = run_in_process(["info"])
self.assertEqual(code, 0, stderr)
self.assertIn("anydoc", stdout)
self.assertIn("0.1.6", stdout)
self.assertEqual(stderr, "")
def test_info_version_prints_exact_version(self):
code, stdout, stderr = run_in_process(["info", "--version"])
self.assertEqual(code, 0, stderr)
self.assertEqual(stdout.strip(), "0.1.6")
self.assertEqual(stderr, "")
def test_convert_missing_input_prevalidation(self):
code, stdout, stderr = run_in_process(
["convert", "/nonexistent/anydoc-input.docx"]
)
self.assertEqual(code, 1)
self.assertEqual(stdout, "")
self.assertIn("/nonexistent/anydoc-input.docx", stderr)
self.assertNotIn("Traceback", stderr)
def test_convert_directory_input_prevalidation(self):
code, stdout, stderr = run_in_process(["convert", str(FIXTURES)])
self.assertEqual(code, 1)
self.assertEqual(stdout, "")
self.assertIn(str(FIXTURES), stderr)
self.assertIn("directory", stderr)
self.assertNotIn("Traceback", stderr)
def test_convert_output_path_is_directory(self):
code, stdout, stderr = run_in_process(
["convert", str(DOCX), "-o", str(FIXTURES)]
)
self.assertEqual(code, 1)
self.assertEqual(stdout, "")
self.assertIn("directory", stderr)
self.assertNotIn("Traceback", stderr)
def test_convert_invalid_format_exit_2(self):
code, stdout, stderr = run_in_process(
["convert", str(DOCX), "-f", "bogus"]
)
self.assertEqual(code, 2)
self.assertEqual(stdout, "")
self.assertIn("invalid format 'bogus'", stderr)
self.assertNotIn("Traceback", stderr)
def test_convert_dry_run_plans_without_executing(self):
with tempfile.TemporaryDirectory() as tmp:
out = Path(tmp) / "out.md"
code, stdout, stderr = run_in_process(
["convert", str(DOCX), "-o", str(out), "--dry-run"]
)
self.assertEqual(code, 0, stderr)
self.assertIn("npx -y " + PINNED, stdout)
self.assertIn(str(DOCX), stdout)
self.assertEqual(stderr, "")
self.assertFalse(out.exists(), "dry-run must not create outputs")
def test_convert_dry_run_json(self):
code, stdout, stderr = run_in_process(
["convert", str(DOCX), "--dry-run", "--json"]
)
self.assertEqual(code, 0, stderr)
doc = json.loads(stdout)
self.assertTrue(doc["dry_run"])
self.assertEqual(doc["command"], "convert")
self.assertIn("npx -y " + PINNED, doc["command_line"])
self.assertEqual(stderr, "")
def test_batch_dry_run_json_plan_and_no_output_dir(self):
with tempfile.TemporaryDirectory() as tmp:
out_dir = Path(tmp) / "out"
code, stdout, stderr = run_in_process(
[
"batch",
str(DOCX),
str(CSV),
"--out-dir",
str(out_dir),
"--dry-run",
"--json",
]
)
self.assertEqual(code, 0, stderr)
doc = json.loads(stdout)
self.assertTrue(doc["dry_run"])
self.assertEqual(doc["command"], "batch")
self.assertEqual(len(doc["plan"]), 2)
for entry in doc["plan"]:
self.assertIn("input", entry)
self.assertIn("output", entry)
self.assertIn("command", entry)
self.assertIn("npx -y " + PINNED, entry["command"])
self.assertEqual(stderr, "")
self.assertFalse(out_dir.exists(), "dry-run must not create out-dir")
def test_batch_dry_run_marks_invalid_inputs(self):
with tempfile.TemporaryDirectory() as tmp:
missing = Path(tmp) / "missing.docx"
code, stdout, stderr = run_in_process(
[
"batch",
str(DOCX),
str(missing),
"--out-dir",
str(Path(tmp) / "out"),
"--dry-run",
"--json",
]
)
self.assertEqual(code, 0, stderr)
doc = json.loads(stdout)
self.assertEqual(len(doc["plan"]), 2)
self.assertFalse(doc["plan"][0]["would_fail"])
self.assertTrue(doc["plan"][1]["would_fail"])
self.assertIn("not found", doc["plan"][1]["error"])
def test_batch_dry_run_defaults_out_dir_to_cwd(self):
with tempfile.TemporaryDirectory() as tmp:
code, stdout, stderr = run_in_process(
[
"batch",
str(DOCX),
"--dry-run",
"--json",
],
)
self.assertEqual(code, 0, stderr)
doc = json.loads(stdout)
self.assertEqual(doc["out_dir"], str(Path.cwd()))
def test_batch_requires_inputs(self):
code, stdout, stderr = run_in_process(["batch"])
self.assertEqual(code, 2)
self.assertEqual(stdout, "")
def test_convert_json_error_stays_parseable(self):
code, stdout, stderr = run_in_process(
["convert", "/nonexistent/anydoc-input.docx", "--json"]
)
self.assertEqual(code, 1)
doc = json.loads(stdout)
self.assertFalse(doc["ok"])
self.assertEqual(doc["exit_code"], 1)
self.assertIn("not found", doc["error"])
self.assertIn("not found", stderr)
def test_node_missing_error_via_minimal_path(self):
tmp, env = minimal_path_env()
try:
result = run_script(["convert", str(DOCX)], env=env)
self.assertEqual(result.returncode, 1)
self.assertEqual(result.stdout, "")
self.assertIn("node", result.stderr.lower())
self.assertIn("20", result.stderr)
self.assertNotIn("Traceback", result.stderr)
finally:
shutil.rmtree(tmp)
def test_node_too_old_error_via_shim(self):
tmp, env = node_shim_env('echo "v18.20.0"')
try:
result = run_script(["convert", str(DOCX)], env=env)
self.assertEqual(result.returncode, 1)
self.assertEqual(result.stdout, "")
self.assertIn("v18.20.0", result.stderr)
self.assertIn("20", result.stderr)
self.assertNotIn("Traceback", result.stderr)
finally:
shutil.rmtree(tmp)
@unittest.skipUnless(shutil.which("node"), "node not on PATH")
def test_npx_missing_error_via_path_with_node(self):
tmp = Path(tempfile.mkdtemp())
try:
bindir = tmp / "bin"
bindir.mkdir()
os.symlink(sys.executable, bindir / "python3")
node_bin = tmp / "bin2"
node_bin.mkdir()
os.symlink(Path(shutil.which("node")), node_bin / "node")
env = os.environ.copy()
env["PATH"] = str(node_bin) + os.pathsep + str(bindir)
result = run_script(["convert", str(DOCX)], env=env)
self.assertEqual(result.returncode, 1)
self.assertEqual(result.stdout, "")
self.assertIn("npx", result.stderr)
self.assertIn(PINNED, result.stderr)
self.assertNotIn("Traceback", result.stderr)
finally:
shutil.rmtree(tmp)
def test_hint_mapping_for_known_error_classes(self):
cases = (
(
"anydoc: unsupported input: PDF has no extractable text "
"(Scanned, 1 pages): OCR is required",
"no-ocr",
("OCR", "Firecrawl Parse", "not retry"),
),
("anydoc: document is encrypted", "encrypted", ("encrypted", "unencrypted")),
(
"anydoc: malformed document: not a readable zip archive: "
"invalid Zip archive: Could not find EOCD",
"malformed",
("malformed", "corrupt", "zip"),
),
(
"anydoc: unsupported input: unrecognized file content and "
"extension: unsupported.xyz",
"unsupported",
("unsupported", "-f"),
),
(
"anydoc: resource limit exceeded (max_entry_bytes): "
"word/document.xml declares 201326759 decompressed bytes",
"resource-limit",
(),
),
# Wrapper pre-validation messages map to the "io" class with no hint.
("input file not found: /x/missing.docx", "io", ()),
("input path is a directory, not a file: /x/dir", "io", ()),
)
for message, expected_class, keywords in cases:
with self.subTest(message=message):
error_class, hint = cli.error_class_hint(message)
self.assertEqual(error_class, expected_class)
if keywords:
self.assertIsNotNone(hint)
for keyword in keywords:
self.assertIn(keyword, hint)
def test_build_cli_command_shape(self):
self.assertEqual(
cli.build_cli_command("report.docx", "out.md", "csv"),
["npx", "-y", PINNED, "report.docx", "-o", "out.md", "-f", "csv"],
)
self.assertEqual(
cli.build_cli_command("report.docx", None, None),
["npx", "-y", PINNED, "report.docx"],
)
# stdin passes through as `-`
self.assertEqual(
cli.build_cli_command("-", None, "csv"),
["npx", "-y", PINNED, "-", "-f", "csv"],
)
# a dash-leading filename places -o/-f BEFORE the `--` separator
# (npx forwards `--` to the CLI, so options after it read as inputs)
self.assertEqual(
cli.build_cli_command("-weird", "o.md", "csv"),
["npx", "-y", PINNED, "-o", "o.md", "-f", "csv", "--", "-weird"],
)
self.assertEqual(
cli.build_cli_command("-weird", None, None),
["npx", "-y", PINNED, "--", "-weird"],
)
@unittest.skipUnless(shutil.which("node") and shutil.which("npx"), "toolchain missing")
def test_runtime_errors_empty_when_toolchain_present(self):
self.assertEqual(cli.runtime_errors(), [])
def test_format_aliases_accepted(self):
code, _stdout, stderr = run_in_process(
["convert", str(DOCX), "-f", "docm", "--dry-run"]
)
self.assertEqual(code, 0, stderr)
# --- CLI timeout: --json must still yield one parseable JSON document ---
def _timeout_side_effect(self):
"""A subprocess.run replacement that raises a TimeoutExpired."""
def _boom(*args, **kwargs):
exc = subprocess.TimeoutExpired(
cmd=args[0], timeout=cli.RUN_TIMEOUT
)
exc.pid = 4242 # set post-construction, as subprocess.run does
raise exc
return _boom
def test_run_cli_timeout_kills_group_and_raises(self):
with mock.patch.object(
cli.subprocess, "run", side_effect=self._timeout_side_effect()
), mock.patch.object(cli.os, "killpg") as mock_kill:
with self.assertRaises(cli.CliTimeoutError):
cli.run_cli(["npx", "-y", cli.PINNED, "x.docx"])
mock_kill.assert_called_once_with(4242, signal.SIGKILL)
def test_convert_timeout_with_json_emits_error_envelope(self):
with mock.patch.object(cli, "runtime_errors", return_value=[]), mock.patch.object(
cli.subprocess, "run", side_effect=self._timeout_side_effect()
), mock.patch.object(cli.os, "killpg") as mock_kill:
code, stdout, stderr = run_in_process(
["convert", str(DOCX), "--json"]
)
self.assertEqual(code, 1)
mock_kill.assert_called_once_with(4242, signal.SIGKILL)
doc = json.loads(stdout) # exactly one parseable JSON document
self.assertFalse(doc["ok"])
self.assertEqual(doc["exit_code"], 1)
self.assertEqual(doc["error_class"], "timeout")
self.assertIn("did not complete within 120 seconds", doc["error"])
self.assertIn("did not complete within 120 seconds", stderr)
self.assertNotIn("Traceback", stderr)
def test_convert_timeout_without_json_uses_stderr(self):
with mock.patch.object(cli, "runtime_errors", return_value=[]), mock.patch.object(
cli.subprocess, "run", side_effect=self._timeout_side_effect()
), mock.patch.object(cli.os, "killpg"):
code, stdout, stderr = run_in_process(["convert", str(DOCX)])
self.assertEqual(code, 1)
self.assertEqual(stdout, "")
self.assertIn("did not complete within 120 seconds", stderr)
self.assertNotIn("Traceback", stderr)
def test_batch_timeout_with_json_emits_error_envelope(self):
with tempfile.TemporaryDirectory() as tmp:
out_dir = Path(tmp) / "out"
with mock.patch.object(
cli, "runtime_errors", return_value=[]
), mock.patch.object(
cli.subprocess, "run", side_effect=self._timeout_side_effect()
), mock.patch.object(cli.os, "killpg"):
code, stdout, stderr = run_in_process(
["batch", str(DOCX), "--out-dir", str(out_dir), "--json"]
)
self.assertEqual(code, 1)
doc = json.loads(stdout)
self.assertFalse(doc["ok"])
self.assertEqual(doc["command"], "batch")
self.assertEqual(doc["error_class"], "timeout")
self.assertIn("did not complete within 120 seconds", stderr)
self.assertNotIn("Traceback", stderr)
class RealCliTests(unittest.TestCase):
"""End-to-end conversions through the pinned CLI; skip when unavailable."""
skip_reason = None
@classmethod
def setUpClass(cls):
if not shutil.which("npx") or not shutil.which("node"):
cls.skip_reason = "npx/node not available"
return
try:
proc = subprocess.run(
["npx", "-y", PINNED, "--version"],
capture_output=True,
text=True,
timeout=120,
)
except (OSError, subprocess.TimeoutExpired):
cls.skip_reason = "pinned CLI unavailable"
return
if proc.returncode != 0 or "0.1.6" not in proc.stdout:
cls.skip_reason = "pinned CLI unavailable"
return
cls.skip_reason = None
def setUp(self):
if self.__class__.skip_reason:
self.skipTest(self.__class__.skip_reason)
def _parse_ok_json(self, result):
"""Assert a successful --json run and return its parsed document."""
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(result.stderr, "")
doc = json.loads(result.stdout)
self.assertTrue(doc["ok"])
self.assertEqual(doc["exit_code"], 0)
return doc
def test_convert_to_stdout(self):
result = run_script(["convert", str(DOCX)])
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(result.stderr, "")
self.assertIn("# ", result.stdout)
self.assertIn("## ", result.stdout)
def test_convert_to_file_silent(self):
with tempfile.TemporaryDirectory() as tmp:
out = Path(tmp) / "out.md"
result = run_script(["convert", str(DOCX), "-o", str(out)])
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(result.stdout, "")
self.assertEqual(result.stderr, "")
self.assertTrue(out.exists())
content = out.read_text(encoding="utf-8")
self.assertIn("## ", content)
def test_convert_silently_overwrites_seeded_file(self):
with tempfile.TemporaryDirectory() as tmp:
out = Path(tmp) / "out.md"
out.write_text("SENTINEL\n", encoding="utf-8")
result = run_script(["convert", str(DOCX), "-o", str(out)])
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(result.stdout, "")
self.assertEqual(result.stderr, "")
content = out.read_text(encoding="utf-8")
self.assertNotIn("SENTINEL", content)
self.assertIn("## ", content)
def test_convert_fresh_cwd_creates_no_stray_files(self):
with tempfile.TemporaryDirectory() as tmp:
result = run_script(["convert", str(DOCX)], cwd=tmp)
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(result.stderr, "")
self.assertFalse((Path(tmp) / "out.md").exists())
self.assertEqual(list(Path(tmp).iterdir()), [])
def test_convert_stdin_csv(self):
result = run_script(
["convert", "-", "-f", "csv"], input_bytes="a,b\n1,2\n"
)
self.assertEqual(result.returncode, 0, result.stderr)
self.assertIn("| a | b |", result.stdout)
self.assertIn("| 1 | 2 |", result.stdout)
def test_convert_empty_stdin_errors_without_hanging(self):
result = run_script(["convert", "-"], input_bytes="")
self.assertNotEqual(result.returncode, 0)
self.assertIn("anydoc", result.stderr)
self.assertNotIn("Traceback", result.stderr)
def test_convert_extensionless_file_with_fmt_csv(self):
with tempfile.TemporaryDirectory() as tmp:
data = Path(tmp) / "data"
data.write_bytes(CSV.read_bytes())
result = run_script(["convert", str(data), "-f", "csv"])
self.assertEqual(result.returncode, 0, result.stderr)
self.assertIn("| Kind | Value | Note |", result.stdout)
def test_convert_scanned_pdf_hint(self):
result = run_script(["convert", str(SCANNED)])
self.assertEqual(result.returncode, 1)
self.assertEqual(result.stdout, "")
self.assertIn("OCR", result.stderr)
self.assertIn("Firecrawl Parse", result.stderr)
self.assertIn("not retry", result.stderr)
self.assertNotIn("Traceback", result.stderr)
def test_convert_encrypted_hint(self):
result = run_script(["convert", str(ENCRYPTED)])
self.assertEqual(result.returncode, 1)
self.assertEqual(result.stdout, "")
self.assertIn("encrypted", result.stderr)
self.assertIn("unencrypted", result.stderr)
self.assertNotIn("Traceback", result.stderr)
def test_convert_malformed_hint(self):
result = run_script(["convert", str(MALFORMED)])
self.assertEqual(result.returncode, 1)
self.assertEqual(result.stdout, "")
self.assertIn("malformed", result.stderr)
self.assertNotIn("Traceback", result.stderr)
def test_convert_unsupported_hint(self):
result = run_script(["convert", str(UNSUPPORTED)])
self.assertEqual(result.returncode, 1)
self.assertEqual(result.stdout, "")
self.assertIn("unsupported", result.stderr)
self.assertIn("-f", result.stderr)
self.assertNotIn("Traceback", result.stderr)
def test_convert_json_success_to_file(self):
with tempfile.TemporaryDirectory() as tmp:
out = Path(tmp) / "out.md"
result = run_script(
["convert", str(DOCX), "-o", str(out), "--json"]
)
doc = self._parse_ok_json(result)
self.assertEqual(doc["output"], str(out))
def test_convert_json_success_embeds_markdown(self):
result = run_script(["convert", str(DOCX), "--json"])
doc = self._parse_ok_json(result)
self.assertIn("## ", doc["markdown"])
def test_convert_json_failure(self):
result = run_script(["convert", str(SCANNED), "--json"])
self.assertEqual(result.returncode, 1)
doc = json.loads(result.stdout)
self.assertFalse(doc["ok"])
self.assertEqual(doc["exit_code"], 1)
self.assertEqual(doc["error_class"], "no-ocr")
self.assertIn("OCR", result.stderr)
def test_batch_mixed_continues_past_failures(self):
with tempfile.TemporaryDirectory() as tmp:
out_dir = Path(tmp) / "out"
result = run_script(
[
"batch",
str(DOCX),
str(ENCRYPTED),
str(TABLES),
"--out-dir",
str(out_dir),
]
)
self.assertEqual(result.returncode, 1)
stdout = result.stdout
self.assertIn("ok %s" % DOCX, stdout)
self.assertIn("FAIL %s" % ENCRYPTED, stdout)
self.assertIn("ok %s" % TABLES, stdout)
self.assertIn("summary: 3 total, 2 succeeded, 1 failed", stdout)
self.assertTrue((out_dir / "fixture-handmade-outline.md").exists())
self.assertTrue((out_dir / "fixture-handmade-tables.md").exists())
self.assertFalse((out_dir / "encrypted--errors.md").exists())
self.assertIn("encrypted", result.stderr)
self.assertIn("hint", result.stderr)
def test_batch_all_valid_exits_zero(self):
with tempfile.TemporaryDirectory() as tmp:
out_dir = Path(tmp) / "out"
result = run_script(
[
"batch",
str(DOCX),
str(CSV),
"--out-dir",
str(out_dir),
]
)
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(result.stderr, "")
self.assertIn("summary: 2 total, 2 succeeded, 0 failed", result.stdout)
self.assertTrue((out_dir / "fixture-handmade-outline.md").exists())
self.assertTrue((out_dir / "fixture-sheet.md").exists())
def test_batch_json_all_valid(self):
with tempfile.TemporaryDirectory() as tmp:
result = run_script(
[
"batch",
str(DOCX),
str(CSV),
"--out-dir",
str(Path(tmp) / "out"),
"--json",
]
)
doc = self._parse_ok_json(result)
self.assertEqual(doc["summary"], {"total": 2, "succeeded": 2, "failed": 0})
self.assertEqual([f["status"] for f in doc["files"]], ["ok", "ok"])
def test_batch_json_mixed_keeps_stdout_parseable(self):
with tempfile.TemporaryDirectory() as tmp:
result = run_script(
[
"batch",
str(DOCX),
str(ENCRYPTED),
"--out-dir",
str(Path(tmp) / "out"),
"--json",
]
)
self.assertEqual(result.returncode, 1)
doc = json.loads(result.stdout)
self.assertFalse(doc["ok"])
self.assertEqual(doc["exit_code"], 1)
self.assertEqual(doc["summary"], {"total": 2, "succeeded": 1, "failed": 1})
self.assertIn("encrypted", result.stderr)
def test_batch_json_failure_entries_share_error_class_shape(self):
with tempfile.TemporaryDirectory() as tmp:
missing = Path(tmp) / "missing.docx"
result = run_script(
[
"batch",
str(ENCRYPTED),
str(missing),
"--out-dir",
str(Path(tmp) / "out"),
"--json",
]
)
self.assertEqual(result.returncode, 1)
doc = json.loads(result.stdout)
by_input = {entry["input"]: entry for entry in doc["files"]}
cli_fail = by_input[str(ENCRYPTED)]
pre_fail = by_input[str(missing)]
self.assertEqual(cli_fail["status"], "failed")
self.assertEqual(cli_fail["error_class"], "encrypted")
self.assertEqual(pre_fail["status"], "failed")
self.assertEqual(pre_fail["error_class"], "io")
self.assertEqual(
set(cli_fail.keys()),
set(pre_fail.keys()),
"all batch failure entries must share the same shape",
)
def test_convert_dash_leading_filename(self):
with tempfile.TemporaryDirectory() as tmp:
(Path(tmp) / "-weird").write_bytes(CSV.read_bytes())
result = run_script(
["convert", "-f", "csv", "--", "-weird"],
cwd=tmp,
)
self.assertEqual(result.returncode, 0, result.stderr)
self.assertIn("| Kind | Value | Note |", result.stdout)
def test_batch_duplicates_convert_per_occurrence(self):
with tempfile.TemporaryDirectory() as tmp:
out_dir = Path(tmp) / "out"
result = run_script(
["batch", str(DOCX), str(DOCX), "--out-dir", str(out_dir)]
)
self.assertEqual(result.returncode, 0, result.stderr)
self.assertIn("summary: 2 total, 2 succeeded, 0 failed", result.stdout)
self.assertEqual(
sorted(p.name for p in out_dir.iterdir()), ["fixture-handmade-outline.md"]
)
def test_batch_same_basename_collision_last_wins(self):
with tempfile.TemporaryDirectory() as tmp:
a_dir = Path(tmp) / "a"
b_dir = Path(tmp) / "b"
a_dir.mkdir()
b_dir.mkdir()
(a_dir / "same.docx").write_bytes(DOCX.read_bytes())
(b_dir / "same.docx").write_bytes(TABLES.read_bytes())
out_dir = Path(tmp) / "out"
result = run_script(
[
"batch",
str(a_dir / "same.docx"),
str(b_dir / "same.docx"),
"--out-dir",
str(out_dir),
]
)
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(
sorted(p.name for p in out_dir.iterdir()), ["same.md"]
)
content = (out_dir / "same.md").read_text(encoding="utf-8")
self.assertIn("| Head A | Head B | Head C |", content)
if __name__ == "__main__":
unittest.main()
+4
View File
@@ -60,6 +60,10 @@ Load this skill when the task involves any of the four formats:
documentation-site conventions, not office documents.
- **Data pipelines** — moving or transforming raw data belongs to
`data-engineering`; Excel here is a *deliverable format*, not a data store.
- **Office documents to Markdown** — converting an existing office document
(docx, xlsx, pptx, pdf, odt, rtf, epub, csv) to GitHub-Flavored Markdown
belongs to the `anydoc` skill; this skill owns generation, editing, and
validation, not document-to-markdown extraction.
## The Shared Workflow
+1
View File
@@ -7,6 +7,7 @@
- [agent-evals-and-observability](agent-evals-and-observability/SKILL.md): Design, run, review, or release framework- and vendor-neutral evaluations and observability for AI agents. Use when defining agent evals, datasets, graders, trajectory review, regression analysis, release gates, production traces, or privacy-aware telemetry. Covers task and trajectory contracts, statistical comparisons, and incident-to-case learning; route framework implementation to pydanticai or langgraph when needed.
- [agent-production-operations](bundles/agent-production-operations/SKILL.md): Operate an evaluated agent with tools and authority in production through a runtime control plane covering versioning, staged rollout, fallback, cost and latency budgets, tool health, human escalation, disablement, and trace-to-eval feedback. Do not use for building agents, designing evals, or general release, SRE, security, or platform engineering — those methods stay with their specialist skills.
- [agent-skills](agent-skills/SKILL.md): Use this skill when creating, reviewing, or editing Agent Skills-format skills, or when implementing skill discovery and loading in an agent client. It covers directory structure, SKILL.md metadata, progressive disclosure, evals, and repository conventions. Do not use this skill for general software work that does not involve the Agent Skills format or lifecycle.
- [anydoc](anydoc/SKILL.md): Convert Word (.doc/.docx/.docm), PowerPoint (.ppt/.pps/.pot/.pptx/.pptm/.ppsx/.ppsm), Excel (.xls/.xlsx/.xlsm/.xlsb), OpenDocument (.odt/.ods/.odp), RTF, EPUB, CSV, and PDF documents to clean GitHub-Flavored Markdown locally with the Any Doc CLI (npx -y @firecrawl/anydoc@0.1.6): headings, GFM tables, slide structure, and footnotes in one pass. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly. Do not use for generating, editing, or validating documents (use documents), for ebook packaging (use epub), or for OCR of scanned or image-only PDFs (anydoc does not OCR; route to OCR tooling).
- [api-design-and-evolution](api-design-and-evolution/SKILL.md): Design, document, review, and evolve consumer-facing APIs and event interfaces. Use when choosing REST/HTTP, GraphQL, RPC, events, webhooks, or streaming; writing OpenAPI or AsyncAPI contracts; defining schemas, pagination, mutations, errors, idempotency, or API compatibility; or planning API versioning, deprecation, and migration. Use secure-software-engineering for a full security lifecycle, ADR authoring for durable architecture decisions, and spec-driven-development for a delivery specification and implementation gates.
- [artifact-pyramids](artifact-pyramids/SKILL.md): Organize durable agent research outputs as summaries, analysis, and evidence dossiers. Use when producing multi-layer research artifacts or coordinating research handoffs.
- [autogen](autogen/SKILL.md): Expert skill for conversational multi-agent AI with Microsoft AutoGen. AssistantAgent, UserProxyAgent, GroupChat, code execution, nested chats, cancellation tokens, tool integration, and MCP support. Use when building conversation-driven multi-agent systems or comparing agent frameworks.