* feat(skill): add anydoc core content and references
Add the anydoc skill content tree: SKILL.md (progressive-disclosure index
with frontmatter per ALLOWED_FIELDS), human-facing README, the five reference
files (formats, cli-reference, errors, workflows, sources), 24 committed
fixtures (valid + error cases), and a fixture-grounded eval manifest with 8
cases. Every documented behavior, exit code, and error message was verified
against the real pinned CLI (npx -y @firecrawl/anydoc@0.1.6); verbatim --help
and error transcripts are reproduced character-for-character.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(skill): add anydoc wrapper script and unit tests
Implements scripts/anydoc, a stdlib-only Python wrapper around the pinned
@firecrawl/anydoc@0.1.6 CLI: convert/batch/info subcommands, global
--json/--dry-run, input and output pre-validation, friendly hints for the
no-OCR/encrypted/malformed/unsupported error classes, Node >= 20 and npx
availability checks, deterministic batch output naming with documented
duplicate/collision behavior, and exit codes 0/1/2. Adds offline unittest
suite (46 tests, real-CLI tests skip when npx is unavailable) and keeps the
wrapper contract documented in cli-reference.md and errors.md.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(skill): ratchet anydoc evals to 14 grounded cases
Verify the pre-authored 8-case manifest and extend it with six
high-signal cases (PDF lower-fidelity pipeline, legacy .ppt table
flattening, ODP same-serializer, RTF, EPUB, CSV header promotion),
each grounded in real pinned-CLI runs against the committed fixtures.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(skill): integrate anydoc into repo catalog and artifacts
Add the sorted anydoc catalog entry to README.md (between agent-skills
and api-design-and-evolution), regenerate the tracked catalog artifacts
(.claude-plugin/marketplace.json, .codex-plugin/plugin.json,
.agents/plugins/marketplace.json, llms.txt) with the ruby generators,
and add a routing note to documents/SKILL.md pointing office-document
to-markdown conversion at the anydoc skill.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* fix(skill): polish anydoc wrapper timeout, JSON shape, and docs
- run_cli raises CliTimeoutError on the 120s timeout; convert/batch with
--json now emit one parseable JSON error envelope (error_class "timeout")
on stdout before exiting, so --json always yields exactly one JSON doc
- batch JSON failure entries (pre-validation and CLI) now carry error_class
("io" for missing/dir inputs, mapped classes for CLI failures), so all
batch failure entries share the same shape
- build_cli_command places -o/-f before the -- separator for dash-leading
filenames, so `convert -f csv -- -weird` converts instead of misparsing
("unexpected second input"); absolute-path inputs unchanged
- workflows.md vault-ingestion recipe globs notes/* instead of docs/* and
warns to run from a temp/vault dir, never touching repo-root docs/
- unit tests: +6 (timeout envelope x4, batch error_class shape,
dash-leading filename); suite grows 46 -> 52
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
---------
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
8.2 KiB
Workflows: recipes for converting documents to markdown
All recipes use the pinned CLI npx -y @firecrawl/anydoc@0.1.6 (ground truth)
and the skill's wrapper scripts/anydoc where it adds value. Commands are
shown relative to the repository root; anydoc/fixtures/... paths can be
replaced with any document path. The vault-ingestion recipe (section 5) is
written to be run from a temp or vault directory holding your own
documents. Each raw-CLI invocation converts exactly one document — there
is no batch mode.
1. Single conversion
# Markdown to stdout
npx -y @firecrawl/anydoc@0.1.6 anydoc/fixtures/fixture-handmade-outline.docx
# Markdown to a file (stdout stays silent; existing file is overwritten)
npx -y @firecrawl/anydoc@0.1.6 anydoc/fixtures/fixture-handmade-outline.docx -o outline.md
# Same jobs through the wrapper
python3 anydoc/scripts/anydoc convert anydoc/fixtures/fixture-handmade-outline.docx
python3 anydoc/scripts/anydoc convert anydoc/fixtures/fixture-handmade-outline.docx -o outline.md
Expected result: exit code 0, empty stderr, and GitHub-Flavored Markdown on
stdout (or written to the -o output file) containing #/##/### heading
lines.
2. Force the input format
# Extensionless or mislabeled file: name the format explicitly
npx -y @firecrawl/anydoc@0.1.6 ./data --format csv
npx -y @firecrawl/anydoc@0.1.6 ./report --format docx
Use --format <name> only when detection cannot work (CSV from stdin, or a
missing/wrong extension). Aliases resolve: --format xls, --format docm,
--format ppsx are accepted. An invalid name exits 2 with
anydoc: invalid format 'bogus'; expected one of: ....
3. Read a document from stdin
# CSV from stdin requires --format csv (no signature, no extension)
printf 'name,role\nAlice,Engineer\n' | npx -y @firecrawl/anydoc@0.1.6 - --format csv
# Any document type can come from stdin; detection reads the bytes
curl -s https://example.com/paper.pdf | npx -y @firecrawl/anydoc@0.1.6 -
The wrapper supports the same: cat data.csv | python3 anydoc/scripts/anydoc convert - -f csv.
Piping notes:
- Markdown goes to stdout only; diagnostics are the single
anydoc: <message>stderr line. - EPIPE is handled: if the downstream pipe closes early
(
... anydoc@0.1.6 big.xlsx | head -n 1), the CLI exits 0 with no stderr noise — piping intoheadis safe and is not a failure.
4. Batch conversion (raw CLI)
The raw CLI takes one document per invocation, so batch with a shell loop:
mkdir -p out
for f in anydoc/fixtures/*.docx; do
npx -y @firecrawl/anydoc@0.1.6 "$f" -o "out/$(basename "${f%.docx}").md"
done
Each failed document (error fixtures, scanned PDFs, encrypted files) exits 1
with its anydoc: <message> on stderr and produces no output file; the loop
continues with the next input. Handle or route those per
errors.md.
Or the wrapper, which is built for this (per-file status, continues past failures, summary, and a non-zero exit when any input failed):
python3 anydoc/scripts/anydoc batch \
anydoc/fixtures/fixture-handmade-outline.docx \
anydoc/fixtures/fixture-sheet.csv \
--out-dir out/
batch --dry-run --json prints the plan (input → output, dry-run marker)
without converting or creating anything:
python3 anydoc/scripts/anydoc batch anydoc/fixtures/fixture-handmade-outline.docx \
anydoc/fixtures/fixture-sheet.csv --out-dir out/ --dry-run --json
5. Vault-ingestion pattern
Convert a folder of mixed office documents to markdown for ingestion into a vault or knowledge base:
-
Collect the documents into a folder (mixed docx/xlsx/pptx/csv/odt/pdf is fine — text-based PDFs only; see the no-OCR caveat in errors.md).
-
Batch-convert with the wrapper into a markdown folder:
python3 anydoc/scripts/anydoc batch notes/*.docx notes/*.xlsx notes/*.csv --out-dir vault/inbox/(or the raw-CLI loop above if you are not using the wrapper).
Run this from a temp or vault directory — never from the agent-skills repo root. The glob matches whatever directory you name, and the repository tracks a top-level
docs/directory (distinct from thedocuments/skill): globbingdocs/*.docxthere, or deleting/cleaning those matches, would damage tracked repository files. Keep the source documents in their own folder (herenotes/) and convert into a separatevault/inbox/folder. -
Verify each output (step 6) — at minimum confirm exit 0 and that the structural markers your formats produce are present (headings for Word/PDF,
|tables for spreadsheets/CSV). -
Failures are per-file: the batch summary names what failed; route those files per errors.md (scanned PDF → OCR tooling, encrypted → unencrypted copy, unsupported → check extension) and re-run only the failures.
6. Output verification
Before treating a conversion as done:
- Exit code 0 — the CLI produced markdown. Exit 1: read the
anydoc: <message>stderr line and match it against errors.md. Exit 2: fix the command (usage error). - Structural markers — check the markers your format actually produces:
- Word / ODT / RTF / text-based PDF:
#/##heading lines (grep -E '^#{1,6} ' out.md). - Spreadsheets (xlsx/xls/ods) and CSV:
## <sheet>headings and|-delimited rows (grep -E '^\|' out.md). - Presentations (pptx/odp): slide titles as plain paragraphs,
>blockquote speaker notes,|table rows (legacy.ppthas no|rows — that is by design, not an error). - EPUB:
#chapter headings and[text](#fragment)internal links.
- Word / ODT / RTF / text-based PDF:
- Tables survived? If the source had tables and the output has no
|rows, check the caveats: PDF and legacy.pptflatten tables by design. - Large outputs: convert with
-o out.mdand inspect the file rather than streaming everything into context.
Use the committed fixtures to sanity-check an environment once:
npx -y @firecrawl/anydoc@0.1.6 anydoc/fixtures/fixture-handmade-outline.docx # headings
npx -y @firecrawl/anydoc@0.1.6 anydoc/fixtures/sheet.xlsx # ## Values + table
npx -y @firecrawl/anydoc@0.1.6 anydoc/fixtures/fixture-text.pdf # headings, no table
7. Large files and resource limits
- Conversion is not streaming — the document is read and processed as a whole, and safety limits protect against decompression and nesting bombs.
- Zip/image bombs are rejected via
max_entry_byteswith exit 1 and the prefixanydoc: resource limit exceeded (max_entry_bytes):(full examples in errors.md). This is by design — do not try to bypass it. -o out.mdis recommended for large documents so the output is written to a reviewable file instead of filling stdout/context; you can then read the parts you need.- Genuinely large real documents (as opposed to bombs) convert normally; the per-document limit only rejects entries whose declared decompressed size exceeds the cap.
- If a resource-limit error fires on a real file, the archive is malformed or hostile — re-export the document rather than disabling the limit.
8. Startup cost and performance
Each npx -y @firecrawl/anydoc@0.1.6 invocation costs roughly 0.33–0.55 s
of warm-cache startup (npm/npx process startup) on top of the conversion
itself, which is a few milliseconds (measured ~5 ms for a PDF, <1 ms for a
DOCX once the process is warm). There is no progress output; conversions are
effectively instant. Plan for ~0.5 s per document in batch loops, and prefer a
single npx process per document (you cannot batch inside one invocation).
9. Offline / cold-cache behavior
- The first
npxrun downloads the package plus the native binary (network required once); later runs use the npm cache. A cold-cache offline run fails with a clear npx fetch error before anydoc executes. - For permanent or fully offline use, install once:
npm install -g @firecrawl/anydoc, then callanydoc <file>directly. - The wrapper always invokes npx with
-y(non-interactive), so it never hangs on npx's install prompt — even on a cold cache it fails fast if the package cannot be fetched.