* feat(skill): add documents family skill (PDF / Word / Excel / PowerPoint) One family skill for PDF, Word (.docx), Excel (.xlsx), and PowerPoint (.pptx) per the family-skill rule (epub precedent): shared workflow in SKILL.md (scope, content model, template, render, validate, deliver) with per-format load-on-demand references, generation templates per format, a stdlib validation script (--json, structural sanity + render check with graceful degradation), one fixture per format, a unittest suite, six output-quality eval cases spanning all four formats, a human README, the README.md index entry, and regenerated catalogs. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> * fix(skill): dispatch PDF renderer args per binary in documents validation The render check passed pdftoppm-only flags (-png/-r/-f/-l) to mutool and ghostscript, which reject them, so a machine with only mutool or gs would false-FAIL valid PDFs. Dispatch per-renderer argument sets (pdftoppm -png; mutool draw -o; gs -sDEVICE=png16m) and cover the dispatch with a unit test. Also: count PDF pages via the /Count page-tree fallback (page objects can hide in compressed ObjStm streams), drop the stale "unsupported input" exit-2 claim from the docstring, and stop labeling skipped files with a FAIL check. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> * style(skill): drop redundant local tempfile import in renderer dispatch test Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> --------- Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
4.9 KiB
PDF — Generation & Validation Reference
Last Updated: 2026-08-03
Load this reference when the target format is PDF — generating a
fixed-layout document, converting content to PDF, or validating a PDF artifact.
It complements the shared workflow in SKILL.md; this file is the PDF-specific
detail for steps 3-5 (template, render, validate).
PDF fundamentals
A PDF file is a linear byte stream, not a container:
- Header —
%PDF-1.xnear the start (x = 2..7 in practice). - Body — numbered indirect objects (
N 0 obj ... endobj): a catalog (/Type /Catalog), page tree (/Type /Pageswith/Kidsand/Count), page objects (/Type /Page), content streams, and font objects. - Cross-reference table (xref) — byte offsets of every object, which lets
readers jump straight to an object; followed by the
trailerwith/Root. startxref— byte offset of the xref table;%%EOFterminates the file.
The validation script checks exactly the properties that break in real life:
the %PDF- header, the %%EOF trailer, and the presence of page objects. A
file that opens in one viewer but not another is almost always a broken xref
or a stream whose /Length does not match its content — see
references/output-quality.md for the cross-format checklist.
Generation paths
PDF is a fixed-layout format: the author, not the reader, decides where every glyph lands. Choose the path by how much layout control you need.
Print-ready HTML/CSS (recommended for reports and memos)
Author the document as HTML with print CSS (@page rules, page breaks,
@media print), then render to PDF with a print-capable engine:
- WeasyPrint (Python, pip installable) — excellent CSS paged-media support;
embed fonts via
@font-face. - Headless Chromium (
--headless --print-to-pdf) — full CSS support, best for complex layouts; pass--no-pdf-header-footerfor clean output.
Keep the content in the content model (step 2 of the shared workflow), fill templates/pdf-template.md, and render. The template is the layout contract; the HTML/CSS is where fonts, margins, and page breaks live.
LaTeX (best for technical and long-form documents)
Write LaTeX source from the content model and compile with a TeX toolchain
(pdflatex, xelatex). Gives precise typography, references, and TOC
control. Costs: a toolchain dependency and a longer render cycle.
Direct PDF construction (small, dependency-free artifacts)
For tiny fixed artifacts (a one-page certificate, a label), a minimal PDF can
be written by hand with stdlib only: build the objects, compute the xref
offsets, and write the trailer. Keep streams short and compute /Length
exactly. This is what the bundled fixture fixtures/sample.pdf does.
What not to do
- Do not fake a PDF by renaming a text file — every PDF must start with the
%PDF-header and end with%%EOF; readers will reject anything else. - Do not generate a PDF that relies on fonts that will not be embedded; unembedded fonts render as garbage or get substituted (see output quality).
- Do not rasterize text to images unless the document is genuinely a scan; text should stay selectable.
Text extraction (reading a PDF)
PDF is a rendering format, so "reading" it means extracting text:
- pypdf (
pip install pypdf) — extract text per page:PageObject.extract_text(). - pdfminer.six — more accurate layout-aware extraction for complex layouts.
- pdftotext (poppler-utils) — fast CLI extraction for simple documents.
Extraction quality varies with how the PDF was produced. Scanned PDFs contain no text layer at all — they are images; extraction requires OCR, which is outside this skill's scope.
Validation specifics
python3 scripts/validate-documents.py --json report.pdf
python3 scripts/validate-documents.py --render-check --json report.pdf
Structural checks the script runs for PDF:
- PDF header —
%PDF-signature near the start. - EOF marker —
%%EOFtrailer near the end. - Page objects — at least one
/Type /Pageobject.
The render check renders page 1 to a raster via pdftoppm (or mutool/gs)
and reports unavailable when no renderer is installed.
Output-quality checklist for PDF
Before delivery, verify:
- Pages render — the render check succeeds; no blank or corrupt pages.
- Text is selectable — a content-stream text marker (BT/ET) is present unless the document is intentionally a scan.
- Fonts are embedded — no
missing glyphboxes; embed via the generator's font options. - Page count matches the scope — no accidental blank trailing pages.
- Links and bookmarks — internal links and the outline are functional.
- Metadata — title/author set where the reader will show them.
See references/output-quality.md for the cross-format version of this checklist.