Files
magnus919_agent-skills/ml-engineering/README.md
T
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
92299e1238 feat(skill): beef up ml-engineering with scripts/templates/evals (#257)
Add a schema-valid eval manifest (6 cases: fine-tuning plan review, eval-set
design, quantization decision, deployment plan, regression triage, training-run
reproducibility), three fillable templates (training-run record, eval regression
table, quantization decision record), a stdlib eval-set overlap/leakage checker
with a unittest suite, routing to the llama-cpp tool skill, and a README Quick
Start documenting the script. Closes #240.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-03 15:25:06 -04:00

2.2 KiB

ML Engineering

Machine learning engineering methodology — model training, fine-tuning (LoRA/QLoRA), evaluation, quantization, deployment, and MLOps pipeline design. Grounded in practical engineering patterns for production ML systems.

Why Install This Skill

Your agent makes informed decisions about fine-tuning approaches, quantization trade-offs, GPU selection, and serving architecture with real VRAM budgets and benchmarks. Fillable templates turn training runs, eval comparisons, and quantization decisions into reviewable records, and the bundled eval-overlap checker catches train/eval contamination before it invalidates a benchmark.

What You Get

Directory Purpose
SKILL.md Core methodology, trigger conditions, reference index
references/ Deep-dive reference files loaded on demand
templates/ Fillable records: training-run record, eval regression table, quantization decision record
scripts/ check-eval-overlap.py — detects test-set leakage between train and eval corpora
evals/ Output-quality eval manifest for the skill's methodology cases

Triggers

Setting up fine-tuning runs, quantizing models, selecting training infrastructure, deploying inference servers, evaluating model quality, or triaging a model regression.

Requirements

Assumes familiarity with PyTorch/HuggingFace ecosystem. References cover vLLM, llama.cpp, TGI, DeepSpeed, and accelerate. The bundled script needs only Python 3 (standard library).

Quick Start

Check an eval corpus for leakage against your training data before trusting any eval score:

python3 ml-engineering/scripts/check-eval-overlap.py --train data/train/ --eval data/eval/

Each eval file is reported with its overlap fraction against the training corpus; an eval file that shares more than 10% of its text with training is flagged LEAK and the script exits 1, so it can gate a CI pipeline. Add --json for machine-readable output, --token-ngram 5 to compare token sequences instead of character shingles, and --max-overlap-fraction 0.05 to tighten the threshold.

Load SKILL.md for the methodology overview and reference table, then load specific references or templates as needed for the task at hand.