Files
magnus919_agent-skills/training-data-annotation/references/annotator-quality.md
T
Magnus HedemarkandGitHub 347c1d3ae7 feat(training-data-annotation): add annotation methodology (#506)
* feat(skill): add training data annotation methodology

* docs(training-data-annotation): register skill and clarify stopping boundary
2026-09-14 17:26:30 -04:00

1.0 KiB

Annotator quality and disagreement

Qualify people against the actual task, language/domain needs, accessibility, and confidentiality requirements. Provide practice items with feedback before production work; measure fatigue, throughput, abstention, and error patterns, not just speed.

Use blinded calibration items and periodic blind repeats. Choose agreement or reliability statistics appropriate to nominal, ordinal, continuous, span, or ranking labels; route the choice and uncertainty calculation to data-scientist. Agreement is evidence about reproducibility under the guide, not proof that the majority is correct.

Classify disagreement as ambiguity, missing evidence, guide defect, annotator mistake, or model-induced anchoring. Adjudication records the decision rule, authority, evidence, and whether already labeled items need re-review. Preserve the original labels and disagreement for audit. If preannotations are visible, compare with a blinded control to estimate anchoring and treat the result as a workflow effect.