Files
magnus919_agent-skills/data-engineering/SKILL.md
T
Magnus HedemarkandGitHub 9917fb455b feat(data): define AI transformation and repair contracts (#505)
* feat(data): validate AI transformation and repair boundaries

* docs(data): expose AI transformation and repair triggers

* docs(data-engineering): preserve the full reference index
2026-09-14 17:22:46 -04:00

5.9 KiB

name, description, license, metadata
name description license metadata
data-engineering Design and operate data infrastructure — database operations (vector, relational, graph, time-series), ETL/ELT pipeline design (dbt patterns, incremental loading), SQL analytical patterns, data quality monitoring, schema migration, and storage infrastructure management, including model-assisted transformation contracts, deterministic acceptance, bounded retries, and sink reconciliation. Do not use for statistical analysis or ML model development. MIT
tags source_repo
data-engineering, etl, dbt, sql, database, graph-db, time-series, vector-db, migration, data-quality, storage, influxdb, neo4j https://github.com/magnus919/hermes-profiles

Data Engineering Methodology

Data engineering is the operational backbone of data-driven systems. This methodology covers running, maintaining, and evolving data infrastructure — from relational databases and vector stores to graph databases, time-series stores, and the transformation pipelines that move data between them.

When a model proposes data transformations, load the AI boundary workflow and use the companion record.

The Data Engineer's Domain

You own You don't own
Database operations — schema management, indexing, backup/recovery, migration across relational, vector, graph, and time-series stores Data modeling and schema design — that's the data architect
Data transformation pipelines — dbt models, ETL/ELT patterns, incremental loading, incremental strategies Statistical analysis and experiments — that's the data scientist
Analytical SQL — window functions, CTEs, query optimization, execution plan analysis, star schema queries Training infrastructure and model deployment — that's the ML engineer
Graph database operations — Neo4j data modeling, Cypher queries, graph algorithms, import/export Application-level data access patterns — that's the developer
Time-series database operations — InfluxDB schema design, downsampling, retention policies, Telegraf Infrastructure provisioning — that's the platform engineer
Data quality monitoring — integrity checks, deduplication, anomaly detection, freshness validation Visual dashboard design — that's the analyst / product-design-and-ux
Storage infrastructure — capacity planning, performance tuning, archival strategies

Reference Files

Reference When to load
references/sql-analytical-patterns.md Writing analytical SQL — window functions, CTEs, execution plan reading, star schema queries, engine-specific optimization (PostgreSQL, DuckDB, ClickHouse, BigQuery, Snowflake)
references/dbt-patterns.md Designing data transformation pipelines with dbt — project structure, modeling layers (staging/intermediate/facts/dimensions), materializations, tests, snapshots, Jinja macros, CI/CD, dbt Mesh
references/etl-pipeline-design.md Building reliable data pipelines — extraction strategies (full, incremental, CDC), transformation layers, validation gates, error handling, idempotency
references/data-quality.md Monitoring data integrity — quality dimensions, validation rule types, anomaly detection, deduplication strategies, pipeline health signals
references/graph-databases.md Working with graph databases — Neo4j data modeling, Cypher query patterns (traversal, aggregation, pathfinding), import strategies, graph algorithms, pipeline integration
references/time-series-databases.md Working with time-series databases — InfluxDB data model (measurements, tags, fields), schema design (cardinality), downsampling, retention, Telegraf ingest, comparison with TimescaleDB/QuestDB/Prometheus
references/vector-db-operations.md Managing vector databases — Milvus, Qdrant, Chroma — index types, collection lifecycle, dimension migrations, backup strategies
references/database-migrations.md Schema evolution — zero-downtime migration patterns, rollback planning, versioned schemas, test-first migrations
references/backup-and-recovery.md Backup strategies per data store type, RPO/RTO planning, WAL archiving, snapshot management, recovery plan template
references/ai-transformation-boundaries.md Defining acceptance, retry, cost, and publication boundaries for model-assisted transformations
  • postgres — operating a PostgreSQL server itself: configuration review, index and query-plan diagnosis, vacuum/bloat management, WAL archiving and point-in-time recovery, replication and failover, upgrades. This skill owns the engine-specific runbooks; data-engineering owns the engine-neutral methodology.
  • supabase — Supabase platform operations: migrations, RLS, Auth, Storage, Functions, and self-hosting. To measure an agent's Supabase task competence, use its agent evals harness reference.

Core Principles

Data without integrity is noise — No pipeline, model, or dashboard is worth more than the quality of the data feeding it. Validate at every boundary.

Design for operability — Every database, pipeline, and store needs monitoring, backup, and recovery procedures defined before it goes to production. If you can't detect failure, you can't recover from it.

Idempotency is a requirement — Every pipeline should produce the same result whether it runs once or twice. Duplicate handling is not optional.

Schema changes are code changes — Every migration needs review, testing, and a rollback plan. Schema drift is technical debt with compounding interest.

Know your storage characteristics — Access patterns, retention requirements, growth rates, and consistency guarantees determine the right storage architecture. Choose based on data, not familiarity.