Files
magnus919_agent-skills/kubernetes/references/observability.md
T

2.3 KiB

Observability and metrics

Minimal health evidence

kubectl get --raw='/readyz?verbose'
kubectl get nodes
kubectl get pods -A
kubectl events -A --types=Warning
kubectl top nodes
kubectl top pods -A --sort-by=memory

kubectl top depends on the resource metrics pipeline and may be unavailable even when the API server is healthy. Distinguish missing metrics, stale metrics, and a measured zero.

Logging

Use bounded kubectl logs with --tail, --since, and --previous. Cluster-level logging is normally implemented by a node agent, sidecar, or application-direct path. Container log rotation and runtime behavior are node/distribution concerns.

Events and audit

Events are short-lived diagnostic evidence and should be correlated by involved object, reason, type, count, and timestamp. Audit logs answer who called the API and at which stage, but require a configured policy and backend. Keep audit policy changes separate from ordinary workload changes.

Metrics selection

Use API-server deprecation metrics to locate deprecated API use. Use controller/node/workload metrics for health and capacity. Prometheus Operator, kube-state-metrics, and provider monitoring are ecosystem overlays; route to their official docs rather than pretending they are Kubernetes core.

Audit policy boundary

Audit logging requires an API-server policy and a backend. The policy levels are None, Metadata, Request, and RequestResponse; the more detailed levels increase sensitivity and cost. Start from the minimum evidence needed, redact or restrict request bodies, and treat audit-policy changes as control-plane changes rather than ordinary workload configuration.

Audit records have RequestReceived, ResponseStarted, ResponseComplete, and Panic stages. The policy is ordered: the first matching rule wins, and an empty rules list is invalid. Backends are currently log files or webhooks. Audit increases API-server memory use, so measure and bound high-volume rules.

Sources