observability-monitoring

Solid

Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting, burn-rate response, and postmortems. Use when asked about monitoring, наблюдаемость, алерты, Prometheus, Grafana, OpenTelemetry, logs, traces, profiles, service health, or incident evidence. Do not use for generic dashboard styling, frontend-only UI work, or unrelated code review.

AI & Automation 138 stars 21 forks Updated yesterday MIT

Install

View on GitHub

Quality Score: 87/100

Stars 20%
71
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
50
License 10%
100
Description 5%
100

Skill Content

# Observability Monitoring Use this skill to turn vague "is it working?" questions into evidence-backed monitoring, alerting, and incident workflows. Start from user or business impact, then move down through the system layers and choose the signal that can prove the current hypothesis. ## Operating rule Do not treat a green dashboard as proof of health. A monitoring claim is complete only when it names: 1. the observed scope and time window; 2. the user, business, or operator outcome being protected; 3. the signal and exact query/probe that supports the claim; 4. the threshold or SLO that defines bad; 5. the next human action and its runbook/evidence link. Keep code review, live runtime proof, UI/render proof, and release readiness as separate verdicts. Investigation is read-only by default. A restart, alert suppression, metric-schema/label change, sampling or retention change, or vendor reconfiguration is a production mutation: involve the responsible owner or incident authority, preserve the relevant evidence first, capture the exact config/command diff, state the rollback condition, and verify the user probe plus SLI after the change. ## Workflow ### 1. Freeze scope and collect live facts Before changing a monitor, alert, host, or service: - Read the repository `AGENTS.md`, relevant rules, runbooks, and deployment docs. - Establish the actual checkout/branch, deployment, process, host, port, proxy/tunnel, and data source. - Record the observation time, environme...

Details

Author
AnastasiyaW
Repository
AnastasiyaW/claude-code-config
Created
4 months ago
Last Updated
yesterday
Language
Python
License
MIT

Integrates with

Bundled in these plugins

Similar Skills

Semantically similar based on skill content — not just same category