observability-and-instrumentation
SolidDesigns or reviews logs, metrics, traces, alerts, correlation, dashboards, and diagnostic instrumentation so failures and performance changes can be detected and explained at component boundaries. Use for observability work, incident readiness, instrumentation, or production diagnostics. Not for fixing a specific incident before reproducing it or for adding noisy logging without an operational question.
Install
Quality Score: 81/100
Skill Content
Details
- Author
- thiientv
- Repository
- thiientv/godmode
- Created
- 4 weeks ago
- Last Updated
- 2 weeks ago
- Language
- Python
- License
- MIT
Similar Skills
Semantically similar based on skill content — not just same category
observability-design
Review or design observability for services and applications using logs, metrics, traces, alerts, and user-impact signals. Use when operators need to detect, explain, and respond to failures without confusing telemetry volume with operational understanding.
observability
Instrument software so production questions get answered from signals, not guesses. Use when adding logging, metrics, tracing, or alerts, or when a system is hard to debug in production.
observability-incident-response
Design and assess application observability and incident response across logs, metrics, traces, correlation, alerting, runbooks, recovery, and post-incident learning. Use for production readiness, reliability improvement, or active incident diagnosis.