Site Reliability Engineer
Observability
Datadog
Unified monitoring, tracing, and log platform for cloud-scale infrastructure.
Grafana
Dashboarding and visualization layer for metrics, logs, and traces.
Prometheus
Open-source metrics collection and alerting system built for reliability.
Incident Response
Opsgenie
Atlassian's incident alerting and on-call management tool.
PagerDuty
On-call scheduling and incident alerting/escalation platform.
incident.io
Incident management platform built around Slack-native workflows.
Chaos Engineering
Chaos Mesh
CNCF chaos engineering platform for Kubernetes.
Gremlin
Managed chaos engineering platform for proactively testing system resilience.
LitmusChaos
CNCF open-source chaos engineering framework for Kubernetes.