StackFollow on WhatAreYouBuilding.AI

Site Reliability Engineer

Observability

Datadog

Unified monitoring, tracing, and log platform for cloud-scale infrastructure.

Grafana

Dashboarding and visualization layer for metrics, logs, and traces.

Prometheus

Open-source metrics collection and alerting system built for reliability.

Incident Response

Opsgenie

Atlassian's incident alerting and on-call management tool.

PagerDuty

On-call scheduling and incident alerting/escalation platform.

incident.io

Incident management platform built around Slack-native workflows.

Chaos Engineering

Chaos Mesh

CNCF chaos engineering platform for Kubernetes.

Gremlin

Managed chaos engineering platform for proactively testing system resilience.

LitmusChaos

CNCF open-source chaos engineering framework for Kubernetes.