Curated knowledge repository of SRE practices, tools, and culture from leading tech organizations including Airbnb, Netflix, Google, and many others.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONAtlantis is a self-hosted Terraform pull request automation tool that enables collaborative infrastructure-as-code workflows.
aiHelpDesk is a governed AI agent system for production database incident response and remediation across VM, Docker/Podman, and Kubernetes deployments. It diagnoses failures, executes fix playbooks under audit trails, and learns from every resolved incident.
Pulse is a batteries-included Spring Boot observability starter providing cardinality firewall, timeout-budget propagation, SLO-as-code, async context, PII masking, and stable error fingerprints with zero agents and one dependency.
A curated collection of resources for Site Reliability Engineer (SRE) interview preparation, covering Linux, networking, containers, Kubernetes, databases, system design, monitoring, incident processes, and interview questions.
Krkn Operator is a Kubernetes-native platform for centralized multi-cluster chaos engineering, enabling teams to orchestrate and manage chaos experiments across Kubernetes and OpenShift clusters from a single control plane.
openstatus is an open-source uptime monitoring and status page platform built around infrastructure as code. Monitors, status pages and notification channels are declared in YAML or Terraform, applied via CLI or CI, and managed through an API, MCP server or AI assistants. It offers 28 global check regions, incident communication and self-hosting options.
A large open-source collection of 2624 DevOps and SRE interview questions and hands-on exercises spanning Linux, Kubernetes, Terraform, AWS, Azure, GCP, Docker, Ansible, CI/CD, networking, databases, observability, and more.