Please let Replit know you found this job on RemoteYeah. This helps us get more companies to post jobs here for you.
Description:
Join Replit's Site Reliability Engineering (SRE) team to ensure the reliability, scalability, and performance of infrastructure serving millions of developers.
Implement automation and establish best practices to maintain high availability and improve system reliability.
Requirements:
8-10 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering).
Strong programming skills in Python or Go, with a focus on high-quality, well-tested code.
Deep understanding of distributed systems and experience with Kubernetes and cloud-native technologies.
Proven track record in designing and maintaining monitoring and observability solutions.
Strong incident management skills and experience leading complex incident responses.
Familiarity with infrastructure as code tools (Terraform, Pulumi) and configuration management.
Excellent communication skills and ability to mentor engineers at various levels.