Please let Runware know you found this job on RemoteYeah. This helps us get more companies to post jobs here for you.
Description:
Runware is building high-performance infrastructure for scalable AI inference across various modalities.
The Site Reliability Engineer will ensure system reliability, performance, and resilience while working across software, infrastructure, and production operations.
Requirements:
Strong experience in SRE, Production Engineering, or similar roles with production systems at scale.
Proficient in debugging distributed systems, APIs, databases, and infrastructure.
Experience with observability systems, SLIs, SLOs, and incident management.
Familiarity with Kubernetes, containers, IaC, and automation using languages like Python, Go, or PHP.
Willingness to participate in an engineering on-call rotation and take ownership of production issues.
Benefits:
Generous paid time off including vacation, sick days, and public holidays.
Meaningful stock options to share in the company's success.
Remote-first work environment with flexible hours.
Paid family leave for maternity, paternity, and caregiver time.
Company retreats held twice a year in inspiring locations.