Please let npv labs know you found this job on RemoteYeah. This helps us get more companies to post jobs here for you.
Description:
Site Reliability Engineer role focused on ensuring reliability of streaming and storage systems (Kafka, Hadoop HDFS, Ceph).
Work within a hybrid infrastructure (bare-metal on-prem, cloud) and manage the full lifecycle of reliability including architecture, deployment, and incident response.
Requirements:
5+ years of experience in running distributed systems reliability at scale in production.
Deep expertise in Kafka, Ceph, or similar distributed infrastructure.
Proven ability to design for scale, reliability, and failure recovery.
Experience mentoring engineers and making technical decisions.
Willingness to work 9am–6pm ET US hours.
Benefits:
Remote work with a high engineering standard and comfortable culture.
Flat hierarchy with easy access to business, product, and operations.
Opportunity to work with a large scale (20PB+) and real growth potential.
Ownership and direct impact on projects with flexibility to shift focus.
Salary up to 150k USD, negotiable for higher figures and EU/UK/US employment.