Please let Mirantis know you found this job on RemoteYeah. This helps us get more companies to post jobs here for you.
Description:
Define reliability metrics for a GPU-accelerated AI platform and own service-level indicators (SLIs) and objectives (SLOs) for the K0rdent Observability Framework (KOF).
Work on hybrid, edge, and air-gapped deployments using the Mirantis K0rdent stack.
Requirements:
5+ years in SRE, platform reliability, or a related software/infrastructure role.
Strong software engineering skills in Go or Python with experience in building and operating APIs or services in production.
Experience defining SLIs/SLOs and error budgets for production systems.
Hands-on experience with observability tools (e.g., Prometheus, OpenTelemetry, Grafana).
Solid understanding of Kubernetes and its emitted signals.
Strong communication skills with technical audiences.
Preferred: Experience with bare-metal and NVIDIA infrastructure, Mirantis K0rdent stack, and high-security environments.
Benefits:
Work with a leading company in cloud infrastructure.
Collaborate with passionate and talented colleagues.
Engage in open-source innovation and professional development.
Participate in conferences, company outings, and hackathons.
Competitive compensation package with strong benefits.