Description:
- Design and build a managed Slurm service on Kubernetes.
- Write clean, reliable, and maintainable Go code.
- Develop scheduling and orchestration capabilities for GPU-intensive and distributed workloads.
- Build observability and automated remediation for GPU, node, network, and control-plane failures using VictoriaMetrics, Grafana, DCGM.
Requirements:
- Hands-on experience using Slurm in production, including workload submission and debugging.
- Strong proficiency in Go, with experience in building production-grade Kubernetes operators and controllers.
- Experience preserving traditional Slurm cluster behavior on Kubernetes.
- Ability to diagnose performance and reliability issues across GPUs and distributed systems.
- A product mindset with strong customer empathy.
- Excellent communication skills and end-to-end ownership of complex distributed-system challenges.
Benefits:
- Competitive compensation.
- Flexible working hours and hybrid or remote options.
- Work from anywhere in the world for up to 45 days per year.
- Private medical insurance for you and your family.*
- Extra paid vacation and sick leave days.*
- Support for lifeβs important moments and celebrations.
- Language courses to help you connect and grow.
- Modern offices with snacks, drinks, and entertainment.*
- Team sports and social activities.*
*Benefits may vary depending on your location.