Remote Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Posted 2 hours ago

Share:

Please let Perplexity know you found this job on RemoteYeah. This helps us get more companies to post jobs here for you.

Description:

  • Build a self-serve compute platform for training and inference workloads.
  • Operate and manage a large GPU fleet across multiple cloud providers.
  • Develop scheduling and placement logic to optimize GPU resource usage.
  • Support both long-running training jobs and production inference services.
  • Manage Kubernetes for GPU orchestration across various clusters.
  • Implement fault tolerance, autoscaling, and observability for the GPU fleet.
  • Collaborate with teams to establish a coherent platform architecture.

Requirements:

  • Deep experience with Kubernetes, including custom operators and multi-cluster federation.
  • Proven experience managing GPU clusters at scale, including NVIDIA hardware and CUDA.
  • Familiarity with orchestrating compute across multiple cloud environments.
  • Strong understanding of distributed systems, scheduling, and resource allocation.
  • Proficient in infrastructure and systems-level programming (Go, Rust, or C++).
  • Experience supporting both training jobs and high-availability inference services.
  • Ability to take ownership of problems and navigate ambiguity.

Benefits:

  • Opportunity to work with cutting-edge GPU technology and cloud infrastructure.
  • Collaborate with a team of experts in AI and machine learning.
  • Contribute to the development of a unified platform for AI workloads.

Job title

Job type

Experience level

Required experience

-

Salary

$250,000—$485,000 / year

Degree requirement

No degree required

Location requirements

Benefits

-

Report this job

Job expired or something else is wrong with this job?

Report job
SerpApi

SerpApi

Scrape Google and other search engines from our fast, easy, and complete API.

RemoteYeah Ads