Description:
- Build and improve the inference layer of the Gcore Inference platform.
- Integrate and operate inference frameworks such as vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM.
- Bring new language and multimodal models into production.
- Improve inference latency, throughput, memory use, GPU utilization, and cost efficiency.
- Debug performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes.
- Collaborate with various teams to turn inference improvements into reliable product features.
- Contribute to open-source inference projects when appropriate.
Requirements:
- 5+ years of experience writing reliable, well-tested production code.
- Strong Python skills and experience designing production systems.
- Hands-on experience with PyTorch and deploying machine learning models.
- Experience with Linux, Docker, and Kubernetes.
- Experience in at least one relevant area: distributed systems, GPU computing, ML runtimes, model optimization, or cluster scheduling.
- Ability to debug complex problems across software, infrastructure, and hardware.
- Strong sense of developer experience and genuine interest in inference engineering.
- Good communication and collaboration skills.
Benefits:
- Competitive compensation.
- Flexible working hours and hybrid or remote options.
- Work from anywhere in the world for up to 45 days per year.
- Private medical insurance for you and your family.*
- Extra paid vacation and sick leave days.*
- Support for lifeβs important moments and celebrations.
- Language courses to help you connect and grow.
- Modern, welcoming offices with snacks, drinks, and entertainment.*
- Team sports and social activities.*
*Benefits may vary depending on your location.