Remote Machine Learning Engineer, Speech - Joint Audio-Video Modeling

Posted 1 hour ago

Share:

Please let Cantina know you found this job on RemoteYeah. This helps us get more companies to post jobs here for you.

Description:

  • Join Cantina Labs as a Research / ML Engineer in the Speech Team to develop advanced speech and audio generation systems focusing on joint audio-video modeling.
  • Own the audio side of multimodal generation, including voice cloning, cinematic dialogue, and adjacent speech tasks.

Requirements:

  • Exceptional experience with large-scale audio models (>8B parameters, >500k hours of data).
  • Deep hands-on experience with diffusion/flow-matching transformers and audio VAEs, neural codecs, and vocoders.
  • Strong software engineering skills, particularly in PyTorch, and experience with multi-node, multi-GPU distributed training.
  • Background in large-scale ML data and experience with voice cloning or expressive speech generation.

Benefits:

  • Competitive salary range of $200,000-$220,000 and generous company equity.
  • Comprehensive medical, dental, and vision insurance with 99.99% of premiums covered.
  • 42 days of paid time off, including PTO, sick days, and company holidays.
  • Additional benefits include generous parental leave, 401(k) plan, lifestyle spending account, and complimentary in-office meals.

Job type

Experience level

Required experience

-

Salary

$200,000—$220,000 / year

Degree requirement

No degree required

Location requirements

Report this job

Job expired or something else is wrong with this job?

Report job
SerpApi

SerpApi

Scrape Google and other search engines from our fast, easy, and complete API.

RemoteYeah Ads