Please let Cantina know you found this job on RemoteYeah. This helps us get more companies to post jobs here for you.
Description:
Join Cantina Labs as a Research / ML Engineer in the Speech Team to develop advanced speech and audio generation systems focusing on joint audio-video modeling.
Own the audio side of multimodal generation, including voice cloning, cinematic dialogue, and adjacent speech tasks.
Requirements:
Exceptional experience with large-scale audio models (>8B parameters, >500k hours of data).
Deep hands-on experience with diffusion/flow-matching transformers and audio VAEs, neural codecs, and vocoders.
Strong software engineering skills, particularly in PyTorch, and experience with multi-node, multi-GPU distributed training.
Background in large-scale ML data and experience with voice cloning or expressive speech generation.
Benefits:
Competitive salary range of $200,000-$220,000 and generous company equity.
Comprehensive medical, dental, and vision insurance with 99.99% of premiums covered.
42 days of paid time off, including PTO, sick days, and company holidays.
Additional benefits include generous parental leave, 401(k) plan, lifestyle spending account, and complimentary in-office meals.