Zyphra

Research Engineer - Audio & Speech Models

San Francisco, California, United StatesFull timePosted 12 days ago
Apply on Zyphra →

Sign into see who you know at Zyphra.

ZYPHRA IS AN ARTIFICIAL INTELLIGENCE COMPANY BASED IN SAN FRANCISCO, CALIFORNIA.

THE ROLE:

As a Research Engineer - Audio & Speech Models, you will be a core contributor on Zyphra’s Audio Team, building the next generation of open-source autoencoders, ASR, TTS, SSL, and speech-to-speech models. You will be deeply involved in the entire model training process, from data gathering and processing to designing novel architectures and training methodologies.

YOU’LL WORK ACROSS:

- Large-scale audio training runs

- Performance optimization of our training stack

- Audio dataset collection, processing, and evaluation

- Architecture and training methodology ablations and improvements

WHAT WE'RE LOOKING FOR / REQUIREMENTS:

- Strong research taste and intuition. The ability to work through a research project from conception to execution to write-up.

- Strong implementation and prototyping ability (can take an idea from conception to experimentation quickly)

- The ability to work well with others in a high-paced research setting

- Excellent communication and collaboration skills, and can work effectively on both research and engineering implementation at scale

QUALIFICATIONS / ADDITIONAL SKILLS:

- Expertise and intuition for training models in the audio domain, including text-to-speech, ASR, speech-to-speech, speech-emotion-recognition, or other models

- Experience in training audio autoencoders

- Understanding of signal processing, especially of audio signals

- Experience with diffusion models, consistency models, or GANs

- Experience with training on large-scale (multi-node) GPU clusters

- Strong grasp of proper experimental methodology for running rigorous ablations and other hypothesis testing

- Understanding of and interest in large-scale, highly parallel data processing pipelines

- Proficiency with PyTorch and Python

- Experience contributing to large pre-existing codebases and rapidly getting up to speed

- Previously...