Machine Learning Researcher, Audio
MACHINE LEARNING RESEARCHER, AUDIO
Location: San Francisco, CA or Remote
ABOUT BLAND
At Bland.com, our mission is to empower enterprises to build AI phone agents at scale. Based in San Francisco, we are a fast-growing team reimagining how customers interact with businesses through voice. We have raised $65 million from leading Silicon Valley investors, including Emergence Capital, Scale Venture Partners, Y Combinator, and founders of Twilio, Affirm, and ElevenLabs.
Voice is quickly becoming the primary interface between businesses and their customers. We are building the models and infrastructure that make those interactions feel natural, reliable, and genuinely human.
THE ROLE: MACHINE LEARNING RESEARCHER, AUDIO
As a Machine Learning Researcher at Bland, you'll be working on foundational research and development across the core components of our voice stack: speech-to-text, large language models, neural audio codecs, and text-to-speech. Your work will define how our agents understand, reason, and speak in real time at enterprise scale.
This is not a narrow research role. You will take ideas from theory to large-scale training to production inference systems serving millions of calls per day. You will design new modeling approaches, validate them with rigorous experimentation, and collaborate with engineering teams to deploy them into real customer environments.
WHAT YOU WILL DO
Build and Scale Next-Generation TTS Systems
- Design and train large scale text-to-speech models capable of expressive, controllable, human-sounding output.
- Develop neural audio codec-based TTS architectures for efficient, high-fidelity generation.
- Improve prosody modeling, question inflection, emotional expression, and multi-speaker robustness.
- Optimize for real-time, low-latency inference in production.
Advance Speech-to-Text Modeling
- Build and fine-tune large scale ASR systems robust to accents, noise, telephony artifacts, and code...