Machine Learning Engineer, Speech/Audio
Software Engineering, Data Science
San Francisco, CA, USA
Machine Learning Engineer, Speech/Audio
San Francisco·Full-time·Machine Learning
You will own the audio modeling stack end to end, from raw signal to production inference. You are the person who ships the models that decide, in real time, who a voice agent is being spoken to.
What you will do
- Own the audio model lifecycle end to end: data, features, architecture, training, evaluation, and the inference path that runs in production.
- Build and improve the models behind addressee detection, deciding whether speech is meant for the agent or for someone else in the room.
- Turn research prototypes into fast, dependable inference that holds up under real world noise, overlap, and accents.
- Design the evaluation harness and datasets that tell us, honestly, whether a change made the product better.
- Work directly with founders and early customers to shape what the model needs to do next.
What we are looking for
- Strong applied machine learning experience with audio, speech, or other signal-heavy models.
- Comfort owning a model from data collection through to something running in production, not just a notebook.
- Fluency with modern deep learning tooling such as PyTorch, and the discipline to measure before and after every change.
- A bias toward shipping, and toward the smallest experiment that answers the question.
- Enough systems sense to care about latency, memory, and how a model behaves under load.
- Bonus: experience with real time audio, source separation, or on-device inference.
About attention labs
attention labs is early: a small team defining a new category at the intersection of speech, cognitive neuroscience, and machine learning. We work from San Francisco, Toronto, and Memphis, and we are remote-friendly for the right person.
Sound like you?
Send a note and tell us why this role fits. We read every application.
Apply for this role