Deepgram
Research Engineer, Machine Learning Systems
About this role
Join Deepgram's Research team as a Machine Learning Engineer to design scalable training systems for speech AI models. You'll prototype novel ideas, build distributed infrastructure for STT/TTS training, and create tools that make ML workflows accessible across the organization while operating in a fast-paced, AI-first environment.
What you'll do
- Architect horizontally scalable systems to accelerate training lifecycles for speech recognition and synthesis models
- Design and implement internal UIs and tools that make ML systems accessible to non-technical stakeholders
- Manage training infrastructure, job orchestration, experiment tracking, and data storage systems
- Optimize data preparation and high-throughput training pipelines for dramatic efficiency gains
- Partner with research scientists to prototype and validate novel modeling approaches
- Build automated evaluation tooling and identify critical experiments to validate ideas quickly
What they're looking for
- Distributed systems and infrastructure design
- Machine learning model training at scale
- Data pipeline and ETL optimization
- Python and systems programming
- Internal tool and UI development
- Experiment tracking and orchestration
- Audio/speech processing experience
- Problem-solving and rapid experimentation
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Deepgram
Deepgram builds voice AI and speech processing technology, offering platforms that power real-time speech recognition and audio intelligence across cloud, edge, and embedded environments. The company is hiring applied ML engineers, backend engineers, embedded AI engineers, and customer-facing roles (customer success and pre-sales engineers) to scale its AI infrastructure and drive enterprise adoption.
- Website
- deepgram.com
Likely interview questions
- Can you describe your experience architecting and scaling distributed training systems? What frameworks and infrastructure have you used, and what bottlenecks did you address?
- Tell us about a time you had to choose between building a quick prototype versus a robust, scalable system. How did you make that decision, and what was the outcome?