Deepgram
ML Ops Infrastructure Engineer
About this role
Deepgram seeks an ML Ops Infrastructure Engineer to build the production systems that bring speech AI models from research to serving millions of API requests. You'll design CI/CD pipelines, deployment infrastructure, monitoring, and testing systems that enable rapid, safe iteration on voice AI models at scale.
What you'll do
- Design and build ML-specific CI/CD pipelines for model validation and deployment across environments
- Architect model deployment systems with A/B testing, versioning, and rollback capabilities
- Implement production monitoring for model performance, drift detection, and regression alerts
- Develop automated retraining pipelines triggered by data changes or performance degradation
- Create build and test environments that mirror production for high-fidelity researcher feedback
- Optimize model serving infrastructure for latency, throughput, and cost efficiency
What they're looking for
- MLOps and ML infrastructure engineering
- Python
- CI/CD systems and pipeline development
- Docker and Kubernetes
- ML model deployment and serving
- Monitoring and observability tools
- Data infrastructure and automation
- Performance optimization
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Deepgram
Deepgram builds voice AI and speech processing technology, offering platforms that power real-time speech recognition and audio intelligence across cloud, edge, and embedded environments. The company is hiring applied ML engineers, backend engineers, embedded AI engineers, and customer-facing roles (customer success and pre-sales engineers) to scale its AI infrastructure and drive enterprise adoption.
- Website
- deepgram.com
Likely interview questions
- Walk us through a CI/CD pipeline you've built for ML model deployment. How did you handle the differences between deploying traditional software versus deploying ML models?
- Describe your experience with model monitoring and drift detection in production. What metrics did you track, and how did you alert on model performance degradation?