Skip to main content

Cohere

Audio Inference Engineer, Model Efficiency

New York (Remote)fulltimemidAdded 1 month ago

About this role

Cohere seeks an Audio Inference Engineer to optimize machine learning model serving for audio workloads, focusing on reducing latency and improving throughput. You'll collaborate with training and infrastructure teams to advance real-time audio inference systems and solve complex performance bottlenecks.

What you'll do

  • Develop and optimize high-performance audio inference systems
  • Identify and resolve bottlenecks in audio model serving pipelines
  • Collaborate with training and serving infrastructure teams on model deployment
  • Advance core metrics including latency, throughput, and quality for audio processing
  • Work on real-time and streaming audio inference architectures
  • Deliver solutions for optimizing audio processing and streaming workloads

What they're looking for

  • C++ and Python programming
  • Deep learning models for audio, speech, or language applications
  • GPU programming and low-level system optimization
  • Machine learning frameworks (PyTorch, TensorFlow, or audio libraries)
  • Inference frameworks (vLLM, SGLang, TensorRT-LLM, or custom systems)
  • Model parallelization across multiple GPUs
  • Real-time streaming architectures
  • Sequence modeling and transformer-based audio systems

Benefits

  • Weekly lunch stipend ($75 or equivalent)
  • Full health, dental, and mental health coverage
  • 6 weeks paid vacation (30 working days)
  • 100% parental leave top-up for up to 6 months
  • Annual enrichment budget for arts, fitness, and professional development
  • Home office stipend and remote-friendly work environment
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Cohere

Cohere builds enterprise AI systems and large language model infrastructure, including audio inference optimization, agentic AI workflows, and petabyte-scale data infrastructure for model training. The company is hiring software engineers, infrastructure specialists, forward-deployed engineers to work with enterprise customers, and IT support staff.

Website
cohere.com
View all jobs at Cohere

Likely interview questions

  • Walk us through a project where you optimized audio or ML inference for latency and throughput. What bottlenecks did you identify and how did you resolve them?
  • Describe your experience with GPU programming and model parallelization. How have you approached optimizing inference across multiple GPUs?