Adaption Labs
Distributed Systems Engineer, Data & Inference Platform
About this role
Build and operate distributed systems that serve LLMs at scale and power large-scale data pipelines. You'll optimize inference services for throughput and cost, debug complex production failures, and partner with researchers to turn experimental workloads into reliable systems.
What you'll do
- Design and operate distributed inference systems for LLMs, managing batching, scheduling, KV cache, and autoscaling across GPU fleets
- Build large-scale data pipelines using frameworks like Ray Data or Spark for training and evaluation datasets
- Identify and resolve production failure modes including stragglers, memory fragmentation, and data corruption
- Define SLOs, build observability infrastructure, and own on-call rotation for production systems
- Partner directly with ML engineers and researchers to scale experimental workloads to production
- Write postmortems and implement durable fixes to prevent recurring incidents
What they're looking for
- Distributed systems design and production operations (5+ years)
- Large-scale data/compute frameworks (Ray, Spark, Flink, Beam, or Dask)
- Python and at least one systems language (Go, Rust, C++)
- GPU/accelerator stack knowledge (CUDA, NCCL, mixed precision, memory layout)
- Kubernetes infrastructure and custom operators/schedulers
- Production incident diagnosis and resolution
- LLM inference engines (vLLM, SGLang, TensorRT-LLM, TGI) — bonus
- Modern lakehouse formats (Iceberg, Delta, Hudi) — bonus
Benefits
- Flexible work with Bay Area collaboration and global team options
- Annual travel stipend (Adaption Passport) to explore new countries
- Weekly meal allowance for take-out or grocery delivery
- Comprehensive medical benefits
- Generous paid time off
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Adaption Labs
Adaption Labs builds machine learning systems and infrastructure that solve real-world customer problems at scale, from efficient AI inference to production ML deployments. The company is hiring Applied ML Engineers, distributed systems engineers, and research-focused technologists to develop adaptive AI solutions and bridge the gap between experimental research and reliable, deployed systems.
View all jobs at Adaption LabsLikely interview questions
- Walk us through a production incident you debugged in a distributed system — what was the failure mode, how did you identify it, and what made the fix durable?
- Describe your experience optimizing inference latency and throughput. What levers did you pull (batching, scheduling, caching, autoscaling), and how did you measure impact?