Skip to main content

Adaption Labs

Distributed Systems Engineer, Data & Inference Platform

San Francisco (Remote)fulltimemidAdded 1 month ago

About this role

Build and operate distributed systems that serve LLMs at scale and power large-scale data pipelines. You'll optimize inference services for throughput and cost, debug complex production failures, and partner with researchers to turn experimental workloads into reliable systems.

What you'll do

  • Design and operate distributed inference systems for LLMs, managing batching, scheduling, KV cache, and autoscaling across GPU fleets
  • Build large-scale data pipelines using frameworks like Ray Data or Spark for training and evaluation datasets
  • Identify and resolve production failure modes including stragglers, memory fragmentation, and data corruption
  • Define SLOs, build observability infrastructure, and own on-call rotation for production systems
  • Partner directly with ML engineers and researchers to scale experimental workloads to production
  • Write postmortems and implement durable fixes to prevent recurring incidents

What they're looking for

  • Distributed systems design and production operations (5+ years)
  • Large-scale data/compute frameworks (Ray, Spark, Flink, Beam, or Dask)
  • Python and at least one systems language (Go, Rust, C++)
  • GPU/accelerator stack knowledge (CUDA, NCCL, mixed precision, memory layout)
  • Kubernetes infrastructure and custom operators/schedulers
  • Production incident diagnosis and resolution
  • LLM inference engines (vLLM, SGLang, TensorRT-LLM, TGI) — bonus
  • Modern lakehouse formats (Iceberg, Delta, Hudi) — bonus

Benefits

  • Flexible work with Bay Area collaboration and global team options
  • Annual travel stipend (Adaption Passport) to explore new countries
  • Weekly meal allowance for take-out or grocery delivery
  • Comprehensive medical benefits
  • Generous paid time off
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Adaption Labs

Adaption Labs builds machine learning systems and infrastructure that solve real-world customer problems at scale, from efficient AI inference to production ML deployments. The company is hiring Applied ML Engineers, distributed systems engineers, and research-focused technologists to develop adaptive AI solutions and bridge the gap between experimental research and reliable, deployed systems.

View all jobs at Adaption Labs

Likely interview questions

  • Walk us through a production incident you debugged in a distributed system — what was the failure mode, how did you identify it, and what made the fix durable?
  • Describe your experience optimizing inference latency and throughput. What levers did you pull (batching, scheduling, caching, autoscaling), and how did you measure impact?