Skip to main content

Dyna Robotics

ML Infrastructure Engineer, Training

Redwood City, CA$220k–$320kfulltimemidAdded 1 month ago

About this role

Dyna Robotics seeks an ML Infrastructure Engineer to design and operate the training systems powering their embodied AI foundation models. You'll architect GPU cluster infrastructure, optimize researcher workflows, handle massive multimodal datasets, and deploy low-latency inference pipelines for real-time robot control.

What you'll do

  • Architect and scale distributed training infrastructure for large GPU clusters with memory optimization techniques
  • Build job scheduling and research codebase systems to enable fast iteration and failure recovery
  • Design high-throughput data pipelines for multimodal robot data (video, proprioception, 3D signals)
  • Develop production inference pipelines with quantization, distillation, and model compilation for real-time robot control
  • Profile and optimize GPU utilization, I/O bottlenecks, and memory fragmentation across compute fleet
  • Own training infrastructure end-to-end to maximize GPU efficiency and reproducibility

What they're looking for

  • PyTorch and distributed training frameworks (DeepSpeed, Accelerate)
  • Cloud GPU environments (GCP/AWS) and Kubernetes orchestration
  • Distributed systems and inter-node communication (NCCL)
  • Memory management and mixed precision training
  • High-performance computing (HPC) systems design
  • Low-latency inference optimization (TensorRT, Triton)
  • Systems profiling and performance optimization
  • Multimodal model architecture (bonus)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Dyna Robotics

Dyna Robotics builds AI-driven robotics platforms that combine embodied AI foundation models with spatial intelligence to enable robots to navigate and operate autonomously across customer sites. The company is hiring software engineers, infrastructure specialists, QA testers, and field deployment engineers to scale its robot deployment systems, training infrastructure, state estimation capabilities, and customer operations.

View all jobs at Dyna Robotics

Likely interview questions

  • Walk us through a time you optimized a distributed training system for GPU utilization. What bottlenecks did you identify, and how did you measure improvement?
  • Describe your experience with PyTorch distributed training frameworks like DeepSpeed or FSDP. How have you approached memory optimization in large-scale models?