Skip to main content

Avride

Software Engineer – ML Platform

Texas, USmidAdded 1 month ago

About this role

Join Avride's ML Platform team to build infrastructure that powers large-scale ML training for autonomous driving. You'll design and optimize the orchestration, distributed compute, and resource governance systems that enable ML teams to train models efficiently at scale on Kubernetes.

What you'll do

  • Build and scale ML compute platform on Kubernetes using Argo Workflows for training and data processing orchestration
  • Design resource governance systems including scheduling, quotas, and policy enforcement across GPU, CPU, memory, and IO
  • Optimize end-to-end training throughput by improving data access patterns, caching, and removing infrastructure bottlenecks
  • Partner with ML teams to debug complex workload issues and implement platform-level solutions
  • Evaluate and integrate open-source tools like Argo Workflows, Ray, and Kubernetes ecosystem components

What they're looking for

  • Python or Go (C++ a plus)
  • Kubernetes architecture and scheduling
  • Distributed systems design and implementation
  • Linux systems debugging and performance optimization
  • Production service operation and observability
  • Networking and storage/IO troubleshooting
  • Argo Workflows, Ray, or similar ML tooling
  • GPU scheduling and distributed training experience
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Avride

Avride develops autonomous vehicle and delivery robot technology, building the computational infrastructure, data systems, and localization platforms that enable autonomous driving development. The company is hiring backend engineers, ML infrastructure specialists, data platform engineers, robotics software engineers, and site infrastructure engineers to scale its simulation, training, logging, and operational systems.

View all jobs at Avride

Likely interview questions

  • Walk us through a complex production incident you've debugged in a distributed system. How did you approach the investigation, and what tools or techniques were critical to finding the root cause?
  • Describe your experience with Kubernetes in production. How have you dealt with scheduling constraints, resource contention, or unexpected pod behavior under high load?