Basis Research Institute
ML Systems Engineer, Infrastructure & Cloud
About this role
Basis seeks an ML Systems Engineer to build and maintain scalable training infrastructure for a nonprofit AI research organization. You'll manage distributed GPU clusters, optimize cloud resources, and ensure reliable, reproducible ML experiments from development through production.
What you'll do
- Own distributed training infrastructure including job launchers, checkpointing, and recovery mechanisms
- Debug and resolve training failures across GPUs, networking, numerics, and data pipelines
- Profile and optimize training performance and resource utilization
- Manage GPU clusters and cloud infrastructure with cost optimization and security best practices
- Build reproducible experiment infrastructure and monitoring systems
- Maintain comprehensive documentation of issues, solutions, and operational lessons learned
What they're looking for
- Distributed training frameworks (PyTorch DDP/FSDP, JAX)
- Cloud administration (AWS/GCP/Azure, Kubernetes, Terraform)
- GPU cluster management and distributed systems debugging
- Mixed precision training, gradient accumulation, and checkpoint/recovery systems
- Full ML stack understanding (hardware to training loops)
- Infrastructure as code and CI/CD practices
- Debugging numerical instabilities and convergence problems
- Knowledge of optimization theory and numerical methods
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Basis Research Institute
Basis Research Institute is a nonprofit AI research organization building trustworthy data infrastructure, scalable ML training systems, custom robotic platforms, and LLM-powered automation tools to support both research and commercial applications. The organization is hiring Data Engineers, ML Systems Engineers, Hardware Systems Engineers, Software Engineers, and Machine Learning Research Engineers to develop its technical infrastructure and translate research into production systems.
View all jobs at Basis Research InstituteLikely interview questions
- Walk us through your experience managing distributed training jobs at scale. What's the largest GPU cluster you've worked with, and what challenges did you encounter with synchronization, communication, or fault tolerance?
- Describe a time you debugged a complex numerical instability or training failure that spanned multiple components of the ML stack. How did you isolate the problem and what was your approach?