Skip to main content

IFM

Machine Learning Engineer – World Model

Sunnyvale, CA$150k–$450kfull-timemidAdded 1 month ago

About this role

Join a research lab at MBZUAI's Institute of Foundation Models to develop ML infrastructure and MLOps systems for cutting-edge world model research. You'll design scalable cloud systems that enable researchers to conduct experiments and advance next-generation AI capabilities.

What you'll do

  • Design and operate scalable cloud infrastructure for ML research and experimentation
  • Build and maintain data pipelines for model training and evaluation
  • Develop MLOps systems to support researcher workflows
  • Ensure reliability, observability, and performance of research infrastructure
  • Collaborate with researchers and engineers on infrastructure requirements
  • Support deployment and scaling of foundation model experiments

What they're looking for

  • Cloud infrastructure and DevOps
  • MLOps and ML systems design
  • Data pipeline engineering
  • Python and software engineering
  • Distributed systems and scaling
  • Infrastructure as Code
  • Monitoring and observability tools
  • Problem-solving and collaboration
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

IFM

Institute of Foundation Models conducts research on large-scale foundation models, diffusion-based language models, and world models, building the distributed training infrastructure and MLOps systems required for cutting-edge AI development. The company is hiring research scientists, ML infrastructure engineers, and systems developers to optimize pre-training frameworks, scale distributed training across multi-GPU clusters, and advance inference and experiment capabilities.

View all jobs at IFM

Likely interview questions

  • Describe your experience designing and deploying ML infrastructure at scale. What challenges have you encountered with distributed training, and how did you solve them?
  • Tell us about a time you built data pipelines for machine learning research. How did you ensure reliability and observability?