IFM
Research Engineer - The Diffusion LLM Team
About this role
Join MBZUAI's Diffusion LLM Team to design and scale cutting-edge diffusion-based language models that match autoregressive quality while enabling faster generation. You'll collaborate with world-class researchers to develop industrial-scale models and advance inference-time scaling for next-generation AI systems.
What you'll do
- Design, train, and scale large language models for research and production deployment
- Lead or contribute to releasing industrial-scale diffusion language models
- Develop and evaluate training strategies and objectives for efficient model scaling
- Collaborate across architecture, training, and infrastructure teams to implement research ideas
- Publish research findings and contribute to open-source model releases
- Work on inference-time scaling to improve sample quality with additional compute
What they're looking for
- Large-scale model training with modern deep learning frameworks
- Transformer architecture design and optimization
- Large-scale optimization techniques
- LLM pre-training and post-training methodologies
- Model scaling and efficiency optimization
- Diffusion models or discrete diffusion (preferred)
- Research publication and open-source contribution
- Independent and collaborative problem-solving
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
IFM
Institute of Foundation Models conducts research on large-scale foundation models, diffusion-based language models, and world models, building the distributed training infrastructure and MLOps systems required for cutting-edge AI development. The company is hiring research scientists, ML infrastructure engineers, and systems developers to optimize pre-training frameworks, scale distributed training across multi-GPU clusters, and advance inference and experiment capabilities.
View all jobs at IFMLikely interview questions
- Can you walk us through your experience training large language models at scale? What frameworks and infrastructure did you use, and what were the key challenges you faced?
- Tell us about a time you optimized training strategies or objectives for model scaling. How did you measure efficiency, and what trade-offs did you navigate?