Lila Sciences
Research Engineer, Frontier Capabilities
About this role
LILA is seeking a Research Engineer to contribute to AI advancements in long-horizon scientific tasks by optimizing training processes and infrastructure for large language models. Candidates can choose to focus on areas such as GPU optimization, stack and infrastructure management, model experimentation, or agentic capabilities development.
What you'll do
- Maximize hardware utilization for RL training runs
- Own and manage post-training infrastructure
- Lead experimentation on reasoning models
- Design and build scientific benchmarks and dashboards
- Train models for planning and tool use in scientific contexts
What they're looking for
- Strong software engineering skills in Python
- Experience with distributed ML training frameworks
- Understanding of large-scale model training techniques
- Experience with cloud or HPC environments
- Ability to communicate technical results
Benefits
- Competitive base compensation with bonus potential
- Generous early-stage equity
- Comprehensive medical, dental, and vision coverage
- Flexible time off with company-wide holidays
- Paid parental leave
- Educational assistance program
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Lila Sciences
Lila Sciences builds AI and automation systems for scientific research, including tools for automated analysis, control systems for lab operations, and large language models for scientific tasks. The company is hiring software engineers, machine learning scientists, research engineers, and automation specialists to develop and maintain these scientific computing platforms.
View all jobs at Lila SciencesLikely interview questions
- Walk us through your experience with distributed ML training frameworks like Megatron-LM, TorchTitan, or DeepSpeed. What performance bottlenecks have you identified and optimized?
- Describe your hands-on experience training models at 100B+ parameters. What were the key challenges and how did you address them?