Anthropic
Research Engineer, Pretraining Scaling
About this role
Anthropic seeks a Research Engineer to own critical aspects of production pretraining pipelines, balancing deep technical work on model training systems with operational responsibilities during launches. You'll debug complex full-stack issues, optimize training efficiency, and collaborate across teams to ensure frontier models train reliably at scale.
What you'll do
- Own production pretraining pipeline including model operations, performance optimization, observability, and reliability
- Debug and resolve issues across hardware, networking, training dynamics, and evaluation infrastructure
- Design and run experiments to improve training efficiency, reduce step time, and enhance model performance
- Respond to on-call incidents during model launches with rapid diagnosis and cross-team coordination
- Build and maintain production logging, monitoring dashboards, and evaluation infrastructure
- Add new capabilities to training codebase such as long context support or novel architectures
What they're looking for
- Large-scale machine learning systems and distributed training
- JAX, TPU, PyTorch, or equivalent ML frameworks at scale
- Full-stack debugging across hardware, networking, and software layers
- Production ML systems and observability tools
- Experimental design and systems optimization
- Clear communication and cross-team collaboration
- LLM pretraining experience
- Systems engineering or operational excellence background
Benefits
- Hands-on experience with some of the largest training runs in the industry
- Work alongside world-class researchers and engineers at a mission-driven company
- Unique learning opportunities and institutional knowledge building
- 5 days per week in-office at San Francisco headquarters
- Involvement in work directly shaping safe and beneficial AI systems
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Anthropic
Anthropic builds Claude, an AI assistant, and is hiring for engineering roles across infrastructure, data systems, and security that support both AI research operations and the company's internal technology needs. The company seeks infrastructure engineers, systems integrators, data scientists, and security specialists to build production-scale systems for training data pipelines, financial operations, developer productivity measurement, research infrastructure, and server firmware security.
- Website
- anthropic.com
Likely interview questions
- Walk us through a time you debugged a complex issue across multiple layers of a system—from hardware to software to training dynamics. How did you approach it?
- Tell us about your hands-on experience training large language models or working at scale with JAX, TPU, or PyTorch. What was the biggest scaling challenge you faced?