Anthropic
Research Engineer, Performance RL (Reinforcement Learning)
About this role
Join Anthropic's Code RL team to advance AI models' ability to write correct, performant code for accelerators using reinforcement learning. You'll design RL environments, conduct experiments, and deliver research into production training runs while collaborating across teams.
What you'll do
- Design and implement RL environments and evaluation systems for accelerator code generation
- Conduct experiments to advance code generation capabilities and shape research roadmap
- Integrate research innovations into model training pipelines
- Collaborate with researchers, engineers, and performance specialists across Anthropic
- Translate accelerator performance knowledge into learnable tasks and reward signals
What they're looking for
- Accelerator programming (CUDA, ROCm, Triton, Pallas)
- ML framework expertise (JAX or PyTorch)
- Full-stack development (kernels, model code, distributed systems)
- Reinforcement learning
- LLM training methodologies
- ML workload optimization and porting
- Research and engineering implementation balance
Benefits
- Annual salary: $350,000–$850,000 USD
- Hybrid work policy (minimum 25% in-office)
- Visa sponsorship available
- Work on cutting-edge AI safety and capability research
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Anthropic
Anthropic builds Claude, an AI assistant, and is hiring for engineering roles across infrastructure, data systems, and security that support both AI research operations and the company's internal technology needs. The company seeks infrastructure engineers, systems integrators, data scientists, and security specialists to build production-scale systems for training data pipelines, financial operations, developer productivity measurement, research infrastructure, and server firmware security.
- Website
- anthropic.com
Likely interview questions
- Walk us through your experience optimizing code for accelerators like CUDA, ROCm, or Triton. What performance metrics did you focus on and how did you measure improvement?
- Describe a time you designed an evaluation or task from scratch. How did you determine what signals were important to measure, and how did you validate your approach?