Anthropic
Research Engineer, Cybersecurity RL (Reinforcement Learning)
About this role
Anthropic seeks a Research Engineer to advance reinforcement learning for cybersecurity applications, combining expertise in secure coding and vulnerability remediation with AI safety. The role bridges research innovation and production engineering within the Horizons team, requiring both novel RL approach development and hands-on implementation.
What you'll do
- Design and implement reinforcement learning environments for cybersecurity tasks
- Conduct experiments and evaluate RL-based approaches for secure coding and vulnerability remediation
- Develop novel RL techniques tailored to defensive cybersecurity domains
- Integrate research outcomes into production AI model training pipelines
- Collaborate with researchers, engineers, and cybersecurity specialists across teams
- Balance exploratory research with practical software engineering implementation
What they're looking for
- Cybersecurity research and domain knowledge
- Reinforcement learning techniques and environments
- Machine learning and deep learning
- Strong software engineering and implementation
- Python or similar programming languages
- LLM training methodologies
- Security engineering and defensive workflows
- Research design and experimentation
Benefits
- Annual salary range: $300,000–$405,000 USD
- Visa sponsorship available
- Hybrid work policy (minimum 25% office time)
- Work on cutting-edge AI safety and beneficial AI systems
- Collaboration with world-class researchers and engineers
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Anthropic
Anthropic builds Claude, an AI assistant, and is hiring for engineering roles across infrastructure, data systems, and security that support both AI research operations and the company's internal technology needs. The company seeks infrastructure engineers, systems integrators, data scientists, and security specialists to build production-scale systems for training data pipelines, financial operations, developer productivity measurement, research infrastructure, and server firmware security.
- Website
- anthropic.com
Likely interview questions
- Can you walk us through a cybersecurity research project or vulnerability you've worked on, and how you'd approach turning that into an RL training signal?
- Describe your experience designing and implementing environments—how would you build an RL environment for secure coding or vulnerability remediation?