Skip to main content

Anthropic

ML/Research Engineer, Safeguards

San Francisco, CA | New York City, NYFrom $500kmidAdded 1 month ago

About this role

Anthropic is hiring ML/Research Engineers to build systems that detect and prevent misuse of AI products, including coordinated attacks and harmful behaviors. You'll develop classifiers, monitoring systems, and defenses while conducting research on adversarial robustness and red-teaming.

What you'll do

  • Develop classifiers and synthetic data pipelines to detect misuse and anomalous behavior at scale
  • Build monitoring systems to identify coordinated harms across multiple interactions
  • Evaluate and improve safety of agentic products, including threat modeling and prompt injection defenses
  • Conduct research on automated red-teaming and adversarial robustness techniques
  • Work across research-to-deployment pipeline from experiments to production systems

What they're looking for

  • Machine learning engineering or applied research (4+ years)
  • Python proficiency
  • ML systems development
  • Language modeling and transformers
  • Anomaly detection and behavioral ML
  • Adversarial machine learning or red-teaming
  • Interpretability research
  • Large-scale ML systems

Benefits

  • Annual compensation: $350,000–$500,000 USD
  • Visa sponsorship available
  • Hybrid work policy (minimum 25% office time)
  • Work on AI safety and responsible scaling initiatives
  • Collaborative research environment with policy and business experts
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Anthropic

Anthropic builds Claude, an AI assistant, and is hiring for engineering roles across infrastructure, data systems, and security that support both AI research operations and the company's internal technology needs. The company seeks infrastructure engineers, systems integrators, data scientists, and security specialists to build production-scale systems for training data pipelines, financial operations, developer productivity measurement, research infrastructure, and server firmware security.

View all jobs at Anthropic

Likely interview questions

  • Walk us through your experience building classifiers or anomaly detection systems at scale. What were the main challenges and how did you evaluate their effectiveness?
  • Describe a time you worked on adversarial machine learning or red-teaming. What attack vectors did you explore and how did you measure robustness?