Skip to main content

Cursor

Software Engineer, Agent Evaluation and Quality

San FranciscofulltimemidAdded 1 month ago

About this role

Build evaluation and quality measurement infrastructure for Cursor's AI coding agent. You'll design datasets, evaluation pipelines, and feedback loops that help the agent improve reliably over time, working across product, data, and engineering teams to turn insights into measurable improvements.

What you'll do

  • Design and build AI evaluation systems with curated datasets, offline replay, scorers, and monitoring dashboards
  • Create feedback loops from real user data to inform model and system improvements
  • Develop analysis tools for debugging agent behavior and identifying failure patterns
  • Define quality metrics and operational guardrails for agent reliability
  • Partner with research, product, and infrastructure teams on quality improvements
  • Build pipelines to analyze agent behavior at scale and surface actionable insights

What they're looking for

  • AI/ML evaluation and measurement systems
  • Data pipeline and infrastructure development
  • Python or similar backend languages
  • SQL and data analysis
  • Metrics design and experimentation
  • Debugging and root cause analysis
  • Cross-team collaboration
  • Knowledge of LLMs and AI agents
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Cursor

Cursor builds an AI-driven code editor used by millions of developers to transform how software is built. The company is hiring for infrastructure engineers, ML systems specialists, enterprise platform builders, security engineers, and customer success roles focused on driving adoption within large organizations.

Website
cursor.com
View all jobs at Cursor

Likely interview questions

  • Describe a time you built an evaluation or measurement system from scratch. How did you decide what metrics to track, and how did you validate they were measuring the right thing?
  • Walk us through how you would design an evaluation system for an AI coding agent. What datasets, metrics, and feedback loops would you prioritize?