Skip to main content

Mercor

Software Engineer, Automation

San Francisco$130k–$500kfulltimemidAdded 1 month ago

About this role

Mercor is seeking a Software Engineer for its Automation team in San Francisco, focused on developing autonomous systems for internal operations. The role involves designing production agents that streamline workflows and enhance operational efficiency for AI development.

What you'll do

  • Build and maintain MCP servers for agent access to Mercor's platform
  • Design evaluation frameworks to assess agent decisions
  • Develop action queues and playbook orchestration systems
  • Instrument systems for outcome capture and feedback
  • Collaborate with operators to identify automation opportunities
  • Debug and enhance agent systems for reliability

What they're looking for

  • Strong backend engineering in Python or similar language
  • Experience with LLM APIs and multi-step reasoning
  • Designing evaluation frameworks
  • Quick system iteration skills
  • Understanding of distributed systems and APIs
  • Full-stack experience is a plus

Benefits

  • Bi-annual performance bonus structure
  • Generous equity grant vested over 4 years
  • Up to $15k relocation bonus
  • $10K housing bonus near the office
  • $1.5K monthly meal stipend
  • Free Equinox membership
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Mercor

Mercor builds a marketplace platform connecting expert talent to AI opportunities, supported by identity infrastructure, matching algorithms, and internal tools for data management. The company is hiring Software Engineers, Machine Learning Engineers, Fullstack Engineers, and Security Engineers to develop backend systems, ML models, cloud infrastructure, and distributed platforms.

Website
mercor.io
View all jobs at Mercor

Likely interview questions

  • Walk us through your experience building with LLM APIs. How have you handled tool use, structured outputs, or multi-step reasoning workflows in production?
  • Describe a time you shipped a system that interacts with the real world. How did you think about failure modes, observability, and rollback?