Skip to main content

Blockit AI

Software Engineer, AI Agents

San FranciscofulltimemidAdded 1 month ago

About this role

Blockit is hiring a Software Engineer to build AI agents that autonomously coordinate meetings and scheduling across multiple people and timezones. You'll own the core intelligence of the platform, designing agent architectures, refining prompts, building evaluation frameworks, and shipping new capabilities as the system handles increasingly complex coordination problems.

What you'll do

  • Design and iterate on agent architectures for autonomous scheduling coordination
  • Write and refine prompts for orchestrator and sub-agent systems
  • Build and maintain evaluation frameworks to measure agent quality and catch regressions
  • Debug agent failures and understand why decisions or interpretations went wrong
  • Implement new agent capabilities as user needs and model improvements evolve
  • Instrument and analyze agent behavior to identify patterns and failure modes

What they're looking for

  • Production software development and ownership (2+ years)
  • Backend engineering across the stack
  • LLM systems and prompt engineering
  • Agent architecture design and evaluation
  • Debugging and systems analysis
  • Clear technical communication
  • Python or similar backend languages
  • Testing and instrumentation
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Blockit AI

Blockit AI builds an autonomous AI agent that revolutionizes scheduling and time coordination across multiple surfaces like Slack and mobile. The company is hiring Software Engineers across Product & Growth, Infrastructure, and AI to develop agent capabilities, backend systems, and core intelligence for autonomous meeting coordination.

View all jobs at Blockit AI

Likely interview questions

  • Walk us through a production LLM system you've built or worked on—what were the biggest failure modes you encountered, and how did you debug them?
  • How would you approach building an evaluation framework to measure the quality of an AI agent that needs to handle complex, multi-party scheduling coordination?