Kalshi
Site Reliability Engineer
About this role
Join Kalshi's engineering team to build infrastructure for a rapidly growing prediction market platform. As a Site Reliability Engineer, you'll design and maintain highly available systems, improve observability, and drive reliability practices across the organization while working on greenfield projects with significant ownership and impact.
What you'll do
- Define and measure key metrics to improve observability, reliability, and service availability
- Build automation and systems to reduce operational toil and burden
- Collaborate with infrastructure engineers to optimize cloud deployments and performance-tune systems
- Identify reliability problems across the stack and implement long-term software improvements
- Debug complex technical issues and participate in on-call rotations for incident response
- Review feature designs for security, scalability, and architectural soundness
What they're looking for
- Production service design and scaling
- System design and performance tuning
- Cloud infrastructure (Docker, Kubernetes, Terraform, EC2)
- Observability and monitoring
- High-quality coding and testing practices
- Debugging and troubleshooting
- Rust, Go, or similar languages (bonus)
- Datadog or similar monitoring tools (bonus)
Benefits
- Salary: $100,000–$250,000 annually plus equity
- Significant ownership of greenfield infrastructure projects
- Opportunity to mentor and influence engineering culture
- Fast-moving, small team environment with visible impact
- Work on financial systems at scale
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Kalshi
Kalshi builds a prediction market trading platform and exchange infrastructure. The company is hiring frontend engineers, backend engineers, site reliability engineers, infrastructure engineers, and mobile engineers to scale its rapidly growing financial trading systems.
View all jobs at KalshiLikely interview questions
- Tell us about a time you designed and scaled a production service to handle high throughput and low latency. What were the key challenges and how did you measure success?
- Describe your experience with observability and incident response. How do you approach debugging complex issues across distributed systems?