AXS
Site Reliability Engineer II
About this role
AXS seeks a Site Reliability Engineer II in Los Angeles to design and maintain scalable infrastructure supporting millions of ticket transactions annually. You'll focus on automation, monitoring, and incident response while collaborating with development teams to ensure high system availability across the ticketing platform.
What you'll do
- Build and scale technology infrastructure to support rapid growth and demand
- Deliver automation solutions for deployment, monitoring, management, and incident response
- Collaborate with cross-functional teams on feature launches and new services
- Lead incident response, root cause analysis, and resolution to minimize downtime
- Develop and improve operational practices and procedures
- Mentor junior engineers on SRE techniques and responsibilities
What they're looking for
- Cloud service operations (monitoring, alerting, data assurance)
- Infrastructure as Code
- Container orchestration platforms
- Continuous Integration/Continuous Delivery (CI/CD)
- Programming or scripting (Python, Bash, Go)
- Cloud-native environment management
- Problem solving and incident management
- System administration and DevOps practices
Benefits
- Medical, dental, and vision insurance
- Paid holidays, vacation, and sick time
- 401(k) plan with 3% employer match
- Parental leave
- Life insurance (company-paid basic and voluntary)
- Flexible spending and health savings account options
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
AXS
AXS builds AI-integrated ticketing and live entertainment platforms that process millions of transactions while leveraging machine learning and LLM capabilities. The company is hiring software engineers, infrastructure specialists, and database engineers to develop scalable backend systems, maintain production infrastructure, and deliver AI-powered user experiences.
View all jobs at AXSLikely interview questions
- Can you walk us through a recent incident you responded to? What was your role, how did you approach root cause analysis, and what preventive measures did you implement?
- Describe your experience with Infrastructure as Code. Which tools have you used (Terraform, CloudFormation, Ansible, etc.) and how did you apply them to solve scalability or deployment challenges?