Skip to main content

Astronomer

Customer Reliability Engineer - Infrastructure

Remote (United States) (Remote)$125k–$130kfulltimemidAdded 1 month ago

About this role

Join Astronomer's Customer Reliability Engineering team as an Infrastructure Specialist to ensure the stability and performance of their managed Airflow platform. You'll handle incident response, manage cloud infrastructure and Kubernetes clusters, and work directly with customers across diverse industries to solve complex reliability challenges.

What you'll do

  • Respond to and triage customer incidents and monitoring alerts across cloud infrastructure and Kubernetes environments
  • Troubleshoot customer production environments and provide technical solutions to ensure platform reliability
  • Participate in on-call rotation including weekend coverage
  • Build and maintain monitoring, alerting, and automation systems to improve operational efficiency
  • Gather customer feedback and communicate infrastructure needs to product development teams
  • Partner with customers on production readiness, SLA management, and documentation

What they're looking for

  • Kubernetes (3+ years)
  • Cloud platforms (AWS, GCP, Azure)
  • Linux system administration
  • Distributed systems troubleshooting and monitoring
  • DevOps and CI/CD practices
  • Python scripting
  • Customer-facing communication
  • Site Reliability Engineering (preferred)

Benefits

  • Competitive salary ($125,000-$130,000 estimated)
  • Equity compensation
  • Comprehensive benefits package
  • Fully remote, distributed team
  • Exposure to cutting-edge multi-cloud technology
  • Direct customer impact and meaningful work
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Astronomer

Astronomer builds a DataOps platform powered by Apache Airflow that helps enterprises manage data workflows at scale. The company is hiring for technical customer-facing roles including Customer Reliability Engineers, Sales Engineers, Field Engineers, and Infrastructure Specialists who support enterprise customers' Airflow adoption and platform reliability.

View all jobs at Astronomer

Likely interview questions

  • Walk us through a time you diagnosed and resolved a critical incident in a production Kubernetes cluster. What was your troubleshooting approach?
  • Describe your experience operating distributed systems at scale across cloud providers. What monitoring and alerting strategies have you found most effective?