Astera
Site Reliability Engineer
About this role
Astera's Neuro-AI program seeks a Site Reliability Engineer to manage and optimize the digital infrastructure supporting cutting-edge AI research. You'll own compute resource access, monitoring, auto-scaling, and automation across a modern cloud-native stack, ensuring researchers have reliable, efficient access to the tools they need.
What you'll do
- Manage compute resource allocation and access control across third-party cloud platforms
- Monitor cluster health and resource utilization with observability tools
- Design and implement auto-scaling solutions based on research demand patterns
- Automate operational processes to increase infrastructure efficiency
- Ensure reproducible and deterministic deployment environments for research
- Maintain clear operational documentation and boundaries for handoff to other engineers
What they're looking for
- Kubernetes and container orchestration
- Infrastructure automation (Ansible or similar)
- Observability and monitoring (Prometheus, Grafana)
- Linux systems administration
- Python scripting
- Cloud networking and distributed systems knowledge
- Docker and containerization
- Access control and security management
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Astera
Astera builds advanced infrastructure for large-scale distributed simulations, neuroscience research tools, and knowledge-sharing platforms aimed at accelerating scientific discovery. The company is hiring software engineers, ML researchers, and scientist-engineers to develop high-performance systems, specialized neuroscience instrumentation, and data-efficient machine learning approaches.
View all jobs at AsteraLikely interview questions
- Walk us through how you've designed or operated a Kubernetes cluster in production. What were the biggest reliability challenges you faced?
- Describe a time when you had to balance operational stability with supporting experimental or research workloads. How did you handle the trade-offs?