Skip to main content

Blaxel

Site Reliability Engineer

San Francisco$175k–$250kfulltimemidAdded 1 month ago

About this role

A Site Reliability Engineer is sought to build and operate the core infrastructure powering an AI compute platform, ensuring ultra-low-latency performance and exceptional reliability at scale. You'll architect observability systems, lead incident response, design automation to eliminate operational toil, and collaborate with infrastructure and development teams to keep AI workloads running smoothly.

What you'll do

  • Architect and operate the 25ms cold-start compute engine and core infrastructure systems
  • Build and evolve observability stacks (metrics, traces, logs) with SLO/SLI monitoring
  • Lead incident response with root cause analysis, post-mortems, and systemic fixes
  • Design self-healing, automated systems to reduce toil and enable scaling
  • Perform stress testing, chaos engineering, and performance benchmarking across compute, networking, and storage layers
  • Own infrastructure-layer security practices including sandboxed compute and network isolation

What they're looking for

  • Go, Rust, or Python programming
  • Linux systems, networking fundamentals, and distributed systems
  • Bare-metal server and datacenter operations (PXE, IPMI, RAID, SR-IOV)
  • Kubernetes or container orchestration
  • Observability tools (Prometheus, Grafana, ELK, Datadog)
  • CI/CD pipeline management (GitHub Actions, GitLab CI, Jenkins)
  • AWS or GCP cloud platforms
  • Incident management and debugging
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Blaxel

Blaxel builds an AI-powered cloud computing platform designed to run AI agent infrastructure at scale with ultra-low-latency performance. The company is hiring founding engineers for product design, customer deployment, and site reliability roles to ship features, support strategic customers, and operate core infrastructure.

View all jobs at Blaxel

Likely interview questions

  • Walk us through a time you debugged a production incident in a distributed system. What was your approach, and how did you prevent it from happening again?
  • Describe your experience optimizing latency-critical systems. How would you approach achieving and maintaining 25ms cold-start times at scale?