Skip to main content

Crusoe

Software Engineer II, Managed Platform Services

San Francisco, CA - USfulltimemidAdded 1 month ago

About this role

Crusoe is seeking a Software Engineer II to design and scale customer-facing cloud platforms and managed services that power AI infrastructure. You'll own foundational infrastructure components, lead cross-functional projects, and drive operational excellence across the platform.

What you'll do

  • Design and build scalable, reliable core infrastructure services for the cloud platform
  • Own end-to-end development and maintenance of specific software modules and service components
  • Lead full lifecycle feature development from business requirements to production deployment
  • Collaborate with engineering, support, and product teams to assess tools and solutions
  • Conduct root cause analysis on production incidents and implement systemic fixes
  • Mentor teammates and provide code review feedback to ensure high-quality standards

What they're looking for

  • Distributed systems design and fault-tolerance
  • Microservices architecture and cloud infrastructure
  • Docker, Kubernetes, Terraform, and CI/CD systems
  • Observability tools (time-series databases, log aggregation, distributed tracing)
  • Performance optimization and system troubleshooting
  • Software design patterns and refactoring
  • Root cause analysis and incident resolution
  • Cross-functional collaboration and mentorship
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Crusoe

Crusoe builds AI infrastructure and data center systems, including cloud platforms, modular facilities, and manufacturing operations. The company is hiring engineers across software, mechanical engineering, facilities management, instrumentation, and CNC programming to design, optimize, and operate its mission-critical infrastructure.

View all jobs at Crusoe

Likely interview questions

  • Walk us through a time you designed and deployed a scalable distributed system or managed cloud service. What were the key architectural decisions you made, and how did you ensure fault tolerance?
  • Tell us about your experience with containerization and orchestration tools like Docker and Kubernetes. How have you used these in production environments?