DoorDash USA
Software Engineer, Spark Platform
About this role
DoorDash is seeking a Software Engineer to build and operate an internal Apache Spark platform serving the entire company. You'll work across runtime optimization, multi-tenant scheduling, cluster automation, and observability for a distributed system running at massive scale.
What you'll do
- Develop and maintain in-house Spark platform runtime, scheduler, and reliability infrastructure
- Design multi-tenant scheduling and executor bin-packing algorithms for efficient resource utilization
- Automate cluster lifecycle management including provisioning, upgrades, and failure handling
- Build observability and incident automation systems to support on-call operations
- Collaborate with platform consumers across the company to solve high-impact problems
- Optimize Spark performance and cost-aware resource placement
What they're looking for
- Apache Spark platform operations at scale
- Kubernetes production systems and multi-tenant cluster management
- Batch or big-data schedulers (YuniKorn, Volcano, Kueue, Spark-on-Kubernetes)
- Observability tools (Prometheus, OpenTelemetry, distributed tracing, structured logging)
- AWS cloud infrastructure and networking
- Python, Go, Scala, or Java programming
- SQL
- SLO/SLI definition and measurement
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
DoorDash USA
DoorDash USA is building autonomous delivery systems including drones and robots, along with internal infrastructure platforms to support large-scale operations. The company is hiring robotics engineers, autonomous systems specialists, infrastructure engineers, and platform software engineers to develop flight control systems, mapping and localization capabilities, and distributed computing platforms.
View all jobs at DoorDash USALikely interview questions
- Tell us about your experience operating Apache Spark at scale — what was the largest deployment you've worked with, and what were the biggest operational challenges you faced?
- Walk us through a time you optimized multi-tenant scheduling or resource allocation on a distributed system. What tradeoffs did you have to navigate?