Discord
Software Engineer, Distributed Systems
About this role
Join Discord's Realtime Infrastructure team to design and maintain critical distributed systems powering millions of daily active users. This role combines backend engineering with infrastructure operations to ensure Discord's reliability and enable new product features.
What you'll do
- Build and operate large-scale, reliable distributed systems
- Collaborate with product teams to develop new features
- Maintain critical tier 0 services and infrastructure
- Write backend code and manage infrastructure operations
- Implement monitoring and alerting systems
- Troubleshoot complex distributed system problems
What they're looking for
- Backend systems design and development
- Distributed systems problem-solving
- Critical service operations and maintenance
- Monitoring and alerting practices
- Open source software knowledge
- Cloud environments (GCP, AWS)
- DevOps tools (Terraform, Kubernetes, Salt)
- Elixir programming (bonus)
Benefits
- Equity compensation
- Competitive salary ($160,000-$180,000 base)
- Health and wellness benefits
- Relocation assistance available
- Impact on one of the world's largest communication platforms
- Work with experienced distributed systems engineers
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Discord
Discord builds a real-time communication platform serving millions of daily active users, with infrastructure handling trillions of messages across gaming and community use cases. The company is hiring backend engineers, infrastructure specialists, and data engineers to develop distributed systems, database infrastructure, developer tools, and large-scale data platforms that power its core platform.
- Website
- discord.com
Likely interview questions
- Tell us about a complex distributed system problem you've solved. What was the challenge, and how did you approach debugging it?
- Describe your experience operating and maintaining critical tier 0 services. How do you handle incidents and ensure reliability?