DRW
Site Reliability Engineer - Algorithmic Trading
About this role
DRW, a leading algorithmic trading firm, seeks a Site Reliability Engineer to ensure their trading, testing, and research systems operate flawlessly. You'll build observability and automation to prevent issues, maintain infrastructure stability, and standardize CI/CD processes across teams.
What you'll do
- Provide 24/7 support for production trading systems, test environments, and research pipelines
- Design and implement monitoring, anomaly detection, and alerting to identify issues proactively
- Maintain and automate the test trading environment to support development teams
- Standardize CI/CD and deployment processes to improve delivery speed and reduce friction
- Conduct stress tests and verify system SLAs under load conditions
- Collaborate with traders, engineers, and infrastructure teams across multiple functions
What they're looking for
- Software development or SRE experience (3-10+ years)
- Observability tools (logging, metrics, tracing)
- Scripting languages (Bash, Python, Ruby)
- Linux TCP/IP stack and network protocols
- Network diagnostics and packet capture analysis
- Application and infrastructure-level troubleshooting
- Automation and CI/CD systems
- AI/ML agents for SRE task acceleration
Benefits
- Annual base salary: $130,000–$225,000 plus discretionary bonus
- Medical, pharmacy, dental, and vision insurance
- 401(k) with discretionary employer match
- Short and long-term disability coverage
- Health savings and flexible spending accounts
- Life and AD&D insurance
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
DRW
DRW is a diversified trading firm that builds and maintains trading technology infrastructure, risk management systems, and systematic trading platforms supporting global 24/7 operations across multiple asset classes. The company is hiring Desktop Systems Engineers, Trade Systems Engineers, and Software Engineers to support endpoint infrastructure, trading system reliability, risk analytics, and full-stack trading platform development.
- Website
- drw.com
Likely interview questions
- Walk us through a time you built observability into a production system—what metrics, logs, or traces did you implement, and how did they help you catch issues before they became critical?
- Describe your experience with Linux networking fundamentals. Have you worked with the TCP-IP stack, multicast networking, or network capture tools like tcpdump? Can you give a specific example?