Anyscale
Distributed LLM Inference Engineer
About this role
Anyscale is seeking a Distributed LLM Inference Engineer to optimize and build high-performance inference systems at scale using Ray. You'll work across the stack integrating Ray Data with LLM engines, collaborate with open-source communities like vLLM, and ship end-to-end solutions for batch and online inference.
What you'll do
- Develop and optimize batch and online inference solutions at scale for Ray users and Anyscale customers
- Integrate Ray Data with LLM engines to achieve cost-effective large-scale ML inference
- Collaborate with open-source projects like vLLM and contribute improvements back to the community
- Implement state-of-the-art techniques from research and open-source communities
- Work across the full stack to balance throughput, latency, and cost for inference workloads
- Partner with product teams to ship end-to-end solutions quickly
What they're looking for
- Large-scale ML inference optimization
- Distributed systems design
- Deep learning frameworks (PyTorch, TensorFlow)
- Python programming
- Ray framework experience
- GPU/CUDA programming
- ML systems knowledge
- LLM inference engines (vLLM, TensorRT-LLM)
Benefits
- Stock options
- Healthcare plans with 99% premium coverage for employees and dependents
- 401k retirement plan
- Education and wellbeing stipend
- Paid parental leave and fertility benefits
- Paid time off
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Anyscale
Anyscale builds Ray, an open-source distributed computing framework and enterprise platform for scaling AI workloads across Kubernetes and cloud providers. The company is hiring forward-deployed engineers to work embedded with customers, software engineers to develop Ray Core, LLM inference specialists, and customer support engineers who combine technical expertise with post-sale success.
- Website
- anyscale.com
Likely interview questions
- Can you describe your experience optimizing ML inference for throughput and latency at scale? What were the key bottlenecks you encountered?
- How familiar are you with vLLM or similar LLM inference engines, and have you contributed to or integrated with open-source ML systems?