Achira
SWE - Distributed
About this role
Achira seeks a Software Engineer to design and build distributed computing infrastructure for machine learning pipelines in drug discovery. You'll optimize large-scale clusters across multiple cloud vendors, focusing on cost efficiency, reliability, and performance for ML training and data generation workflows.
What you'll do
- Architect and implement distributed compute infrastructure for ML data processing and model training
- Optimize cluster observability, scheduling, and resource utilization across CPU/GPU/TPU
- Research and deploy cost-efficient compute solutions using spot instances and multi-cloud strategies
- Develop monitoring and debugging tools for large-scale ML workloads
- Collaborate with ML engineers to reduce training bottlenecks and accelerate pipelines
- Evaluate and integrate emerging distributed computing technologies into the platform
What they're looking for
- Distributed computing frameworks (Ray, Dask, Celery)
- Parallel computing and job scheduling
- Performance profiling and bottleneck identification
- Cloud platforms (AWS, GCP, Azure)
- Cluster orchestration (Kubernetes, Slurm)
- ML frameworks (PyTorch, TensorFlow, JAX)
- MLOps and GPU performance monitoring
- System reliability and cost optimization
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Achira
Achira builds AI-driven drug discovery platforms using deep learning and foundation models for molecular simulation. The company is hiring Machine Learning Research Engineers and Software Engineers to optimize GPU-accelerated model implementations, design scalable distributed infrastructure, and manage large-scale ML pipelines across cloud environments.
View all jobs at AchiraLikely interview questions
- Walk us through your experience building or optimizing distributed computing systems. What frameworks have you worked with (Ray, Dask, Celery, etc.) and what was your role in those projects?
- Describe a time when you had to debug performance bottlenecks in a distributed system. How did you identify the root cause and what was your approach to fixing it?