Skip to main content

Achira

SWE - Distributed

San Francisco Office (Remote)$164.6k–$259kfulltimemidAdded 1 month ago

About this role

Achira seeks a Software Engineer to design and build distributed computing infrastructure for machine learning pipelines in drug discovery. You'll optimize large-scale clusters across multiple cloud vendors, focusing on cost efficiency, reliability, and performance for ML training and data generation workflows.

What you'll do

  • Architect and implement distributed compute infrastructure for ML data processing and model training
  • Optimize cluster observability, scheduling, and resource utilization across CPU/GPU/TPU
  • Research and deploy cost-efficient compute solutions using spot instances and multi-cloud strategies
  • Develop monitoring and debugging tools for large-scale ML workloads
  • Collaborate with ML engineers to reduce training bottlenecks and accelerate pipelines
  • Evaluate and integrate emerging distributed computing technologies into the platform

What they're looking for

  • Distributed computing frameworks (Ray, Dask, Celery)
  • Parallel computing and job scheduling
  • Performance profiling and bottleneck identification
  • Cloud platforms (AWS, GCP, Azure)
  • Cluster orchestration (Kubernetes, Slurm)
  • ML frameworks (PyTorch, TensorFlow, JAX)
  • MLOps and GPU performance monitoring
  • System reliability and cost optimization
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Achira

Achira builds AI-driven drug discovery platforms using deep learning and foundation models for molecular simulation. The company is hiring Machine Learning Research Engineers and Software Engineers to optimize GPU-accelerated model implementations, design scalable distributed infrastructure, and manage large-scale ML pipelines across cloud environments.

View all jobs at Achira

Likely interview questions

  • Walk us through your experience building or optimizing distributed computing systems. What frameworks have you worked with (Ray, Dask, Celery, etc.) and what was your role in those projects?
  • Describe a time when you had to debug performance bottlenecks in a distributed system. How did you identify the root cause and what was your approach to fixing it?