Instawork
ML Engineer - Instawork Robotics
San Francisco, California, United States$170k–$200kmidAdded 1 month ago
About this role
Instawork Robotics seeks an ML Engineer to develop and scale data pipelines for training physical AI and robotics foundation models. You'll design labeling systems, optimize data workflows, and collaborate with robotics teams to deliver high-quality datasets for frontier AI labs.
What you'll do
- Design and maintain data labeling and enrichment pipelines for robotics training datasets
- Identify and implement improvements to data pipeline efficiency, scalability, and quality
- Develop methodologies to measure dataset quality and ML model performance
- Monitor academic research and industry best practices in robotics learning
- Collaborate cross-functionally with robotics, data ops, QA, and engineering teams
- Translate frontier research techniques into production systems
What they're looking for
- Machine learning and computer vision (robotics or autonomous vehicles focus)
- Distributed systems and cloud computing (AWS)
- Large-scale data pipeline architecture
- Python or similar ML languages
- Production system design and deployment
- Data quality measurement and analytics
- Cross-functional technical communication
- Problem-solving with incomplete information
Benefits
- Salary: $170,000–$200,000 (CA-based)
- Stock option equity
- Medical, dental, and vision coverage
- Flexible paid time off and 8+ paid holidays
- 401(k), HSA contributions, and flexible spending plans
- Phone and commuter stipends
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Instawork
Instawork Robotics develops physical AI and robotics foundation models powered by machine learning and high-quality datasets. The company is hiring ML Engineers to build data pipelines, labeling systems, and data workflows that enable their robotics AI capabilities.
- Website
- instawork.com
Likely interview questions
- Describe your experience building and optimizing data pipelines for machine learning models. What distributed systems or cloud technologies (particularly AWS) have you used to handle large-scale datasets?
- Walk us through a project where you translated academic research or frontier lab techniques into production ML systems. What were the key challenges?