Field AI
Agentic AI/ML Engineer, Multimodal
About this role
FieldAI seeks an AI/ML engineer to develop multimodal foundation models powering a global fleet of autonomous robots. You'll work across computer vision, vision-language models, and agentic AI systems, taking models from research through production deployment in real-world field conditions.
What you'll do
- Design and develop multimodal foundation models for scene understanding and video analysis
- Curate datasets and fine-tune models for real-world robot deployment
- Implement agentic AI features including tool use, memory systems, and retrieval-augmented generation
- Optimize model inference for production environments
- Evaluate model performance and iterate based on field deployment feedback
- Contribute to broader perception and insight initiatives across the company
What they're looking for
- Computer vision and image processing
- Vision-language models (VLMs)
- Multimodal machine learning
- Agentic AI and tool-use systems
- Model fine-tuning and evaluation
- ML inference optimization
- Video understanding and long-context analysis
- Full-cycle ML development and deployment
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Field AI
Field AI develops embodied AI and autonomous robotics systems for real-world deployment in industrial environments like oil & gas and mining. The company is hiring software engineers to build web-based systems, perception and validation pipelines, test infrastructure, ROS-based robotic software, and customer-facing products that integrate AI with field-deployed hardware.
View all jobs at Field AILikely interview questions
- Walk us through your experience with multimodal models—specifically vision-language models (VLMs). How have you approached fine-tuning or evaluating them on custom datasets?
- Describe a time you optimized ML model inference for production deployment. What trade-offs did you consider between accuracy, latency, and resource constraints?