Code Metal
Applied AI Research Engineer
About this role
Code Metal is seeking an Applied AI Research Engineer to develop advanced AI systems that support military planning through human-machine collaboration. The role involves designing, building, and evaluating AI systems and workflows that enhance decision-making capabilities in critical operational contexts.
What you'll do
- Design and build agentic AI systems for human-machine teaming
- Develop AI pipelines integrating models and external tools
- Conduct experiments to evaluate AI performance and trust
- Fine-tune and evaluate foundation models for planning tasks
- Build datasets and automated benchmarks for ongoing research
- Collaborate with engineers and domain experts to implement AI solutions
What they're looking for
- Strong Python programming
- Experience with AI and machine learning systems
- Proficiency in PyTorch
- Understanding of multi-agent systems
- Knowledge of simulation and evaluation techniques
- Ability to work collaboratively in small teams
- Expertise in data processing and retrieval pipelines
- Experience in developing explainable AI systems
Benefits
- Opportunity to work on impactful mission-driven projects
- Engagement in hands-on AI development beyond demos
- Involvement in innovative research on human-machine collaboration
- Rapid prototyping and development in a small team setting
- Visibility and feedback from real users in operational roles
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Code Metal
Code Metal builds AI systems and algorithm-to-hardware platforms for mission-critical defense, aerospace, and automotive applications, with a focus on formal methods verification and human-machine collaboration. The company is hiring applied AI researchers, formal methods specialists, forward deployed engineers, and IT support staff to develop and deploy these advanced systems.
View all jobs at Code MetalLikely interview questions
- Walk us through a time you built an agentic AI system or multi-agent workflow. What was the planning or decision-support task, and how did you integrate language models with external tools or simulation?
- Describe your experience fine-tuning or post-training foundation models. What evaluation metrics did you use, and how did you measure reliability and trustworthiness?