Meridial
ML Engineer Specialist - Freelance AI Trainer Project
About this role
Freelance ML specialist role focusing on hands-on machine learning challenges to improve large language models. You'll solve applied ML problems, evaluate research writing, debug model behavior, design benchmarks, and refine training datasets to strengthen next-generation AI systems.
What you'll do
- Solve applied ML tasks iterating on model reasoning and optimization workflows
- Evaluate scientific writing using review rubrics for clarity, correctness, and impact
- Debug advanced LLM behavior on optimization, fairness, transformers, and NLP
- Design and build new benchmark tasks with datasets and evaluation scripts
- Improve training datasets through structured feedback on model outputs
- Document failure modes and provide reference implementations
What they're looking for
- Machine learning and deep learning
- Natural Language Processing (NLP)
- Python programming
- ML frameworks and cloud workflows
- Supervised/unsupervised learning
- Reinforcement learning
- Technical writing and communication
- Research methodology
Benefits
- Remote, flexible contract work with schedule autonomy
- $30-$50 per hour based on experience and location
- Work on cutting-edge AI training initiatives
- Direct impact on next-generation language models
- Research-style problem-solving at scale
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Meridial
Meridial trains and improves advanced AI models through expert evaluation and feedback across infrastructure, software engineering, machine learning, and language domains. The company hires experienced freelance specialists—including software engineers, ML experts, and language specialists—to test AI reasoning, identify failure modes, and provide detailed training data to enhance model capabilities.
View all jobs at MeridialLikely interview questions
- Walk us through a recent ML project where you iterated on a model to improve performance. What metrics did you track, and how did you decide when to stop iterating?
- Describe your experience debugging large language models or neural networks. Can you give an example of a failure mode you uncovered and how you documented it?