NewsBreak
Machine Learning Engineer, LLM Post-Training
About this role
NewsBreak is seeking a Machine Learning Engineer to lead the post-training of large language models, primarily focusing on reinforcement learning techniques. This role involves collaborating with product teams to translate requirements into effective training strategies while managing data preparation and evaluation processes.
What you'll do
- Oversee post-training processes involving continuous pre-training, supervised fine-tuning, and reinforcement learning.
- Design and manage data preparation tailored to training needs.
- Collaborate with business and product teams to create training plans.
- Conduct large-scale training on GPU clusters with distributed training techniques.
- Build evaluation and reward pipelines to ensure model quality.
- Stay updated with post-training research and implement new techniques.
What they're looking for
- Hands-on experience with LLM post-training
- Strong data engineering for machine learning
- Large-scale GPU training expertise
- Proficient in PyTorch
- Understanding of tokenization and attention mechanisms
- Fast iteration and business impact orientation
- Communication skills for cross-team collaboration
Benefits
- 100% health, dental, and vision coverage for employees
- Top-tier 401(K) plan with company matching
- Paid time off and holidays
- Flexible spending and health savings accounts
- Team activity budget
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
NewsBreak
NewsBreak operates a news and local marketplace platform serving 40M+ monthly active users, leveraging machine learning and data infrastructure to power recommendations, advertising, and matching between consumers and businesses. The company is hiring software engineers, ML engineers, and data specialists to build and optimize backend systems, ML infrastructure, recommendation engines, and algorithms that drive user engagement and monetization.
- Website
- newsbreak.com
Likely interview questions
- Walk us through a time you personally implemented and debugged an RL algorithm like PPO or DPO for LLM post-training. What were the key challenges and how did you overcome them?
- Describe your experience designing and curating datasets for different post-training stages (SFT, preference pairs, reward signals). How do you decide what data to collect for a specific product use case?