Brooklyn Sports & Entertainment
LLMOps Engineer
About this role
Brooklyn Sports & Entertainment seeks an LLMOps Engineer to build and maintain production-ready AI infrastructure that ensures safety and reliability as the organization scales its AI capabilities. You'll manage deployment pipelines, monitoring systems, and operational processes that support the company's AI platform across its sports and entertainment properties.
What you'll do
- Design and maintain LLM deployment infrastructure for production environments
- Develop monitoring and observability systems for AI model performance and safety
- Build CI/CD pipelines and automation for model versioning and rollouts
- Establish operational procedures and runbooks for AI system management
- Collaborate with data science and engineering teams on AI platform requirements
- Ensure security, compliance, and reliability of AI systems in production
What they're looking for
- LLM deployment and orchestration platforms
- DevOps and infrastructure-as-code practices
- Python or similar programming languages
- Cloud platforms (AWS, GCP, or Azure)
- Monitoring and logging tools
- CI/CD pipeline development
- Container technologies (Docker, Kubernetes)
- Model versioning and experiment tracking
Opens the official application on the employer’s site. No login required.
Brooklyn Sports & Entertainment
Brooklyn Sports & Entertainment builds AI-powered platforms and software solutions that enhance decision-making and operations across its sports and entertainment properties. The company is hiring for AI infrastructure, full-stack development, analytics, and engineering roles focused on deploying production-ready AI systems, agentic applications, and operational tools used by coaches, scouts, and business teams.
View all jobs at Brooklyn Sports & EntertainmentLikely interview questions
- Can you walk us through your experience building or maintaining LLM infrastructure in production? What were the key safety and reliability challenges you faced?
- How have you approached monitoring and evaluating LLM model performance over time, and what metrics did you track to ensure safety during updates or redeployments?