Baseten
Data Engineer
About this role
Baseten is seeking a Data Engineer to build and scale their internal data platform, transforming product and business data into reliable datasets for cross-functional decision-making. You'll design data models, pipelines, and analytics infrastructure while working with AI inference and observability data to support Product, Engineering, Finance, Marketing, and Sales teams.
What you'll do
- Design and maintain core data models and semantic layers
- Develop and orchestrate batch and streaming data pipelines using Apache Beam, Kafka, Airflow, or similar
- Analyze inference and infrastructure telemetry from observability tools
- Define and maintain company-wide metrics for product usage, performance, and customer lifecycle
- Enable self-service analytics through agents and tools with structured semantic layers
- Ensure data reliability and quality through testing, documentation, and governance
What they're looking for
- Data pipeline orchestration (Airflow, Beam, Kafka)
- Data modeling and semantic layer design
- Batch and streaming data processing
- Observability tools (OpenTelemetry, Grafana)
- SQL and data warehousing
- Inference and ML infrastructure metrics
- B2B SaaS analytics
- Forecasting and predictive modeling
Benefits
- Competitive compensation with meaningful equity
- 100% medical, dental, and vision insurance coverage for employee and dependents
- Flexible PTO with company-wide Winter Break
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Baseten
Baseten builds an AI inference platform that enables companies to deploy and manage machine learning models in production at scale. The company is hiring software engineers, infrastructure specialists, and product engineers to develop observability systems, frontend experiences, reliability infrastructure, developer tools, and enterprise deployment solutions.
- Website
- baseten.com
Likely interview questions
- Walk us through how you'd design a data model to track inference metrics like latency, throughput, and token usage across multiple AI models in production.
- Describe your experience building batch and streaming pipelines. Which orchestration tool have you used most, and what challenges did you face at scale?