Skip to main content

Baseten

Data Engineer

San Francisco (Remote)$180k–$250kfulltimemidAdded 1 month ago

About this role

Baseten is seeking a Data Engineer to build and scale their internal data platform, transforming product and business data into reliable datasets for cross-functional decision-making. You'll design data models, pipelines, and analytics infrastructure while working with AI inference and observability data to support Product, Engineering, Finance, Marketing, and Sales teams.

What you'll do

  • Design and maintain core data models and semantic layers
  • Develop and orchestrate batch and streaming data pipelines using Apache Beam, Kafka, Airflow, or similar
  • Analyze inference and infrastructure telemetry from observability tools
  • Define and maintain company-wide metrics for product usage, performance, and customer lifecycle
  • Enable self-service analytics through agents and tools with structured semantic layers
  • Ensure data reliability and quality through testing, documentation, and governance

What they're looking for

  • Data pipeline orchestration (Airflow, Beam, Kafka)
  • Data modeling and semantic layer design
  • Batch and streaming data processing
  • Observability tools (OpenTelemetry, Grafana)
  • SQL and data warehousing
  • Inference and ML infrastructure metrics
  • B2B SaaS analytics
  • Forecasting and predictive modeling

Benefits

  • Competitive compensation with meaningful equity
  • 100% medical, dental, and vision insurance coverage for employee and dependents
  • Flexible PTO with company-wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Baseten

Baseten builds an AI inference platform that enables companies to deploy and manage machine learning models in production at scale. The company is hiring software engineers, infrastructure specialists, and product engineers to develop observability systems, frontend experiences, reliability infrastructure, developer tools, and enterprise deployment solutions.

View all jobs at Baseten

Likely interview questions

  • Walk us through how you'd design a data model to track inference metrics like latency, throughput, and token usage across multiple AI models in production.
  • Describe your experience building batch and streaming pipelines. Which orchestration tool have you used most, and what challenges did you face at scale?