Arlo
Data Engineer
About this role
Arlo, an AI-driven health insurance company, seeks a Data Engineer to build and maintain data pipelines that power underwriting and pricing models. You'll own data ingestion, transformation, quality monitoring, and collaboration with data science teams to ensure clean, reliable data flows across the organization.
What you'll do
- Build and maintain ingestion pipelines for diverse healthcare data sources including TPA feeds, claims, eligibility, and enrollment records
- Design dbt models and transformation logic that create authoritative data tables for underwriting, pricing, and reporting
- Implement pipeline orchestration using tools like Dagster or Airflow with monitoring, retries, and alerting
- Build data quality monitoring systems to catch duplicates, ID mismatches, enrollment gaps, and reporting delays
- Partner with data science teams to develop and maintain feature pipelines for ML models
- Document data quality limitations and work with engineering to prioritize fixes for upstream issues
What they're looking for
- Python and SQL (production-quality code)
- Pipeline orchestration tools (Dagster, Airflow, Prefect, or similar)
- dbt or equivalent data transformation frameworks
- Cloud data environments (AWS, GCP, or Azure)
- Columnar and analytical databases
- Data quality and observability practices
- Healthcare data or claims experience (nice to have)
- ML feature pipeline support (nice to have)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Arlo
Arlo is an AI-driven health insurance company using artificial intelligence to reduce healthcare costs and improve operations. The company is hiring AI engineers, data engineers, and product engineers to build AI agents, data pipelines, and member-facing features that transform underwriting, claims, member support, and pricing across their platform.
View all jobs at ArloLikely interview questions
- Walk us through a complex data pipeline you built from scratch. What data sources did you ingest, what transformations did you apply, and how did you ensure it stayed reliable in production?
- Tell me about a time you discovered a data quality issue that was silently propagating downstream. How did you identify it, fix it, and prevent similar issues in the future?