Basis Research Institute
Data Engineer, Platform
About this role
Basis, a nonprofit AI research organization, seeks a Data Engineer to build reliable data pipelines with strong provenance tracking and quality controls. This platform role involves curating datasets for ML training, managing cross-project data infrastructure, and enabling both commercial products and internal research with trustworthy data systems.
What you'll do
- Design and build data pipelines with comprehensive provenance, lineage tracking, and quality gates
- Curate documented datasets for model training, evaluation, and experimentation
- Develop data quality frameworks and governance systems for internal and external use
- Support scaling of data infrastructure to handle medium-scale models and multiple teams
- Coordinate shared datasets across Platform and Research teams to prevent duplication
- Ensure data systems enable reproducible ML experiments and research outcomes
What they're looking for
- Expert SQL and Python for data processing
- Distributed computing frameworks (Spark, Dask)
- Workflow orchestration tools (Airflow, Dagster, Prefect)
- Cloud data platforms (Snowflake, BigQuery, Redshift, S3)
- ML data requirements and feature engineering
- Data quality, validation, and governance implementation
- Data modeling and schema design optimization
- Batch and streaming systems (Kafka, Kinesis, Flink)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Basis Research Institute
Basis Research Institute is a nonprofit AI research organization building trustworthy data infrastructure, scalable ML training systems, custom robotic platforms, and LLM-powered automation tools to support both research and commercial applications. The organization is hiring Data Engineers, ML Systems Engineers, Hardware Systems Engineers, Software Engineers, and Machine Learning Research Engineers to develop its technical infrastructure and translate research into production systems.
View all jobs at Basis Research InstituteLikely interview questions
- Walk us through a time you built a data pipeline for ML model training or evaluation. How did you approach data quality and what validation frameworks did you implement?
- Describe your experience with data provenance and lineage tracking. How have you ensured data transparency and reproducibility in past projects?