Cohere
Data Engineer, Data Foundations
About this role
Cohere seeks a Data Engineer to build foundational data infrastructure supporting the company's enterprise AI platform. You'll design and maintain production-grade data systems that power analytics, customer experiences, and strategic decision-making across the organization.
What you'll do
- Design and implement production-grade data processing systems handling unstructured and structured data at scale
- Build end-to-end data pipelines and ETL workflows using distributed processing frameworks
- Collaborate with researchers, engineers, and business teams to translate data insights into product and strategy recommendations
- Develop performant datasets in relational databases and cloud storage systems
- Maintain and optimize modern analytics infrastructure and tooling
- Own implementations from conception through delivery and measurable outcomes
What they're looking for
- Python and SQL
- Apache Beam, Spark, or Flink
- BigQuery, Airflow, or dbt
- Large-scale data system design
- Data transformation and pipeline development
- Java or Golang (preferred)
- Kubernetes (preferred)
- AI/ML domain knowledge
Benefits
- $75 weekly lunch stipend
- Comprehensive health, dental, and mental health coverage
- 6 weeks paid vacation (30 working days)
- 100% parental leave top-up for up to 6 months
- Education and learning stipend for professional development
- Home office stipend ($500) and remote-friendly work environment
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Cohere
Cohere builds enterprise AI systems and large language model infrastructure, including audio inference optimization, agentic AI workflows, and petabyte-scale data infrastructure for model training. The company is hiring software engineers, infrastructure specialists, forward-deployed engineers to work with enterprise customers, and IT support staff.
- Website
- cohere.com
Likely interview questions
- Walk us through a production data processing system you've built. What framework did you use (Spark, Beam, Flink), and how did you handle scale?
- Describe your experience transforming unstructured data into performant relational or blob storage datasets. What were the key challenges?