Glean
Software Engineer, Data Foundations
About this role
Join Glean's Data Foundations team to build the data ingestion and management layer powering enterprise search, AI assistants, and AI agents. You'll develop connectors across hundreds of SaaS platforms, transform unstructured content into structured knowledge, and ensure data quality at scale across billions of documents.
What you'll do
- Build and scale connectors to SaaS and on-premises systems like Google Workspace, Microsoft 365, Slack, Salesforce, and ServiceNow
- Implement full syncs, incremental updates via webhooks/APIs, rate-limiting, and complex authentication flows
- Transform raw enterprise content into structured, permission-aware representations optimized for search and LLM reasoning
- Design document schemas and enrichment pipelines including entity extraction and access control propagation
- Develop advanced datasource capabilities such as actions, live-fetch, and query language support
- Enable AI product capabilities through deep integrations that automate tasks and enhance indexed data with live information
What they're looking for
- API integration and connector development
- Data ingestion and ETL pipeline design
- Unstructured data processing and transformation
- Authentication and security protocols
- Distributed systems and scalability
- Schema design and data modeling
- Backend software engineering
- Working with enterprise SaaS platforms
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Glean
Glean builds a Work AI platform that helps enterprises access and leverage their internal data intelligently. The company is hiring backend engineers, infrastructure specialists, fullstack engineers, and billing platform leads to develop scalable features, robust data infrastructure, consumption-based billing systems, and enterprise-grade storage and analytics capabilities.
- Website
- glean.com
Likely interview questions
- Walk us through your experience building or maintaining data connectors to external APIs. How have you handled rate limiting, authentication complexity, and incremental sync strategies?
- Describe a time you transformed unstructured data into a structured format optimized for search or ML consumption. What challenges did you encounter?