Skip to main content

EXA

Software Engineer, Distributed Data Systems

San Francisco, California$180k–$350kfulltimemidAdded 1 month ago

About this role

Exa, a venture-backed AI search company, is seeking a Software Engineer to design and build large-scale distributed data infrastructure supporting web crawling, embedding model training, and vector database retrieval. You'll architect data systems handling hundreds of petabytes while maintaining high reliability and performance.

What you'll do

  • Design lakehouse architectures to handle massive web crawl datasets (100+ PB scale)
  • Build and operate distributed data processing pipelines for billions of daily documents
  • Architect data layers for embedding training infrastructure using Ray and similar frameworks
  • Develop streaming pipelines for real-time indexing and search
  • Scale analytical query systems (ClickHouse) across petabyte-scale search logs
  • Ensure system reliability and operational excellence

What they're looking for

  • Lakehouse architecture (Delta Lake, Iceberg, Hudi)
  • Distributed data processing (Spark, Ray, ClickHouse)
  • Streaming systems (Kafka, Flink)
  • Large-scale infrastructure design and operations
  • Vector database and storage formats (bonus: Lance)
  • GPU-accelerated data processing (bonus: RAPIDS, cuDF)
  • Systems reliability and observability
  • Web-scale data engineering

Benefits

  • Premium healthcare (medical, dental, vision)
  • Fertility benefits
  • 16 weeks fully paid parental leave
  • Monthly wellness stipend
  • Visa sponsorship available (STEM OPT, H1B, O1, E3)
  • In-person role in San Francisco
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

EXA

Exa builds AI-powered search infrastructure that powers major products and serves hundreds of thousands of developers through APIs, SDKs, and web crawling technology. The company is hiring full-stack engineers, research engineers, distributed systems architects, customer-facing deployment engineers, and growth engineers to expand its platform and customer base.

View all jobs at EXA

Likely interview questions

  • Walk us through a large-scale distributed data pipeline you've built in production. What were the biggest challenges with reliability and how did you solve them?
  • You need to design a lakehouse architecture to handle 100+ petabytes of web crawl data. Which format would you choose—Delta Lake, Iceberg, or Hudi—and why?