Babel Street
Software Engineer III
About this role
Babel Street seeks a Software Engineer III to architect and maintain advanced automated data extraction systems that harvest business intelligence from complex web environments while circumventing anti-bot security measures. This role ensures reliable, high-quality data flow into internal databases and is based in Reston, VA or Somerville, MA with remote work potential.
What you'll do
- Design and develop automated web scraping systems (spiders) for extracting critical business intelligence
- Overcome technical barriers and sophisticated anti-bot security measures
- Maintain continuous, reliable flow of high-quality data into internal databases
- Ensure end-to-end responsibility for data extraction solutions
- Architect scalable solutions for complex web environments
- Monitor and troubleshoot extraction system performance
What they're looking for
- Web scraping and data extraction
- Python or similar programming languages
- Handling anti-bot security systems
- Database integration
- System architecture and design
- Automated testing
- Problem-solving in complex technical environments
- Data quality assurance
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Babel Street
Babel Street develops intelligence platforms that extract, match, and analyze complex data from diverse sources using advanced AI techniques including natural language processing, computer vision, and web data harvesting. The company is hiring software engineers across multiple specializations—from data infrastructure and NLP to computer vision—to build and maintain the core systems powering their identity intelligence and data analytics capabilities.
- Website
- babelstreet.com
Likely interview questions
- Walk us through your experience designing and maintaining web scraping or data extraction systems. How have you handled anti-bot measures and sophisticated technical barriers?
- Describe a complex data pipeline you've built. How did you ensure data quality and reliability in a production environment?