Web Scraping Specialist
Who We Are:
We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.
We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.
We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.
The Role.
We are seeking a Web Scraping Specialist who is proficient and brings significant experience in data extraction and web scraping techniques. You will join a small, specialized team and lead efforts to gather and analyze data, optimize scraping processes, and support our vision for a future where Grass plays a crucial role in transforming internet data accessibility.
Who You Are.
- Demonstrated ability to extract data from complex websites with minimal supervision, with a portfolio or examples of past projects.
- Proficiency in languages such as Python or JavaScript, with strong skills in libraries and frameworks like BeautifulSoup, Scrapy, or Selenium.
- Knowledge of asynchronous programming, multithreading, and distributed scraping.
- In-depth knowledge of HTML, CSS, JavaScript, and the Document Object Model (DOM).
- Experience with NoSQL databases (MongoDB, Cassandra), capable of designing efficient storage solutions and managing data integrity.
- Ability to apply machine learning algorithms for data cleaning, categorization, or predictive analysis adds significant value.
- Experience with cloud services (AWS, Google Cloud, Azure) for deploying and managing scraping jobs at scale.
- Active participation in open-source projects related to web scraping, data pro...