brozzler - distributed browser-based web crawler
Role in this project:
Full-stack Developer Contributions:61 reviews, 717 commits, 200 PRs in 6 years 6 months
Contributions summary:Barbara made various contributions to the `brozzler` repository, including browser-related configurations, behavior implementations for specific websites, and updates to the core crawling logic. The user added features for handling user logins, implemented JavaScript behavior for Instagram, and adjusted the browser's network settings. Furthermore, the user addressed issues related to video downloads and made changes to the project's documentation.
web-crawler
Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.
Role in this project:
Back-end Developer Contributions:1 review, 71 commits, 39 PRs in 5 years 11 months
Contributions summary:Barbara primarily contributed to the Heritrix3 web crawler project by implementing features related to AMQP integration, specifically around queuing and processing URLs received via AMQP. They developed components like `AMQPUrlWaiter` and `AMQPUrlReceiver` to manage the flow of URLs. The user also worked on `ExtractorYoutubeDL`, configuring the tool to extract video metadata, and applying filters. These changes show an understanding of the crawler's architecture and its interaction with external systems.
heritrixweb-crawlerjavawebcrawlingwarc