Colin Rosenthal

IT Consultant at The Royal Danish Library (formerly Statsbiblioteket)

Central Denmark Region, Denmark
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

👤
Senior
🎓
Top School
Colin Rosenthal is an IT Consultant and software engineer with 17 years of industry experience and a strong background in research-grade science, holding a PhD in Astrophysics and an MSc in Software Engineering. He blends technical leadership and hands-on Java backend development with practical expertise in version control, CI/CD, integration testing and containerisation, and has contributed resilience fixes to the well-known Heritrix web crawler. Comfortable in both Danish and English workplaces, he has an international track record including work in the USA and Norway and long-term service at the Royal Danish Library. Equally at ease translating complex scientific models into robust software, he brings disciplined project management instincts from his academic grant and research roles to production engineering.
code17 years of coding experience
job5 years of employment as a software developer
bookcand.it. (MSc), Software Engineering, cand.it. (MSc), Software Engineering at Aarhus University
bookPhD, Astrophysics, PhD, Astrophysics at Cambridge University
github-logo-circle

Github Skills (11)

crawler10
javas10
web-crawler10
webscraper10
webscraping10
heritrix10
java10
crawling10
regex9
junit8
jtest8

Programming languages (6)

JavaJinjaShellJavaScriptHTMLPython

Github contributions (5)

github-logo-circle
internetarchive/heritrix3

Oct 2018 - May 2022

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.
Role in this project:
userBack-end Developer
Contributions:10 commits, 5 PRs, 11 comments in 3 years 8 months
Contributions summary:Colin primarily focused on improving the Heritrix3 web crawler project. They implemented a timeout feature for regular expression matching within the crawl logic, enhancing the crawler's resilience. They also fixed a potential null pointer exception in the `hashCode` method, improving the functionality of the browse beans feature. Furthermore, the user added code to filter out embedded images.
web-crawlerinternet-archivescalearchivecommoncrawl
netarchivesuite/crawlrss

Oct 2016 - Oct 2024

Contributions:14 pushes in 8 years 1 month
rsscrawlfeedadd-onheritrix
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial