Michael Collado is a Principal Engineer with nearly two decades of experience building durable, large-scale data and platform systems and leading teams from architecture through GA launches. He currently leads Snowflake’s Open Catalog product and the Horizon–Polaris integration, and helped found Apache Polaris where he remains a PMC member shaping both technical roadmap and governance. His background includes designing ingestion and analytics platforms at Amazon that processed hundreds of billions of events per day and streaming petabytes of vehicle data at Cruise, demonstrating repeated delivery at extreme scale. Michael is an active open-source steward beyond Polaris—serving on the OpenLineage TSC and contributing meaningful Spark and lineage integrations to OpenLineage and Marquez. He blends hands-on engineering with cross-team coordination and mentorship, routinely turning ambitious ideas into production-ready, community-aligned platforms. Based in Seattle, he pairs a rare mix of product-minded systems design and long-term operational ownership informed by deep open-source collaboration.
10 years of coding experience
19 years of employment as a software developer
Bachelor of Arts - BA, English Language and Literature/Letters, Bachelor of Arts - BA, English Language and Literature/Letters at Missouri State University
Contributions:4 releases, 334 reviews, 176 commits in 1 year 8 months
Contributions summary:Michael primarily focused on improving the OpenLineage project's Spark integration. Their contributions involved adding logging statements, updating and creating new features, and refactoring existing code. They implemented features such as metric facets, and support for logical RDDs. Additionally, they made improvements to the build process, addressed sporadic integration test failures, and corrected issues related to JDBC connections and data parsing.
Collect, aggregate, and visualize a data ecosystem's metadata
Role in this project:
Back-end Developer & Data Engineer
Contributions:3 releases, 199 reviews, 103 commits in 1 year 10 months
Contributions summary:Michael primarily focused on enhancing the data lineage capabilities of the Marquez project, specifically related to Apache Spark integration. They implemented logging statements to capture and track lineage events, including failures and plan executions. The user also contributed to integrating dataset metrics, like row counts, and updating code to provide a more comprehensive data lineage picture within the Spark environment.
datavisualizedata-catalogdata-dictionarymetadata
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.