Data ingestion library for Amundsen to build graph and search index
Role in this project:
Back-end Developer Contributions:40 releases, 2 reviews, 76 commits in 1 year 4 months
Contributions summary:Jin primarily focused on enhancing the data ingestion pipeline and its related components. They implemented a shutdown hook for a closer object and improved the file system Neo4j CSV loader by adding features to manage temporary directories, which included options to force creation or deletion. The user also worked on deleting stale nodes and relations within Neo4j, adding timestamp-based expiration to improve data management.
data-ingestiongraphamundsen
Role in this project:
Back-end & Data Engineer Contributions:12 commits, 11 PRs, 35 comments in 5 months
Contributions summary:Jin primarily focused on enhancing the Azkaban plugins, particularly those related to Gobblin integration. Their contributions included implementing support for the HDFS to MySQL data flow and refactoring code related to Gobblin presets. The user also added functionalities to support Hive (ORC) formats and addressed code refactoring related to Teradata. Furthermore, the user addressed the removal of password from printed job properties in Gobblin.
azkaban