Deep Learning Pipelines for Apache Spark
Role in this project:
ML Engineer Contributions:9 reviews, 14 commits, 16 PRs in 1 year 9 months
Contributions summary:Weichen primarily contributed to the deprecation and removal of legacy APIs and functionalities within the `sparkdl` library, which focuses on deep learning pipelines for Apache Spark. This included deprecating Keras model-related utilities and transformers, and migrating to more modern approaches such as Pandas UDFs. Furthermore, the user updated the codebase to align with a new version (2.0.0) and updated documentation related to the Xgboost API. This work streamlined the library and modernized its core functionalities.
apache-sparkdeep-learning
Role in this project:
ML Engineer Contributions:7 commits, 8 PRs, 36 comments in 8 months
Contributions summary:Weichen's contributions primarily involve adding and benchmarking machine learning algorithms within the Spark SQL performance testing framework. They added performance tests for various ML algorithms, including LinearSVC, OneHotEncoder, VectorAssembler, StringIndexer, Tokenizer, FPGrowth, MinHashLSH, BucketedRandomProjectionLSH, and Word2Vec. Furthermore, the user enhanced the framework by introducing additional methods tests and refining existing code, such as fixing `df.drop` in VectorAssembler and adding the `associationRules` method in `FPGrowth`.