Weichen Xu

Software Engineer at Databricks

Singapore
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Weichen Xu is a software engineer with a decade of experience building distributed ML infrastructure and GenAI systems, currently driving work at Databricks and serving as an Apache Spark committer. He specializes in scalable ML platforms, Ray and Spark integrations, and reliability improvements across projects like MLflow, Horovod, XGBoost, and TensorFlow/PyTorch tooling. A prolific open-source contributor, he has enhanced core Spark ML features, Ray-on-Spark autoscaling, and artifact robustness in MLflow while modernizing deep-learning pipelines for production. Comfortable shipping back-end fixes, performance benchmarks, and cross-language APIs, he combines production-grade engineering with research-caliber ML understanding. Notably, his contributions span both low-level runtime stability (e.g., daemon scripts, AccumulatorV2 migrations) and higher-level model lifecycle improvements, making him effective at bridging infrastructure and ML workflows.
code10 years of coding experience
bookMaster of Science - MS Computer Science, Master of Science - MS Computer Science at Beihang University
github-logo-circle

Github Skills (74)

gbm10
spark-sql10
apache-spark10
xgboost10
parallel10
python10
testing10
pyspark10
performance-testing10
hyperparameter-optimization10
javas10
deep-learning10
ray10
data-pipeline10
distributed-training10

Programming languages (10)

TypeScriptJavaShellC++BatchfileScalaJavaScriptHTML

Github contributions (5)

github-logo-circle
Deep Learning Pipelines for Apache Spark
Role in this project:
userML Engineer
Contributions:9 reviews, 14 commits, 16 PRs in 1 year 9 months
Contributions summary:Weichen primarily contributed to the deprecation and removal of legacy APIs and functionalities within the `sparkdl` library, which focuses on deep learning pipelines for Apache Spark. This included deprecating Keras model-related utilities and transformers, and migrating to more modern approaches such as Pandas UDFs. Furthermore, the user updated the codebase to align with a new version (2.0.0) and updated documentation related to the Xgboost API. This work streamlined the library and modernized its core functionalities.
apache-sparkdeep-learning
databricks/spark-sql-perf

Sep 2017 - May 2018

Role in this project:
userML Engineer
Contributions:7 commits, 8 PRs, 36 comments in 8 months
Contributions summary:Weichen's contributions primarily involve adding and benchmarking machine learning algorithms within the Spark SQL performance testing framework. They added performance tests for various ML algorithms, including LinearSVC, OneHotEncoder, VectorAssembler, StringIndexer, Tokenizer, FPGrowth, MinHashLSH, BucketedRandomProjectionLSH, and Word2Vec. Furthermore, the user enhanced the framework by introducing additional methods tests and refining existing code, such as fixing `df.drop` in VectorAssembler and adding the `associationRules` method in `FPGrowth`.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial