Susan Li is a Chief Data Scientist with nine years of hands-on experience translating customer- and product-focused problems into scalable data science solutions across startups and research-driven firms. She specializes in customer analytics, predictive modeling, NLP and time-series forecasting, and has led projects from CLV and churn modeling to price-anomaly detection and recommender systems. A passionate data literacy advocate and prolific writer, she explains complex ML and NLP concepts for both technical and general audiences and maintains practical NLP notebooks on GitHub showcasing text preprocessing with Pandas, NLTK, SpaCy and Gensim. Her background spans analytics, product integration and business development—bringing a rare blend of commercial acumen and deep technical craft to help organizations operationalize data.
9 years of coding experience
15 years of employment as a software developer
Statistics with R Specialization Statistics, Statistics with R Specialization Statistics at Coursera
Nanodegree Data Science, Nanodegree Data Science at Udacity
University of California, Irvine
Digital Marketing Management Certificate, Digital Marketing Management Certificate at University of Toronto
Bachelor of Arts (B.A.) English Language and Literature General, Bachelor of Arts (B.A.) English Language and Literature General at Dalian University of Foreign Languages
Statistics Computer Science Text Mining, Statistics Computer Science Text Mining at Massive Online Courses
MBA, MBA at Maastricht University
Certificate Natural Language Processing with Deep Learning, Certificate Natural Language Processing with Deep Learning at Stanford University School of Engineering
University of Georgia
Digital Analytics - UBC/DAA Award of Achievement Program, Digital Analytics - UBC/DAA Award of Achievement Program at The University of British Columbia
Micro Master Certificate Artificial Intelligence, Micro Master Certificate Artificial Intelligence at Columbia University
Scikit-Learn, NLTK, Spacy, Gensim, Textblob and more
Role in this project:
Back-end Developer
Contributions:1 release, 114 commits, 96 pushes in 4 years 4 months
Contributions summary:Susan added several notebooks demonstrating the use of the Pandas library for text data processing. The code examples within the notebooks involved tasks like tokenizing text, removing punctuation, cleaning stopwords, stemming and lemmatizing and using regular expressions. The changes suggest work related to processing and manipulating textual data for a project focused on natural language processing.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.