Sujan Dutta is a PhD candidate and research-focused machine learning engineer with eight years of experience blending NLP, computational social science, and applied AI across academia and industry. Based in San Jose, he has contributed to Siri understanding and privacy-preserving ML at Apple and explored agentic AI during an applied research internship at Adobe, while driving NLP research as a Graduate Research Assistant at RIT. He teaches machine learning and data science to a global audience through Normalized Nerd and open-source projects that implement core algorithms and NLP pipelines from scratch. Comfortable with LLMs, word embeddings, and model interpretability, Sujan bridges rigorous research with practical engineering—often reimplementing foundational algorithms to deepen understanding and explainability.
8 years of coding experience
4 years of employment as a software developer
Doctor of Philosophy - PhD Computing and Information Sciences, Doctor of Philosophy - PhD Computing and Information Sciences at Rochester Institute of Technology
Bachelor of Technology Computer Science, Bachelor of Technology Computer Science at Kalyani Government Engineering College
Implementation of basic ML algorithms from scratch in python...
Role in this project:
Data Scientist
Contributions:36 commits, 35 pushes, 1 branch in 1 year 11 months
Contributions summary:Sujan primarily contributed to implementing and updating machine learning models within the repository. Their work involved generating datasets and implementing various algorithms from scratch, including linear regression, logistic regression, and decision trees. They also added code related to K-means clustering and Naive Bayes, suggesting a focus on a range of fundamental machine learning techniques. The user also provided implementations of stochastic gradient descent.
Contributions:72 commits, 72 pushes, 1 branch in 1 year 7 months
Contributions summary:Sujan primarily worked on data analysis and machine learning tasks, demonstrated by the implementation of evaluation metrics using scikit-learn and the creation of an NLP model to classify text data using techniques like Bag of Words and TF-IDF. They also experimented with word embeddings using Word2Vec and GloVe models. The user's contributions also extend to model interpretability, as seen by analyzing the performance metrics to understand the effectiveness of these models.
pythonyoutube-channelvideosjavascriptyoutube
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.