Sean Owen

Austin, Texas, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
award
Top expert inAndroid and Kotlin Development
Sean Owen is a seasoned data science and open-source leader with 18 years of experience building large-scale ML and data platforms, currently driving applied research on improving open LLMs at Mosaic/Databricks. He blends hands-on engineering — as an Apache Spark committer and primary author of projects like zxing and Oryx — with product and field-facing leadership from roles at Databricks, Cloudera and as a founder. His work spans real-time, lambda-style architectures, scalable model training, and pragmatic data curation for production ML. Sean has a track record of translating research into customer impact, leading global DS teams and advising product engineering on ML architecture. He also spent time as an early-stage software investor and holds an MBA from London Business School and a BA in Computer Science from Harvard. A practical polyglot contributor, he’s comfortable optimizing both cluster-scale pipelines and the nitty-gritty of model training on specific hardware.
code18 years of coding experience
stackoverflow-logo

Stackoverflow

Stats
66,451reputation
8.9mreached
1,444answers
13questions
Badges
equals
top-5%
encoding
top-5%
hadoop
top-1%
java
top-1%
tomcat
top-5%
regex
top-5%
github-logo-circle

Github Skills (108)

transformers10
spark-sql10
apache-spark10
python10
mapreduce10
hadoop10
security10
javas10
eclipse10
refactoring10
clustering10
scikit10
dataframes10
xml-parsing10
hadoop-mapreduce10

Programming languages (11)

JavaShellC++RustScalaJavaScriptHTMLJupyter Notebook

Github contributions (5)

github-logo-circle
databrickslabs/dolly

Mar 2023 - Sep 2023

Databricks’ Dolly, a large language model trained on the Databricks Machine Learning Platform
Role in this project:
userML Engineer
Contributions:29 reviews, 30 PRs, 13 pushes in 5 months
Contributions summary:Sean made several contributions related to training and generation within the Dolly large language model project. They modified training parameters, specifically increasing evaluation steps. The user also added and updated NVIDIA library installations, and optimized model generation for specific hardware (A10, V100). Finally, they refactored the code to leverage the Hugging Face dataset and correctly set flags for bf16 training on specified hardware.
chatbotdatabricksdollygpt
sryza/aas

Aug 2014 - Jun 2022

Code to accompany Advanced Analytics with Spark from O'Reilly Media
Role in this project:
userData Scientist
Contributions:5 releases, 4 reviews, 193 commits in 7 years 11 months
Contributions summary:Sean primarily contributed to the implementation of clustering algorithms, specifically focusing on K-means and its variations. Their contributions included initial K-means code, along with iterations exploring scoring methods and normalization techniques. They also developed code to visualize the clustering results using R and incorporated categorical features. Furthermore, they implemented an anomaly detection system and a decision forest model, demonstrating a focus on applying machine learning techniques.
analyticso-reillyspark
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial
Sean Owen