Principal Member Of Technical Staff at Oracle Labs
Medford, Massachusetts, United States
Join Prog.AI to see contacts
Join Prog.AI to see contacts
Summary
🤩
Rockstar
🎓
Top School
John Sullivan is a Principal Member of Technical Staff with 14 years of experience building scalable NLP, ML, and data-processing systems, currently based at Oracle Labs in Virginia. He bridges research and production engineering, having driven large-scale entity resolution, probabilistic modeling, and distributed MCMC systems during graduate research and later productionized ML components at Oracle and Cambridge Semantics. A frequent open-source contributor, he has deep back-end expertise in JVM and Scala ecosystems—contributing to FACTORIE and Tribuo with data loaders, metadata-aware iterators, and XGBoost feature support—plus numerical work in the Clojure-based Incanter. His work consistently emphasizes robust data integration, metadata handling, and model-to-production transitions, not just novel algorithms. He pairs a CS master's from UMass Amherst with practical experience delivering ETL, annotation, and document-processing pipelines for industrial and government stakeholders. Colleagues describe him as a pragmatic researcher-engineer who surfaces subtle data-quality and metadata issues early, saving downstream debugging and deployment time.
14 years of coding experience
7 years of employment as a software developer
BA, Religious Studies, BA, Religious Studies at Brown University
FACTORIE is a toolkit for deployable probabilistic modeling, implemented as a software library in Scala. It provides its users with a succinct language for creating relational factor graphs, estimating parameters and performing inference.
Role in this project:
Back-end Developer
Contributions:296 commits, 11 PRs, 56 pushes in 3 years 9 months
Contributions summary:John's commits primarily involve adding and modifying code related to loading and processing data, specifically shallow parsing data from the Conll2000 shared task. They implemented a data loader in Scala for the Conll2000 dataset and defined a BIOChunkDomain for handling the sentence chunks. Additionally, they made changes to the coreference model, adding a new feature. This indicates contributions focused on data integration and model enhancement.
Contributions:204 reviews, 10 commits, 8 PRs in 2 years 3 months
Contributions summary:John primarily contributed to the Java-based machine learning library, focusing on refactoring and improving data handling components. They refactored the `ColumnarIterator` and related classes to support complex metadata and example weight calculations. They also fixed bugs in data reading and processing, specifically in `ResultSetIterator` and `ColumnarDataSource`, and added features, such as full support for XGBoost feature importance metrics and multi-output support to the `ResponseProcessor`.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.