Yury Zemlyanskiy is a machine learning and NLP researcher-engineer with 13 years of experience, currently a Member of Technical Staff at Anthropic and a Ph.D. student at USC under Fei Sha. He has driven production ML systems at Facebook and Google, built translation and moderation pipelines, and contributed to prominent open-source projects like Caffe2 and OpenCV—improving RNN APIs, memory optimization, and optical-flow algorithms. His work spans dataset engineering for dialogue (ParlAI/OpenSubtitles), memory-augmented models, and practical system optimizations, reflecting a blend of deep research and production-grade engineering. Based in Menlo Park, he combines academic rigor with a track record of shipping scalable ML components and non-obvious strengths in refactoring complex C++/Python codebases for maintainability and performance.
13 years of coding experience
15 years of employment as a software developer
Doctor of Philosophy - PhD, Computer Science, Doctor of Philosophy - PhD, Computer Science at University of Southern California
Mathematics and Computer Science, Mathematics and Computer Science at Saint Petersburg State University
Caffe2 is a lightweight, modular, and scalable deep learning framework.
Role in this project:
ML Engineer
Contributions:48 commits in 4 months
Contributions summary:Yury focused on enhancing the Caffe2 deep learning framework by simplifying the RNN (Recurrent Neural Network) API, improving its flexibility, and addressing limitations in handling complex network structures. They introduced a wrapper for the RecurrentNetworkOp and refactored the API to make it easier to use. The user also worked on memory optimization and gradient propagation for RecurrentNetworkOp.
A framework for training and evaluating AI models on a variety of openly available dialogue datasets.
Role in this project:
ML Engineer
Contributions:7 commits, 1 PR, 7 comments in 5 days
Contributions summary:Yury primarily focused on data processing and dataset creation for the ParlAI framework, specifically for the OpenSubtitles dataset. They implemented a new data source for OpenSubtitles2016, addressing deduplication, tokenization issues, and parallelization for efficient processing. The user also modified the existing OpenSubtitles data loaders, including merging 2009 and 2018 versions, and added support for dialogue history. The changes include modifying the data processing pipeline to support various dialog formats.
nlpdeep-learningdatasetmachine-learningtraining
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.