Marek Novotný is a Principal Software Engineer based in Prague with 13 years of experience building scalable data and machine learning platforms. He leads core backend work at H2O.ai, driving the Sparkling Water integration and H2O-3 platform while contributing Java/Scala/Python fixes that improve model APIs, serialization, and cross-platform compatibility. Marek has a strong track record in open-source—contributions to Apache Spark, H2O-3 and Sparkling‑Water include new array/map functions, stability fixes, and exposing important ML parameters (e.g., weightCol for multiple algorithms). His background in Big Data R&D at Barclays and early full-stack/.NET roles gives him a pragmatic blend of research, production engineering and architecture skills. Colleagues rely on him for technical leadership, test-stability improvements, and thoughtful refactors that ease long-term maintenance.
13 years of coding experience
13 years of employment as a software developer
Master's degree, Computer Science – Software Systems, Master's degree, Computer Science – Software Systems at Charles University in Prague, Faculty of Mathematics and Physics
Electronic Computer Systems, Electronic Computer Systems at High School of Electrical Engineering in Hluboká nad Vltavou
Sparkling Water provides H2O functionality inside Spark cluster
Role in this project:
Backend Developer & ML Engineer
Contributions:336 reviews, 1357 commits, 1201 PRs in 3 years 10 months
Contributions summary:Marek made several commits improving the hint descriptions for broadcast joins, fixing tests related to the use of non-native language, and implementing an annotation for deprecating legacy methods. A significant part of the work involves exposing the 'weightCol' parameter for DL, GBM, GLM, and XGBoost in the Python API and also renaming the 'predictionCol' parameter to 'labelCol'. Additionally, the user has made changes to the H2O GLRM model and updated code related to the SW backend.
Contributions:120 commits, 9 PRs, 7 pushes in 1 year 8 months
Contributions summary:Marek's contributions primarily focused on refactoring and improving the persistence layer of the `spline` repository, which tracks and visualizes data lineage. They separated the persistence layer into distinct modules, including API and Mongo, and updated the data model to accommodate the new features. They also added an Atlas persistence layer and adjusted the configuration to use the added persistence layers.
data-lineagevisualizationsparklineagetracking
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.