Walid Gara is a Senior Data Engineer based in Paris with nine years of experience building and optimizing large-scale data platforms across telecom and utilities. He combines hands-on Spark optimization and data architecture with leadership roles at SFR and Carrefour while co-founding a startup, demonstrating both enterprise and entrepreneurial impact. A strong open-source contributor, Walid has implemented key pandas-on-Spark features in the popular Koalas project and added online learning algorithms to River and MOA, bridging batch and streaming ML. His background in telecom networks and formal data science training from Télécom Paris and Université Paris-Saclay informs pragmatic, performance-focused designs. Comfortable across DevOps, CI/CD and production pipelines, he often tackles infrastructure automation and containerization to make ML systems reproducible. Notably, his contributions span low-level API work to deployment automation, reflecting a rare blend of algorithmic depth and production engineering.
9 years of coding experience
4 years of employment as a software developer
Engineer's degree Computer Systems Networking and Telecommunications, Engineer's degree Computer Systems Networking and Telecommunications at SUP'COM
Mathématiques Physique, Mathématiques Physique at IPEIT - Institut Préparatoire aux Etudes d'Ingénieurs de Tunis
Engineer's degree Data Science, Engineer's degree Data Science at Télécom Paris
M2 Data & Knowledge (D&K) Big Data, M2 Data & Knowledge (D&K) Big Data at Université Paris-Saclay
MOA is an open source framework for Big Data stream mining. It includes a collection of machine learning algorithms (classification, regression, clustering, outlier detection, concept drift detection and recommender systems) and tools for evaluation.
Role in this project:
DevOps Engineer
Contributions:3 reviews, 13 commits, 3 PRs in 3 years 6 months
Contributions summary:Walid primarily focuses on infrastructure and deployment aspects, evident through the creation and modification of Dockerfiles. Their contributions include building Docker images for development and GUI environments, integrating with the MOA project. They also addressed CICD pipeline issues, incorporating changes to Travis CI and improving the Docker setup. These changes highlight their role in automating the build, testing, and deployment processes for the project.
A machine learning package for streaming data in Python. The other ancestor of River.
Role in this project:
ML Engineer
Contributions:3 reviews, 43 commits, 9 PRs in 1 year 8 months
Contributions summary:Walid contributed to the development and implementation of machine learning models within the scikit-multiflow repository. Their work involved adding new methods to existing machine learning models, specifically within the Hoeffding Tree implementation, including the Hoeffding Anytime Tree. This involved creating new classes and methods related to tree structure, node behavior, and split evaluation. The user's contributions expand the capabilities of the library, enabling more advanced and flexible stream learning techniques.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.