Hari Shreedharan

Principal Engineer at The Apache Software Foundation

Sunnyvale, California, United States
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Hari Shreedharan is a Principal Engineer with 14 years of experience designing and shipping large-scale, fault-tolerant distributed systems, currently building cloud-native streaming infrastructure at Confluent. He combines deep backend expertise with pragmatic product delivery from roles at Snowflake, StreamSets, Cloudera and Yahoo, and holds multiple patents in progressive error handling and data drift mitigation for resilient data pipelines. An active Apache committer and PMC member—formerly chair—he has contributed to high-profile projects like Apache Flink-adjacent Flume, Sqoop, and Spark (notably improving Spark–Flume integration and streaming receivers). Hari pairs hands-on engineering with mentorship across open-source communities and has repeatedly solved tricky production issues such as reliable Kafka/Flume sinks and HBase file-channel durability. Based in Sunnyvale, he brings both academic rigor from Cornell and a track record of turning complex ingestion challenges into dependable, testable systems.
code14 years of coding experience
job15 years of employment as a software developer
bookMasters Computer Science, Masters Computer Science at Cornell University
bookBachelors IT, Bachelors IT at Malaviya National Institute of Technology Jaipur
bookCKMNSS
stackoverflow-logo

Stackoverflow

Stats
136reputation
14kreached
3answers
0questions
github-logo-circle

Github Skills (33)

hbase10
api-rest10
apache-spark10
spark10
flume10
back-end-development10
api-design10
restful-api10
testing10
big-data10
kafka10
java10
scala10
javas10
sqoop10

Programming languages (4)

JavaScalaGroovyPython

Github contributions (5)

github-logo-circle
apache/logging-flume

May 2012 - Feb 2016

Apache Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log-like data
Role in this project:
userBack-end Developer & DevOps Engineer
Contributions:587 commits, 26 comments in 3 years 9 months
Contributions summary:Hari's contributions primarily involve improvements to the HBase sink within the Apache Flume project, enhancing its stability and features. They addressed issues related to file handling, including directory creation failures, potential data loss in the file channel, and the correct closing of file streams. The user also contributed to the overall performance of the Kafka sink by handling re-connections and improving batching efficiency. Furthermore, the user implemented solutions to manage failures within the system and ensure that events are still written to the channel.
apachebig-dataflumejavaapache-flume
cloudera/livy

Nov 2015 - Mar 2016

Livy is an open source REST interface for interacting with Apache Spark from anywhere
Role in this project:
userBack-end Developer
Contributions:10 commits, 24 PRs, 20 pushes in 4 months
Contributions summary:Hari primarily focused on enhancing the Livy project by implementing and integrating a remote Spark client. This involved adding a new class to manage sessions and clients, standardizing dependencies on SLF4J, and introducing functionalities to create, manage and submit jobs. Further contributions extended to include a servlet for handling Livy client sessions, enabling session creation, teardown, and the submission of both synchronous and asynchronous jobs. The user also added support for file and jar uploads.
emrapacheanywheresparkscala
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial