DataGenerator is a Java library for systematically producing large volumes of data. DataGenerator frames data production as a modeling problem, with a user providing a model of dependencies among variables and the library traversing the model to produce relevant data sets.
Role in this project:
Back-end & DevOps Engineer Contributions:27 commits, 16 PRs, 14 comments in 9 months
Contributions summary:Brijesh primarily contributed to the `dg-spark` module, focusing on integrating DataGenerator with Apache Spark. Their work involved implementing a `SparkDistributor` to parallelize data generation across Spark clusters, and modifying existing engine components for Spark compatibility. They also addressed various syntax errors and code updates within the project, ensuring compatibility and fixing functional errors. These changes include updates to existing documentation.