Tomoki Hayashi

Postdoctoral Researcher at 名古屋大学

Nagoya, Japan
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts
email-iconphone-icongithub-logolinkedin-logotwitter-logostackoverflow-logofacebook-logo
Join Prog.AI to see contacts

Summary

🤩
Rockstar
🎓
Top School
Tomoki Hayashi is a postdoctoral researcher and entrepreneur based in Nagoya with eight years’ experience in statistical speech and audio signal processing. He is a main developer of the widely used ESPnet end-to-end speech toolkit, contributes backend implementations to high-profile neural vocoder projects like ParallelWaveGAN, and serves as COO of Human Dataware Lab while holding a research position at Nagoya University. His work bridges rigorous academic research (PhD in information science) and production-ready open-source engineering, earning him the IEEE SPS Japan Young Author Best Paper Award and ASJ Itakura Award. Less obvious: he combines hands-on implementation of core model components (residual/up-sampling modules and generator design) with tooling and compatibility fixes that improve reproducibility and installer robustness for the speech research community.
code8 years of coding experience
book五条高等学校
book博士, 情報科学研究科 メディア科学専攻, 博士, 情報科学研究科 メディア科学専攻 at 名古屋大学
github-logo-circle

Github Skills (12)

neural-network10
pytorch10
audio-processing10
convolutional-neural-networks10
python10
scripting9
bash9
model-building9
version-control9
script9
git9
software-packaging7

Programming languages (11)

TypeScriptC++ShellJavaScriptGoLuaHTMLJupyter Notebook

Github contributions (5)

github-logo-circle
espnet/espnet

Jun 2018 - Jan 2023

End-to-End Speech Processing Toolkit
Role in this project:
userBack-end Developer
Contributions:50 releases, 252 reviews, 3190 commits in 4 years 7 months
Contributions summary:Tomoki's commits focus on updating the version number in the project's setup file, adding compatibility features, and fixing bugs related to the installation process. The contributions include adding fill_missing_args for compatibility, fixing issues related to pretrained models, and making adjustments to the scripting. The work primarily involves modifying existing code files, specifically relating to setup and utility scripts.
speech-recognitionspeech-separationchainerspoken-language-understandingspeech-processing
kan-bayashi/ParallelWaveGAN

Oct 2019 - Jan 2023

Unofficial Parallel WaveGAN (+ MelGAN & Multi-band MelGAN & HiFi-GAN & StyleMelGAN) with Pytorch
Role in this project:
userBack-end Developer
Contributions:22 releases, 28 reviews, 1150 commits in 3 years 3 months
Contributions summary:Tomoki's contributions focus on implementing and integrating fundamental audio processing building blocks for a speech generation model. Specifically, the user implemented a residual block module and an upsampling module, both crucial components for the generation process within the WaveGAN framework. The user also built the generator class, which combines and configures different audio layers to build a usable model.
realtimebandneural-vocodermelganparallel-wavenet
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.
Request Free Trial