Nandan Thakur is a Ph.D. candidate at the University of Waterloo specializing in information retrieval and NLP, with eight years of industry and research experience spanning academia and tech labs. He focuses on heterogeneous benchmarking and robust, efficient neural retrieval methods for low-resource and specialized domains, and makes his work accessible by releasing easy-to-use code. Nandan contributed substantially to BEIR, a widely adopted open-source IR benchmark (NeurIPS 2021) with strong GitHub traction, and has implemented retrieval pipelines using Elasticsearch, DPR, DPR-style binary retrievers, and cross-encoder rerankers. His internships at Google Research, Databricks/MosaicML, and collaborations with Vectara and Huawei reflect a blend of production-minded engineering and rigorous evaluation. Notably, he brings cross-disciplinary experience from computational biology to large-scale enterprise systems, enabling practical solutions that bridge research and deployable tooling.
8 years of coding experience
4 years of employment as a software developer
BITS Pilani, Birla Institute of Technology and Science
High School, High School at Modern School, Barakhamba Road
Doctor of Philosophy - PhD, Computer Science, Doctor of Philosophy - PhD, Computer Science at University of Waterloo
A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
Role in this project:
Back-end Developer & Data Scientist
Contributions:8 releases, 6 reviews, 344 commits in 2 years
Contributions summary:Nandan contributed primarily to the development of a retrieval system, including implementing search functionality using Elasticsearch and incorporating various models like DPR and a custom model for generating questions and improving retrieval performance. They also implemented a Binary Passage Retriever and applied a cross-encoder model for document reranking. Their work involves building the retrieval components and improving performance.
INCOME: An Easy Repository for Training and Evaluation of Index Compression Methods in Dense Retrieval. Includes BPR and JPQ.
Contributions:43 commits, 6 PRs, 23 pushes in 8 months
compression
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.