Douglas Wu is a Principal Scientist in computational biology with 11 years of experience building production-ready genomics pipelines and R&D methods that scale for clinical and commercial testing. He blends deep wet-lab-informed genomics expertise (PhD in Molecular Genetics) with practical software engineering—designing RESTful services, automating LIMS workflows, and refactoring Python microservices to support >1,000 samples/month. His work spans assay development for hard-to-sequence genes, integration of new sequencing platforms (Illumina, PacBio, RP-PCR), and quantitative methods for benchmarking and reproducible data pipelines. An active open-source contributor, he has improved pysam internals for better coverage and intron detection, underscoring a focus on performant, accurate genomic data handling. Colleagues know him for turning R&D prototypes into validated production systems that materially increased revenue per sample. Based in Rockville, MD, he combines hands-on algorithm development with operational rigor to close the gap between cutting-edge genomics and clinical deployment.
11 years of coding experience
13 years of employment as a software developer
Doctor of Philosophy (PhD) Molecular Genetics, Doctor of Philosophy (PhD) Molecular Genetics at The University of Texas at Austin
Open Science Grid User Summer School
Bachelor of Science (BS) Biochemistry, Bachelor of Science (BS) Biochemistry at University of Illinois Urbana-Champaign
Pysam is a Python package for reading, manipulating, and writing genomics data such as SAM/BAM/CRAM and VCF/BCF files. It's a lightweight wrapper of the HTSlib API, the same one that powers samtools, bcftools, and tabix.
Role in this project:
Back-end Developer
Contributions:22 commits, 2 PRs, 10 comments in 19 days
Contributions summary:Douglas primarily focused on improving the `pysam` library's functionality for genomics data processing. They made several commits addressing base quality handling within the `count_coverage` function and fixed related warnings. The user's work included refining the quality thresholding mechanism and optimizing the `find_introns` function for improved performance. These changes indicate a focus on data analysis and manipulation within the context of genomic sequencing data.
Find and Hire Top DevelopersWe’ve analyzed the programming source code of over 60 million software developers on GitHub and scored them by 50,000 skills. Sign-up on Prog,AI to search for software developers.