M.Tech Bioinformatics student at Delhi Technological University (graduating 2027). I build and validate bioinformatics pipelines for genomics, transcriptomics and drug discovery, and I test them against ground truth.
- clingenomics: ACMG/AMP germline and AMP/ASCO/CAP somatic variant interpretation engine in Python, with 62 tests, CI and a ClinVar concordance harness.
- ngs-variant-calling-pipeline: Snakemake and Docker pipeline benchmarking GATK against DeepVariant on NA12878 chr20 with the GIAB truth set (DeepVariant F1 99.02% for SNPs, 97.35% for indels).
- rnaseq-analysis-suite: bulk, single-cell and long-read RNA-seq pipelines scored against simulated ground truth, with the real-data drop reported.
- ad-target-dossier: harmonization of three Alzheimer's cohorts (494 samples) to ontology terms, with a Neo4j knowledge graph and Open Targets scores.
- lung-adenocarcinoma-deg-analysis: differential expression in 107 lung adenocarcinoma samples and survival testing in 497 TCGA-LUAD patients.
- g4-ligand-classifier: G-quadruplex ligand classifier with scaffold-grouped validation (CV AUC 0.94 on 367 compounds).
Python, R, SQL, Snakemake, Nextflow, Docker, GATK, DeepVariant, DESeq2, Scanpy, RDKit, scikit-learn, Neo4j