AI for life science — multi-omics, molecular modelling, and small structural tools
I work on the boundary between computation and wet-lab biology: multi-omics data pipelines, representation learning for high-dimensional assays, and small self-contained utilities for structural questions. Currently based in Hangzhou.
- Multi-omics data engineering: automated feature annotation and isomer resolution for multi-omics and especially lipidomics — the parts nobody enjoys, but the initiation of a story.
- Representation learning: autoencoders for high-dimensional, low-sample-size omics data, and the question of what to transfer when the feature columns do not travel.
- Structural tools: small utilities that answer one question from a structure alone. No simulation when a cheaper answer exists.
- Interpretability: gradCAM, SHAP, and a general preference for models whose behaviour can be explained to a chemist.
- tc-coupling-profile — a structure-derived indicator (block total correlation of the Gaussian Network Model) for where assuming independent residue motion costs the most. Pure numpy, no simulation, no training.
- omics-ae-classifier — latent-space representation learning for high-dimensional low-sample-size omics: pre-train an autoencoder on a discovery cohort, freeze the encoder, adapt the head.
- Languages: Python, R
- Deep / machine learning: PyTorch, scikit-learn, LightGBM, autoencoders, graph neural networks
- Omics: mass-spectrometry data, spatial omics, single-cell tooling
- Computational chemistry: Schrödinger, AutoDock, FEP workflows, molecular dynamics analysis
- Interpretability: gradCAM, SHAP
Despite daily encounters with complex algorithms, I still consider myself a coding novice on a lifelong learning journey. Away from the keyboard I am an enthusiast of palaeontology and entomology — endlessly fascinated by the structural evolution of life across geological time.