I work on gene expression data, and on whether you can trust what it tells you.
PhD Bioinformatics, University of West London (2026). I run the MSc Bioinformatics programme there and teach the computational modules.
What is here. Analysis work in R and Python: RNA-seq differential expression, variant calling, eQTL, ChIP-seq, DNA methylation, and some machine learning outside genomics. The numbered repositories are MSc coursework and are archived, so they are read-only.
What I am working on. Writing up two results about evaluation on small gene expression datasets. One is how much feature-selection leakage inflates cross-validated performance when marker genes are chosen on the whole dataset instead of inside each fold. The other is that on an imbalanced cohort, a model can reach 90% accuracy at a Cohen's kappa of exactly zero, which is to say it has learned nothing. Both came out of re-examining my own thesis pipeline.
Python · R · C++ · Bash · Git · Linux · Docker · Snakemake · Bioconductor · scikit-learn
