Big data in genomics, anything sequencing. From ultra-long reads (Oxford Nanopore, PacBio) to short reads (Illumina), and the bioinformatics that ties it together.
Which structural variants shape the traits we breed for?
What do long reads reveal that short reads can't?
How do we turn petabytes of sequence into decisions?
Which recessive variants hide in healthy carriers?
Can a pipeline stay reproducible at population scale?
What story is the genome trying to tell?
Which methylation marks track with health and fertility?
How well can we impute rare variants from cheap SNP data?
What does the microbiome add to the host's genome?
Your question here — what should we go find out together?
01
Reading the genome from every angle. Nanopore and PacBio ultra-long reads to assemble genomes and resolve structural variation at population scale, the repetitive regions short reads can't reach, plus bulk and single-cell RNA-seq to see which genes switch on, in which cells, and when. Earlier work built de novo transcriptomes and receptor repertoires for non-model species.
02
Finding signal in large genomic datasets: GWAS and genomic prediction, sequence imputation, and mapping quantitative traits.
03
Building scalable, reproducible pipelines in Nextflow, plus the small scripts and RNA-seq guidelines that quietly hold a project together.
High-performance computing with SLURM/PBS and a lot of Bash scripting.
Advanced. Magic tricks with tidyr, dplyr, Bioconductor, and plenty of ggplot2 styling.
Intermediate. Quick tools and data analysis scripts, here and there.
Workflow management for scalable bioinformatics pipelines.
Version control for everything: code, pipelines and the occasional 3am fix.
Cloud computing with EC2 and S3 when the laptop isn't enough.
Increasingly applying LLMs and GenAI to bioscience: agentic assistants for genomic workflows, literature and variant triage, and explainable AI (XAI) over big datasets. Also dabbling in Rust for when speed really matters (well, just a lil bit).
Three chapters, one thread: reading DNA and RNA to answer real questions, whether the subject has fins, a shell, or four stomachs.
Research for the Australian dairy industry, moving from SNP arrays to long-read sequencing to find important genetic variants and map complex traits. This is where population-scale genomics meets structural variation: imputation accuracy, GWAS for milk and blood urea nitrogen, carrier frequencies of recessive defects, and fine-mapping fertility QTL.
A PhD in aquaculture genomics at the University of the Sunshine Coast, my first deep dive into genomes. De novo transcriptomes and neuropeptide/GPCR repertoires across crayfish, prawns and lobsters (Cherax, Macrobrachium, Penaeus, Nephrops, Panulirus), chasing the genes behind reproduction, sexual development and social behaviour in non-model species.
Where it all started: molecular genetics and biotechnology at International University, Ho Chi Minh City National University, including transcriptomic work on salinity adaptation in striped catfish (Pangasianodon hypophthalmus). The first time I saw a genome tell its story.