Statistical Genomics 2026
1 Course Description
High-throughput ’omics studies generate ever larger datasets and, as a consequence, complex data interpretation challenges. This course focusses on statistical concepts involved in preprocessing, quantification and differential analysis of high-throughput ’omics data. The core focus will be on shotgun proteomics and (bulk and single-cell) RNA-sequencing. Experimental design is essential to allow for correct interpretation in all ’omics studies, and we will cover how to design a statistically efficient experiment, as well as discuss the impact experimental design has on how we model ’omics data, introducing concepts such as blocking. The course will rely exclusively on free and user-friendly open-source tools in R/Bioconductor. We hope that this will provide a solid basis for beginners, but will also bring new perspectives to those already familiar with standard data analysis workflows for proteomics and next-generation sequencing applications.
2 Target Audience
This course is oriented towards biologists and bioinformaticians with a particular interest in differential analysis for quantitative ’omics data.
3 Prerequisites
The prerequisites for the Statistical Genomics Analysis course are the successful completion of a basic course of statistics that covers topics on data exploration and descriptive statistics, statistical modeling, and inference: linear models, confidence intervals, t-tests, F-tests, anova, chi-squared test. The basis concepts may be revisited in the online course at https://statomics.github.io/PSLS/ (English) and in https://statomics.github.io/sbc/ (Dutch).
In addition, knowledge of programming in R is preferred. A primer to R and Data visualization in R can be found at:
RBasics: https://dodona.ugent.be/nl/courses/335/RData Exploration: https://dodona.ugent.be/nl/courses/345/
4 Software and data
For the course you need to install R 4.6 or higher from cran.
We also recommend you to use rstudio (optional)
To install all the necessary package, please use R 4.6 or higher and execute the code chunk below in the RStudio Console/R console:
if (!requireNamespace("BiocManager", quietly = TRUE))
install.packages("BiocManager")
BiocManager::install(c(
"arrow",
"BiocParallel",
"BiocFileCache",
"ComplexHeatmap",
"dplyr",
"ExploreModelMatrix",
"ggpattern",
"ggplot2",
"ggrepel",
"impute",
"MsDataHub",
"patchwork",
"scater",
"tidyr",
"bookdown",
"iq",
"QFeatures",
"msqrob2",
"kableExtra",
"data.table",
"ggcorrplot",
"ggpubr",
"callr",
"curl",
"httpuv"
))5 Detailed Program
Introduction (Week 1)
Module I: Proteomics Data Analysis (Week 2-5)
1. Bioinformatics for proteomics
- Lecture MS-Basics: PDF
- Lecture Identification: PDF
- Tutorials: Identification
2. Statistics for Proteomics Data Analysis
2.1 Data Processing
This parts introduces the user to the key concepts for data processing in differential proteomics data analysis and provides extensive description of the code. While this part is conceptual, the concepts are illustrated using a real spike-in study.
2.2 Statistical Inference & Design concepts
License
This material is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. You are free to share (copy and redistribute the material in any medium or format) and adapt (remix, transform, and build upon the material) for any purpose, even commercially, as long as you give appropriate credit and distribute your contributions under the same license as the original.