Chapter 25: Bioconductor
Shih Chien University
2026-07-20
Most of this book is field-agnostic; this chapter is the exception, focused on bioinformatics. Bioconductor is an open-source project of R packages for analyzing high-throughput genomic data — microarrays, next-generation sequencing, and more. It has its own release cycle, its own repository (separate from CRAN), and shared data structures that let its hundreds of packages interoperate.
The example analyzes a breast-cancer microarray study (GEO accession GSE2034, Affymetrix HG-U133A platform). Raw Affymetrix data lives in .CEL files, read in a batch with ReadAffy from the affy package:
ReadAffy returns an AffyBatch — unprocessed probe-level intensities plus the chip-definition (cdf) that maps probes to the 22,283 probe sets on HG-U133A.
Raw intensities must be quality-checked, background-corrected, normalized, summarized by probe set, and log-transformed before analysis:
rma (and relatives vsnrma, expresso) collapse the AffyBatch into an ExpressionSet — the central Bioconductor container.
Often the processed expression set is already public — GEOquery::getGEO downloads it directly:
No .CEL files needed — the analysis can start from a single download.
An ExpressionSet bundles everything one experiment needs, with accessor functions:
| Component | Accessor | Contents |
|---|---|---|
| Expression matrix | exprs() |
genes (rows) × samples (columns) |
| Phenotype data | pData() / phenoData() |
per-sample metadata (treatment, outcome…) |
| Feature data | fData() / featureData() |
per-gene annotation |
| Experiment info | experimentData() |
study-level description |
| Annotation | annotation() |
the platform, e.g. "GPL96" |
Real studies need clinical/phenotype data joined to the expression matrix. Replace or merge the pData:
Aligning samples between the two sources is the fiddly part — the book spends real effort getting the row order to match.
Once you have an ExpressionSet with matched phenotypes, the standard next step is finding differentially expressed genes — package limma (linear models for microarrays):
Tip
Why a separate ecosystem? Genomic data is huge, structured (genes × samples × annotation), and method-rich. Bioconductor’s shared ExpressionSet (and successors like SummarizedExperiment) let dozens of packages — normalization, testing, annotation, visualization — plug together. The R you learned transfers directly; only the data structures and the installer are special.
Copyright. These slides are adapted from R in a Nutshell: A Desktop Quick Reference (2nd ed.) by Joseph Adler, O’Reilly Media. All rights reserved by the original author and publisher.
Non-commercial use only. These materials are strictly for educational purposes and may not be used for commercial gain.
Attribution. Any reproduction, distribution, or use of these materials must properly credit the original source.
R in a Nutshell: A Desktop Quick Reference