Programming for Applications

Chapter 25: Bioconductor

Yu-You Liou (NTU)

Shih Chien University

2026-07-20

R for Bioinformatics

What Bioconductor Is

Most of this book is field-agnostic; this chapter is the exception, focused on bioinformatics. Bioconductor is an open-source project of R packages for analyzing high-throughput genomic data — microarrays, next-generation sequencing, and more. It has its own release cycle, its own repository (separate from CRAN), and shared data structures that let its hundreds of packages interoperate.

Note

Installation differs from CRAN. Bioconductor packages install through its own installer — historically biocLite, today BiocManager:

install.packages("BiocManager")
BiocManager::install(c("affy", "GEOquery", "limma"))

A Worked Example

Loading Raw Expression Data

The example analyzes a breast-cancer microarray study (GEO accession GSE2034, Affymetrix HG-U133A platform). Raw Affymetrix data lives in .CEL files, read in a batch with ReadAffy from the affy package:

library(affy)
GSE2034 <- ReadAffy()                          # reads all .CEL in the working dir
# or name them explicitly:
GSE2034.smpl <- ReadAffy(filenames=paste0("GSM", 36777:36800, ".CEL"))
GSE2034                                         # an AffyBatch object

ReadAffy returns an AffyBatch — unprocessed probe-level intensities plus the chip-definition (cdf) that maps probes to the 22,283 probe sets on HG-U133A.

From Raw Data to Expression

Raw intensities must be quality-checked, background-corrected, normalized, summarized by probe set, and log-transformed before analysis:

library(simpleaffy)
qc(GSE2034)                          # quality-control report

# normalize + summarize in one step (several methods available):
eset <- rma(GSE2034)                 # robust multi-array average -> ExpressionSet
# or: vsnrma(GSE2034), expresso(GSE2034, ...)

rma (and relatives vsnrma, expresso) collapse the AffyBatch into an ExpressionSet — the central Bioconductor container.

Loading Preprocessed Data from GEO

Often the processed expression set is already public — GEOquery::getGEO downloads it directly:

library(GEOquery)
GSE2034.geo <- getGEO("GSE2034")     # a list of ExpressionSet objects
class(GSE2034.geo[[1]])
## [1] "ExpressionSet"

No .CEL files needed — the analysis can start from a single download.

The ExpressionSet

Anatomy of the Central Data Structure

An ExpressionSet bundles everything one experiment needs, with accessor functions:

Component Accessor Contents
Expression matrix exprs() genes (rows) × samples (columns)
Phenotype data pData() / phenoData() per-sample metadata (treatment, outcome…)
Feature data fData() / featureData() per-gene annotation
Experiment info experimentData() study-level description
Annotation annotation() the platform, e.g. "GPL96"
dim(exprs(eset))            # genes x samples
head(pData(eset))           # sample metadata

Matching Phenotype Data

Real studies need clinical/phenotype data joined to the expression matrix. Replace or merge the pData:

# build phenotype data and attach it
pData(eset) <- new("AnnotatedDataFrame", data=clinical.df)
# or merge GEO phenotype with your own:
merged <- merge(pData(GSE2034.geo[[1]]), my.clinical, by="sample")
pData(GSE2034.geo[[1]]) <- merged

Aligning samples between the two sources is the fiddly part — the book spends real effort getting the row order to match.

Downstream: Differential Expression

Once you have an ExpressionSet with matched phenotypes, the standard next step is finding differentially expressed genes — package limma (linear models for microarrays):

library(limma)
design <- model.matrix(~ pData(eset)$outcome)
fit <- lmFit(eset, design)           # fit a linear model per gene
fit <- eBayes(fit)                   # empirical-Bayes moderation
topTable(fit, number=20)             # top differentially expressed genes

Tip

Why a separate ecosystem? Genomic data is huge, structured (genes × samples × annotation), and method-rich. Bioconductor’s shared ExpressionSet (and successors like SummarizedExperiment) let dozens of packages — normalization, testing, annotation, visualization — plug together. The R you learned transfers directly; only the data structures and the installer are special.