Chapter 21: Classification Models
Shih Chien University
2026-07-20
A classification model predicts a categorical response — spam or not, disease or healthy, which of several classes. The toolkit parallels Chapter 20, but the response is a factor. The chapter’s running example is Spambase: 4,601 emails with 57 numeric features, each labeled spam or not.
We move from linear methods (logistic regression, discriminant analysis) to trees and ensembles (bagging, boosting, random forests) to neural networks and SVMs.
The workhorse for a binary response: a glm with family=binomial, modeling the log-odds of the outcome:
summary reports coefficients, z-values, p-values, plus null and residual deviance and AIC (not R²).type="response" returns probabilities; threshold them (e.g., at 0.5) to get class labels.nnet::multinom (multinomial log-linear model), or ordered outcomes via MASS::polr.Discriminant analysis finds the (linear) combinations of predictors that best separate the classes — MASS::lda (and qda for quadratic boundaries):
Richer variants in package mda:
fda — flexible discriminant analysis (nonlinear boundaries via regression);mda — mixture discriminant analysis (each class a mixture of Gaussians).LDA assumes equal class covariances and roughly normal predictors; QDA relaxes the equal-covariance assumption.
For counts in a contingency table, log-linear models test associations among categorical variables — stats::loglin(table, margin, ...) (and the friendlier formula interface MASS::loglm):
Arguments to loglin: the table, a list of margin vectors to fit, an optional starting estimate, and maxit (iteration cap for the IPF algorithm).
rpart builds a tree by recursive partitioning — repeatedly splitting on the feature that best separates classes:
The printcp/plotcp output (CP, nsplit, rel error, xerror) guides pruning: pick the complexity parameter that minimizes cross-validated error.
Single trees are unstable; ensembles of trees are far more accurate:
ipred::bagging or adabag::bagging: fit many trees to bootstrap samples, average/vote. Reduces variance.ada::ada or adabag::boosting (AdaBoost), gbm::gbm (gradient boosting): fit trees sequentially, each focusing on the previous one’s errors. Reduces bias and variance.These are embarrassingly parallel — each tree fits independently (Chapter 26).
randomForest bags trees and randomizes the features considered at each split, decorrelating them — usually the strongest off-the-shelf classifier:
Key knobs: ntree (number of trees), mtry (features tried per split). The out-of-bag (OOB) error is a free, honest accuracy estimate — no separate test set needed.
nnet::nnet fits a single-hidden-layer feed-forward network:
size sets the hidden units, decay is weight regularization (curbing overfitting), maxit caps iterations. (Modern deep learning lives in keras/torch, but the idea is the same, scaled up.)
e1071::svm finds the maximum-margin boundary, kernelized for nonlinearity:
type="C-classification" is the standard classifier; kernel may be "linear", "polynomial", "radial" (RBF), or "sigmoid". Tune cost and kernel parameters with tune. The same package’s naiveBayes gives a fast probabilistic baseline.
The simplest classifier of all — class::knn(train, test, cl, k) — labels each point by majority vote of its k nearest neighbors; no model is “fit” at all.
Tip
No universal best. Logistic regression is interpretable; LDA needs few assumptions checked; trees are readable; random forests and boosting usually win on raw accuracy; SVMs shine in high dimensions; kNN is a quick baseline. Compare honestly with a held-out test set or cross-validation, and judge with a confusion matrix, not just overall accuracy.
Copyright. These slides are adapted from R in a Nutshell: A Desktop Quick Reference (2nd ed.) by Joseph Adler, O’Reilly Media. All rights reserved by the original author and publisher.
Non-commercial use only. These materials are strictly for educational purposes and may not be used for commercial gain.
Attribution. Any reproduction, distribution, or use of these materials must properly credit the original source.
R in a Nutshell: A Desktop Quick Reference