Programming for Applications

Chapter 21: Classification Models

Yu-You Liou (NTU)

Shih Chien University

2026-07-20

Predicting Categories

Regression’s Categorical Cousin

A classification model predicts a categorical response — spam or not, disease or healthy, which of several classes. The toolkit parallels Chapter 20, but the response is a factor. The chapter’s running example is Spambase: 4,601 emails with 57 numeric features, each labeled spam or not.

We move from linear methods (logistic regression, discriminant analysis) to trees and ensembles (bagging, boosting, random forests) to neural networks and SVMs.

Linear Classification Models

Logistic Regression

The workhorse for a binary response: a glm with family=binomial, modeling the log-odds of the outcome:

spam.model <- glm(is_spam ~ ., data=spambase, family=binomial(link="logit"))
summary(spam.model)
predict(spam.model, newdata, type="response")   # predicted probabilities
  • summary reports coefficients, z-values, p-values, plus null and residual deviance and AIC (not R²).
  • type="response" returns probabilities; threshold them (e.g., at 0.5) to get class labels.
  • “Number of Fisher Scoring iterations” reports how quickly the IRLS fit converged.
  • Multi-class extension: nnet::multinom (multinomial log-linear model), or ordered outcomes via MASS::polr.

Linear and Flexible Discriminant Analysis

Discriminant analysis finds the (linear) combinations of predictors that best separate the classes — MASS::lda (and qda for quadratic boundaries):

library(MASS)
iris.lda <- lda(Species ~ ., data=iris)
predict(iris.lda, iris)$class          # predicted classes

Richer variants in package mda:

  • fda — flexible discriminant analysis (nonlinear boundaries via regression);
  • mda — mixture discriminant analysis (each class a mixture of Gaussians).

LDA assumes equal class covariances and roughly normal predictors; QDA relaxes the equal-covariance assumption.

Log-Linear Models

For counts in a contingency table, log-linear models test associations among categorical variables — stats::loglin(table, margin, ...) (and the friendlier formula interface MASS::loglm):

library(MASS)
loglm(~ Sex + Class + Survived, data=Titanic)   # mutual independence
loglm(~ Sex*Class + Survived, data=Titanic)     # Sex,Class associated

Arguments to loglin: the table, a list of margin vectors to fit, an optional starting estimate, and maxit (iteration cap for the IPF algorithm).

Trees and Ensembles

Classification Trees: rpart

rpart builds a tree by recursive partitioning — repeatedly splitting on the feature that best separates classes:

library(rpart)
spam.tree <- rpart(is_spam ~ ., data=spambase)
spam.tree                                  # text description of the splits
printcp(spam.tree)                         # CP table: complexity vs. error
plot(spam.tree); text(spam.tree)           # draw it
spam.tree.pruned <- prune(spam.tree, cp=0.01)   # prune to avoid overfitting

The printcp/plotcp output (CP, nsplit, rel error, xerror) guides pruning: pick the complexity parameter that minimizes cross-validated error.

Bagging and Boosting

Single trees are unstable; ensembles of trees are far more accurate:

  • Bagging (bootstrap aggregating) — ipred::bagging or adabag::bagging: fit many trees to bootstrap samples, average/vote. Reduces variance.
  • Boosting — ada::ada or adabag::boosting (AdaBoost), gbm::gbm (gradient boosting): fit trees sequentially, each focusing on the previous one’s errors. Reduces bias and variance.
library(adabag)
spam.bag  <- bagging(is_spam ~ ., data=spambase)
spam.boost <- boosting(is_spam ~ ., data=spambase)
predict(spam.boost, newdata)$class

These are embarrassingly parallel — each tree fits independently (Chapter 26).

Random Forests

randomForest bags trees and randomizes the features considered at each split, decorrelating them — usually the strongest off-the-shelf classifier:

library(randomForest)
spam.rf <- randomForest(is_spam ~ ., data=spambase, ntree=500)
spam.rf                                # OOB error estimate + confusion matrix
importance(spam.rf)                    # variable importance
predict(spam.rf, newdata)

Key knobs: ntree (number of trees), mtry (features tried per split). The out-of-bag (OOB) error is a free, honest accuracy estimate — no separate test set needed.

Neural Networks and SVMs

Neural Networks

nnet::nnet fits a single-hidden-layer feed-forward network:

library(nnet)
spam.nn <- nnet(is_spam ~ ., data=spambase, size=10, decay=0.01, maxit=200)
predict(spam.nn, newdata, type="class")

size sets the hidden units, decay is weight regularization (curbing overfitting), maxit caps iterations. (Modern deep learning lives in keras/torch, but the idea is the same, scaled up.)

Support Vector Machines

e1071::svm finds the maximum-margin boundary, kernelized for nonlinearity:

library(e1071)
spam.svm <- svm(is_spam ~ ., data=spambase,
                type="C-classification", kernel="radial")
predict(spam.svm, newdata)
naiveBayes(is_spam ~ ., data=spambase)   # also in e1071: naive Bayes

type="C-classification" is the standard classifier; kernel may be "linear", "polynomial", "radial" (RBF), or "sigmoid". Tune cost and kernel parameters with tune. The same package’s naiveBayes gives a fast probabilistic baseline.

k-Nearest Neighbors, and Choosing a Model

The simplest classifier of all — class::knn(train, test, cl, k) — labels each point by majority vote of its k nearest neighbors; no model is “fit” at all.

Tip

No universal best. Logistic regression is interpretable; LDA needs few assumptions checked; trees are readable; random forests and boosting usually win on raw accuracy; SVMs shine in high dimensions; kNN is a quick baseline. Compare honestly with a held-out test set or cross-validation, and judge with a confusion matrix, not just overall accuracy.