Programming for Applications

Chapter 3: A Short R Tutorial

Yu-You Liou (NTU)

Shih Chien University

2026-08-03

Learning by Typing

How to Use This Chapter

This chapter is a hands-on tour of R, packed with examples. The best way to learn is to start R and type along:

  • enter exactly what the slides show — or change it slightly and see what happens (the book’s suggestion: if the example says 3 + 4, try 3 - 4);
  • everything happens at the console, so if terms like prompt or expression feel foreign, revisit Chapter 2 first.

By the end you will have touched every major idea in the course: vectors, functions, variables, data structures, classes, models, charts, and the help system.

Basic Operations in R

Expressions and Evaluation

Type an expression at the console, press Enter, and R evaluates it and prints the result (if there is one). Simple math works exactly as expected, including operator precedence and parentheses:

1 + 2 + 3
[1] 6
1 + 2 * 3
[1] 7
(1 + 2) * 3
[1] 9

The interactive interpreter automatically prints whatever object an expression returns — that is why you see the answers without asking for them.

The Mysterious [1]: Everything Is a Vector

Why does every answer start with [1]? Because any number you type is a vector — an ordered collection of numbers. The bracketed number is the index of the first element shown on that row. Each result above is a vector with a single element, hence [1].

Longer vectors are built with the c(...) function (c for combine):

c(0, 1, 1, 2, 3, 5, 8)   # first seven Fibonacci numbers
[1] 0 1 1 2 3 5 8

Indices Across Multiple Lines

The sequence operator : produces consecutive integers — handy for seeing how row labels work when output wraps:

1:50
 [1]  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25
[26] 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50

Each row begins with the index of its first element (e.g., [23], [45]), so you can locate any element at a glance.

Vector Arithmetic Is Pairwise

Operations on two vectors are matched element by element, returning a new vector:

c(1, 2, 3, 4) + c(10, 20, 30, 40)
[1] 11 22 33 44
c(1, 2, 3, 4) * c(10, 20, 30, 40)
[1]  10  40  90 160
c(1, 2, 3, 4) - c(1, 1, 1, 1)
[1] 0 1 2 3

This is the single most characteristic habit of R: you rarely loop over elements — you operate on whole vectors at once.

Recycling: Unequal Lengths

When the vectors differ in length, R recycles the shorter one, repeating it as many times as needed:

c(1, 2, 3, 4) + 1
[1] 2 3 4 5
1 / c(1, 2, 3, 4, 5)
[1] 1.0000000 0.5000000 0.3333333 0.2500000 0.2000000
c(1, 2, 3, 4) + c(10, 100)
[1]  11 102  13 104
c(1, 2, 3, 4, 5) + c(10, 100)
Warning in c(1, 2, 3, 4, 5) + c(10, 100): 較長的物件長度並非較短物件長度的倍數
[1]  11 102  13 104  15

Warning

Watch the warning. If the longer length is not a multiple of the shorter, R still computes an answer but warns you — usually a sign of a mistake in your code.

Character Vectors

Expressions can involve text too:

"Hello world."
[1] "Hello world."
c("Hello world", "Hello R interpreter")
[1] "Hello world"         "Hello R interpreter"
  • A quoted string is a character vector of length 1; combining two gives length 2.
  • Terminology alert for C programmers: in C, a “character” is a single letter and an ordered set is a “string”; in R, what C calls a string is a single character value.

Comments

Anything after a pound sign (#) on a line is ignored — even mid-expression:

# a comment at the beginning of a line
1 + 2 + # and one in the middle
  3
[1] 6

Comment generously: your future self is the most frequent reader of your code.

Functions

Calling Functions

In R, the operations that do the work are functions — just like math class. Most calls take the form f(argument1, argument2, ...):

exp(1)
[1] 2.718282
cos(3.141593)
[1] -1
log2(1)
[1] 0

Each of these takes a single argument. Many functions take several.

Named and Positional Arguments

Arguments can be supplied by name, or by position when given in the default order:

log(x = 64, base = 4)
[1] 3
log(64, 4)
[1] 3

Both calls are identical. Named arguments shine when a function has many parameters and you want to set only a few.

Operators Are Functions Too

Not every function looks like f(...). Some appear as operators:

17 + 2
[1] 19
2 ^ 10
[1] 1024
3 == 4
[1] FALSE
  • + is ordinary addition; ^ is exponentiation (note: not commutative); == tests equality and returns a Boolean (TRUE/FALSE) — yes, R has a logical data type.
  • Under the hood, the interpreter translates every operator into an equivalent function call — a fact we exploit in Chapter 5.

Variables

Assignment: x <- 1, Read “x gets 1”

R lets you attach names to values. The assignment operator is <-, pronounced “gets” — pleasingly close to algorithm pseudocode:

x <- 1
y <- 2
z <- c(x, y)
z
[1] 1 2

Substitution Happens at Assignment Time

The value is substituted when the assignment is made, not when the variable is later evaluated. Change y after building z, and z is unmoved:

y <- 4
z
[1] 1 2

The subtleties of how variables are evaluated (environments, lazy evaluation) wait until Chapter 8.

Indexing a Vector by Position

R offers several ways to refer to members of a vector:

b <- c(1,2,3,4,5,6,7,8,9,10,11,12)
b[7]          # the 7th item
[1] 7
b[1:6]        # items 1 through 6
[1] 1 2 3 4 5 6
b[c(1,6,11)]  # items 1, 6, and 11
[1]  1  6 11
b[c(8,4,9)]   # out of order: returned in the order referenced
[1] 8 4 9

The index vector can be any integer vector — contiguous, scattered, even shuffled.

Indexing with a Logical Vector

You can also select elements with a vector of TRUE/FALSE values. Example: keep only multiples of 3, i.e., elements congruent to 0 (mod 3):

b %% 3 == 0          # the logical mask
 [1] FALSE FALSE  TRUE FALSE FALSE  TRUE FALSE FALSE  TRUE FALSE FALSE  TRUE
b[b %% 3 == 0]       # apply it
[1]  3  6  9 12

This filter by condition idiom is the workhorse of data analysis in R — you will use it constantly.

Two More Assignment Operators

  • = also assigns, left to right, like most languages. Use it if you prefer — but the book (and these slides) stick with <- for readability. Never confuse = (assign) with == (test equality):
one <- 1; two <- 2
one = two    # assigns! one is now 2
one
[1] 2
one <- 1; two <- 2
one == two   # compares
[1] FALSE
  • -> assigns to the right — occasionally handy when you typed a long expression and only then remembered to save it:
3 -> three
three
[1] 3

Warning

Inside a function call, f(x <- 3) assigns x in your workspace and then passes the value — almost never what you intend. Use = to bind arguments: f(x = 3).

Functions Are Objects with Names

A function is just another object bound to a symbol. Define your own and call it like a built-in:

f <- function(x, y) { c(x + 1, y + 1) }
f(1, 2)
[1] 2 3

A useful trick follows: typing the name alone (no parentheses) prints the function’s source code:

f
function (x, y) 
{
    c(x + 1, y + 1)
}

This works on most built-in functions too — R is open source all the way down.

Introduction to Data Structures

Arrays

An array is a multidimensional vector. Internally, vectors and arrays are stored identically — an array is simply a vector wearing a dimension attribute — but arrays display and index differently:

a <- array(c(1,2,3,4,5,6,7,8,9,10,11,12), dim = c(3, 4))
a
     [,1] [,2] [,3] [,4]
[1,]    1    4    7   10
[2,]    2    5    8   11
[3,]    3    6    9   12
a[2, 2]      # row 2, column 2
[1] 5

Compare the same contents as a plain vector:

v <- c(1,2,3,4,5,6,7,8,9,10,11,12)
v
 [1]  1  2  3  4  5  6  7  8  9 10 11 12

Matrices and Higher Dimensions

A matrix is just a two-dimensional array:

m <- matrix(data = c(1,2,3,4,5,6,7,8,9,10,11,12), nrow = 3, ncol = 4)
m
     [,1] [,2] [,3] [,4]
[1,]    1    4    7   10
[2,]    2    5    8   11
[3,]    3    6    9   12

Arrays may have more than two dimensions:

w <- array(1:18, dim = c(3, 3, 2))
w[1, 1, 1]
[1] 1

(Printed in full, w appears as two 3×3 slices labeled , , 1 and , , 2.)

Slicing Arrays

R’s indexing syntax for parts of an array is clean: one index per dimension, separated by commas — and omitting an index selects everything in that dimension:

a[1, 2]        # single cell
[1] 4
a[1:2, 1:2]    # submatrix
     [,1] [,2]
[1,]    1    4
[2,]    2    5
a[1, ]         # first row only
[1]  1  4  7 10
a[, 1]         # first column only
[1] 1 2 3
a[1:2, ]       # a range of rows
     [,1] [,2] [,3] [,4]
[1,]    1    4    7   10
[2,]    2    5    8   11
a[c(1, 3), ]   # even a noncontiguous set of rows
     [,1] [,2] [,3] [,4]
[1,]    1    4    7   10
[2,]    3    6    9   12

Lists: Mixing Types

Everything so far held a single data type. Lists hold a heterogeneous collection of objects, with optionally named components — subtly different from “lists” in other languages:

e <- list(thing = "hat", size = "8.25")
e
$thing
[1] "hat"

$size
[1] "8.25"

Items are accessible by name or by position — note the three distinct forms:

e$thing    # by name: the element itself
[1] "hat"
e[1]       # single bracket: a sublist
$thing
[1] "hat"
e[[1]]     # double bracket: the element itself
[1] "hat"

Lists Inside Lists

A list can even contain other lists:

g <- list("this list references another list", e)
g
[[1]]
[1] "this list references another list"

[[2]]
[[2]]$thing
[1] "hat"

[[2]]$size
[1] "8.25"

This nesting ability is what lets R represent arbitrarily complex objects — fitted models, for instance, are big nested lists.

Data Frames

A data frame is a list of named vectors, all the same length — think spreadsheet or database table. It is the natural container for experimental data. The book’s example: 2008 win/loss records of the National League East:

teams <- c("PHI","NYM","FLA","ATL","WSN")
w <- c(92, 89, 84, 72, 59)
l <- c(70, 73, 77, 90, 102)
nleast <- data.frame(teams, w, l)
nleast
  teams  w   l
1   PHI 92  70
2   NYM 89  73
3   FLA 84  77
4   ATL 72  90
5   WSN 59 102

Picking Values Out of a Data Frame

Columns are reached with the $ operator (just like list components):

nleast$w
[1] 92 89 84 72 59

To find a specific value — say, losses by the Florida Marlins — build a logical vector and use it as a filter:

nleast$teams == "FLA"
[1] FALSE FALSE  TRUE FALSE FALSE
nleast$l[nleast$teams == "FLA"]
[1] 77

Chapter 11 shows how to import data frames from files and databases; beyond lists, R also offers formal class definitions via S4 objects for heterogeneous data.

Objects and Classes

Every Object Has a Class

R is an object-oriented language: every object has a type and belongs to a class. We have met several classes already — query them with class():

class(teams)
[1] "character"
class(w)
[1] "numeric"
class(nleast)
[1] "data.frame"
class(class)
[1] "function"

Note the last line: a function is an object of class function.

Methods and Generic Functions

  • Functions tied to a specific class are called methods. (R’s class system is far less formal than Java’s — not every function belongs to a class.)
  • Methods for different classes may share one name: such names are generic functions. They buy you two things:
    1. easy guessing — the function name you know probably works on the new class;
    2. code reuse — one piece of code can handle many types.

Example: + is generic. It adds numbers, but also knows what to do with a date and a number:

17 + 6
[1] 23
as.Date("2009-09-08") + 7
[1] "2009-09-15"

The Hidden print Call

When you evaluate an expression at the console, the interpreter silently calls the generic function print on the result:

x <- 1 + 2 + 3 + 4
x          # actually runs print(x)
[1] 10
  • Define a new class, define a print method for it, and you control how its objects display.
  • Some packages exploit this: lattice plotting functions return objects without drawing; the console’s automatic print is what plots them. Call a lattice function inside another function or script and nothing appears unless you print it explicitly — a classic gotcha.

Objects return in depth in Chapter 7; classes in Chapter 10.

Models and Formulas

What Is a Model?

To statisticians, a model concisely describes a set of data, usually via a mathematical formula. Two typical goals:

  • predictive — train on data, predict new values;
  • descriptive — understand the data better.

For a linear model predicting \(y\) from \(x_1, x_2, \ldots, x_n\) (the dependent and independent variables):

\[y = c_0 + c_1 x_1 + c_2 x_2 + \cdots + c_n x_n + \varepsilon\]

R expresses the relationship as y ~ x1 + x2 + ... + xn — a formula object.

Fitting a Linear Model: lm

The cars data set (in the base distribution, collected in the 1920s) records car speeds and stopping distances. Assume stopping distance is a linear function of speed; the formula is dist ~ speed, and lm estimates the parameters:

cars.lm <- lm(formula = dist ~ speed, data = cars)
cars.lm

Call:
lm(formula = dist ~ speed, data = cars)

Coefficients:
(Intercept)        speed  
    -17.579        3.932  

Printing an lm object shows the original call (so you can see data and formula) and the estimated coefficients.

More Detail: summary

summary(cars.lm)

Call:
lm(formula = dist ~ speed, data = cars)

Residuals:
    Min      1Q  Median      3Q     Max 
-29.069  -9.525  -2.272   9.215  43.201 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept) -17.5791     6.7584  -2.601   0.0123 *  
speed         3.9324     0.4155   9.464 1.49e-12 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 15.38 on 48 degrees of freedom
Multiple R-squared:  0.6511,    Adjusted R-squared:  0.6438 
F-statistic: 89.57 on 1 and 48 DF,  p-value: 1.49e-12

The summary reports the call, the residual distribution, the coefficient table (estimates, standard errors, \(t\)-values, \(p\)-values), and fit statistics (\(R^2\), \(F\)-statistic).

To Name or Not to Name?

You may also fit-and-inspect in one breath, without saving the model:

lm(dist ~ speed, data = cars)
summary(lm(dist ~ speed, data = cars))

Convenient — but naming the model is usually wiser:

  • follow-up analyses (residual plots, extra statistics, adding/dropping variables) need the object;
  • refitting a complex model on big data is slow; a saved object spares you;
  • named objects make code easier to read and modify.

Charts and Graphics

The Graphics Toolbox

R ships several visualization systems: graphics (the classic base system), grid, and lattice — with graphics and lattice the most used in the book’s era (and ggplot2, Chapter 15, dominating today).

  • Everything Excel can draw — column, bar, line, pie, scatter — R draws too, with far better automation and customization.
  • And then hundreds of chart types Excel never dreamed of.

To make this concrete, the book uses real data: every field goal attempted in the NFL in 2005. (A field goal: kicking the ball between the goalposts for 3 points; a miss hands possession to the other team at the spot of the kick.)

Loading the Example Data

The book loads data with library(nutshell); that package has left CRAN, so this course fetches the same data sets from the CRAN GitHub mirror:

temp <- tempfile(fileext = ".rda")
download.file("https://raw.githubusercontent.com/cran/nutshell/master/data/field.goals.rda",
              temp, mode = "wb")
load(temp)
names(field.goals)
 [1] "home.team"    "week"         "qtr"          "away.team"    "offense"     
 [6] "defense"      "play.type"    "player"       "yards"        "stadium.type"

names() lists the columns: home and away team, week, quarter, offense, defense, play type, yards, stadium type, player.

A First Histogram

The hist function shows a distribution in one line:

hist(field.goals$yards)

(Your colors may differ from the book’s — the author tweaked graphical parameters for print.)

Tuning the Histogram: breaks

Wanting more detail at different distances, increase the number of bins via the breaks argument:

hist(field.goals$yards, breaks = 35)

A single argument transforms the picture — typical of base graphics: sensible defaults, endless knobs.

Tabulating and Strip Charts

How many kicks were blocked? The table function tabulates a categorical variable:

table(field.goals$play.type)

FG aborted FG blocked    FG good      FG no 
         8         24        787        163 

A strip chart plots one point per observation along an axis. Select only blocked kicks, jitter the points so they don’t overprint, and restyle them with pch:

stripchart(field.goals[field.goals$play.type == "FG blocked", ]$yards,
           pch = 19, method = "jitter")

Scatterplots: the cars Data Again

The cars data set has 50 observations of speed and stopping distance:

dim(cars)
[1] 50  2
names(cars)
[1] "speed" "dist" 
summary(cars)
     speed           dist       
 Min.   : 4.0   Min.   :  2.00  
 1st Qu.:12.0   1st Qu.: 26.00  
 Median :15.0   Median : 36.00  
 Mean   :15.4   Mean   : 42.98  
 3rd Qu.:19.0   3rd Qu.: 56.00  
 Max.   :25.0   Max.   :120.00  

Plot the relationship, with labeled axes:

plot(cars, xlab = "Speed (mph)", ylab = "Stopping distance (ft)",
     las = 1, xlim = c(0, 25))

At a glance: stopping distance grows roughly in proportion to speed.

Lattice Graphics and Factors

Lattice excels at comparing groups. It is not loaded by default — calling a lattice function first requires:

library(lattice)

The example data: American food consumption, 1980–2005, in the consumption data set (48 observations: Amount consumed, Food type, Year, in 5-year steps).

temp <- tempfile(fileext = ".rda")
download.file("https://raw.githubusercontent.com/cran/nutshell/master/data/consumption.rda",
              temp, mode = "wb")
load(temp)

Two columns are numeric (Amount, Year); Food is a new type: a factor — R’s compact representation of categorical values, created with factor() and central to modeling functions.

A Conditioned Dot Plot

We want Amount against Year, drawn separately for each Food — lattice expresses this with the formula Amount ~ Year | Food:

dotplot(Amount ~ Year | Food, consumption)   # default settings

The book’s author found the defaults hard to read (labels too big, shared scales, awkward stacking) and tuned two options:

dotplot(Amount ~ Year | Food, data = consumption,
        aspect = "xy", scales = list(relation = "sliced", cex = .4))
  • aspect adjusts panel aspect ratios so changes bank toward 45° (easiest to perceive);
  • scales controls axis drawing — "sliced" gives each panel its own slice of the range.

Chapter 14 treats lattice in full.

Getting Help

The Help System

R documents every function in every installed package:

help(glm)      # full help page for glm
?glm           # exactly equivalent shorthand
?`+`           # operators need backquotes
example(glm)   # auto-run the examples from a help file

Searching for Help

Can’t remember the function’s name? Search by topic:

help.search("regression")   # list relevant topics
??regression                # shorthand for the same

For a whole package’s documentation, the help option of library is more complete than a single page:

library(help = "grDevices")

Vignettes

Some packages — especially from Bioconductor — include vignettes: short documents that walk through using the package, examples included.

vignette("affy")        # view a specific vignette (if installed)
vignette(all = FALSE)   # vignettes in attached packages
vignette(all = TRUE)    # vignettes in all installed packages

Tip

Habit to build now: when a function surprises you, read ?fun before searching the web — R’s built-in documentation is unusually good, and the examples actually run.