Programming for Applications

Chapter 7: R Objects

Yu-You Liou (NTU)

Shih Chien University

2026-08-03

The Building Blocks

Types and Classes

All objects in R rest on a basic set of built-in objects:

  • an object’s type defines how it is stored;
  • its class defines what information it contains and how it may be used.

This chapter covers the built-in objects themselves — the object-oriented machinery (class definitions, inheritance, methods) waits until Chapter 10. Note that “object-oriented programming” means more than “programming with objects”!

A Map of the Primitive Types

The built-in object types sort into a few categories:

  • Basic vectors — single-type collections: integers, doubles, complex, character, logical, raw.
  • Compound objects — containers of named objects: lists, pairlists, S4 objects, environments.
  • Special objects — any, NULL, ...: each means something important in context, but you never create one yourself.
  • R language objects — objects representing R code, evaluable to other objects.
  • Functions — the workhorses: take arguments, return outputs, sometimes cause side effects (plots, files, network).
  • Internal / bytecode — types defined by R but normally invisible to user-level code.

Primitive Types (1): Basic Vectors

Type Description Example
integer Whole numbers; produced naturally by sequences 5:5, integer(5)
double Floating-point (8 bytes on modern platforms); the default for numeric values 1, -1, 2^50
complex Complex numbers; the imaginary part carries an i suffix (a bare 3i is valid) 2+3i, exp(0+1i*pi)
character Text strings (a “string” in other languages) "Hello world."
logical Boolean values TRUE, FALSE
raw Raw bytes; for encoding objects from outside R raw(8), charToRaw("Hello")

Primitive Types (2): Compound Objects

Type Description Example
list A possibly heterogeneous, optionally named collection; data frames are built on lists list(1, 2, "hat")
pairlist Name–value pairs, mainly internal; deprecated in user code (plain lists are as efficient and more flexible) .Options
S4 Objects supporting modern OO: inheritance, methods (Chapter 10) —
environment The set of symbol–value pairs in a context, plus a pointer to an enclosing environment .GlobalEnv, new.env()

Primitive Types (3): Special and Language

Type Description Example
any “Any type is OK” — prevents coercion; used in S4 slots and generic signatures representation(data="ANY")
NULL “There is no object”; can have no attributes NULL
... Variable-length argument lists, especially arguments passed through to other functions —
symbol A language object referring to another object as.name("x"), quote(x)
promise Evaluated on first use, not creation; implements delayed loading in packages delayedAssign("v", c(x,y,z))
language Objects representing R code itself quote(function(x) {x+1})
expression An unevaluated expression; create with expression(), run with eval() expression(1 + 2)

Primitive Types (4): Functions and Internal

Type Description Example
closure Functions written in R: user-defined, most of base R, most packages function(x) {x+1}, print
special Internal functions whose arguments are not necessarily evaluated if, [
builtin Internal functions that evaluate their arguments +, ^
bytecode Compiled R functions from the compiler package cmpfun(function(x) {x^2})
char Scalar “string”; character vectors are made of these (users never touch them) —
externalptr External pointer, used in C code —
weakref Weak reference (internal only) —

Vectors

Building Vectors: c and Coercion

The six basic vector types appear constantly. The simplest constructor is c, which combines its arguments — and coerces them all to a single type:

v <- c(.295, .300, .250, .287, .215)
v
[1] 0.295 0.300 0.250 0.287 0.215
v <- c(.295, .300, .250, .287, "zilch")   # one character value...
v                                          # ...and everything is character
[1] "0.295" "0.3"   "0.25"  "0.287" "zilch"

Recursive Combination

c can assemble a vector from nested structures with recursive=TRUE — but beware: a list argument without it gives you back a list:

v <- c(.295, .300, .250, .287, list(.102, .200, .303), recursive=TRUE)
v
[1] 0.295 0.300 0.250 0.287 0.102 0.200 0.303
typeof(v)
[1] "double"
v <- c(.295, .300, .250, .287, list(1, 2, 3))
typeof(v)
[1] "list"
class(v)
[1] "list"

Sequences and the length Attribute

Two more vector builders — the : operator and the more flexible seq:

1:10
 [1]  1  2  3  4  5  6  7  8  9 10
seq(from=5, to=25, by=5)
[1]  5 10 15 20 25

A vector’s length can be manipulated directly — shrinking discards, expanding fills with NA:

w <- 1:10
length(w) <- 5
w
[1] 1 2 3 4 5
length(w) <- 10
w
 [1]  1  2  3  4  5 NA NA NA NA NA

Lists

Ordered, Optionally Named, Heterogeneous

A list is an ordered collection of objects, indexable by position — recall the [ vs. [[ distinction:

l <- list(1, 2, 3, 4, 5)
l[1]      # a sublist
[[1]]
[1] 1
l[[1]]    # the element
[1] 1

Named elements model real things naturally. A physical parcel — destination New York, 2×6×9 inches, $12.95 postage — mixes three data types in one object:

parcel <- list(destination="New York", dimensions=c(2, 6, 9), price=12.95)
parcel$price
[1] 12.95

Lists are the building block for heterogeneous structures — data frames are built on them.

Other Objects

Matrices

A matrix extends a vector to two dimensions, holding two-dimensional data of one type. The clean constructor is matrix, here with named rows and columns:

m <- matrix(data=1:12, nrow=4, ncol=3,
            dimnames=list(c("r1", "r2", "r3", "r4"),
                          c("c1", "c2", "c3")))
m
   c1 c2 c3
r1  1  5  9
r2  2  6 10
r3  3  7 11
r4  4  8 12
  • Other structures convert via as.matrix.
  • Implementation note: a matrix is stored as a vector, not as a vector of vectors — subscripts are an access convention, not the storage layout. And unlike most classes, matrices carry no explicit class attribute (attributes: later this chapter).

Arrays

An array extends a vector to any number of dimensions:

a <- array(data=1:24, dim=c(3, 4, 2))
a
, , 1

     [,1] [,2] [,3] [,4]
[1,]    1    4    7   10
[2,]    2    5    8   11
[3,]    3    6    9   12

, , 2

     [,1] [,2] [,3] [,4]
[1,]   13   16   19   22
[2,]   14   17   20   23
[3,]   15   18   21   24

Like matrices, arrays are stored as plain vectors underneath, and have no explicit class attribute.

Factors: Categorical Values

Categorical data could live in a character vector — but repeating long strings is wasteful. A factor is an ordered collection of items whose possible values are called levels:

eye.colors <- factor(c("brown", "blue", "blue", "green",
                       "brown", "brown", "brown"))
levels(eye.colors)
[1] "blue"  "brown" "green"
eye.colors
[1] brown blue  blue  green brown brown brown
Levels: blue brown green

Note the print format: no quotes, and the levels listed explicitly — visibly not a character vector.

Ordered Factors

Sometimes level order matters. A survey asks for reactions to “melon is delicious with an omelet”: Strongly Disagree … Strongly Agree. Coding these 1–5 implies equal spacing and meaningful averages — is Disagree + Agree really = Neutral? An ordered factor captures the ranking without the arithmetic claims:

survey.results <- factor(
  c("Disagree", "Neutral", "Strongly Disagree",
    "Neutral", "Agree", "Strongly Agree",
    "Disagree", "Strongly Agree", "Neutral",
    "Strongly Disagree", "Neutral", "Agree"),
  levels=c("Strongly Disagree", "Disagree", "Neutral",
           "Agree", "Strongly Agree"),
  ordered=TRUE)
survey.results
 [1] Disagree          Neutral           Strongly Disagree Neutral          
 [5] Agree             Strongly Agree    Disagree          Strongly Agree   
 [9] Neutral           Strongly Disagree Neutral           Agree            
5 Levels: Strongly Disagree < Disagree < Neutral < ... < Strongly Agree

Factors Under the Hood

Internally a factor is an integer vector plus a levels attribute mapping integers to labels (integers are small and fixed-size — efficient). Strip the class and the implementation shows:

eye.colors
[1] brown blue  blue  green brown brown brown
Levels: blue brown green
class(eye.colors)
[1] "factor"
eye.colors.integer.vector <- unclass(eye.colors)
eye.colors.integer.vector
[1] 2 1 1 3 2 2 2
attr(,"levels")
[1] "blue"  "brown" "green"
class(eye.colors.integer.vector)
[1] "integer"

Restore the class attribute, and it’s a factor again:

class(eye.colors.integer.vector) <- "factor"
eye.colors.integer.vector
[1] brown blue  blue  green brown brown brown
Levels: blue brown green

Data Frames

A data frame represents a table: scientific observations with several measurements each, or a database table of rows and typed columns. Each column (“variable”) may have a different type, but all must have the same length:

data.frame(a=c(1, 2, 3, 4, 5), b=c(1, 2, 3, 4))
Error in `data.frame()`:
! 引數值意味着不同的列數: 5, 4

The book’s example — where users search most for the word “bacon” (Google Insights, 2004–2009):

top.bacon.searching.cities <- data.frame(
  city = c("Seattle", "Washington", "Chicago", "New York", "Portland",
           "St Louis", "Denver", "Boston", "Minneapolis", "Austin",
           "Philadelphia", "San Francisco", "Atlanta", "Los Angeles",
           "Richardson"),
  rank = c(100, 96, 94, 93, 93, 92, 90, 90, 89, 87, 85, 84, 82, 80, 80)
)
head(top.bacon.searching.cities, 8)
        city rank
1    Seattle  100
2 Washington   96
3    Chicago   94
4   New York   93
5   Portland   93
6   St Louis   92
7     Denver   90
8     Boston   90

Data Frames Are Lists

typeof(top.bacon.searching.cities)
[1] "list"
class(top.bacon.searching.cities)
[1] "data.frame"

A data frame is implemented as a list with class data.frame — so list methods work unchanged: top.bacon.searching.cities$rank extracts the rank column.

Formulas

To express a relationship between variables — for a plot or a model — R provides the formula class:

sample.formula <- as.formula(y ~ x1 + x2 + x3)
class(sample.formula)
[1] "formula"
typeof(sample.formula)
[1] "language"

This reads “y is a function of x1, x2, and x3.” Chapter 3’s lattice example, Amount ~ Year | Food, reads “Amount as a function of Year, conditioned on Food.”

Formula Vocabulary

Item Meaning
variable names the variables themselves
~ separates response (left) from stimulus variables (right)
+ a linear relationship between variables
0 added to a formula: omit the intercept, e.g. y~u+w+v+0
| conditioning variables (lattice formulas, Chapter 14)
I() interpret the enclosed expression arithmetically: a+b means both variables; I(a+b) means their sum
* interactions: y~(u+v)*w ≡ y~u+v+w+u:w+v:w
^ crossing to a degree: y~(u+w)^2 ≡ y~(u+w)*(u+w)
f(x) a function of variables as a term, e.g. y~log(u)+sin(v)+w

Some functions add their own vocabulary (e.g., s() for smoothing splines in gam). Formulas return in Chapters 14 and 20.

Time Series

Many problems track a variable over time; class ts represents this, feeding regression functions (ar, arima) and specialized plot methods.

ts(data = NA, start = 1, end = numeric(0), frequency = 1,
   deltat = 1, ts.eps = getOption("ts.eps"), class = , names = )
Argument Description
data the observations (vector or matrix)
start / end time of first/last observation: one number, or (unit, offset)
frequency observations per unit of time
deltat sampling fraction between observations (frequency = 1/deltat)
ts.eps comparison tolerance for frequencies
class "ts" for one series; c("mts","ts") for several
names names of each series in a multiseries object

Eight quarters starting Q2 2008 — the print method knows about quarters and months:

ts(1:8, start=c(2008, 2), frequency=4)
     Qtr1 Qtr2 Qtr3 Qtr4
2008         1    2    3
2009    4    5    6    7
2010    8               

A Real Time Series: Turkey Prices

The USDA tracks retail meat prices (supermarkets covering ~20% of the US market, averaged by month). The book packages monthly turkey prices as turkey.price.ts:

temp <- tempfile(fileext = ".rda")
download.file("https://raw.githubusercontent.com/cran/nutshell/master/data/turkey.price.ts.rda",
              temp, mode = "wb")
load(temp)
turkey.price.ts
      Jan  Feb  Mar  Apr  May  Jun  Jul  Aug  Sep  Oct  Nov  Dec
2001 1.58 1.75 1.63 1.45 1.56 2.07 1.81 1.74 1.54 1.45 0.57 1.15
2002 1.50 1.66 1.34 1.67 1.81 1.60 1.70 1.87 1.47 1.59 0.74 0.82
2003 1.43 1.77 1.47 1.38 1.66 1.66 1.61 1.74 1.62 1.39 0.70 1.07
2004 1.48 1.48 1.50 1.27 1.56 1.61 1.55 1.69 1.49 1.32 0.53 1.03
2005 1.62 1.63 1.40 1.73 1.73 1.80 1.92 1.77 1.71 1.53 0.67 1.09
2006 1.71 1.90 1.68 1.46 1.86 1.85 1.88 1.86 1.62 1.45 0.67 1.18
2007 1.68 1.74 1.70 1.49 1.81 1.96 1.97 1.91 1.89 1.65 0.70 1.17
2008 1.76 1.78 1.53 1.90                                        

Utility functions read off the structure:

start(turkey.price.ts)
[1] 2001    1
end(turkey.price.ts)
[1] 2008    4
frequency(turkey.price.ts)
[1] 12
deltat(turkey.price.ts)
[1] 0.08333333

Shingles

A shingle generalizes a factor to a continuous variable: a numeric vector plus a set of intervals which — like roof shingles — may overlap. They let a continuous variable serve as a conditioning or grouping variable, and are used extensively in the lattice package (Chapter 14).

Dates and Times

Three classes represent moments in time:

  • Date — dates without times;
  • POSIXct — date-times as seconds since 1970-01-01 00:00;
  • POSIXlt — date-times as separate vectors: sec (0–61, allowing leap seconds!), min, hour, mday, mon (0–11), year (since 1900), wday, yday, isdst.

Store dates as date objects, not strings or numbers — the classes support arithmetic, and many plot functions require them:

date.I.started.writing <- as.Date("2/13/2009", "%m/%d/%Y")
date.I.started.writing
[1] "2009-02-13"
today <- Sys.Date()
today - date.I.started.writing
Time difference of 6380 days

Connections

Connections move data between R and the outside world — like file pointers in C or filehandles in Perl. Targets include files, URLs, zip/gzip/bzip-compressed files, pipes, network sockets, FIFOs, even the system clipboard.

The lifecycle: create → open → use → close. Loading a saved (gzip-compressed) data file explicitly:

consumption.connection <- gzfile(description="consumption.RData", open="r")
load(consumption.connection)
close(consumption.connection)

Usually you skip all this: save, load, read.table open connections implicitly when handed a filename or URL. Explicit connections earn their keep with nonstandard sources. See ?connection.

Attributes

Properties Attached to Objects

Objects carry named properties — attributes — explaining what the object represents and how to interpret it. Often, two similar objects differ only in their attributes. (Why attributes rather than lists or S4? History: attributes predate R’s modern object systems.)

Attribute Description
class the class of the object
comment a comment, often a description
dim dimensions
dimnames names along each dimension
names names of elements (columns of a data frame, etc.)
row.names names of rows (related to dimnames)
tsp start time (time series)
levels levels of a factor

For attribute a of object x: query with a(x), set with a(x) <- value. (Changes affect the current environment only.)

Inspecting Attributes

All attributes at once via attributes; individually via accessor functions:

m <- matrix(data=1:12, nrow=4, ncol=3,
            dimnames=list(c("r1", "r2", "r3", "r4"),
                          c("c1", "c2", "c3")))
attributes(m)
$dim
[1] 4 3

$dimnames
$dimnames[[1]]
[1] "r1" "r2" "r3" "r4"

$dimnames[[2]]
[1] "c1" "c2" "c3"
dim(m)
[1] 4 3
colnames(m)
[1] "c1" "c2" "c3"
rownames(m)
[1] "r1" "r2" "r3" "r4"

Attributes Define the Object

Change the attributes, change the class — remove dim and a matrix collapses to a vector:

dim(m) <- NULL
m
 [1]  1  2  3  4  5  6  7  8  9 10 11 12
class(m)
[1] "integer"
typeof(m)
[1] "integer"

Same Data, Different Attributes

An array and a vector with identical contents:

a <- array(1:12, dim=c(3, 4))
b <- 1:12
a[2, 2]
[1] 5
b[2, 2]      # error: b has no dimensions
Error in `b[2, 2]`:
! 維度數目不正確

Are they “the same”? == compares cell-wise (and answers in the shape of a):

a == b
     [,1] [,2] [,3] [,4]
[1,] TRUE TRUE TRUE TRUE
[2,] TRUE TRUE TRUE TRUE
[3,] TRUE TRUE TRUE TRUE

all.equal and identical

all.equal compares data and attributes, explaining any differences; identical answers strictly yes or no:

all.equal(a, b)
[1] "Attributes: < Modes: list, NULL >"                   
[2] "Attributes: < Lengths: 1, 0 >"                       
[3] "Attributes: < names for target but not for current >"
[4] "Attributes: < current is not list-like >"            
[5] "target is matrix, current is numeric"                
identical(a, b)
[1] FALSE

Give b the missing attribute and the difference evaporates:

dim(b) <- c(3, 4)
b[2, 2]
[1] 5
all.equal(a, b)
[1] TRUE
identical(a, b)
[1] TRUE

Class Is an Attribute

For simple objects, class and type align; for compound objects they diverge. Query with class and typeof (for matrices/arrays the class is implicit):

x <- c(1, 2, 3)
typeof(x)
[1] "double"
class(x)
[1] "numeric"

Because class is just an attribute, you can build a factor by hand from an integer vector:

v <- as.integer(c(1, 1, 1, 2, 1, 2, 2, 3, 1))
levels(v) <- c("what", "who", "why")
class(v) <- "factor"
v
[1] what what what who  what who  who  why  what
Levels: what who why

(No guarantee the internal implementation of factors never changes — treat this trick as a demonstration, not a technique.)

Quoting and the Limits of Inspection

To ask about a symbol itself rather than what it points to, quote it:

class(quote(x))
[1] "name"
typeof(quote(x))
[1] "symbol"

Some types resist inspection entirely: there is no way to isolate an any, ..., char, or promise object — checking a promise’s type would force its evaluation, converting it into an ordinary object.