Programming for Applications

Chapter 6: R Syntax

Yu-You Liou (NTU)

Shih Chien University

2026-08-03

Writing Valid R

Sugar over Function Calls

Chapter 5 showed that every R expression can be rewritten as a function call. Yet R provides special syntax — syntactic sugar — so that assignments, lookups, and arithmetic read naturally. (Without it, R code would look like LISP, all parentheses; no coincidence, since S was inspired by LISP.)

This chapter is a readable, not exhaustive, guide to writing valid R expressions:

  • constants: numbers, characters, symbols;
  • operators and their precedence;
  • assignments, grouping, control structures;
  • the indexing notation for data structures;
  • and a few words on code style.

Constants

Numeric Vectors

Numbers are interpreted literally — and hexadecimal works with the 0x prefix:

1.1
[1] 1.1
2
[1] 2
2^1023
[1] 8.988466e+307
0x1
[1] 1
0xFFFF
[1] 65535

By default every number — even a plain integer — is stored as a double-precision floating-point value. To get a true integer, use sequence notation or as:

typeof(1)
[1] "double"
typeof(1:1)
[1] "integer"
typeof(as(1, "integer"))
[1] "integer"

Precision and Size Limits

a:b yields the integers from a to b; arbitrary sets combine with c:

v <- c(173, 12, 1.12312, -93)

Flexibility has limits — doubles cannot distinguish or represent everything:

(2^1023 + 1) == 2^1023   # limits of precision
[1] TRUE
2^1024                   # limits of size
[1] Inf

In practice this rarely matters: the databases your data comes from can’t represent such numbers either.

Complex Numbers

Complex values are written as real part + imaginary part i:

0+1i ^ 2
[1] -1+0i
sqrt(-1+0i)
[1] 0+1i
exp(0+1i * pi)
[1] -1+1.224647e-16i
sqrt(-1)        # NaN! sqrt returns the same type it was given
Warning in sqrt(-1): 產生了 NaNs
[1] NaN

Note the last line: handed the numeric -1, sqrt answers in numerics — and the numeric answer to √−1 is NaN, with a warning. Handed a complex number, it answers in complex.

Character Vectors

A character object is the text between a pair of quotes — double or single, your choice, or escaped with backslashes:

"hello"
[1] "hello"
'hello'
[1] "hello"
identical("\"hello\"", '"hello"')
[1] TRUE
identical('\'hello\'', "'hello'")
[1] TRUE

Single quotes are convenient when the text itself contains double quotes (and vice versa). Longer vectors, as always, come from c:

numbers <- c("one", "two", "three", "four", "five")
numbers
[1] "one"   "two"   "three" "four"  "five" 

Symbols

A symbol is a constant that refers to another object — the name of a variable. x <- 1 means “map the symbol x to the value 1 in the current environment.”

Names that start with a letter and contain letters, digits, periods, and underscores can be typed directly — and case matters:

x1 <- 1
X1 <- 2
x1
[1] 1
X1
[1] 2
x1.1 <- 1
x1.1_1 <- 1

Backquotes and Reserved Words

Symbols with special syntax are reached by enclosing them in backquotes — that is how you get help on the assignment operator (?`<-`), and, if you insist, how you create perverse names:

`1+2=3` <- "hello"
`1+2=3`
[1] "hello"

Some words are reserved and cannot be used as symbols: if, else, repeat, while, function, for, in, next, break, TRUE, FALSE, NULL, Inf, NaN, NA, NA_integer_, NA_real_, NA_complex_, NA_character_, ..., and ..1 through ..9.

Redefining Built-ins

Primitive functions not on the reserved list can be redefined — even c:

c
function (...)  .Primitive("c")
c <- 1
c
[1] 1
v <- c(1, 2, 3)   # and yet the combine function still works!
v
[1] 1 2 3

R can tell from context that c(1, 2, 3) needs a function named c, and finds the original. Clever — but shadowing built-in names is still a recipe for confusion; avoid it in real code.

Operators

Binary and Unary Operators

An operator is a function of one or two arguments that can be written without parentheses:

1 + 19      # addition
[1] 20
5 * 4       # multiplication
[1] 20
41 %% 21    # modulus
[1] 20
20 ^ 1      # exponent
[1] 20
21 %/% 2    # integer division
[1] 10

Unary operators take a single operand — negation (-7) and help (?) are familiar examples.

Defining Your Own Operator

A user-defined binary operator is any function of two variables assigned to a name of the form %text%:

`%myop%` <- function(a, b) { 2*a + 2*b }
1 %myop% 1
[1] 4
1 %myop% 2
[1] 6

Some core language constructs are themselves binary operators: assignment (symbol on the left, value on the right), indexing (symbol left, index right), and even function calls (function left, arguments right — don’t be fooled by the closing bracket or parenthesis):

x <- c(1, 2, 3, 4, 5)
x[3]
[1] 3
max(1, 2)
[1] 2

Order of Operations

Just as in school math (1 + 2 × 5 multiplies first), R resolves ambiguity with fixed precedence rules. From highest to lowest, roughly: function calls and grouping; indexing and lookup; arithmetic; comparison; formulas; assignment; help.

Operators (by priority) Description
( { Function calls, grouping
:: ::: Namespace access
$ @ Component / slot extraction
[ [[ Indexing
^ Exponentiation (right to left)
- + Unary minus, plus
: Sequence
%any% Special operators
* / Multiply, divide
+ - Binary add, subtract
< > <= >= == != Ordering, comparison
! Negation
& && And
| || Or
~ Formulas
-> ->> Rightward assignment
<- <<- Assignment (right to left)
= Assignment (right to left)
? Help

The authoritative, current list: help(syntax).

Assignments

Plain and Property Assignments

Most assignments simply bind an object to a symbol:

x <- 1
y <- list(shoes="loafers", hat="Yankees cap", shirt="white")
z <- function(a, b, c) { a ^ b / c }
v <- c(1, 2, 3, 4, 5, 6, 7, 8)

But R also allows a function on the left-hand side — replacing an object with a modified version of itself:

dim(v) <- c(2, 4)        # v is now a 2x4 matrix
v[2, 2] <- 10
formals(z) <- alist(a=1, b=2, c=3)

Replacement Functions

The magic behind the scenes: a statement of the form fun(sym) <- val is syntactic sugar for

`fun<-`(sym, val)

which replaces the object bound to sym in the current environment. By convention fun names a property of the object. Define your own method named myproperty<-, and R lets you write myproperty(x) <- value — your own left-hand-side function.

Grouping Expressions

Semicolons and Parentheses

Expressions go on separate lines, or on one line separated by semicolons:

x <- 1; y <- 2; z <- 3

Parentheses return the result of the expression inside — precedence-wise the same as a function call. Indeed, grouping is equivalent to calling an identity function:

2 * (5 + 1)
[1] 12
f <- function(x) x
2 * f(5 + 1)        # the same thing
[1] 12
2 * 5 + 1           # versus the default order of operations
[1] 11

Curly Braces

Curly braces evaluate a series of expressions (newline- or semicolon-separated) and return only the last:

f <- function() { x <- 1; y <- 2; x + y }
f()
[1] 3
{ x <- 1; y <- 2; x + y }    # braces work outside functions too
[1] 3

Braces Do Not Create Environments

A crucial difference: a function call creates a new environment; braces do not — their contents evaluate in the current environment:

f <- function() { u <- 1; v <- 2; u + v }
f()
[1] 3
u                       # not defined out here!
Error:
! 找不到物件 'u'
{ u <- 1; v <- 2; u + v }
[1] 3
u                       # defined: braces ran in the current environment
[1] 1
v
[1] 2

Internally, brace notation is a call to the `{` function — though its arguments are not evaluated like a standard function’s. (Scope and environments: Chapter 8.)

Control Structures

Conditional Statements

The forms are if (condition) true_expression else false_expression and if (condition) expression. Because the branches are not always evaluated, if has type special:

typeof(`if`)
[1] "special"
if (FALSE) "this will not be printed"
if (FALSE) "this will not be printed" else "this will be printed"
[1] "this will be printed"
x <- 10
if (is(x, "numeric")) x/2 else print("x is not numeric")
[1] 5

Conditions Must Be Scalar; ifelse Is the Vector Version

An if condition is not a vector operation. In the book’s era a longer condition used the first element with a warning; modern R (≥ 4.2) makes it an error:

x <- 10
y <- c(8, 10, 12, 3, 17)
if (x < y) x else y
Error in `if (x < y) ...`:
! 條件的長度 > 1

For an element-wise choice, use ifelse:

a <- c("a", "a", "a", "a", "a")
b <- c("b", "b", "b", "b", "b")
ifelse(c(TRUE, FALSE, TRUE, FALSE, TRUE), a, b)
[1] "a" "b" "a" "b" "a"

switch

Returning different values for different inputs via chained else ifs is verbose:

switcheroo.if.then <- function(x) {
  if (x == "a")      "alligator"
  else if (x == "b") "bear"
  else if (x == "c") "camel"
  else               "moose"
}

switch does the same job tidily — named arguments map inputs to results; the unnamed argument is the default:

switcheroo.switch <- function(x) {
  switch(x, a="alligator", b="bear", c="camel", "moose")
}
switcheroo.if.then("a"); switcheroo.switch("a")
[1] "alligator"
[1] "alligator"
switcheroo.if.then("f"); switcheroo.switch("f")
[1] "moose"
[1] "moose"

Loops: repeat

R has three looping constructs. The simplest, repeat, just repeats an expression — forever, unless you break (and next skips to the next iteration). Multiples of 5 up to 25:

i <- 5
repeat { if (i > 25) break else { print(i); i <- i + 5 } }
[1] 5
[1] 10
[1] 15
[1] 20
[1] 25

Forget the break and you have an infinite loop — occasionally useful for interactive applications, otherwise a bug.

Loops: while

while (condition) expression repeats as long as the condition holds; break and next work here too:

i <- 5
while (i <= 25) { print(i); i <- i + 5 }
[1] 5
[1] 10
[1] 15
[1] 20
[1] 25

Loops: for

for (var in list) expression iterates over the elements of a vector or list:

for (i in seq(from=5, to=25, by=5)) print(i)
[1] 5
[1] 10
[1] 15
[1] 20
[1] 25

Two properties to remember:

for (i in seq(from=5, to=25, by=5)) i   # 1. no printing unless you print()
i                                        # 2. the loop variable leaks into the calling environment
[1] 25

Like conditionals, repeat, while, and for have type special — their body is not necessarily evaluated.

Looping Extensions: iterators and foreach

Note

Missing Java-style iterators or foreach? CRAN add-on packages supply both — and they make code easier to parallelize.

iterators: iter(obj, checkFunc=..., recycle=...) builds an iterator; nextElem fetches values (filtered by checkFunc) until it stops with “StopIteration”:

library(iterators)
onetofive <- iter(1:5)
nextElem(onetofive); nextElem(onetofive); nextElem(onetofive)
[1] 1
[1] 2
[1] 3
nextElem(onetofive); nextElem(onetofive)
[1] 4
[1] 5
nextElem(onetofive)    # exhausted
Error:
! StopIteration

foreach: builds a loop object, evaluated by the %do% (serial) or %dopar% (parallel — Chapter 26) operators:

library(foreach)
sqrts.1to5 <- foreach(i=1:5) %do% sqrt(i)
sqrts.1to5[1:3]
[[1]]
[1] 1

[[2]]
[1] 1.414214

[[3]]
[1] 1.732051

Accessing Data Structures

The Four Access Operators

Syntax Objects Description
x[i] vectors, lists Elements described by i — an integer, character (names), or logical vector. No partial matching. On a list returns a list; on a vector, a vector.
x[[i]] vectors, lists A single element matching i (integer or character of length 1). Partial name matching with exact=FALSE.
x$n lists The element named n.
x@n S4 objects The value in slot n.

Single vs. double brackets, three differences: [[ returns exactly one element while [ may return several; by name, [ matches exactly while [[ allows partial matches; on lists, [ returns a list but [[ returns the element itself.

Indexing by Integer Vector

v <- 100:119
v[5]
[1] 104
v[1:5]
[1] 100 101 102 103 104
v[c(1, 6, 11, 16)]
[1] 100 105 110 115
v[[3]]              # double brackets: single element
[1] 102

Negative integers exclude the named positions:

v[-15:-1]           # everything except elements 1..15
[1] 115 116 117 118 119

Integer Indexing of Lists

The same notation applies to lists:

l <- list(a=1, b=2, c=3, d=4, e=5, f=6, g=7, h=8, i=9, j=10)
l[1:3]
$a
[1] 1

$b
[1] 2

$c
[1] 3
l[-7:-1]
$h
[1] 8

$i
[1] 9

$j
[1] 10

Integer Indexing of Matrices

For multidimensional structures, give one index per dimension — or omit one to take everything along it:

m <- matrix(data=c(101:112), nrow=3, ncol=4)
m
     [,1] [,2] [,3] [,4]
[1,]  101  104  107  110
[2,]  102  105  108  111
[3,]  103  106  109  112
m[3]        # single index: positions in the underlying vector
[1] 103
m[3, 4]
[1] 112
m[1:2, 1:2]
     [,1] [,2]
[1,]  101  104
[2,]  102  105
m[1:2, ]
     [,1] [,2] [,3] [,4]
[1,]  101  104  107  110
[2,]  102  105  108  111
m[, 3:4]
     [,1] [,2]
[1,]  107  110
[2,]  108  111
[3,]  109  112

Dimension Dropping and drop=FALSE

Subsets are automatically coerced to the fewest sensible dimensions: a matrix-shaped subset returns a matrix, a vector-shaped one a vector. Disable with drop=FALSE:

a <- array(data=c(101:124), dim=c(2, 3, 4))
class(a[1, 1, ])
[1] "integer"
class(a[1, , ])
[1] "matrix" "array" 
class(a[1:2, 1:2, 1:2])
[1] "array"
class(a[1, 1, 1, drop=FALSE])
[1] "array"

Replacing and Extending by Index

The same notation assigns — single cells, blocks, even positions beyond the current end (gaps become NA):

m[1] <- 1000
m[1:2, 1:2] <- matrix(c(1001:1004), nrow=2, ncol=2)
m
     [,1] [,2] [,3] [,4]
[1,] 1001 1003  107  110
[2,] 1002 1004  108  111
[3,]  103  106  109  112
v <- 1:12
v[15] <- 15
v
 [1]  1  2  3  4  5  6  7  8  9 10 11 12 NA NA 15

A factor used as an index is interpreted as an integer vector.

Indexing by Logical Vector

A logical vector selects the TRUE positions — and is recycled if shorter than the target:

v <- 100:119
v[rep(c(TRUE, FALSE), 10)]
 [1] 100 102 104 106 108 110 112 114 116 118
v[(v == 103)]          # compute the mask from the vector itself
[1] 103
v[(v %% 3 == 0)]       # multiples of three
[1] 102 105 108 111 114 117
v[c(TRUE, FALSE, FALSE)]   # recycled mask
[1] 100 103 106 109 112 115 118

Lists too:

l[(l > 7)]
$h
[1] 8

$i
[1] 9

$j
[1] 10

Indexing by Name

List elements may be named; index them with $, or by name vectors in single brackets:

l <- list(a=1, b=2, c=3, d=4, e=5, f=6, g=7, h=8, i=9, j=10)
l$j
[1] 10
l[c("a", "b", "c")]
$a
[1] 1

$b
[1] 2

$c
[1] 3

Double brackets take a single name — and uniquely allow partial matching with exact=FALSE:

dairy <- list(milk="1 gallon", butter="1 pound", eggs=12)
dairy$milk
[1] "1 gallon"
dairy[["milk"]]
[1] "1 gallon"
dairy[["mil"]]                  # no exact match: NULL
NULL
dairy[["mil", exact=FALSE]]     # partial match allowed
[1] "1 gallon"

Indexing Nested Lists

For lists of lists, [[ accepts a vector and walks down the levels:

fruit <- list(apples=6, oranges=3, bananas=10)
shopping.list <- list(dairy=dairy, fruit=fruit)
shopping.list[[c("dairy", "milk")]]
[1] "1 gallon"
shopping.list[[c(1, 2)]]
[1] "1 pound"

R Code Style Standards

Readable Code Is Maintainable Code

Style isn’t syntax, but it decides whether others (and future you) can read your code. The book follows Google’s R Style Guide; the highlights:

  • Indentation: two spaces, never tabs; inside parentheses, align to the innermost.
  • Spacing: single spaces around binary operators and after commas; none between a function name and its argument list.
  • Blocks: opening brace { not on its own line; closing brace } on its own line; indent inner blocks two spaces.
  • Semicolons: omit when optional.
  • Naming: objects in lowercase words separated by periods; function names in CapWords and preferably verbs.
  • Assignment: use <-, not =.

And no — you are not obliged to name your objects field.goals or top.bacon.searching.cities. It’s just convention. (The modern community mostly follows the tidyverse style guide, which prefers snake_case — pick one style and stay consistent.)