Programming for Applications

Chapter 26: R and Hadoop

Yu-You Liou (NTU)

Shih Chien University

2026-07-20

Big Data with R

Why Hadoop

Some problems are too big for one machine. The fix: break the work into pieces, solve them in parallel, recombine — the laundry analogy: ten washers finish ten loads in the time one washer does one. Many statistical methods are naturally parallel — bagging, boosting, and random forests (Chapters 20–21) fit each tree independently.

In the book’s era, Hadoop was the de facto standard for big data: store enormous data sets, run computations across a cluster, tolerate machine failures. This chapter shows how to drive Hadoop from R.

Warning

A snapshot of 2012. Hadoop’s dominance has faded — today’s big-data work leans on Spark (R interface sparklyr), cloud data warehouses, and arrow/duckdb for larger-than-memory data on one machine. The concepts below — Map/Reduce, distributed storage — remain foundational, and several still power modern systems.

How Hadoop Works

Map/Reduce

Hadoop’s programming model splits every job into two steps:

  • Map — process each record independently, emitting key–value pairs. (Parse each web-log line; emit (location, 1).)
  • Reduce — gather all values for each key and combine them. (Sum the counts per location.)

Because every map runs independently, the work spreads across as many machines as you have. Classic examples from the book: a web-traffic report, traffic by location, and predicting user behavior — all expressible as map then reduce.

The Two Pillars

Beyond Map/Reduce, Hadoop provides:

  • Distributed data storage — HDFS (Hadoop Distributed File System): splits files into blocks, replicates them across machines for fault tolerance, and presents one logical file system.
  • Cluster management — schedules jobs across servers, restarts failed tasks, moves computation to where the data lives.

The framework itself is written in Java, but you need not write Java — R can drive it.

Driving Hadoop from R

Three Routes

The book describes three ways to use Hadoop from R:

  • RHadoop — the most mature, best-integrated set of R packages;
  • Segue — runs R lapply-style jobs on Amazon Elastic MapReduce;
  • Hadoop streaming — feed any R script to Hadoop as a mapper/reducer via standard input and output.

RHadoop

RHadoop is three packages:

  • rmr — the core: write Map/Reduce jobs in R, move data in and out;
  • rhdfs — manage files on HDFS from R (most ordinary file operations);
  • rhbase — interface to HBase, Hadoop’s database.

Setup, then a job (rmr syntax, shown unevaluated — it needs a Hadoop cluster):

library(rmr2)
# count deaths by sex: a map step and a reduce step, both R functions
result <- mapreduce(
  input = "deaths.csv",
  map    = function(k, v) keyval(v$sex, 1),
  reduce = function(sex, counts) keyval(sex, sum(counts)))
from.dfs(result)                 # pull results back into R

The book’s four-node example tallies US deaths by sex for 2009 — a job that runs in parallel across the cluster. to.dfs/from.dfs move data between R and HDFS.

Hadoop in the Cloud

You needn’t own a cluster — rent one. The book walks through Amazon Elastic MapReduce (EMR): bootstrap a cluster, configure it from a template, run the job, tear it down — paying only for the minutes used. Segue wraps this so an R user can run a parallel lapply on EMR with almost no Hadoop knowledge.

# Segue-style: a parallel lapply on a rented cluster (conceptual)
library(segue)
cluster <- createCluster(numInstances=4)
results <- emrlapply(cluster, my.list, my.function)
stopCluster(cluster)

Putting It in Perspective

Tip

The enduring lessons. Hadoop the product is fading, but its ideas are everywhere: (1) partition data and computation; (2) express work as map (independent, parallel) then reduce (combine); (3) move compute to the data, not the reverse; (4) rent elastic capacity in the cloud. Modern R reaches big data through sparklyr, arrow, duckdb, and future/furrr for parallelism — but you’ll recognize Map/Reduce inside all of them.

This chapter closes R in a Nutshell — from installing R (Chapter 1) to running it across a cluster. The throughline: the same language, the same objects, the same functions, scaling from one console to thousands of machines.