Chapter 26: R and Hadoop
Shih Chien University
2026-07-20
Some problems are too big for one machine. The fix: break the work into pieces, solve them in parallel, recombine — the laundry analogy: ten washers finish ten loads in the time one washer does one. Many statistical methods are naturally parallel — bagging, boosting, and random forests (Chapters 20–21) fit each tree independently.
In the book’s era, Hadoop was the de facto standard for big data: store enormous data sets, run computations across a cluster, tolerate machine failures. This chapter shows how to drive Hadoop from R.
Warning
A snapshot of 2012. Hadoop’s dominance has faded — today’s big-data work leans on Spark (R interface sparklyr), cloud data warehouses, and arrow/duckdb for larger-than-memory data on one machine. The concepts below — Map/Reduce, distributed storage — remain foundational, and several still power modern systems.
Hadoop’s programming model splits every job into two steps:
(location, 1).)Because every map runs independently, the work spreads across as many machines as you have. Classic examples from the book: a web-traffic report, traffic by location, and predicting user behavior — all expressible as map then reduce.
Beyond Map/Reduce, Hadoop provides:
The framework itself is written in Java, but you need not write Java — R can drive it.
The book describes three ways to use Hadoop from R:
lapply-style jobs on Amazon Elastic MapReduce;RHadoop is three packages:
rmr — the core: write Map/Reduce jobs in R, move data in and out;rhdfs — manage files on HDFS from R (most ordinary file operations);rhbase — interface to HBase, Hadoop’s database.Setup, then a job (rmr syntax, shown unevaluated — it needs a Hadoop cluster):
The book’s four-node example tallies US deaths by sex for 2009 — a job that runs in parallel across the cluster. to.dfs/from.dfs move data between R and HDFS.
You needn’t own a cluster — rent one. The book walks through Amazon Elastic MapReduce (EMR): bootstrap a cluster, configure it from a template, run the job, tear it down — paying only for the minutes used. Segue wraps this so an R user can run a parallel lapply on EMR with almost no Hadoop knowledge.
Tip
The enduring lessons. Hadoop the product is fading, but its ideas are everywhere: (1) partition data and computation; (2) express work as map (independent, parallel) then reduce (combine); (3) move compute to the data, not the reverse; (4) rent elastic capacity in the cloud. Modern R reaches big data through sparklyr, arrow, duckdb, and future/furrr for parallelism — but you’ll recognize Map/Reduce inside all of them.
This chapter closes R in a Nutshell — from installing R (Chapter 1) to running it across a cluster. The throughline: the same language, the same objects, the same functions, scaling from one console to thousands of machines.
Copyright. These slides are adapted from R in a Nutshell: A Desktop Quick Reference (2nd ed.) by Joseph Adler, O’Reilly Media. All rights reserved by the original author and publisher.
Non-commercial use only. These materials are strictly for educational purposes and may not be used for commercial gain.
Attribution. Any reproduction, distribution, or use of these materials must properly credit the original source.
R in a Nutshell: A Desktop Quick Reference