1.5 Web Scraping
1.5 Web Scraping
Learning objectives
By the end of this chapter, you can:
- model HTML through tags, classes, and IDs
- extract content using four rvest tools:
read_html(),html_elements(),html_text2(), andhtml_table() - locate and verify selectors using browser developer tools
- apply the polite sequence
bow()→scrape()→nod(), respecting robots.txt and request limits - judge legal and ethical boundaries: terms of service, personal data, copyright, and request frequency
- pivot to API/JSON endpoints when JavaScript renders a page
Prerequisite check (≤5 minutes)
Answer these two questions independently. If the first is unfamiliar, review list navigation in 1.4 §2.
- Given
x <- list(a = list(b = c("p", "q"))), how do you extract"q"? - What do
.price,#main, andli amatch? It is fine to guess: §1 explains them.
2. Four rvest tools: read, find, extract text, extract tables
library(rvest)
page <- read_html(html) # 1. Read a string, local file, or URL
page |> html_elements("h2") # 2. Find all matches
page |> html_elements(".price") |>
html_text2() # 3. Extract text: c("28", "22")
page |> html_element("a") |>
html_attr("href") # Extract an attribute: "menu-full.html"The fourth tool handles tables: much of the data you want is already inside a <table> element.
page2 <- read_html('
<table>
<tr><th>Dish</th><th>Price</th></tr>
<tr><td>Kung pao chicken</td><td>28</td></tr>
<tr><td>Mapo tofu</td><td>22</td></tr>
</table>')
page2 |> html_element("table") |> html_table()
# A 2-row, 2-column tibble: dish and priceThe plural version returns all matches, possibly none. The singular version takes the first. Use the singular form with html_table() for one table; use the plural form for text from many elements. A typical symptom of choosing incorrectly is one row or zero results when the page visibly contains dozens.
html_text2() cleans up whitespace such as indentation and line breaks, making it preferable for readable page text. html_text() retains the raw text and can help with debugging.
For the html string in §1, write the selector first, then verify it: ① extract both dish names without their prices; ② extract the text of every li inside the div; ③ turn both prices into the numeric vector c(28, 22) (hint: as.numeric()).
3. Find selectors with browser developer tools
Real pages are much messier than teaching strings. Use this workflow:
- Open the target page, right-click the content, and choose Inspect.
- Right-click the highlighted element, then Copy → Copy selector, but do not use it uncritically.
- Verify it in R with
read_html(url) |> html_elements(sel) |> html_text2(). Check both the number of results and their contents.
Copy selector often gives a positional path such as #root > div:nth-child(3) > div > span:nth-child(2). Moving a single div can break it. Prefer semantic selectors: the element’s own class or ID, or a nearby stable container. If you see a chain of :nth-child selectors, look again for a meaningful class name.
4. The polite sequence: bow() → scrape() → nod()
For a real website, replace direct read_html(url) calls with the three-step polite workflow. polite checks robots.txt, manages request delays, and maintains a session.
library(polite)
library(rvest)
session <- bow(
"https://rvest.tidyverse.org/", # First, bow to the host
user_agent = "posit-conf-2026 course exercise (contact: [email protected])"
)
session # Print robots.txt permissions and delay in seconds
page <- scrape(session) # Retrieve the page after the required delay
page |> html_elements("a") |> html_attr("href") |> head()
session <- nod(session, "reference/") # Change path on the same host with nod()
scrape(session) |> html_elements("a") |> length()Use bow() once per host; call it again for a new host. scrape() waits as required before making a request. nod() changes the path while retaining the session and rate limits. Fetch multiple pages sequentially, not concurrently.
Each semester I tell students that 90% of “I need a scraper” problems hide an undiscovered API or download button. Before writing a scraper, spend ten minutes looking for Data, Download, or API links. Scraping should be a last resort, not the first reaction. That advice can save a week of debugging.
5. Law and ethics: four questions before scraping
| Question | What to assess |
|---|---|
| Does robots.txt allow access? | Etiquette: polite checks it, but permission here is not permission to do anything you wish |
| Do the terms of service prohibit scraping? | Contractual restrictions: some sites explicitly prohibit it; violations may lead to account suspension or other consequences |
| Does the data concern people? | Personal information is protected by laws such as GDPR and PIPL; a lawful basis is needed |
| Is the content copyrighted? | Facts can be organized; republishing passages or images requires attention to permission |
One more rule about frequency: keep requests at a pace comparable to a person browsing manually. The server and its bandwidth bill belong to someone else.
For research, record source URLs and access dates, collect only what is needed, cite the source in outputs, and obtain appropriate ethics review for data about people, including anonymization where needed.
6. When scraping fails: JavaScript and finding the API first
The symptom: content is visible in the browser but absent from HTML retrieved with read_html(). The page may be a shell whose data is rendered by JavaScript in the browser.
To diagnose this, press Ctrl+U to inspect page source: the original content sent by the server. If the data is absent there, ordinary rvest HTML retrieval will not see it.
Work through these options in order:
- Find the data endpoint: open developer tools → Network → XHR/Fetch, refresh, and find a response containing JSON. This is the endpoint the website uses. Copy its URL and parameters and return to the httr2 workflow in Chapter 1.4. JSON is often much cleaner than HTML.
- Find an official export: a download page, official API, or open dataset takes priority over scraping.
- Last resort: a headless browser, such as real Chrome controlled by chromote. This is slower, more fragile, and more expensive to maintain. This course only introduces the option.
First check whether the content exists in the original HTML, using page source. Only then question the selector. Otherwise, you can spend an entire evening debugging the wrong thing.
Use the html string in §1 to create a two-column tibble of dish names and prices. Hint: use html_elements(".dish"), then html_element(".price") within each element; rvest selects the first match per node in a node set. Alternatively, extract two vectors and combine them with tibble(). Dish names must contain neither price digits nor extra whitespace.
Write polite_text(url, selector, delay = 5) using bow() and scrape() internally, returning matching text as a character vector. Test two selectors, such as "h1" and "a", on https://rvest.tidyverse.org/. In comments, record the robots.txt conclusion and request delay printed by the session.
Round 1 (AI off): choose a public data page, preferably a statistical Wikipedia article with a <table>. Follow the complete polite workflow: check robots.txt → bow() → scrape() → html_table() → clean into a tidy tibble. Record source and access date. Round 2 (AI allowed): show your code to Posit Assistant and ask only: “If the site is redesigned tomorrow, what breaks first, and how can that failure be more useful?” Revise it and record which suggestions you accepted or rejected, and why.
Capstone
Task: become a data scout. Find a source in your field that requires scraping; first establish that it has no API or download option, which is part of the assignment. Deliver: ① a compliance record with the robots.txt conclusion, relevant terms-of-service excerpts, and request-frequency plan; ② a polite session extracting at least one table and one list, cleaned into tidy tibbles; ③ a one-page Quarto report with data, sources, access dates, and an ethics statement.
| Dimension | Meets expectations | Good | Excellent |
|---|---|---|---|
| Responsible access | Records robots.txt | Includes terms and frequency plan | Proactively minimizes collection |
| Extraction | Table agrees with the page | Clean tidy tibble | Handles missing cells and footnotes |
| Robustness | Runs successfully | Validates key selectors | Failure messages support concrete action |
| Ethics reporting | Cites sources | Includes access dates | Explains personal-data and copyright considerations |
SOURCES · Source mapping
| Section | Material | Use |
|---|---|---|
| §2–§3: rvest functions and selectors | Official rvest documentation | Referenced |
| §4: polite workflow | Official polite documentation | Referenced |
| Miniature page, selector practice, ethics framework, exercises, capstone, and rubric | This project | Original |
This chapter is published under CC-BY-SA 4.0.