Datasets¶
The course datasets live in ALL CSV FILES - 2nd Edition/ and are distributed
by the textbook authors at statlearning.com for
use with the book.
How the notebooks resolve data
The labs load datasets straight from the ISLP package wherever possible;
the four the package does not ship (Advertising, Heart, Income1,
Income2) stream from the book’s official site, and the bundled CSVs act as an
offline fallback. You should never have to download anything by hand.
What’s bundled¶
File |
Rows |
Cols |
Contents |
Mainly used in |
|---|---|---|---|---|
|
200 |
5 |
Sales against TV / radio / newspaper budgets |
Ch 2–3 |
|
397 |
9 |
Fuel economy and specs of 1970s–80s cars |
Ch 3, 5, 8 |
|
8 645 |
16 |
Hourly bike rentals in Washington DC (counts) |
Ch 4 (Poisson) |
|
506 |
14 |
Housing values in Boston suburbs |
Ch 3, 8 |
|
88 |
9 |
Survival times of brain-cancer patients |
Ch 11 |
|
5 822 |
86 |
Caravan-insurance purchases (highly imbalanced) |
Ch 4 |
|
400 |
11 |
Child car-seat sales at 400 stores |
Ch 3, 8 |
|
999 |
40 |
Gene-expression data for the Ch 12 exercise |
Ch 12 |
|
777 |
19 |
Statistics for 777 US colleges |
Ch 6 |
|
400 |
11 |
Credit-card balances and customer attributes |
Ch 3, 6 |
|
10 000 |
4 |
Credit-card default (the classification workhorse) |
Ch 4 |
|
50 |
2 001 |
Monthly excess returns of 2 000 fund managers |
Ch 13 |
|
303 |
15 |
Heart-disease diagnosis |
Ch 8 |
|
322 |
20 |
Baseball salaries and career statistics |
Ch 6 |
|
30 |
3 |
Income vs. years of education (simulated) |
Ch 2 |
|
30 |
4 |
Income vs. education and seniority (simulated) |
Ch 2 |
|
1 070 |
18 |
Orange-juice brand purchases |
Ch 8–9 |
|
100 |
2 |
Two asset returns — the bootstrap example |
Ch 5 |
|
244 |
10 |
Time to publication of clinical trials |
Ch 11 |
|
1 250 |
9 |
Daily S&P 500 returns, 2001–2005 |
Ch 4 |
|
3 000 |
11 |
Wages of Mid-Atlantic male workers |
Ch 1, 7 |
|
1 089 |
9 |
Weekly S&P 500 returns, 1990–2010 |
Ch 4 |
Auto.data is the original whitespace-delimited version of Auto.csv, kept for
the exercise that demonstrates reading a non-CSV file.
Loading a dataset¶
from ISLP import load_data
Boston = load_data("Boston")
Boston.head()
import pandas as pd
from pathlib import Path
DATA = Path("ALL CSV FILES - 2nd Edition") # relative to the repo root
Boston = pd.read_csv(DATA / "Boston.csv", index_col=0)
Several files carry an unnamed index column (Boston, College, Heart,
Income1, Income2, Bikeshare, BrainCancer, Publication, Fund) —
pass index_col=0 for those.
import pandas as pd
url = "https://www.statlearning.com/s/Advertising.csv"
Advertising = pd.read_csv(url, index_col=0)
Missing values
Auto contains a handful of rows with ? in horsepower; ISLP’s load_data
drops them, so a CSV fallback should do the same
(pd.read_csv(..., na_values="?").dropna()) for results to match the slides.
Attribution¶
The datasets are © the textbook authors and are distributed by them at statlearning.com for use with the book. They are bundled here only so the course runs offline — see Citation & licence.
Where to go next¶
Lab notebooks — where each dataset is loaded and used.
Python environment — how
ISLPresolves the data.Lecture slides — the figures computed from these files.