Getting Through Chapter 1 Without Losing Your Mind
Chapter 1 of The Practice Of Statistics 5th Edition is called "Exploring Data," and it is where the entire course gets its vocabulary. You learn the difference between a statistic and a parameter. You learn what categorical and quantitative variables are. You learn to draw dotplots, stemplots, histograms, and modified boxplots. That sounds straightforward until you actually sit down with the problem set and realize half the points are being taken off for labeling choices you didn't even know existed. I took this class and then TA'd it for two semesters. The grading patterns are consistent enough that I can predict where students will bleed points before they even make the mistake.
The Practice Of Statistics 5th Edition Chapter 1
Here is the core of what the chapter covers and how it actually plays out in homework and exams. You need to correctly classify every variable you encounter. Categorical puts data into groups. Quantitative takes numerical operations. That is the simple version. The hard version shows up when a problem gives you a variable like "ZIP code" or "number of siblings in a household." ZIP code looks numeric but is categorical. It is a label. Doing math on it is meaningless. Number of siblings is quantitative because adding and averaging it actually tells you something useful. I remember a student who lost three points on a midterm for calling "student ID number" quantitative. It is categorical. The numbers are assigned codes, not measurements. Once you understand that distinction, you stop second-guessing these questions.
Describing Distributions
When you describe a distribution, you need to cover shape, center, spread, and unusual features. Shape means symmetric, skewed left, skewed right, or roughly symmetric. Center means median or mean depending on the situation. Spread means range, IQR, or standard deviation. Unusual features means outliers, gaps, or clusters. If you leave any of those four out, your answer is incomplete and the rubric will reflect that. Skewed right distributions pull the mean above the median. Skewed left pulls the mean below the median. Symmetric distributions keep them close. This matters when you choose between mean and median as the measure of center. Use median for skewed data or data with outliers. Use mean for roughly symmetric data without extreme outliers. You will lose points if you pick the wrong one and do not explain why.
Get the Full Details

Graphs and When to Use Each One
Dotplots work for small datasets where you want to see individual values. Stemplots are useful when you have moderate data and want to preserve the actual numbers while seeing the shape. Histograms handle larger datasets and show distribution shape clearly. You choose the graph based on the size of the dataset and what information you need to convey. A stemplot with three-digit numbers gets unwieldy fast. Do not use one for that. A histogram with too few bins hides the shape. Too many bins creates noise. Five to twenty bins is usually reasonable, but you should adjust based on the range and count of your data. Modified boxplots are the default in this textbook. They show outliers explicitly as separate points. A standard boxplot folds outliers into the whiskers, which distorts the scale. Always check whether the question asks for a modified or standard boxplot before you draw. They look different and are graded differently.
Five-Number Summary and Boxplots
You need to find the minimum, Q1, median, Q3, and maximum. That is the five-number summary. From there you build the boxplot. The interquartile range is Q3 minus Q1. Any value below Q1 minus 1.5 times the IQR or above Q3 plus 1.5 times the IQR is flagged as an outlier in this textbook's convention. Values beyond 3 times the IQR are extreme outliers. You mark outliers as individual points past the whiskers. I ran into a case once where a dataset had a cluster of values at exactly the outlier boundary. The calculator gave the same IQR whether you included the boundary point or not, but whether you classified it as an outlier depended on the strictness of the inequality. The textbook uses strict inequalities, so a value exactly at Q1 minus 1.5 times IQR is not an outlier. Make sure you know which convention your instructor expects. Ask directly. I have seen two professors in the same department grade this differently on the same exam.
Back-to-Back Stemplots and Comparing Distributions
When you compare two distributions, back-to-back stemplots are efficient. You share the stem and mirror the leaves. Write out the key. Without a key, a grader cannot verify your work. Compare shape, center, and spread explicitly. Do not just say "one is higher." Say "the median of group A is 42 compared to 35 for group B." Specific numbers earn points. Vague language does not. Sometimes you are given raw data and asked to build a frequency table. Count the occurrences in each class interval. Relative frequency is the count divided by the total. Cumulative relative frequency adds each relative frequency to the sum of the previous ones. These tables feed directly into histograms and ogives. Getting the class width wrong throws off every subsequent calculation. Decide on the number of classes first, then divide the range by that number and round up to a convenient value. Do not round down unless the data range divides perfectly. Reading scale on graphs is where most mistakes happen. If the y-axis does not start at zero, the visual comparison is distorted. A histogram bar height represents frequency or relative frequency, not density, unless the bin widths are unequal. If bin widths are unequal, you need relative frequency divided by bin width for the bar height to be meaningful. This comes up occasionally in AP-level courses and on exams. The Practice Of Statistics 5th Edition Chapter 1 introduces histograms with equal bins, but later chapters and cumulative review problems mix things up.

Another issue is misreading stemplot leaves. If the leaf unit is 0.1 and you write "3 | 5", the value is 3.5, not 35. Label the leaf unit every time. It takes three seconds and prevents confusion.
What This Chapter Does Not Cover Well
It does not give you much on calculating standard deviation by hand efficiently. You will learn the formula, but the arithmetic is tedious. Know how to use your calculator's stat functions. On TI-84, enter data into L1, run 1-Var Stats, and read out the values. It cuts calculation time from fifteen minutes to under a minute. On exams where calculators are allowed, this is essential. Where they are not, practice the formula by hand until it stops being painful. There is also almost no discussion of how sampling method affects the distributions you analyze. The chapter focuses on describing data you already have. It assumes the data came from a reasonable source. In real work, that assumption is often wrong. You might be handed a convenience sample and asked to generalize. This textbook does not address that tension in Chapter 1. It arrives later. For now, just know that descriptive statistics do not fix bad data collection.
Practical Advice for the Coursework
Do the calculator work early. Learn StatCrunch or your calculator's statistical functions in the first week. It saves hours over the semester. Practice labeling graphs until it becomes automatic. Every graph needs a title, labeled axes, and a consistent scale. Missing any of those costs points consistently. Work through the chapter exercises in order. The difficulty ramps up gradually, and skipping ahead leaves gaps in your understanding of notation that later chapters depend on. If you are using this book for self-study, the online resources and solution manuals are uneven. Some problems have full solutions. Many do not. Do not rely on finding every answer online. The practice itself is what builds competence. Focus on understanding why you choose a median over a mean, why you mark certain points as outliers, and why a particular graph type fits a particular dataset. Those decisions are what the exams test, not memorization of definitions.
