Stem and Leaf Plots Are Underused and People Still Get Them Wrong

I remember working on a regression analysis project back in 2018 where I had to present data distributions to a group of clients who had zero statistics background. I tried a histogram first and they glazed over. A stem and leaf plot took about forty-five seconds to draw on a whiteboard and suddenly everyone understood the spread, the skew, and the outliers at a glance. That is the actual utility of this method. It is not a toy from middle school math. The construction is straightforward. You split each data point into a stem and a leaf. The stem holds all digits except the final one. The leaf is the last digit. If your dataset contains test scores like 72, 78, 79, 81, and 85, the stems are 7 and 8. The leaves under stem 7 would be 2, 8, 9 and under stem 8 you would write 1 and 5. You list the stems in ascending order down the left side and then write the leaves horizontally next to their corresponding stem, also in ascending order. The trick people miss is deciding what constitutes the stem when your data spans a wide range. I once had a dataset of monthly temperatures across a decade that ran from negative twelve to forty-one degrees Celsius. Using single digits as stems was useless because you ended up with twenty-three rows and nothing readable. I switched to tens as stems instead, which compressed the plot into five rows and made the bimodal distribution obvious within seconds.

Here is a quick example. Say your data is: 15, 17, 21, 23, 23, 28, 34, 39, 42, 42, 47. The plot looks like this: 1 | 5 7
2 | 1 3 3 8
3 | 4 9
4 | 2 2 7

You read it by combining the stem and leaf. The second row, third entry is 23. Every number in the plot maps back to an actual data point. That is the main advantage over a histogram, which bins values and loses the individual observations. With a stem and leaf you can still recover the raw data if you need to. There is a nuance with decimal data that most guides skip. If your values are 3.2, 3.7, 4.1, 4.8, and 5.0, you treat the digit before the decimal as the stem and the first digit after as the leaf. The plot becomes 3 | 2 7, 4 | 1 8, 5 | 0. You do not invent extra decimal places or try to force two-digit leaves unless the precision of your measurement actually demands it. Adding false precision makes the plot unreadable and introduces errors during interpretation. Another thing beginners consistently get wrong is sorting the leaves. I have seen plots where leaves are written in the order the data appears in the raw dataset. That defeats the entire purpose. Unsorted leaves make it impossible to quickly identify the median, range, or mode at a glance. Always sort both stems and leaves in ascending order before you consider the plot finished.

Get the Full Details

How to Read a Stem and Leaf Plot: 3 Easy Steps
How to Read a Stem and Leaf Plot: 3 Easy Steps

The plot also has a real limitation. When you exceed roughly two hundred data points, the leaves become too dense to read. The whole method collapses into a wall of numbers that is worse than just looking at a table. I hit this ceiling when trying to plot survey responses from a sample of three hundred people across twelve Likert-scale items. I switched to a frequency polygon instead and saved myself about twenty minutes of headache. If your stems produce more than eight or nine rows, you should consider splitting them. A split stem means each stem value appears twice. The first row gets leaves from zero through four and the second row gets leaves from five through nine. This stretches the distribution vertically and reveals clustering that a single-stem version would hide. I use this technique regularly when working with exam scores that bunch tightly around a particular range. A common pitfall involves datasets with very few distinct values, like a set of responses that only takes integers one through five. The plot becomes a thin vertical line with no shape. In those cases a stem and leaf plot adds nothing you could not see from a simple frequency table. Don't force the method where it doesn't fit.

When presenting these plots, always include a key. Something as simple as 3 | 4 means 34. Without that, anyone unfamiliar with how you constructed the stems will interpret the numbers differently and your entire visualization becomes misleading. I learned this the hard way during a peer review where a colleague interpreted my stems as hundreds instead of tens, which completely changed their reading of the distribution. The method is fast to produce by hand for small datasets and equally fast to generate in any spreadsheet software. Excel does not have a built-in stem and leaf function, but a quick pivot with text-to-columns separating the tens and units digits takes about three minutes. Google Sheets requires the same manual step. For anything above three hundred observations, move directly to a box plot or a density curve and skip this method entirely.