The median is not just a middle number, and most people explain it wrong.
You take a dataset, sort it from smallest to largest, and pick the value that splits it in half. Half the observations fall below it, half above. If there is an even number of observations, you average the two middle values. That is the procedure. The reason it exists is entirely separate from that procedure, and confusing the two is where most students and practitioners get sloppy. The Meaning Of Median In Math is best understood through its behavior under stress. The mean is fragile. One extreme value drags it in whatever direction that value lies. The median refuses to move much. That property—the robustness to outliers—is why the median dominates income reporting, housing market summaries, and any domain where a few massive values would otherwise distort the picture. The median does not care about magnitude beyond ordering. It only cares about rank position. Calculating it is trivial for raw data. Sort, locate the middle, done. The real work starts when you leave the textbook behind. I once built a compensation benchmarking model for a firm that used salary self-reports from job seekers. The raw mean was completely useless because the top 2 percent of respondents reported six-figure salaries while the bulk of the sample was clustered between 35 and 60 thousand. The mean was lying to the client. The median sat near 51 thousand and reflected what most actual applicants earned. I stopped reporting the mean for that dataset entirely after that.
Here is the mechanical part people skip and regret later. For an odd count, the median lands at position (n + 1) / 2 in the sorted list. For an even count, you average the values at positions n / 2 and n / 2 + 1. That is it. Nothing more complicated. But once you move to grouped frequency data, the mechanics change enough that you need a real method instead of guesswork.
Grouped data requires interpolation, not eyeballing
When your data is presented as class intervals with frequencies, you cannot just look at the midpoint of a bin and call it the median. You have to locate the median class first by finding where the cumulative frequency reaches half the total, then interpolate within that class. The formula uses the lower boundary of the median class, the cumulative frequency before it, the frequency inside it, and the class width. Apply it consistently and you will get a defensible estimate. Skip it and your answer will drift enough to matter in any regression or comparison. I deal with this constantly when working with survey data binned into age brackets or income ranges. There is no way around it. You either interpolate properly or you accept that your central tendency estimate is loose. Loose is fine for internal sense-checking. It is not fine when you are publishing results or feeding them into a model that will amplify the error.
Get the Full Details

What the median hides and what it does not
The median discards information about the distance between values. Two datasets can share an identical median while looking completely different. One could be tightly clustered and the other wildly dispersed. If you only report the median, you are hiding the shape of the distribution. I always pair it with the interquartile range or at least the 10th and 90th percentiles when I present it to anyone who will read past the headline number. Another thing people miss. The median minimizes the sum of absolute deviations from the center point. The mean minimizes the sum of squared deviations. That mathematical distinction explains why the median resists outliers more than the mean does. Squaring penalties large deviations into the stratosphere. Absolute deviations treat them linearly. The difference is not subtle. It is the entire reason these two measures diverge in skewed distributions.
When the median is the wrong tool
The median fails when your analysis requires mathematical tractability. It does not play nicely with algebra. You cannot decompose it across subgroups the way you can with the mean. Weighted combinations of medians do not equal the median of the combined data. That limitation matters if you are building predictive models or aggregating metrics across segments. The median also loses interpretability with very small samples. With n = 3, the median is just the middle value. Saying it represents central tendency is technically correct and practically meaningless. I worked on a project once where the client insisted on comparing median response times across five support teams with sample sizes ranging from eight to twenty tickets per team. The variance between teams was enormous and the samples were too small for the median to stabilize. The medians looked different but the differences were noise. I pushed for a bootstrap confidence interval around each median instead. It took longer to compute but the conclusion was honest. Without it, we would have made decisions based on apparent differences that did not exist. Discrete data with heavy ties also flattens the median. When half your observations share the same value, the median is that value and tells you almost nothing about the rest of the distribution. In those cases, reporting the mode alongside the median and adding a percentile breakdown gives you actual information instead of a placeholder number.
A practical workflow I use now
I sort the data first. Always. I check the count and determine odd or even. I calculate the median position. For grouped data, I build the cumulative frequency table before touching the interpolation formula. I then compute the IQR and identify any extreme outliers separately rather than letting them sit unexamined next to the median. When presenting results, I report the median with the IQR and sample size together. That gives readers enough signal to assess reliability without requiring a full distribution plot. The median is a rank-based summary statistic. It is simple to compute and brutally honest about where the center of your data sits when the mean would be compromised. It is not a universal answer. It does not replace distributional analysis, and it should not be reported in isolation. Use it when robustness matters. Do not use it when you need analytical convenience or subgroup aggregation. That distinction is where experience actually shows up.