Describing your data before you test anything
Mean and SD, or median and IQR — and how to tell which.
Mean ± standard deviation summarises data that are roughly symmetric. Median with the interquartile range summarises data that are skewed or ordinal, because the median is not dragged about by a few extreme values. Frequency and percentage summarise categorical data.
Why it matters
The first table of every results chapter describes the participants, and it is the table an examiner reads most carefully. Reporting a mean for skewed data is the commonest error in it — and it is self-contradictory when the same chapter later justifies a non-parametric test on the grounds that the data are not normal.
An example
Correct: 'The mean age was 58.4 ± 7.2 years. Median length of stay was 4 days (IQR 3–7). Thirty-one participants (73.8%) were female.'
A mean is quoted with its standard deviation, never alone — 'the mean pain score was 4.2' tells the reader nothing about spread.
Common mistakes
- A mean with no measure of spread.
- Mean and SD for length of stay, income, or anything else with a long tail.
- Percentages to two decimal places on a sample of 42. Report '31 (73.8%)', not '73.81%'.
Read next
- Normality, and what to do when it fails — What the assumption actually means, how it is checked, and why failing it is not a disaster.
- Writing a statistical result the way a journal expects — Test statistic, degrees of freedom, p-value, effect size — in that order, in a sentence.
Chavery Research Companion applies this to your own study: it asks the questions in plain language, checks the assumptions against your data, recommends the test, and writes the sentence that reports it. Start free — planning and the master chart cost nothing.