Why does the standard deviation formula have a square root?
The square root is there to return the answer to the units of the data. The formula squares each distance from the mean, which turns centimetres into square centimetres. Taking the square root at the end undoes that, so a standard deviation of heights is in centimetres, like the heights themselves.
Why is there a square root in the standard deviation equation?
s = √( Σ(xᵢ − x̄)² ÷ (n − 1) )
Read it from the inside out: find each value's deviation from the mean, square it, add the squares, divide by n − 1 to get the variance, then take the square root. Every step before the root works in squared units. The root is the only step that brings the result back to the scale you measured on.
If you met this as a multiple-choice question, the answer is “to return the answer to the units of the quantity you were measuring”. The other usual options are wrong: the root does not solve for time, does not force the answer above 1 (a standard deviation can be 0.03), and does nothing about outliers.
A worked example with units
Five heights, in cm: 160, 165, 170, 175, 180. The mean is 170 cm.
| Height xᵢ (cm) | Deviation xᵢ − x̄ (cm) | Squared (cm²) |
|---|---|---|
| 160 | −10 | 100 |
| 165 | −5 | 25 |
| 170 | 0 | 0 |
| 175 | +5 | 25 |
| 180 | +10 | 100 |
| Sum | 0 | 250 |
- Variance: 250 ÷ 4 = 62.5 cm². Square centimetres are an area, which says nothing obvious about how tall people are.
- Standard deviation: √62.5 = 7.91 cm. Back in centimetres, and directly comparable with the heights: most are within about 8 cm of 170.
Why square the deviations in the first place?
Look at the deviation column: it adds up to 0. That is always true, for any data set, because the mean is the balance point. Values above it exactly offset values below it, so the average deviation is always zero and useless as a measure of spread. Squaring makes every term positive (−10 and +10 both become 100), so they cannot cancel.
Squaring also weights large deviations more heavily: a value 10 cm out contributes four times as much as one 5 cm out. That is why a single outlier can inflate the standard deviation so much.
Why not use absolute values instead?
Taking |xᵢ − x̄| also stops the cancelling, and it needs no square root because the units never change. That measure exists: the mean absolute deviation, which is 6 cm for the heights above, a little smaller than the SD of 7.91 cm.
The squared version became the standard for practical reasons. Variances of independent quantities add: if two errors are independent, the variance of their sum is the sum of their variances, which has no simple equivalent for absolute deviations. Squares are smooth, so they work with calculus and least squares regression. And σ is the parameter that defines the width of the normal curve, so the SD links straight to the 68–95–99.7 rule.
Why divide by n − 1 and not n?
For a sample, the deviations are measured from the sample mean, which sits closer to the sample's own values than the true population mean does. Dividing by n would underestimate the population variance on average; n − 1 (Bessel's correction) fixes that. For a whole population, divide by N. The sample vs population page compares the two.
Variance and standard deviation: the same information
The variance and the SD carry the same information: square the SD to get the variance, take the root of the variance to get the SD. Use the variance for algebra, such as adding independent sources of variation, and the SD for describing data to people. The variance calculator gives both, with every step, and standard deviation vs variance covers when to use each.
Related calculators
-
Standard deviation formula
Every symbol in the sample and population formulas.
-
Standard deviation vs variance
When to use each, and how they convert.
-
Mean absolute deviation calculator
The spread you get with absolute values instead of squares.
Common questions
Which of the following could explain why the equation for standard deviation has a square root?
To return the answer to the units of the quantity you were measuring. Squaring the deviations turns cm into cm²; the square root turns the variance back into cm. It does not make the answer bigger than 1, solve for time or remove outliers.
Why do we square the deviations instead of just adding them?
Because the plain deviations always add up to zero: the values above the mean exactly balance those below. Squaring makes every deviation positive, so they can no longer cancel.
What is the square of the standard deviation?
The variance. Squaring the standard deviation undoes the square root and gives back the variance, s² for a sample and σ² for a population.
Why not use the absolute value of the deviations instead?
You can, and the result is the mean absolute deviation (MAD). The squared version won out because variances of independent quantities add up, it is easy to work with in algebra and calculus, and it is the natural scale parameter of the normal distribution.