Learn
The standard deviation formula
There are two standard deviation formulas, and they differ in exactly one character. This page takes each apart term by term, explains why every operation in them is there, and shows the arithmetic on real numbers. The calculator underneath applies them to your data.
The two formulas
Sample standard deviation — use this when your numbers are a subset drawn from a larger group:
s = √[ Σ(x − x̄)² / (n − 1) ]
Population standard deviation — use this when your numbers are the entire group:
σ = √[ Σ(x − μ)² / N ]
Read either from the inside out: find the mean, measure how far each value sits from it, square those distances, average them, then undo the squaring. The only difference is that the sample version divides by n − 1 where the population version divides by N. If you are working from a sample — which in coursework you almost always are — use the n − 1 version.
Two shortcuts worth knowing before the detail. Standard deviation is simply
the square root of the variance, so if somebody has already handed you s²
you are one operation from the answer: s = √(s²). And in a spreadsheet the whole
thing collapses to =STDEV.S(A1:A100) for a sample or
=STDEV.P(A1:A100) for a population.
Separate numbers with commas, spaces or new lines, or paste a spreadsheet column.
Decimals and negatives are fine; write 10:3 for a value that occurs 3 times.
Your results will appear here: the standard deviation first, then the rest of the summary and a chart.
Standard deviation (sample)
3.04432
Your values typically sit about 3.04 above or below their mean of 15.88, in the same units as your data. 5 of 8 values (63%) fall between 12.83 and 18.92, within one standard deviation of the mean; for normally distributed data about 68% would.
Population SD (σ): 2.8477, if these values are the whole group.
- Count (n)
- 8
- Mean (x̄)
- 15.875
- Variance (s²)
- 9.26786
- Standard error
- 1.07633
- Minimum
- 12
- Q1 (25%)
- 13.75
- Median
- 15.5
- Q3 (75%)
- 17.5
- Maximum
- 21
- Range
- 9
More statistics (5)
- Relative SD (%RSD)
- 19.1768%
- Coefficient of variation
- 0.191768
- Sum (Σx)
- 127
- Sum of squares, Σ(x − x̄)²
- 64.875
- IQR (Q3 − Q1)
- 3.75
Data distribution
Shaded bands mark ±1, ±2 and ±3 SD from the mean. 5 of 8 values (63%) fall within ±1 SD.
Chart as text
Mean 15.875, sample standard deviation s = 3.04432, from 8 values between 12 and 21.
- Within ±1 SD (12.83 to 18.92): 5 of 8 values (63%). About 68% for normal data.
- Within ±2 SD: 8 (100%). About 95% for normal data.
- Within ±3 SD: 8 (100%). About 99.7% for normal data.
Show the working, step by step
Every term, explained
| Term | Name | What it does |
|---|---|---|
| x | Each value | Stands in turn for every individual number in your data. |
| x̄ / μ | Mean | The average — sample mean is written x̄ ("x-bar"), population mean μ ("mu"). |
| x − x̄ | Deviation | How far one value sits from the mean. Negative below it, positive above. |
| (x − x̄)² | Squared deviation | Makes every term positive so they cannot cancel, and penalises far-out values more. |
| Σ | Sum | "Add all of these up." |
| n / N | Count | How many values. Lower case n for a sample, capital N for a population. |
| n − 1 | Degrees of freedom | Bessel's correction — see below. |
| √ | Square root | Returns the answer to the original units after the squaring. |
Why square the deviations at all?
Try averaging the raw deviations on 2, 4, 6. The mean is 4, so the deviations are
−2, 0 and +2, and they sum to zero. That is not a coincidence of this example: the deviations
from the mean sum to zero for every dataset, by the definition of the mean. An
unsquared average of them would report zero spread for all data, which is useless.
Squaring solves that by making every term non-negative. It also has the side effect of weighting a deviation of 10 as a hundred times more significant than a deviation of 1, rather than ten times — which is why a single outlier can move a standard deviation so sharply. Taking the square root at the end undoes the inflation of scale, but not that reweighting.
Why n − 1 for a sample?
A sample's mean is calculated from that sample, so it sits as close to those particular numbers as any single point can. The true population mean is almost always somewhere slightly else, and further from your data on average. Measuring spread around the sample's own too-flattering centre therefore understates the population's real spread every time — the bias is systematic, not random.
Dividing by n − 1 instead of n inflates the result by exactly enough to cancel that bias out. The correction is named for Friedrich Bessel, and its size depends entirely on n:
| n | n / (n − 1) | Effect on the variance |
|---|---|---|
| 3 | 1.500 | 50% larger |
| 5 | 1.250 | 25% larger |
| 10 | 1.111 | 11% larger |
| 30 | 1.034 | 3.4% larger |
| 100 | 1.010 | 1% larger |
| 1000 | 1.001 | 0.1% larger |
For a large sample the choice barely matters. For a small one it matters a great deal, which is exactly when people are most likely to be working by hand and most likely to pick wrong.
A worked example, both ways
Data: 2, 4, 4, 4, 5, 5, 7, 9.
- Sum = 40, count = 8, so the mean is 5.
- Deviations: −3, −1, −1, −1, 0, 0, 2, 4.
- Squared: 9, 1, 1, 1, 0, 0, 4, 16. Their sum, Σ(x − x̄)², is 32.
-
As a population: 32 ÷ 8 = 4, and √4 = σ = 2 exactly.
As a sample: 32 ÷ 7 = 4.5714, and √4.5714 = s = 2.1381.
This is the standard textbook example precisely because the population answer comes out as a whole number. Paste it into the calculator above and toggle the mode to watch only the denominator change.
The shortcut formula, and why not to use it
Older textbooks give a one-pass rearrangement that avoids computing the mean first:
s² = [ Σx² − (Σx)² / n ] / (n − 1)
Algebraically it is identical. Numerically it is not. It subtracts two large, nearly equal
quantities, and in floating-point arithmetic that step throws away most of the significant
digits. On the values 1000000.1, 1000000.2, 1000000.3 — whose true variance is
exactly 0.01 — the shortcut can return zero or even a negative number, which is impossible
for a variance. The formulas at the top of this page, computed with
Welford's algorithm, do not have this failure mode.
Every standard deviation formula, in one table
The same idea rearranged for whatever you happen to have — raw values, a frequency table, several groups, or a variance somebody already computed.
| You want | Formula | Calculator |
|---|---|---|
| Sample SD | s = √[ Σ(x − x̄)² / (n − 1) ] | Standard deviation |
| Population SD | σ = √[ Σ(x − μ)² / N ] | Standard deviation |
| Variance | s² = Σ(x − x̄)² / (n − 1) | Variance |
| SD from a variance | s = √(s²) | Variance |
| Sum of squares | SS = Σ(x − x̄)² | Sum of squares |
| Standard error of the mean | SE = s / √n | Standard error |
| Relative SD (%RSD) | %RSD = 100 × s / |x̄| | Relative SD |
| Z-score | z = (x − μ) / σ | Z-score |
| Pooled SD | sp = √[ Σ(nᵢ − 1)sᵢ² / Σ(nᵢ − 1) ] | Pooled SD |
| Weighted SD | sw = √[ Σwᵢ(xᵢ − x̄w)² / (Σwᵢ − 1) ] | Weighted SD |
| Grouped frequency data | s = √[ Σf(x − x̄)² / (Σf − 1) ] | Grouped data |
| Mean absolute deviation | MAD = Σ|x − x̄| / n | MAD |
| Root mean square | RMS = √( Σx² / n ) | RMS |
| Binomial SD | σ = √( np(1 − p) ) | Binomial |
| Poisson SD | σ = √λ | Poisson |
| Two-asset portfolio SD | σp = √( w₁²σ₁² + w₂²σ₂² + 2w₁w₂ρσ₁σ₂ ) | Portfolio SD |
The formula for grouped data
When the data arrives as a frequency table rather than a list, you never see the individual values — only that some midpoint x occurred f times. Each squared deviation is therefore weighted by its frequency, and the count becomes Σf:
s = √[ Σf(x − x̄)² / (Σf − 1) ] where x̄ = Σfx / Σf
The answer is an approximation, because treating every value in a class as sitting exactly at the midpoint is a fiction — real values are spread across the interval. Expect it to run slightly low for heavily skewed classes. The grouped-data calculator shows the fx and fx² columns as it goes.
The formula to copy and paste (Word, Google Docs)
Plain Unicode, one line each — safe in Word, Google Docs, email or a comment:
s = √(Σ(x − x̄)² / (n − 1))σ = √(Σ(x − μ)² / N)
In LaTeX:
s = \sqrt{\frac{\sum (x_i - \bar{x})^2}{n-1}}\sigma = \sqrt{\frac{\sum (x_i - \mu)^2}{N}}
For a properly typeset version, both editors have an equation tool that accepts this syntax.
In Google Docs choose Insert → Equation and type
\sigma, \sqrt or \sum followed
by a space; each turns into its symbol. In Word, press
Alt+= to open an equation and paste the LaTeX line above — current
versions of Word convert it when you choose LaTeX in the Equation ribbon.
The notation guide has the individual characters and how to type each one.
The same formula in Excel, Python and R
Every one of these applies the formulas at the top of this page; they differ only in which denominator they pick by default, and that default is the single most common source of a wrong answer.
| Tool | Sample (n − 1) | Population (N) | Default |
|---|---|---|---|
| Excel, Sheets | =STDEV.S(A1:A100) | =STDEV.P(A1:A100) | Neither — you choose |
| NumPy | np.std(a, ddof=1) | np.std(a) | Population |
| pandas | s.std() | s.std(ddof=0) | Sample |
| Python statistics | statistics.stdev(a) | statistics.pstdev(a) | Neither — separate functions |
| R | sd(x) | sd(x) * sqrt((n − 1) / n) | Sample |
NumPy and pandas disagreeing on the default is a genuine trap: the same column run through
both returns two different numbers unless you set ddof explicitly. R has no
population function at all, so you rescale sd() by hand. The
Excel,
Python and
R guides work through each in full.
Related calculators
-
Standard deviation calculator
Apply the formula to your own numbers and see each step.
-
Symbols and notation
What σ, s, μ, x̄, n and Σ each stand for.
-
Sample vs population
Which denominator to use, and why it matters most for small n.
-
Variance calculator
The same formula without the final square root.
Common questions
What is the standard deviation formula?
For a sample: s = √[ Σ(x − x̄)² / (n − 1) ]. For a population:
σ = √[ Σ(x − μ)² / N ]. The two are identical except for the denominator —
a sample divides by n − 1, a population by N.
Why are the deviations squared?
Because deviations above and below the mean always sum to exactly zero, so averaging them raw would give 0 for every dataset. Squaring makes every term positive so they cannot cancel. It also weights large deviations more heavily than small ones, which is usually what you want from a measure of spread.
Taking absolute values instead would also stop the cancelling — that gives the mean absolute deviation, a real and occasionally preferable statistic. Squaring won out because it is differentiable everywhere and because variances of independent quantities add, which makes the whole of statistical theory tractable.
What is the computational or "shortcut" formula?
s² = [ Σx² − (Σx)²/n ] / (n − 1). It gives the same answer in exact
arithmetic and needs only one pass through the data, which mattered when people worked with
mechanical calculators.
On a computer it is a bad idea. It subtracts two large and nearly equal numbers, which destroys precision — on the data 1000000.1, 1000000.2, 1000000.3 it can return a negative variance. This site uses Welford's algorithm instead.
What does the Σ symbol mean in the formula?
Σ is the Greek capital sigma and means "add up all of these". Σ(x − x̄)²
instructs you to work out (x − x̄)² for every value x in the dataset and total the results.
The full notation guide is here.
Is the standard deviation formula the same as the variance formula?
Almost. The variance is everything under the square root; the standard deviation is the square root of it. So σ = √(σ²), and σ² = σ². The only practical difference is units: variance is in squared units, standard deviation is in the original units.