The Normal Distribution
Learning goals
- Read probability as area under a density curve
- Note that a single value has probability zero
- Describe the standard curve, symmetric and fast-decaying
- Apply the 68-95-99.7 rule as rounded percentages
- Convert with to use one table
- Walk back to from a percentage
A shape that keeps turning up
Start with something you already built. In the Binomial Theorem lesson you learned that flipping a coin times produces different ways to get exactly heads. You also learned that those counts are the rows of Pascal’s triangle. Take . There are equally likely sequences of heads and tails, so the probability of exactly heads is
Now draw those eleven probabilities as bars. The binomial coefficients are large in the middle and tiny at the ends. The plain reason is that there are many ways to get five heads and only one way to get ten. The bars pile into a hump.
Coin flips are not special. What they are is a sum of many small independent effects: ten little pushes, each one either up or down, added together. Human height is the same kind of thing, a total of many genetic and nutritional nudges. A measurement error is a total of many tiny disturbances in the instrument and the air. A Galton board, a peg-covered board down which a ball bounces left or right at each row, is that idea made physical. The balls really do pile up in a bell at the bottom, because the number of paths to each slot is a binomial coefficient.
That is the pattern: add up many independent small effects and the total tends toward this one shape, provided no single one of them dominates the total. The claim holds whatever the individual effects look like. That proviso is not a technicality, and we will come back to it. When one ingredient is allowed to run away with the sum, the bell never forms, which is exactly why some famous quantities are not normal. The theorem that makes all of this precise is the Central Limit Theorem. It is one of the deepest results in mathematics, it is the reason the bell turns up in so many unrelated places, and proving it is far beyond this course. Name it, believe the Galton board, and move on. What this lesson can do honestly is describe the curve and work with it.
Probability as area, and one thing we cannot do
Coin flips are discrete: there are finitely many outcomes, each with its own probability, and the probabilities add to . Height is not like that. A person can be cm tall, or cm, or cm. There are infinitely many possible values and no way to hand each one a probability that still totals .
So the whole set-up has to change. For a continuous quantity we do not attach probability to points. We draw a curve, called a density curve, and we agree on one rule:
Probability is area under the curve. The area between and is , and the total area under the whole curve is .
Everything follows from that rule, including a genuine surprise. Ask for the probability that an adult is exactly cm tall, to infinite precision. The region above the single point has width zero, so it has no area at all:
That is not a rounding convention. For a continuous quantity, every individual value has probability zero, and probability lives only in intervals. A first reaction is that this makes the model absurd (surely somebody is exactly cm). The resolution is that “exactly cm” is a claim about infinitely many decimal places, which no person and no ruler ever satisfies. Real questions are always intervals: between and , above , below .
This has a consequence you should test yourself with. In the discrete probability you met earlier in this chapter, and differ whenever is a possible outcome, because the second one also counts it. For a continuous quantity, the outcome contributes zero, so
The endpoint simply does not matter. Include it or exclude it, the area is the same.
Check your understanding
is a continuous quantity with a density curve. Which statement is true?
Probability is area, and the region sitting above the single point has zero width, so it has zero area.
Adding the endpoint therefore adds nothing, which makes and equal. The height of the curve at is not a probability at all; only an area is. And no probability can exceed another by "being easier to satisfy" when the two regions have the same area.
The curve, and the debt
The standard normal curve is the density curve given by
You can read every part of this. The exponent is never positive, and since it gives the same value at and at , so the curve is symmetric about . It is largest when the exponent is , that is when , so the peak is at the centre. And grows fast, so collapses fast: at the height is of the peak, and at it is of the peak, about a ninetieth. That is why the tails look empty. The constant out front is not decoration: it is exactly the number that makes the total area come out to . Where that number comes from is not something this course can show you.
Which brings us to the honest part.
We cannot compute these areas. Finding the area under a curve is the central problem of integral calculus, and you have not met calculus yet. That would be a fair enough excuse on its own, but the truth is worse and more interesting: this particular curve has no elementary antiderivative. The area function of is the formula that would hand you the area from out to any . That function is not built from powers, roots, exponentials, logarithms and trigonometric functions at all. Get the direction of that claim right, because it is easy to state backwards. The difficulty is not in writing down, which we just did, but in writing down the area it encloses. This is not a gap in anybody’s cleverness; it is a theorem. So even a student who finishes calculus cannot write the area between and in closed form, because no such closed form exists.
What people do instead is compute the areas numerically, to as many decimal places as they like, by adding up huge numbers of very thin strips. They then write the answers down in a table. The table is not a definition and it is not a cheat. It is the record of a computation nobody can shortcut. Every number you are about to look up was earned that way, and knowing that is the difference between using the table and understanding it.
The 68-95-99.7 rule
Because the areas have been computed once and for all, we can state the headline results. For the normal curve, measuring distance from the centre in standard deviations:
For any normal distribution with mean and standard deviation :
- about of the values lie within standard deviation of the mean, between and ,
- about lie within standard deviations, between and ,
- about lie within standard deviations, between and .
This is the 68-95-99.7 rule, sometimes called the empirical rule. Be clear about its status: it is not something we derived, it is the numerical answer rounded off. The computed areas are , and . The rule is a memory aid, accurate to within half a percentage point, and when a question wants more precision you go to the table instead.
Two useful facts fall straight out of symmetry. The two halves of the central band are equal, so each holds about . And what is left outside the -standard-deviation band is only about , split between two tails. So roughly of values sit more than standard deviations above the mean. Rare, but not impossible.
Keep the rule’s numbers and the computed numbers apart, because they are not the same and a careful student will notice. The band figures , and are the rule’s own bookkeeping, chosen so that they add up to , and . The computed band areas are , then , then , with left in each tail beyond standard deviations. The two agree to within a fraction of a percentage point, which is all the rule ever promised. So when a question says to use the 68-95-99.7 rule, use the rule’s numbers. When it says to use the table, use the table, and do not expect the last decimal place to match.
Worked example 1 Resting heart rates with the 68-95-99.7 rule
Adult resting heart rates are roughly normal with mean beats per minute and standard deviation . What fraction of adults have a resting heart rate between and beats per minute?
First locate the two endpoints in standard deviations from the mean. Since , the value sits exactly standard deviation below the mean. Since , the value sits exactly standard deviations above it. So the question asks for the area from to in standard-deviation units.
Read that off the band picture. From to is , from to is another , and from to is :
So about of adults fall in that range. The table, which we build next, gives for the same question. The gap of about a third of a percentage point is exactly the price of using a rounded rule. That gap is usually a price worth paying for an answer you can get in your head.
The z-score, and why one table is enough
There are infinitely many normal distributions, one for every choice of and . Heights in centimetres, test scores out of , loaf weights in grams: all different curves. It would be hopeless to tabulate them all. We do not have to, and the reason is a piece of chapter 2.
Define the z-score of a value :
Read it as a unit conversion. The numerator measures how far is from the mean, in the original units (centimetres, points, grams). Dividing by re-expresses that distance in standard deviations. A z-score of means “one and a half standard deviations above the mean”, whatever the original quantity was, and that sentence carries no units at all. That is the whole trick: z-scores are the common currency in which every normal distribution can be compared.
Now here is why the trick is allowed, which is a stronger claim than saying it is convenient. A normal distribution with mean and standard deviation has the density curve
and the claim is that this is not a new curve at all. It is , shifted and rescaled, which are exactly the graph transformations you studied in chapter 2.
Every normal curve is the standard curve in different units#
Start from the general density and substitute the z-score. If , then squaring gives , so the exponent in is
Substituting that back, and pulling the out to the front,
So is built from by exactly two moves from the transformations lesson. Replacing the input by shifts the graph right by and multiplies every horizontal distance from the centre by . Those two effects are what put the peak at and set the width. Chapter 2 names a horizontal scaling by the input multiplier, never by its reciprocal, and the input multiplier here is . So in that lesson’s vocabulary this move is a horizontal compression by . It is the same move under both names, since dividing the inputs by is multiplying them by . The geometric phrasing is the one that matters below, and it has the advantage of staying correct whether is bigger or smaller than . Multiplying the output by then divides every height by , which is what keeps a wide curve from also being a tall one.
Those two moves cancel each other as far as area is concerned, and that is the heart of it. Cut the region under into thin vertical strips, each one close to a rectangle of width and height . The horizontal move multiplies every width by ; the vertical move divides every height by . A strip that had area now has area
the very same area. Since every strip keeps its area, so does the whole region assembled from them.
Two honest remarks about that step. The strips only approximate the region, and the claim that the approximation becomes exact is one we are taking on trust. The table already asked the very same trust of you a moment ago, when its numbers turned out to come from adding up huge numbers of very thin strips. And notice what the argument does and does not do. It never computes an area. It shows only that two areas are equal, because a map that multiplies widths by and divides heights by leaves every area alone. That is why the argument is allowed even though computing either area is not: equality is cheap here, and evaluation is what costs calculus.
The total area under is therefore still , which is what makes the correct constant, and, more usefully, corresponding pieces of area match up:
Any probability question about is therefore the same question about the standard curve, once each value is rewritten as a z-score. One table can serve every normal distribution because, up to a shift and a rescale, there is only one normal curve.
Check your understanding
Jars of peanut butter are filled to a mean of grams with standard deviation grams. One jar weighs grams. What is its z-score?
Subtract the mean to find the distance from the centre, then divide by the standard deviation to re-express that distance in standard deviations.
The jar is standard deviations above the mean, so the z-score is positive. Dividing by instead of by would give , and alone is the distance in grams, not in standard deviations.
Reading the standard normal table
The table records one thing: the area to the left of a given z-score. Write it
Because a single point has no area, and are the same number, so you never have to worry about which inequality the table means.
Here is the table for from to . Two columns of pairs, to keep it compact.
The table starts at and never goes negative, which looks like half a table. It is a full one, because symmetry supplies the rest.
Negative z-scores: #
The standard curve is symmetric about , since . Reflecting the picture in the vertical axis therefore leaves the curve unchanged. That reflection also carries the region to the left of exactly onto the region to the right of . Two regions that are reflections of one another have the same area, so
The total area under the curve is , and the region to the right of is everything that is not to its left, so . Putting those two sentences together,
So a table of positive z-scores is a table of all of them. For example .
Three question types cover almost everything, and each is a subtraction away from the table:
- Left of : read directly.
- Right of : the area to the right is everything not to the left, so .
- Between and : take the area left of and remove the area left of , giving .
Check your understanding
Using the table, what is ?
The table gives areas to the left, and this question asks for the area to the right, so subtract from the total area of .
The value is the table entry itself, which answers instead. The value is , the area from the centre out to . That is the answer you get by subtracting from instead of from , and so forgetting the whole left half of the curve.
Worked example 2 Comparing two scores from different tests
A student scores on a biology test where the class mean was with , and on a chemistry test where the class mean was with . The raw score is higher. Which performance was stronger relative to the class?
The raw scores cannot be compared, because they are measured on different scales. Convert each to a z-score, which is the same scale for both.
For biology,
For chemistry,
The biology score is standard deviations above its class mean, while the chemistry score is only above its own. So the biology performance was the stronger one, even though its raw number was nine points lower.
The table turns those into percentiles: and , so the student beat about of the biology class and about of the chemistry class. This is what z-scores buy you: a way to compare performances that were never measured in the same units.
Worked example 3 A filling machine, with the table
A machine fills bags of flour. The weights are normal with mean grams and standard deviation grams.
Part (a): what fraction of bags weigh less than grams? Standardize the cutoff first:
The table has no negative entries, so use symmetry: . About of bags are under grams.
Part (b): what fraction weigh between and grams? Standardize both endpoints:
For an area between two values, take the area left of the upper one and remove the area left of the lower one:
So about of bags land in that range. Notice that both parts were solved on the same table, even though the flour weights are nowhere near and nowhere near a spread of . The z-score did all the translating.
Going backwards, from a percentage to a value
Every question so far handed you a value and asked for an area. The reverse question is just as common, and just as easy once you read the table in the other direction. You are given an area and you want the value that produces it.
The tool is the z-score formula, solved for . From , multiply both sides by and add :
In words: start at the mean and walk standard deviations. The procedure is to convert the percentage into an area to the left, find the z-score whose table entry matches that area, and then walk.
Worked example 4 Cutoffs for the top of a distribution
Test scores are normal with and .
Part (a): a scholarship goes to the top . What is the cutoff score? If of the area lies above the cutoff, then the area below it is
Now search the column for . It appears exactly, at . So the cutoff sits standard deviations above the mean, and
A score of or better makes the top .
Part (b): an interview goes to the top . The area below the cutoff is , and now the table does not cooperate: and , and falls between them. Our table is only printed in steps of , so the honest answer from this table is “a little below ”, giving a cutoff of about
A finer table, printed in steps of , gives , which is as close to as makes no difference, and the standard cutoff is
Three z-scores are worth memorising because they are asked for constantly: cuts off the top , cuts off the top , and cuts off the top . That last one is why of a normal distribution lies within standard deviations of the mean. It is worth noticing that this is not the same as the in the 68-95-99.7 rule. The rule’s is the rounded area out to standard deviations, which is really . Exactly needs .