The Normal Distribution Advanced. This lesson goes beyond core Algebra II. You can skip it.
Learning goals
- Read probability as area under a density curve
- Note that a single value has probability zero
- Describe the standard curve, symmetric and fast-decaying
- Apply the 68-95-99.7 rule as rounded percentages
- Convert with to use one table
- Walk back to from a percentage
A shape that keeps turning up
Start with something you already built. In the Binomial Theorem lesson you learned that flipping a coin times produces different ways to get exactly heads. You also learned that those counts are the rows of Pascal’s triangle. Take . There are equally likely sequences of heads and tails, so the probability of exactly heads is
Now draw those eleven probabilities as bars. The binomial coefficients are large in the middle and tiny at the ends. The plain reason is that there are many ways to get five heads and only one way to get ten. The bars pile into a hump.
Coin flips are not special. What they are is a sum of many small independent effects: ten little pushes, each one either up or down, added together. Human height is the same kind of thing, a total of many genetic and nutritional nudges. A measurement error is a total of many tiny disturbances in the instrument and the air. A Galton board, a peg-covered board down which a ball bounces left or right at each row, is that idea made physical. The balls really do pile up in a bell at the bottom, because the number of paths to each slot is a binomial coefficient.
That is the pattern: add up many independent small effects and the total tends toward this one shape, provided no single one of them dominates the total and a few other technical conditions hold that this course will not state precisely. That proviso is not a technicality, and we will come back to it. When one ingredient is allowed to run away with the sum, the bell may not form, which is part of why some famous quantities are not normal. The theorem that makes all of this precise is the Central Limit Theorem. It is one of the deepest results in mathematics, it is the reason the bell turns up in so many unrelated places, and proving it is far beyond this course. What this lesson can do honestly is describe the curve and work with it.
Probability as area, and one thing we cannot do
Coin flips are discrete: there are finitely many outcomes, each with its own probability, and the probabilities add to . Height is not like that. A person can be cm tall, or cm, or cm. There are infinitely many possible values and no way to hand each one a probability that still totals .
So the whole set-up has to change. For a continuous quantity we do not attach probability to points. We draw a curve, called a density curve, and we agree on one rule. Let name the numerical result of the experiment, the height of a randomly chosen adult, say. Then a statement like is itself an event in the sense of the previous lesson: the set of every outcome whose measured value is below , so is just that event’s probability, written in this new notation:
Probability is area under the curve. The curve never dips below the horizontal axis, the area between and is , and the total area under the whole curve is .
Everything follows from that rule, including a genuine surprise. Ask for the probability that an adult is exactly cm tall, to infinite precision. The region above the single point has width zero, so it has no area at all:
That is not a rounding convention. For a continuous quantity, every individual value has probability zero, and probability lives only in intervals. A first reaction is that this makes the model absurd (surely somebody is exactly cm). It does not. Probability zero does not mean impossible; it means that one single point gets no share of the curve’s area, because area needs width and a point has none. In practice this rarely matters anyway, since “exactly cm” would mean matching infinitely many decimal places, which no measurement ever pins down. Real questions are always intervals: between and , above , below .
This has a consequence you should test yourself with. In the discrete probability you met earlier in this chapter, and differ whenever is a possible outcome, because the second one also counts it. For a continuous quantity, the outcome contributes zero, so
The endpoint simply does not matter. Include it or exclude it, the area is the same.
Check your understanding
is a continuous quantity with a density curve. Which statement is true?
Probability is area, and the region sitting above the single point has zero width, so it has zero area.
Adding the endpoint therefore adds nothing, which makes and equal. The height of the curve at is not a probability at all; only an area is. And no probability can exceed another by "being easier to satisfy" when the two regions have the same area.
The curve, and the debt
The standard normal curve is the density curve given by
You can read every part of this. The exponent is never positive, and since it gives the same value at and at , so the curve is symmetric about . It is largest when the exponent is , that is when , so the peak is at the center. And grows fast, so collapses fast: at the height is of the peak, and at it is of the peak, about a ninetieth. That is why the tails look empty. The constant out front is not decoration: it is exactly the number that makes the total area come out to . Where that number comes from is not something this course can show you.
We cannot compute these areas by hand. Finding the area under a curve is what integral calculus is for, and you have not met calculus yet. Even worse: mathematicians have proven that this curve’s area cannot be reached by the ordinary algebra formulas that calculus normally produces, no matter how much calculus you eventually learn. That is not a gap in anybody’s cleverness; it is a theorem.
What people do instead is compute the areas numerically, to as many decimal places as they like, by adding up huge numbers of very thin strips. They then write the answers down in a table. The table is not a definition and it is not a cheat. It is the record of a computation nobody can shortcut. Every number you are about to look up was earned that way, and knowing that is the difference between using the table and understanding it.
Check your understanding
Put these three heights of the standard normal curve in order, tallest to shortest: , , .
Three facts about together decide this. The curve is symmetric about , so only the distance from matters, not the sign: ranks exactly where would. The peak sits at , so is tallest no matter what. And the curve decays as grows, so between the two remaining points the one closer to is taller: is distance from , is distance , so . Putting the three together gives .
The 68-95-99.7 rule
Because the areas have been computed once and for all, we can state the headline results. For the normal curve, measuring distance from the center in standard deviations:
For any normal distribution with mean and standard deviation :
- about of the values lie within standard deviation of the mean, between and ,
- about lie within standard deviations, between and ,
- about lie within standard deviations, between and .
This is the 68-95-99.7 rule, sometimes called the empirical rule. Be clear about its status: it is not something we derived, it is the numerical answer rounded off. The computed areas are , and . The rule is a memory aid, accurate to within half a percentage point, and when a question wants more precision you go to the table instead.
Two useful facts fall straight out of symmetry. The two halves of the central band are equal, so each holds about . And using the rule’s own rounded bookkeeping, what is left outside the -standard-deviation band is about , split evenly between two tails, so about of values sit more than standard deviations above the mean. Rare, but not impossible.
Keep the rule’s numbers and the computed numbers apart, because they are not the same and a careful student will notice. The band figures , and are the rule’s own bookkeeping, chosen so that they add up to , and . The computed band areas are , then , then , with left in each tail beyond standard deviations. The two agree to within a fraction of a percentage point, which is all the rule ever promised. So when a question says to use the 68-95-99.7 rule, use the rule’s numbers. When it says to use the table, use the table, and do not expect the last decimal place to match.
Worked example 1 Resting heart rates with the 68-95-99.7 rule
Adult resting heart rates are roughly normal with mean beats per minute and standard deviation . What fraction of adults have a resting heart rate between and beats per minute?
First locate the two endpoints in standard deviations from the mean. Since , the value sits exactly standard deviation below the mean. Since , the value sits exactly standard deviations above it. So the question asks for the area from to in standard-deviation units.
Read that off the band picture. From to is , from to is another , and from to is :
So about of adults fall in that range. The table, which we build next, gives for the same question. The gap of about a third of a percentage point is exactly the price of using a rounded rule. That gap is usually a price worth paying for an answer you can get in your head.
Check your understanding
For a normal distribution, about what percent of values lie between and ?
This range covers two right-hand bands: from to is , and from to is another .
The value alone only reaches out to , missing the second band. The value is the whole two-sided band out to , not just the right half of it, and is the whole two-sided band out to .
Try it yourself. In the figure below, drag the two marks to one standard deviation each side of the mean and read the percentage: about 68. Now slide the mean without touching anything else. The marks stay where you put them, so they are no longer one spread each side, and the percentage falls away. Walk them back out to one spread each side of the new mean and it reads 68 again. That is the point: the answer depends only on how many standard deviations out the marks are, never on where the mean sits or how wide the curve is, which is why a single table can serve every normal distribution. Then slide the spread and watch the curve flatten as it widens, because the area underneath is always the whole of it.
How far out, in standard deviations?
Mean 100, standard deviation 15. Between 85 and 115 lies about 68% of the area. That is one standard deviation each side of the mean, wherever the mean is.
The z-score, and why one table is enough
There are infinitely many normal distributions, one for every choice of and . Heights in centimeters, test scores out of , loaf weights in grams: all different curves. It would be hopeless to tabulate them all. We do not have to, and the reason is a piece of chapter 2.
Define the z-score of a value :
Read it as a unit conversion. The numerator measures how far is from the mean, in the original units (centimeters, points, grams). Dividing by re-expresses that distance in standard deviations. A z-score of means “one and a half standard deviations above the mean”, whatever the original quantity was, and that sentence carries no units at all. That is the whole trick: z-scores are the common currency in which every normal distribution can be compared. Applying that same conversion to the whole quantity, not just one value, gives it a name: is the standardized version of , the same random result relabeled in standard deviations, so a statement like is just written in the new units.
Now here is why the trick is allowed, which is a stronger claim than saying it is convenient. A normal distribution with mean and standard deviation has the density curve
and the claim is that this is not a new curve at all. It is , shifted and rescaled, exactly the graph transformations you studied in chapter 2: shifted right by , and stretched horizontally, with its height correspondingly shrunk, by a factor of .
Here is why that is enough to justify one shared table. Stretch a shape sideways by a factor of while shrinking its height by that same factor, and its area does not change: a strip that gets times wider and times shorter has the same width times height as before. So shifting and rescaling into leaves matching pieces of area exactly equal,
which is the whole reason a single table can serve every normal distribution. Up to a shift and a rescale, there is only one normal curve.
Check your understanding
Jars of peanut butter are filled to a mean of grams with standard deviation grams. One jar weighs grams. What is its z-score?
Subtract the mean to find the distance from the center, then divide by the standard deviation to re-express that distance in standard deviations.
The jar is standard deviations above the mean, so the z-score is positive. Flipping the ratio to instead of would give , and alone is the distance in grams, not in standard deviations.
Reading the standard normal table
The table records one thing: the area to the left of a given z-score. Write it
Because a single point has no area, and are the same number, so you never have to worry about which inequality the table means.
Here is the table for from to . Two columns of pairs, to keep it compact.
The table starts at and never goes negative, which looks like half a table. It is a full one, because symmetry supplies the rest.
Negative z-scores: #
The standard curve is symmetric about , since . Reflecting the picture in the vertical axis therefore leaves the curve unchanged. That reflection also carries the region to the left of exactly onto the region to the right of . Two regions that are reflections of one another have the same area, so
The total area under the curve is , and the region to the right of is everything that is not to its left, so . Putting those two sentences together,
So a table of positive z-scores is a table of all of them. For example .
Three question types cover almost everything, and each is a subtraction away from the table:
- Left of : read directly.
- Right of : the area to the right is everything not to the left, so .
- Between and : take the area left of and remove the area left of , giving .
Check your understanding
Using the table, what is ?
The table gives areas to the left, and this question asks for the area to the right, so subtract from the total area of .
The value is the table entry itself, which answers instead. The value is , the area from the center out to . That is the answer you get by subtracting from instead of subtracting from , and so forgetting the whole left half of the curve.
Worked example 2 Comparing two scores from different tests
A student scores on a biology test where the class mean was with , and on a chemistry test where the class mean was with . Both classes’ scores are roughly normal, which is what makes a table lookup valid below. The raw score is higher. Which performance was stronger relative to the class?
The raw scores cannot be compared, because they are measured on different scales. Convert each to a z-score, which is the same scale for both.
For biology,
For chemistry,
The biology score is standard deviations above its class mean, while the chemistry score is only above its own. So the biology performance was the stronger one, even though its raw number was nine points lower.
The table turns those into percentiles: and , so the student beat about of the biology class and about of the chemistry class. This is what z-scores buy you: a way to compare performances that were never measured in the same units.
Worked example 3 A filling machine, with the table
A machine fills bags of flour. The weights are normal with mean grams and standard deviation grams.
Part (a): what fraction of bags weigh less than grams? Standardize the cutoff first:
The table has no negative entries, so use symmetry: . About of bags are under grams.
Part (b): what fraction weigh between and grams? Standardize both endpoints:
For an area between two values, take the area left of the upper one and remove the area left of the lower one:
So about of bags land in that range. Notice that both parts were solved on the same table, even though the flour weights are nowhere near and nowhere near a spread of . The z-score did all the translating.
Going backwards, from a percentage to a value
Every question so far handed you a value and asked for an area. The reverse question is just as common, and just as easy once you read the table in the other direction. You are given an area and you want the value that produces it.
The tool is the z-score formula, solved for . From , multiply both sides by and add :
In words: start at the mean and walk standard deviations. The procedure is to convert the percentage into an area to the left, find the z-score whose table entry matches that area, and then walk.
Worked example 4 Cutoffs for the top of a distribution
Test scores are normal with and .
Part (a): a scholarship goes to the top . What is the cutoff score? If of the area lies above the cutoff, then the area below it is
Now search the column for . It appears exactly, at . So the cutoff sits standard deviations above the mean, and
A score of or better makes the top .
Part (b): an interview goes to the top . The area below the cutoff is , and now the table does not cooperate: and , and falls between them. Our table is only printed in steps of , so the honest answer from this table is “a little below ”, giving a cutoff of about
A finer table, printed in steps of , gives , which is as close to as makes no difference, and the standard cutoff is
Three z-scores come up constantly and are worth recognizing: cuts off the top , cuts off the top , and cuts off the top . That last one is why of a normal distribution lies within standard deviations of the mean. It is worth noticing that this is not the same as the in the 68-95-99.7 rule. The rule’s is the rounded area out to standard deviations, which is really . Exactly needs .
Check your understanding
Delivery times are normal with minutes and minutes. The table gives . What delivery time is the th percentile?
means has of the area to its left, so start at the mean and walk standard deviations.
The value walks the wrong direction, subtracting instead of adding. The value is the z-score itself, not a delivery time, and comes from walking a full standard deviation instead of half of one.