Introduction to Probability
Learning goals
- Place a probability on the scale from to , as a fraction, decimal, or percent
- Find a probability by dividing favorable outcomes by equally likely total outcomes, after checking that assumption holds
- Apply the complement rule, since all outcome probabilities add to one
- Distinguish experimental probability from theoretical, and say how more independent trials bring them closer
- Count a two-stage experiment's outcomes, then apply the same formula
A number from 0 to 1
Probability is measured on a fixed scale from to . A theoretical probability of means the event is impossible, and a theoretical probability of means it is certain; getting zero successes in a finite experiment does not prove an event is impossible, since chance alone can produce a run with none. Everything else falls in between, and the larger the number, the more likely the event. A probability of sits exactly in the middle and means the event is as likely to happen as not, the “fifty-fifty” of everyday speech.
Because every probability in this chapter is a ratio of whole numbers, a fraction of equally likely outcomes, you can write it three equivalent ways. The three forms are a fraction, a decimal, and a percent from to . A one-in-four chance is , which is , which is , all naming the same spot on the scale. Use whichever form is clearest for the problem. Be ready to convert between them, since these are the fractions, decimals, and percents from earlier chapters doing familiar work.
No probability is ever negative or larger than , because an event cannot happen less often than never or more often than always. If a calculation hands you or , you have made a mistake somewhere.
Outcomes, and what equally likely means
To compute a probability we have to be careful about the setup. An experiment is any action with an uncertain result, such as flipping a coin or rolling a die. Each possible result is an outcome, and the full list of outcomes is the sample space. An event is the part of the sample space you care about, a chosen collection of outcomes. Rolling a die has the sample space , and “roll an even number” is the event .
The method of this lesson rests on one assumption: that the outcomes are equally likely, meaning each one has the same chance as any other. A fair coin, a fair die, a spinner with equal sectors, and a marble drawn at random are the standard ways of promising this. The words fair, equal, and at random are not decoration; they are what let you find a probability just by counting.
Equally likely is an assumption you must check, not something that is always true. Tomorrow’s weather is either rain or no rain, but those two outcomes are usually not equally likely. So the chance of rain is not automatically . A bent coin might favor one side, and a spinner whose sectors have different sizes lands more often on the big ones. When the outcomes are not equally likely, counting them is not enough, and you need data or extra information instead.
The probability of an event
Now the payoff. Take the coin flip from the opening: heads or tails, equally likely outcomes. Since they are equally likely, they split the one whole chance into equal shares, so each outcome, including tails, gets . That is exactly the “fifty-fifty” you already know.
The same reasoning works for any number of equally likely outcomes, not just . Suppose an experiment has equally likely outcomes. Something must happen, so the chances of all outcomes together make up the whole, named (certain). Since the outcomes are equal, that one whole splits into equal shares, so each single outcome has probability .
An event is just some of those outcomes. If the event contains of them, it collects of the equal shares, so its probability is . That is the central formula of the lesson:
“Favorable” is only a name for the outcomes in your event; it does not mean good, just counted. The denominator is the size of the whole sample space, which is exactly the total you practiced counting in the previous lesson.
Worked example 1 The probability of tails
A fair coin is flipped once. What is the probability that it lands tails?
A flip has two equally likely outcomes, heads and tails, so the sample space has outcomes. Exactly one of them, tails, is favorable.
The word “fair” is what lets us count this way. It promises the two sides are equally likely, so each one takes half of the whole.
Worked example 2 The probability of an even roll
A fair die is rolled once. What is the probability of rolling an even number?
The sample space is the six faces , all equally likely, so the total is . The event “even” is the set , the three faces shaded in the die figure above, so there are favorable outcomes.
For a final answer, simplify the fraction when you can. Here and are the same number, and is the simplest way to write it, though keeping the unreduced form is what shows the favorable and total counts.
Check your understanding
A fair die is rolled once. What is the probability of rolling a ?
The die has equally likely faces, and only one of them is a , so there is favorable outcome out of .
Every single face of a fair die has this same probability, .
More single-event probabilities
The formula does not care whether the experiment is a coin, a die, a spinner, or a bag of marbles. The work is always the same: count the favorable outcomes, count the total equally likely outcomes, and divide. Here it is on a spinner and on a random draw.
Worked example 3 Reading a probability off a spinner
The spinner above has equal sectors, and of them are shaded. If you spin it once, what is the probability of landing on a shaded sector?
Because the sectors are equal, the sectors are equally likely, so the sample space has outcomes. The shaded sectors are the favorable ones.
Equal sectors are what make this work. If the shaded part were one wide sector and the rest were narrow, the outcomes would not be equally likely. In that case you could not just count sectors.
An experiment with exactly equally likely outcomes is worth a closer look. The reason is that the counting and the three ways of writing a probability all show up in one picture. Picture a drum of raffle tickets, mixed and drawn at random. Each square below is one ticket, and shading a square marks it as a favorable outcome. So the count of shaded squares is the numerator, while the squares are always the denominator.
Counting favorable outcomes out of a hundred
25 squares of 100 outcomes are favorable. So the probability is 25/100, or 1/4 in lowest terms. As a percent, 25%. As a decimal, 0.25.
Give yourself some targets. Shade squares and one probability is reported three ways at once, as , as , and as , all naming the one-in-four chance. Shade none and you are at , the impossible end of the scale; shade all and you are at , certain. Then find the shading where the favorable squares and the ones left unshaded are equal in number, and . At that split each event, favorable or not, has probability , the even chance in the middle.
Worked example 4 Drawing a marble at random
A bag holds red, blue, and green marbles, all the same size and weight. You draw one marble without looking. What is the probability it is red?
The marbles are all the same size and weight and the bag is well mixed, so drawing one without looking makes every marble equally likely to be the one you get. The total is the whole bag. Count it with the addition principle from the last lesson:
The favorable outcomes are the red marbles, so
The same draw gives and , and notice the three probabilities add to , the whole bag.
Why the probabilities add to one
That last observation, that the probabilities of all the outcomes add to , is not a coincidence of that bag. It is always true, and it is worth a careful statement, because it is also the engine behind a shortcut called the complement rule.
The outcome probabilities add to , and #
Take an experiment with equally likely outcomes, so each outcome has probability . Add up the probabilities of all of them. There are shares, each equal to , so the total is
The whole sample space has probability , which is just the statement that something is certain to happen.
Now take any event built from of these outcomes, so . Its complement, written “not ”, is everything else in the sample space, the other outcomes, so . Adding the two,
because every outcome is either in or not in , never both and never neither. Subtracting from each side gives the complement rule:
The complement rule
The complement rule earns its own line because it often saves work. When the outcomes you want are awkward to count but the leftover outcomes are easy, count the leftovers and subtract from . Problems that ask for “at least one” of something are the classic case, since “at least one” is the complement of “none”.
Worked example 5 Using the complement rule
Using the same bag of marbles ( red, blue, green), what is the probability of drawing a marble that is not red?
You could count the favorable outcomes directly. Not red means blue or green, which is marbles, so . The complement rule reaches the same answer faster, starting from :
Both routes agree, which is the whole point. A marble is either red or not red, so those two probabilities must fill up the one whole.
Here is where the shortcut actually saves work. Flip two fair coins together, so one cannot affect the other, and ask for the probability of at least one head. Counting directly means listing every way to get one head or two: HH, HT, TH, that is of the equally likely outcomes. The complement is faster, because “at least one head” fails only for TT, no heads at all:
Both answers agree, , but the complement found it by ruling out just one outcome instead of counting three.
Check your understanding
A fair die is rolled once. What is the probability of not rolling a ?
From the earlier checkpoint, . 'Not a ' is the complement, so subtract from .
As a check, the faces that are not a are , which is of the faces, again .
Theoretical and experimental probability
Everything so far has been theoretical probability: you reason about equally likely outcomes and compute a fraction before ever running the experiment. There is a second kind. Experimental probability, also called relative frequency, comes from actually doing the experiment many times and recording what happens:
If you flip a fair coin times and get heads, the experimental probability of heads is . That happens even though the theoretical value is . The two are different numbers, and that is expected, because a short run of flips rarely splits exactly evenly.
Here is the key link between them. This works when the trials do not affect each other, such as flipping the same coin over and over: one more trial does not have to move the experimental probability closer to the theoretical value, since a single flip can push the running total either way and short runs bounce around. What is true over the long run: as you repeat more and more independent trials, that running proportion becomes increasingly likely to sit close to the theoretical probability. Twenty flips might give heads, but ten thousand flips will almost certainly land very near . This pattern, that relative frequency settles near the true probability as the number of trials grows, is called the law of large numbers. It is why an insurer can predict the long-run rate of claims closely even though any single policy is uncertain. It is also the bridge that makes a theoretical something you can actually watch happen.
Worked example 6 Experimental probability of heads
A fair coin is flipped times and lands heads times. What is the experimental probability of heads from this data, and how does it compare with the theoretical value?
Experimental probability divides the number of times the event happened by the number of trials.
The theoretical probability for a fair coin is . The experimental value is close but not exactly equal, which is normal for only flips. Flip the same coin independently ten thousand times and the fraction of heads would almost certainly sit much nearer to .
A gentle step to two-stage experiments
Single events are the heart of probability, but the same formula stretches to experiments built from two stages, as long as you keep counting carefully. Rolling two dice, flipping two coins, or spinning twice are all sequences of stages, so the multiplication principle from the last lesson counts their total outcomes: each stage still offers the same number of results next, no matter what came before. The probability formula needs two things more than a correct count, though. Each stage’s own outcomes must already be equally likely on their own, such as a fair die or a fair coin, and the stages must not affect each other’s results, called independent trials, such as rolling two dice together or flipping two coins together. When both hold, as they do for every two-stage experiment in this lesson, the combined outcomes are equally likely too. The probability formula does not change at all: it is still favorable outcomes over total equally likely outcomes. Only the counting gets a little bigger.
Worked example 7 Rolling a double with two dice
Two fair dice, one red and one blue, are rolled together. What is the probability that they show the same number, a double?
First count the total. The red die has results and the blue die has , and one die’s result does not change the other’s. The roll is a sequence of two independent stages, so the multiplication principle gives
equally likely outcomes, one per cell of the grid above. The doubles are , which is favorable outcomes, the shaded diagonal.
We found the denominator by counting, exactly the tool from the previous lesson. The pairs stay equally likely because each die is fair on its own and the two rolls do not affect each other.
Check your understanding
Two fair coins are flipped together, so one does not affect the other. What is the probability that both land heads?
Each coin has outcomes, so by the multiplication principle there are equally likely results: HH, HT, TH, TT. Only one of them, HH, is two heads.