Introduction to Probability

Learning goals

  • Place a probability on the scale from 00 to 11, as a fraction, decimal, or percent
  • Find a probability by dividing favorable outcomes by equally likely total outcomes, after checking that assumption holds
  • Apply the complement rule, since all outcome probabilities add to one
  • Distinguish experimental probability from theoretical, and say how more independent trials bring them closer
  • Count a two-stage experiment's outcomes, then apply the same formula

A number from 0 to 1

Probability is measured on a fixed scale from 00 to 11. A theoretical probability of 00 means the event is impossible, and a theoretical probability of 11 means it is certain; getting zero successes in a finite experiment does not prove an event is impossible, since chance alone can produce a run with none. Everything else falls in between, and the larger the number, the more likely the event. A probability of 12\frac{1}{2} sits exactly in the middle and means the event is as likely to happen as not, the “fifty-fifty” of everyday speech.

Because every probability in this chapter is a ratio of whole numbers, a fraction of equally likely outcomes, you can write it three equivalent ways. The three forms are a fraction, a decimal, and a percent from 0%0\% to 100%100\%. A one-in-four chance is 14\frac{1}{4}, which is 0.250.25, which is 25%25\%, all naming the same spot on the scale. Use whichever form is clearest for the problem. Be ready to convert between them, since these are the fractions, decimals, and percents from earlier chapters doing familiar work.

The theoretical probability scale from 0 to 1A horizontal scale from 0 (impossible) through one-half (even chance) to 1 (certain) for theoretical probability, each point shown as a fraction, a decimal, and a percent.01/41/23/4100.250.50.7510%25%50%75%100%ImpossibleUnlikelyEven chanceLikelyCertain
The theoretical probability scale runs from 0 (impossible) to 1 (certain). The same value can be written as a fraction, a decimal, or a percent, so an even chance is one-half, 0.5, or 50 percent.

No probability is ever negative or larger than 11, because an event cannot happen less often than never or more often than always. If a calculation hands you 1.31.3 or −0.2-0.2, you have made a mistake somewhere.

Outcomes, and what equally likely means

To compute a probability we have to be careful about the setup. An experiment is any action with an uncertain result, such as flipping a coin or rolling a die. Each possible result is an outcome, and the full list of outcomes is the sample space. An event is the part of the sample space you care about, a chosen collection of outcomes. Rolling a die has the sample space {1,2,3,4,5,6}\{1, 2, 3, 4, 5, 6\}, and “roll an even number” is the event {2,4,6}\{2, 4, 6\}.

The six faces of a fair die, with the even faces shadedThe sample space of a die: faces 1 through 6, with 2, 4, and 6 shaded as the event “roll an even number”.
A fair die has six equally likely faces. The three shaded faces are the even numbers 2, 4, and 6, so the event 'roll an even number' covers 3 of the 6 outcomes.

The method of this lesson rests on one assumption: that the outcomes are equally likely, meaning each one has the same chance as any other. A fair coin, a fair die, a spinner with equal sectors, and a marble drawn at random are the standard ways of promising this. The words fair, equal, and at random are not decoration; they are what let you find a probability just by counting.

Equally likely is an assumption you must check, not something that is always true. Tomorrow’s weather is either rain or no rain, but those two outcomes are usually not equally likely. So the chance of rain is not automatically 12\frac{1}{2}. A bent coin might favor one side, and a spinner whose sectors have different sizes lands more often on the big ones. When the outcomes are not equally likely, counting them is not enough, and you need data or extra information instead.

The probability of an event

Now the payoff. Take the coin flip from the opening: heads or tails, 22 equally likely outcomes. Since they are equally likely, they split the one whole chance into 22 equal shares, so each outcome, including tails, gets 12\frac{1}{2}. That is exactly the “fifty-fifty” you already know.

The same reasoning works for any number of equally likely outcomes, not just 22. Suppose an experiment has NN equally likely outcomes. Something must happen, so the chances of all NN outcomes together make up the whole, named 11 (certain). Since the outcomes are equal, that one whole splits into NN equal shares, so each single outcome has probability 1N\frac{1}{N}.

An event is just some of those outcomes. If the event contains kk of them, it collects kk of the equal shares, so its probability is k×1N=kNk \times \frac{1}{N} = \frac{k}{N}. That is the central formula of the lesson:

P(event)=number of favorable outcomestotal number of equally likely outcomes.P(\text{event}) = \frac{\text{number of favorable outcomes}}{\text{total number of equally likely outcomes}}.

“Favorable” is only a name for the outcomes in your event; it does not mean good, just counted. The denominator is the size of the whole sample space, which is exactly the total you practiced counting in the previous lesson.

Worked example 1 The probability of tails

A fair coin is flipped once. What is the probability that it lands tails?

A flip has two equally likely outcomes, heads and tails, so the sample space has 22 outcomes. Exactly one of them, tails, is favorable.

P(tails)=12=0.5=50%.P(\text{tails}) = \frac{1}{2} = 0.5 = 50\%.

The word “fair” is what lets us count this way. It promises the two sides are equally likely, so each one takes half of the whole.

Worked example 2 The probability of an even roll

A fair die is rolled once. What is the probability of rolling an even number?

The sample space is the six faces {1,2,3,4,5,6}\{1, 2, 3, 4, 5, 6\}, all equally likely, so the total is 66. The event “even” is the set {2,4,6}\{2, 4, 6\}, the three faces shaded in the die figure above, so there are 33 favorable outcomes.

P(even)=36=12=0.5=50%.P(\text{even}) = \frac{3}{6} = \frac{1}{2} = 0.5 = 50\%.

For a final answer, simplify the fraction when you can. Here 36\frac{3}{6} and 12\frac{1}{2} are the same number, and 12\frac{1}{2} is the simplest way to write it, though keeping the unreduced form is what shows the favorable and total counts.

Check your understanding

A fair die is rolled once. What is the probability of rolling a 33?

Answer choices

More single-event probabilities

The formula does not care whether the experiment is a coin, a die, a spinner, or a bag of marbles. The work is always the same: count the favorable outcomes, count the total equally likely outcomes, and divide. Here it is on a spinner and on a random draw.

A fair spinner with 8 equal sectors, 3 shadedA circle cut into 8 equal sectors, with 3 of them shaded, illustrating a probability of 3 out of 8.
A fair spinner with 8 equal sectors. Because the sectors are equal, each is equally likely. Three sectors are shaded, so the probability of landing on a shaded sector is 3 out of 8.

Worked example 3 Reading a probability off a spinner

The spinner above has 88 equal sectors, and 33 of them are shaded. If you spin it once, what is the probability of landing on a shaded sector?

Because the sectors are equal, the 88 sectors are equally likely, so the sample space has 88 outcomes. The shaded sectors are the 33 favorable ones.

P(shaded)=38=0.375=37.5%.P(\text{shaded}) = \frac{3}{8} = 0.375 = 37.5\%.

Equal sectors are what make this work. If the shaded part were one wide sector and the rest were narrow, the outcomes would not be equally likely. In that case you could not just count sectors.

An experiment with exactly 100100 equally likely outcomes is worth a closer look. The reason is that the counting and the three ways of writing a probability all show up in one picture. Picture a drum of 100100 raffle tickets, mixed and drawn at random. Each square below is one ticket, and shading a square marks it as a favorable outcome. So the count of shaded squares is the numerator, while the 100100 squares are always the denominator.

Counting favorable outcomes out of a hundred

25 squares of 100 outcomes are favorable. So the probability is 25/100, or 1/4 in lowest terms. As a percent, 25%. As a decimal, 0.25. One hundred equal squares arranged 10 across and 10 down, filling from the top left. Use the controls below the figure to change how many are shaded.
Squares shaded

25 squares of 100 outcomes are favorable. So the probability is 25/100, or 1/4 in lowest terms. As a percent, 25%. As a decimal, 0.25.

One hundred equally likely outcomes drawn as a hundred-square. Shade the outcomes that count as favorable, and read the probability as a count, a percent, a fraction, and a decimal.

Give yourself some targets. Shade 2525 squares and one probability is reported three ways at once, as 25100\frac{25}{100}, as 0.250.25, and as 25%25\%, all naming the one-in-four chance. Shade none and you are at 00, the impossible end of the scale; shade all 100100 and you are at 11, certain. Then find the shading where the favorable squares and the ones left unshaded are equal in number, 5050 and 5050. At that split each event, favorable or not, has probability 12\frac{1}{2}, the even chance in the middle.

Worked example 4 Drawing a marble at random

A bag holds 33 red, 55 blue, and 22 green marbles, all the same size and weight. You draw one marble without looking. What is the probability it is red?

The marbles are all the same size and weight and the bag is well mixed, so drawing one without looking makes every marble equally likely to be the one you get. The total is the whole bag. Count it with the addition principle from the last lesson:

3+5+2=10 marbles.3 + 5 + 2 = 10 \text{ marbles}.

The favorable outcomes are the 33 red marbles, so

P(red)=310=0.3=30%.P(\text{red}) = \frac{3}{10} = 0.3 = 30\%.

The same draw gives P(blue)=510=12P(\text{blue}) = \frac{5}{10} = \frac{1}{2} and P(green)=210=15P(\text{green}) = \frac{2}{10} = \frac{1}{5}, and notice the three probabilities add to 310+510+210=1\frac{3}{10} + \frac{5}{10} + \frac{2}{10} = 1, the whole bag.

Why the probabilities add to one

That last observation, that the probabilities of all the outcomes add to 11, is not a coincidence of that bag. It is always true, and it is worth a careful statement, because it is also the engine behind a shortcut called the complement rule.

The outcome probabilities add to 11, and P(not A)=1−P(A)P(\text{not } A) = 1 - P(A)#

Take an experiment with NN equally likely outcomes, so each outcome has probability 1N\frac{1}{N}. Add up the probabilities of all NN of them. There are NN shares, each equal to 1N\frac{1}{N}, so the total is

1N+1N+⋯+1N⏟N terms=N×1N=1.\underbrace{\frac{1}{N} + \frac{1}{N} + \cdots + \frac{1}{N}}_{N \text{ terms}} = N \times \frac{1}{N} = 1.

The whole sample space has probability 11, which is just the statement that something is certain to happen.

Now take any event AA built from kk of these outcomes, so P(A)=kNP(A) = \frac{k}{N}. Its complement, written “not AA”, is everything else in the sample space, the other N−kN - k outcomes, so P(not A)=N−kNP(\text{not } A) = \frac{N - k}{N}. Adding the two,

P(A)+P(not A)=kN+N−kN=NN=1,P(A) + P(\text{not } A) = \frac{k}{N} + \frac{N - k}{N} = \frac{N}{N} = 1,

because every outcome is either in AA or not in AA, never both and never neither. Subtracting P(A)P(A) from each side gives the complement rule:

P(not A)=1−P(A).P(\text{not } A) = 1 - P(A).

The complement rule

The complement rule earns its own line because it often saves work. When the outcomes you want are awkward to count but the leftover outcomes are easy, count the leftovers and subtract from 11. Problems that ask for “at least one” of something are the classic case, since “at least one” is the complement of “none”.

Worked example 5 Using the complement rule

Using the same bag of 1010 marbles (33 red, 55 blue, 22 green), what is the probability of drawing a marble that is not red?

You could count the favorable outcomes directly. Not red means blue or green, which is 5+2=75 + 2 = 7 marbles, so 710\frac{7}{10}. The complement rule reaches the same answer faster, starting from P(red)=310P(\text{red}) = \frac{3}{10}:

P(not red)=1−P(red)=1−310=710=0.7=70%.P(\text{not red}) = 1 - P(\text{red}) = 1 - \frac{3}{10} = \frac{7}{10} = 0.7 = 70\%.

Both routes agree, which is the whole point. A marble is either red or not red, so those two probabilities must fill up the one whole.

Here is where the shortcut actually saves work. Flip two fair coins together, so one cannot affect the other, and ask for the probability of at least one head. Counting directly means listing every way to get one head or two: HH, HT, TH, that is 33 of the 44 equally likely outcomes. The complement is faster, because “at least one head” fails only for TT, no heads at all:

P(at least one head)=1−P(no heads)=1−14=34.P(\text{at least one head}) = 1 - P(\text{no heads}) = 1 - \frac{1}{4} = \frac{3}{4}.

Both answers agree, 34\frac{3}{4}, but the complement found it by ruling out just one outcome instead of counting three.

Check your understanding

A fair die is rolled once. What is the probability of not rolling a 33?

Answer choices

Theoretical and experimental probability

Everything so far has been theoretical probability: you reason about equally likely outcomes and compute a fraction before ever running the experiment. There is a second kind. Experimental probability, also called relative frequency, comes from actually doing the experiment many times and recording what happens:

Pexp(event)=number of times the event happenednumber of trials.P_{\text{exp}}(\text{event}) = \frac{\text{number of times the event happened}}{\text{number of trials}}.

If you flip a fair coin 2020 times and get 1111 heads, the experimental probability of heads is 1120=0.55\frac{11}{20} = 0.55. That happens even though the theoretical value is 12\frac{1}{2}. The two are different numbers, and that is expected, because a short run of flips rarely splits exactly evenly.

Here is the key link between them. This works when the trials do not affect each other, such as flipping the same coin over and over: one more trial does not have to move the experimental probability closer to the theoretical value, since a single flip can push the running total either way and short runs bounce around. What is true over the long run: as you repeat more and more independent trials, that running proportion becomes increasingly likely to sit close to the theoretical probability. Twenty flips might give 55%55\% heads, but ten thousand flips will almost certainly land very near 50%50\%. This pattern, that relative frequency settles near the true probability as the number of trials grows, is called the law of large numbers. It is why an insurer can predict the long-run rate of claims closely even though any single policy is uncertain. It is also the bridge that makes a theoretical 12\frac{1}{2} something you can actually watch happen.

Worked example 6 Experimental probability of heads

A fair coin is flipped 5050 times and lands heads 2727 times. What is the experimental probability of heads from this data, and how does it compare with the theoretical value?

Experimental probability divides the number of times the event happened by the number of trials.

Pexp(heads)=2750=0.54=54%.P_{\text{exp}}(\text{heads}) = \frac{27}{50} = 0.54 = 54\%.

The theoretical probability for a fair coin is 12=50%\frac{1}{2} = 50\%. The experimental value is close but not exactly equal, which is normal for only 5050 flips. Flip the same coin independently ten thousand times and the fraction of heads would almost certainly sit much nearer to 50%50\%.

A gentle step to two-stage experiments

Single events are the heart of probability, but the same formula stretches to experiments built from two stages, as long as you keep counting carefully. Rolling two dice, flipping two coins, or spinning twice are all sequences of stages, so the multiplication principle from the last lesson counts their total outcomes: each stage still offers the same number of results next, no matter what came before. The probability formula needs two things more than a correct count, though. Each stage’s own outcomes must already be equally likely on their own, such as a fair die or a fair coin, and the stages must not affect each other’s results, called independent trials, such as rolling two dice together or flipping two coins together. When both hold, as they do for every two-stage experiment in this lesson, the combined outcomes are equally likely too. The probability formula does not change at all: it is still favorable outcomes over total equally likely outcomes. Only the counting gets a little bigger.

The 36 outcomes of two dice, with the 6 doubles shadedA 6 by 6 grid showing 36 equally likely outcomes; the diagonal cells, where both dice show the same number, are shaded.Die 2Die 1123456123456
Rolling two fair dice together, so one cannot affect the other, gives 6 times 6 = 36 equally likely outcomes, one per cell. The 6 shaded cells down the diagonal are the doubles, where both dice match, so the probability of a double is 6 out of 36.

Worked example 7 Rolling a double with two dice

Two fair dice, one red and one blue, are rolled together. What is the probability that they show the same number, a double?

First count the total. The red die has 66 results and the blue die has 66, and one die’s result does not change the other’s. The roll is a sequence of two independent stages, so the multiplication principle gives

6×6=366 \times 6 = 36

equally likely outcomes, one per cell of the grid above. The doubles are (1,1),(2,2),(3,3),(4,4),(5,5),(6,6)(1,1), (2,2), (3,3), (4,4), (5,5), (6,6), which is 66 favorable outcomes, the shaded diagonal.

P(double)=636=16≈0.167≈16.7%.P(\text{double}) = \frac{6}{36} = \frac{1}{6} \approx 0.167 \approx 16.7\%.

We found the denominator by counting, exactly the tool from the previous lesson. The 3636 pairs stay equally likely because each die is fair on its own and the two rolls do not affect each other.

Check your understanding

Two fair coins are flipped together, so one does not affect the other. What is the probability that both land heads?

Answer choices

Common mistakes

Practice

Multiple Choice Questions (MCQ)

Progressively harder sets of questions. Each opens on its own page.

Core practice

Practice problems at the level of the course, to be worked out on paper. Hints one at a time, then the answer or the full worked solution, with your progress kept in this browser.

Core practice Work it out on paper 10 problems Start →
More practice (optional)

Extra sets, as hard as the Challenge set. Each one opens on its own page.

More resources (optional)

Other explanations of this lesson, if you want a second take.

A bit of history (optional)

People have gambled with dice for thousands of years. For most of that time, hardly anyone tried to pin the chances down with a number. Players went on hunches instead.

One of the earliest careful counts of them that survives is by Gerolamo Cardano, an Italian doctor of the fifteen hundreds who also wrote a short book on games of chance. In it he did what you have just done. He treated the six faces of a fair die as six equal cases. Then he read the chance of an event as the cases that win it, set against every case there is.

His book had almost no effect, because it went unread. Nobody printed it until 1663, nearly ninety years after he died.

The word his method turns on is equal. Cardano knew it, and the same book warns you about crooked dice. Every problem here still leans on that one condition. A fair coin, a fair die, equal sectors, a marble drawn from a well-mixed bag: take the promise away, and counting alone cannot tell you the probability.