Compute g∘f, reading it right to left with f acting first
Decide whether g∘f equals f∘g for a given pair
Find the domain of g∘f from the range of f
Decompose a function as g∘f, knowing the split is a choice
Write f2 for f∘f, never for the squared output
Reading a composition right to left
Give the chaining its symbol. The composition of f and g, written g∘f, is the function
(g∘f)(x)=g(f(x)).
Check it against the numbers above: with f(x)=x+1 and g(x)=x2,
(g∘f)(2)=g(f(2))=g(3)=9,
exactly the 2→3→9 chain from the opening. You feed x to f, and then feed the result f(x) to
g; the circle is the operation sign, exactly as + is the sign for addition.
The one thing to fix in your reading, because it causes more confusion than anything else in this lesson,
is the order. In g∘f the function on the right, f, acts first, and the function on the
left, g, acts second. The notation runs right to left, opposite to the way you read a sentence. The
reason is built into the definition: g(f(x)) has f in the inner parentheses, and the inner operation
always happens before the outer one. So g∘f means “do f, then g,” never “do g, then f.”
The composite g of f as a two-machine pipeline. The input x enters f and comes out as f(x); that value enters g and comes out as g(f(x)). The machines run left to right, but we write g of f because g acts on the output of f, so the name reads right to left and f goes first.
Worked example 1Both orders of the same two functions
Let f(x)=x+1 and g(x)=x2. Compute g∘f and f∘g, keeping track of which function
acts first.
For g∘f, the inner function is f, so substitute f(x)=x+1 into g:
(g∘f)(x)=g(f(x))=(x+1)2=x2+2x+1.
For f∘g, the inner function is now g, so substitute g(x)=x2 into f:
(f∘g)(x)=f(g(x))=x2+1.
The two results, x2+2x+1 and x2+1, are different functions. Swapping the order changed the
answer, which is the first sign that composition, unlike addition, does not always let you reverse the two
inputs.
Check your understanding
For f(x)=2x and g(x)=x−3, which expression is (g∘f)(x)?
In g∘f the right function f acts first, then g acts on the result.
(g∘f)(x)=g(f(x))=g(2x)=2x−3
The trap 2x−6 is the other order, (f∘g)(x)=f(x−3)=2(x−3).
Order matters, but not always
The one law that addition has and composition often lacks is commutativity: for numbers, a+b=b+a
always. For functions, g∘f and f∘g are usually different, and Worked Example 1 already
showed why: running the same two machines in the opposite order gave x2+2x+1 instead of x2+1.
These differ by 2x, which is zero only at x=0.
But “usually different” is not “never equal,” and the difference is worth stating exactly, because
overshooting it is the classic error. Some pairs genuinely commute. Take the two translations f(x)=x+1 and g(x)=x+2. Then
(f∘g)(x)=(x+2)+1=x+3=(g∘f)(x),
and the same works for any two translations f(x)=x+a and g(x)=x+b, since both orders give x+a+b and addition of the shifts does not care about order. So the correct statement names both cases:
composition is not always commutative, but it is not never commutative either; commuting is the
exception, not the rule.
Composition has two more things in common with addition, each worth a name. It is associative: for any
three functions, grouping does not matter, (h∘g)∘f=h∘(g∘f), so you can write
h∘g∘f with no parentheses.
And it has an identity: a function id(x)=x that changes nothing when composed with
f, the way 0 changes nothing under addition. Which set the identity is defined on matters, so when we
need to be specific we write idS for the identity on the set S. Throughout this lesson,
when a function’s codomain is not stated separately, take it to be that function’s range, the values it
actually produces. Every domain and identity claim below relies on that convention.
Check your understanding
For which pair do the two compositions agree, (f∘g)(x)=(g∘f)(x) for every x?
Two translations always commute, because you can add the two shifts in either order.
(f∘g)(x)=(x−2)+5=x+3=(g∘f)(x)
Each other pair fails: for f(x)=x2 and g(x)=x+1, (g∘f)(x)=x2+1 but (f∘g)(x)=(x+1)2; for f(x)=2x and g(x)=x+1, the orders give 2x+2 and 2x+1; for f(x)=x2 and g(x)=2x, they give 4x2 and 2x2.
The domain of a composition
Here is the idea the substitution procedure hides, and it is the most important thing in the lesson. When
you build (g∘f)(x)=g(f(x)), two things must be true for the answer to exist: f must accept x,
and g must accept the number f(x) that comes out.
By definition (g∘f)(x)=g(f(x)), so this value exists exactly when both steps inside it are legal.
First f(x) must be defined, which requires x∈domf. Then g is applied to the
number f(x), which is legal exactly when f(x)∈domg. So
dom(g∘f)={x∈domf:f(x)∈domg}.
This is notdomf∩domg. The second condition constrains f(x),
an output of f, so testing it needs the range of f, not its domain. A value of x can lie in both
domf and domg yet be excluded, because f(x) escapes
domg. Or it can lie outside domg and still be fine, as long as x∈domf and f(x)∈domg.
∎
Worked example 2A square root inside a reciprocal
Let f(x)=x and g(x)=x−21. Find a formula for g∘f and its domain.
The formula is a direct substitution:
(g∘f)(x)=g(x)=x−21.
Now the domain. First f must accept x, and x needs x≥0. Second g must accept the output
x, and g rejects only the input 2, so we need x=2. Since x=2 exactly
when x=4,
dom(g∘f)={x≥0 and x=4}.
Compare that with the wrong shortcut. The intersection domf∩domg
would keep x=4 and throw out x=2, the value g forbids. The truth is the reverse: x=2 is fine,
because 2=2, while x=4 is excluded, because 4=2 is the one input g cannot
take. The excluded input is found through the range of f, not by intersecting the two domains.
The domain of g of f is the part of dom f whose image lands where g is defined. Here f(x) = sqrt(x) sends the inputs x greater than or equal to 0 (top, and 4 is one of them) into the values g accepts (bottom, everything except 2). The inputs 0 and 9 map to 0 and 3 and survive; the input 4 is in dom f but maps to the barred value 2 (dashed), so 4 is removed from dom(g of f), not from dom f. Hence dom(g of f) is x greater than or equal to 0 with x not equal to 4.
Check your understanding
Let f(x)=x and g(x)=x−31. What is the domain of g∘f?
You need x∈domf, so x≥0, and also f(x)∈domg, so x=3.
x=3⟺x=9
So remove x=9, leaving x≥0 with x=9. The tempting x=3 is the intersection domf∩domg, the wrong rule; x=3 is allowed since 3=3.
Worked example 3Same formula, smaller domain
Let f(x)=x and g(x)=x2. Find g∘f, and check its domain before deciding whether it
is an identity function.
Substituting gives a formula that simplifies all the way down:
(g∘f)(x)=g(x)=(x)2=x.
The formula alone looks like the identity. Check the domain before trusting that. First f needs x≥0; then g accepts every real number that comes out, so nothing more is excluded. So
dom(g∘f)={x≥0}.
So g∘f is the rule x↦x with domain x≥0: that is id[0,∞),
the identity on [0,∞), but not idR, the identity on all of
R. The two share a formula but not a domain, so by the equality test they are different
functions. The simplified x threw away the record of where x refused to run; only the
domain remembers it.
The other order behaves differently still. Here (f∘g)(x)=x2=∣x∣, defined for
every real x, so dom(f∘g) is all of R while dom(g∘f) is only x≥0, a reminder that dom(f∘g) and dom(g∘f)
are in general different sets.
Decomposing a function, and why the answer is not unique
Composition also runs backward as a question. Given a function h, can you write it as g∘f for
simpler pieces f and g? This is decomposition, and it lets you see a complicated rule as a short
pipeline of easy steps.
Take h(x)=3x+1. The natural split peels off the outermost operation, the square root, and
calls the inside a separate step: let f(x)=3x+1 and g(x)=x. Then
(g∘f)(x)=g(3x+1)=3x+1=h(x).
But this is not the only way. A different split absorbs the 3 elsewhere: f(x)=3x and
g(x)=x+1 also gives g(f(x))=3x+1=h(x). Both are correct, since each composes to
h. So “decompose h” has more than one right answer. Among the valid splits there is usually a most
natural one, peel off the outermost operation, but no split is the unique correct answer.
Worked example 4Decomposing a cube
Write h(x)=(2x−5)3 as a composition g∘f.
Peel off the outermost operation, the cube. Let f(x)=2x−5 be the inside and g(x)=x3 be the cube:
(g∘f)(x)=g(2x−5)=(2x−5)3=h(x).
Check by recomposing: substitute f‘s formula into g and confirm it rebuilds h exactly, which it does.
As with the square-root example, this is one valid decomposition among several; the checkpoint below asks
you to find another.
Check your understanding
The lesson decomposes h(x)=(2x−5)3 as f(x)=2x−5, g(x)=x3. Which pair gives a genuinely different valid decomposition, h=g∘f?
Check by composing: g(f(x))=g(2x)=(2x−5)3=h(x), so this pair works.
The second pair is the one already shown in the lesson, not a different one. The third gives (3x−5)3=h(x), and the fourth gives (2x−10)3+5=h(x): both compose to the wrong function.
Composing a function with itself
Because composition is associative, composing a function with itself repeatedly is unambiguous, and it has
its own notation. Write f2=f∘f, the small superscript counting how many times f is applied.
Read f2(x)=f(f(x)): apply f, then apply f again to the result.
Here is the notation warning. That superscript clashes with the other thing a superscript can mean,
squaring the output, (f(x))2=f(x)⋅f(x). The two are different functions. For f(x)=2x−1,
f2(x)=f(f(x))=2(2x−1)−1=4x−3,
while
(f(x))2=(2x−1)2=4x2−4x+1,
and these are different functions: one is linear, the other is not. Which meaning f2 carries is a matter
of convention, and this lesson uses f2 for the composition f∘f; when we mean the square of the
output we write (f(x))2. One more warning for the next lesson: there f−1 will mean the inverse
of f under composition, never f1. In this lesson’s convention, a superscript on a function’s
name is about composition, not multiplication.
Check your understanding
For f(x)=3x+2, which expression equals f2(x), meaning f∘f?
f2(x)=f(f(x))=f(3x+2)=3(3x+2)+2=9x+8.
The trap 9x2+12x+4 is (f(x))2=(3x+2)2, the squared output, not the composition f∘f.
Transformations were compositions all along
The previous lesson wrote every transformed graph as y=af(b(x−h))+k, split into an
inside that acts on the input before f and an outside that acts on f‘s output afterward. Name
those two pieces as functions, an input map and an output map. The transformed rule is then exactly a
composition of three functions: input map, then f, then output map. That is also why “stretch then
shift” gave a different graph from “shift then stretch”: the two output maps did not commute, the same
reason g∘f=f∘g in general.
One more distinction is worth a sentence. Composition chains two functions, feeding one’s output into
the other’s input; adding, subtracting, multiplying, or dividing two functions instead combines their
outputs at the same input, (f+g)(x)=f(x)+g(x). They are different operations that happen to both
start from two functions.
Common mistakes
Practice
Multiple Choice Questions (MCQ)
Progressively harder sets of questions. Each opens on its own page.
Practice problems at the level of the course, to be worked out on paper. Hints one at a
time, then the answer or the full worked solution, with your progress kept in this browser.
Let f, g, and h be functions, and consider the two ways to group their composition, (h∘g)∘f and h∘(g∘f). Two functions are equal exactly when they have the same domain,
codomain, and values, the equality test from Function Notation, so we check the values, the domains, and
the codomains.
For the values, apply each grouping to an input x and unfold the definition one step at a time. The left
grouping gives
((h∘g)∘f)(x)=(h∘g)(f(x))=h(g(f(x))),
and the right grouping gives
(h∘(g∘f))(x)=h((g∘f)(x))=h(g(f(x))).
Both collapse to the single expression h(g(f(x))), so the two functions agree wherever they are defined.
For the domains, read off what each grouping requires. Building h(g(f(x))) is legal exactly when x∈domf so that f(x) exists, then f(x)∈domg so that g(f(x)) exists,
then g(f(x))∈domh so that the outer h applies. These three conditions are the same
for both groupings, so the two functions share the same domain. Both also share the same codomain, whatever
codomain we declared for h, since both groupings end by applying h. By the equality test they are the
same function.
∎
Because the grouping never matters, we drop the parentheses and write h∘g∘f.
The identity function, defined carefully
The main lesson uses the identity function informally. Here is why composing with it changes nothing, done
carefully: the equality test demands the same domain, codomain, and values, not just a matching formula.
Let f:A→B be a function. For any set S, let idS(x)=x be the identity function
on S, the function with domain S, codomain S, and rule “return the input unchanged.” We show
f∘idA=f and idB∘f=f, taking the identity on the correct
set each time.
For every x∈A,
(f∘idA)(x)=f(idA(x))=f(x),
and f∘idA has domain A and codomain B, matching f exactly. So f∘idA=f.
Likewise, for every x∈A,
(idB∘f)(x)=idB(f(x))=f(x),
since f(x)∈B, the domain of idB, so the composition is legal for every x∈A.
And idB∘f has domain A and codomain B, again matching f. So
idB∘f=f.
The two sides are not symmetric. On the left, the identity’s codomain is forced: idC∘f=f only when C=B exactly, since a composite’s codomain always matches its outer function’s, and
that has to equal B. On the right, f∘idA=f, and so does f∘idT for any set T that contains A. The domain of that composite is {x∈T:x∈A}=A either way, since f‘s own domain restriction already confines things to A. So
idA is simply the smallest, most natural choice on that side, not the only one that
works; the care that matters is on the left.
∎
Composition also has a name for what these two facts give it, together with the fact that it is not
generally commutative. When f, g, and h all map one fixed set X back into itself, composing any two
of them gives another function X→X, idX serves as the identity, and composition of
such functions is associative. Associative, with an identity, but not generally commutative and not
guaranteed an inverse for every element: that structure is what algebraists call a monoid. The
self-maps of X under composition form one such monoid. The missing piece, an inverse that undoes a
function, is what the next lesson supplies for the functions that have one.
Writing a transformation as a composition
The main lesson states the connection in one paragraph. Here it is in full, with the input and output maps
named and a worked example.
Write B(x)=b(x−h) for the input map and A(y)=ay+k for the output map, so B acts on the input
before f and A acts on f‘s output afterward. Then
y=A(f(B(x)))=(A∘f∘B)(x).
Reading right to left, B acts first, then f, then A, exactly the story the previous lesson told: the
inside happens before f, and the outside after. Its rule that input maps “run backward because you solve
for x” is also explained: in A∘f∘B the rightmost map B acts before f. So to find which
x produces a given input to f, you must undo B, that is, solve B(x)=u for x.
Worked example 5A transformation read as a composition
Write y=2f(3x−6)+1 in the form A∘f∘B, and identify A and B.
Factor the inside so it matches b(x−h), and read the outside directly:
3x−6=3(x−2),soB(x)=3(x−2),A(y)=2y+1.
Check that composing them rebuilds the rule:
(A∘f∘B)(x)=A(f(3(x−2)))=2f(3x−6)+1.
So the transformation is the composition A∘f∘B with B(x)=3(x−2) acting on the input and
A(y)=2y+1 acting on the output.
Iterating further: f3 and a general pattern
The main lesson computes f2 for f(x)=2x−1. Here is f3, and the general pattern it starts.
Apply f once more, now to f2(x)=4x−3:
f3(x)=f(f2(x))=2(4x−3)−1=8x−7.
A pattern is visible: each application doubles the coefficient of x and adjusts the constant to match,
giving fn(x)=2nx−(2n−1), which reads 2x−1, 4x−3, 8x−7 for n=1,2,3. The pattern
continues because each new application does the same two things to whatever came before it: double the
result and subtract 1. So the coefficient of x doubles again, and the constant term absorbs one more
copy of that same rule. None of these equal the output square (f(x))2=4x2−4x+1, which is not even
linear: iterating a function and squaring its value are unrelated operations that happen to share a symbol.
Combining functions without chaining them
Composition is not the only way to build a new function from two others. You can add, subtract, multiply,
or divide two functions by combining their outputs at each input:
These evaluate both functions at the same x and combine the two numbers, so both functions see the
original input. Composition is different in kind: it chains them, feeding the output of one as the
input of the other, so only f sees x and g sees f(x). The domain rules differ to match. A sum or
product is defined where both functions are, domf∩domg, the very
intersection that is the wrong answer for composition. A quotient gf additionally removes the
inputs where g(x)=0.
A bit of history (optional)
For most of the nineteenth century, mathematicians who studied groups only ever meant one thing: groups of
permutations, the different ways to rearrange a list. Combining two rearrangements meant doing one after
the other, which is composition and nothing else. Nobody thought of it as an example of something more
general, because nothing else looked like it.
Arthur Cayley, an English mathematician who earned his living as a lawyer, changed that in 1854. His short
papers on groups kept the laws an operation obeys and let go of the objects obeying them. In place of a
recipe for shuffling a list, he printed a table of results. The same laws turned out to describe things
with nothing else in common, such as square blocks of numbers under their own rule for multiplying.
This lesson takes the identical step with composition, a procedure you already knew. It names the laws
that procedure obeys: associative, carrying an identity, and, among functions from one set to itself, not
generally commutative.