A composition changes in stages. The inner function transforms the original input, and the outer function responds to that transformed value. The chain rule connects those stages without pretending that either one acts alone. It is the central differentiation rule for nested expressions, implicit relationships, and rates that pass through several variables. This lesson builds the rule from local change so that its factors have meaning rather than appearing as a memorized pattern.
The rule is also a model of causal bookkeeping. If time changes position and position changes energy, then time changes energy through both links. Each derivative measures the sensitivity of one quantity to the quantity immediately before it. Multiplying those sensitivities produces the overall rate. That interpretation will help you choose factors, track units, and detect missing layers.
Learning goals and prerequisite check
By the end of this lesson, you should be able to recognize a composition before differentiating it. You should be able to name its inner and outer functions in a consistent order. You will derive the chain rule from a difference quotient and from local linearity. You will apply it to powers, radicals, exponentials, logarithms, and trigonometric compositions. You will also use it in implicit differentiation, inverse derivatives, and physical rate models.
Recall that a composition is written . The small circle means “apply first and then apply ,” not multiplication. For example, if and , then . The letter is a temporary input name for the outer function. Replacing by reveals the nested structure.
You should also recall the derivative as a limit of average rates. At , the derivative is when this limit exists. The symbol denotes a nonzero input change that approaches zero. The numerator is the corresponding output change. The quotient therefore asks for output change per unit input change at progressively smaller scales.
Recognize the layers before using a rule
An expression needs the chain rule when one nontrivial function is used as the input of another. In , the power of five is the outer operation and is the inner expression. In , sine is outer and squaring is inner. In , the exponential is outer and cosine is inner. The visible order of symbols often runs opposite to the order in which inputs are processed.
A reliable diagnostic is to ask what single quantity you would replace with a box. Replacing by turns into . The boxed quantity is the inner function, and the remaining operation is the outer function. This test works for radicals, denominators, logarithms, and trigonometric functions. It also prevents you from mistaking an ordinary sum such as for a composition.
Some expressions contain several nested layers. For , the input is squared, sine acts on that result, and the exponential acts last. A dependency list is . Every arrow whose operation changes with its input contributes a derivative factor. Writing this list before calculating is often faster than repairing a missing factor afterward.
State the chain rule precisely
Suppose and . If is differentiable at and is differentiable at , then . In function notation, the same statement is . The factor is the derivative of the outer function evaluated at the unchanged inner output. The factor is the derivative of the inner function.
The hypotheses matter because a composition can fail to be differentiable at either layer. The inner function must have a well-defined local linear approximation at the chosen input. The outer function must have one at the inner output, not merely somewhere else in its domain. These conditions explain why checking domains before and after differentiation is necessary. A symbolic rule cannot manufacture differentiability where the original relationship lacks it.
Leibniz notation makes the dependency structure visible, but the symbols do not literally cancel like ordinary numbers. The expression is a product of two derivatives justified by a limit theorem. Its middle symbol suggests the matching intermediate variable. That suggestion is valuable for arranging factors and checking units. The proof, however, depends on differentiability rather than algebraic cancellation of infinitesimals.
Understand why the factors multiply
Near an input , differentiability lets us approximate the inner change by . The symbol means a finite change, so is the input increment and is the resulting inner increment. Near , the outer change is . Substituting the first approximation into the second gives . The coefficient of is therefore the overall derivative.
This argument is local linearity in action. The inner derivative converts a small change in into a small change in . The outer derivative then converts that change in into a small change in . Successive conversions multiply in the same way that successive scale factors multiply. The chain rule is therefore a composition rule for local linear models.
A limit derivation makes the same logic exact. Insert the factor into the composition’s difference quotient when the inner difference is nonzero. The quotient becomes an outer difference quotient multiplied by an inner difference quotient. As approaches zero, continuity from differentiability makes the intermediate input approach . The two limits become and , producing their product, with the zero-inner-change cases handled by the differentiability remainder form.
Apply the outer-then-inner procedure
Differentiate by preserving the inner expression. Let , so . The outer derivative is , while the inner derivative is . Multiplication gives . Simplifying constants yields .
Notice what remains unchanged during the outer step. Differentiating the fifth power lowers its exponent to four, but the complete inner expression stays inside the fourth power. Only after that step do we multiply by the derivative of the inside. Replacing the inside prematurely often destroys the composition’s structure. The verbal pattern is “derivative of the outside, evaluated at the inside, times derivative of the inside.” Say that pattern aloud while pointing to the matching factors until the structure becomes automatic.
Now differentiate . The fractional exponent is the outer operation, and is inner. The power rule gives , and the inner derivative is . Thus . This derivative is finite only where the radicand is positive, even though the original square root also exists at a zero radicand.
Differentiate exponential and logarithmic compositions
For , the exponential function is outer and is inner. Because the derivative of with respect to is , the outer factor is . The inner derivative is . Therefore . The original exponential factor remains because its rate of change equals its value with respect to its own input.
For a general positive base , . The symbol is the natural logarithm of the constant base and belongs to the outer derivative. For example, . The factor three comes from the inner function . Omitting either constant changes the local rate.
For , the logarithm rule gives the reciprocal of its input times the input’s derivative. Thus . The original real-valued logarithm requires . That inequality means . A correct derivative statement preserves this domain rather than reporting only a formula.
Differentiate trigonometric compositions
For , sine is outer and the square is inner. The outer derivative is cosine evaluated at the same inner expression, giving . The inner derivative is . Therefore . This standard formula assumes angles are measured in radians, because the familiar sine derivative is a radian-based limit result.
For , the outer derivative contributes a negative sine. Keeping the inside unchanged gives . The inner derivative contributes the factor four. Hence . The negative sign and the inner factor arise from different layers and should be checked separately.
For , conventional notation means , not . The cube is outer, tangent is the next layer, and is the innermost input. Differentiation gives . The expression is defined only where . Parsing notation correctly is part of recognizing the composition.
Manage multiple layers systematically
Consider . Begin with the outer exponential, whose derivative reproduces . Next differentiate the sine layer to obtain . Finally differentiate the square layer to obtain . Multiplying gives . Each factor corresponds to one arrow in the dependency chain.
Intermediate variables make a long chain easier to audit. Set , , and . Then . Substitution produces and then restores and . This notation helps ensure that no variable is skipped or differentiated with respect to the wrong input.
The same method works for . The layers are square root, addition of one, squaring, logarithm, and the original input. Constant addition contributes a derivative of one and does not create a visible factor. The result is . Its real domain requires , and the denominator remains nonzero throughout that domain, including at where .
Track units through composed rates
Suppose position is measured in meters, time in seconds, and energy in joules. The sensitivity has units of joules per meter. The velocity has units of meters per second. Their product has units . These are the required units for the energy rate .
Consider a spherical balloon with . If radius depends on time, the chain rule gives . At and , the rate is . This equals . Cubic meters per second confirm that the result is a volume rate.
Units can expose a missing chain factor. The geometric derivative alone has square-meter units, not cubic meters per second. Multiplication by supplies meters per second and completes the dimensions. A numerically plausible answer with wrong dimensions is still structurally wrong. Treat dimensional analysis as an independent verification step.
Use the chain rule in implicit differentiation
An implicit equation relates variables without isolating one of them. For , regard as an unknown function of . Differentiating with respect to gives . The factor appears because differentiating requires the chain rule. Solving yields where .
The restriction belongs to the solved slope formula. At points and , the circle has vertical tangents, so a finite derivative with respect to does not exist. Differentiating with respect to instead gives information about . This behavior is geometric rather than an algebraic defect. Domain and tangent orientation must accompany the formula.
For , use both the product rule and chain rule. Differentiation gives . Collecting derivative terms produces . Therefore when . Naming which rule creates each term prevents the common omission of from the derivative of .
Derive derivatives of inverse functions
If has a differentiable local inverse, then . Differentiate both sides with respect to . The chain rule gives . Solving produces . The denominator must be nonzero at the corresponding original input.
Suppose and . Because , the inverse derivative at five is . The input five belongs to the inverse function, while two is the matching input of the original function. Confusing those coordinates is a frequent error. Thinking of inverse graphs as reflections across helps swap their roles correctly.
Local invertibility also matters. The function has no single inverse on all real numbers because two inputs can share one output. Restricting it to produces the inverse . At the original input zero, , so the reciprocal formula predicts no finite inverse derivative. Correspondingly, the square-root graph has a vertical tangent at zero.
Diagnose errors and verify results
The most common error is differentiating the outer function but forgetting the inner derivative. Writing misses the factor . Another error is replacing the inside with its derivative, producing . The correct outer step keeps the original inside intact. Only multiplication by its derivative follows afterward.
A second class of errors comes from misidentifying structure. A product such as needs the product rule, while needs the chain rule. A quotient can contain compositions and therefore require both quotient and chain rules. Parentheses and a dependency tree reveal which rule governs each layer. Name the outermost operation before performing any algebra.
Verify a result in three ways. First, compare units whenever variables represent measured quantities. Second, check a numerical difference quotient for a small safe value of . Third, inspect signs and rough magnitudes from the graph near . Agreement among symbolic, numerical, and graphical evidence is stronger than repeated symbolic manipulation alone.
Guided synthesis and connection forward
Differentiate . The logarithm is outer, so its derivative contributes the reciprocal . The inner derivative is . Thus . Because is always positive for real , both the original function and derivative are defined everywhere.
Now differentiate . The square is outer, cosine is the middle layer, and is inner. The factors are , , and . Their product is . The equivalent form follows from a double-angle identity, but the unsimplified form displays the chain more clearly.
For independent practice, differentiate , where defined, and . For each expression, draw or write the dependency chain before calculating. Label every derivative factor with the layer that created it. State the real domain and check one value numerically. The next applications lessons will use this same sensitivity logic in related rates, optimization models, and motion.
Sources and further study
The principal mathematical reference for this lesson is OpenStax Calculus, Volume 1. Its differentiation chapters develop the chain rule from composite functions. They also connect the rule to implicit differentiation and inverse functions. Compare its notation with the dependency notation used here. Translating between equivalent notations is a useful form of retrieval practice.
The AP Calculus AB course overview provides the curriculum alignment for these skills. It places composition analysis and the chain rule within the broader study of differentiation. Its framework also emphasizes interpreting derivatives in context. Use its topic sequence to identify which prerequisite rules need further review. Curriculum placement does not replace the mathematical hypotheses stated in this lesson.
When consulting another source, test each example by naming its layers before reading the solution. Then predict the number and kind of derivative factors that should appear. Check whether the source states relevant domain restrictions. Verify one worked result with a numerical difference quotient. Active comparison turns a reference into practice rather than passive rereading.