How to Actually Use the Chain Rule When It Gets Complicated
I keep seeing people struggle with this in grad courses, and honestly most of the struggle comes from trying to memorize formulas without visualizing the dependency structure first. Here's how I actually approach it when a problem has five or six nested layers. Start with the dependency diagram. Draw it before touching any symbols. If z depends on x and y, and x and y both depend on s and t, you're looking at two paths from z back to each independent variable. For each path, multiply along the edges. Then add the path products. Mathematically, dz/ds = z/x · dx/ds + z/y · dy/ds. That's the one you memorize. But the real version of the Chain Rule For Multivariable gets messier when your variables depend on each other too, not just through separate independent inputs. That's where most people lose points on exams and waste hours on homework.
The Tree Method (And Why You Need It)
Every time I encounter a problem with more than three layers, I build a dependency tree. Write z at the top. Branch down to whatever variables z directly depends on. From each of those, branch further down to whatever they depend on. Keep going until you hit independent variables. Count the branches. Each complete path from z to an independent variable gives you one product term in the final derivative. I remember a specific problem once where z was defined implicitly as a function of u and v, and u and v were both functions of x and y, but x and y themselves were coupled through a constraint equation g(x, y) = 0. Standard textbook examples don't cover this case cleanly. What I ended up doing was treating x as the dependent variable and y as independent (or vice versa), using implicit differentiation to find dx/dy from the constraint, and then applying the chain rule only after substituting that relationship in. Took about 45 minutes to set up, five minutes to differentiate once it was clean.
Common Pitfalls I See Repeatedly
Mixing up partial and total derivatives. This is the biggest one. z/x means "hold everything else constant." dz/ds means "account for every way s affects z." If you write a partial when you need a total, your answer will be wrong even if the algebra is perfect. The notation tells you what to do; respect it. Dropping a path in the tree. If you have three independent variables feeding into z through some intermediate layer, you need three terms in your final sum. Students routinely write two. Check your tree. Count your paths. Forgetting that intermediate variables can depend on each other. The clean version assumes x and y are independent of each other. They aren't always. If x depends on y, the simple two-path formula is incomplete. You need an extra term: z/x · x/y when you're differentiating with respect to y. This comes up in thermodynamics constantly, and textbooks gloss over it.
Get the Full Details

When You Need the Jacobian Instead
For systems with many variables, writing out individual partial products gets tedious and error-prone. The Jacobian matrix packages everything. If you have a vector of outputs y = f(x) where x is a vector of n inputs, the Jacobian J is an m×n matrix where J_ij = y_i/x_j. To chain two transformations together, you multiply their Jacobians. J(gf) = J(g) · J(f). Matrix multiplication handles the summation of paths automatically. This is how computational frameworks like PyTorch do automatic differentiation under the hood. The Jacobian approach also makes it obvious when your transformation is singular. If the Jacobian determinant is zero at a point, the mapping locally collapses dimensions. You can't invert it there, and the chain rule still works but you lose information. This matters a lot if you're doing coordinate transformations in physics or optimization.
Practical Walkthrough of the Chain Rule For Multivariable
Let's say w = x²y + ysin(x), where x = rcos() and y = rsin(). Find dw/dr and dw/d. First, the tree: w branches to x and y. Both x and y branch to r and . Two paths from w to r. Two paths from w to . dw/dr = w/x · x/r + w/y · y/r. Compute each piece: w/x = 2xy + y²cos(x), wait no, w/x = 2xy + y(-sin(x))? Let me be careful. w = x²y + ysin(x), so w/x = 2xy + ycos(x). w/y = x² + sin(x). x/r = cos(). y/r = sin(). So dw/dr = (2xy + ycos(x))cos() + (x² + sin(x))sin(). Then substitute x = rcos() and y = rsin() at the end if you need the answer in terms of r and alone.
That's it. Ten lines of calculation. The mistake zone is in computing w/x correctly, which means treating y as a constant. Easy to forget when y itself contains r and , but that's precisely why the partial derivative notation matters. In w/x, y is held constant regardless of what y depends on elsewhere.

The Hard Case: Implicit Dependencies
Here's where things get genuinely annoying. Suppose F(x, y, z) = 0 defines z implicitly as a function of x and y, and both x and y depend on a parameter t. You want dz/dt. You can't just write dz/dt = z/x · dx/dt + z/y · dy/dt without first finding z/x and z/y, which requires implicit differentiation: z/x = -F_x/F_z and z/y = -F_y/F_z. Then plug those into the chain rule formula. I had a problem last semester where F involved exponential and trigonometric terms in all three variables, and F_z happened to be nearly zero at the evaluation point. The implicit derivative formula still works, but numerically it's unstable. I ended up solving for z explicitly using a numerical root finder at each step instead of relying on the analytical implicit formula. Not elegant, but it gave me correct results where the analytical approach was numerically unreliable.
What This Method Doesn't Handle Well
The chain rule assumes differentiability. If any function in the chain has a discontinuity or a kink, the whole thing breaks. Absolute value functions, piecewise definitions, floor functions — these show up in machine learning loss functions all the time, and the chain rule simply doesn't apply at the non-differentiable points. You need subgradients there, and that's a different topic entirely. Another limitation: the chain rule gives you derivatives, not integrals. There's no clean "reverse chain rule" for multivariable substitution in multiple dimensions beyond the Jacobian determinant in change of variables for integrals, and even that has conditions (the transformation needs to be C¹ and injective on the domain, among other things). Don't try to force the chain rule logic into integration problems without checking those conditions. If you're working with discrete variables or data that's only available numerically, analytical chain rule application is impossible. You'd use finite differences or automatic differentiation tools instead. For manual calculation, stick to smooth analytic functions and make sure your dependency tree is complete before you start differentiating.