Why Is Gravitational Potential Energy Negative?
On reference points, the sign in U = −GMm/r, and the three different ways I've come across it over the years.
Ever since I started prepping for IPhO, I got in the habit of deriving everything from scratch instead of just trusting the formula sheet. This was before even 11th grade, so when school got to a concept my textbook covered with a simplified, calculus-free explanation, I'd often get stuck on the gaps — the parts left out because the full treatment needed calculus, or was just judged too hard for the chapter. I distinctly remember burning over 20 minutes of my physics teacher's online lecture during COVID, holding up the class, because his explanation (and the book's) for why gravitational potential energy carries a negative sign was full of holes. To this day I have to give credit where it's due: my 11th-grade KPK textbook actually pulled off a rigorous, calculus-free derivation of it. It still kind of baffles me that they managed it — most sources don't.
I'd probably have forgotten all this if the question didn't keep resurfacing. I've re-derived this same result several times over the years, each time in a more general way — from the sum-of-small-steps approximation the textbooks use, to ordinary single-variable calculus, to the vector calculus version that doesn't care what path you take. What finally got me to sit down and write it all out properly was answering this exact question, again, for a high-school student in a community I mentor in — on the spot, in whatever plaintext math you can type into a chat box. This is that, done right, with equations instead of typed-out approximations.
$ reference points, typeset
Take a uniform field near Earth's surface, where \(g\) doesn't meaningfully change over the heights you care about. Pick ground level, \(h = 0\), as the zero of potential energy. Then a mass \(m\) at height \(h\) has
\[ U(h) = mgh \]
Nothing sacred about \(h=0\) as the zero point — it's just convenient. If instead we declared the zero to sit at \(h = 10\), the formula shifts to
\[ U(h) = mg(h - 10) \]
and now a mass sitting right at the ground, \(h=0\), has \(U = -10mg\). That's not the object secretly carrying negative energy — it's the object sitting below whatever point we arbitrarily decided to call zero. Move the zero, and the number changes; nothing physical does.
Watch that this is actually true by finding the speed of a mass dropped from height \(h\), once with each reference point. First, the natural choice, zero at the ground:
\[ \begin{aligned} U_1 &= mgh, \quad U_2 = 0 \\ KE_1 + U_1 &= KE_2 + U_2 \\ 0 + mgh &= \tfrac{1}{2}mv^2 + 0 \\ v &= \sqrt{2gh} \end{aligned} \]
Now the awkward choice, zero at \(h=10\):
\[ \begin{aligned} U_1 &= mg(h-10), \quad U_2 = mg(0-10) = -10mg \\ KE_1 + U_1 &= KE_2 + U_2 \\ 0 + mg(h-10) &= \tfrac{1}{2}mv^2 - 10mg \\ mgh &= \tfrac{1}{2}mv^2 \\ v &= \sqrt{2gh} \end{aligned} \]
Same answer. It has to be — only differences in potential energy are physical, and the constant offset between the two conventions (\(10mg\), here) cancels on both sides of energy conservation no matter what it is. The only rule is that you can't switch reference points mid-calculation, or that cancellation breaks.
$ gravity at a finite distance
\(U = mgh\) only holds close to the surface, where the field is roughly uniform. Once you're far enough out that \(g\) itself changes with distance, you need the real, inverse-square version. A mass \(m\) at distance \(r\) from a planet of mass \(M\) feels
\[ F(r) = -\frac{GMm}{r^2} \]
where the minus sign just says the force points inward, back toward the planet, for every \(r\). \(g\) itself is just this force divided by \(m\), and it's clearly not constant — it falls off as \(1/r^2\). So — same logic as before — pick a reference point and go. The one most textbooks skip over as an option: use the planet's own surface, \(r = R\), as the zero.
Potential energy is defined so that moving an object costs (or releases) exactly the work done against the force field:
\[ U(r) - U(R) = -\int_R^r F(r')\, dr' \]
Carrying out the integral,
\[ -\int_R^r \left(-\frac{GMm}{r'^2}\right) dr' \;=\; GMm\int_R^r \frac{dr'}{r'^2} \;=\; GMm\left[-\frac{1}{r'}\right]_R^r \;=\; GMm\left(\frac{1}{R} - \frac{1}{r}\right) \]
and with \(U(R) = 0\),
\[ U(r) = GMm\left(\frac{1}{R} - \frac{1}{r}\right) \]
exactly the formula I gave her. Nothing negative in sight, and it's perfectly usable — you just have to carry the planet's radius \(R\) around in every calculation. That's the real reason this convention loses out, more than any elegance argument: \(R\) is a property of this specific lump of rock, not of the two-body problem you're actually trying to solve. Swap the planet for a different one, or ask about the potential energy between two point masses that don't have a "surface" at all, and the formula stops making sense as written.
Send the reference point to \(r \to \infty\) instead — a choice that costs nothing, since gravity is negligible there for any \(M\) — and the formula stops caring about any particular body's size:
\[ U(r) - U(\infty) = -\int_\infty^r F(r')\, dr' = GMm\left[-\frac{1}{r'}\right]_\infty^r = GMm\left(\frac{1}{r} - 0\right) = \frac{GMm}{r} \]
and with \(U(\infty) = 0\),
\[ U(r) = -\frac{GMm}{r} \]
There's the sign. And notice the two formulas agree with each other up to a constant, exactly the way the \(mgh\) example did:
\[ \underbrace{-\frac{GMm}{r}}_{\text{ref. at }\infty} \;-\; \underbrace{\left(-\frac{GMm}{R}\right)}_{\text{ref. at }\infty,\ \text{evaluated at } R} \;=\; GMm\left(\frac{1}{R}-\frac{1}{r}\right) \;=\; \underbrace{U(r)}_{\text{ref. at }R} \]
Same physics, two constants apart — the shift is exactly \(U(R)\) measured on the infinity-referenced scale, which is precisely the reference-point rule from the last section.
Now, the actual question — is it still negative at some large but finite distance out in space, and does that trace back to gravity being attractive? Both yes, and they're really the same fact. \(U(r) = -GMm/r\) is negative for every finite \(r > 0\); it only reaches \(0\) in the limit \(r \to \infty\). Think of it operationally: \(U(r)\) is the work some external agent would have to do to hold the mass steady while lowering it, quasi-statically, from infinity down to \(r\). Because gravity pulls inward and the object is moving inward, gravity itself does positive work over that trip — which means the restraining external force, pointing the opposite way, does negative work. Negative work in equals negative potential energy out, for every finite \(r\), not just the specific one you happen to be standing at.
Equivalently: whatever the sign of the force, \(U(r)\) is the potential energy gained in falling from the reference point down to \(r\). Falling toward something that attracts you is a one-way trip toward lower energy, so \(U\) has to come out negative. Flip gravity to a hypothetical repulsive force, \(F(r) = +GMm/r^2\), redo the same integral, and the sign flips with it:
\[ U(r) = \frac{GMm}{r} \]
positive at every finite \(r\), because now the external agent has to push the object inward against the repulsion, doing positive work the whole way. This isn't hypothetical for long, either — it's exactly Coulomb's law between two charges of the same sign, \(U = kQq/r > 0\), while opposite charges attract and give \(U = -k|Qq|/r\), the same story as gravity. The sign was never really about gravity specifically. It's about attraction versus repulsion, with "attractive" as the special case that happens to describe every mass in the universe.
$ three ways to get here
I've now derived \(U(r) = -GMm/r\) at three different levels of machinery over the years, each one subsuming the last. It's worth walking through all three, because the first is the one that actually impressed me.
1. Without calculus — sum of small steps
This is the method my 11th-grade textbook uses, and it's a genuinely clever piece of pre-calculus reasoning: no derivatives, no integral sign, just algebra and a well-chosen approximation.
Cut the trip from \(r_1\) out to some far point \(r_N\) into \(N\!-\!1\) small hops of size \(\Delta r\). Over any one hop, from \(r_i\) to \(r_{i+1}\), the force isn't constant — it's an inverse-square law — but over a small enough step it's close enough to constant that you can approximate the work as force times distance, \(\Delta W = F_{av}\, \Delta r\), using some suitable "average" force for that step.
The trick is in picking \(F_{av}\). Take it at the geometric mean distance, \(r_{av}\) defined by \(r_{av}^2 = r_i\, r_{i+1}\), rather than the arithmetic mean. You can get there from the arithmetic mean with nothing but algebra: write \(r_{av} = (r_i + r_{i+1})/2\), substitute \(r_{i+1} = r_i + \Delta r\), square it, and throw away the \(\Delta r^2\) term as negligible for a small enough step — the same "small enough that its square doesn't matter" move calculus formalizes with limits, done here by hand instead. What falls out is \(r_{av}^2 \approx r_i\, r_{i+1}\), and that's exactly the substitution that makes everything else telescope. With \(F_{av} = GMm / r_{av}^2\),
\[ \Delta W_{i \to i+1} = \frac{GMm}{r_i\, r_{i+1}}(r_{i+1} - r_i) = GMm\left(\frac{1}{r_i} - \frac{1}{r_{i+1}}\right) \]
Write that out for every hop and add them all up. Every interior term appears once with a plus and once with a minus, and cancels:
\[ \Delta W_{1 \to N} = GMm\left[\left(\frac{1}{r_1}-\frac{1}{r_2}\right)+\left(\frac{1}{r_2}-\frac{1}{r_3}\right)+\cdots+\left(\frac{1}{r_{N-1}}-\frac{1}{r_N}\right)\right] = GMm\left(\frac{1}{r_1} - \frac{1}{r_N}\right) \]
Send the far endpoint to infinity, \(r_N \to \infty\), so \(1/r_N \to 0\), and set the near endpoint to the planet's surface, \(r_1 = R\):
\[ \Delta W_{R \to \infty} = \frac{GMm}{R} \]
— the work needed to remove the mass from the surface out to infinity. The book then makes the same move that trips everyone up at first: this is work done against gravity, so by the definition of potential energy it's recorded with a minus sign, giving \(U = -GMm/R\) at the surface, and \(U(r) = -GMm/r\) in general. No limits, no derivatives — just a clever choice of average and a sum that collapses on itself.
2. Ordinary calculus
Shrink \(\Delta r \to 0\) and let \(N \to \infty\), and the sum above becomes exactly the integral from the section before — the geometric-mean trick was doing, by hand, precisely what the limit does automatically:
\[ U(r) - U(r_0) = -\int_{r_0}^{r} F(r')\, dr' \]
This is the version I met properly once calculus was on the table, and it's the one most textbooks jump straight to — skipping the sum-of-steps picture entirely, which is a shame, since it's the sum that actually explains where the integral comes from instead of just asserting it.
3. Vector calculus — the general case
Both versions above quietly assumed the mass moves straight along the radial line. The fully general statement doesn't need that. A force field \(\vec F(\vec r)\) is conservative — meaning potential energy is even well-defined for it in the first place — when it can be written as the negative gradient of some scalar field:
\[ \vec F(\vec r) = -\nabla U(\vec r) \]
Gravity, being central and spherically symmetric, is conservative, and the line integral of \(\vec F\) between two points comes out identical no matter which path you take between them — straight in, a wide spiral, doesn't matter:
\[ U(\vec r) - U(\vec r_0) = -\int_{\vec r_0}^{\vec r} \vec F(\vec r\,') \cdot d\vec l \]
which is exactly the earlier integral, just no longer restricted to a radial path. Run it in reverse and differentiate \(U(r) = -GMm/r\) to check the whole thing is consistent:
\[ -\nabla\!\left(-\frac{GMm}{r}\right) = -\frac{GMm}{r^2}\,\hat r = \vec F(r) \]
— Newton's law of gravitation, recovered from the potential you started with. This is the version I've come back to the most since — the same negative sign, the same reference point at infinity, just stated in a way that no longer cares what path, or even what force law, you're working with. Every level here is the same idea; each one just stops assuming something the last one needed.