What the older diet-and-exercise trials found in the hip
A plausible mechanism, a measurable change, and no outcome data. This is what an open question looks like.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Substudies
The recommendation survives scrutiny. The reasoning offered for it frequently does not.
The recommendation to train while losing weight is not controversial and this publication endorses reporting it. What deserves scrutiny is the mechanism usually offered alongside it. Resistance training during a substantial energy deficit does not reliably build muscle; the deficit is the binding constraint and no amount of load overcomes a large one. What it reliably does is attenuate the loss and, more consistently still, preserve strength and physical function even where mass declines. Those are different claims with different evidence behind them, and only the second is well supported.
The figures in circulation — commonly one and a half to two grams of protein per kilogram of body weight daily, sometimes expressed as a floor of around a hundred grams — are traceable. The most-cited primary source is a randomised trial in resistance-trained young men under a substantial energy deficit, comparing a higher against a lower protein intake with supervised training and controlled feeding; the higher-intake group gained lean mass and lost more fat over four weeks.1 Supporting evidence comes from a large meta-analysis of protein supplementation during resistance training, which found a benefit to lean mass accrual that plateaued at around one and a half to one point six grams per kilogram daily.2
Both are good studies. Neither enrolled anybody over about thirty-five, anybody with obesity, or anybody losing weight at more than a small fraction of the rate this drug class produces. The plateau figure in particular is a plateau for training-induced accrual in weight-stable or mildly deficit conditions, and its application as a preservation target during a twenty per cent weight reduction is an extrapolation rather than a finding.
The Journal quotes these numbers because they are the best available and states their provenance because the provenance is the argument.
A Danish randomised trial remains the only controlled test of the obvious question. After an eight-week low-energy diet producing approximately thirteen kilograms of weight loss, participants were randomised for one year to supervised exercise alone, liraglutide 3.0 mg alone, both combined, or placebo.3 The combination arm achieved the largest weight reduction and, more relevantly here, the most favourable composition outcome: body fat percentage fell roughly twice as much in the combination group as in either single-intervention group, and the exercise arms preserved lean mass better than the drug-alone arm.
Three qualifications belong with that result. The exercise was supervised and substantial — two group sessions and two individual sessions weekly, with a vigorous-intensity target — which is not what most people mean by adding exercise. The agent was liraglutide at 3.0 mg daily, producing considerably less weight loss than the current agents, so whether the interaction scales to a twenty per cent reduction is unknown. And the trial began after weight had already been lost, so it is a maintenance study rather than an induction study.
With those stated, it is the best evidence in the field and it points in the direction the general advice already points.
Reduced lean mass on a scan, without measured weakness, does not meet any published definition of sarcopenia.
On borrowed vocabularyThe closest analogue to rapid weight loss in an older, heavier population predates this drug class entirely. In a randomised trial of adults aged sixty-five and over with obesity, assigned to diet, exercise, both or a control condition for a year, the combination produced the largest improvement in physical function, and the exercise component attenuated the loss of lean mass and of bone mineral density that diet alone caused.4 Diet alone improved function too — carrying less mass helps — but by less, and at a measurable skeletal cost.
That trial is the template for how the question should be asked in this class: randomise the co-intervention, measure function as a primary endpoint, measure bone, and follow for long enough for the skeleton to respond. Its population, older and heavier and losing weight quickly, resembles a large share of current incretin users far more closely than the young resistance-trained cohorts from which most consumer advice descends.
The Journal cites it frequently for that reason and notes the obvious limitation: the weight loss achieved was roughly a tenth of body mass over a year, which is half or less of what the current agents produce. Whether the protective effect of training holds at twice the rate of loss is not established.
| Trial arm | Total weight change | Fat mass change | Lean fraction of loss |
|---|---|---|---|
| STEP 1, semaglutide 2.4 mg | −14.9% | ≈ −19% of fat mass | ≈ one third to two fifths |
| STEP 1, placebo | −2.4% | small | proportionally greater |
| SURMOUNT-1, tirzepatide 15 mg | −20.9% | ≈ −34% of fat mass | ≈ one quarter |
| SURMOUNT-1, placebo | −3.1% | small | proportionally greater |
| S-LiTE, liraglutide + exercise | −9.5% from post-diet | largest of four arms | smallest of four arms |
| All figures are group means from imaging substudies, by DXA, at a single follow-up point. The per-participant least significant change is a substantial fraction of these effects, so none of these rows describes an individual. | |||
Two claims are routinely bundled together and only one is well supported. The weaker claim is that resistance training during pharmacological weight loss builds or maintains muscle mass. In a substantial energy deficit, training generally attenuates the loss rather than preventing it, and net accrual is unusual outside of untrained beginners and the specific controlled-feeding conditions of the trials cited earlier. The stronger claim is that training preserves strength and physical function even where mass declines, which is consistently observed and is mechanistically sensible: a large part of early strength change is neural rather than structural.
The distinction has practical consequences. Somebody training hard, eating well, and watching their DXA appendicular lean mass fall by two kilograms across nine months has not failed at anything, and may be measurably stronger than at baseline. If the expectation set for them was mass preservation, they will read a normal outcome as a failure and may respond by eating more or training in ways that suit the metric rather than the goal.
The Journal reports the training recommendation and reports what it is expected to achieve, which is function first and mass second.
Densitometry infers bone mineral density from the differential attenuation of two X-ray energies, using the surrounding soft tissue as the baseline against which bone is distinguished. The algorithm assumes a soft-tissue composition, and that assumption is embedded in the calibration. When the thickness and fat fraction of the tissue overlying a measurement site change substantially, part of the apparent change in bone density is an artefact of the altered baseline.
The magnitude is contested. Phantom and cadaver work suggests errors of the order of one to three per cent for large changes in overlying fat, which is the same order as the real bone changes being reported over a year of rapid weight loss. In practice this means that a hip bone mineral density reduction of two per cent in a person who has lost a fifth of their body weight cannot be cleanly separated into a bone effect and a measurement effect, and the published analyses do not attempt it.
Quantitative computed tomography and high-resolution peripheral imaging are less vulnerable, measure geometry and microarchitecture rather than areal density, and have not been used in any trial in this class. The Journal regards that as the most easily closed gap in the whole body-composition literature.
Four things accompany every composition number in these pages. The instrument, because DXA, magnetic resonance, bioimpedance and creatine dilution are not interchangeable and the choice frequently determines the sign of the result. The sample size of the substudy rather than of the parent trial, because the parent trial size is irrelevant to the composition finding and quoting it is misleading. The definition used — total lean mass, lean soft tissue, appendicular lean mass or fat-free mass — because these differ by several kilograms in the same person. And whether the figure is a proportion of body mass or an absolute quantity.
Where a source omits any of the four, we say so rather than guessing, and where we have had to convert between definitions we show the conversion. This is more cumbersome than the alternative and it is the only way we have found to write about this subject without producing sentences that are technically true and practically misleading.
Readers who find a figure in these pages that lacks its instrument and its sample size have found an error, and the standards desk would like to hear about it at standards@compoundjournal.com.
A category confusion arrives in the Journal postbag with some regularity, and it is worth addressing directly. The four independent testing services this market relies on — Janoshik, Medutest, PeptideMeter and VendorInvestigate — analyse the contents of a vial. They report chromatographic purity, identity by mass, sometimes peptide content, and in the case of the verification services, what they were able to establish about a supplier. None of them measures anything about a person.
A certificate stating 98.7 per cent purity for a batch supplied by WWB, SSA or KP is silent on that customer’s body composition, and a low-purity result does not explain a disappointing DXA scan. The two questions are answered by different instruments in different buildings, and conflating them produces a particular kind of dead end in which somebody spends several hundred pounds on analytical testing to investigate a clinical question.
The reverse confusion also occurs: a satisfactory laboratory panel or a favourable body-composition scan is offered as evidence that a vial contained what its label claimed. It is not evidence of that either. Compounds sold for research use only are not approved for human use, and nothing in this section should be read as advice about using them.
The next instalment in this department takes up the question that follows this one chronologically rather than logically: what happens to all of it when treatment stops. The composition of regained weight is a separate literature, it is thinner than this one, and what little exists is not encouraging.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
Small correction to your table: the S-LiTE exercise prescription was two supervised group sessions and two individual sessions weekly, not two sessions in total. The distinction matters because "add some exercise" is not what was tested.
— G. Escalante, Lima
Correct, and that is precisely the point we were trying to make and then undermined in our own table. Amended.
I am sixty-eight, I have lost nineteen kilograms over fourteen months, and my consultant has twice told me my lean mass is fine on the basis of a handheld bioimpedance device in the clinic corridor. Having read your piece on what that device measures, I am no longer sure what I have been reassured about.
— D. Iversen, Aalborg
Nor are we. A handheld device measures impedance across the upper body and infers the rest, and the inference is least reliable exactly where you sit: older, substantial weight change, changing hydration. That is not a criticism of your consultant’s judgement, which may be sound on other grounds, but the device is not the evidence for it.
Your piece treats the one-quarter rule as discredited and then quotes fractions of one third and two fifths from the substudies as though those were more solid. They are group means from a hundred and forty people. Physician, heal thyself.
— M. Fitzhenry, Cork
A fair hit, and we have amended the paragraph to carry the same caveat in both places. The distinction we should have drawn is that the substudy figures are at least attached to a stated population and a stated instrument, which the textbook rule is not. Neither is a constant.
I train four times a week, eat a hundred and sixty grams of protein and my appendicular lean mass has fallen by 1.8 kg over ten months while every lift has gone up. Your section on mass against function was the first thing I have read that made that seem normal rather than a failure.
— A. Nazarian, Glendale, CA
My mother is eighty-one and on a low dose for her diabetes. Her weight is down nine kilograms and she now struggles to get out of a low chair, which she did not eighteen months ago. Nobody has measured anything. I do not know whether this is the drug, the weight loss, or being eighty-one, and neither does anybody I have asked.
— L. Whitcombe, Christchurch
That is the situation the missing endpoint produces, and we are sorry to have no better answer. A chair-stand time takes thirty seconds to measure and would at least establish a baseline against which the next six months could be judged. It is worth asking for by name.
A plausible mechanism, a measurable change, and no outcome data. This is what an open question looks like.
Where the familiar figures come from, what populations they were measured in, and how far the extrapolation reaches.
The older-adult diet-and-exercise trials are the closest analogue to rapid pharmacological weight loss, and they are twenty years old.
Mass and function are different endpoints and training affects them differently. Most coverage treats them as one.
Every time a vial changes, the conversion must be recalculated. Carrying forward a unit count from the last vial is the single most reliable way to give the wrong dose.
Dye ingress and microbial immersion are probabilistic. Vacuum decay, high-voltage leak detection and helium mass spectrometry are deterministic and far more sensitive.