SURPASS-3 was built to answer the stopping question, and it did
Three randomised withdrawal designs have tested what happens when treatment stops. Their results are consistent and they are consistently misreported.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Skeletal health
The gap between a defensible recommendation and a confident one is where most of the harm in this subject lives.
The Journal is not arguing that the popular advice is wrong. Eating adequate protein and loading the skeleton during rapid weight loss are sensible on general physiological grounds, carry essentially no risk, and are what this publication would do. The argument is narrower and, we think, more useful: the confidence with which these things are asserted, and the specificity of the numbers attached to them, are borrowed from literatures that did not study this drug class, this rate of loss, or this population, and the borrowing should be visible.
Clinical teaching has long held that approximately twenty-five per cent of the mass lost during weight reduction is fat-free tissue. The figure appears in textbooks, in review articles and in a great deal of consumer material, usually without a citation and always without an interval.
A critical review published in 2014 traced the rule to a limited number of older studies, examined the variation across the wider literature, and concluded that treating one-quarter as a constant is not defensible.1 The fraction of loss that is fat-free tissue varies systematically with baseline adiposity — heavier people lose proportionally more fat — and with the rate of loss, the protein intake, the activity pattern and the measurement method. Reported values span from well under fifteen per cent to above thirty-five.
This matters for the current argument in a specific way. Both the reassuring and the alarming readings of the incretin substudy data are constructed by comparing an observed fat-free fraction against the one-quarter benchmark. If the benchmark is a loose average rather than an expectation, both comparisons are weaker than they appear, and the honest statement is that the observed fractions sit within the range that dietary weight loss has always produced.
There is a rhetorical move available to both sides of this argument and it works by choosing a denominator. Report lean mass as a proportion of total body mass and it rises during successful treatment, because fat is falling faster; the treatment looks composition-improving, which it is. Report lean mass in absolute kilograms and it falls; the treatment looks muscle-costing, which it also is. Both statements can be made from the same scan pair without either being false.
The Journal reports both, in that order, and thinks anybody presenting only one should be asked why. The proportional figure is the right one for questions about metabolic quality: a body with a higher lean fraction handles glucose better and carries less ectopic fat. The absolute figure is the right one for questions about function and reserve, because a hip fracture at seventy-eight is not prevented by a favourable ratio.
The two framings also diverge most sharply exactly where the stakes are highest. A person losing twenty-five per cent of their body weight will show an excellent proportional result and the largest absolute lean-mass reduction in the cohort. Selecting the framing selects the conclusion, which is why the trade has settled on whichever one suits it.
Reduced lean mass on a scan, without measured weakness, does not meet any published definition of sarcopenia.
On borrowed vocabularyThe clinical question is not how many kilograms of lean tissue a person has. It is whether they can climb stairs, rise from a chair without using their arms, carry shopping, and recover from an illness that keeps them in bed for a week. Those are measurable — grip strength, gait speed, chair-stand time, stair-climb power, the short physical performance battery — and they are measured routinely in geriatrics and sports science. Not one phase 3 trial in this drug class has reported them as a pre-specified endpoint.
That absence is the strongest available criticism of the programmes, and it has been made in the general medical literature by authors who are otherwise unsympathetic to muscle-loss alarmism.2 Their argument is worth stating precisely: the concern about lean-mass loss is plausible but unquantified, the instrument used to assess it is a poor proxy for the tissue of interest, and the endpoints that would settle whether it matters are cheap, validated and were simply not collected.
Where function has been measured during substantial weight loss by other routes, the results are mostly reassuring: physical performance usually improves, because carrying less mass is itself a functional benefit. That is a reasonable prior and it is not a substitute for the measurement.
| Method | Directly measured | Muscle mass estimate | Typical CV | Practical limit |
|---|---|---|---|---|
| DXA | X-ray attenuation at two energies | By subtraction; appendicular proxy | 1.0–1.5% | Soft-tissue and hydration assumptions |
| Bioimpedance | Electrical impedance | By population regression | 2–5% | Tracks body water, not tissue |
| Magnetic resonance | Tissue volumes | Segmented, near-direct | <1% | Cost, throughput, analysis time |
| D3-creatine dilution | Creatine pool size | Direct, whole-body muscle | ≈5% | Timed urine plus mass spectrometry |
| Air displacement | Body volume and density | Two-compartment only | 1–2% | No regional data at all |
| Coefficients of variation are for repeated measurement on the same device with a consistent operator. Cross-device comparison degrades all of them and is not recoverable by calibration. | ||||
Two claims are routinely bundled together and only one is well supported. The weaker claim is that resistance training during pharmacological weight loss builds or maintains muscle mass. In a substantial energy deficit, training generally attenuates the loss rather than preventing it, and net accrual is unusual outside of untrained beginners and the specific controlled-feeding conditions of the trials cited earlier. The stronger claim is that training preserves strength and physical function even where mass declines, which is consistently observed and is mechanistically sensible: a large part of early strength change is neural rather than structural.
The distinction has practical consequences. Somebody training hard, eating well, and watching their DXA appendicular lean mass fall by two kilograms across nine months has not failed at anything, and may be measurably stronger than at baseline. If the expectation set for them was mass preservation, they will read a normal outcome as a failure and may respond by eating more or training in ways that suit the metric rather than the goal.
The Journal reports the training recommendation and reports what it is expected to achieve, which is function first and mass second.
The word sarcopenia has migrated from clinical medicine into consumer discussion of this drug class and lost its definition in transit. In the working definitions used by the European and Asian consensus groups, sarcopenia requires low muscle strength, with low muscle quantity or quality confirming it and poor physical performance indicating severity. Strength is the entry criterion. Reduced lean mass on a scan, in the absence of measured weakness, does not meet any published definition of sarcopenia.
This matters because the borrowed term imports a prognosis. Sarcopenia in its clinical sense is associated with falls, fractures, hospitalisation and mortality, and those associations were established in older adults with measured weakness, frequently in the context of illness or immobility. Applying the label to a forty-two-year-old whose DXA appendicular lean mass has fallen by one and a half kilograms while their strength has increased is not a cautious extrapolation; it is a category error with a frightening prognosis attached.
The related term sarcopenic obesity has the same problem in a more acute form, since it requires both criteria to be met and is frequently used to mean nothing more than a low lean fraction. The Journal uses both terms only in their defined sense and asks correspondents who use them to say which criteria they mean.
Readers should be sceptical of any body-composition figure quoted without its instrument, and sceptical of their own scans taken less than six months apart on different machines. The measurement error in this field is not a technicality; it is comparable in size to the effects being discussed, and it is the reason the same substudy tables support opposite conclusions in different hands.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
I train four times a week, eat a hundred and sixty grams of protein and my appendicular lean mass has fallen by 1.8 kg over ten months while every lift has gone up. Your section on mass against function was the first thing I have read that made that seem normal rather than a failure.
— D. Mazzarella, Catania
Your piece treats the one-quarter rule as discredited and then quotes fractions of one third and two fifths from the substudies as though those were more solid. They are group means from a hundred and forty people. Physician, heal thyself.
— A. Mbeki, Lusaka
A fair hit, and we have amended the paragraph to carry the same caveat in both places. The distinction we should have drawn is that the substudy figures are at least attached to a stated population and a stated instrument, which the textbook rule is not. Neither is a constant.
You write that no trial has measured strength. There are observational cohorts with grip strength data. Why do you insist on randomised measurement?
— R. Whitlam, Adelaide, SA
Because grip strength in an observational cohort of people who chose to take a drug, and who differ from those who did not in age, motivation and comorbidity, cannot separate the drug effect from the selection. We report those cohorts and we do not treat them as answering the question.
As a DXA technologist of twenty-two years I would add one thing to your precision section: the largest source of error in practice is not the machine, it is positioning. A patient scanned with their arms two centimetres further from their trunk will report different regional values. We are trained to a protocol and the protocol is not always followed.
— C. Rautenbach, Pretoria
We should have said this and did not. It also argues for what you presumably practise: same device, same technologist, same protocol, and a note in the record when any of those changes.
The soft-tissue artefact point in your bone section is underplayed. In a patient losing twenty per cent of body mass the change in overlying tissue is well outside the range the calibration was validated over, and the published analyses do not report a sensitivity analysis for it. That is not a caveat, it is a gap.
— D. Ramkissoon, Port of Spain
We accept the escalation and have strengthened the wording. The absence of any published sensitivity analysis is, as you say, the more damaging observation.
Three randomised withdrawal designs have tested what happens when treatment stops. Their results are consistent and they are consistently misreported.
Micronutrient guidance for this drug class is borrowed almost entirely from post-bariatric surveillance, where the anatomy is different and the deficiency mechanisms are not…
A design note rather than a result: what the comparator was, and what that permits you to conclude.
Real-world persistence figures, with their definitions stated, because the definitions are doing most of the work.
Three different explanations for the same abnormal number, and how to tell them apart.
The variance around the mean regain trajectory is large and unexplained, exactly as it is for the weight loss.