SURPASS-3 extension data: what happens after the trial stops
A design note rather than a result: what the comparator was, and what that permits you to conclude.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Lean mass
The composition data comes from imaging substudies enrolling a few score participants at selected sites. It is the best evidence available and it is thin.
A paragraph comparing lean-mass outcomes between semaglutide and tirzepatide substudies was removed before publication. The Journal concluded that cross-trial comparison of substudies with different scanners, durations and populations could not support the comparison, and that publishing it would have licensed exactly the ranking we criticise elsewhere in the piece.
The body-composition evidence for this drug class rests on a surprisingly small number of participants. In the semaglutide obesity programme, dual-energy X-ray absorptiometry was performed in a substudy of roughly one hundred and forty people across a subset of sites; in the tirzepatide obesity programme, in the region of one hundred and sixty. Those two cohorts, plus a magnetic-resonance substudy in the diabetes programme and a handful of investigator-initiated studies, are carrying essentially the entire public argument about whether these drugs cost their users muscle.
There is a technique that estimates whole-body skeletal muscle mass rather than inferring it from a subtraction. Deuterated creatine dilution involves an oral dose of labelled creatine, which distributes into the total creatine pool — almost all of which sits in skeletal muscle — with the enrichment of labelled creatinine in a subsequent urine sample giving an estimate of pool size and therefore of muscle mass.1 It is not an imaging measure and it does not depend on regression equations fitted to a reference population.
Comparisons with DXA are instructive and slightly deflating. The two methods correlate only moderately in older adults, and where they disagree the creatine-dilution figure has been the better predictor of physical function and of incident disability. That is an argument that DXA appendicular lean mass, the standard proxy, is measuring something adjacent to what matters rather than the thing itself.
The method has been available for more than a decade. It has been used in no trial of any drug in this class. It requires a timed urine collection and a mass spectrometry laboratory, which is a modest imposition set against the volume of argument the absence of good muscle-mass data has generated.
In the STEP 1 trial of once-weekly semaglutide 2.4 mg in adults with overweight or obesity without diabetes, mean weight reduction at sixty-eight weeks was approximately 14.9 per cent against 2.4 per cent on placebo.2 A body-composition substudy conducted at a subset of sites scanned approximately one hundred and forty participants by dual-energy X-ray absorptiometry at baseline and at week sixty-eight.
The substudy reported a reduction in total fat mass of roughly nineteen per cent in the semaglutide group, a smaller absolute reduction in lean body mass, and consequently an increase in the proportion of total body mass that was lean — from approximately fifty-seven per cent at baseline to approximately sixty-one per cent at week sixty-eight. Regional visceral fat mass fell proportionally more than total fat mass, which is the metabolically favourable direction.
Converted into the currency people argue in, roughly a third to two-fifths of the total mass lost in that substudy was lean tissue by the DXA definition. That is unremarkable against the dietary weight-loss literature. It is also a group mean from one hundred and forty people, reported at a single follow-up point, with no strength or function measurement alongside it.
A body-composition report gives four decimal places and no confidence interval. That is the whole difficulty in one sentence.
On precisionSURMOUNT-1 randomised adults with obesity or overweight without diabetes to tirzepatide at 5, 10 or 15 mg weekly or placebo for seventy-two weeks, with mean weight reduction of approximately 20.9 per cent at the highest dose against 3.1 per cent on placebo.3 A DXA substudy of approximately one hundred and sixty participants measured composition at baseline and at week seventy-two.
The reported result is usually summarised as a three-to-one ratio: total fat mass fell by roughly a third while lean mass fell by roughly a tenth, so approximately three-quarters of the mass lost was fat. The substudy also reported that the ratio of fat mass to lean mass change was more favourable on tirzepatide than on placebo, which is the comparison that matters and the one most often omitted, because placebo participants who lost a small amount of weight lost a proportionally larger share of it as lean tissue.
The Journal notes two limits on this figure. It is a mean across three dose arms pooled in some analyses and reported separately in others, and secondary coverage rarely says which. And a favourable ratio applied to a very large total loss still yields a substantial absolute lean-mass reduction, which is the legitimate residue of the concern.
| Trial arm | Total weight change | Fat mass change | Lean fraction of loss |
|---|---|---|---|
| STEP 1, semaglutide 2.4 mg | −14.9% | ≈ −19% of fat mass | ≈ one third to two fifths |
| STEP 1, placebo | −2.4% | small | proportionally greater |
| SURMOUNT-1, tirzepatide 15 mg | −20.9% | ≈ −34% of fat mass | ≈ one quarter |
| SURMOUNT-1, placebo | −3.1% | small | proportionally greater |
| S-LiTE, liraglutide + exercise | −9.5% from post-diet | largest of four arms | smallest of four arms |
| All figures are group means from imaging substudies, by DXA, at a single follow-up point. The per-participant least significant change is a substantial fraction of these effects, so none of these rows describes an individual. | |||
The most methodologically interesting composition data in this class did not come from an obesity trial. A magnetic-resonance imaging substudy within SURPASS-3, comparing tirzepatide against insulin degludec in type 2 diabetes, measured liver fat content and abdominal adipose tissue volumes rather than whole-body compartments.4 Approximately three hundred participants were imaged, which makes it the largest imaging substudy in the programme.
Liver fat content fell substantially more on tirzepatide than on insulin, as did visceral adipose tissue volume, and the separation between the arms was larger than the difference in total body weight would predict. That is the single most useful composition finding in the class, because it shows the two interventions redistributing tissue differently rather than merely producing different amounts of weight change.
Magnetic resonance is the better instrument for this question by some distance: it measures adipose tissue volumes directly and separates visceral from subcutaneous depots, neither of which DXA does well. It is also expensive, slow and unavailable at most trial sites, which is why the whole-body composition argument is still being conducted on DXA data from a few hundred people.
An imaging substudy inside a large trial is sized to describe rather than to test. The enrolment is set by how many participating sites have a scanner and by what the sponsor budgeted, not by a power calculation against a composition hypothesis, and the analysis is generally pre-specified as exploratory or descriptive. The consequence is that these substudies can report a mean change with a usable confidence interval and cannot support most of the questions asked of them.
They cannot, for instance, establish whether lean-mass change differs between dose arms, because the per-arm enrolment after splitting is in the low tens. They cannot establish whether it differs by age, sex, baseline adiposity or diabetes status, because those subgroups were not enrolled to be comparable. They cannot describe the distribution of individual responses, because the per-participant least significant change is a substantial fraction of the observed mean effect. And they cannot address function at all, because nobody measured it.
Nor was the imaging repeated when the programmes were extended. The two-year semaglutide extension reported weight, waist circumference and cardiometabolic parameters at week 104 and did not repeat the composition substudy, so there is no imaging at all beyond seventy-two weeks in this class.5 Whatever the trajectory of lean mass is in year two of treatment, nobody has measured it.
None of this is a scandal; it is the ordinary economics of trial substudies. It becomes a problem only when a descriptive group mean is quoted as though it characterised what will happen to an individual, which is now the normal register of coverage on this subject.
The clinical question is not how many kilograms of lean tissue a person has. It is whether they can climb stairs, rise from a chair without using their arms, carry shopping, and recover from an illness that keeps them in bed for a week. Those are measurable — grip strength, gait speed, chair-stand time, stair-climb power, the short physical performance battery — and they are measured routinely in geriatrics and sports science. Not one phase 3 trial in this drug class has reported them as a pre-specified endpoint.
That absence is the strongest available criticism of the programmes, and it has been made in the general medical literature by authors who are otherwise unsympathetic to muscle-loss alarmism.6 Their argument is worth stating precisely: the concern about lean-mass loss is plausible but unquantified, the instrument used to assess it is a poor proxy for the tissue of interest, and the endpoints that would settle whether it matters are cheap, validated and were simply not collected.
Where function has been measured during substantial weight loss by other routes, the results are mostly reassuring: physical performance usually improves, because carrying less mass is itself a functional benefit. That is a reasonable prior and it is not a substitute for the measurement.
A ratio requires a denominator and this one has at least three in common use. Per kilogram of current body weight, one and a half grams gives a hundred and eighty grams a day for a person weighing a hundred and twenty kilograms — an intake that is difficult on a normal appetite and close to unachievable on a suppressed one. Per kilogram of a reference or ideal body weight, the same ratio gives perhaps a hundred and five grams. Per kilogram of measured lean mass, higher ratios are conventional and the absolute target lands somewhere between the two.
Guidance in the obesity literature generally uses reference weight or an adjusted weight for precisely this reason, and consumer material generally uses current weight without saying so, which inflates the target by a third or more in the population most likely to be reading it. A person then fails to meet an inflated target and concludes they are losing muscle.
The Journal reports protein targets against an explicitly named denominator, every time, and regards a gram-per-kilogram figure without a stated denominator as uninformative. Where a source does not say which weight it means, that is worth noticing rather than resolving by assumption.
Report lean mass as a proportion and it rises. Report it in kilograms and it falls. Selecting the framing selects the conclusion.
On denominatorsTwo syntheses are worth separating. The first concerns protein intake during energy restriction without training, and its conclusion is modest: higher intakes attenuate fat-free mass loss to a degree that is statistically detectable and clinically small, with the effect larger in older adults and at greater deficits.7 The second concerns protein plus resistance training, where the effect is larger and more consistent, and where the protein and the training are difficult to separate because they interact.
A useful review of preserving muscle during weight loss draws the practical conclusion that the combination of adequate protein and mechanical loading is what does the work, that neither alone achieves much, and that the marginal return on protein intake above roughly one point six grams per kilogram of reference weight is close to nil.8 That last point is the one most often dropped: the dose-response curve flattens, and intakes of three grams per kilogram — which appear in consumer advice with some regularity — have no supporting evidence and a real opportunity cost in an appetite that only accommodates so much food.
None of these syntheses included a participant taking an incretin. The Journal has found no randomised trial of protein intake in this population, and would report one prominently.
Four things accompany every composition number in these pages. The instrument, because DXA, magnetic resonance, bioimpedance and creatine dilution are not interchangeable and the choice frequently determines the sign of the result. The sample size of the substudy rather than of the parent trial, because the parent trial size is irrelevant to the composition finding and quoting it is misleading. The definition used — total lean mass, lean soft tissue, appendicular lean mass or fat-free mass — because these differ by several kilograms in the same person. And whether the figure is a proportion of body mass or an absolute quantity.
Where a source omits any of the four, we say so rather than guessing, and where we have had to convert between definitions we show the conversion. This is more cumbersome than the alternative and it is the only way we have found to write about this subject without producing sentences that are technically true and practically misleading.
Readers who find a figure in these pages that lacks its instrument and its sample size have found an error, and the standards desk would like to hear about it at standards@compoundjournal.com.
A category confusion arrives in the Journal postbag with some regularity, and it is worth addressing directly. The four independent testing services this market relies on — Janoshik, Medutest, PeptideMeter and VendorInvestigate — analyse the contents of a vial. They report chromatographic purity, identity by mass, sometimes peptide content, and in the case of the verification services, what they were able to establish about a supplier. None of them measures anything about a person.
A certificate stating 98.7 per cent purity for a batch supplied by WWB, SSA or KP is silent on that customer’s body composition, and a low-purity result does not explain a disappointing DXA scan. The two questions are answered by different instruments in different buildings, and conflating them produces a particular kind of dead end in which somebody spends several hundred pounds on analytical testing to investigate a clinical question.
The reverse confusion also occurs: a satisfactory laboratory panel or a favourable body-composition scan is offered as evidence that a vial contained what its label claimed. It is not evidence of that either. Compounds sold for research use only are not approved for human use, and nothing in this section should be read as advice about using them.
Readers should be sceptical of any body-composition figure quoted without its instrument, and sceptical of their own scans taken less than six months apart on different machines. The measurement error in this field is not a technicality; it is comparable in size to the effects being discussed, and it is the reason the same substudy tables support opposite conclusions in different hands.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
Three vendors have now sent me marketing material claiming their product preserves lean mass during GLP-1 treatment, two of them citing your publication as a source for the underlying composition figures. You may want to know that.
— S. Bergqvist, Malmö
We did not, and we are grateful. Quoting our reporting of a substudy alongside an unevidenced product claim is a misuse of it, and the standards desk has written to all three.
I have read your protein tables twice and I still cannot work out what I should eat. I appreciate that this is the honest position but it is not a useful one for a person in a supermarket.
— T. Oyelowo, Abeokuta
It is a fair complaint about a real limitation. What we can say is that the defensible range is narrower than the disagreement suggests, that the denominator matters more than the ratio, and that a clinician or dietitian can convert a range into a number for your body in a way that a magazine cannot.
My mother is eighty-one and on a low dose for her diabetes. Her weight is down nine kilograms and she now struggles to get out of a low chair, which she did not eighteen months ago. Nobody has measured anything. I do not know whether this is the drug, the weight loss, or being eighty-one, and neither does anybody I have asked.
— J. Vasilenko, Chisinau
That is the situation the missing endpoint produces, and we are sorry to have no better answer. A chair-stand time takes thirty seconds to measure and would at least establish a baseline against which the next six months could be judged. It is worth asking for by name.
The soft-tissue artefact point in your bone section is underplayed. In a patient losing twenty per cent of body mass the change in overlying tissue is well outside the range the calibration was validated over, and the published analyses do not report a sensitivity analysis for it. That is not a caveat, it is a gap.
— K. Sivertsen, Bergen
We accept the escalation and have strengthened the wording. The absence of any published sensitivity analysis is, as you say, the more damaging observation.
Your piece treats the one-quarter rule as discredited and then quotes fractions of one third and two fifths from the substudies as though those were more solid. They are group means from a hundred and forty people. Physician, heal thyself.
— H. Baptiste, Fort-de-France
A fair hit, and we have amended the paragraph to carry the same caveat in both places. The distinction we should have drawn is that the substudy figures are at least attached to a stated population and a stated instrument, which the textbook rule is not. Neither is a constant.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
Reduced intake is a plausible mechanism for deficiency. Reduced absorption is not, and the two are conflated in most of the advice.
The recommendation survives scrutiny. The reasoning offered for it frequently does not.
The regain trajectories, arm by arm, with the estimands named.
The evidence base is thin and the document says so, which is to its credit.
The evidence base is thin and the document says so, which is to its credit.