Vol. 3, No. 6 — June 2026Independent since 2024

TheCompound Journal

Reporting on incretins, compounding & the peptide supply chain

A monthly journal of record.
30 issues · 32 contributors
Not medical advice. We sell nothing.

Skeletal health

Reading the exenatide composition data without the press release

A substudy powered to describe a mean is not a substudy powered to detect a clinically meaningful individual change.

An imaging substudy is a secondary exercise. It is not the reason the trial was funded, it does not determine whether the trial succeeded, its sites are chosen for having a scanner rather than for representing the population, and its sample size is set by what the sponsor was willing to pay for rather than by a power calculation against a composition hypothesis. None of that makes the results wrong. All of it should temper the confidence with which single decimal places from those tables are quoted eighteen months later.

What the semaglutide DXA substudy measured

In the STEP 1 trial of once-weekly semaglutide 2.4 mg in adults with overweight or obesity without diabetes, mean weight reduction at sixty-eight weeks was approximately 14.9 per cent against 2.4 per cent on placebo.1 A body-composition substudy conducted at a subset of sites scanned approximately one hundred and forty participants by dual-energy X-ray absorptiometry at baseline and at week sixty-eight.

The substudy reported a reduction in total fat mass of roughly nineteen per cent in the semaglutide group, a smaller absolute reduction in lean body mass, and consequently an increase in the proportion of total body mass that was lean — from approximately fifty-seven per cent at baseline to approximately sixty-one per cent at week sixty-eight. Regional visceral fat mass fell proportionally more than total fat mass, which is the metabolically favourable direction.

Converted into the currency people argue in, roughly a third to two-fifths of the total mass lost in that substudy was lean tissue by the DXA definition. That is unremarkable against the dietary weight-loss literature. It is also a group mean from one hundred and forty people, reported at a single follow-up point, with no strength or function measurement alongside it.

The tirzepatide substudy, and the ratio it reported

SURMOUNT-1 randomised adults with obesity or overweight without diabetes to tirzepatide at 5, 10 or 15 mg weekly or placebo for seventy-two weeks, with mean weight reduction of approximately 20.9 per cent at the highest dose against 3.1 per cent on placebo.2 A DXA substudy of approximately one hundred and sixty participants measured composition at baseline and at week seventy-two.

The reported result is usually summarised as a three-to-one ratio: total fat mass fell by roughly a third while lean mass fell by roughly a tenth, so approximately three-quarters of the mass lost was fat. The substudy also reported that the ratio of fat mass to lean mass change was more favourable on tirzepatide than on placebo, which is the comparison that matters and the one most often omitted, because placebo participants who lost a small amount of weight lost a proportionally larger share of it as lean tissue.

The Journal notes two limits on this figure. It is a mean across three dose arms pooled in some analyses and reported separately in others, and secondary coverage rarely says which. And a favourable ratio applied to a very large total loss still yields a substantial absolute lean-mass reduction, which is the legitimate residue of the concern.

Reduced lean mass on a scan, without measured weakness, does not meet any published definition of sarcopenia.

On borrowed vocabulary

The magnetic-resonance substudy in the diabetes programme

The most methodologically interesting composition data in this class did not come from an obesity trial. A magnetic-resonance imaging substudy within SURPASS-3, comparing tirzepatide against insulin degludec in type 2 diabetes, measured liver fat content and abdominal adipose tissue volumes rather than whole-body compartments.3 Approximately three hundred participants were imaged, which makes it the largest imaging substudy in the programme.

Liver fat content fell substantially more on tirzepatide than on insulin, as did visceral adipose tissue volume, and the separation between the arms was larger than the difference in total body weight would predict. That is the single most useful composition finding in the class, because it shows the two interventions redistributing tissue differently rather than merely producing different amounts of weight change.

Magnetic resonance is the better instrument for this question by some distance: it measures adipose tissue volumes directly and separates visceral from subcutaneous depots, neither of which DXA does well. It is also expensive, slow and unavailable at most trial sites, which is why the whole-body composition argument is still being conducted on DXA data from a few hundred people.

What each instrument measures, and what it costs in precision
MethodDirectly measuredMuscle mass estimateTypical CVPractical limit
DXAX-ray attenuation at two energiesBy subtraction; appendicular proxy1.0–1.5%Soft-tissue and hydration assumptions
BioimpedanceElectrical impedanceBy population regression2–5%Tracks body water, not tissue
Magnetic resonanceTissue volumesSegmented, near-direct<1%Cost, throughput, analysis time
D3-creatine dilutionCreatine pool sizeDirect, whole-body muscle≈5%Timed urine plus mass spectrometry
Air displacementBody volume and densityTwo-compartment only1–2%No regional data at all
Coefficients of variation are for repeated measurement on the same device with a consistent operator. Cross-device comparison degrades all of them and is not recoverable by calibration.

What the substudies were never powered to detect

An imaging substudy inside a large trial is sized to describe rather than to test. The enrolment is set by how many participating sites have a scanner and by what the sponsor budgeted, not by a power calculation against a composition hypothesis, and the analysis is generally pre-specified as exploratory or descriptive. The consequence is that these substudies can report a mean change with a usable confidence interval and cannot support most of the questions asked of them.

They cannot, for instance, establish whether lean-mass change differs between dose arms, because the per-arm enrolment after splitting is in the low tens. They cannot establish whether it differs by age, sex, baseline adiposity or diabetes status, because those subgroups were not enrolled to be comparable. They cannot describe the distribution of individual responses, because the per-participant least significant change is a substantial fraction of the observed mean effect. And they cannot address function at all, because nobody measured it.

Nor was the imaging repeated when the programmes were extended. The two-year semaglutide extension reported weight, waist circumference and cardiometabolic parameters at week 104 and did not repeat the composition substudy, so there is no imaging at all beyond seventy-two weeks in this class.4 Whatever the trajectory of lean mass is in year two of treatment, nobody has measured it.

None of this is a scandal; it is the ordinary economics of trial substudies. It becomes a problem only when a descriptive group mean is quoted as though it characterised what will happen to an individual, which is now the normal register of coverage on this subject.

The endpoint nobody measured

The clinical question is not how many kilograms of lean tissue a person has. It is whether they can climb stairs, rise from a chair without using their arms, carry shopping, and recover from an illness that keeps them in bed for a week. Those are measurable — grip strength, gait speed, chair-stand time, stair-climb power, the short physical performance battery — and they are measured routinely in geriatrics and sports science. Not one phase 3 trial in this drug class has reported them as a pre-specified endpoint.

That absence is the strongest available criticism of the programmes, and it has been made in the general medical literature by authors who are otherwise unsympathetic to muscle-loss alarmism.5 Their argument is worth stating precisely: the concern about lean-mass loss is plausible but unquantified, the instrument used to assess it is a poor proxy for the tissue of interest, and the endpoints that would settle whether it matters are cheap, validated and were simply not collected.

Where function has been measured during substantial weight loss by other routes, the results are mostly reassuring: physical performance usually improves, because carrying less mass is itself a functional benefit. That is a reasonable prior and it is not a substitute for the measurement.

1.81.30.90.400.1Total mass (s…1Appendicular …1.5Whole-body le…1.6Total fat0.4Visceral fatkilograms
Figure. Least significant change between two same-device DXA scans, by compartment, expressed in kilograms for a representative 110 kg adult. Any difference smaller than the bar is not distinguishable from measurement noise.

What the testing services can and cannot tell you here

A category confusion arrives in the Journal postbag with some regularity, and it is worth addressing directly. The four independent testing services this market relies on — Janoshik, Medutest, PeptideMeter and VendorInvestigate — analyse the contents of a vial. They report chromatographic purity, identity by mass, sometimes peptide content, and in the case of the verification services, what they were able to establish about a supplier. None of them measures anything about a person.

A certificate stating 98.7 per cent purity for a batch supplied by WWB, SSA or KP is silent on that customer’s body composition, and a low-purity result does not explain a disappointing DXA scan. The two questions are answered by different instruments in different buildings, and conflating them produces a particular kind of dead end in which somebody spends several hundred pounds on analytical testing to investigate a clinical question.

The reverse confusion also occurs: a satisfactory laboratory panel or a favourable body-composition scan is offered as evidence that a vial contained what its label claimed. It is not evidence of that either. Compounds sold for research use only are not approved for human use, and nothing in this section should be read as advice about using them.

Readers should be sceptical of any body-composition figure quoted without its instrument, and sceptical of their own scans taken less than six months apart on different machines. The measurement error in this field is not a technicality; it is comparable in size to the effects being discussed, and it is the reason the same substudy tables support opposite conclusions in different hands.

References

  1. Wilding JPH, Batterham RL, Calanna S, et al. “Once-Weekly Semaglutide in Adults with Overweight or Obesity.” New England Journal of Medicine. 2021;384(11):989–1002.
  2. Jastreboff AM, Aronne LJ, Ahmad NN, et al. “Tirzepatide Once Weekly for the Treatment of Obesity.” New England Journal of Medicine. 2022;387(3):205–216.
  3. Gastaldelli A, Cusi K, Fernández Landó L, et al. “Effect of tirzepatide versus insulin degludec on liver fat content and abdominal adipose tissue in people with type 2 diabetes (SURPASS-3 MRI): a substudy of a randomised, open-label, parallel-group, phase 3 trial.” Lancet Diabetes & Endocrinology. 2022;10(6):393–406.
  4. Garvey WT, Batterham RL, Bhatta M, et al. “Two-year effects of semaglutide in adults with overweight or obesity: the STEP 5 trial.” Nature Medicine. 2022;28(10):2083–2091.
  5. Conte C, Hall KD, Klein S. “Is Weight Loss–Induced Muscle Mass Loss Clinically Relevant?” JAMA. 2024;332(1):9–10.

Letters to the Editor

5 printed

Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.

I am sixty-eight, I have lost nineteen kilograms over fourteen months, and my consultant has twice told me my lean mass is fine on the basis of a handheld bioimpedance device in the clinic corridor. Having read your piece on what that device measures, I am no longer sure what I have been reassured about.

N. Halvorsen, Trondheim

The Journal replies

Nor are we. A handheld device measures impedance across the upper body and infers the rest, and the inference is least reliable exactly where you sit: older, substantial weight change, changing hydration. That is not a criticism of your consultant’s judgement, which may be sound on other grounds, but the device is not the evidence for it.

Small correction to your table: the S-LiTE exercise prescription was two supervised group sessions and two individual sessions weekly, not two sessions in total. The distinction matters because "add some exercise" is not what was tested.

T. Elorriaga, San Sebastián

The Journal replies

Correct, and that is precisely the point we were trying to make and then undermined in our own table. Amended.

I train four times a week, eat a hundred and sixty grams of protein and my appendicular lean mass has fallen by 1.8 kg over ten months while every lift has gone up. Your section on mass against function was the first thing I have read that made that seem normal rather than a failure.

R. Hollenbeck, Spokane, WA

Your piece treats the one-quarter rule as discredited and then quotes fractions of one third and two fifths from the substudies as though those were more solid. They are group means from a hundred and forty people. Physician, heal thyself.

H. Baptiste, Fort-de-France

The Journal replies

A fair hit, and we have amended the paragraph to carry the same caveat in both places. The distinction we should have drawn is that the substudy figures are at least attached to a stated population and a stated instrument, which the textbook rule is not. Neither is a constant.

You write that no trial has measured strength. There are observational cohorts with grip strength data. Why do you insist on randomised measurement?

K. Sivertsen, Bergen

The Journal replies

Because grip strength in an observational cohort of people who chose to take a drug, and who differ from those who did not in age, motivation and comorbidity, cannot separate the drug effect from the selection. We report those cohorts and we do not treat them as answering the question.

Related coverage