Side effects, cost, supply, target: four reasons with four trajectories
Real-world persistence figures, with their definitions stated, because the definitions are doing most of the work.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Measurement
A body-composition report gives four decimal places and no confidence interval. That is the whole difficulty in one sentence.
The Journal has asked four separate imaging physicists the same question over the past year: given the best clinical DXA in routine use, what is the smallest change in appendicular lean mass you would report to a patient as real? The answers clustered between six hundred grams and one and a half kilograms, depending on the machine, the operator, the positioning protocol and whether the two scans were performed on the same device. Nobody said less than half a kilogram. That figure should be printed at the top of every body-composition report and is printed on none of them.
Every widely used body-composition instrument partitions the body into compartments, and the compartment names do more work than they should. In the standard three-compartment DXA output, a body consists of fat mass, bone mineral content and lean soft tissue. The third of those is defined by subtraction: it is what remains once fat and bone are accounted for. It therefore includes skeletal muscle, cardiac and smooth muscle, the liver, kidneys, gut and other viscera, the skin, the blood, and all extracellular and intracellular water.
The water term is the one that causes the most confusion in the first weeks of treatment. Muscle glycogen binds water at roughly three grams per gram, so a shift in glycogen stores produces a change in lean mass measurement several times its own size. Reduced food intake, reduced carbohydrate intake and reduced training volume all lower glycogen. A person who reads a two-kilogram fall in lean mass across the first month of treatment may have lost very little muscle and a good deal of water, and no instrument in routine use can tell them which.
This is not a pedantic distinction. It determines whether an early reading is alarming or unremarkable, and it is the reason the Journal treats composition measurements taken inside the first eight weeks of treatment as close to uninterpretable.
Bioelectrical impedance analysis passes a small alternating current through the body and measures the opposition to it. Lean tissue, being largely water and electrolyte, conducts; fat does not. From the measured impedance, a height term, a weight term and a set of population-derived regression equations, the device produces a fat mass figure. The impedance is measured. The body composition is computed from an equation fitted to somebody else.
The consequences are well documented. Agreement with DXA at the group level is often reasonable; agreement at the individual level is not, with limits of agreement for fat mass frequently spanning several kilograms in either direction, and the disagreement growing at higher body mass index — precisely the population of interest here.1 Worse for our purposes, the measurement is sensitive to hydration status, recent exercise, recent meals, ambient temperature, skin moisture and time of day, all of which are changing during incretin treatment. A device that reads fat mass as a function of body water, used in a person whose body water is unstable, will report composition changes that are hydration changes. The Journal does not report BIA-derived composition changes from consumer devices, and would not treat them as evidence of anything.
Report lean mass as a proportion and it rises. Report it in kilograms and it falls. Selecting the framing selects the conclusion.
On denominatorsThere is a technique that estimates whole-body skeletal muscle mass rather than inferring it from a subtraction. Deuterated creatine dilution involves an oral dose of labelled creatine, which distributes into the total creatine pool — almost all of which sits in skeletal muscle — with the enrichment of labelled creatinine in a subsequent urine sample giving an estimate of pool size and therefore of muscle mass.2 It is not an imaging measure and it does not depend on regression equations fitted to a reference population.
Comparisons with DXA are instructive and slightly deflating. The two methods correlate only moderately in older adults, and where they disagree the creatine-dilution figure has been the better predictor of physical function and of incident disability. That is an argument that DXA appendicular lean mass, the standard proxy, is measuring something adjacent to what matters rather than the thing itself.
The method has been available for more than a decade. It has been used in no trial of any drug in this class. It requires a timed urine collection and a mass spectrometry laboratory, which is a modest imposition set against the volume of argument the absence of good muscle-mass data has generated.
| Endpoint | Measured in a randomised trial? | Where |
|---|---|---|
| Areal BMD, hip and spine | Yes, as a secondary analysis | S-LiTE bone analysis |
| Bone turnover markers | Yes, small studies | Investigator-initiated |
| Bone geometry or microarchitecture | No | — |
| Incident fracture | No | — |
| Falls | No | — |
| Absence from this table means the Journal could not find a pre-specified randomised measurement, not that no observational data exists. Observational fracture data in weight loss is confounded in both directions. | ||
Clinical teaching has long held that approximately twenty-five per cent of the mass lost during weight reduction is fat-free tissue. The figure appears in textbooks, in review articles and in a great deal of consumer material, usually without a citation and always without an interval.
A critical review published in 2014 traced the rule to a limited number of older studies, examined the variation across the wider literature, and concluded that treating one-quarter as a constant is not defensible.3 The fraction of loss that is fat-free tissue varies systematically with baseline adiposity — heavier people lose proportionally more fat — and with the rate of loss, the protein intake, the activity pattern and the measurement method. Reported values span from well under fifteen per cent to above thirty-five.
This matters for the current argument in a specific way. Both the reassuring and the alarming readings of the incretin substudy data are constructed by comparing an observed fat-free fraction against the one-quarter benchmark. If the benchmark is a loose average rather than an expectation, both comparisons are weaker than they appear, and the honest statement is that the observed fractions sit within the range that dietary weight loss has always produced.
There is a rhetorical move available to both sides of this argument and it works by choosing a denominator. Report lean mass as a proportion of total body mass and it rises during successful treatment, because fat is falling faster; the treatment looks composition-improving, which it is. Report lean mass in absolute kilograms and it falls; the treatment looks muscle-costing, which it also is. Both statements can be made from the same scan pair without either being false.
The Journal reports both, in that order, and thinks anybody presenting only one should be asked why. The proportional figure is the right one for questions about metabolic quality: a body with a higher lean fraction handles glucose better and carries less ectopic fat. The absolute figure is the right one for questions about function and reserve, because a hip fracture at seventy-eight is not prevented by a favourable ratio.
The two framings also diverge most sharply exactly where the stakes are highest. A person losing twenty-five per cent of their body weight will show an excellent proportional result and the largest absolute lean-mass reduction in the cohort. Selecting the framing selects the conclusion, which is why the trade has settled on whichever one suits it.
Densitometry infers bone mineral density from the differential attenuation of two X-ray energies, using the surrounding soft tissue as the baseline against which bone is distinguished. The algorithm assumes a soft-tissue composition, and that assumption is embedded in the calibration. When the thickness and fat fraction of the tissue overlying a measurement site change substantially, part of the apparent change in bone density is an artefact of the altered baseline.
The magnitude is contested. Phantom and cadaver work suggests errors of the order of one to three per cent for large changes in overlying fat, which is the same order as the real bone changes being reported over a year of rapid weight loss. In practice this means that a hip bone mineral density reduction of two per cent in a person who has lost a fifth of their body weight cannot be cleanly separated into a bone effect and a measurement effect, and the published analyses do not attempt it.
Quantitative computed tomography and high-resolution peripheral imaging are less vulnerable, measure geometry and microarchitecture rather than areal density, and have not been used in any trial in this class. The Journal regards that as the most easily closed gap in the whole body-composition literature.
What would change our reporting is a single trial: current agent, pre-specified strength and physical-function endpoints, randomised co-intervention, bone imaging that is not confounded by soft-tissue change, and a follow-up long enough for the skeleton to respond. It would cost a fraction of what the parent programmes cost. Its absence, four years into the largest voluntary weight-loss experiment in medical history, is the finding this department keeps returning to.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
As a DXA technologist of twenty-two years I would add one thing to your precision section: the largest source of error in practice is not the machine, it is positioning. A patient scanned with their arms two centimetres further from their trunk will report different regional values. We are trained to a protocol and the protocol is not always followed.
— E. Sørheim, Stavanger
We should have said this and did not. It also argues for what you presumably practise: same device, same technologist, same protocol, and a note in the record when any of those changes.
You write that no trial has measured strength. There are observational cohorts with grip strength data. Why do you insist on randomised measurement?
— H. Steinmetz, Basel
Because grip strength in an observational cohort of people who chose to take a drug, and who differ from those who did not in age, motivation and comorbidity, cannot separate the drug effect from the selection. We report those cohorts and we do not treat them as answering the question.
Your piece treats the one-quarter rule as discredited and then quotes fractions of one third and two fifths from the substudies as though those were more solid. They are group means from a hundred and forty people. Physician, heal thyself.
— C. Tremonti, Palermo
A fair hit, and we have amended the paragraph to carry the same caveat in both places. The distinction we should have drawn is that the substudy figures are at least attached to a stated population and a stated instrument, which the textbook rule is not. Neither is a constant.
I train four times a week, eat a hundred and sixty grams of protein and my appendicular lean mass has fallen by 1.8 kg over ten months while every lift has gone up. Your section on mass against function was the first thing I have read that made that seem normal rather than a failure.
— G. Thorbjørnsen, Tromsø
Real-world persistence figures, with their definitions stated, because the definitions are doing most of the work.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
The evidence base is one secondary analysis, several small studies and a large amount of extrapolation from bariatric surgery.
Cost is the modal reason for discontinuation in every dataset we have seen, and it is absent from the clinical literature.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
The assay is not the problem. The interpretation of a lagging integral as a current measurement is the problem.