STEP 1 extension data: what happens after the trial stops
A design note rather than a result: what the comparator was, and what that permits you to conclude.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Skeletal health
The recommendation survives scrutiny. The reasoning offered for it frequently does not.
Only one randomised trial has done the obvious experiment. In a Danish study of weight-loss maintenance, participants who had already lost weight on a low-energy diet were randomised to exercise alone, a GLP-1 receptor agonist alone, both together, or placebo, and followed for a year with body composition measured throughout. The combination group did better than either component on weight, on fat mass and on the proportion of the loss that was fat. It is a single trial, it used liraglutide rather than a current agent, and it remains the best evidence anybody has for the proposition that training changes the composition of pharmacological weight loss.
Every widely used body-composition instrument partitions the body into compartments, and the compartment names do more work than they should. In the standard three-compartment DXA output, a body consists of fat mass, bone mineral content and lean soft tissue. The third of those is defined by subtraction: it is what remains once fat and bone are accounted for. It therefore includes skeletal muscle, cardiac and smooth muscle, the liver, kidneys, gut and other viscera, the skin, the blood, and all extracellular and intracellular water.
The water term is the one that causes the most confusion in the first weeks of treatment. Muscle glycogen binds water at roughly three grams per gram, so a shift in glycogen stores produces a change in lean mass measurement several times its own size. Reduced food intake, reduced carbohydrate intake and reduced training volume all lower glycogen. A person who reads a two-kilogram fall in lean mass across the first month of treatment may have lost very little muscle and a good deal of water, and no instrument in routine use can tell them which.
This is not a pedantic distinction. It determines whether an early reading is alarming or unremarkable, and it is the reason the Journal treats composition measurements taken inside the first eight weeks of treatment as close to uninterpretable.
A Danish randomised trial remains the only controlled test of the obvious question. After an eight-week low-energy diet producing approximately thirteen kilograms of weight loss, participants were randomised for one year to supervised exercise alone, liraglutide 3.0 mg alone, both combined, or placebo.1 The combination arm achieved the largest weight reduction and, more relevantly here, the most favourable composition outcome: body fat percentage fell roughly twice as much in the combination group as in either single-intervention group, and the exercise arms preserved lean mass better than the drug-alone arm.
Three qualifications belong with that result. The exercise was supervised and substantial — two group sessions and two individual sessions weekly, with a vigorous-intensity target — which is not what most people mean by adding exercise. The agent was liraglutide at 3.0 mg daily, producing considerably less weight loss than the current agents, so whether the interaction scales to a twenty per cent reduction is unknown. And the trial began after weight had already been lost, so it is a maintenance study rather than an induction study.
With those stated, it is the best evidence in the field and it points in the direction the general advice already points.
A body-composition report gives four decimal places and no confidence interval. That is the whole difficulty in one sentence.
On precisionThe closest analogue to rapid weight loss in an older, heavier population predates this drug class entirely. In a randomised trial of adults aged sixty-five and over with obesity, assigned to diet, exercise, both or a control condition for a year, the combination produced the largest improvement in physical function, and the exercise component attenuated the loss of lean mass and of bone mineral density that diet alone caused.2 Diet alone improved function too — carrying less mass helps — but by less, and at a measurable skeletal cost.
That trial is the template for how the question should be asked in this class: randomise the co-intervention, measure function as a primary endpoint, measure bone, and follow for long enough for the skeleton to respond. Its population, older and heavier and losing weight quickly, resembles a large share of current incretin users far more closely than the young resistance-trained cohorts from which most consumer advice descends.
The Journal cites it frequently for that reason and notes the obvious limitation: the weight loss achieved was roughly a tenth of body mass over a year, which is half or less of what the current agents produce. Whether the protective effect of training holds at twice the rate of loss is not established.
| Target | Population it was established in | Duration | Denominator used |
|---|---|---|---|
| 0.8 g/kg/day | General adult requirement, nitrogen balance | Weeks | Current body weight |
| 1.2–1.5 g/kg/day | Older adults, energy restriction | 6–12 months | Current or adjusted weight |
| 1.6 g/kg/day | Resistance training, plateau of accrual | 8–16 weeks | Current body weight |
| 2.4 g/kg/day | Resistance-trained young men, large deficit | 4 weeks | Current body weight |
| 1.5 g/kg reference weight | Obesity management guidance | Not trial-derived | Reference or ideal weight |
| No target in this table was established in anybody taking a GLP-1 receptor agonist. The denominator column is the reason the same ratio produces targets differing by a third or more. | |||
Two claims are routinely bundled together and only one is well supported. The weaker claim is that resistance training during pharmacological weight loss builds or maintains muscle mass. In a substantial energy deficit, training generally attenuates the loss rather than preventing it, and net accrual is unusual outside of untrained beginners and the specific controlled-feeding conditions of the trials cited earlier. The stronger claim is that training preserves strength and physical function even where mass declines, which is consistently observed and is mechanistically sensible: a large part of early strength change is neural rather than structural.
The distinction has practical consequences. Somebody training hard, eating well, and watching their DXA appendicular lean mass fall by two kilograms across nine months has not failed at anything, and may be measurably stronger than at baseline. If the expectation set for them was mass preservation, they will read a normal outcome as a failure and may respond by eating more or training in ways that suit the metric rather than the goal.
The Journal reports the training recommendation and reports what it is expected to achieve, which is function first and mass second.
The word sarcopenia has migrated from clinical medicine into consumer discussion of this drug class and lost its definition in transit. In the working definitions used by the European and Asian consensus groups, sarcopenia requires low muscle strength, with low muscle quantity or quality confirming it and poor physical performance indicating severity. Strength is the entry criterion. Reduced lean mass on a scan, in the absence of measured weakness, does not meet any published definition of sarcopenia.
This matters because the borrowed term imports a prognosis. Sarcopenia in its clinical sense is associated with falls, fractures, hospitalisation and mortality, and those associations were established in older adults with measured weakness, frequently in the context of illness or immobility. Applying the label to a forty-two-year-old whose DXA appendicular lean mass has fallen by one and a half kilograms while their strength has increased is not a cautious extrapolation; it is a category error with a frightening prognosis attached.
The related term sarcopenic obesity has the same problem in a more acute form, since it requires both criteria to be met and is frequently used to mean nothing more than a low lean fraction. The Journal uses both terms only in their defined sense and asks correspondents who use them to say which criteria they mean.
Four things accompany every composition number in these pages. The instrument, because DXA, magnetic resonance, bioimpedance and creatine dilution are not interchangeable and the choice frequently determines the sign of the result. The sample size of the substudy rather than of the parent trial, because the parent trial size is irrelevant to the composition finding and quoting it is misleading. The definition used — total lean mass, lean soft tissue, appendicular lean mass or fat-free mass — because these differ by several kilograms in the same person. And whether the figure is a proportion of body mass or an absolute quantity.
Where a source omits any of the four, we say so rather than guessing, and where we have had to convert between definitions we show the conversion. This is more cumbersome than the alternative and it is the only way we have found to write about this subject without producing sentences that are technically true and practically misleading.
Readers who find a figure in these pages that lacks its instrument and its sample size have found an error, and the standards desk would like to hear about it at standards@compoundjournal.com.
A category confusion arrives in the Journal postbag with some regularity, and it is worth addressing directly. The four independent testing services this market relies on — Janoshik, Medutest, PeptideMeter and VendorInvestigate — analyse the contents of a vial. They report chromatographic purity, identity by mass, sometimes peptide content, and in the case of the verification services, what they were able to establish about a supplier. None of them measures anything about a person.
A certificate stating 98.7 per cent purity for a batch supplied by WWB, SSA or KP is silent on that customer’s body composition, and a low-purity result does not explain a disappointing DXA scan. The two questions are answered by different instruments in different buildings, and conflating them produces a particular kind of dead end in which somebody spends several hundred pounds on analytical testing to investigate a clinical question.
The reverse confusion also occurs: a satisfactory laboratory panel or a favourable body-composition scan is offered as evidence that a vial contained what its label claimed. It is not evidence of that either. Compounds sold for research use only are not approved for human use, and nothing in this section should be read as advice about using them.
Readers should be sceptical of any body-composition figure quoted without its instrument, and sceptical of their own scans taken less than six months apart on different machines. The measurement error in this field is not a technicality; it is comparable in size to the effects being discussed, and it is the reason the same substudy tables support opposite conclusions in different hands.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
Your piece treats the one-quarter rule as discredited and then quotes fractions of one third and two fifths from the substudies as though those were more solid. They are group means from a hundred and forty people. Physician, heal thyself.
— L. Whitcombe, Christchurch
A fair hit, and we have amended the paragraph to carry the same caveat in both places. The distinction we should have drawn is that the substudy figures are at least attached to a stated population and a stated instrument, which the textbook rule is not. Neither is a constant.
I train four times a week, eat a hundred and sixty grams of protein and my appendicular lean mass has fallen by 1.8 kg over ten months while every lift has gone up. Your section on mass against function was the first thing I have read that made that seem normal rather than a failure.
— J. Mbatha, Durban
As a DXA technologist of twenty-two years I would add one thing to your precision section: the largest source of error in practice is not the machine, it is positioning. A patient scanned with their arms two centimetres further from their trunk will report different regional values. We are trained to a protocol and the protocol is not always followed.
— C. Nightingale, Plymouth
We should have said this and did not. It also argues for what you presumably practise: same device, same technologist, same protocol, and a note in the record when any of those changes.
You write that no trial has measured strength. There are observational cohorts with grip strength data. Why do you insist on randomised measurement?
— J. Kettleborough, Nottingham
Because grip strength in an observational cohort of people who chose to take a drug, and who differ from those who did not in age, motivation and comorbidity, cannot separate the drug effect from the selection. We report those cohorts and we do not treat them as answering the question.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
Dose reduction is not withdrawal, and the trials that tested withdrawal cannot be read as testing it.
The gap between a defensible recommendation and a confident one is where most of the harm in this subject lives.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
The regulated world states the method, the instrument, the convention and the tolerance. This market states a verdict.
Chromatographic purity is cheap, fast and comparable-looking. Those three properties, and not its usefulness, explain why it became the industry’s single figure of merit.