SURMOUNT-3 misses on a secondary endpoint the coverage has not mentioned
A design note rather than a result: what the comparator was, and what that permits you to conclude.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Stopping
The commonest real-world strategy in this drug class is the least studied one.
What can be reasoned from the pharmacology is limited but not nothing. Exposure at steady state scales roughly with dose for these agents, and the dose-response curve for weight effect is shallow at the upper end — the increment from the second-highest to the highest dose is consistently smaller than the increments below it. That geometry suggests a reduced maintenance dose would retain most of the effect, which is a hypothesis with a mechanism, not a result. Nobody has randomised it, so nobody knows the shape of the curve on the way back down, and hysteresis is entirely possible.
STEP 4 is the cleanest test of continuation in the semaglutide programme. All participants took semaglutide through a twenty-week escalation to 2.4 mg weekly, achieving a mean reduction of approximately 10.6 per cent. They were then randomised two to one to continue semaglutide or to switch to placebo for a further forty-eight weeks, with lifestyle support maintained in both arms.1
Those who continued lost a further 7.9 per cent, reaching roughly 17.4 per cent below their original baseline at week 68. Those switched to placebo regained approximately 6.9 per cent, ending near 5 per cent below baseline. The between-group difference of about fifteen percentage points is the effect of continuing treatment for a year, measured in a population that had already demonstrated a response.
The design detail that matters most is that lifestyle support continued in the placebo arm. This is not a comparison of drug against nothing; it is a comparison of drug plus support against support alone, in people who had lost weight on the drug. The regain observed is therefore what happens with the behavioural intervention still running, which makes it a more conservative estimate of the drug contribution rather than a less one.
SURMOUNT-4 applied the same architecture to tirzepatide with a longer lead-in. Participants escalated over thirty-six weeks of open-label treatment to their maximum tolerated dose of 10 or 15 mg weekly, achieving a mean reduction of approximately 20.9 per cent, and were then randomised one to one to continue or to switch to placebo for fifty-two weeks.2
Continuation produced a further mean reduction of about 5.5 per cent, for a total near 25.3 per cent at week 88. Withdrawal produced a mean regain of about 14 per cent of body weight, leaving that arm approximately 9.9 per cent below original baseline. The between-arm difference of roughly fifteen percentage points is similar in magnitude to STEP 4 despite the much larger initial loss.
The steeper regain in absolute terms is the expected consequence of a larger loss rather than evidence of anything peculiar to the agent. It is nonetheless the figure most often quoted without its denominator, and a fourteen-point regain from a twenty-one-point loss is a materially different statement from a fourteen-point regain from a ten-point loss. Both arms in this trial ended below where they began, and the arm that stopped ended roughly where the continued arm of the semaglutide programme did.
Every withdrawal trial compared a full dose against nothing. The comparison almost every patient actually faces has never been randomised.
On the maintenance gapMean regain trajectories conceal a distribution as wide as the one that characterises the initial response. In the reported withdrawal arms, some participants returned to within a percentage point of their original baseline within a year while others retained most of their loss with no pharmacological support whatever. The interquartile ranges in the published figures are broad, and the trials were not designed to explain them.
Nothing measured at randomisation predicts an individual regain trajectory usefully. Baseline body mass index, the magnitude of the initial loss, age, sex and diabetes status all shift the mean modestly and leave the spread largely intact. The behavioural variables that plausibly matter — what a person was eating and doing during the loss, and whether any of it persisted — were not measured with the granularity required to test them.
This is the same structural gap that runs through the whole field: the mean effects are well characterised and the variance is not. It has a specific practical consequence here. A person deciding whether to stop cannot be given a personal probability of holding their weight, because no such probability has been established, and any source offering one is offering a group mean dressed as a prediction.
| Reason | Randomised evidence on outcome | Typical notice | Resumption likely? |
|---|---|---|---|
| Protocol-driven withdrawal | Three designs | Planned | Not applicable |
| Reached target weight | None | Planned | Sometimes |
| Intolerable side effects | Discontinuation rates only | Days | Sometimes, lower dose |
| Cost or coverage loss | None | Weeks or none | Often, when coverage returns |
| Supply interruption | None | None | Usually, at reset tolerability |
| Discontinuation rates for adverse events are reported in every pivotal trial; outcomes after discontinuation for the other reasons are not, because the trials did not enrol people who stopped for them. | |||
Set the three withdrawal trials side by side and a conspicuous absence appears. All three compared a full maintenance dose against placebo. None compared a full dose against a reduced one. The comparison that the great majority of successfully treated people actually face — can I take less of this and hold what I have — has not been randomised at any dose, in any programme, for any agent in this class.
The commercial explanation is straightforward and the Journal states it without much comment: a trial demonstrating that a third of the dose maintains most of the effect would reduce the revenue per treated patient by roughly the same fraction, and sponsors are not obliged to run trials against their own interest. The regulatory explanation is that maintenance dosing falls outside the approved label question, which is whether the product is effective at the studied dose.
The result is that an enormous amount of clinical practice is being conducted on inference. What can be inferred is that the dose-response curve for weight effect flattens at the top of the range, which suggests a step down would cost less than proportionally. Whether the curve is the same shape descending as ascending is unknown, and hysteresis in either direction would not be surprising.
It is worth noting what the one head-to-head weight trial in this class did and did not do. It compared two agents at their respective licensed doses and reported the difference in weight outcome; it did not establish dose equivalence between them, and it cannot be used to convert a maintenance dose of one into a maintenance dose of the other.3 Pharmacies asked to substitute during the shortage period had no equivalence basis to work from, whatever the conversion tables in circulation implied.
The Journal has asked clinicians in four jurisdictions how they manage maintenance and received a broadly consistent description that appears in no guideline. Reduce by one escalation step once the weight has been stable for a period; hold for eight to twelve weeks, which is long enough for the new exposure to reach steady state and for a trend to become visible; if the weight rises by more than a small threshold, return to the previous step. Some reduce again after a further stable interval; most do not go below the second step.
Two things recommend this approach and neither is evidence. It follows the pharmacokinetics, in that eight to twelve weeks is comfortably longer than the four to five weeks required to reach steady state at the new dose, so the observation is not being made on a still-changing exposure. And it is reversible, which a decision to stop is not in the same easy way.
The Journal reports this as description, not endorsement. It is not a dosing recommendation, no trial supports it, and the appropriate person to design a maintenance strategy is a clinician who knows the patient. We report it because a practice this widespread deserves to be described accurately rather than left to circulate in fragments.
A supply gap is a discontinuation with no notice, no plan and no taper. It differs from every other route to stopping in that it is imposed on both the patient and the prescriber, its duration is unknown at the outset, and it frequently ends as abruptly as it began. The shortage listings of recent years produced these events at population scale, and they have not been studied as a clinical exposure.
Three features make them distinctive. The patient cannot plan a maintenance strategy around an interruption of unknown length. Substitution — to a different agent, a different dose, or a compounded preparation — happens under time pressure and often without a dose-equivalence basis, since no head-to-head equivalence data exists between agents in this class. And the resumption problem described above applies in full, because the gaps were typically long enough to reset tolerability.
The Journal reported these events as they occurred and continues to think they represent the largest uncontrolled interruption experiment in the history of the class. What nobody collected was outcome data: how much weight was regained during the gaps, how many people never resumed, and what happened to the glycaemic control of those taking the drugs for diabetes rather than for weight.
Restarting after months away is well tolerated in general and the response is broadly reproducible: people who lost weight on an agent and stopped generally lose weight again on resuming, at a similar rate. There is no established phenomenon of a diminished second response in this class, and the withdrawal trials that re-offered treatment after their observation periods did not report one.
Three practical features recur. Escalation has to start again from a low dose for tolerability reasons, which means several weeks before the previous maintenance exposure is re-established. The nausea of a second escalation is frequently reported as worse than the first, for which the Journal has seen no mechanistic explanation and would not rule out reporting bias. And the weight trajectory on restarting begins from wherever the person now is, so a second course is a longer project than the first if regain was substantial.
None of this constitutes advice about whether to restart, which is a clinical decision. It is offered as a description of what the trial reports and the correspondence describe, and readers should note that no trial has been designed to study re-initiation as its primary question.
That a treatment for a chronic condition stops working when it is stopped is not a finding about the treatment. It is a finding about the condition.
On how the withdrawal trials were receivedEvery trial in this class delivers a behavioural intervention alongside the drug: energy-restriction targets, activity targets, and regular contact with a study team. That contact is itself an intervention of measurable effect, which is why placebo arms in these programmes lose two to three per cent of body weight rather than nothing. Where the behavioural component was deliberately intensified, the placebo arm lost around 5.7 per cent over sixty-eight weeks, which is a useful upper bound on what contact and counselling alone achieved in these populations.4
It matters for the withdrawal question in a way that is usually elided. The semaglutide off-treatment extension withdrew the drug and the lifestyle support together, so its regain figure describes the removal of a package.5 The STEP 4 and SURMOUNT-4 withdrawal arms kept the lifestyle component running, so their regain figures describe the removal of a molecule with support maintained.12 Those are different experiments and the second is the more conservative.
Anybody comparing regain figures across the three should therefore expect the extension to look worse, and it does. The Journal states which withdrawal design a figure comes from every time it quotes one, because the alternative is pooling two different experiments into a single number that describes neither. The same caution applies to the frequent comparison with dietary weight-loss regain, where the behavioural intervention is the whole of the treatment.
| Time since last dose | Approx. residual exposure | What is measurable |
|---|---|---|
| 1 week | ≈50% | Little change in appetite reported |
| 2 weeks | ≈25% | Appetite return commonly reported; fasting glucose rising |
| 4 weeks | ≈3–6% | Gastric emptying normalised; tolerability reset |
| 8 weeks | <1% | Weight trajectory established; HbA1c partially reflects change |
| 12 weeks | nil | HbA1c reflects the post-cessation period |
| Residual exposure assumes a 7-day half-life and first-order elimination. The observations in the third column are drawn from trial reports and correspondence and are not measurements from a single study. | ||
The Journal’s position is that three trials would resolve almost everything currently argued about in this area, and that all three are straightforward. The first is a dose-reduction design: after a lead-in to target, randomise to full dose, one step down, two steps down, or placebo, and follow for a year with weight as the primary endpoint. It would establish the shape of the descending dose-response curve and would cost a fraction of a pivotal programme.
The second is an interval design: after a lead-in, randomise to weekly, fortnightly and three-weekly administration at the same nominal dose. It would answer the intermittent-schedule question directly and would settle whether the exposure pattern matters independently of average exposure.
The third is a taper design: randomise abrupt cessation against a stepped reduction over twelve weeks, with appetite, eating behaviour and weight measured for a year afterwards. It would test the only argument for tapering that is worth testing.
None of the three is under way as far as the Journal can establish. Readers who know otherwise should write to letters@compoundjournal.com; a registered protocol for any of them would be news in this department.
Four things accompany every regain number in these pages. Which withdrawal design it comes from, because an off-treatment extension and a randomised placebo switch are different experiments. Whether the lifestyle intervention continued in the arm being described. What the denominator is — regain as a percentage of body weight, as a percentage of the weight lost, or as a final position relative to original baseline, three quantities that are routinely quoted interchangeably. And the follow-up duration, because the regain curve decelerates and a figure at six months is not a figure at a year.
The third of those is where most of the misreporting happens. A statement that participants regained two-thirds is a proportion of loss; a statement that they regained eleven per cent is a proportion of body weight; a statement that they finished 5.6 per cent below baseline is a final position. All three can describe the same arm and they are not interchangeable.
Where a source we are quoting has not stated its denominator, we say that rather than inferring it. Readers who find a regain figure in these pages without its design and its denominator have found an error, and the standards desk would like to hear about it at standards@compoundjournal.com.
This is reporting on a body of trial evidence and it is not advice about whether or how to stop taking a medicine. The decision to discontinue an agent prescribed for glycaemic control, cardiovascular risk or kidney disease is materially different from the decision to discontinue one prescribed for weight, and in every case it belongs with a clinician who has seen the person and knows why the drug was started.
Two further notes. Compounds sold for research use only are not approved for human use in any jurisdiction, and nothing here should be read as guidance about using them or about stopping their use. And where this piece describes what clinicians report doing about maintenance dosing, that is description of practice and not a schedule anybody should adopt from a magazine.
The Journal takes correspondence on this subject at letters@compoundjournal.com and factual challenges at standards@compoundjournal.com. Letters describing a personal experience of stopping are read with attention and are published, where they are published, as accounts rather than as evidence — a distinction this department tries hard to preserve in both directions.
Two practical items follow from the pharmacology rather than from the trials, and only two. An interruption long enough to clear the drug is long enough to reset tolerability, so resumption is a fresh escalation and should be planned as one. And a laboratory panel drawn less than three months after stopping will not yet show the full glycaemic consequence, whatever it turns out to be.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
A withdrawal trial answers a narrower question than it appears to. This piece states which question.
The document is four pages. Three of them are about dates.
We work through the residual-exposure table so the decision can be made from numbers rather than from feel.
Dose reduction is not withdrawal, and the trials that tested withdrawal cannot be read as testing it.
Efficacy was never the question in this appraisal. Duration of treatment was.