A design note on withdrawal trials, and why the run-in matters
The regain trajectories, arm by arm, with the estimands named.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Discontinuation
A survey of the maintenance evidence, which is shorter than the survey of the withdrawal evidence.
Here is the gap. Every randomised withdrawal trial in this class compared continued treatment at the full dose against placebo. Not one has compared continued treatment at the full dose against continued treatment at a reduced dose, which is the comparison that a person who has reached their target weight and would like to spend less money, take less drug, or feel fewer effects actually needs. The most common maintenance strategy in clinical practice is therefore the strategy with the least evidence behind it, and the disparity is not close.
STEP 4 is the cleanest test of continuation in the semaglutide programme. All participants took semaglutide through a twenty-week escalation to 2.4 mg weekly, achieving a mean reduction of approximately 10.6 per cent. They were then randomised two to one to continue semaglutide or to switch to placebo for a further forty-eight weeks, with lifestyle support maintained in both arms.1
Those who continued lost a further 7.9 per cent, reaching roughly 17.4 per cent below their original baseline at week 68. Those switched to placebo regained approximately 6.9 per cent, ending near 5 per cent below baseline. The between-group difference of about fifteen percentage points is the effect of continuing treatment for a year, measured in a population that had already demonstrated a response.
The design detail that matters most is that lifestyle support continued in the placebo arm. This is not a comparison of drug against nothing; it is a comparison of drug plus support against support alone, in people who had lost weight on the drug. The regain observed is therefore what happens with the behavioural intervention still running, which makes it a more conservative estimate of the drug contribution rather than a less one.
Set the three withdrawal trials side by side and a conspicuous absence appears. All three compared a full maintenance dose against placebo. None compared a full dose against a reduced one. The comparison that the great majority of successfully treated people actually face — can I take less of this and hold what I have — has not been randomised at any dose, in any programme, for any agent in this class.
The commercial explanation is straightforward and the Journal states it without much comment: a trial demonstrating that a third of the dose maintains most of the effect would reduce the revenue per treated patient by roughly the same fraction, and sponsors are not obliged to run trials against their own interest. The regulatory explanation is that maintenance dosing falls outside the approved label question, which is whether the product is effective at the studied dose.
The result is that an enormous amount of clinical practice is being conducted on inference. What can be inferred is that the dose-response curve for weight effect flattens at the top of the range, which suggests a step down would cost less than proportionally. Whether the curve is the same shape descending as ascending is unknown, and hysteresis in either direction would not be surprising.
It is worth noting what the one head-to-head weight trial in this class did and did not do. It compared two agents at their respective licensed doses and reported the difference in weight outcome; it did not establish dose equivalence between them, and it cannot be used to convert a maintenance dose of one into a maintenance dose of the other.2 Pharmacies asked to substitute during the shortage period had no equivalence basis to work from, whatever the conversion tables in circulation implied.
Regain as a share of loss, as a share of body weight, and as a final position relative to baseline are three numbers. They are quoted as one.
On denominatorsThe Journal has asked clinicians in four jurisdictions how they manage maintenance and received a broadly consistent description that appears in no guideline. Reduce by one escalation step once the weight has been stable for a period; hold for eight to twelve weeks, which is long enough for the new exposure to reach steady state and for a trend to become visible; if the weight rises by more than a small threshold, return to the previous step. Some reduce again after a further stable interval; most do not go below the second step.
Two things recommend this approach and neither is evidence. It follows the pharmacokinetics, in that eight to twelve weeks is comfortably longer than the four to five weeks required to reach steady state at the new dose, so the observation is not being made on a still-changing exposure. And it is reversible, which a decision to stop is not in the same easy way.
The Journal reports this as description, not endorsement. It is not a dosing recommendation, no trial supports it, and the appropriate person to design a maintenance strategy is a clinician who knows the patient. We report it because a practice this widespread deserves to be described accurately rather than left to circulate in fragments.
| Time since last dose | Approx. residual exposure | What is measurable |
|---|---|---|
| 1 week | ≈50% | Little change in appetite reported |
| 2 weeks | ≈25% | Appetite return commonly reported; fasting glucose rising |
| 4 weeks | ≈3–6% | Gastric emptying normalised; tolerability reset |
| 8 weeks | <1% | Weight trajectory established; HbA1c partially reflects change |
| 12 weeks | nil | HbA1c reflects the post-cessation period |
| Residual exposure assumes a 7-day half-life and first-order elimination. The observations in the third column are drawn from trial reports and correspondence and are not measurements from a single study. | ||
Restarting after months away is well tolerated in general and the response is broadly reproducible: people who lost weight on an agent and stopped generally lose weight again on resuming, at a similar rate. There is no established phenomenon of a diminished second response in this class, and the withdrawal trials that re-offered treatment after their observation periods did not report one.
Three practical features recur. Escalation has to start again from a low dose for tolerability reasons, which means several weeks before the previous maintenance exposure is re-established. The nausea of a second escalation is frequently reported as worse than the first, for which the Journal has seen no mechanistic explanation and would not rule out reporting bias. And the weight trajectory on restarting begins from wherever the person now is, so a second course is a longer project than the first if regain was substantial.
None of this constitutes advice about whether to restart, which is a clinical decision. It is offered as a description of what the trial reports and the correspondence describe, and readers should note that no trial has been designed to study re-initiation as its primary question.
The Journal’s position is that three trials would resolve almost everything currently argued about in this area, and that all three are straightforward. The first is a dose-reduction design: after a lead-in to target, randomise to full dose, one step down, two steps down, or placebo, and follow for a year with weight as the primary endpoint. It would establish the shape of the descending dose-response curve and would cost a fraction of a pivotal programme.
The second is an interval design: after a lead-in, randomise to weekly, fortnightly and three-weekly administration at the same nominal dose. It would answer the intermittent-schedule question directly and would settle whether the exposure pattern matters independently of average exposure.
The third is a taper design: randomise abrupt cessation against a stepped reduction over twelve weeks, with appetite, eating behaviour and weight measured for a year afterwards. It would test the only argument for tapering that is worth testing.
None of the three is under way as far as the Journal can establish. Readers who know otherwise should write to letters@compoundjournal.com; a registered protocol for any of them would be news in this department.
Two practical items follow from the pharmacology rather than from the trials, and only two. An interruption long enough to clear the drug is long enough to reset tolerability, so resumption is a fresh escalation and should be planned as one. And a laboratory panel drawn less than three months after stopping will not yet show the full glycaemic consequence, whatever it turns out to be.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
As a prescriber I would push back on your framing of the maintenance gap. We are not practising without evidence; we are practising on pharmacological inference, which is what clinicians do in every field where the trial has not been run. Calling it unevidenced makes reasonable practice sound reckless.
— H. Fitzmaurice, Preston
A fair objection and we have adjusted the wording. Our intention was to locate the absence with the people who could have funded the trial rather than with the clinicians managing without it, and on rereading the original paragraph did not achieve that.
You say no dose-equivalence data exists between agents in this class. During the shortage my pharmacy substituted one for another on the basis of a conversion table they had printed from somewhere. Where would such a table have come from?
— B. Sundqvist, Turku
Almost certainly from cross-trial comparison of weight-loss percentages, which is not an equivalence basis. There is no head-to-head dose-titration study permitting conversion between these agents, and STEP 8 — the only head-to-head weight trial we know of — compared two agents at their own licensed doses rather than establishing equivalence between them.
The regain trajectories, arm by arm, with the estimands named.
Efficacy was never the question in this appraisal. Duration of treatment was.
Reported from the sessions, and from the two hours afterwards.
The recommendation survives scrutiny. The reasoning offered for it frequently does not.
The condition on arrival is recorded by some services and not others, and it is one of the more informative lines in a report.
Lot-level verification with a public report is a real and achievable thing. It exists in this market, on a minority of listings.