The out-of-range flag is a statistical convention with a printer attached
A twelve-analyte panel in a perfectly healthy person has roughly even odds of producing at least one flagged result.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Assay behaviour
Timing is the whole of the post-cessation panel: draw it too early and it measures the treatment period.
The strongest argument for a laboratory panel before treatment begins is not that it will find anything. Usually it will not, and a publication that promised otherwise would be selling tests rather than reporting on them. The argument is that a baseline establishes the person’s own values, which converts every subsequent result from an interval comparison into a delta comparison — and the delta is by some distance the more informative of the two, because within-person biological variation is smaller than between-person variation for almost every analyte on a routine panel. A result of 78 means one thing in somebody whose previous value was 62 and something else entirely in somebody who has never been measured.
Two results in the same person differ for three reasons: the analyte genuinely changed, the assay is imprecise, and the analyte varies within the person from day to day. The last two are quantified in the biological variation literature as the analytical coefficient of variation and the within-subject coefficient of variation, and databases of the latter have been maintained for decades.1
The reference change value combines them: approximately 2.77 times the square root of the sum of their squares, for a two-sided ninety-five per cent probability that a difference is real. The results are instructive. Sodium, with tiny biological variation, has a reference change value of around three per cent. Creatinine is about fourteen per cent. Alanine aminotransferase, with within-subject variation above twenty per cent, requires something like a sixty per cent change. Triglycerides, more variable still, require more.
Apply that to a routine monitoring situation. An ALT moving from 28 to 41 units per litre — a rise of forty-six per cent that crosses no threshold and is unlikely to be flagged — sits inside the reference change value and may be nothing at all. An ALT moving from 28 to 62 has moved. Nothing on the report distinguishes the two cases, and the distinction is the entire question.
The tirzepatide monotherapy trial in type 2 diabetes reported HbA1c reductions of approximately 1.87 to 2.07 percentage points across its dose arms against approximately 0.04 for placebo, from a baseline of around 7.9 per cent.2 The head-to-head against semaglutide 1 mg reported reductions of approximately 2.01, 2.24 and 2.30 percentage points across tirzepatide doses against 1.86 for semaglutide, from a baseline near 8.3 per cent.3
Three things are worth extracting from those figures for a laboratory-medicine readership. The baseline value governs the achievable reduction — a trial recruiting at 8.3 per cent will report a larger fall than one recruiting at 7.9, and cross-trial comparison without baselines is uninterpretable. The reductions are far larger than the reference change value for HbA1c, so these are unambiguous signals rather than statistical artefacts. And the placebo arms moved barely at all, which tells you something about how stable the analyte is in the absence of an intervention.
Reductions of two percentage points are at the upper end of what any glucose-lowering therapy has achieved, and the Journal reports them as such while noting that they are group means from populations selected for baseline control in a defined range.
Half of an HbA1c comes from the preceding month. A panel drawn four weeks after stopping is measuring the treatment period.
On the lagSustained energy restriction produces a characteristic and benign change in thyroid function tests: triiodothyronine falls, reverse triiodothyronine rises, thyroxine changes little and thyroid-stimulating hormone falls modestly or remains unchanged. This is the low-T3 pattern of adaptation to reduced energy availability, it is not hypothyroidism, and treating it as such is an error that predates this drug class by decades.
The relevant point for monitoring is that a thyroid panel drawn during rapid weight loss will frequently show a low or low-normal free T3, and that this does not indicate thyroid disease, does not require treatment, and reverses when energy balance is restored. Thyroid-stimulating hormone remains the appropriate first-line test for suspected thyroid dysfunction; adding free T3 to a panel during active weight loss reliably generates a result that requires explaining.
Separately and unrelatedly, this class carries a boxed warning in some jurisdictions derived from rodent thyroid C-cell findings. Serum calcitonin monitoring is not recommended for that purpose, and pharmacoepidemiological work examining thyroid cancer incidence in treated populations has not established the association the rodent data raised as a possibility.4 The Journal reports the boxed warning as what it is: a precaution derived from a rodent finding whose human relevance remains unestablished.
| Measurement | Integration window | Weighting |
|---|---|---|
| Fasting glucose | Hours | Instantaneous, high day-to-day variation |
| Glycated albumin | 2–3 weeks | Roughly even |
| Fructosamine | 2–3 weeks | Roughly even |
| HbA1c | ≈120 days | ≈50% from the preceding month |
| Continuous glucose metrics | The wear period | Direct, minute by minute |
| The weighting column is why HbA1c measured monthly produces overlapping windows rather than independent observations, and why the pivotal trials scheduled it quarterly. | ||
Post-bariatric micronutrient surveillance is well founded and specific. Roux-en-Y gastric bypass bypasses the duodenum and proximal jejunum, which are the principal absorption sites for iron, calcium and several B vitamins; reduced gastric acid impairs the release of food-bound B12 and the reduction of ferric iron; and the intrinsic-factor pathway is compromised by the reduction in parietal cell mass. Each of those is an identified mechanism supporting a specific test at a specific interval.
None of them applies to a receptor agonist. The gastrointestinal tract is anatomically intact, acid secretion is broadly preserved, and no absorption site is bypassed. The mechanism that does apply is reduced intake, which predicts deficiency in proportion to dietary inadequacy rather than in the bariatric pattern. Those two predictions differ: a person eating half as much of a varied diet is at different risk from a person whose duodenum has been bypassed, and the appropriate surveillance is not obviously the same.
The Journal has looked for a cohort study characterising micronutrient status in this population at twelve months or beyond and has not found one. Until one exists, monitoring schedules for this drug class are precautionary extrapolation. That is a defensible thing to do and it should be described accurately rather than presented as protocol.
A baseline panel rarely finds anything. Its value is almost entirely in what it makes possible later: within-person comparison, which for nearly every analyte on a routine panel is a more sensitive instrument than comparison against a reference interval, because within-subject biological variation is smaller than between-subject variation.
The arithmetic behind that is worth stating. For an analyte where the within-subject coefficient of variation is substantially smaller than the between-subject value — a condition satisfied by creatinine, the liver enzymes, HbA1c, the thyroid hormones and most of the electrolytes — a person’s own previous result is a better comparator than the population interval. The index of individuality formalises this, and for the analytes in question it says clearly that population intervals are relatively insensitive to change in an individual.
The practical consequence is that a person with a baseline creatinine of 62 whose value is now 78 has information that a person presenting with 78 and no baseline does not, even though both results sit inside every reference interval in use. That is the whole argument for the baseline panel, and it is a stronger argument than the one usually offered, which is that the panel might find an undiagnosed problem.
Frequent monitoring in a person doing well is a reliable generator of work. Each comprehensive panel carries a substantial probability of at least one flagged result; the flags are mostly noise; each requires explanation, repetition or investigation; and the cumulative effect over a year of monthly panels is several investigations and no additional information about the person.
There is also a specific problem with monitoring an analyte more frequently than its own window. HbA1c integrates three months. Measuring it monthly produces overlapping windows in which two-thirds of the data is shared between consecutive results, so the apparent trend is smoother than the underlying glycaemia and the independent information per measurement is low. The trials in this class scheduled it quarterly for exactly this reason.
The counter-argument deserves stating fairly: monitoring during escalation, when tolerability problems and their metabolic consequences are most likely, is a different proposition from monitoring during stable maintenance, and the case for closer observation in the first three months is reasonable. What the Journal has not seen is any evidence that a fixed frequent schedule during maintenance detects anything that a symptom-prompted panel would miss. Readers who know of such evidence should write to standards@compoundjournal.com.
Timing is almost the whole of this. A panel drawn four weeks after a final injection is largely measuring the treatment period, because HbA1c integrates three months and the drug was present for most of them. A panel drawn at twelve weeks reflects the post-cessation period for HbA1c and reflects it fully at sixteen. Fasting glucose responds within days to weeks and is therefore the earlier indicator, at the cost of much larger within-person variation.
The other analytes have their own timescales. Alanine aminotransferase responds over weeks to months as hepatic fat returns with weight. Triglycerides respond quickly and noisily. Creatinine drifts back as lean mass is regained, which means estimated glomerular filtration rate falls during regain for the same non-renal reason it rose during loss. Blood pressure, which is not a laboratory measurement but travels with these panels, reverts over weeks.
The commonest misreading the Journal encounters in correspondence is a person concluding from a reassuring panel at four to six weeks after stopping that the metabolic consequences of cessation are smaller than they were told to expect. At that interval the panel cannot have shown them. The finding at twelve weeks is frequently different, and it is the one worth waiting for.
A lipase at twice the upper limit in an asymptomatic person is a common finding with no established significance.
On pancreatic enzymesA distinction has to be drawn firmly because the postbag suggests it frequently is not. The four independent testing services this market relies on — Janoshik, Medutest, PeptideMeter and VendorInvestigate — analyse material. They report chromatographic purity, identity by mass, peptide content where it is measured, and in the case of the verification services what could be established about a supplier. A clinical laboratory analyses a person. The two produce documents that superficially resemble each other and answer entirely unrelated questions.
A purity certificate reporting 99.1 per cent for a batch from WWB, CPC or QYB tells you nothing about anybody liver enzymes. A normal panel does not confirm that a vial contained what its label claimed, and an abnormal one does not establish that it did not. Where a person suspects a supply problem, the instrument for that is analytical testing of the material; where a person has an abnormal laboratory result, the instrument is clinical assessment. Substituting one for the other is a reliable way to spend money and learn nothing.
Compounds sold for research use only are not approved for human use in any jurisdiction, and nothing in this department should be read as guidance about using them or about monitoring their use.
Five things accompany a laboratory number in these pages. The units, because international and conventional units differ for several analytes and the same value means different things in each. The reference interval used, with a note where the interval is contested, as it is for alanine aminotransferase. The baseline, because a change of 1.8 percentage points in HbA1c from a starting value of 8.3 is a different claim from the same change from 9.5. The estimand where the figure comes from a trial. And the reference change value where we are discussing an individual delta rather than a group mean.
We also state the assay method where it matters, which is more often than one would like: HbA1c in the presence of a haemoglobin variant, thyroid function in the presence of interfering antibodies, and creatinine measured by enzymatic against Jaffe methods all behave differently, and a comparison across methods is not a comparison.
This is a heavier apparatus than most publications carry and it exists because the alternative, in our experience, is a stream of technically accurate figures that lead readers to conclusions the data does not support. Errors in this apparatus should be reported to standards@compoundjournal.com; the correction log records what came of each one.
The Laboratory Notebook reports what tests measure, how they behave, and what has been found using them. It does not recommend monitoring schedules, interpret readers’ results, or advise on treatment. A laboratory result belongs in a conversation with a clinician who has the rest of the picture, and this publication is emphatically not that conversation.
Two standing notes. Several compounds discussed in these pages are sold for research use only and are not approved for human use in any jurisdiction; the Journal reports on them as commodities and as analytical problems, not as therapies. And where we describe what the pivotal trials monitored, that is reporting on trial protocols and not a template anybody should adopt from a magazine.
Correspondence is welcome at letters@compoundjournal.com. The Journal receives a steady flow of letters containing readers’ own panel results with a request for interpretation, and we do not provide it — not from caution but because a panel without a history, an examination and a reason for ordering it cannot be interpreted by anybody, including us.
The Journal’s position on monitoring is that the panel is the easy part and the interpretation is the hard part, and that almost all of the available effort goes into the first. A baseline panel is worth having because it converts every subsequent result from a population comparison into a personal one. Beyond that, the frequency at which most people are being tested generates flags faster than information, and there is no evidence that it detects anything a symptom-prompted panel would miss.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
My HbA1c was 6.1 before starting and 6.0 after four months. My clinic recorded this as no improvement in glycaemic control. My continuous monitor says my average glucose fell by a fifth over the same period. Which is measuring what?
— H. Barreto, Recife
Both are measuring correctly and the discrepancy is worth pursuing with your clinic rather than with us. A 0.1-point change is well inside the reference change value for HbA1c, so the assay has not detected a change; whether that is because the change is genuinely small or because something is affecting your glycation is not answerable from the numbers alone.
You describe the low-T3 pattern during energy restriction as benign adaptation. There is a body of opinion holding that it represents a genuine hypometabolic state requiring treatment. I do not hold that view but your readers will encounter it.
— C. Wilcoxson, Des Moines, IA
They will, and we should have named it in order to say why we do not report it. There is no randomised evidence that treating the low-T3 pattern of energy restriction improves any outcome, and there is a mechanistic argument that suppressing an adaptive response is unlikely to help. We report it as adaptation for those reasons and would report a trial that changed the picture.
You give the reference change value for ALT as about sixty per cent, which strikes me as so large as to make routine monitoring of it pointless. Is that your position?
— M. Ferrari, Trieste
Not quite. It makes monitoring for small movements pointless, which is different. A doubling is well outside the RCV and is a real signal; a rise from 28 to 41 is not. The value of the test lies in detecting the former, and much of the anxiety it generates comes from acting on the latter.
As a biomedical scientist I would add one point to your reference-interval section: many laboratories do not derive their own intervals at all. They adopt the manufacturer interval for the platform, which was established in a population that may have nothing to do with the one being tested. The interval on the report can be a document about a different country.
— K. Rautio, Tampere
This is correct, common, and something we should have stated. We have added it, and it strengthens rather than weakens the argument for within-person comparison.
Your piece assumes readers have a clinician ordering panels for them. A great many of us are buying our own, from services that supply an interval, a flag and nothing else. The apparatus you describe is not available to us at all.
— N. Fairweather, Hamilton
It is available, with effort: published biological variation databases are open, the reference change value is one line of arithmetic, and your own previous result is the comparator that does most of the work. But you are right that the market supplying these panels has no incentive to include any of it, and that is worth stating.
A twelve-analyte panel in a perfectly healthy person has roughly even odds of producing at least one flagged result.
The panel drawn during a week of vomiting is measuring the vomiting.
A substantial proportion of the abnormal results generated during rapid weight loss are consequences of the weight loss rather than findings about the person.
A flag is a probability statement about a population. It is not a statement about the person holding the printout.
A pharmacist catches most of these in licensed practice. In this market there is no pharmacist, so the checks have to be structural.
Sharps disposal is a legal obligation in most jurisdictions and a safety obligation everywhere. Household waste is not a route.