The limits of mechanism: why receptor data will not tell you who responds
The mechanism is well described. The variance is not.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Panels
Two sources of noise sit under every number: how reproducible the assay is, and how much the analyte varies within the same person on the same day.
The Journal reports laboratory findings from the trials constantly and has come to regard the interpretation of an individual panel as the most consistently mishandled subject in this whole field. The reason is structural rather than educational: the printout gives a number, an interval and a flag, and gives no imprecision estimate, no within-person variation figure and no reference change value. Everything required to interpret the result correctly is omitted from the document that reports it.
A reference interval is an empirical statement about a population. A laboratory recruits a reference group meeting defined health criteria, measures the analyte, and reports the central ninety-five per cent of the resulting distribution, usually as the 2.5th to 97.5th percentiles. Everything about that construction has consequences. The interval is specific to the assay and platform used to derive it. It is specific to the reference population — its age structure, sex distribution, ethnicity and, for some analytes, its diet and altitude. And it deliberately excludes one in twenty healthy people at each end by design.
Two further points follow. The interval is not a target: for several analytes the optimal value on outcome grounds sits well inside it or below it, and low-density lipoprotein cholesterol is the standard example. And it is not a diagnostic threshold: decision limits, which are what clinical guidelines actually use, are derived from outcome data rather than from a healthy distribution, which is why the diagnostic cut-off for diabetes is not the upper limit of a reference interval.
Laboratories that report both a reference interval and a decision limit are doing the reader a service. Most report one number and one flag, and leave the distinction to be inferred.
If an analyte’s reference interval excludes five per cent of healthy people, and if analytes were independent, the probability that a healthy person produces at least one flagged result on a panel of n analytes is one minus 0.95 to the power of n. For a twelve-analyte panel that is approximately 46 per cent. For twenty analytes, approximately 64 per cent. For a thirty-analyte comprehensive panel with lipids and thyroid included, approximately 79 per cent.
Analytes are not independent — electrolytes covary, liver enzymes covary, so the true figures are somewhat lower — but the direction and rough magnitude hold. The practical implication is uncomfortable and rarely stated: on a comprehensive panel, the flagged result is the normal outcome, and treating each flag as requiring explanation is a commitment to explaining noise.
This is the strongest single argument for ordering panels against questions rather than by habit. A panel assembled because each analyte answers something the clinician wants to know produces flags that mean something. A panel assembled because it comes as a bundle produces a document in which the interesting result, if there is one, is hidden among four uninteresting ones. The Journal makes this point in a publication whose readers frequently order their own panels privately, and it applies with more force there rather than less.
A lipase at twice the upper limit in an asymptomatic person is a common finding with no established significance.
On pancreatic enzymesTwo results in the same person differ for three reasons: the analyte genuinely changed, the assay is imprecise, and the analyte varies within the person from day to day. The last two are quantified in the biological variation literature as the analytical coefficient of variation and the within-subject coefficient of variation, and databases of the latter have been maintained for decades.1
The reference change value combines them: approximately 2.77 times the square root of the sum of their squares, for a two-sided ninety-five per cent probability that a difference is real. The results are instructive. Sodium, with tiny biological variation, has a reference change value of around three per cent. Creatinine is about fourteen per cent. Alanine aminotransferase, with within-subject variation above twenty per cent, requires something like a sixty per cent change. Triglycerides, more variable still, require more.
Apply that to a routine monitoring situation. An ALT moving from 28 to 41 units per litre — a rise of forty-six per cent that crosses no threshold and is unlikely to be flagged — sits inside the reference change value and may be nothing at all. An ALT moving from 28 to 62 has moved. Nothing on the report distinguishes the two cases, and the distinction is the entire question.
| Analyte | Analytical CV | Within-subject CV | Reference change value |
|---|---|---|---|
| Sodium | 0.8% | 0.7% | ≈3% |
| HbA1c | 2.0% | 1.7% | ≈7% relative |
| Creatinine | 2.5% | 4.5% | ≈14% |
| Alanine aminotransferase | 5% | 20% | ≈57% |
| Triglycerides | 3% | 12% | ≈34% |
| Thyroid-stimulating hormone | 6% | 17% | ≈50% |
| Ferritin | 4% | 13% | ≈38% |
| Coefficients are representative values from published biological variation databases and differ between laboratories and platforms. The RCV column is calculated as 2.77 times the root sum of squares and is rounded. | |||
Most clinical laboratories report an upper limit of normal for alanine aminotransferase somewhere between about 40 and 55 units per litre, with a modest sex difference or none. Those intervals were derived from reference populations that were screened for viral hepatitis and heavy alcohol use but not, in most cases, for hepatic steatosis — which was neither commonly diagnosed nor considered when many of the intervals were established.
Work redefining the healthy range in a large population of prospective blood donors, screened for viral markers, alcohol intake and metabolic risk factors, arrived at substantially lower limits: in the region of 30 units per litre for men and around 19 for women.2 Those figures have been influential in hepatology and have largely not propagated into general laboratory reporting.
The consequence for this population is direct. A person starting treatment with an ALT of 44 has a flagged result by a strict standard and an unflagged one by their laboratory interval; a fall to 31 during treatment represents normalisation by one standard and continued abnormality by the other. Neither reading is wrong. The Journal reports ALT against both where it can, and regards a laboratory report giving only the wider interval as incomplete rather than incorrect.
Frequent monitoring in a person doing well is a reliable generator of work. Each comprehensive panel carries a substantial probability of at least one flagged result; the flags are mostly noise; each requires explanation, repetition or investigation; and the cumulative effect over a year of monthly panels is several investigations and no additional information about the person.
There is also a specific problem with monitoring an analyte more frequently than its own window. HbA1c integrates three months. Measuring it monthly produces overlapping windows in which two-thirds of the data is shared between consecutive results, so the apparent trend is smoother than the underlying glycaemia and the independent information per measurement is low. The trials in this class scheduled it quarterly for exactly this reason.
The counter-argument deserves stating fairly: monitoring during escalation, when tolerability problems and their metabolic consequences are most likely, is a different proposition from monitoring during stable maintenance, and the case for closer observation in the first three months is reasonable. What the Journal has not seen is any evidence that a fixed frequent schedule during maintenance detects anything that a symptom-prompted panel would miss. Readers who know of such evidence should write to standards@compoundjournal.com.
Readers who order their own panels privately, which a substantial fraction of this publication’s readership does, are in a specific position: they have the data and not the interpretive apparatus, and the apparatus is where the value is. The reference change value for the analyte in question, and the person’s own previous result, do more interpretive work than any reference interval on the printout.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
A small thing but it matters in practice: your table gives amylase and lipase rising as a finding of unclear significance. In my laboratory we no longer report amylase at all for suspected pancreatitis, because lipase is more sensitive and more specific and having both invites the wrong one to be acted on.
— H. Baptiste, Fort-de-France
A reasonable position and increasingly the standard one. We report amylase because the trial data reported it, not because we think it should be ordered.
My ferritin fell from 118 to 34 across a year on treatment and I was told this was expected because inflammation had fallen. Six months later I was clearly iron deficient. I appreciate the section on this being ambiguous, but the ambiguity was resolved in one direction and nobody looked.
— R. Hollenbeck, Spokane, WA
That is the failure mode the ambiguity produces, and it is the reason a transferrin saturation alongside costs almost nothing and resolves the question. We are sorry it went that way.
My HbA1c was 6.1 before starting and 6.0 after four months. My clinic recorded this as no improvement in glycaemic control. My continuous monitor says my average glucose fell by a fifth over the same period. Which is measuring what?
— T. Elorriaga, San Sebastián
Both are measuring correctly and the discrepancy is worth pursuing with your clinic rather than with us. A 0.1-point change is well inside the reference change value for HbA1c, so the assay has not detected a change; whether that is because the change is genuinely small or because something is affecting your glycation is not answerable from the numbers alone.
You describe the low-T3 pattern during energy restriction as benign adaptation. There is a body of opinion holding that it represents a genuine hypometabolic state requiring treatment. I do not hold that view but your readers will encounter it.
— N. Halvorsen, Trondheim
They will, and we should have named it in order to say why we do not report it. There is no randomised evidence that treating the low-T3 pattern of energy restriction improves any outcome, and there is a mechanistic argument that suppressing an adaptive response is unlikely to help. We report it as adaptation for those reasons and would report a trial that changed the picture.
You give the reference change value for ALT as about sixty per cent, which strikes me as so large as to make routine monitoring of it pointless. Is that your position?
— S. Bergqvist, Malmö
Not quite. It makes monitoring for small movements pointless, which is different. A doubling is well outside the RCV and is a real signal; a rise from 28 to 41 is not. The value of the test lies in detecting the former, and much of the anxiety it generates comes from acting on the latter.
The mechanism is well described. The variance is not.
The dose-response curve in this class flattens near its top. That has direct consequences for whether the final rung is worth climbing.
The commonest real-world strategy in this drug class is the least studied one.
What the trials measured, which in the case of micronutrients is very little.
What the Journal would want measured before treating this as settled in either direction.
Water is a reactant in hydrolysis, a plasticiser that mobilises the amorphous matrix, and a carrier for everything else. All three matter.