An interlaboratory comparison nobody had run
The Journal submitted split samples from single lots to three assay services, under names unconnected to this publication, and published each method alongside each result.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Institutions
What a verification mark would have to carry to be checkable: a date, a lot, a method, a submitter and a link to the report.
Follow a laboratory report from the bench to a product page and watch the information fall away. The report states a method, a date, a sample identifier, a submitter and a result. The vendor extracts the result. The listing extracts the fact that a result exists. The badge extracts the word verified. At each step the claim gets shorter, more general and more durable, until what began as a measurement on one vial on one day has become a permanent property of a product line, asserted in a graphic that carries no date.
Since the unpublished results are invisible, any estimate of the selection effect has to be constructed rather than measured, and the Journal offers the following as an illustration rather than a finding. Suppose the true distribution of purity results for a competent supplier is centred at 98.0% with a standard deviation of 0.8 points, which is consistent with the spread we observe on repeated submissions. Suppose the supplier publishes results above 98.0% and files the rest.
The published mean is then approximately 98.6%, the published minimum is 98.0%, and the apparent variability is roughly halved. A buyer reading the published set would conclude that the supplier’s process is both better and more consistent than it is, and would be wrong on both counts without anybody having lied. Increase the publication threshold to 98.5% and the published mean rises to 99.0% while the true mean is unchanged.
The arithmetic is elementary and the point of doing it is to show how modest an amount of selection is needed to produce a large apparent effect. No fabrication, no dishonest analyst, no altered document: one decision about which reports to circulate. Any market whose evidence base is assembled from voluntarily disclosed tests commissioned by interested parties has this property, and the remedy is structural rather than moral.
A frequently overlooked asymmetry: none of these services has any authority over a vendor. A laboratory that finds a submitted sample at 91% purity cannot compel a recall, cannot require a retest, cannot publish the finding over the client’s objection without breaching confidentiality, and cannot prevent the vendor from continuing to advertise a figure obtained on a different lot. It can decline further business, which is a real sanction and a slow one.
This is not a shortcoming of the services. Confidentiality to the client is a requirement of the accreditation framework, not an indulgence, and a laboratory that published clients’ results unilaterally would be a worse institution, not a better one. But it means the word verification is doing something the underlying arrangement cannot support: verification in ordinary usage implies a check that can fail with consequences, and here the only consequence of a bad result is that nobody hears about it.
The one structural exception is the public archive. Where a service records that a submission occurred, a vendor cannot quietly discard an unfavourable result, because the fact of the test is on the record even if the detail is not. That is why this article treats the archive as the most important product feature in the sector, and why the Journal’s standing request to all four services is that the existence of a submission be public even where the result is confidential.
A badge is not evidence. It is an assertion that evidence exists, offered without the particulars that would let anybody assess it.
The Journal’s standing positionA verification badge on a product listing typically asserts, in one or two words and a graphic, that the product has been tested by a named service. Consider what that claim leaves open. Which lot was tested. When. Who submitted the sample and how it was selected. What was measured — purity alone, or identity, or content. By what method. Whether the lot currently on sale is the lot that was tested. Whether the report is available to the reader.
Every one of those is material and every one is absent. A badge is therefore not evidence; it is an assertion that evidence exists, offered without the particulars that would let anybody assess it. The Journal does not treat badges as evidence in its coverage, states so wherever it reports one, and will not cite a badge as support for a claim about material.
The remedy is four data points the vendor already possesses: the lot number, the date of analysis, the service, and a link to the report. Several listings in this market already carry them, which establishes both that it is possible and that it is not commercially fatal. What it requires is that a badge become perishable — attached to a lot rather than to a product line — and perishability is precisely the property a marketing asset is designed not to have.
| Element | Cheapest tier | Mid tier | Most detailed tier |
|---|---|---|---|
| Compound and lot as declared | Yes | Yes | Yes |
| Date of receipt | Sometimes | Yes | Yes |
| Condition on arrival | No | Sometimes | Yes |
| Instrument identified | No | Sometimes | Yes |
| Column identified | No | Sometimes | Yes |
| Gradient stated | No | Yes | Yes |
| Detection wavelength | Sometimes | Yes | Yes |
| Integration threshold | No | Sometimes | Yes |
| Chromatogram reproduced | No | Yes | Yes |
| Named analyst | No | Sometimes | Yes |
| “Sample as received” statement | Yes | Yes | Yes |
| Aggregated across the three assay services rather than attributed, because tier names and boundaries differ between them and a service-by-service table would invite comparison of products that are not comparable. Every element listed appears on at least one service’s standard output. | |||
The decay is worth tracing precisely, because at no step does anybody say anything false. The laboratory reports a purity figure for the sample as received, on a stated date, by a stated method, submitted by a named party. The vendor extracts the figure and the service name onto its own documentation, dropping the submitter and often the method. A reseller reproduces the vendor’s documentation, dropping the date. A listing summarises the whole chain as tested by a named service. A badge reduces it to verified.
Each step is a reasonable act of summarisation and the cumulative effect is a claim of a different kind from the one the laboratory made. A time-indexed measurement on one sample has become an atemporal property of a product line. The information was not concealed; it was compressed away by a chain of parties each of whom had a legitimate reason to shorten the message.
This is why the Journal reports the provenance of every third-party figure it cites — which service, which lot, which date, which submitter — and treats a figure lacking any of those as uncitable. It makes our coverage sparser than the market’s. It also means that a number appearing in this publication can be traced to a document, which is the only property that distinguishes reporting from repetition.
First, publication of submissions rather than only of results: the date, the vendor as named by the submitter, and the compound, with the result confidential where the client requires it. This defeats most of the selection effect and costs nothing.
Second, submitter type on every report — vendor, buyer, publication or reseller — which the laboratory knows and which determines what the result can support. Third, lot-linked badges carrying a lot number, a date and a link, so that a verification claim expires with the lot it describes. Fourth, gradient and integration threshold printed on every purity report, which is the only way two figures from different services can be compared at all.
None of the four requires a regulator, new legislation, or any change in analytical practice. Three of them require a laboratory to print information it already holds; the fourth requires a vendor to accept that a badge should perish. The Journal has put all four to each of the services covered here. Responses have been mixed and mostly constructive, and are printed in our correspondence pages as they arrive. We will publish an annual note on which have been adopted, because the alternative is making the same request indefinitely without recording the answer.
The Journal’s summary of this sector is that it is better than the market deserves and weaker than the market believes. Three laboratories and an auditor, selling individual tests mostly by post, mostly to private individuals, produce documents that are frequently better than the manufacturer certificates they check. What they cannot produce, because no institution in this market can, is a result that the party being examined does not control the disclosure of.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
Your selection-effect model assumes a supplier publishes results above a fixed threshold. Real behaviour is surely more complicated: a supplier might publish a poor result on a batch it has withdrawn, or publish everything for a period to establish credibility and then stop. The arithmetic is fine and the behavioural assumption is a cartoon.
— G. Thorbjørnsen, Tromsø
Agreed, and the figure caption now says illustrative arithmetic rather than model. The point survives the simplification, which is that a small amount of selection produces a large apparent effect, but we should not have dressed a demonstration as an estimate.
As a buyer I found the section on who chose the vial genuinely clarifying and slightly deflating. I have been treating vendor-published reports as equivalent to my own submissions for two years, and on your account they are not equivalent by an amount that cannot be measured.
— C. Tremonti, Palermo
That is the correct reading, and the unmeasurable part is the honest part. We would add only that vendor-published reports are not worthless — a vendor willing to commission testing at all is behaving better than one that will not — they are simply weaker in a specific way.
On the archive point: a public record of submissions by vendor would be gamed within a month. Vendors would submit under the names of resellers, or through intermediaries, and the archive would show a distribution as selected as the current one but with a veneer of completeness.
— H. Steinmetz, Basel
Probably true in part, and it is the strongest argument against our proposal. Our answer is that gaming requires effort and leaves traces, which the present arrangement does not, and that a partially gamed record is more informative than no record. We would not claim more than that.
VendorInvestigate does not measure anything and you have grouped it with three laboratories under the heading independent testing. That is exactly the conflation your article says the market makes.
— E. Sørheim, Stavanger
A fair hit. The tag under which this coverage sits predates the distinction we now draw, and we have added the distinction to the second paragraph and to the table. The department name will follow at the next reorganisation of the site.
I run analytical services and I object to the framing of your blind comparison. You bought our cheapest tier, published the number it produced alongside a competitor’s most thorough package, and called the result a spread. It is not a spread. It is three different products, priced accordingly, and your own table says so two columns to the right of the headline figure.
— N. Zangwill, Manchester
This is the objection we thought was most likely and we think it is partly right. The comparison is of standard products at standard prices, which is what buyers actually purchase, and we said so. But the presentation invites the reading you object to, and the figure caption now states the tier alongside each result rather than leaving it to the method columns.
The Journal submitted split samples from single lots to three assay services, under names unconnected to this publication, and published each method alongside each result.
Two signatures — one who performed the work, one who approved its release — are the ordinary regulated convention and are almost unknown here.
This piece takes the measurement apart into the decisions it is made of, because each decision moves the answer.
The Journal submitted split samples from single lots to three assay services, under names unconnected to this publication, and published each method alongside each result.
Reported from the analysis, not from a warning notice.
Follow the resin, not the catalogue.