PeptideMeter result on a Zhuhai mazdutide lot lands 2.7 points below the supplier’s figure
The result is unremarkable. What the report omits is not.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Verification
Every step between the laboratory report and the product page removes information, and the badge is the last step.
The Journal’s position is that a verification badge, in the form it currently takes across this market, is not evidence. It is a claim that evidence exists somewhere, and it withholds every particular a reader would need to assess it. This is not a criticism of the laboratories, which do not design vendors’ product pages, and it is not usually a criticism of the vendors either, since the convention long predates any of them and buyers demonstrably respond to the graphic.
Since the unpublished results are invisible, any estimate of the selection effect has to be constructed rather than measured, and the Journal offers the following as an illustration rather than a finding. Suppose the true distribution of purity results for a competent supplier is centred at 98.0% with a standard deviation of 0.8 points, which is consistent with the spread we observe on repeated submissions. Suppose the supplier publishes results above 98.0% and files the rest.
The published mean is then approximately 98.6%, the published minimum is 98.0%, and the apparent variability is roughly halved. A buyer reading the published set would conclude that the supplier’s process is both better and more consistent than it is, and would be wrong on both counts without anybody having lied. Increase the publication threshold to 98.5% and the published mean rises to 99.0% while the true mean is unchanged.
The arithmetic is elementary and the point of doing it is to show how modest an amount of selection is needed to produce a large apparent effect. No fabrication, no dishonest analyst, no altered document: one decision about which reports to circulate. Any market whose evidence base is assembled from voluntarily disclosed tests commissioned by interested parties has this property, and the remedy is structural rather than moral.
A frequently overlooked asymmetry: none of these services has any authority over a vendor. A laboratory that finds a submitted sample at 91% purity cannot compel a recall, cannot require a retest, cannot publish the finding over the client’s objection without breaching confidentiality, and cannot prevent the vendor from continuing to advertise a figure obtained on a different lot. It can decline further business, which is a real sanction and a slow one.
This is not a shortcoming of the services. Confidentiality to the client is a requirement of the accreditation framework, not an indulgence, and a laboratory that published clients’ results unilaterally would be a worse institution, not a better one. But it means the word verification is doing something the underlying arrangement cannot support: verification in ordinary usage implies a check that can fail with consequences, and here the only consequence of a bad result is that nobody hears about it.
The one structural exception is the public archive. Where a service records that a submission occurred, a vendor cannot quietly discard an unfavourable result, because the fact of the test is on the record even if the detail is not. That is why this article treats the archive as the most important product feature in the sector, and why the Journal’s standing request to all four services is that the existence of a submission be public even where the result is confidential.
A laboratory cannot fail a vendor. It can only issue a report to whoever paid for it.
On the limits of the word verificationAcross the four services the Journal has catalogued eleven elements that appear on some reports and not others: the compound and lot as declared by the submitter; the date of receipt; the condition of the sample on arrival; the instrument; the column; the gradient; the detection wavelength; the integration threshold; the reproduced chromatogram; a named analyst; and an explicit statement that results apply to the sample as received.
No service omits all of these and none includes all of them on every tier. The best reports in the sector carry nine or ten and are genuinely good documents by any standard — better, in several respects, than the manufacturer certificates they are checking. The thinnest carry three, and a three-element report is a number with a letterhead.
The reporting requirements in the accreditation standard are the natural benchmark here, and they are not demanding: identification of the items tested, the date of receipt, the methods used, the results with units, and a clear statement of what the results apply to.1 A report meeting that list is checkable. The Journal’s standing request to all four services is a single addition beyond it — print the gradient and the integration threshold — because those two numbers are what make a purity figure comparable with anybody else’s.
| Element | Cheapest tier | Mid tier | Most detailed tier |
|---|---|---|---|
| Compound and lot as declared | Yes | Yes | Yes |
| Date of receipt | Sometimes | Yes | Yes |
| Condition on arrival | No | Sometimes | Yes |
| Instrument identified | No | Sometimes | Yes |
| Column identified | No | Sometimes | Yes |
| Gradient stated | No | Yes | Yes |
| Detection wavelength | Sometimes | Yes | Yes |
| Integration threshold | No | Sometimes | Yes |
| Chromatogram reproduced | No | Yes | Yes |
| Named analyst | No | Sometimes | Yes |
| “Sample as received” statement | Yes | Yes | Yes |
| Aggregated across the three assay services rather than attributed, because tier names and boundaries differ between them and a service-by-service table would invite comparison of products that are not comparable. Every element listed appears on at least one service’s standard output. | |||
A verification badge on a product listing typically asserts, in one or two words and a graphic, that the product has been tested by a named service. Consider what that claim leaves open. Which lot was tested. When. Who submitted the sample and how it was selected. What was measured — purity alone, or identity, or content. By what method. Whether the lot currently on sale is the lot that was tested. Whether the report is available to the reader.
Every one of those is material and every one is absent. A badge is therefore not evidence; it is an assertion that evidence exists, offered without the particulars that would let anybody assess it. The Journal does not treat badges as evidence in its coverage, states so wherever it reports one, and will not cite a badge as support for a claim about material.
The remedy is four data points the vendor already possesses: the lot number, the date of analysis, the service, and a link to the report. Several listings in this market already carry them, which establishes both that it is possible and that it is not commercially fatal. What it requires is that a badge become perishable — attached to a lot rather than to a product line — and perishability is precisely the property a marketing asset is designed not to have.
The decay is worth tracing precisely, because at no step does anybody say anything false. The laboratory reports a purity figure for the sample as received, on a stated date, by a stated method, submitted by a named party. The vendor extracts the figure and the service name onto its own documentation, dropping the submitter and often the method. A reseller reproduces the vendor’s documentation, dropping the date. A listing summarises the whole chain as tested by a named service. A badge reduces it to verified.
Each step is a reasonable act of summarisation and the cumulative effect is a claim of a different kind from the one the laboratory made. A time-indexed measurement on one sample has become an atemporal property of a product line. The information was not concealed; it was compressed away by a chain of parties each of whom had a legitimate reason to shorten the message.
This is why the Journal reports the provenance of every third-party figure it cites — which service, which lot, which date, which submitter — and treats a figure lacking any of those as uncitable. It makes our coverage sparser than the market’s. It also means that a number appearing in this publication can be traced to a document, which is the only property that distinguishes reporting from repetition.
First, publication of submissions rather than only of results: the date, the vendor as named by the submitter, and the compound, with the result confidential where the client requires it. This defeats most of the selection effect and costs nothing.
Second, submitter type on every report — vendor, buyer, publication or reseller — which the laboratory knows and which determines what the result can support. Third, lot-linked badges carrying a lot number, a date and a link, so that a verification claim expires with the lot it describes. Fourth, gradient and integration threshold printed on every purity report, which is the only way two figures from different services can be compared at all.
None of the four requires a regulator, new legislation, or any change in analytical practice. Three of them require a laboratory to print information it already holds; the fourth requires a vendor to accept that a badge should perish. The Journal has put all four to each of the services covered here. Responses have been mixed and mostly constructive, and are printed in our correspondence pages as they arrive. We will publish an annual note on which have been adopted, because the alternative is making the same request indefinitely without recording the answer.
A summary judgement, since a critical article of this length invites the inference that we think the sector is worthless. We do not. Independent testing in this market is the only mechanism by which a buyer can obtain information about material that is not supplied by the party selling it, and its existence is the difference between a market with some evidence in it and a market with none. Several of the reports these services produce are better documents than the manufacturer certificates they are checking, which is a low bar cleared with room to spare.
The three criticisms we would press are narrow. A sample is not a batch, and the trade cites samples as batches. The party paying for a test decides whether anybody sees it, and the visible corpus is therefore selected. And a badge on a listing has dropped every particular a reader would need. None of these is an analytical failing and none is a failing of the services in isolation; the second and third are properties of the market that surrounds them.
What we would tell a reader is this. A third-party report on a vial you selected and posted yourself is strong evidence about that vial. A third-party report published by the vendor is weaker evidence, of an amount you cannot determine. A badge is not evidence. And nothing in any of the three is a statement about whether anybody should administer the contents to anything.
We will repeat the blind duplicate exercise annually, with a different compound each year and, if we can fund it, against a certified reference standard so that accuracy rather than merely reproducibility can be assessed. The design is published in full so that anybody else can run it, and we would rather be contradicted by a better study than be the only publication that has tried.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
As a buyer I found the section on who chose the vial genuinely clarifying and slightly deflating. I have been treating vendor-published reports as equivalent to my own submissions for two years, and on your account they are not equivalent by an amount that cannot be measured.
— P. Vuković, Split
That is the correct reading, and the unmeasurable part is the honest part. We would add only that vendor-published reports are not worthless — a vendor willing to commission testing at all is behaving better than one that will not — they are simply weaker in a specific way.
Your selection-effect model assumes a supplier publishes results above a fixed threshold. Real behaviour is surely more complicated: a supplier might publish a poor result on a batch it has withdrawn, or publish everything for a period to establish credibility and then stop. The arithmetic is fine and the behavioural assumption is a cartoon.
— L. Marulanda, Medellín
Agreed, and the figure caption now says illustrative arithmetic rather than model. The point survives the simplification, which is that a small amount of selection produces a large apparent effect, but we should not have dressed a demonstration as an estimate.
VendorInvestigate does not measure anything and you have grouped it with three laboratories under the heading independent testing. That is exactly the conflation your article says the market makes.
— R. Sundaresan, Coimbatore
A fair hit. The tag under which this coverage sits predates the distinction we now draw, and we have added the distinction to the second paragraph and to the table. The department name will follow at the next reorganisation of the site.
On the archive point: a public record of submissions by vendor would be gamed within a month. Vendors would submit under the names of resellers, or through intermediaries, and the archive would show a distribution as selected as the current one but with a veneer of completeness.
— R. Mothibi, Gaborone
Probably true in part, and it is the strongest argument against our proposal. Our answer is that gaming requires effort and leaves traces, which the present arrangement does not, and that a partially gamed record is more informative than no record. We would not claim more than that.
The result is unremarkable. What the report omits is not.
A reminder that a purity figure is the output of a method, and that methods differ.
Follow the resin, not the catalogue.
Documentation practice is the only part of vendor quality a buyer can assess before purchase.
A result on one vial generalises to a batch only if the vial was drawn in a way that makes it representative — and nobody records how it was drawn.
Efficacy was never the question in this appraisal. Duration of treatment was.