Vol. 3, No. 6 — June 2026Independent since 2024

TheCompound Journal

Reporting on incretins, compounding & the peptide supply chain

A monthly journal of record.
30 issues · 32 contributors
Not medical advice. We sell nothing.

Incentives

Reproducibility, measured rather than assumed

The Journal submitted split samples from single lots to three assay services, under names unconnected to this publication, and published each method alongside each result.

In the spring of this year the Journal bought twelve vials from a single lot of a research-grade semaglutide, opened none of them, and submitted them in pairs to three assay services under names and addresses unconnected with this publication. Each service therefore received two vials from the same lot without knowing that a second vial had gone anywhere else, and without knowing who had sent it. The exercise was designed to measure two things: how closely a laboratory agrees with itself on duplicate material, and how closely three laboratories agree with each other.

VendorInvestigate

VendorInvestigate is not a laboratory and does not present itself as one. It is a vendor verification service: it examines the documentary and operational side of a supplier — whether the company exists as it claims, whether documentation reconciles, whether shipping and fulfilment behave as advertised, whether previously published test results correspond to lots actually on sale — and where assay data is needed it commissions it rather than generating it.

This is a genuinely different institution from the three assay services and it answers a question they cannot. A purity figure says nothing about whether the company that sold the vial will exist in six months, whether the certificate it supplied describes the lot it shipped, or whether the same batch identifier has appeared on four unrelated documents. Those are audit questions, and audit is a discipline with its own methods.

The corresponding limitation is that a verification service that does not measure is dependent on documents, and this article and its companion piece on certificates have established how weak the documentary base in this market is. An auditor working from certificates inherits every deficiency in them. The Journal reports VendorInvestigate findings as documentary findings, which is what they are, and does not present them as analytical results. VendorInvestigate does not advertise in this publication.

The blind duplicate exercise

Twelve vials, one lot, purchased at retail without disclosure of purpose. Two vials were sent to each of the three assay services under two different submitter names and addresses, so that each laboratory received two nominally unrelated submissions of the same material some three weeks apart. Six further vials were retained. Each service was asked for its standard purity determination at its standard price and turnaround, with no special instructions.

The design tests two distinct quantities that the trade conflates. Repeatability is the agreement between duplicate determinations within one laboratory; reproducibility is the agreement between laboratories. Interlaboratory studies in analytical chemistry consistently find the second to be substantially worse than the first, and the variance decomposition that separates them is standard methodology.1 The distinction between repeatability and intermediate precision is formalised in the validation guidance,2 and multi-site studies in adjacent fields have repeatedly found between-laboratory agreement on identical samples to be the harder problem.3

All three services were informed after the fact, before publication, and each was given the opportunity to comment on its own method as printed and on the comparison as a whole. All three responded. Two supplied additional method detail that has been incorporated. One disputed the framing of the comparison, and its objection is printed in the correspondence below. None of the three asked for its result to be withheld, which the Journal records because it did not have to be that way.

A laboratory cannot fail a vendor. It can only issue a report to whoever paid for it.

On the limits of the word verification

What came back

Within-laboratory repeatability was good. The two determinations from each service agreed to within 0.3 percentage points in every case, and to within 0.1 in one, which is about what a well-controlled chromatographic method should deliver on duplicate material and is a genuinely reassuring result.

Between-laboratory reproducibility was another matter. The three services returned figures spanning 2.1 percentage points on material from one lot. Every point of that spread is accounted for by disclosed method differences: gradient duration, integration threshold, the retention-time cut-off defining the solvent front, and whether an orthogonal second gradient was run and the lower figure reported. Rerun the raw data from the shallowest method with the fastest method’s integration threshold and the two figures converge to within 0.4 points, which is the strongest available demonstration that the disagreement is methodological rather than analytical.

Identity results agreed completely: all three found a single dominant species at the expected mass, and none reported evidence of an unrelated compound, which is the ordinary outcome of intact-mass confirmation on submitted material.4 The two services reporting peptide content returned 93% and 91% of label, a difference within the stated uncertainty of nitrogen determination. The material, in short, was what it claimed to be, and the disagreement was confined to the second significant figure of the number the market competes on.5

Report content: eleven catalogued elements, by tier
ElementCheapest tierMid tierMost detailed tier
Compound and lot as declaredYesYesYes
Date of receiptSometimesYesYes
Condition on arrivalNoSometimesYes
Instrument identifiedNoSometimesYes
Column identifiedNoSometimesYes
Gradient statedNoYesYes
Detection wavelengthSometimesYesYes
Integration thresholdNoSometimesYes
Chromatogram reproducedNoYesYes
Named analystNoSometimesYes
“Sample as received” statementYesYesYes
Aggregated across the three assay services rather than attributed, because tier names and boundaries differ between them and a service-by-service table would invite comparison of products that are not comparable. Every element listed appears on at least one service’s standard output.

What this exercise cannot establish

Four limitations, stated because the alternative is letting readers over-read a small study. First, one lot of one compound from one supplier is not a sample from which the performance of these services in general can be inferred; it is an existence proof about method-driven spread. Second, three services is too few for any statistical treatment beyond the descriptive; published round-robin studies of peptide purity use seven or more participants for exactly that reason.6

Third, and most important, the exercise tested reproducibility, not accuracy. All three could be equally wrong: without a certified reference standard of known purity, there is no true value against which to score them, and the compendial approach to validating a purity procedure requires exactly such a reference to establish accuracy rather than mere agreement.7 What we measured is dispersion around an unknown centre, and the same constraint applies to any quantitation attempted without a matched standard.8

Fourth, blind submission tests a laboratory’s ordinary process, which is the point, but it also means we bought the cheapest standard product from each service rather than the most thorough. A comparison of each service’s best available package would be a different and probably more flattering study, and it would tell a buyer less, because almost nobody buys the best available package.

The Journal will repeat the exercise annually with a different compound and, funding permitting, against a certified reference standard. The design is published so that others can run it.

How a claim decays between the bench and the listing

The decay is worth tracing precisely, because at no step does anybody say anything false. The laboratory reports a purity figure for the sample as received, on a stated date, by a stated method, submitted by a named party. The vendor extracts the figure and the service name onto its own documentation, dropping the submitter and often the method. A reseller reproduces the vendor’s documentation, dropping the date. A listing summarises the whole chain as tested by a named service. A badge reduces it to verified.

Each step is a reasonable act of summarisation and the cumulative effect is a claim of a different kind from the one the laboratory made. A time-indexed measurement on one sample has become an atemporal property of a product line. The information was not concealed; it was compressed away by a chain of parties each of whom had a legitimate reason to shorten the message.

This is why the Journal reports the provenance of every third-party figure it cites — which service, which lot, which date, which submitter — and treats a figure lacking any of those as uncitable. It makes our coverage sparser than the market’s. It also means that a number appearing in this publication can be traced to a document, which is the only property that distinguishes reporting from repetition.

111835528098.9S1 · A99S1 · B98.1S2 · A98.4S2 · B96.8S3 · A96.9S3 · Bper cent area
Figure. Blind duplicate exercise: reported purity on material from a single lot. Bars in pairs are duplicate submissions to the same service three weeks apart. Within-laboratory agreement is tight; the spread across services is 2.1 percentage points and is fully explained by disclosed method differences.

Our own position, stated in full

Janoshik Analytical and PeptideMeter both advertise in The Compound Journal. Both relationships are disclosed by name on our funding page, together with every other sponsor. No advertiser sees editorial copy before publication, no advertiser has any role in commissioning or reviewing coverage, and the analytical-chemistry desk is contractually barred from consulting for any vendor, testing service or compounding pharmacy. This article was edited by the standards desk under the same rules as every other piece in the department.

The Journal also pays these services. We have submitted samples to three of the four on commercial terms, at list prices, and the blind duplicate exercise described above was funded from editorial budget. We are therefore simultaneously a customer of the institutions we are reporting on and a recipient of advertising revenue from two of them. Readers are entitled to weigh that, and the only useful response we can offer is to state it plainly and to publish objections.

Our position on the substance is unchanged by any of it. All four services are legitimate operations and we have no evidence of dishonesty by any of them. The problems this article describes are structural — who commissions testing, who decides what is published, and what a sample can support about a batch — and they would persist unchanged if every person working at all four organisations were beyond reproach. Correspondence to standards@compoundjournal.com.

What the Journal thinks these services are worth

A summary judgement, since a critical article of this length invites the inference that we think the sector is worthless. We do not. Independent testing in this market is the only mechanism by which a buyer can obtain information about material that is not supplied by the party selling it, and its existence is the difference between a market with some evidence in it and a market with none. Several of the reports these services produce are better documents than the manufacturer certificates they are checking, which is a low bar cleared with room to spare.

The three criticisms we would press are narrow. A sample is not a batch, and the trade cites samples as batches. The party paying for a test decides whether anybody sees it, and the visible corpus is therefore selected. And a badge on a listing has dropped every particular a reader would need. None of these is an analytical failing and none is a failing of the services in isolation; the second and third are properties of the market that surrounds them.

What we would tell a reader is this. A third-party report on a vial you selected and posted yourself is strong evidence about that vial. A third-party report published by the vendor is weaker evidence, of an amount you cannot determine. A badge is not evidence. And nothing in any of the three is a statement about whether anybody should administer the contents to anything.

We will repeat the blind duplicate exercise annually, with a different compound each year and, if we can fund it, against a certified reference standard so that accuracy rather than merely reproducibility can be assessed. The design is published in full so that anybody else can run it, and we would rather be contradicted by a better study than be the only publication that has tried.

References

  1. “Variance components in interlaboratory studies: separating repeatability from reproducibility in chromatographic assays.” Analytical Chemistry. 2020;92(7):5024–5033.
  2. International Council for Harmonisation. Q2(R2): Validation of Analytical Procedures. 2023. Sections on accuracy, precision and the distinction between repeatability and intermediate precision.
  3. “Interlaboratory reproducibility of quantitative measurements on identical samples: lessons from multi-site studies.” Molecular & Cellular Proteomics. 2017;16(4):648–661.
  4. “Identity confirmation of submitted peptide samples in contract analysis: practice, reporting and limitations.” Rapid Communications in Mass Spectrometry. 2018;32(14):1121–1130.
  5. “Interlaboratory comparison of reversed-phase purity determination for synthetic peptides: sources of between-laboratory variance.” Journal of Chromatography A. 2021;1642:462024.
  6. “A round-robin study of purity and content determination for synthetic peptides across seven laboratories.” Journal of Peptide Science. 2019;25(9):e3196.
  7. United States Pharmacopeia. General chapter ⟨1225⟩, Validation of Compendial Procedures. USP–NF. On the reference materials required to establish accuracy as distinct from precision.
  8. “Quantitation by mass spectrometry without a matched reference standard: what can and cannot be claimed.” Journal of the American Society for Mass Spectrometry. 2019;30(6):985–996.

Letters to the Editor

5 printed

Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.

I submitted a vial to one of these services last year, got a result three points below what the vendor advertised, and did not know what to do with it. Your article explains why: I had one measurement on one vial, no method comparison, and no way to know if my vial was representative. I still do not know what to do with it, but I understand the shape of not knowing.

H. Baptiste, Fort-de-France

The Journal replies

That is a better summary of this article’s practical content than our own closing paragraphs. The one thing we would add is that your result is worth publishing wherever you can, because buyer-submitted results are the scarcest and most informative category in the entire corpus.

I run analytical services and I object to the framing of your blind comparison. You bought our cheapest tier, published the number it produced alongside a competitor’s most thorough package, and called the result a spread. It is not a spread. It is three different products, priced accordingly, and your own table says so two columns to the right of the headline figure.

R. Hollenbeck, Spokane, WA

The Journal replies

This is the objection we thought was most likely and we think it is partly right. The comparison is of standard products at standard prices, which is what buyers actually purchase, and we said so. But the presentation invites the reading you object to, and the figure caption now states the tier alongside each result rather than leaving it to the method columns.

You note that no service offers sterility testing and that it takes fourteen days. Worth saying more plainly: a purity certificate and a sterility assurance are not merely different tests, they are different disciplines with different facilities, and no amount of chromatography will ever bear on it.

T. Elorriaga, San Sebastián

The Journal replies

Correct and worth the emphasis. We have said it in the certificates piece and should say it here too: nothing any of these four services sells addresses sterility, endotoxin or container closure integrity, and no combination of their reports adds up to one.

The price table is the most useful thing you have published this year and also the thing most likely to be quoted out of context by somebody selling a comparison service. You might consider a note.

N. Halvorsen, Trondheim

The Journal replies

There is one, and we have strengthened it. The figures are what this publication was invoiced at list rates and are not quotations a reader should expect; volume submitters pay materially less.

VendorInvestigate does not measure anything and you have grouped it with three laboratories under the heading independent testing. That is exactly the conflation your article says the market makes.

S. Bergqvist, Malmö

The Journal replies

A fair hit. The tag under which this coverage sits predates the distinction we now draw, and we have added the distinction to the second paragraph and to the table. The department name will follow at the next reorganisation of the site.

Related coverage