Medutest result on a Shanghai mazdutide lot lands 1.2 points above the supplier’s stated figure
The Journal submitted the sample and paid for the analysis. The vendor was told in advance.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Institutions
Two years ago we ran an anonymised version of this comparison and promised a named one. This is it, with every method printed in full.
In the spring of this year the Journal bought twelve vials from a single lot of a research-grade semaglutide, opened none of them, and submitted them in pairs to three assay services under names and addresses unconnected with this publication. Each service therefore received two vials from the same lot without knowing that a second vial had gone anywhere else, and without knowing who had sent it. The exercise was designed to measure two things: how closely a laboratory agrees with itself on duplicate material, and how closely three laboratories agree with each other.
Janoshik Analytical is a Czech laboratory and the service whose reports circulate most widely in this market. Its core offering is chromatographic purity determination and peptide content by elemental nitrogen determination, with mass-spectrometric identity confirmation available, on samples submitted by post. Reports are issued to the submitter and carry the method, the gradient and integration detail, which places them at the more informative end of what this market produces.
Janoshik advertises in this publication. The relationship is disclosed by name on our funding page, no advertiser sees editorial copy before publication, and this section was written and edited under the same rules as every other in this department. Readers who consider the arrangement disqualifying are entitled to say so and several have; the correspondence is printed.
The Journal’s substantive observations are two. First, the practice of reporting content alongside purity is the single most useful thing any service in this market does, because content is the number that determines how much peptide a nominal mass represents and almost nobody else reports it. Second, the reports describe submitted samples, they say so, and the trade nonetheless quotes them as batch verification. That second point is a criticism of citation practice rather than of the laboratory, and it applies to every service in this article.
The most important sentence in every third-party report in this market is the one identifying what was tested, and it always says the same thing: the sample submitted. That phrasing is precise and correct, and the entire trade reads past it. A result obtained on one vial extends to a batch only if the vial is representative, and representativeness is a property of how the vial was selected, not of how carefully it was analysed.
In regulated practice, sampling is a controlled activity in its own right: the accreditation standard treats it as part of the laboratory activity, requiring a documented sampling plan and records of how the portion tested was obtained.1 Where a laboratory receives a sample it did not draw, the standard expects the report to make clear that the results apply to the sample as received. All four services do this. The market quotes them anyway as though the batch had been tested, which inverts the pharmacopoeial convention that a result on a sample is evidence about a batch only under a stated sampling assumption.2
The practical significance depends on the fill. A batch filled in a single session from a homogeneous bulk solution is likely to be uniform, and a single vial is decent evidence about it. A batch assembled from subdivided bulk, filled across sessions, or blended from more than one synthesis is not, and a single vial is evidence about a vial. Nothing on a report tells a reader which situation applies, because the laboratory does not know either.
Selective publication is a harder problem than shading, and no laboratory is in a position to detect it. Four genuine reports and one published best figure produce a page in which every document is authentic and the set is the misrepresentation.
A badge is not evidence. It is an assertion that evidence exists, offered without the particulars that would let anybody assess it.
The Journal’s standing positionReduce the incentive problem to its mechanism and it is very simple. The party who pays for a test receives the report. The report is that party’s property. Nothing obliges them to publish it. Therefore the set of third-party results visible in this market is not the set of results obtained; it is the subset that somebody with a commercial interest chose to make visible.
The size of the resulting distortion is unknown and unknowable from published data, which is precisely the difficulty. If a vendor submits twelve batches over a year and publishes eight, the four unpublished results are invisible, and no amount of scrutiny applied to the eight will recover them. This is the mechanism the research literature calls publication bias, and it is well characterised: the distortion it produces is larger when the number of tests is small, when the cost of a test is high relative to the value of a favourable result, and when nobody records that a test was commissioned.
The accreditation standard for testing laboratories requires a laboratory to identify and manage risks to its impartiality arising from its commercial relationships, and to be able to demonstrate that it has done so.3 That obligation sits on the laboratory and it is the right place for it. What the standard cannot reach is the client’s filing cabinet, and the filing cabinet is where this problem lives.
| Analysis | Quoted turnaround | Observed turnaround | Price band (EUR, single sample) |
|---|---|---|---|
| Purity, generic gradient | 3–5 working days | 4–9 days | 55–90 |
| Purity, extended gradient | 5–10 working days | 7–16 days | 110–180 |
| Purity + orthogonal confirmation | 2–3 weeks | 15–31 days | 190–320 |
| Identity by intact mass | 3–7 working days | 5–12 days | 45–110 |
| Peptide content by nitrogen | 1–2 weeks | 9–22 days | 160–280 |
| Water by Karl Fischer | 1 week | 6–11 days | 70–130 |
| Peptide mapping / sequence | 3–5 weeks | 26–38 days | 480–950 |
| Prices are the amounts actually invoiced to this publication at list rates between the second quarter of 2024 and the first quarter of 2026, converted where necessary at the rate on the invoice date, and are not quotations any reader should expect. Volume submitters pay materially less. Turnaround is measured from posting to receipt of the report. | |||
Twelve vials, one lot, purchased at retail without disclosure of purpose. Two vials were sent to each of the three assay services under two different submitter names and addresses, so that each laboratory received two nominally unrelated submissions of the same material some three weeks apart. Six further vials were retained. Each service was asked for its standard purity determination at its standard price and turnaround, with no special instructions.
The design tests two distinct quantities that the trade conflates. Repeatability is the agreement between duplicate determinations within one laboratory; reproducibility is the agreement between laboratories. Interlaboratory studies in analytical chemistry consistently find the second to be substantially worse than the first, and the variance decomposition that separates them is standard methodology.4 The distinction between repeatability and intermediate precision is formalised in the validation guidance,5 and multi-site studies in adjacent fields have repeatedly found between-laboratory agreement on identical samples to be the harder problem.6
All three services were informed after the fact, before publication, and each was given the opportunity to comment on its own method as printed and on the comparison as a whole. All three responded. Two supplied additional method detail that has been incorporated. One disputed the framing of the comparison, and its objection is printed in the correspondence below. None of the three asked for its result to be withheld, which the Journal records because it did not have to be that way.
Interlaboratory spread of a point or so on a chromatographic purity is unremarkable by the standards of proficiency testing in other sectors. It is also consequential when a stated specification sits inside that spread, and both things are true at once.
Within-laboratory repeatability was good. The two determinations from each service agreed to within 0.3 percentage points in every case, and to within 0.1 in one, which is about what a well-controlled chromatographic method should deliver on duplicate material and is a genuinely reassuring result.
Between-laboratory reproducibility was another matter. The three services returned figures spanning 2.1 percentage points on material from one lot. Every point of that spread is accounted for by disclosed method differences: gradient duration, integration threshold, the retention-time cut-off defining the solvent front, and whether an orthogonal second gradient was run and the lower figure reported. Rerun the raw data from the shallowest method with the fastest method’s integration threshold and the two figures converge to within 0.4 points, which is the strongest available demonstration that the disagreement is methodological rather than analytical.
Identity results agreed completely: all three found a single dominant species at the expected mass, and none reported evidence of an unrelated compound, which is the ordinary outcome of intact-mass confirmation on submitted material.7 The two services reporting peptide content returned 93% and 91% of label, a difference within the stated uncertainty of nitrogen determination. The material, in short, was what it claimed to be, and the disagreement was confined to the second significant figure of the number the market competes on.8
Four limitations, stated because the alternative is letting readers over-read a small study. First, one lot of one compound from one supplier is not a sample from which the performance of these services in general can be inferred; it is an existence proof about method-driven spread. Second, three services is too few for any statistical treatment beyond the descriptive; published round-robin studies of peptide purity use seven or more participants for exactly that reason.9
Third, and most important, the exercise tested reproducibility, not accuracy. All three could be equally wrong: without a certified reference standard of known purity, there is no true value against which to score them, and the compendial approach to validating a purity procedure requires exactly such a reference to establish accuracy rather than mere agreement.10 What we measured is dispersion around an unknown centre, and the same constraint applies to any quantitation attempted without a matched standard.11
Fourth, blind submission tests a laboratory’s ordinary process, which is the point, but it also means we bought the cheapest standard product from each service rather than the most thorough. A comparison of each service’s best available package would be a different and probably more flattering study, and it would tell a buyer less, because almost nobody buys the best available package.
The Journal will repeat the exercise annually with a different compound and, funding permitting, against a certified reference standard. The design is published so that others can run it.
A verification badge on a product listing typically asserts, in one or two words and a graphic, that the product has been tested by a named service. Consider what that claim leaves open. Which lot was tested. When. Who submitted the sample and how it was selected. What was measured — purity alone, or identity, or content. By what method. Whether the lot currently on sale is the lot that was tested. Whether the report is available to the reader.
Every one of those is material and every one is absent. A badge is therefore not evidence; it is an assertion that evidence exists, offered without the particulars that would let anybody assess it. The Journal does not treat badges as evidence in its coverage, states so wherever it reports one, and will not cite a badge as support for a claim about material.
The remedy is four data points the vendor already possesses: the lot number, the date of analysis, the service, and a link to the report. Several listings in this market already carry them, which establishes both that it is possible and that it is not commercially fatal. What it requires is that a badge become perishable — attached to a lot rather than to a product line — and perishability is precisely the property a marketing asset is designed not to have.
Publish the fact of the submission and keep the result confidential. A vendor with nine submissions and three published results is visibly a vendor with six it withheld.
The narrower request we now make of all four servicesThe decay is worth tracing precisely, because at no step does anybody say anything false. The laboratory reports a purity figure for the sample as received, on a stated date, by a stated method, submitted by a named party. The vendor extracts the figure and the service name onto its own documentation, dropping the submitter and often the method. A reseller reproduces the vendor’s documentation, dropping the date. A listing summarises the whole chain as tested by a named service. A badge reduces it to verified.
Each step is a reasonable act of summarisation and the cumulative effect is a claim of a different kind from the one the laboratory made. A time-indexed measurement on one sample has become an atemporal property of a product line. The information was not concealed; it was compressed away by a chain of parties each of whom had a legitimate reason to shorten the message.
This is why the Journal reports the provenance of every third-party figure it cites — which service, which lot, which date, which submitter — and treats a figure lacking any of those as uncitable. It makes our coverage sparser than the market’s. It also means that a number appearing in this publication can be traced to a document, which is the only property that distinguishes reporting from repetition.
| Element | Cheapest tier | Mid tier | Most detailed tier |
|---|---|---|---|
| Compound and lot as declared | Yes | Yes | Yes |
| Date of receipt | Sometimes | Yes | Yes |
| Condition on arrival | No | Sometimes | Yes |
| Instrument identified | No | Sometimes | Yes |
| Column identified | No | Sometimes | Yes |
| Gradient stated | No | Yes | Yes |
| Detection wavelength | Sometimes | Yes | Yes |
| Integration threshold | No | Sometimes | Yes |
| Chromatogram reproduced | No | Yes | Yes |
| Named analyst | No | Sometimes | Yes |
| “Sample as received” statement | Yes | Yes | Yes |
| Aggregated across the three assay services rather than attributed, because tier names and boundaries differ between them and a service-by-service table would invite comparison of products that are not comparable. Every element listed appears on at least one service’s standard output. | |||
Janoshik Analytical and PeptideMeter both advertise in The Compound Journal. Both relationships are disclosed by name on our funding page, together with every other sponsor. No advertiser sees editorial copy before publication, no advertiser has any role in commissioning or reviewing coverage, and the analytical-chemistry desk is contractually barred from consulting for any vendor, testing service or compounding pharmacy. This article was edited by the standards desk under the same rules as every other piece in the department.
The Journal also pays these services. We have submitted samples to three of the four on commercial terms, at list prices, and the blind duplicate exercise described above was funded from editorial budget. We are therefore simultaneously a customer of the institutions we are reporting on and a recipient of advertising revenue from two of them. Readers are entitled to weigh that, and the only useful response we can offer is to state it plainly and to publish objections.
Our position on the substance is unchanged by any of it. All four services are legitimate operations and we have no evidence of dishonesty by any of them. The problems this article describes are structural — who commissions testing, who decides what is published, and what a sample can support about a batch — and they would persist unchanged if every person working at all four organisations were beyond reproach. Correspondence to standards@compoundjournal.com.
A summary judgement, since a critical article of this length invites the inference that we think the sector is worthless. We do not. Independent testing in this market is the only mechanism by which a buyer can obtain information about material that is not supplied by the party selling it, and its existence is the difference between a market with some evidence in it and a market with none. Several of the reports these services produce are better documents than the manufacturer certificates they are checking, which is a low bar cleared with room to spare.
The three criticisms we would press are narrow. A sample is not a batch, and the trade cites samples as batches. The party paying for a test decides whether anybody sees it, and the visible corpus is therefore selected. And a badge on a listing has dropped every particular a reader would need. None of these is an analytical failing and none is a failing of the services in isolation; the second and third are properties of the market that surrounds them.
What we would tell a reader is this. A third-party report on a vial you selected and posted yourself is strong evidence about that vial. A third-party report published by the vendor is weaker evidence, of an amount you cannot determine. A badge is not evidence. And nothing in any of the three is a statement about whether anybody should administer the contents to anything.
One thing this article has deliberately not done is rank the four services. We do not think the evidence supports a ranking, we do not think a ranking would be used carefully, and two of the four advertise here, which is a reason for additional caution rather than a reason to pretend the relationship does not exist. What we have tried to publish instead is the material a reader would need to rank them for their own purposes.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
Anonymity for the purchaser and anonymity for the supplier are different protections and they are frequently confused. The first keeps the sample honest; the second is a courtesy that mostly benefits the party being tested.
— C. Bąkowski, Łódź
The suppliers who cooperate with testing are, almost by definition, the ones with least to fear. A programme built on cooperation therefore samples the top of the market and reports the result as though it described the whole of it.
— C. Farquharson, Aberdeen
The weakest link is almost always the domestic leg. A vial delivered to a residential address, kept in a kitchen for a week and then posted onward has lost any custody claim it had, and this is the ordinary case rather than an unusual one.
— H. Baptiste, Fort-de-France
You note that no service offers sterility testing and that it takes fourteen days. Worth saying more plainly: a purity certificate and a sterility assurance are not merely different tests, they are different disciplines with different facilities, and no amount of chromatography will ever bear on it.
— L. Dziedzic, Wrocław
Correct and worth the emphasis. We have said it in the certificates piece and should say it here too: nothing any of these four services sells addresses sterility, endotoxin or container closure integrity, and no combination of their reports adds up to one.
The Journal submitted the sample and paid for the analysis. The vendor was told in advance.
Documentation practice is the only part of vendor quality a buyer can assess before purchase.
Follow the resin, not the catalogue.
Documentation practice is the only part of vendor quality a buyer can assess before purchase.
A reminder that a purity figure is the output of a method, and that methods differ.
Two signatures — one who performed the work, one who approved its release — are the ordinary regulated convention and are almost unknown here.