One category, put to the test

Same abstracts. Same question. Six answers.

Adverse reactions recovered from 600 adalimumab abstracts and matched to a 108-concept key adjudicated by a pharmacovigilance expert. Each model read the same corpus and the same 24-article evidence package, with the same schema, parser and matcher.

ArcaScience
10496.3% · missed 4
Gemini 3.6 Flash
7064.8% · missed 38
Claude Opus 5
6963.9% · missed 39
Perplexity Sonar
6661.1% · missed 42
Kimi K3
5853.7% · missed 50
GPT-5.6 Sol
5147.2% · missed 57

Exact McNemar tests at concept level, Holm-adjusted over the fixed family of five. Largest adjusted p: 3.9 × 10⁻⁸. The four ArcaScience misses are named in the paper: caesarean section, gait disturbance, malignant disease, onycholysis.

Against the strongest model
+34 concepts

31.5 points of recall

Of the concepts where the two systems disagree, 37 were found by ArcaScience alone and 3 by Gemini alone. Against GPT-5.6 the gap widens to 53 concepts.

Evidence, kept
100% cited

Every concept has a source

A reaction without a corpus citation does not count. ArcaScience processes each article on its own, so the passage behind every concept survives into the output.

Measured in production
$0.0162

Per article, no cache

Paid-model ceiling recorded on 3,819 safety articles in July 2026, repriced at catalog rates with prompt caching disabled.

The control run

And that was on a drug they had memorised.

Asked about adalimumab with no document at all, the frontier models recite between 12.9% and 44.8% of the 2002 label. Adalimumab has been on the market for over two decades, so that label sits in their training data. Even with that head start, they lost the benchmark.

Given the corpus and required to cite it, the same models score 13.4% to 18.5% against that label. Every BRB-C score of record comes from source-grounded calls with valid corpus citations. Your candidate has no label to recite. For it, only what a system can find and cite in the evidence exists.

Claude Opus 5
44.8%
GPT-5.6 Sol
43.1%
Kimi K3
34.5%
Perplexity Sonar
17.7%
Gemini 3.6 Flash
12.9%
Share of the 232-item 2002 adalimumab label recited with no sources given. First recorded sourceless run per model; Perplexity's repeat scored 9.9%.
Scale economics

The full risk profile, at the scale of the whole literature.

Sixteen categories over 30 million articles, costed at catalog prices with prompt caching and negotiated discounts left out. This is what building a risk profile from everything published would cost each way.

Frontier models, separate passes
$1.36M to $18.5M

Perplexity Sonar at the low end, Claude Opus 5 at the high end, each rereading every article for every target family.

Recall held in the scenario: 45.4% to 62.0%, the measured single-target score of each model.
ArcaScience workflow
$486K to $1.79M

From complete reuse of the instrumented July 2026 workflow to a proportional marginal workload for the families without direct telemetry.

Benchmark recall carried at 96.3%. Anchor: $61.88 no-cache ceiling over 3,819 production articles.

At 3,000 articles the same scenario prices the pipeline at $0.0162 per article against $0.0362 for Perplexity, $0.0399 for GPT-5.6, $0.0620 for Gemini, $0.2011 for Kimi and $0.4944 for Claude. Deterministic retrieval and filtering remove most raw snippets before any paid call: 533,891 in, 6,513 out.

Designed so that ArcaScience could lose.

The corpus was frozen before scoring, the answer key was built by dated adjudication, and the misses of every system, ArcaScience included, are published by name.

Frozen corpus600 records drawn uniformly, seed 20260807, from 1,488 abstract-bearing references prequalified for an earlier client project.
Adjudicated key689 spans annotated in Label Studio, collapsed to 122 concepts, ruled down to 108 by a pharmacovigilance expert. Every ruling dated and published.
Matched evidenceA fixed 24-article full-text package, selected before any recovery was known, supplied identically to every model.
Open releaseProtocol, score vectors, prompt, run files, deviations log, cost model and production manifests. Rescoring needs no model call.

See your candidate's risk profile.

The pipeline scored here is the one that builds ArcaScience benefit-risk assessments: the literature read in full, sixteen categories of evidence structured, every finding traceable to its source, delivered along FDA BRF, BRAT and CIOMS XII.