The platform

Verify every claim in a paper, end to end.

Upload a manuscript and Alethic decomposes it into atomic claims, retrieves each cited source, and audits citation fidelity, statistical consistency, and figure-text alignment, with full evidence trails behind every verdict.

Verification report · 2024

Efficacy of CRISPR-Cas9 in Mammalian Gene Therapy

Chen, L. · Nakamura, K. · Patel, R. · Williams, D. · Nature Biotechnology · 47/47 refs resolved
Sections
Abstract6
Introduction11
Methods18
Results34
Discussion34
Contradicted9Misstated10Needs a citation7Supported26
Introduction11 claims

Cas9 cuts both DNA strands three base pairs upstream of the PAM, leaving blunt ends.

12factual
Supported

Homology-directed repair is absent in post-mitotic cells, ruling out precise correction in neurons.

341 paper unavailablefactual
Misstated

One AAV vector can carry SpCas9, its guide RNA and a strong ubiquitous promoter.

67methodological
Contradicted
Contradicted

SpCas9, its guide RNA and a strong ubiquitous promoter fit in a single AAV vector.

0 of 2 cited papers we could read state this

[6]
Packaging limits of AAV vectors for in vivo genome editingContradicts

“SpCas9 with a ubiquitous promoter exceeds the 4.7 kb a single AAV genome can hold.”

[7]
Delivery routes for CRISPR nucleases in vivoDoesn't cover

Compares viral and lipid delivery; says nothing about AAV packaging limits.

Verification engine

Every claim. Every citation.
Every figure. Verified.

Document-level scores hide what matters. Alethic decomposes a paper into its atomic claims, retrieves the underlying evidence, and verifies each one with a full audit trail.

The pipeline
Multi-agent · live
Three specialists, in sequence.
01  Extraction
47 claims identified
02  Verification
Resolving evidence 41 / 47
03  Report
Compiling output
~25%1
of citations contain errors
κ = 0.182
peer-review agreement (NIH grants)
~50%3
of psychology papers have p-value errors
32 mo4
median time to retraction
1Lukić et al. 2004, J Bone Joint Surg Am/2Pier et al. 2018, PNAS/3Nuijten et al. 2016, Behav Res Methods/4Steen et al. 2013, PLOS ONE
statisticalcitation
Partially Verified

Multi-head attention with 8 heads provides a clear improvement over single-head attention on the WMT 2014 English-to-German task27.

What the paper measured

25.8→26.4+0.6

BLEU · §3.2.2 / Table 3

Why partial

BLEU figure of 26.4 matches Table 3 verbatim. The cited paper supports the qualitative direction of the claim, but the +0.6 BLEU gain is smaller than what “clear improvement” typically denotes in this literature.

Source

“Multi-head attention allows the model to jointly attend to information from different representation subspaces.”
Attention Is All You Need · Vaswani et al., NeurIPS 2017 · §3.2.2
Verified · full text
01Atomic claims

We don't score papers. We verify their claims.

Alethic decomposes a paper into individual claims, classifies each by type, and verifies each against retrieved evidence. Eight verdict types, full rationale, source quotations — every step traceable.

8 verdictsPer-claim rationaleSource quotations
02Multimodal review

Figures that don't support their captions.

A vision-language pass evaluates whether each figure actually supports the claims made about it. Detects truncated axes, misleading scales, and visual–textual misalignment — with the same evidence trail as text claims.

Axis detectionCaption alignmentVerbatim quotes
Figure check · Figure 2Disputed
Editing efficiency by cell liney-axis truncated at 85
100959085HEK293TiPSCT-cellHepNeuron

Caption claim

“Figure 2 shows dramatic gains across cell lines with CRISPR-Cas9 optimization.”

Y-axis spans 85–100; Δ HEK293T → T-cell = 5pp. The visual “dramatic gain” reflects the truncated scale, not the underlying difference.

Statistical recomputation · §4.2Disputed

Reported

Recomputed

mean
3.47
n=21 · scale 1–7
closest possible
3.43 or 3.52
Impossible
t(20)
2.84
t(20)
2.84Matches
p
0.010
p
0.0103Rounds to 0.010

GRIM check failed. With n=21 on a 1–7 integer scale, no sum of responses can produce a mean of 3.47. Either n or the reported mean is wrong. t and p are internally consistent.

03Statistical integrity

When the numbers don't add up.

Alethic recomputes test statistics from reported values, runs GRIM/SPRITE consistency checks, and flags effect-size implausibility. Surfaces claims that don't hold up to their own data.

GRIM / SPRITEp-value recomputeEffect-size sanity
Security & privacy

Built for the institutions
research lives in.

Manuscripts and data are processed transiently in isolated environments. Nothing is retained, nothing is used for model training, and on-prem deployment is available for institutions that need the perimeter inside.

Start verifying today

Every claim,
verified.

Join researchers at 12+ institutions already using Alethic to catch what peer review misses.