Method
The algorithm, in public, with its failure modes.
Everything below is arithmetic on publication metadata. No model is asked whether a paper is good. Where the rubric can be wrong, this page says how.
§04 · THE RUBRIC
Two scores, shown side by side.
A strong study in the wrong population and a weak study in the right one are different failures. You need to see which one you have, so the two numbers never collapse into one.
- EEvidence strength
- How much the study design, sample size, and registry rigor earn. Computed from publication metadata alone. No model touches it.
- RPersonal relevance
- How closely the study population resembles the profile in your vault. Starts at 50, meaning no reason to think it does or doesn't apply. Never a filter.
| Grade | Band | Badge | What it means |
|---|---|---|---|
| A | E 80–100 | Strong human evidenceSystematic review or meta-analysis of randomized trials, in people. | |
| B | E 60–79 | Moderate human evidenceA randomized trial, or a large prospective cohort, in people. | |
| C | E 40–59 | Limited human evidenceNon-randomized trial, smaller cohort, or a study with real design problems. | |
| D | E 20–39 | Weak or preliminaryCase-control, cross-sectional, case series, or a pilot that was underpowered by design. | |
| E | E 0–19 | Not human evidenceRodent, cell culture, or opinion. Kept visible, never treated as support. |
Why a rat RCT can't outrank a human case series
The design base is multiplied by species before anything else happens. Human 1.00. Mixed 0.95. Animal 0.25. Cell culture 0.15. A randomized trial scores 75 at base, so an animal RCT lands at 18.75, just under a human case series at 15. That caps rodent work firmly outside the believe-this range while keeping it on screen where you can see it.
The five verdicts
- SUPPORTEDTwo or more studies at tier B or better, at least 70% agreeing on direction, and at least one human RCT or systematic review.
- MIXEDTwo or more at tier B or better, but agreement on direction is under 70%.
- CONTRADICTEDThe modal direction is no effect and at least two tier A or B studies agree on it.
- HARM_SIGNALAt least one tier A or B study reports the intervention harmed.
- INSUFFICIENTEverything else. It's the default, it fires often, and that's the point.
INSUFFICIENT is borrowed from the USPSTF's I statement. Most supplement questions genuinely deserve it. Not having that bucket is a large part of why AI literature tools mislead people.
It is CEBM-inspired. It is not GRADE.
The spine is the Oxford CEBM 2011 levels, layered with SORT's patient-oriented outcome axis. GRADE rates a whole body of evidence across five downgrade domains that simply aren't derivable from metadata, so emitting a GRADE label would be a costume. The rubric ships versioned as cebm-x-1.0 and you can read every term of it.
§04.1 · SPECIES
The MeSH query that swallows humans.
Species is read off the MeSH XML of each record, never asked for in the query. The reason is one number, and getting it wrong would quietly corrupt every answer the tool gives.
The trap
Read off the MeSH XML, never asked for in the query. "animals[mh]" explodes to 29.1M records and silently swallows Humans; unexploded it's 7.95M.
Why it matters
creatine AND humans[mh] returns 43,033 records. creatine AND animals[mh] NOT humans[mh] returns 21,672. About a third of the corpus for a supplement people have studied for thirty years is rodent and cell work. For obscure compounds the ratio is much worse, and a naive search hands it to you undifferentiated.
§04.2 · WORKED EXAMPLE
CYP2C19 and clopidogrel, end to end.
The five records behind the homepage panel, with the reason each landed in the tier it did — including the one that contradicts the other four, and the one whose numbers had to be inverted to share an axis.
The call
CYP2C19 *2/*17
Intermediate Metabolizer
- *17
- Increased function
- *2
- No function
- CPIC level
- A
- ClinPGx level
- 1A
- Testing
- Actionable PGx
One no-function allele and one increased-function allele do not cancel. A tool that averaged them would get this wrong, which is why it is the example.
This result signifies that the patient has one copy of a normal function allele and one copy of a no function allele OR one copy of an increased function and one copy of a no function allele. Based on the genotype result this patient is predicted to be an intermediate metabolizer of CYP2C19 substrates. This patient may be at risk for an adverse or poor response to medications that are metabolized by CYP2C19. To avoid an untoward drug response, dose adjustments or alternative therapeutic agents may be necessary for medications metabolized by CYP2C19. Please consult a clinical pharmacist for more information about how CYP2C19 metabolic status influences drug selection and dosing.
Why ancestry moves the answer
| Allele | rsID | Function | East Asian | European | Callset |
|---|---|---|---|---|---|
| CYP2C19*2 | rs4244285 | No function | 30.23%9,559 / 31,626 | 14.6133%152,480 / 1,043,430 | exome |
| CYP2C19*3 | rs4986893 | No function | 9.23%3,662 / 39,672 | 0.0101%112 / 1,111,770 | exome |
| CYP2C19*17 | rs12248560 | Increased function | 0.95%49 / 5,174 | 22.0006%14,949 / 67,948 | genome |
No exome data exists for this variant — it is upstream of the coding region and outside exome capture. gnomAD returns exome: null. Genome callset only (AN ~152k).
gnomAD puts *17 at 0.95% in East Asians; CPIC’s literature table puts it at 2.05% — about 2× apart. Different sampling frames (population sequencing vs. pooled published cohorts), and gnomAD’s East Asian genome AN is only ~5,174. Report the range, not one number, if this appears as a standalone figure.
And here is the fact that cuts against the simple story, which is exactly why it is on the page: *2/*17 specifically is rarer in East Asians — 1.2% against 6.3% in Europeans — because *17 is nearly absent in East Asia. The phenotype class is far more common; this particular route into it is not.
The five records
Relative risk of the study’s ischemic endpoint for a CYP2C19 loss-of-function carrier on clopidogrel, versus that study’s comparator. Greater than 1 = worse on clopidogrel.
East Asian LOF carriers: 2× major adverse cardiac events after stenting
OR 1.99 (95% CI 1.64–2.42)- Design
- Systematic review + meta-analysis, 20 studies
- Population
- East Asian (China, Korea, Japan), PCI with stent implantation
- Endpoint
- MACE (cardiovascular death + myocardial infarction)
- Compared
- carriers of at least 1 CYP2C19 LOF allele (*2 and/or *3) vs non-carriers, all on clopidogrel
Why this tier. Systematic review with meta-analysis — top of the design ramp — across 20 studies and 15,056 patients, ancestry-matched to the example genotype, with prespecified subgroups by loading dose and by nationality. Independently corroborated: the CPIC 2022 guideline (PMID 35034351) cites this paper for its East Asian effect estimates. Discounted from a pure RCT-meta-analysis because the component studies are predominantly observational cohorts; the large N and the concordance with row 2 carry it.
What cuts against it. Also reports stent thrombosis OR 4.77 (2.84–8.01) and a LOWER bleeding risk in carriers, OR 0.66 (0.46–0.96) — the mechanism cuts both ways and the site should say so. The CPIC guideline quotes this paper’s intermediate-metabolizer-specific figure as OR 1.92 (1.34–2.76); that subgroup value is not in the abstract, so it is stored separately in CPIC_QUOTED_EAST_ASIAN_EFFECTS rather than used here.
Carrier risk is 1.9× in Asians but 1.2× in whites — same drug, same procedure
RR 1.91 (95% CI 1.61–2.27)- Design
- Systematic review + meta-analysis, 24 studies, ancestry-stratified
- Population
- Asian and white strata analysed separately; quoted estimate is the Asian PCI stratum (n=10,017)
- Endpoint
- Major cardiovascular outcomes
- Compared
- carriers of at least 1 CYP2C19 LOF allele vs non-carriers, Asians undergoing PCI (n=10,017)
Why this tier. Systematic review with meta-analysis, 36,076 participants, and the primary analysis was restricted a priori to studies with at least 500 participants — an explicit small-study-bias guard, which is the exact failure row 5 attacks. It is also the only record here that formally TESTS the ancestry modifier rather than assuming it: heterogeneity between strata P<0.001. This is the load-bearing citation for the hero’s ancestry claim.
What cuts against it. Same paper, same analysis, other strata: whites undergoing PCI RR 1.20 (1.10–1.31), n=19,016; whites NOT undergoing PCI RR 0.99 (0.84–1.17), n=7,043. Indication matters as much as ancestry — do not quote the Asian figure without the white one.
A single reduced-function allele was enough to raise events and stent thrombosis
HR 1.55 (95% CI 1.11–2.17)- Design
- Collaborative meta-analysis of 9 investigator-contributed cohorts
- Population
- Predominantly European ancestry; 91.3% underwent PCI, 54.5% had ACS
- Endpoint
- Composite of cardiovascular death, myocardial infarction, or stroke
- Compared
- carriers of 1 reduced-function CYP2C19 allele vs non-carriers, all on clopidogrel
Why this tier. Pooled analysis of 9 studies with individual hazard ratios contributed by the original investigators — stronger than any one cohort, but it is not a registered systematic review and it is treatment-only (everyone got clopidogrel), so it cannot separate a prognostic effect of the genotype from a predictive drug-response effect. Held at MODERATE for that reason, not for its size.
What cuts against it. Two reduced-function alleles: HR 1.76 (1.24–2.50). Stent thrombosis was the sharper signal: HR 2.67 (1.69–4.22) for one allele, 3.97 (1.75–9.02) for two.
Randomised in Chinese carriers: ticagrelor beat clopidogrel, 6.0% vs 7.6% stroke
HR 0.77 (95% CI 0.64–0.94)- Design
- Randomised, double-blind, placebo-controlled trial, 202 centres
- Population
- China, 98.0% Han Chinese; all enrolled patients were CYP2C19 LOF carriers with minor stroke or TIA
- Endpoint
- New stroke within 90 days (191/3,205 = 6.0% vs 243/3,207 = 7.6%)
- Compared
- ticagrelor vs clopidogrel, among CYP2C19 LOF carriers
Why this tier. The single strongest individual study here — randomised, double-blind, 6,412 patients, and it enrolled ONLY loss-of-function carriers, so the genotype is not a subgroup afterthought. Held at MODERATE rather than STRONG because it is one trial in one country and one indication, and because it compares two drugs rather than two genotypes — it demonstrates that switching helps carriers, which is adjacent to, not identical to, the claim that carriers do worse on clopidogrel.
What cuts against it. Bleeding went the other way: any bleeding 5.3% on ticagrelor vs 2.5% on clopidogrel, though severe or moderate bleeding was 0.3% in both arms. NEJM’s own summary calls the stroke reduction "modest".
Axis transform. Published as ticagrelor-vs-clopidogrel HR 0.77 (0.64–0.94). Inverted to put clopidogrel in the numerator so it shares the axis with rows 1–3 and 5: point 1/0.77, interval [1/0.94, 1/0.64]. Exact arithmetic on a ratio measure. Captions must quote the published 0.77 (0.64–0.94), not this value.
In the largest review, the association vanished once small-study bias was removed
RR 0.97 (95% CI 0.86–1.09)- Design
- Systematic review + meta-analysis, 32 studies (6 randomised)
- Population
- Mixed, predominantly European ancestry; PCI and non-PCI indications pooled
- Endpoint
- Cardiovascular events
- Compared
- carriers of at least 1 reduced-function allele vs non-carriers, restricted to studies with at least 200 events
Why this tier. The largest evidence synthesis in this set — 42,016 patients, 3,545 cardiovascular events — and it reaches the opposite conclusion to rows 1–4. It found the expected association in naive analysis (RR 1.18, 1.09–1.28) but also found significant small-study bias (Harbord test P=.001); restricting to studies with at least 200 events collapsed the estimate to 0.97 (0.86–1.09), and in the 6 randomised effect-modification studies there was no genotype × treatment interaction (P>.05). Tiered CONFLICTING, not WEAK: its design is strong, its finding is contested. Strength and agreement are different axes, which is the whole point of the categorical tiers.
What cuts against it. Predates CHANCE-2 (2021), TAILOR-PCI (2020) and both East Asian meta-analyses above, and pooled PCI with non-PCI indications — which row 2 later showed is where the effect genuinely disappears (whites, non-PCI: RR 0.99). The contradiction is real but it is partly a question about WHO and WHICH INDICATION, not only about whether.
CPIC Level A with a Strong recommendation for ACS/PCI, but the two randomised East Asian trials here studied a neurovascular indication, where CPIC’s own recommendation strength drops to Moderate — and the largest synthesis in the set (n=42,016) finds the association attenuates to null once small-study bias is removed.
Avoid standard dose (75 mg) clopidogrel if possible. Use prasugrel or ticagrelor at standard dose if no contraindication.
Consider an alternative P2Y12 inhibitor at standard dose if clinically indicated and no contraindication.
Those two paragraphs are written for prescribers, not for you. This is the kind of output you take to an appointment. Information for a future prescription, not a dosing instruction. CPIC guidance is written for prescribers. Not medical advice.
The CPIC guideline also quotes an intermediate-metabolizer-specific estimate — OR 1.92 (1.34–2.76) for MACE. It is verified only as a quotation inside the guideline, not from the source paper’s own abstract, so it is not used as a row above.
§04.3 · VERIFICATION
What we could not verify.
Every claim on this site was pulled from a primary source and carries the URL it came from. This is the list of things a reader could reasonably expect here that are not verified, and three citations that were wrong until they were fetched.
The unverified log, exactly as written
Entry 7 has since been resolved: the WEAK and ANIMAL ONLY badges are demonstrated in the grade table further up this page, which is where a reader can see them without a study being padded into the homepage stack to justify one.
Three citations that were wrong until they were fetched
Recalled from memory while assembling the example above. Each would have shipped a citation that resolves to a real paper about something else entirely, which is the most damaging failure available to a tool like this.
| Meant | Recalled | What that PMID actually is | Correct |
|---|---|---|---|
| CHANCE-2 trial, NEJM 2021 | 34551253 | Interactions between Brassica Biofumigants and Soil Microbiota. J Agric Food Chem 2021. | 34708996 |
| Holmes MV, CYP2C19 systematic review, JAMA 2011 | 22110105 | Urinary sodium and potassium excretion and risk of cardiovascular events. JAMA 2011. | 22203539 |
| TAILOR-PCI, JAMA 2020 | 32805007 | Effects of Dietary Glycemic Index and Glycemic Load ... Polycystic Ovary Syndrome. Adv Nutr 2021. | 32840598 |
Endpoints hit
The screening count in the verdict line is a live PubMed query and it will drift. Query string:
("CYP2C19"[Title/Abstract]) AND ("clopidogrel"[Title/Abstract]) AND (humans[Filter]) AND ("cardiovascular"[Title/Abstract] OR "stent thrombosis"[Title/Abstract] OR "MACE"[Title/Abstract] OR "stroke"[Title/Abstract] OR "myocardial infarction"[Title/Abstract]) AND (meta-analysis[pt] OR randomized controlled trial[pt] OR systematic review[pt])