How every number on this site is computed.
Trust is a product feature, not boilerplate. This page renders the exact constants shipped in the scoring engine (score-v0.2.0), served machine-readably at /methodology.json. Coverage is a curated watchlist of 39 products across 6 workflows — not a census. Composite scores are secondary navigation. Eligibility gates are not scored.
Evidence rows: 142 · confirmed by the gate: 42 (30%) · unverified: 100 (70%).
Verification status is computed exclusively by the executable evidence gate (scripts/verify-evidence.mjs) — never hand-authored. A row becomes excerpt-matched only when its stored excerpt hash binds to text retrieved live from the source. Unverified rows remain visible but carry reduced weight everywhere evidence is scored.
Current ≤45d · Aging 46–120d · Historical >120d (excluded from default rankings) · Insufficient = no dated evidence.
Clock: 2026-08-27T04:05Z. Every age on this site is computed against this seed timestamp — not the wall clock — so published numbers cannot silently drift. Future-dated changelog or evidence records are rejected at seed.
| Dimension | Weight | Deterministic inputs |
|---|---|---|
| Evidence strength | 25% | Best-row-floored blend of source-strength × recency (120-day half-life) × confirmation multiplier, plus a depth term saturating at 5 confirmed rows. |
| Deployment traction | 20% | Named-customer counts extracted from gate-confirmed claims (sqrt-scaled) plus confirmed expansion events within the trailing 365 days. |
| Interoperability | 15% | Named EHR integrations (capped), FHIR/API availability, SSO support, marketplace presence. |
| Trust & governance | 15% | BAA availability, security attestations, active FDA disclosure where applicable. Conditional caps: 85 without FDA, 92 with. |
| Economic value evidence | 15% | Count of independently verifiable ROI claims ONLY (peer-reviewed / regulatory / independent report). Vendor pages and uncorroborated cases contribute nothing. |
| Market durability | 10% | Log-scaled disclosed funding plus institutional investor count. |
Governing principles (formulas behind this disclosure)
- Composite is weight-normalized over available dimensions; dimensions without evidence render UNKNOWN and do not default to zero.
- Coverage penalty: composite × (1 − 0.15 × missing-weight-share). A profile missing most dimensions cannot look great.
- Freshness penalty: when median evidence age passes 45 days, subtract 1 point per 15 days, capped at 12 points (saturating at 225 days). Order: coverage multiplies first, freshness subtracts second. Inside the quality blend, evidence value also decays with a 120-day half-life — a separate mechanism.
- Confidence: High requires ≥95% weight coverage AND at least 2 rows CONFIRMED by the evidence gate citing strong sources (peer-reviewed, regulatory, or independent report classes). An unverified strong-class row never earns High.
- Economic value counts ONLY independently verifiable claims (peer-reviewed / regulatory / independent report); vendor pages add nothing.
- Trust & governance caps at 85 without an active FDA disclosure and 92 with one; “pending” grants nothing.
Source strength multiplies each row's contribution; an unrecognized source type aborts scoring rather than defaulting. Unconfirmed rows take an additional confirmation discount (see /methodology.json: confirmed_statuses). Material score movements require stronger evidence classes or second-key review.
Mean Brier across resolved questions: 0.076 (6 resolved). Calibration is provisional until 20 resolved questions are on file. No qualitative grade is published below that sample size.
Binary questions use recency-weighted Brier scoring at resolution (the same tested function ships in the API). Crowd, expert, and model-assisted forecasts are recorded separately before any blending. Resolution follows the predeclared source-of-truth rule in each question's resolution clause; outcome checks against that rule are logged with the resolution record.
| Workflow | Target median age | Current median | Next refresh |
|---|---|---|---|
| Ambient documentation | 90d | 210d | Wave R1 (Sep 2026) |
| Revenue cycle / prior auth | 90d | 190d | Wave R1 (Sep 2026) |
| Imaging & diagnostics | 120d | 260d | Wave R2 (Oct 2026) |
| Patient access | 120d | 240d | Wave R2 (Oct 2026) |
| Clinical decision support | 120d | 180d | Wave R2 (Oct 2026) |
| Enterprise agents | 120d | 150d | Wave R1 (Sep 2026) |
benchmark-bands-v1.0 — four dimensions (deployment traction · evidence strength · interoperability · economic value) computed only from gated evidence rows. Output is a band, never a bare 0–100. UNKNOWN means the evidence is not on file — it is always rendered with why + closes.
- Strong — ≥3 qualifying rows with ≥1 independent source
- ≥3 qualifying rows but no independent source — volume without independent corroboration caps at Moderate
- Moderate — 2 qualifying rows
- Thin — 1 qualifying row
- UNKNOWN — 0 qualifying rows; always rendered with why + closes
- A qualifying row is a gated evidence row whose verification_status was computed by the gate as verified / excerpt_matched / second_key_confirmed. unverified, link_ok and stale rows never count.
- Independent source = source_type with SOURCE_STRENGTH ≥ 0.85 (peer_reviewed_study, regulatory_filing, independent_report).
- economic_value rows must additionally carry a quantified parse (money, %, hours/FTE, patient/claim counts, or explicit ROI/payback) to qualify.
- A vendor is rankable only when all four dimensions are non-UNKNOWN. Rank key = Σ band weight (Strong 3 … UNKNOWN 0); tiebreak = total qualifying rows, then name.
- Calibration is provisional until a vendor has ≥20 qualifying gated rows (reuses calibrationLabel from lib/intel.ts).
| Dimension | Band rule | UNKNOWN why | Closes when |
|---|---|---|---|
| Deployment traction | Qualifying deployment/traction rows (named customers, go-lives, expansions). Strong ≥3 w/ ≥1 independent · Moderate 2 · Thin 1 · UNKNOWN 0. | No deployment/traction evidence row (named customers, go-lives, expansions) has passed the evidence gate. | Excerpt-verified deployment claims — named customer announcements, go-live reports, or expansion counts from any source class. |
| Evidence strength | Qualifying evidence-quality rows (studies, evaluations, benchmarks, regulatory filings); highest source tier named in why. Strong ≥3 w/ ≥1 independent · Moderate 2 · Thin 1 · UNKNOWN 0. | No evidence-quality row (study, evaluation, benchmark, regulatory filing) has passed the evidence gate. | A peer-reviewed study, independent evaluation, or regulatory filing with a gate-matched excerpt. |
| Interoperability | Qualifying FHIR/EHR/marketplace rows only. Product-profile flags (EHR integrations, FHIR/API) corroborate only and can never create evidence. Strong ≥3 w/ ≥1 independent · Moderate 2 · Thin 1 · UNKNOWN 0. | No interoperability evidence row (FHIR/EHR/marketplace) has passed the evidence gate. Product-profile flags (EHR integrations, FHIR/API) corroborate only — they cannot create evidence. | Excerpt-verified integration claims: marketplace listings, FHIR/API documentation, named EHR integrations. |
| Economic value evidence | Qualifying economic_value rows that carry a quantified parse (money, %, hours/FTE, counts, ROI/payback). Strong ≥3 w/ ≥1 independent · Moderate 2 · Thin 1 · UNKNOWN 0. | No quantified economic-value row (ROI, payback, cost/throughput numbers) has passed the evidence gate. | An audited implementation report or peer-reviewed study with payback/throughput numbers. |
Ranking exists only among vendors with all four dimensions non-UNKNOWN (rank key = Σ band weight, Strong 3 … UNKNOWN 0; tiebreak total qualifying rows). Calibration is provisional until a vendor reaches 20 qualifying rows. Every n is traceable to evidence row ids, exposed in the API and the /benchmark UI.