
Research Intelligence Methodology
How HealthAI Central collects, categorizes, and synthesizes data across the health AI lifecycle. We combine deterministic pipeline ingestion with AI-assisted classification to track innovation from lab bench to bedside billing.
Data Sources & Ingestion
Our data pipelines run daily, pulling from primary federal and academic registries. We use strict boolean search strategies to isolate health AI records from the broader medical literature. Every source below documents the exact search terms, filters, and inclusion decisions used in production — expand any source to see its full methodology.
Science & Discovery
- PubMedIngested daily via NCBI E-utilities (ESearch/EFetch). Two boolean landscape queries isolate health AI literature, plus a company-affiliation pass for tracked companies.
View exact queries & filters
API / AccessNCBI E-utilities API (ESearch → EFetch), daily incremental runs. Records deduplicated by PMID.
Exact landscape queries (run verbatim)("artificial intelligence" OR "machine learning" OR "deep learning") AND (health OR medical OR clinical OR healthcare OR biomedical)("neural network" OR "natural language processing") AND (health OR medical OR clinical OR diagnosis OR patient)"{company name}"[Affiliation] — run for every tracked company in our directoryFilters & ScopePublication dates from January 2000 forward. Incremental runs re-query with a moving minimum date so only new records are fetched.
Inclusion DecisionsA publication is included if it matches either landscape query (AI/ML terms co-occurring with health context terms) or is affiliated with a tracked health AI company. MeSH terms assigned by the National Library of Medicine are stored with every record and drive clinical-area categorization downstream.
- PreprintsSourced from bioRxiv and medRxiv via the bioRxiv API. Every new preprint is screened at insert time against 69 word-boundary regex patterns across five keyword families.
View exact queries & filters
API / Accessapi.biorxiv.org, both servers (bioRxiv and medRxiv), daily. The API has no server-side search, so all filtering happens at insert time. Deduplicated by DOI.
Keyword screen — title OR abstract must match ≥1 of 69 case-insensitive regex patterns, grouped in five familiesCore AI/ML (11): artificial intelligence · machine learning · deep learning · neural network · natural language processing · computer vision · large language model · foundation model · transformer model · generative AI / generative artificial intelligenceArchitectures & methods (16): convolutional neural · recurrent neural · random forest · gradient boosting · support vector machine · reinforcement learning · federated learning · transfer learning · self-supervised · semi-supervised · unsupervised learning · supervised learning · graph neural · attention mechanism · vision transformer · diffusion modelHealth-AI applications (20): clinical decision support · computer-aided detection/diagnosis · radiomics · pathomics · digital pathology · medical imaging AI · AI-assisted · AI-powered · ML-based · predictive model · predictive analytics · clinical prediction · risk prediction · automated diagnosis · automated detection · automated segmentation · image segmentation · object detection · EHR/electronic health record co-occurring with predict/model/algorithmDigital health (10): digital health · digital therapeutics · mHealth · wearable + sensor/device/monitor · remote patient monitoring · telehealth + AI/algorithm/model · chatbot · GPT · ChatGPT · LLMBioinformatics AI (6): protein structure prediction · AlphaFold · drug discovery + AI/ML/deep/neural/model · genomic, single-cell, or multi-omics co-occurring with deep learning / neural / machine learningFilters & ScopeAll patterns use word boundaries and case-insensitive matching to avoid false positives (e.g., "LLM" must be a standalone token).
Inclusion DecisionsA preprint is stored only if it matches at least one pattern. The specific matched keywords are saved on every record for auditability, so any inclusion can be traced back to the exact terms that triggered it.
- NIH GrantsIngested from the NIH RePORTER Projects API. Nine landscape search terms are run against project titles and indexed terms, fiscal years 2000–2026, plus an organization pass for tracked companies.
View exact queries & filters
API / AccessNIH RePORTER Projects API v2 (POST /v2/projects/search), paginated per fiscal year to stay under the API result cap. Deduplicated by application ID.
Exact search terms — each run as an advanced text search (operator AND) against project title + indexed project terms, per fiscal yearartificial intelligencemachine learningdeep learningneural networknatural language processingcomputer vision healthclinical decision support systemcomputer-aided detectioncomputer-aided diagnosisFilters & ScopeFiscal years 2000–2026 for the full backfill; nightly incremental runs cover the current and prior fiscal year. A separate organization pass queries org_names for every tracked company.
Inclusion DecisionsA grant is included if any landscape term matches its title or NIH-indexed terms, or if its recipient organization is a tracked health AI company. Project abstracts, funding amounts, PI details, and administering institute are stored in full.
- CitationsCitation counts for all 188K+ publications are fetched from OpenAlex nightly, powering Most Cited Articles rankings.
View exact queries & filters
API / AccessOpenAlex API (free, keyless, polite pool via contact email), nightly incremental. Publications matched by PMID in batches.
Inclusion DecisionsChosen over Google Scholar, which has no API and blocks automation. cited_by_count is refreshed for new publications each night; existing counts are backfilled periodically.
- People & InstitutionsAggregated across all sources (authors, PIs, trial investigators). Enriched via monthly AI-assisted scraping of 40+ top university health AI labs.
Clinical Validation
- ClinicalTrials.govIngested daily via the ClinicalTrials.gov API v2 using a single boolean landscape query, with a last-update date filter for incremental runs (~141 new/updated trials per day).
View exact queries & filters
API / AccessClinicalTrials.gov API v2, daily incremental. Deduplicated by NCT ID; updated trials are refreshed in place.
Exact queries (run verbatim as query.term)"artificial intelligence" OR "machine learning" OR "deep learning"Incremental filter: AREA[LastUpdatePostDate]RANGE[{last-run-date}, MAX]"{company name}" — sponsor/collaborator pass for every tracked companyFilters & ScopeThe landscape query searches all indexed trial text (interventions, conditions, titles, summaries). Incremental runs fetch only trials whose registry record changed since the last run.
Inclusion DecisionsA trial is included if it matches the AI/ML landscape query or names a tracked company. Phase, enrollment, status, sponsor class, design fields (allocation, masking, age groups), and MeSH condition/intervention terms are stored in full.
- Trial Status & Stop ReasonsTracks phase, enrollment, and completion status. Free-text "why stopped" reasons on stopped trials are LLM-classified nightly into seven categories.
View exact queries & filters
API / AccessStatus fields come directly from the registry record. Stop-reason classification runs nightly on new stopped trials.
Inclusion DecisionsMany AI trials are observational or device-based and do not use traditional drug phases (marked N/A). The ~21,500 stopped trials carrying a free-text stop reason are classified by an LLM into seven buckets — enrollment, funding, business decision, safety/efficacy, COVID-19, data/results, other — which powers the "Why Trials Stop" analysis on Clinical Validation pages. Classifications are stored per trial and never overwrite registry data.
Regulatory Clearance
- FDA AI DevicesAnchored to the official FDA AI-Enabled Medical Device List (1,500+ devices). Enriched with full 510(k), De Novo, and PMA endpoint data via openFDA.
View exact queries & filters
API / AccessThe official FDA AI-Enabled Medical Device List CSV is downloaded from fda.gov and upserted daily. Each device is then enriched via exact submission-number lookups on openFDA (device/510k by K-number, device/pma, device/classification).
Inclusion DecisionsThe FDA list itself is the inclusion criterion — we do not decide what counts as an "AI-enabled device"; the FDA does. Clinical areas are assigned in two passes: the FDA review panel deterministically sets the primary area, then an LLM reads the device description and intended use to add one to three secondary areas the panel label alone would miss.
- Safety & RecallsCross-referenced against FDA device recalls and MAUDE adverse event reports using product codes, with a 14-day daily lookback window.
View exact queries & filters
API / AccessopenFDA device/recall + device/enforcement (recalls) and device/event (MAUDE adverse event reports), daily incremental with a 14-day lookback window to catch late-arriving records.
Inclusion DecisionsRecalls and adverse events are linked to AI devices by FDA product code — the same code family the device was cleared under. Both AI-linked and broader device-safety records are retained so AI device safety can be benchmarked against the wider device landscape.
Market Adoption
- Reimbursement CodesA strictly curated, source-verified registry of AI-specific CPT/HCPCS codes (currently 50: 5 Category I, 45 Category III). Excludes unverified draft codes. Mapped to AMA Appendix S taxonomy.
View exact queries & filters
API / AccessHand-curated registry maintained in code with every entry source-verified against primary CMS and AMA documents.
Inclusion DecisionsInclusion requires verification against primary sources (CMS Medicare Physician Fee Schedule, CMS billing articles, AMA CPT releases); draft or rumored codes are excluded. Each entry carries the code, official descriptor, code type (Category I / Category III / HCPCS), AMA Appendix S taxonomy tier (assistive / augmentative / autonomous), clinical area, example devices, effective date, status, and source notes. A small number of non-AI "context codes" (e.g., the 92227–92229 retinal imaging family) are retained and flagged for trend comparison.
- Medicare UtilizationNational utilization and payment data for every code in our AI registry, from CMS Medicare Physician & Other Practitioners datasets, 2013–2024.
View exact queries & filters
API / Accessdata.cms.gov "Medicare Physician & Other Practitioners — by Provider and Service" datasets, one fixed dataset per calendar year (2013–2024), refreshed when CMS publishes a new annual release.
Filters & ScopeOnly rows whose HCPCS code appears in our curated AI reimbursement code registry are ingested — utilization is measured strictly for verified AI codes.
Inclusion DecisionsService counts, provider counts, and average payment amounts are aggregated per code per year. Because CMS data lags roughly two years, the most recent year shown is the latest CMS release, not the current calendar year.
- Government ContractsSourced from USAspending.gov via an 8-keyword × 3-agency search matrix, then screened by an LLM for genuine health AI relevance (~21% pass rate).
View exact queries & filters
API / AccessUSAspending API v2 (spending_by_award), contract award types A–D, award period 2016 to present, paginated and deduplicated by award ID.
Exact search matrix — each of 8 AI keywords crossed with each of 3 awarding agencies (24 queries)Keywords: artificial intelligence · machine learning · deep learning · neural network · computer vision · natural language processing · predictive analytics · clinical decision supportAgencies: Department of Health and Human Services · Department of Veterans Affairs · Department of DefenseFilters & ScopeKeyword search casts a wide net by design; a second-stage LLM relevance screen then evaluates every contract description individually.
Inclusion DecisionsA contract is kept only if it is genuinely about AI/ML applied to healthcare, medicine, clinical care, biomedical research, or health administration. Explicitly rejected: generic IT infrastructure, non-health AI, staffing contracts without AI substance, and non-AI health services. Roughly 21% of fetched contracts pass (346 of 1,650 at last count); passing contracts are tagged to one to three clinical areas with confidence scores, and rejected ones are retained but excluded from dashboards.
- Public CompaniesFinancials for 59 tracked publicly traded health AI companies, pulled from SEC XBRL company facts.
View exact queries & filters
API / AccessSEC XBRL companyconcept API, queried per company CIK with quarterly refresh cadence.
US-GAAP concepts pulledRevenue: RevenueFromContractWithCustomerExcludingAssessedTax (falling back to Revenues)R&D expense: ResearchAndDevelopmentExpenseNet income: NetIncomeLossInclusion DecisionsOnly companies in our tracked public-company list are pulled. Values come directly from company XBRL filings with no restatement or adjustment on our side.
Clinical Area Categorization
We organize all research data into a unified taxonomy of 30 areas (22 clinical specialties and 8 cross-cutting technologies). Because data comes from different sources, we use three distinct categorization methods.
1. MeSH-Based Mapping (Publications & Trials)
For PubMed publications and ClinicalTrials.gov records, we rely on the National Library of Medicine's Medical Subject Headings (MeSH).
- We processed the full MeSH 2026 descriptor tree; 12,287 descriptors map into our 30 clinical areas via tree-number branch rules (e.g.,
C04.*maps to Oncology). - Specialty branches like
M01(Persons/Pediatrics) andC13(Pregnancy/OBGYN) are explicitly mapped to ensure comprehensive coverage. - When a paper or trial is ingested, its assigned MeSH terms automatically link it to the corresponding clinical areas.
2. Source-Specific Mapping (Grants & Preprints)
For sources without MeSH indexing, we map native metadata to our taxonomy:
- NIH Grants: Categorized based on the administering NIH Institute (e.g., NCI maps to Oncology, NICHD maps to Pediatrics).
- Preprints: Categorized using the native subject categories provided by bioRxiv and medRxiv.
3. AI-Assisted Classification (FDA Devices & Contracts)
For sources with unstructured text or generic regulatory categories, we use Large Language Models (LLMs) to assign clinical areas:
- FDA Devices: The FDA review panel deterministically sets the primary clinical area; an LLM then analyzes the device description and intended use to add 1–3 secondary areas the panel label alone would miss.
- Government Contracts: An LLM evaluates each contract description to confirm genuine health AI relevance and assign 1–3 clinical areas with confidence scores.
Multi-Tagging & Primary Areas
A single record (e.g., a trial for a pediatric oncology AI tool) can be tagged with multiple clinical areas. To prevent double-counting in top-level summaries, one area is designated as the Primary Area based on a strict hierarchy (clinical specialties take precedence over cross-cutting technologies).
Maturity Slider & Momentum Score
Each clinical area page features a Maturity Slider and a Momentum Score. These are composite indicators designed to summarize where an area sits in the health AI lifecycle and how quickly it is progressing.
Maturity Slider (Position)
The slider places each area on a four-stage continuum: Research-Dominant → Translational → Regulated → Reimbursed. The position is calculated as a weighted composite of the area's relative standing across lifecycle stages.
Calculation
- For each area, we compute a Stage Index for four dimensions: Research (publications), Trials (clinical trials), Regulatory (FDA devices), and Reimbursement (CPT codes).
- Each Stage Index = (area's count in that stage) / (average count across all 30 areas). A value of 1.0× means average; 3.0× means three times the average.
- The Position is a weighted sum:
position = 0 × w_research + 1 × w_trials + 2 × w_regulatory + 3 × w_reimbursement, where each weight is the normalized Stage Index for that dimension. - The result maps to a 0–3 scale: 0 = pure Research-Dominant, 3 = fully Reimbursed. The slider dot is placed at
(position / 3) × 100%along the gradient bar.
Stage Label Assignment
- Research-Dominant (position 0.0–0.75): Significant publication activity but minimal trials, devices, or codes.
- Translational (position 0.75–1.5): Active clinical trials; some devices may exist but reimbursement is absent or minimal.
- Regulated (position 1.5–2.25): Multiple FDA-cleared devices; reimbursement infrastructure is emerging.
- Reimbursed (position 2.25–3.0): Established CPT codes with measurable Medicare payment activity.
Momentum Score (0–100)
The Momentum Score quantifies how quickly an area is growing across all lifecycle stages. It is displayed as a circular gauge on each area's page.
Calculation
- Start with a base score of 50 (neutral).
- Add points for year-over-year (YoY) growth in each lane:
- Publications: up to +15 points (YoY% / 2, capped at 15)
- Clinical Trials: up to +15 points (YoY% / 2, capped at 15)
- FDA Devices: up to +10 points (YoY%, capped at 10)
- NIH Grants: up to +10 points (YoY% / 2, capped at 10)
- Add a +10 acceleration bonus if the fastest-growing lane exceeds +20% YoY.
- Clamp the final score to the 0–100 range.
YoY Calculation
Year-over-year growth compares the last two complete calendar years (the current year is excluded as partial). For example, if 2024 had 1,248 publications and 2023 had 1,021, the YoY is ((1248 - 1021) / 1021) × 100 = +22%.
Interpretation
- 70–100: Accelerating — strong growth across multiple stages.
- 40–69: Steady — moderate or mixed growth signals.
- 0–39: Decelerating — declining activity in key stages.
Lane Acceleration Indicators
Each lifecycle lane on the Momentum panel displays an acceleration arrow:
- ▲ Accelerating: YoY growth exceeds +15%.
- → Steady: YoY growth is between -5% and +15%.
- ▼ Decelerating: YoY growth is below -5%.
- · Insufficient data: Fewer than 2 complete years of data available.
Lane Divergence Detection
When publications are surging (>+20% YoY) while trials are declining (<0% YoY), the system flags a lane divergence. This signals a widening translation backlog: new science is being produced faster than it enters clinical validation.
Editorial Standards & Summarization
Our editorial pipeline synthesizes raw data into actionable intelligence while maintaining strict sourcing standards.
- Source Verification: Every curated record (like Reimbursement Codes) and every generated news brief must include a direct, clickable link to the primary source.
- Duplicate Detection: Our pipeline uses semantic deduplication to prevent the same story from being published multiple times, even if reported by different outlets.
- No Funding Rounds: By editorial policy, we do not track or cover venture capital funding rounds. We focus exclusively on scientific, clinical, regulatory, and operational market adoption.
Data Accuracy Notice: Research intelligence on HealthAI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.
© 2026 HealthAI Central. All rights reserved.
Source: https://healthaicentral.com/research/methodology