BioTransfer · Methods

How disease briefings are built

Every figure on a briefing page comes from a public API on a fixed schedule. This page says which ones, what each number means, what is deliberately excluded, and where the method is weak.

Where every number comes from

No figure on a briefing is written by hand or produced by a language model. Each is fetched, computed, and stamped with the date it was retrieved. Prose on the page may only describe values already in that fetched set.

SectionSourceRefreshed
Targets, drugs, tractabilityOpen Targets Platformweekly
Trials, cell therapyClinicalTrials.gov (US NIH/NLM)weekly
DatasetsNCBI GEOweekly
Dataset reuse countsEurope PMCweekly
Citation impact, translation potentialNIH iCiteweekly
Research momentumPubMed E-utilitiesmonthly
FundingNIH RePORTERquarterly
ApprovalsopenFDA drug labels (current SPL text)weekly
Disease burdenSEER Cancer Stat Facts; ACS and CBTRUS where SEER publishes no pageannually
Mutation landscapecBioPortal public studiesmonthly
Gapsderived from the above — no separate sourceevery build

Fifty-three briefings, not fifty-three populations

Each briefing is built under one disease definition, stated at the top of the page and in its JSON. The set overlaps: lung cancer and small cell lung cancer, kidney cancer and Wilms tumour, melanoma and uveal melanoma, non-Hodgkin lymphoma subtypes beside each other. Their cases, dollars, trials and papers must not be added across pages; a grant, a trial or a paper can sit on several of them by design.

What a briefing is not

These pages are for research orientation and target prioritisation. They are not clinical guidance, not prescribing information, and not investment advice. A drug listed as approved is approved for the indications named beside it — which, on most pages, is a different disease from the one the page is about.

What "reused" means

This is the number most particular to these pages, and the one most worth understanding before trusting a ranking.

Reuse count

The number of papers in Europe PMC whose full text names the dataset's accession — for example GSE49710. A paper that downloads and reanalyses deposited data has to name the accession; a paper that merely agrees with the original study's conclusions does not.

So reuse measures whether anyone has successfully worked with the files, which is a different question from whether the paper was influential.

Why it is shown separately from citations

On the neuroblastoma page, GSE26494 carries 697 citations and a very high impact score — and two reuses. The study shaped the field; its deposited data did not travel. Blending both into a single relevance score would hide exactly the distinction a researcher choosing a dataset needs.

Reuse is also a harsh filter. Of 235 neuroblastoma studies, 11 clear ten reuses. Most clear none.

The two citation measures

MeasureMeaningRead it as
RCRRelative Citation Ratio, from NIH iCite — citations normalised for both research field and paper age1.0 is field average; 19.7 is far above it. Fair to compare a 2006 paper with a 2021 one
APTApproximate Potential to Translate, from NIH iCite — a model estimate of whether a paper will be cited by clinical work0 to 1. Higher means more likely to reach clinical literature. An estimate, not an outcome

How datasets are ranked

GEO's own relevance ranking is keyword matching. It cannot tell a 498-patient tumour cohort from a six-well cell-line experiment, so we score the things it does not expose. The formula is published in full so it can be checked rather than trusted.

score = 3.0 · log(1 + reuse)      how often the data was actually reused
      + 2.0 · log(1 + 4 · RCR)   field- and age-normalised citation impact
      + 4.0 · APT                translational potential
      + 1.2 · log(1 + samples)   cohort size
      + 0.8 · assay_tier         spatial / single-cell 4 · sequencing 3 · array 2
      + 2.5 · patient_cohort     patient > xenograft > cell line
      + 1.5 · clinical_fields    survival · stage or risk group · driver status
      + 0.7 · target_overlap     genes shared with the page's target table

There are no learned weights and no hidden features. Every component that contributed to a row is shown on that row as a chip, so a ranking can be argued with.

Two rails, not one list

Citation and reuse signals do not exist for anything published in the last eighteen months, so a single ranking would permanently bury new data. ESTABLISHED is scored with those signals; RECENT is scored without them.

One study, one row

A single study often deposits several GEO SubSeries. These are collapsed to one row by their linked publication, so a five-part study cannot occupy five of the top ten places.

How approvals are read

A drug counts as labelled for the disease when the current FDA label's INDICATIONS AND USAGE section names the disease in a positive indication clause. Naming it is not enough.

The clause is read, not the label

Labels name diseases under Limitations of Use ("efficacy in GIST has not been demonstrated"), in contraindications, in risk statements ("a greater predisposition to endometrial carcinoma") and inside other names ("non-Hodgkin lymphoma", "melanoma-associated antigen A4"). Each clause that names the disease is classed on its own words; a label whose only clauses are negative is listed on the page under Labels that name the disease without indicating for it, with the clause quoted, and is not counted.

Four roles

Labelled here — a drug whose label positively indicates it for the disease. Backbone — a broad cytotoxic whose label lists many tumours. Supportive — given in the disease without acting on it: stem-cell mobilisers, growth factors, rescue agents, antiemetics, bone agents, thyroid hormone after thyroidectomy. Not a therapy — a diagnostic or contrast agent. Only the first is what "approved for this disease" means on the page and in every cross-disease count.

What the label text cannot tell us

openFDA serves the current SPL. An indication withdrawn from the market while the SPL is still served (idelalisib and duvelisib in follicular lymphoma, panobinostat, copanlisib) is removed by a curated list in the profile, with the reason printed. Salt forms, biosimilars and co-formulated hyaluronidase collapse to one molecule. A label that drops out of openFDA's index for a week is kept for thirty days and marked; openFDA's index is not stable day to day and the changelog should record withdrawals, not indexing weather. Approval history, regimen, line of therapy and biomarker restrictions are in the quoted clause, not in the count.

How trials are counted

ClinicalTrials.gov is searched by condition and intervention. Its condition search is loose — a search for Wilms tumour returns WT1 vaccine trials in leukaemia, a search for cervical cancer returns head-and-neck trials — so every returned trial is re-checked.

Disease-eligible

A trial counts when one of its registered conditions or keywords names the disease (the same regex the briefing uses for literature), or when every registered condition is a tumour-agnostic basket ("solid tumours", "advanced cancer"; for haematological diseases also "lymphoma", "leukaemia" or "haematological malignancy" as the profile allows). A trial whose conditions name a different disease is dropped and listed in the JSON under dropped_off_disease with the conditions that excluded it. Across the 53 briefings this removed about a fifth of what the condition search returned.

Cell therapy

A trial is cell therapy when its title names an administered cell product (CAR, TCR, NK, TIL, adoptive) and not an antibody-fusion, bispecific, vaccine or peptide that merely mentions T cells. It counts for a target only when the title names that target or one of its synonyms: a HER2 CAR-T trial that adds pembrolizumab is a HER2 trial, not a PD-1 trial. The targets searched are the profile's list, which includes drug-class axes as well as antigens; the page says "busiest among those searched".

Registered is not treated

Counts are registered interventional studies, deduplicated by NCT id. Active means a recruiting, enrolling or active-not-recruiting status on the date shown. A withdrawn trial never enrolled anyone and is counted separately. None of this says how many patients were dosed or what happened to them.

How research momentum is computed

Two windows of MeSH-indexed publications — typically 2015–2018 against 2021–2025 — compared as share of the disease's literature, not as raw counts. A topic registers as rising only if it grew faster than the field grew.

Papers are restricted to those indexed with the disease as a major topic, so incidental mentions are not counted.

Three things we deliberately do not do

We do not rank terms that were absent from the earlier window. Their apparent growth is often a MeSH vocabulary term that did not exist yet, not new science. Their fold change is null in the JSON and printed as "new"; the rising table needs at least five papers in the earlier window.

We do not draw a trend line to the present. MeSH indexing lags publication by roughly a year, so recent years are always undercounted. On the neuroblastoma page, indexed papers appear to fall from 894 in 2023 to 333 in 2026 — that is indexing lag, not a collapse in research.

We do not report terms below twelve papers in the current window. Below that, share ratios are noise.

How funding is computed

NIH awards whose title, abstract or terms name the disease, summed per fiscal year. Award dollars are the obligations recorded in that year, not the lifetime value of a grant.

Dollars are always shown against project counts

"Funding doubled" is true of neuroblastoma and misleading on its own. Obligations rose 108% between 2013 and 2025 while the number of distinct funded projects stayed flat near 200 and the median award grew 72%. The money bought larger awards at a concentrated set of institutions, not more independent groups. Both numbers are therefore always shown together.

Project counts use distinct core project numbers, so supplements and multi-year records do not inflate them.

NIH only, by agency code

RePORTER also carries VA, FDA, CDC/NIOSH and AHRQ projects. Only records whose agency code is NIH are counted; the others are tallied in the JSON under other_agency_excluded so the exclusion can be seen.

A zero year is a search result

Where the newest fiscal year has no award that passes the briefing's topic filter, the page reports the last year that has one and names the empty years. That is a statement about the filter and the query, both printed, not a statement that no such research is funded.

How burden is recorded

Cases, deaths, survival and median age come from SEER Cancer Stat Facts where SEER publishes a page for the disease, and from ACS or CBTRUS where it does not. Every figure carries its mapping: direct when the source counts the disease itself, proxy when it counts a broader category, derived proxy when a share of a parent was applied, with the share and its source printed.

Units are the source's units

Basal and cutaneous squamous cell carcinoma counts are lesions, not people, because that is what the source counts; breast cancer is female invasive breast cancer; hepatocellular carcinoma is 70% of liver-and-intrahepatic-bile-duct cancer by NCI's statement, an estimate of an estimate. The unit and the scope are printed beside every figure, and a derived death count that came from an incidence share is labelled as such.

Years of life lost is a proxy, not the WHO measure

The page's figure is (1 − five-year relative survival) × (78.4 − median age at diagnosis) × cases. WHO and GBD years of life lost use deaths by age with a life table; this one uses diagnosis age and treats five-year survival as cure. It separates a disease from one ten times its size, not from a neighbour, and is comparable across these pages only. Its assumptions are listed beside it on every briefing.

How the mutation landscape is built

For each disease with public cBioPortal cohorts, patient-level non-silent mutation counts and discrete copy-number calls, per cohort, never pooled. The genes shown are the briefing's curated targets plus the forty most frequently altered genes in the reference cohort, with known passengers (TTN, MUC16, olfactory receptors) left out of the ranked set but not the data.

Denominators are the panel

A gene absent from a targeted panel is not assayed, not zero. Percentages are over the patients whose panel carried the gene; copy-number percentages are over the cohort's copy-number roster, which is often smaller. Hypermutated patients (more than ten times the cohort median, or over 100 mutations) are counted separately and the figures are given with and without them.

What it is not

Not a meta-analysis: cohorts differ in assay, era and selection, and a figure from a 2012 exome study is not the same quantity as one from a 2024 clinical-panel cohort. The cross-disease index takes one reference cohort per disease and names it; the landscape page keeps every cohort apart.

Across diseases

The cross-disease pages — burden, targets, drugs, antigens, mutations, gaps, changelog — are derived from the briefings on every build and carry no source of their own. A gene page lists each briefing that carries the gene, whether a drug against it is labelled for that disease, and its reference-cohort alteration frequency. The index JSON at /disease/api.json describes every endpoint and its row fields.

What we exclude, and why

Public biomedical databases return a surprising amount of material that mentions a disease without being about it. Every briefing states what was removed and how much. Two patterns account for nearly all of it.

Cell lines used as models in other fields

Neuroblastoma's SH-SY5Y line is widely used as a stand-in for neurons in neurodegeneration research. Those papers carry the Neuroblastoma index term. Left in, they made Alzheimer's disease 5.4% and Parkinson's disease 4.3% of "neuroblastoma" literature and dominated the list of rising topics.

Different diseases sharing a name

Olfactory neuroblastoma — esthesioneuroblastoma — is a sinonasal tumour of adults, unrelated to the childhood tumour. It contributed a further 4.6% of literature and a similar share of grants.

Together these accounted for roughly 14% of the literature, 13% of grant records and 20% of candidate datasets on the neuroblastoma page. Each briefing carries a per-disease exclusion list, reviewed by a person rather than generated, and reports the percentage removed alongside every figure it affects.

Exclusions are published, not hidden

Every section names what it removed. Where a record was removed by hand rather than by rule — GSE16716 on the neuroblastoma page — that is stated too. A filtered number you can check beats an unfiltered number you cannot.

Known limitations

Stated plainly, because a briefing that hides these is worth less than one that does not.

Machine access

These pages are built to be read by software as much as by people. Nothing is behind a click, a tab, or a script — the full content of a briefing is in the delivered HTML.

ResourceAddress
Briefing page/disease/<slug>
Same figures as JSON/disease/<slug>.json
Mutation landscape/disease/<slug>/mutations/ and .json
Cross-disease index/disease/api.json — lists every endpoint
Structured markupschema.org MedicalWebPage + Dataset

The JSON is the same object the page renders from, so what a machine reads and what a reader sees cannot drift apart. Each section carries its own source, retrieved_at, query and excluded fields.

Citing a figure

Quote the value with its retrieval date and the section it came from — for example: "GSE49710 has been named in 208 Europe PMC full-text papers (BioTransfer, retrieved 2026-09-05)." Values change on the schedules listed at the top of this page.