Reading guide Start here → understand the question and method (§00) → see how the experiments performed (§01 trace, §02 table) → learn which biological features drove the best score (§03 weights + feature guide) → explore the top candidate genes (§04) → check the gold standard template (§05) → see how all genes distribute (§06 histogram).

Experiment Results

An autonomous agent ran 10 scoring experiments on orphan GWAS genes across 6 diseases, searching for the formula that best predicts which genes resemble known drug targets. Each experiment rewrote scorer.py with a different weight vector, measured Spearman correlation against the gold standard centroid, and kept or discarded the change.

The question, method, and metric

Which GWAS-implicated genes — currently without any approved drug — are most likely to become drug targets? We use 13 curated gene-drug pairs as a biological template and ask: can a scoring formula rank the 169 orphan genes in the same order that their feature similarity to known drug targets would suggest?

🎯

The question

Out of hundreds of genes implicated by GWAS but lacking approved drugs, which have biological profiles most similar to proven drug targets? Answering this prioritises the most tractable candidates for drug discovery.

⚙️

The method

13 gold-standard gene-drug pairs (GWAS-confirmed + approved drug for the same disease) define a reference centroid in 10-dimensional feature space. Each orphan gene's cosine similarity to that centroid is its "ground-truth" score. The agent optimises a weighted linear formula to reproduce that ordering.

📐

The metric — Spearman r

Spearman correlation measures whether the scorer's ranking matches the cosine-similarity ranking. Range: −1 (inverted) to +1 (perfect). A value above 0.85 means the scorer reproduces the gold-centroid ordering very well — experiment-to-experiment changes reveal which features matter most.

STEP 1

Feature enrichment

Each gene gets 10 features from Open Targets (OT score, druggability, tractability), GWAS Catalog (p-value, study count), and PubMed (publication count). Gold genes use the same pipeline.

STEP 2

Gold centroid

The 13 gold genes are normalized together with orphan genes and averaged to a centroid vector. This centroid encodes what "approved drug target" looks like biologically.

STEP 3

Cosine similarity

Each orphan gene's normalized vector is compared to the centroid. The resulting similarity score is the ground-truth ranking the agent tries to reproduce.

STEP 4

Agent loop

The agent tries 10 weight configurations. After each, Spearman r is measured between the scorer ranking and cosine-similarity ranking. Better formulas are kept.

—Best Spearman r
—Baseline r
—Improvement
—Experiments kept
—Orphan genes
—Gold standard pairs

Spearman r over 10 iterations

Green filled dots = formula kept (improved). Red hollow dots = discarded. Dashed line = baseline equal-weight formula.

What the agent tried

Every experiment, its description, Spearman r, delta from baseline, and outcome. Best row highlighted.

#What the agent triedSpearman rvs baselineResult
Loading…

The winning scoring formula

Feature weights from the best-performing experiment, sorted by importance. Higher weight = feature was more useful for reproducing the gold-centroid ordering.

Each bar above shows how much weight the best experiment gave to that feature. Below is what every feature actually measures — grouped by what biological question it answers.

Group 1 — Genetic causality: is this gene actually causing disease?

gwas_pval_log10

GWAS signal strength. −log₁₀(p-value) of the strongest GWAS association. 8 = genome-wide significance threshold. 20–30 = very strong, highly replicated signal. Higher = more confident genetic link to disease.

n_gwas_studies

Replication count. How many independent GWAS studies found this gene at genome-wide significance. One hit could be noise. Ten independent hits across different populations means the signal is real and robust.

mr_z_score

Mendelian randomization z-score. Estimates whether the gene is causally upstream of disease — not just correlated. Uses genetic variants as natural experiments to test directionality. Higher = stronger causal evidence. This is the most important feature for target validation.

eqtl_effect

eQTL effect size. Does the disease-linked variant also change how much of this protein the cell makes? A high eQTL effect means the GWAS signal acts through gene expression — making the mechanism clearer and the target more interpretable.

Group 2 — Biological context: is this gene in the right place doing the right thing?

tissue_specificity

Tissue specificity. Is this gene particularly active in the disease-relevant tissue (e.g. brain for neurological diseases)? A gene expressed everywhere is a harder target — more side effects. A gene expressed mainly where you need it is more tractable.

ppi_degree

Protein interaction degree. How many other proteins does this protein physically touch? Hub proteins are biologically central but often too connected to drug safely. Peripheral proteins with focused roles make cleaner targets.

pubmed_count_5yr

Recent publication count. Papers mentioning this gene in PubMed over the last 5 years. Proxy for scientific momentum. Very high counts mean competitive space; very low counts may signal an underexplored opportunity.

Group 3 — Druggability: can we actually make a medicine for this gene?

druggability_score

Small-molecule tractability. Does this protein have a binding pocket a drug can fit into? From Open Targets tractability assessment. 0.35 = no evidence. 0.70 = advanced clinical candidate. 0.90 = approved drug already exists for this protein.

open_targets_score

Open Targets composite score. A 0–1 aggregate combining genetic associations, somatic mutations, animal models, literature mining, and functional genomics across all diseases. The broadest single measure of how well-supported this gene is as a target.

burden_daly_m

Disease burden (DALYs). DALYs from GBD 2021. Used in two ways: as a normalized feature in the linear formula, and as a post-hoc multiplier (score × (1 + DALYs/200)) when BURDEN_MULTIPLIER is on. CAD = 182M DALYs; IBD = 3M. Genes in higher-burden diseases rank higher all else equal.

Leave-one-out cross-validation

Each gold standard gene was temporarily removed from the reference set, added to the orphan candidate pool, and scored by the winning formula. Recovery rate = fraction of gold genes that rank in the top 20.

Loading loo.json…

The agent's top 20 candidates

Ranked by the best scoring formula. Similarity bar shows cosine distance to the gold centroid — how "drug-target-like" the gene's biological profile is.

Loading…

The 13 curated gene-drug pairs

Each pair was selected because: (1) the gene has genome-wide significant GWAS associations for that specific disease, and (2) an approved drug targets the same gene for the same indication. These genes define what a successful drug target looks like in our feature space.

GeneDiseaseApproved DrugAreaGenetics score
Loading…

How orphan genes distribute against the gold centroid

Histogram of cosine similarity scores across all orphan genes. Genes far right are biologically most similar to the gold standard drug targets.