Most genes implicated by GWAS studies have no approved drug. An autonomous agent searches for which of these "orphan" genes most closely resemble proven drug targets — using only public APIs and 10 biological features.
GWAS studies have identified thousands of gene variants linked to disease. But the vast majority of the genes they implicate have no approved therapeutic. The bottleneck is prioritisation: which genes are worth investing in?
Genetic associations are causal evidence. Genes identified by GWAS have been directly implicated in disease biology by human variation — stronger prior evidence than expression or network data alone.
Genes with approved drugs are already "solved." The value lies in the 88% that remain unapproached — especially those that biologically resemble proven targets.
Alzheimer's, Parkinson's, CAD, type 2 diabetes, schizophrenia, and IBD span neurological, cardiovascular, metabolic, psychiatric, and immune biology with rich GWAS datasets.
Each pair requires: (1) genome-wide significant GWAS association for the disease, and (2) an approved drug targeting the gene for that specific indication. These 13 genes define what a successful drug target looks like in our 10-dimensional feature space.
| Gene | Disease | Approved Drug | Area | Genetics score |
|---|---|---|---|---|
| Loading… | ||||
The pipeline runs entirely from public APIs — no hardcoded gene lists, no proprietary databases. Everything from the gold standard to the candidate features is discovered programmatically.
13 gene-drug pairs are curated where the gene has genome-wide significant GWAS associations and an approved drug for the same disease. These 13 genes are enriched with real biological features from Open Targets and PubMed APIs to define a reference centroid in 10-dimensional feature space.
The GWAS Catalog is queried for each disease. Associations with p < 5×10⁻⁸ are extracted. Genes in the gold standard are separated. The top 30 orphan genes per disease by GWAS significance are taken forward.
Each orphan gene gets 10 features: Open Targets association score and druggability tractability, PubMed publication count (5 years), plus proxy features for eQTL effect size and Mendelian randomization z-score derived from GWAS evidence depth.
Every gene gets 10 numbers pulled from public databases. Together they answer one question: does this gene look like the kind of gene that becomes a drug target?
The features fall into three groups. Genetic causality answers: is this gene actually causing the disease, or just nearby? Biological context answers: is it active in the right place, and how well studied is it? Druggability answers: can we actually make a medicine for it?
The strength of the genetic association, expressed as −log₁₀(p-value). A value of 8 means p = 10⁻⁸ — the standard genome-wide significance threshold. Values of 20–30 indicate very strong, highly replicated signals. Plain English: how confidently does genetic variation at this gene predict disease risk?
How many independent GWAS studies have implicated this gene at genome-wide significance. One study could be a false positive. Ten studies from different populations mean the signal is real. Plain English: has this finding been repeated by multiple independent groups?
A z-score estimating whether the gene is causally upstream of the disease — not just correlated with it. MR uses genetic variants as natural experiments to test directionality. Higher score = stronger causal evidence. Plain English: does changing this gene's activity actually cause disease, or is the association coincidental?
Expression quantitative trait locus effect — does the genetic variant that hits this gene also change how much of the protein the cell makes? A high eQTL effect means the GWAS signal likely acts through gene expression, which is often easier to drug. Plain English: does the disease-linked variant work by turning this gene up or down?
Is this gene particularly active in the tissue most relevant to the disease? A neurological disease gene that is highly expressed in brain tissue is a more plausible therapeutic target than one expressed everywhere equally. Values range 0 (ubiquitous) to 1 (highly specific). Plain English: is this gene doing its work in the right organ?
How many other proteins does this gene's protein physically interact with? Hub proteins (high degree) sit at the center of biological networks — biologically important, but harder to drug without side effects. Low-degree proteins are more tractable targets. Plain English: how connected is this protein in the cell's machinery?
Number of PubMed papers mentioning this gene in the last 5 years. Highly published genes are better understood — but may also be more competitive commercially. Very low counts may signal an underexplored opportunity. Plain English: how much scientific attention has this gene received recently?
From Open Targets tractability data. Does this protein have a binding pocket that a small molecule could fit into? Ranges from 0.35 (no evidence of druggability) to 0.90 (an approved small-molecule drug already exists for it). Plain English: could a pill be made to hit this target?
A 0–1 aggregate score from Open Targets Platform, combining genetic association evidence, somatic mutations, animal models, literature mining, and functional genomics data across all diseases. Plain English: across all the evidence types Open Targets tracks, how well-supported is this gene as a disease target overall?
DALYs (disability-adjusted life years) from GBD 2021 — how much suffering this disease causes globally. Used in two ways: as a normalized feature in the scoring formula and as a post-hoc multiplier (score × (1 + DALYs/200)) when enabled. Cardiovascular disease scores 182M DALYs; immune diseases score 3M. Plain English: genes in higher-burden diseases get a larger score boost, all else equal.
The metric is the Spearman rank correlation between two rankings of orphan genes: (1) the scorer's ranking from the weighted formula, and (2) the cosine similarity ranking to the gold standard centroid. A perfect scorer would rank orphan genes in exactly the same order as their biological distance to proven drug targets.
Orphan + gold gene features are min-max normalized together so all genes share the same scale across all 10 features.
The 13 normalized gold gene vectors are averaged to one centroid — the "ideal drug target profile" in feature space.
Each orphan gene's similarity to the centroid is computed. This is the ground truth the agent tries to reproduce.
Spearman r measures whether the scorer ranking agrees with the cosine-similarity ranking. Higher r = better formula.
Our baseline equal-weight formula achieves r ≈ 0.87, meaning even before optimisation the biological features naturally align with the gold standard. The agent's job is to find which features to emphasise to push this higher.
After data collection, agent.py runs 10 pre-defined experiments. Each modifies the scoring weights, measures Spearman r, and decides whether to keep or discard the change — then auto-commits results to GitHub.
Each iteration auto-commits and pushes results to GitHub — the full experiment trail is in git history
Early experiments up-weight causal features (MR z-score, druggability) and down-weight popularity proxies (PubMed count, PPI degree) to test whether the gold centroid is driven by causal biology or citation count.
Mid experiments add multiplicative interaction terms — MR × druggability, eQTL × tissue specificity — testing whether combining causal and functional signals predicts drug-likeness better than either alone.
Late experiments add a disease-burden multiplier (DALYs from GBD 2021) to test whether genes in higher-burden diseases should rank higher even with identical biological profiles.
Results update after every experiment run. The agent commits each outcome to GitHub, so the full trail is public.
Autonomous prioritisation of GWAS candidates sits at the intersection of genomics, AI, and drug discovery. The pipeline demonstrates several things at once.
Public GWAS data contains most of the signal needed to rank orphan candidates by drug-likeness. This pipeline can run daily as new GWAS results are published, continuously updating the priority list without manual curation.
The agent shows that a simple autoresearch loop — modify, evaluate, keep/discard — can systematically explore a feature-weight search space. No LLM reasoning required; the metric provides the signal. This pattern generalises to any ranking problem with a defined ground truth.
The entire pipeline — data, features, scoring, results — runs on free public APIs and is published on GitHub. Any lab can fork it, change the disease scope, and run it in an afternoon. The barrier to genomics-informed drug target prioritisation is now near zero.
The full pipeline is open source. Runs on any machine with Python 3 and internet access. No API keys, no subscriptions.
~8 minutes to build data · 10 experiments · auto-pushes results to GitHub Pages