Gold Standard — Discovered from Open Targets API
Experiment Trace — recall@K over iterations
Ranked Candidate Genes
Methodology
Phase 1 — Gold Standard Discovery (Open Targets)
For each of 6 diseases, the Open Targets GraphQL API is queried for gene-disease associations.
Genes with
genetic_association or genetic_literature evidence AND
an approved drug (maxClinicalStage == APPROVAL) are selected.
The top 2–3 per disease by genetics evidence score form the gold standard ground truth.
Nothing is hardcoded — the gold standard is fully discovered at runtime.
Phase 2 — Orphan Gene Discovery (GWAS Catalog)
GWAS Catalog is queried for genome-wide significant associations (p < 5×10⁻⁸) per disease.
Gene symbols are extracted from
authorReportedGenes in each locus.
Any gene already in the gold standard is excluded. Up to 30 orphan genes per disease
(no approved drug, but GWAS signal) are retained.
Phase 3 — Feature Enrichment (Open Targets + PubMed)
Each orphan gene is enriched with:
open_targets_score— max association score across top 3 diseasesdruggability_score— SM tractability tier from Open Targetspubmed_count_5yr— PubMed publication count (last 5 years)eqtl_effect— proxy: (gwas_p/50)×0.5 + ot_score×0.5mr_z_score— proxy: (gwas_p/10)×(1+n_studies/20)tissue_specificity— disease-area lookupppi_degree— binned from PubMed count
Autoresearch Loop (agent.py)
agent.py runs 10 pre-defined experiments. Each experiment rewrites
scorer.py with new feature weights and interaction terms,
evaluates recall@K against the discovered gold standard, and keeps
the change only if recall improves. Results and scorer snapshots are
committed to GitHub after every iteration.