Controlled-access repository
Population-scale genomics, responsibly shared
1.4 million whole genomes and 3.1 million exomes from 41 contributing cohorts, available to approved researchers through a data access committee and an audited enclave.
How the archive works
Managed access, not open access
Every dataset has a data access committee. Applications are reviewed against the consent under which the samples were collected, and decisions are published in aggregate.
Analysis in place
The enclave puts compute next to the data. Most projects never export a single CRAM — they export a results table.
The enclave
Egress is audited
Where data does leave, it leaves through a reviewed export with a record of what, when, by whom and under which approval.
Standard formats
CRAM 3.1 against GRCh38 and T2T-CHM13, gVCF per sample, and joint-called VCF per cohort. No bespoke containers.
Beacon and cohort browser
Query allele presence without access, then apply for the cohorts that matter. Beacon v2 with the filtering terms extension.
Provenance and versioning
Every callset records the pipeline version, reference build and the exact sample manifest. Old callsets are never deleted.
Volume
Why we would rather you did not download it
A single 30× whole genome is a 40 GB CRAM. A modest study — five thousand cases and five thousand controls — is 400 TB before you have written a line of analysis.
- One 30× WGS CRAM: 38–44 GB
- One cohort of 10,000: ~410 TB
- Joint-called VCF for the same cohort: 2.1 TB
- The same analysis run in the enclave: 4 GB of results