This inventory separates three facts that were previously conflated: when an upstream source was refreshed, when the local PostgreSQL table was observed, and when a replaceable snapshot was copied to the cloud. Local observation is the source of truth for research queries. See the immutable run manifests for the exact inputs used by an individual result.
—
Registered sources
—
Observed tables
—
Populated tables
—
Local rows
—
Local table data
—
Snapshot copied
How to read this inventory
“Observed” proves the local table existed and records its estimated row count, physical size, and schema hash. “Upstream” is only populated when an ingestion job records a real source refresh. “Snapshot” is operational availability, not scientific freshness. Cloud copies are intentionally limited to replaceable tables; provenance and research artifacts use immutable content-addressed storage instead.
Query the copied lake
Search bounded, read-only snapshots across genes, variants, drugs, diseases, pathways, phenotypes, pharmacogenomics, and structures. For a precise lookup, use the documented table and column catalogue.
FAIRdata.ai Collections is the discovery and qualification layer; this data lake is the execution and query layer. Collections pass explicit metadata, licence, curation-evidence, and resolvable-file gates before they are candidates for ingestion.
Loading version, licence, completeness, and validation contracts…
What each database can rigorously prove?
The data lake isn't just storage — each database enables a specific rigorous analysis. Cross-database evidence from independent families is triangulated by construction.