Skip to content

Datasets

Seven Perturb-seq datasets have been processed through PerturbScape. Each was scored against the traits most relevant to the cell system it comes from, so the trait list differs substantially between them.

Source citations pending

The descriptions below cover the cell system and what was tested. The primary reference for each underlying screen is still to be added.

HepG2

Hepatocellular carcinoma line

Liver-derived cells, scored against circulating liver-function and lipid biomarkers.

Perturbations
255
Traits
10
Contexts
1

Jurkat

T lymphocyte leukemia line

T-cell system scored against a broad panel of autoimmune conditions and blood cell indices.

Perturbations
184
Traits
19
Contexts
1

K562-GWPS

Genome-wide Perturb-seq

By far the largest screen here, covering most expressed genes, scored against erythroid indices.

Perturbations
7,991
Traits
2
Contexts
1

Monocytes

Knockdown and knockout

The only dataset with two perturbation modalities, analyzed and plotted separately.

Perturbations
193
Traits
23
Contexts
2

PerturbAI

Brain, four neuronal subclasses

Scored against autism, with results reported separately per neuronal subclass.

Perturbations
1,855
Traits
1
Contexts
4

T2D

Pancreatic islet differentiation

Islet-directed differentiation sampled across stages, scored against glycemic traits.

Perturbations
33
Traits
2
Contexts
6

TeloHAEC

Aortic endothelial cells

Telomerase-immortalized endothelium, scored against cardiovascular traits.

Perturbations
577
Traits
3
Contexts
1

Traits tested per dataset

The trait panel was chosen to match the tissue, which is why HepG2 carries liver biomarkers and TeloHAEC carries blood pressure and coronary artery disease.

Alanine aminotransferase, Alkaline phosphatase, Aspartate aminotransferase, Cholesterol, Gamma glutamyltransferase, HDL cholesterol, LDL direct, Lipoprotein A, Total bilirubin, Triglycerides

All auto-immune, Allergy eczema diagnosed, Alzheimers, Asthma diagnosed, Celiac, Crohn's disease, Eosinophil count, Hypothyroidism self rep., IBD, Lupus, Lymphocyte count, Monocyte count, Multiple sclerosis, Primary biliary cirrhosis, Psoriasis, Rheumatoid arthritis, Type 1 diabetes, Ulcerative colitis, White count

Mean corpuscular hemoglobin, Red blood cell count

All auto-immune, Allergy eczema diagnosed, Alzheimers, Asthma diagnosed, Celiac, Crohn's disease, Eosinophil count, Hypothyroidism self rep., IBD, Lupus, Lymphocyte count, Mean corpuscular hemoglobin, Mean platelet vol., Monocyte count, Multiple sclerosis, Platelet count, Primary biliary cirrhosis, Psoriasis, Red blood cell count, Rheumatoid arthritis, Type 1 diabetes, Ulcerative colitis, White count

Autism

HbA1C, Type 2 diabetes

Coronary artery disease, Diastolic blood pressure, Systolic blood pressure

Context values

context subdivides a dataset. What it subdivides by depends on the dataset, which is why the column carries a neutral name.

Dataset Context values Meaning
Monocytes Monocytes KD, Monocytes KO Perturbation modality — knockdown versus knockout
PerturbAI 005_L4-5_IT_CTX_Glut, 052_Pvalb_Gaba, 151_TH_Prkcd_Grin2c_Glut, 155_MB_Glut Neuronal subclass
T2D D3, D7, D11, D18, D18 Endo, D18 Non-endo Differentiation stage, with day 18 additionally split by endocrine status
HepG2, Jurkat, K562-GWPS, TeloHAEC none No subdivision; context repeats the dataset name

For the four datasets without a subdivision, context is set to the dataset name so the column is never empty. The explorer hides the context selector for those.

Each context is scored and plotted independently: selecting Monocytes KD gives a different UMAP from Monocytes KO, and each T2D stage has its own.

The All entry

Every dataset also carries an entry named All that is not a perturbation: the programs learned for each perturbation, combined into a single meta-program and scored the same way. It is excluded from the perturbation counts above. See All is not a perturbation.

Notes on trait naming

Trait names come from the GWAS panel used for each dataset and are kept as they appear in the source results. The build script merges names that differ only in capitalization onto their most frequent spelling; the current source files are already consistent, so no merges are being applied.

Names are not merged across different wordings. If two panels describe a closely related quantity differently, both entries stand, because merging them would misrepresent the source.

Comparing across datasets

TRS is standardized by total trait heritability, so values are comparable across annotations and across traits. Comparisons across datasets still warrant care:

  • Screens differ enormously in size. K562-GWPS tests 7,991 perturbations against 2 traits; T2D tests 33 against 2. Multiple-testing burden is not comparable.
  • The variant-to-gene links used in annotation are restricted by tissue, so the SNP universe reachable by a program differs between a liver and a brain dataset.
  • Trait panels barely overlap, so most cross-dataset comparisons are between different traits as well as different cell systems.

UMAP coordinates are computed per dataset and trait. Positions are meaningful only within a single plot — a point at the same coordinates in two different UMAPs means nothing.