Datasets¶
Seven Perturb-seq datasets have been processed through PerturbScape. Each was scored against the traits most relevant to the cell system it comes from, so the trait list differs substantially between them.
Source citations pending
The descriptions below cover the cell system and what was tested. The primary reference for each underlying screen is still to be added.
HepG2¶
Hepatocellular carcinoma line
Liver-derived cells, scored against circulating liver-function and lipid biomarkers.
- Perturbations
- 255
- Traits
- 10
- Contexts
- 1
Jurkat¶
T lymphocyte leukemia line
T-cell system scored against a broad panel of autoimmune conditions and blood cell indices.
- Perturbations
- 184
- Traits
- 19
- Contexts
- 1
K562-GWPS¶
Genome-wide Perturb-seq
By far the largest screen here, covering most expressed genes, scored against erythroid indices.
- Perturbations
- 7,991
- Traits
- 2
- Contexts
- 1
Monocytes¶
Knockdown and knockout
The only dataset with two perturbation modalities, analyzed and plotted separately.
- Perturbations
- 193
- Traits
- 23
- Contexts
- 2
PerturbAI¶
Brain, four neuronal subclasses
Scored against autism, with results reported separately per neuronal subclass.
- Perturbations
- 1,855
- Traits
- 1
- Contexts
- 4
T2D¶
Pancreatic islet differentiation
Islet-directed differentiation sampled across stages, scored against glycemic traits.
- Perturbations
- 33
- Traits
- 2
- Contexts
- 6
TeloHAEC¶
Aortic endothelial cells
Telomerase-immortalized endothelium, scored against cardiovascular traits.
- Perturbations
- 577
- Traits
- 3
- Contexts
- 1
Traits tested per dataset¶
The trait panel was chosen to match the tissue, which is why HepG2 carries liver biomarkers and TeloHAEC carries blood pressure and coronary artery disease.
Alanine aminotransferase, Alkaline phosphatase, Aspartate aminotransferase, Cholesterol, Gamma glutamyltransferase, HDL cholesterol, LDL direct, Lipoprotein A, Total bilirubin, Triglycerides
All auto-immune, Allergy eczema diagnosed, Alzheimers, Asthma diagnosed, Celiac, Crohn's disease, Eosinophil count, Hypothyroidism self rep., IBD, Lupus, Lymphocyte count, Monocyte count, Multiple sclerosis, Primary biliary cirrhosis, Psoriasis, Rheumatoid arthritis, Type 1 diabetes, Ulcerative colitis, White count
Mean corpuscular hemoglobin, Red blood cell count
All auto-immune, Allergy eczema diagnosed, Alzheimers, Asthma diagnosed, Celiac, Crohn's disease, Eosinophil count, Hypothyroidism self rep., IBD, Lupus, Lymphocyte count, Mean corpuscular hemoglobin, Mean platelet vol., Monocyte count, Multiple sclerosis, Platelet count, Primary biliary cirrhosis, Psoriasis, Red blood cell count, Rheumatoid arthritis, Type 1 diabetes, Ulcerative colitis, White count
Autism
HbA1C, Type 2 diabetes
Coronary artery disease, Diastolic blood pressure, Systolic blood pressure
Context values¶
context subdivides a dataset. What it subdivides by depends on the dataset, which is why the column carries a neutral name.
| Dataset | Context values | Meaning |
|---|---|---|
| Monocytes | Monocytes KD, Monocytes KO | Perturbation modality — knockdown versus knockout |
| PerturbAI | 005_L4-5_IT_CTX_Glut, 052_Pvalb_Gaba, 151_TH_Prkcd_Grin2c_Glut, 155_MB_Glut | Neuronal subclass |
| T2D | D3, D7, D11, D18, D18 Endo, D18 Non-endo | Differentiation stage, with day 18 additionally split by endocrine status |
| HepG2, Jurkat, K562-GWPS, TeloHAEC | none | No subdivision; context repeats the dataset name |
For the four datasets without a subdivision, context is set to the dataset name so the column is never empty. The explorer hides the context selector for those.
Each context is scored and plotted independently: selecting Monocytes KD gives a different UMAP from Monocytes KO, and each T2D stage has its own.
The All entry¶
Every dataset also carries an entry named All that is not a perturbation: the programs learned for each perturbation, combined into a single meta-program and scored the same way. It is excluded from the perturbation counts above. See All is not a perturbation.
Notes on trait naming¶
Trait names come from the GWAS panel used for each dataset and are kept as they appear in the source results. The build script merges names that differ only in capitalization onto their most frequent spelling; the current source files are already consistent, so no merges are being applied.
Names are not merged across different wordings. If two panels describe a closely related quantity differently, both entries stand, because merging them would misrepresent the source.
Comparing across datasets¶
TRS is standardized by total trait heritability, so values are comparable across annotations and across traits. Comparisons across datasets still warrant care:
- Screens differ enormously in size. K562-GWPS tests 7,991 perturbations against 2 traits; T2D tests 33 against 2. Multiple-testing burden is not comparable.
- The variant-to-gene links used in annotation are restricted by tissue, so the SNP universe reachable by a program differs between a liver and a brain dataset.
- Trait panels barely overlap, so most cross-dataset comparisons are between different traits as well as different cell systems.
UMAP coordinates are computed per dataset and trait. Positions are meaningful only within a single plot — a point at the same coordinates in two different UMAPs means nothing.