Skip to content

Documentation

PerturbScape has two stages that run separately.

Stage 1 - program discovery is a Snakemake workflow of eight modules, built for a SLURM cluster. It takes a Perturb-seq .h5ad and produces gene programs for every perturbation.

Stage 2 - disease enrichment runs outside Snakemake. It takes those programs plus GWAS resources and produces a Trait Relevance Score (TRS) - a standardized heritability effect size - for every perturbation-trait pair.

flowchart LR
    subgraph S1["Stage 1 — Snakemake"]
        A[".h5ad"] --> B["8 modules"] --> C["Gene programs"]
    end
    subgraph S2["Stage 2 — run separately"]
        C --> D["Program selection"] --> E["Meta-programs"] --> F["S-LDSC"] --> G["TRS"]
    end

Start here

01

Getting Started

Install Snakemake and the cluster shim, check your .h5ad has what the pipeline needs, and run a first module end to end.

02

Configuration

The master config, how module configs override it, analysis modes, and how to size SLURM resources.

03

Running the Pipeline

Every run_modules.sh action and flag, where the logs are, and what to do when a job fails.

04

Outputs

Where results land and what every file produced by every module contains.

Understand the methods

Methods overview

Why contrastive analysis, and how the three layers of the pipeline fit together.

Contrastive Embeddings

The subtractive and ratio objectives, automatic selection of the contrast strength, and the kernel variants.

Gene Programs and Hotspot

Turning an embedding into interpretable gene modules, and why the graph is built on the contrastive space.

Disease Enrichment

Program selection, meta-program construction, variant annotation, S-LDSC, and TRS - with the code for each step.

Module reference

The eight modules, grouped by what they share:

Group Modules Contrastive Hotspot
Linear contrastive cpca, contrapc yes yes
Kernel contrastive kcpca, kcontrapc yes yes
Non-contrastive baselines pca, cnmf no yes
ContrastiveVI contrastivevi yes yes
DE and DGCA de-dgca yes no

Hotspot is not a module of its own - it runs as a second rule inside each of the seven modules that produce an embedding.

Requirements

Requirement Notes
HPC cluster with SLURM run_modules.sh wraps sbatch; each module ships a SLURM profile
Conda or Miniconda Every module builds its own isolated environment
Snakemake 8+ Earlier versions lack the executor-plugin architecture
A Perturb-seq .h5ad Must contain control cells - see Getting Started

Stage 2 additionally needs MAGMA gene-level results, GWAS summary statistics, an LD reference panel, and variant-to-gene maps. See Disease Enrichment.

Not a laptop workload

Default requests are 128 GB for the main jobs and 64 GB for Hotspot, rising to 256 GB for the kernel modules. Running locally is realistic only for very small test data.