Contrastive Embeddings¶
All contrastive modules share the same setup. Let \(X_t\) be the scaled expression matrix for target (perturbed) cells and \(X_b\) the scaled matrix for background (control) cells, with covariances
Both matrices are standardized independently before the covariances are formed, and the pipeline skips scaling if a block already has mean \(\approx 0\) and standard deviation \(\approx 1\).
The methods differ in how they combine \(C_t\) and \(C_b\) into a single matrix whose eigenvectors become the embedding.
The two objectives¶
Used by cpca and its kernel form kcpca.
Directions are penalized in proportion to how much variance they carry in the background. At \(\alpha = 0\) this is plain PCA on the target cells; as \(\alpha\) grows, shared structure is progressively subtracted away.
The search grid is logarithmic:
where \(A\) is the max_log_alpha setting.
Used by contrapc and its kernel form kcontrapc.
With the regularized background eigendecomposition \(C_b + \gamma I = V \Lambda V^\top\),
Instead of subtracting the background, this whitens by it. The interpolation is bounded:
- \(\alpha = 0\) → plain PCA on the target
- \(\alpha = 1\) → full background whitening, \(C_b^{-1} C_t\)
The search grid is therefore linear and closed:
\(\gamma\) (gamma, default 0.001) is a ridge term that keeps \(C_b\) invertible. Background covariances in single-cell data are routinely rank-deficient, so this is load-bearing, not cosmetic.
Why both?
Subtraction and whitening fail differently. cPCA's \(\alpha\) is unbounded and its scale depends on the relative magnitudes of the two covariances, so the useful range varies between datasets. ContraNorm's \(\alpha\) is bounded in \([0, 1]\) and comparable across datasets, but it needs \(C_b\) to be well-conditioned and pays for the eigendecomposition.
Choosing α automatically¶
With transform_mode: "auto" (the default), the pipeline does not pick a single \(\alpha\). It picks four, by spectral clustering over the whole grid:
flowchart TB
A["Evaluate Σ(α) for all 40 α values"] --> B[Each α gives a subspace<br/>of top eigenvectors]
B --> C[Affinity matrix:<br/>pairwise subspace similarity]
C --> D[Spectral clustering<br/>into 4 clusters]
D --> E[One exemplar α per cluster]
E --> F[Return 4 representative α<br/>and their embeddings] The idea is that as \(\alpha\) sweeps its range, the resulting subspace does not change smoothly — it stays roughly fixed, then snaps to a different regime. Clustering the grid by subspace similarity finds those regimes and returns one representative from each, so you get a spread of qualitatively distinct embeddings rather than 40 near-duplicates.
The two objectives handle the endpoints differently:
| cPCA | ContraNorm | |
|---|---|---|
| Grid | [0] + logspace(-1, max_log_alpha, 40) | linspace(0, 1, 40) |
| Cluster containing \(\alpha=0\) | Skipped | Skipped |
| Forced into the result | — | \(\alpha = 0\) and \(\alpha = 1\) |
| Final selection | 4 cluster exemplars | [0, exemplar₁, exemplar₂, 1] |
Output columns are named for the \(\alpha\) that produced them, e.g. alpha_0.271_pc_1, so every component remains traceable to its regularization level.
Manual mode¶
To fix a single \(\alpha\):
This runs one decomposition and skips the clustering entirely. Useful when you have already identified a good \(\alpha\) and want reproducible, cheaper runs across many perturbations.
Kernel variants¶
kcpca and kcontrapc use the same two objectives, but replace the gene-gene covariances with RBF kernels between genes, so two genes count as similar when their expression profiles across cells are close, not only when they are linearly correlated.
Each group is first min-max scaled per cell, a step skipped if its values already lie in \([0, 1]\). Every gene \(g\) is then represented by its profile across that group's \(n\) cells, \(x_g \in \mathbb{R}^{n}\), and mapped to \(D\) random features \(\phi(x_g)\) with the Fastfood transform (sklearn_extra.kernel_approximation.Fastfood), chosen so that \(\phi(x_g)^\top \phi(x_h)\) approximates the RBF kernel \(k(x_g, x_h)\). Stacking the features into \(\Phi_t \in \mathbb{R}^{D \times G}\), one column per gene, gives
both \(G \times G\), with \(D = 2n\) features per bandwidth. \(K_t\) and \(K_b\) take the place of \(C_t\) and \(C_b\) in both objectives, so the eigenvectors are still gene loadings, and cells are projected onto them as in the linear modules, using the min-max scaled expression. When choosing \(\alpha\), candidate subspaces are compared on the random features \(\Phi_t\) rather than on the cells.
| Parameter | Default | Meaning |
|---|---|---|
sigma_list_target | null | RBF bandwidth for the target kernel; null → \(\sqrt{n_t}\) |
sigma_list_background | null | RBF bandwidth for the background kernel; null → \(\sqrt{n_b}\) |
kernel_seed | 42 | Seed for the Fastfood random features |
Giving several bandwidths sums the kernels, one per value. Kernel modules search 20 \(\alpha\) values rather than 40, since each evaluation is more expensive.
Preprocessing and guards¶
Applied inside every contrastive module before fitting:
| Step | Behavior |
|---|---|
use_de_genes: true | Restricts to genes with \(p \le 0.05\) between target and background |
| Minimum group size | Hard error if target or background has < 25 cells |
| Non-finite values | Replaced with 0, with a warning naming the count |
| Scaling | Linear modules standardize each gene; kernel modules min-max scale each cell. Target and background are scaled separately |
No cell or gene filtering happens at this stage, so filter the .h5ad beforehand if you need it.
The 25-cell floor is independent of min_cells
min_cells filters targets at the Snakemake level, before jobs are submitted. The 25-cell check happens inside the fit. If you set min_cells: 0, perturbations with fewer than 25 cells will be submitted and then fail at runtime. Setting min_cells: 25 or higher avoids burning cluster time on jobs that cannot succeed.
Projection scope¶
project_which controls which cells are projected into the learned space:
| Value | Projects | Use when |
|---|---|---|
"target" (default) | Perturbed cells only | Studying structure within perturbed cells |
"background" | Control cells only | Sanity check — should show little structure |
"both" | All cells | Visualizing separation between the two groups |
Hotspot consumes the projection, so this choice determines which cells the gene programs are learned from.