Kernel Contrastive Modules¶
kcpca kcontrapc
The linear contrastive objectives with the gene-gene covariances replaced by RBF kernels between genes, capturing nonlinear relationships between genes that a covariance cannot represent.
flowchart LR
A["Expression<br/>cells x genes"] --> B["Min-max scale<br/>each cell"]
B --> C["Fastfood features<br/>per gene"]
C --> D["Gene x gene kernels<br/>K_t, K_b"]
D --> E["Same contrast objective"]
E --> F["Eigendecomposition<br/>gene loadings"]
F --> G["Project cells"] Each gene is described by its expression profile across the cells of a group, and Fastfood (sklearn_extra.kernel_approximation.Fastfood) maps that profile to random features whose inner products approximate an RBF kernel. The target and background kernels \(K_t\) and \(K_b\) are \(G \times G\), so the eigenvectors are gene loadings, as in the linear modules, and cells are projected onto them. See Contrastive Embeddings for the full construction.
Both modules search 20 alpha values instead of 40, since each evaluation is more expensive.
Shared kernel parameters¶
| Parameter | Type | Default | Description |
|---|---|---|---|
sigma_list_target | list or null | null | RBF bandwidth for the target kernel; null uses \(\sqrt{n_t}\), the square root of the target cell count |
sigma_list_background | list or null | null | RBF bandwidth for the background kernel; null uses \(\sqrt{n_b}\) |
kernel_seed | int | 42 | Seed for the Fastfood random features |
Giving several bandwidths sums the kernels, one per value.
Reproducibility¶
Results are deterministic given a fixed kernel_seed. Changing the seed changes the random features and therefore the embedding. To check that a finding is not an artefact of one draw, rerun with a different seed and confirm the gene programs are stable.
Cost¶
The heaviest modules in the pipeline: each group's random features form a matrix of twice its cell count by the number of genes, and every alpha needs a full \(G \times G\) eigendecomposition. Suggested mem_mb: 256000, and a bigmem partition if available.
kcpca¶
The subtractive objective on the gene kernels:
with \(\alpha \in \{0\} \cup \mathrm{logspace}(-1,\ A,\ 20)\), where \(A\) is max_log_alpha.
# kcpca/config.yaml
sigma_list_target: null
sigma_list_background: null
kernel_seed: 42
use_de_genes: false
max_log_alpha: 3
transform_mode: "auto"
manual_alpha: null
n_components: 10
project_which: "target"
Outputs¶
| File | Contents |
|---|---|
kcpca_projection.csv | Cell coordinates, columns alpha_<a>_pc_<i> |
kcpca_eigenvalues.csv | Eigenvalues per alpha and component |
kcpca_eigenvectors.csv | Gene loadings per alpha and component |
kcpca_best_alphas.txt | The alpha values selected |
kcpca_metadata.txt | Run parameters, including the seed |
hotspot_kcpca_* | See Hotspot |
kcontrapc¶
The ratio objective on the gene kernels. With \(K_b + \gamma I = V \Lambda V^\top\):
Everything from contrapc carries over - the bounded alpha range, the forced inclusion of 0 and 1, and the role of gamma - with the kernels in place of the covariances.
# kcontrapc/config.yaml
sigma_list_target: null
sigma_list_background: null
kernel_seed: 42
use_de_genes: false
gamma: 0.001
transform_mode: "auto"
manual_alpha: null
n_components: 10
project_which: "target"
gamma matters more here
RBF kernel matrices have fast-decaying spectra, so \(K_b\) is often worse conditioned than the linear background covariance. If high-alpha components look like noise, raise gamma - this is the first thing to try, not the last.
Outputs¶
| File | Contents |
|---|---|
kcontrapc_projection.csv | Cell coordinates, columns alpha_<a>_pc_<i> |
kcontrapc_eigenvalues.csv | Eigenvalues per alpha and component |
kcontrapc_eigenvectors.csv | Gene loadings per alpha and component |
kcontrapc_best_alphas.txt | The alpha values selected, including 0.0 and 1.0 |
kcontrapc_metadata.txt | Run parameters, including the seed |
hotspot_kcontrapc_* | See Hotspot |
See also¶
- Linear Contrastive - the linear forms of both objectives
- Contrastive Embeddings - full mathematical treatment