Skip to content

Kernel Contrastive Modules

kcpca kcontrapc

The linear contrastive objectives with the gene-gene covariances replaced by RBF kernels between genes, capturing nonlinear relationships between genes that a covariance cannot represent.

flowchart LR
    A["Expression<br/>cells x genes"] --> B["Min-max scale<br/>each cell"]
    B --> C["Fastfood features<br/>per gene"]
    C --> D["Gene x gene kernels<br/>K_t, K_b"]
    D --> E["Same contrast objective"]
    E --> F["Eigendecomposition<br/>gene loadings"]
    F --> G["Project cells"]

Each gene is described by its expression profile across the cells of a group, and Fastfood (sklearn_extra.kernel_approximation.Fastfood) maps that profile to random features whose inner products approximate an RBF kernel. The target and background kernels \(K_t\) and \(K_b\) are \(G \times G\), so the eigenvectors are gene loadings, as in the linear modules, and cells are projected onto them. See Contrastive Embeddings for the full construction.

Both modules search 20 alpha values instead of 40, since each evaluation is more expensive.

Shared kernel parameters

sigma_list_target: null
sigma_list_background: null
kernel_seed: 42
Parameter Type Default Description
sigma_list_target list or null null RBF bandwidth for the target kernel; null uses \(\sqrt{n_t}\), the square root of the target cell count
sigma_list_background list or null null RBF bandwidth for the background kernel; null uses \(\sqrt{n_b}\)
kernel_seed int 42 Seed for the Fastfood random features

Giving several bandwidths sums the kernels, one per value.

Reproducibility

Results are deterministic given a fixed kernel_seed. Changing the seed changes the random features and therefore the embedding. To check that a finding is not an artefact of one draw, rerun with a different seed and confirm the gene programs are stable.

Cost

The heaviest modules in the pipeline: each group's random features form a matrix of twice its cell count by the number of genes, and every alpha needs a full \(G \times G\) eigendecomposition. Suggested mem_mb: 256000, and a bigmem partition if available.


kcpca

The subtractive objective on the gene kernels:

\[ \Sigma(\alpha) = K_t - \alpha\, K_b \]

with \(\alpha \in \{0\} \cup \mathrm{logspace}(-1,\ A,\ 20)\), where \(A\) is max_log_alpha.

# kcpca/config.yaml
sigma_list_target: null
sigma_list_background: null
kernel_seed: 42

use_de_genes: false
max_log_alpha: 3
transform_mode: "auto"
manual_alpha: null
n_components: 10
project_which: "target"

Outputs

File Contents
kcpca_projection.csv Cell coordinates, columns alpha_<a>_pc_<i>
kcpca_eigenvalues.csv Eigenvalues per alpha and component
kcpca_eigenvectors.csv Gene loadings per alpha and component
kcpca_best_alphas.txt The alpha values selected
kcpca_metadata.txt Run parameters, including the seed
hotspot_kcpca_* See Hotspot

kcontrapc

The ratio objective on the gene kernels. With \(K_b + \gamma I = V \Lambda V^\top\):

\[ \Sigma(\alpha) = V \Lambda^{-\alpha} V^\top K_t, \qquad \alpha \in [0, 1] \]

Everything from contrapc carries over - the bounded alpha range, the forced inclusion of 0 and 1, and the role of gamma - with the kernels in place of the covariances.

# kcontrapc/config.yaml
sigma_list_target: null
sigma_list_background: null
kernel_seed: 42

use_de_genes: false
gamma: 0.001
transform_mode: "auto"
manual_alpha: null
n_components: 10
project_which: "target"

gamma matters more here

RBF kernel matrices have fast-decaying spectra, so \(K_b\) is often worse conditioned than the linear background covariance. If high-alpha components look like noise, raise gamma - this is the first thing to try, not the last.

Outputs

File Contents
kcontrapc_projection.csv Cell coordinates, columns alpha_<a>_pc_<i>
kcontrapc_eigenvalues.csv Eigenvalues per alpha and component
kcontrapc_eigenvectors.csv Gene loadings per alpha and component
kcontrapc_best_alphas.txt The alpha values selected, including 0.0 and 1.0
kcontrapc_metadata.txt Run parameters, including the seed
hotspot_kcontrapc_* See Hotspot

See also