Graphs and clustering¶
find_neighbors builds both graphs Seurat builds — the directed KNN graph and
the shared-nearest-neighbour graph on top of it — and find_clusters runs
community detection over the SNN.
For Louvain, Louvain with multilevel refinement and smart local moving,
find_clusters runs Seurat's own modularity optimiser: a translation of the C++
behind FindClusters, with the same restarts and the same random-number stream.
Given the same graph it returns Seurat's partition, label for label. Leiden runs
through leidenalg.
Four details here were wrong before the integration tutorial went looking, and
each mattered downstream: the KNN graph is stored directed rather than
symmetrized, the SNN keeps its diagonal and is computed in float64, run_umap
zeroes the diagonal when handed a graph, and singletons are folded into their
nearest community the way GroupSingletons does. They are noted in the
docstrings because each one changes cluster assignments, not just internals.
Neighbour graphs¶
find_neighbors
¶
find_neighbors(seurat, dims: Optional[Union[list[int], range]] = None, k_param: int = 20, assay: Optional[str] = None, reduction: str = 'pca', graph_name: Optional[str] = None, return_neighbor: bool = False, compute_snn: Optional[bool] = None, prune_snn: float = 1 / 15, seed: int = 42) -> None
Build KNN and SNN graphs from a low-dimensional embedding.
Mirrors R's FindNeighbors(pbmc, dims = 1:10). Stores Graph objects in
seurat.graphs[f"{graph_name}_nn"] and seurat.graphs[f"{graph_name}_snn"].
Parameters:
-
dims(Optional[Union[list[int], range]], default:None) –which PCs to use (0-indexed; default all available)
-
k_param(int, default:20) –number of nearest neighbors, including the cell itself
-
reduction(str, default:'pca') –which reduction to use ('pca' by default)
-
graph_name(Optional[str], default:None) –prefix for graph names (defaults to active assay name). With
return_neighbor=Trueit names theNeighboroutright, rather than acting as a prefix. -
return_neighbor(bool, default:False) –store the raw KNN result — indices and distances — as a
Neighborinseurat.neighborsinstead of building graphs. Seurat'sreturn.neighbor. -
compute_snn(Optional[bool], default:None) –build the SNN graph. Defaults to
not return_neighbor, as in Seurat, which cannot do both. -
prune_snn(float, default:1 / 15) –edges with Jaccard index below this are pruned (Seurat default 1/15)
Notes
return_neighbor=True stores a Neighbor under f"{assay}.nn" — a
dot, where the graphs use an underscore (RNA.nn against RNA_nn /
RNA_snn). That is Seurat's naming, and the separator is the only thing
distinguishing the two in a printout.
The stored indices are 0-based, where R's Indices() are 1-based.
Everything else — the k columns, self first at distance 0, the ordering — is
the same.
Source code in truecell/neighbors.py
18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 | |
find_multi_modal_neighbors
¶
find_multi_modal_neighbors(seurat, reduction_list: Sequence[str] = ('pca', 'apca'), dims_list: Optional[Sequence[Sequence[int]]] = None, k_nn: int = 20, l2_norm: bool = True, knn_graph_name: str = 'wknn', snn_graph_name: str = 'wsnn', knn_range: int = 200, prune_snn: float = 1 / 15, sd_scale: float = 1.0, cross_constant: Optional[float] = None, smooth: bool = False, seed: int = 42) -> None
Compute weighted-nearest-neighbour graphs across modalities.
Mirrors R's FindMultiModalNeighbors(obj, reduction.list =
list("pca","apca"), dims.list = list(1:30, 1:18)). Stores:
seurat.graphs[knn_graph_name]— joint KNN graphseurat.graphs[snn_graph_name]— joint SNN graph (feed tofind_clusters(graph_name=...)orrun_umap(graph=...))seurat.meta_data["<modality>.weight"]— per-cell modality weights
Parameters:
-
reduction_list(Sequence[str], default:('pca', 'apca')) –reductions to combine, one per modality
-
dims_list(Optional[Sequence[Sequence[int]]], default:None) –dims (0-indexed) to use from each reduction; default all
-
k_nn(int, default:20) –number of joint neighbours to keep per cell
-
l2_norm(bool, default:True) –L2-normalise each embedding first (R's
l2.norm) -
knn_range(int, default:200) –candidate neighbours each modality nominates before the joint re-ranking (R's
knn.range) -
prune_snn(float, default:1 / 15) –Jaccard prune threshold for the joint SNN graph
-
sd_scale(float, default:1.0) –scaling on the per-cell kernel bandwidth (R's
sd.scale) -
cross_constant(Optional[float], default:None) –denominator guard in the modality score; default 1e-4
-
smooth(bool, default:False) –average each cell's modality score over its neighbours
Notes
The weights are stored under each reduction's assay_used name, so an RNA
"pca" and an ADT "apca" produce RNA.weight and ADT.weight —
the same columns Seurat writes, and readable by the plotting functions as if
they were features.
Source code in truecell/multimodal.py
47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 | |
Community detection¶
find_clusters
¶
find_clusters(seurat, resolution: float | Sequence[float] = 0.8, algorithm: int = 1, graph_name: str | None = None, random_seed: int = 0, n_iter: int = 10, group_singletons: bool = True, cluster_name: str | Sequence[str] | None = None, modularity_fxn: int = 1, n_start: int = 10, optimizer: str = 'seurat', n_iterations: int | None = None) -> None
Cluster the cells on a neighbour graph, as Seurat's FindClusters does.
Mirrors R's FindClusters(pbmc), including its vector form
FindClusters(pbmc, resolution = c(0.4, 0.8, 1.2)).
Algorithms 1–3 run Seurat's own modularity optimiser, translated from the C++
behind RunModularityClusteringCpp. It makes n_start random starts from
one random stream, iterates each up to n_iter times, and keeps the
partition with the highest modularity. Given the same graph and settings,
truecell returns Seurat's partition label for label. Algorithm 4 is Leiden,
run by leidenalg.
Each resolution is written to its own metadata column, named
{graph_name}_res.{resolution} as Seurat names it. seurat_clusters and
the active identities are set from the last resolution in the sequence —
last as given, not largest, which is what Seurat does.
Parameters:
-
resolution(float | Sequence[float], default:0.8) –higher values give more / finer clusters. A sequence runs each in turn, as R's
resolution = c(...)does. -
algorithm(int, default:1) –1 = Louvain, 2 = Louvain with multilevel refinement, 3 = smart local moving (SLM), 4 = Leiden.
-
graph_name(str | None, default:None) –SNN graph to use (defaults to '{assay}_snn')
-
random_seed(int, default:0) –seeds the optimiser. Leiden needs a seed above 0, so there a seed of 0 or below becomes 1, with a warning, as in Seurat.
-
n_iter(int, default:10) –iterations per random start. Louvain stops early once an iteration changes nothing; SLM runs all of them. For Leiden it is leidenalg's iteration count, and a negative value runs until the partition is stable.
-
group_singletons(bool, default:True) –absorb size-1 clusters into their best-connected neighbour, as Seurat's
GroupSingletonsdoes. WithFalsethey are all pooled into one"singleton"cluster instead — again matching Seurat. -
cluster_name(str | Sequence[str] | None, default:None) –override the generated column name(s); one name, or one per resolution. Seurat's
cluster.name.seurat_clustersis still written either way. -
modularity_fxn(int, default:1) –1 = standard modularity; 2 = Seurat's alternative, in which every cell weighs 1 and
resolutionmust be at most 1. -
n_start(int, default:10) –random starts; the partition with the highest modularity wins.
-
optimizer(str, default:'seurat') –"seurat"runs Seurat's optimiser for algorithms 1–3."igraph"runs one pass of igraph'scommunity_multilevelfor algorithm 1 or 2, which is how truecell clustered up to 1.2; it reads neithern_startnorn_iter. -
n_iterations(int | None, default:None) –deprecated name for
n_iter, accepted for one more release.
Notes
Seurat's partition can depend on the processor. Built for arm64, its C++ fuses a multiply-add in the rule that decides whether a cell moves, and that flips moves whose gain is within one rounding of zero. Such near-ties need edge weights that tie exactly, such as small fractions; they did not occur on PBMC 3k's shared nearest-neighbour graph. truecell uses plain IEEE arithmetic, which is Seurat as built for x86_64.
Each resolution is clustered from the same seed rather than from a running
RNG stream, so a partition does not depend on which resolutions preceded it
or on the order they were given in. Verified against Seurat 5.5.1, where
resolution 0.8 gives the same partition alone, in c(0.4, 0.8, 1.2), and
in c(1.2, 0.8, 0.4).
Source code in truecell/clustering.py
37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 | |