The Pattern Classification module predicts RNA subcellular localization patterns for each cell-gene pair.
Command:
subcellfeatA feature-only prediction command is also available:
subcellfeat-patternSpatial transcriptomics data contain both gene identity and RNA molecule coordinates. The same gene may show different spatial distributions across cells. The Pattern module converts the spatial distribution of each cell-gene pair into an interpretable localization pattern label.
The final output column is:
pattern
The primary 8-class model supports:
Foci
Nuclear
Cytoplasmic
Nuclear edge
Cell edge
Protrusion
Radial
Random
The fallback 7-class model supports:
Nuclear
Cytoplasmic
Nuclear edge
Cell edge
Protrusion
Radial
Random
The 7-class model is retained for compatibility and controlled experiments only. The command-line defaults now use the primary 8-class model and keep Foci predictions.
PKL bundle
↓
Read data_df / cell_boundary / nuclear_boundary
↓
Standard prefiltering and cell synchronization
↓
Bento RNAforest13 feature extraction
↓
SPRAWL4 feature extraction
↓
Merge into 17-dimensional feature table
↓
Primary 8-class XGBoost prediction
↓
Keep primary 8-class prediction
↓
Save final pattern table
The classifier uses 17 spatial features.
bento_cell_inner_proximity
bento_nucleus_inner_proximity
bento_nucleus_outer_proximity
bento_cell_inner_asymmetry
bento_nucleus_inner_asymmetry
bento_nucleus_outer_asymmetry
bento_point_dispersion_norm
bento_nucleus_dispersion_norm
bento_l_max
bento_l_max_gradient
bento_l_min_gradient
bento_l_monotony
bento_l_half_radius
sprawl_peripheral
sprawl_central
sprawl_punctate
sprawl_radial
Primary model:
models/multiclass_xgb_8class_prop075_final_from_cv.joblib
Fallback model:
models/multiclass_xgb_7class_no_foci_final_from_cv.joblib
Current default:
Foci fallback disabled
The fallback mechanism can still be enabled explicitly from the Python API for compatibility checks, but it is no longer part of the public command-line default.
subcellfeat \
--pkl data/simulated_data_dict.pkl \
--out results/simulated_pattern.parquet \
--profileFor large datasets, SPRAWL punctate and radial can be approximated by reducing the sampling parameters:
subcellfeat \
--pkl data/simulated_data_dict.pkl \
--out Pattern/simulated_pattern.parquet \
--profile \
--sprawl-iterations 20 \
--sprawl-pairs 2To compute only the 17-dimensional feature table:
subcellfeat \
--pkl data/simulated_data_dict.pkl \
--out results/simulated_features.parquet \
--features-only \
--profileThen predict patterns from a feature table:
subcellfeat-pattern \
--input results/simulated_features.parquet \
--output results/simulated_pattern.parquet \
--profileTypical output includes:
cell
gene
17 feature columns
p_Foci
p_Nuclear
p_Cytoplasmic
p_Nuclear edge
p_Cell edge
p_Protrusion
p_Radial
p_Random
pattern
pattern_prob
n_transcripts
low_abundance
Default outputs include p_Foci. If an older result file lacks p_Foci, it was likely generated by the historical fallback workflow or by a non-default API call.
n_transcripts and low_abundance require the transcript count to be known, so they are
present when the pattern table is produced from a PKL bundle by subcellfeat. Running
subcellfeat-pattern on a feature table generated before these columns existed will
simply omit them.
The classifier assigns the arg-max class, pattern = argmax_k P(y=k|x). This is a
standardized computational annotation of the observed spatial distribution, not a
determination of biological localization, and it should be filtered before use.
Out-of-fold accuracy from the five-fold GroupKFold cross-validation of the shipped 8-class model (12,000 simulated cell-gene samples, 1,500 per class; macro F1 = 0.957):
pattern_prob |
share of held-out samples | held-out accuracy |
|---|---|---|
| < 0.5 | 0.5% | 35.1% |
| 0.5-0.6 | 1.4% | 54.9% |
| 0.6-0.7 | 1.5% | 48.6% |
| 0.7-0.8 | 1.7% | 66.0% |
| 0.8-0.9 | 2.8% | 79.1% |
| >= 0.9 | 92.2% | 98.4% |
Accuracy rises with probability apart from the 0.6-0.7 bin, which holds only 179
samples. Cumulatively, calls at pattern_prob >= 0.9 are 98.4% correct and calls below
that are 63.6% correct, so the probability is informative enough to filter on.
Two limits matter. First, this mapping was measured on held-out simulated data; real
data have no ground truth, so it cannot be claimed as a calibration for real datasets.
Second, the probability distribution itself shifts on real data: 92.2% of held-out
simulated samples reach pattern_prob >= 0.9, whereas in real STRAND datasets only
30-54% of cell-gene pairs do. That gap is the practical reason to filter on
pattern_prob rather than to trust every label equally.
low_abundance marks pairs whose transcript count sits at the low end of the supported
range (default n_transcripts < 10; the classifier never sees pairs below 6). It is a
descriptive marker. Low counts do not reliably imply low probability — median
pattern_prob for the lowest versus highest abundance bin across five datasets:
| dataset | n_transcripts 6-7 |
n_transcripts >= 50 |
|---|---|---|
| Molecular Cartography cardiomyocytes | 0.810 | 0.991 |
| CosMx HEK293T | 0.817 | 0.878 |
| seqFISH+ fibroblast | 0.906 | 0.911 |
| Xenium ovarian cancer | 0.832 | 0.862 |
| MERFISH U2OS | 0.890 | 0.752 |
The MERFISH U2OS trend is reversed because at high transcript counts the calls concentrate in the diffuse classes (Radial and Cytoplasmic), which are genuinely harder to separate from each other and from Random. What does hold across datasets is that the fraction of very low-probability calls shrinks as counts rise.
Treat n_transcripts and pattern_prob as two independent axes and filter on both;
they are deliberately not collapsed into a single quality score.
- The official final class column is
pattern. pattern_top1is not used as the final output column.- The default workflow includes prefiltering.
- Use
--no-prefilteronly for debugging or reproducing external workflows. --no-foci-fallbackis kept as a compatibility flag; the CLI already disables fallback by default.