Skip to content

Latest commit

 

History

History
281 lines (212 loc) · 6.72 KB

File metadata and controls

281 lines (212 loc) · 6.72 KB

Pattern Classification Module

The Pattern Classification module predicts RNA subcellular localization patterns for each cell-gene pair.

Command:

subcellfeat

A feature-only prediction command is also available:

subcellfeat-pattern

1. Purpose

Spatial transcriptomics data contain both gene identity and RNA molecule coordinates. The same gene may show different spatial distributions across cells. The Pattern module converts the spatial distribution of each cell-gene pair into an interpretable localization pattern label.

The final output column is:

pattern

2. Supported classes

The primary 8-class model supports:

Foci
Nuclear
Cytoplasmic
Nuclear edge
Cell edge
Protrusion
Radial
Random

The fallback 7-class model supports:

Nuclear
Cytoplasmic
Nuclear edge
Cell edge
Protrusion
Radial
Random

The 7-class model is retained for compatibility and controlled experiments only. The command-line defaults now use the primary 8-class model and keep Foci predictions.


3. Main workflow

PKL bundle
   ↓
Read data_df / cell_boundary / nuclear_boundary
   ↓
Standard prefiltering and cell synchronization
   ↓
Bento RNAforest13 feature extraction
   ↓
SPRAWL4 feature extraction
   ↓
Merge into 17-dimensional feature table
   ↓
Primary 8-class XGBoost prediction
   ↓
Keep primary 8-class prediction
   ↓
Save final pattern table

4. Feature system

The classifier uses 17 spatial features.

4.1 Bento RNAforest13 features

bento_cell_inner_proximity
bento_nucleus_inner_proximity
bento_nucleus_outer_proximity
bento_cell_inner_asymmetry
bento_nucleus_inner_asymmetry
bento_nucleus_outer_asymmetry
bento_point_dispersion_norm
bento_nucleus_dispersion_norm
bento_l_max
bento_l_max_gradient
bento_l_min_gradient
bento_l_monotony
bento_l_half_radius

4.2 SPRAWL features

sprawl_peripheral
sprawl_central
sprawl_punctate
sprawl_radial

5. Default models

Primary model:

models/multiclass_xgb_8class_prop075_final_from_cv.joblib

Fallback model:

models/multiclass_xgb_7class_no_foci_final_from_cv.joblib

Current default:

Foci fallback disabled

The fallback mechanism can still be enabled explicitly from the Python API for compatibility checks, but it is no longer part of the public command-line default.


6. Basic usage

subcellfeat \
  --pkl data/simulated_data_dict.pkl \
  --out results/simulated_pattern.parquet \
  --profile

For large datasets, SPRAWL punctate and radial can be approximated by reducing the sampling parameters:

subcellfeat \
  --pkl data/simulated_data_dict.pkl \
  --out Pattern/simulated_pattern.parquet \
  --profile \
  --sprawl-iterations 20 \
  --sprawl-pairs 2

7. Feature-only mode

To compute only the 17-dimensional feature table:

subcellfeat \
  --pkl data/simulated_data_dict.pkl \
  --out results/simulated_features.parquet \
  --features-only \
  --profile

Then predict patterns from a feature table:

subcellfeat-pattern \
  --input results/simulated_features.parquet \
  --output results/simulated_pattern.parquet \
  --profile

8. Output columns

Typical output includes:

cell
gene
17 feature columns
p_Foci
p_Nuclear
p_Cytoplasmic
p_Nuclear edge
p_Cell edge
p_Protrusion
p_Radial
p_Random
pattern
pattern_prob
n_transcripts
low_abundance

Default outputs include p_Foci. If an older result file lacks p_Foci, it was likely generated by the historical fallback workflow or by a non-default API call.

n_transcripts and low_abundance require the transcript count to be known, so they are present when the pattern table is produced from a PKL bundle by subcellfeat. Running subcellfeat-pattern on a feature table generated before these columns existed will simply omit them.


9. Interpreting the annotation

The classifier assigns the arg-max class, pattern = argmax_k P(y=k|x). This is a standardized computational annotation of the observed spatial distribution, not a determination of biological localization, and it should be filtered before use.

9.1 What pattern_prob means

Out-of-fold accuracy from the five-fold GroupKFold cross-validation of the shipped 8-class model (12,000 simulated cell-gene samples, 1,500 per class; macro F1 = 0.957):

pattern_prob share of held-out samples held-out accuracy
< 0.5 0.5% 35.1%
0.5-0.6 1.4% 54.9%
0.6-0.7 1.5% 48.6%
0.7-0.8 1.7% 66.0%
0.8-0.9 2.8% 79.1%
>= 0.9 92.2% 98.4%

Accuracy rises with probability apart from the 0.6-0.7 bin, which holds only 179 samples. Cumulatively, calls at pattern_prob >= 0.9 are 98.4% correct and calls below that are 63.6% correct, so the probability is informative enough to filter on.

Two limits matter. First, this mapping was measured on held-out simulated data; real data have no ground truth, so it cannot be claimed as a calibration for real datasets. Second, the probability distribution itself shifts on real data: 92.2% of held-out simulated samples reach pattern_prob >= 0.9, whereas in real STRAND datasets only 30-54% of cell-gene pairs do. That gap is the practical reason to filter on pattern_prob rather than to trust every label equally.

9.2 What low_abundance does and does not mean

low_abundance marks pairs whose transcript count sits at the low end of the supported range (default n_transcripts < 10; the classifier never sees pairs below 6). It is a descriptive marker. Low counts do not reliably imply low probability — median pattern_prob for the lowest versus highest abundance bin across five datasets:

dataset n_transcripts 6-7 n_transcripts >= 50
Molecular Cartography cardiomyocytes 0.810 0.991
CosMx HEK293T 0.817 0.878
seqFISH+ fibroblast 0.906 0.911
Xenium ovarian cancer 0.832 0.862
MERFISH U2OS 0.890 0.752

The MERFISH U2OS trend is reversed because at high transcript counts the calls concentrate in the diffuse classes (Radial and Cytoplasmic), which are genuinely harder to separate from each other and from Random. What does hold across datasets is that the fraction of very low-probability calls shrinks as counts rise.

Treat n_transcripts and pattern_prob as two independent axes and filter on both; they are deliberately not collapsed into a single quality score.


10. Notes

  1. The official final class column is pattern.
  2. pattern_top1 is not used as the final output column.
  3. The default workflow includes prefiltering.
  4. Use --no-prefilter only for debugging or reproducing external workflows.
  5. --no-foci-fallback is kept as a compatibility flag; the CLI already disables fallback by default.