A UK DRI fork of nf-core/scdownstream.
Important
This pipeline began life as a fork of nf-core/scdownstream and has since diverged substantially. It is now developed independently by UK DRI Informatics and follows its own release and configuration conventions.
Please raise issues and questions in this repository. For the original, nf-core community-supported pipeline, see nf-co.re/scdownstream.
The original authors and contributors are credited in
ACKNOWLEDGEMENTS.md.
UK DRI scdownstream is a Nextflow pipeline for the downstream analysis of processed single-cell RNA-seq data. It takes per-sample count matrices (h5ad, 10x h5, SingleCellExperiment / Seurat RDS, or CSV) and carries them through quality control, ambient RNA correction, doublet detection, cell type annotation, scVI integration, dimensionality reduction and clustering, then on through marker genes, gene set enrichment and cell–cell communication, and finally pseudobulk differential expression.
The input is ideally the output of nf-core/scrnaseq or a similar
per-sample count matrix object files. We recommend supplying it as AnnData (.h5ad) objects: AnnData is
part of the scverse ecosystem that this pipeline is built on, and it scales
to large count matrices without the size limits that SingleCellExperiment and Seurat objects run into
in R.
The pipeline is organised as three sequential stages, each its own Nextflow entry point, so that long-running analyses can be checkpointed and re-run independently.
It inherits the learnings and implementations of the following pipelines (alphabetical), via nf-core/scdownstream:
The pipeline is always invoked with an explicit -entry. The three stages are sequential: each
consumes the previous stage's finalized .h5ad via --base_adata.
| Stage | Entry point | Required input | Main output |
|---|---|---|---|
| 1 | -entry qc_clustering |
--input samplesheet.csv |
<name>_qc_clustering.h5ad, MultiQC report, QC/clustering HTML report |
| 2 | -entry downstream |
--base_adata <stage 1 h5ad> |
<name>_downstream.h5ad, <name>_downstream_markers.json.gz, analysis HTML report |
| 3 | -entry differential_genes |
--base_adata <stage 2 h5ad> + --diffgenes_contrasts contrasts.tsv |
per-contrast DE tables, pseudobulk h5ad, DE HTML report |
samplesheet.csv
│
▼ -entry qc_clustering
<name>_qc_clustering.h5ad ── QC, integration (scVI), UMAP, Leiden clustering
│
▼ -entry downstream --base_adata
<name>_downstream.h5ad ── marker genes, enrichment, LIANA+ cell–cell communication
│
▼ -entry differential_genes --base_adata --diffgenes_contrasts
differential_genes/ ── decoupler pseudobulk → PyDESeq2 per group × contrast
Important
Always pass -entry, so it is explicit which stage runs and on what input. Without it the
pipeline falls back to stage 1 (qc_clustering) and prints a warning; it never runs downstream
or differential_genes on its own. The upstream single-pass workflow has been removed.
- Load and convert inputs to h5ad (h5ad, 10x h5, RDS, CSV), add a
samplecolumn toobs, and optionally per-sample metadata from a TSV (--metadata) - Per-sample quality control
- QC metrics for raw counts (
MultiQC) - Doublet detection — scrublet
(doublets are annotated, not removed unless
--doublet_removalis set) - Ambient RNA correction — decontX (default), soupX, CellBender, scAR
- Cell filtering — fixed thresholds, or automatic N-MAD outlier thresholds via
--automatic_cell_filtering; with--filtering_keep_outliersthe outliers are kept and marked inobs["outlier"]instead of removed
- QC metrics for raw counts (
- Cell type annotation — CellTypist and/or SingleR with celldex references
- Harmonise gene symbol, batch and cell type columns across samples, resolving duplicate gene symbols
- Merge and integrate — scVI
- Embeddings and clustering — log-normalisation, HVG selection, PCA, neighbours,
UMAP,
Leiden at every
resolution in
--clustering_resolutions - Reports — a Quarto QC/clustering report and a
MultiQCreport
--qc_only stops after step 3, producing per-sample QC objects without merging or integration.
- Marker genes per cluster —
scanpy.tl.rank_genes_groups(Wilcoxon) on the clustering named by--selected_clustering - Gene set enrichment over those markers
- Cell–cell communication — LIANA+ rank aggregation, using HCOP orthologs for non-human species
- Marker export to
<name>_downstream_markers.json.gz, and h5ad/RDS finalisation - A Quarto analysis report
- Pseudobulk aggregation per sample × group label — decoupler
(
countslayer, sum mode) - Drop pseudobulk samples below
--diffgenes_min_counts/--diffgenes_min_cells - Split into one object per group label (cluster or cell type,
--diffgenes_group_col) - Differential expression per group label × contrast —
PyDESeq2, design
~ variable [+ blocking…] - A Quarto differential expression report across all results
Contrasts are supplied as a TSV following the
nf-core/differentialabundance definition — see
assets/contrasts.tsv and the
usage documentation.
This fork focuses on a curated set of tools — the approaches we have validated on UK DRI data:
| Step | Supported |
|---|---|
| Doublet detection | scrublet |
| Ambient RNA correction | decontx (default), soupx, cellbender, scar, or none |
| Integration | scvi |
| Clustering | Leiden, at multiple resolutions |
| Cell type annotation | CellTypist, SingleR / celldex |
| Differential expression | PyDESeq2 on decoupler pseudobulk |
Further tools will be added as they are curated and validated. Until then, please use the values above — see Supported tool choices for the details.
You need:
- Nextflow ≥ 24.10.5, which requires Java 17 or later.
- Apptainer (or Singularity). This is the recommended way to run the pipeline. Docker also works, but see the Docker Hub pull-limit note under Container images.
- An x86_64 Linux machine or cluster. The container images are published for
linux/amd64only. - Internet access from the machine running the tasks: images are pulled on first use, CellTypist models are downloaded at run time, and gene set enrichment queries the g:Profiler web service.
UK DRI users: see the UK DRI Informatics wiki for how to run the pipeline on the cluster.
Cell–cell communication (LIANA+, in -entry downstream) uses a human ligand–receptor resource.
Mapping it onto another species needs an HCOP ortholog table, which you download once:
-
Set
--specieson every stage. Supported values arehuman(default) andmouse. -
Download the HCOP table for your species from the HGNC HCOP downloads. Use the fifteen-column file, and keep its original name, because LIANA+ looks for
<directory>/human_<species>_hcop_fifteen_column.txt.gz:mkdir -p hcop curl -o hcop/human_mouse_hcop_fifteen_column.txt.gz \ https://storage.googleapis.com/public-download-files/hcop/human_mouse_hcop_fifteen_column.txt.gz -
Pass the directory to
-entry downstreamwith--ortholog_hcop_directory:nextflow run UKDRI/scdownstream -r dev_ukdri -entry downstream \ -profile apptainer \ --base_adata results/qc_clustering/my_study_qc_clustering.h5ad \ --name my_study \ --species mouse \ --ortholog_hcop_directory hcop \ --outdir results/downstream
Without it, -entry downstream stops with an error for a non-human --species. Human data does not
need it. See Reference data for details.
Note
Install the requirements listed under Prerequisites first. If you are unsure
about the filtered / unfiltered distinction in the samplesheet, see
Filtered and unfiltered matrices.
Prepare a samplesheet describing your per-sample matrices:
sample,unfiltered
sample1,/absolute/path/to/sample1.h5ad
sample2,/absolute/path/to/sample2.h5
sample3,relative/path/to/sample3.rds
sample4,/absolute/path/to/sample4.csvEach entry is an h5ad, 10x h5, RDS or CSV file. RDS files may contain any object convertible to a
SingleCellExperiment via Seurat's as.SingleCellExperiment.
CSV files should hold a matrix with genes as columns and cells as rows, the first column being cell
barcodes.
Stage 1 — QC, integration and clustering:
nextflow run UKDRI/scdownstream -r dev_ukdri -entry qc_clustering \
-profile apptainer \
--input samplesheet.csv \
--metadata sample_metadata.tsv \
--umap_color_by diagnosis,sex \
--name my_study \
--species human \
--outdir results/qc_clustering--metadata is optional: a tab-separated file with one row per sample, whose columns (donor,
diagnosis, sex, ...) are added to every cell's obs. See
Per-sample metadata. --umap_color_by plots such columns on
the PCA and scVI UMAPs in the reports, i.e. before and after integration. See
UMAP colourings.
Stage 2 — downstream analysis:
nextflow run UKDRI/scdownstream -r dev_ukdri -entry downstream \
-profile apptainer \
--base_adata results/qc_clustering/my_study_qc_clustering.h5ad \
--name my_study \
--species human \
--outdir results/downstreamStage 3 — differential expression:
nextflow run UKDRI/scdownstream -r dev_ukdri -entry differential_genes \
-profile apptainer \
--base_adata results/downstream/my_study_downstream.h5ad \
--diffgenes_contrasts contrasts.tsv \
--diffgenes_group_col cell_type \
--diffgenes_sample_col sample \
--name my_study \
--outdir results/differential_genesWarning
Provide pipeline parameters on the command line or via -params-file. Custom config files passed
with -c can supply any Nextflow configuration except parameters.
Every process runs from a public container image; nothing has to be built by hand. Nextflow pulls
each image on first use, converting it to a .sif in $NXF_SINGULARITY_CACHEDIR under the
singularity or apptainer profile. Two images are specific to this fork; they are hosted on
Docker Hub and pulled the same way. Their Dockerfiles are kept in the repository only as the recipe
they were built from:
| Image (Docker Hub) | Required by | Recipe |
|---|---|---|
docker.io/nhecker/scanpy-report:1.11.4-coreinf0.4 |
SCANPY_GENERATE_REPORT, SCANPY_GENERATE_REPORT_QC, PYDESEQ2_GENERATE_REPORT, SCANPY_ENRICH |
modules/local/scanpy/report/Dockerfile |
docker.io/nhecker/pydeseq2:0.1 |
DIFFERENTIAL_GENES_PER_CONTRAST |
modules/local/pydeseq2/differential_genes/Dockerfile |
Both extend the public gcfntnu/scanpy:1.11.4 image, which DECOUPLER_PSEUDOBULK, FILTER_PSEUDOBULK and
SCANPY_EXPORT_MARKERS use directly. The images are published for linux/amd64 only.
Important
We recommend running with -profile apptainer (or singularity), not -profile docker. These
images are hosted on Docker Hub, which limits anonymous pulls on free accounts. Apptainer pulls each
image once and reuses the .sif from the cache for every later run. Set
NXF_SINGULARITY_CACHEDIR to a shared directory so every node and every run uses the same copy.
Docker pulls through each machine's own daemon, so a multi-node or cloud run can repeat the same
pull many times and hit the limit mid-run. If you must use Docker, run docker login first, or
pre-pull the images on each host.
This pipeline is work in development. The following are known and, for now, expected behaviours. Several of them silently affect results, so please read before interpreting output.
- Doublets are flagged, not removed, by default. scrublet writes its scores and predictions
into the object. Set
--doublet_removalto remove them right after detection, with--doublet_detection_thresholdmethods having to agree (default 1). - Keep
scviin--integration_methods. Everything after the merge is built on the scVI latent space; a selection that omitsscviproduces no output and no error. - Per-sample QC overrides in the samplesheet are ignored. The per-sample
min_genes,max_mito_percentage,automatic_cell_filteringand related columns are read into the sample metadata but the global--min_*/--max_*parameters are what actually reach the filtering step. Set filters globally. --prep_cellxgeneis no longer supported. Leave it at its default.-profile test_offlineis no longer supported. Use-profile testinstead.- Set
--speciesexplicitly. It defaults tohuman, and mouse data analysed under the human default produces wrong enrichment and cell–cell communication results without any error. See Running the pipeline on non-human species. - The upstream single-pass workflow has been removed. Always pass
-entry; without it the pipeline runsqc_clustering(see the note above). --unify_gene_symbolsis no longer supported. HUGO-based gene symbol unification only applies to human data and is not reliable enough to recommend. Gene symbols are still harmonised across samples without it.- Several other inherited parameters are also not currently supported:
--skip_enrichment,--skip_liana,--skip_rankgenesgroups,--pseudobulk*,--cluster_per_label,--cluster_global, and theexclude_samples_col/exclude_samples_valuescolumns of the contrasts file. - MultiQC coverage is partial — the Quarto reports are the more complete view of a run.
- Usage — samplesheet format, all parameters, per-entry-point guidance
- Output — the files each stage produces and how to read them
- UK DRI Informatics wiki — running the pipeline on the UK DRI cluster
Parameters are defined in nextflow_schema.json; nextflow run … --help
lists them.
This pipeline is derived from nf-core/scdownstream, originally written by
Nico Trummer, and would not exist without the work of its authors and
contributors. Full attribution — original authors, upstream contributors, the nf-core framework and
the pipelines this work builds on — is in ACKNOWLEDGEMENTS.md.
The fork is maintained by UK DRI Informatics.
Contributions are welcome via pull request against the dev_ukdri branch of
UKDRI/scdownstream. For help, open an issue in this
repository or contact UK DRI Informatics.
An extensive list of references for the tools used by the pipeline can be found in
CITATIONS.md.
This pipeline is no longer part of nf-core, but it was built on the nf-core framework and template, which you may wish to cite:
The nf-core framework for community-curated bioinformatics pipelines.
Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.
Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x.