Skip to content
 
 

Latest commit

 

History

1,081 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

UK DRI scdownstream

A UK DRI fork of nf-core/scdownstream.

Important

This pipeline began life as a fork of nf-core/scdownstream and has since diverged substantially. It is now developed independently by UK DRI Informatics and follows its own release and configuration conventions.

Please raise issues and questions in this repository. For the original, nf-core community-supported pipeline, see nf-co.re/scdownstream.

The original authors and contributors are credited in ACKNOWLEDGEMENTS.md.

Nextflow run with apptainer License: MIT

Introduction

UK DRI scdownstream is a Nextflow pipeline for the downstream analysis of processed single-cell RNA-seq data. It takes per-sample count matrices (h5ad, 10x h5, SingleCellExperiment / Seurat RDS, or CSV) and carries them through quality control, ambient RNA correction, doublet detection, cell type annotation, scVI integration, dimensionality reduction and clustering, then on through marker genes, gene set enrichment and cell–cell communication, and finally pseudobulk differential expression.

The input is ideally the output of nf-core/scrnaseq or a similar per-sample count matrix object files. We recommend supplying it as AnnData (.h5ad) objects: AnnData is part of the scverse ecosystem that this pipeline is built on, and it scales to large count matrices without the size limits that SingleCellExperiment and Seurat objects run into in R.

The pipeline is organised as three sequential stages, each its own Nextflow entry point, so that long-running analyses can be checkpointed and re-run independently.

It inherits the learnings and implementations of the following pipelines (alphabetical), via nf-core/scdownstream:

Entry points

The pipeline is always invoked with an explicit -entry. The three stages are sequential: each consumes the previous stage's finalized .h5ad via --base_adata.

Stage Entry point Required input Main output
1 -entry qc_clustering --input samplesheet.csv <name>_qc_clustering.h5ad, MultiQC report, QC/clustering HTML report
2 -entry downstream --base_adata <stage 1 h5ad> <name>_downstream.h5ad, <name>_downstream_markers.json.gz, analysis HTML report
3 -entry differential_genes --base_adata <stage 2 h5ad> + --diffgenes_contrasts contrasts.tsv per-contrast DE tables, pseudobulk h5ad, DE HTML report
samplesheet.csv
      │
      ▼  -entry qc_clustering
<name>_qc_clustering.h5ad  ── QC, integration (scVI), UMAP, Leiden clustering
      │
      ▼  -entry downstream        --base_adata
<name>_downstream.h5ad  ── marker genes, enrichment, LIANA+ cell–cell communication
      │
      ▼  -entry differential_genes  --base_adata  --diffgenes_contrasts
differential_genes/    ── decoupler pseudobulk → PyDESeq2 per group × contrast

Important

Always pass -entry, so it is explicit which stage runs and on what input. Without it the pipeline falls back to stage 1 (qc_clustering) and prints a warning; it never runs downstream or differential_genes on its own. The upstream single-pass workflow has been removed.

Pipeline steps

Stage 1 — qc_clustering

  1. Load and convert inputs to h5ad (h5ad, 10x h5, RDS, CSV), add a sample column to obs, and optionally per-sample metadata from a TSV (--metadata)
  2. Per-sample quality control
    1. QC metrics for raw counts (MultiQC)
    2. Doublet detection — scrublet (doublets are annotated, not removed unless --doublet_removal is set)
    3. Ambient RNA correction — decontX (default), soupX, CellBender, scAR
    4. Cell filtering — fixed thresholds, or automatic N-MAD outlier thresholds via --automatic_cell_filtering; with --filtering_keep_outliers the outliers are kept and marked in obs["outlier"] instead of removed
  3. Cell type annotation — CellTypist and/or SingleR with celldex references
  4. Harmonise gene symbol, batch and cell type columns across samples, resolving duplicate gene symbols
  5. Merge and integrate — scVI
  6. Embeddings and clustering — log-normalisation, HVG selection, PCA, neighbours, UMAP, Leiden at every resolution in --clustering_resolutions
  7. Reports — a Quarto QC/clustering report and a MultiQC report

--qc_only stops after step 3, producing per-sample QC objects without merging or integration.

Stage 2 — downstream

  1. Marker genes per cluster — scanpy.tl.rank_genes_groups (Wilcoxon) on the clustering named by --selected_clustering
  2. Gene set enrichment over those markers
  3. Cell–cell communication — LIANA+ rank aggregation, using HCOP orthologs for non-human species
  4. Marker export to <name>_downstream_markers.json.gz, and h5ad/RDS finalisation
  5. A Quarto analysis report

Stage 3 — differential_genes

  1. Pseudobulk aggregation per sample × group label — decoupler (counts layer, sum mode)
  2. Drop pseudobulk samples below --diffgenes_min_counts / --diffgenes_min_cells
  3. Split into one object per group label (cluster or cell type, --diffgenes_group_col)
  4. Differential expression per group label × contrast — PyDESeq2, design ~ variable [+ blocking…]
  5. A Quarto differential expression report across all results

Contrasts are supplied as a TSV following the nf-core/differentialabundance definition — see assets/contrasts.tsv and the usage documentation.

Tool choices

This fork focuses on a curated set of tools — the approaches we have validated on UK DRI data:

Step Supported
Doublet detection scrublet
Ambient RNA correction decontx (default), soupx, cellbender, scar, or none
Integration scvi
Clustering Leiden, at multiple resolutions
Cell type annotation CellTypist, SingleR / celldex
Differential expression PyDESeq2 on decoupler pseudobulk

Further tools will be added as they are curated and validated. Until then, please use the values above — see Supported tool choices for the details.

Prerequisites

You need:

  • Nextflow ≥ 24.10.5, which requires Java 17 or later.
  • Apptainer (or Singularity). This is the recommended way to run the pipeline. Docker also works, but see the Docker Hub pull-limit note under Container images.
  • An x86_64 Linux machine or cluster. The container images are published for linux/amd64 only.
  • Internet access from the machine running the tasks: images are pulled on first use, CellTypist models are downloaded at run time, and gene set enrichment queries the g:Profiler web service.

UK DRI users: see the UK DRI Informatics wiki for how to run the pipeline on the cluster.

Running the pipeline on non-human species

Cell–cell communication (LIANA+, in -entry downstream) uses a human ligand–receptor resource. Mapping it onto another species needs an HCOP ortholog table, which you download once:

  1. Set --species on every stage. Supported values are human (default) and mouse.

  2. Download the HCOP table for your species from the HGNC HCOP downloads. Use the fifteen-column file, and keep its original name, because LIANA+ looks for <directory>/human_<species>_hcop_fifteen_column.txt.gz:

    mkdir -p hcop
    curl -o hcop/human_mouse_hcop_fifteen_column.txt.gz \
        https://storage.googleapis.com/public-download-files/hcop/human_mouse_hcop_fifteen_column.txt.gz
  3. Pass the directory to -entry downstream with --ortholog_hcop_directory:

    nextflow run UKDRI/scdownstream -r dev_ukdri -entry downstream \
       -profile apptainer \
       --base_adata results/qc_clustering/my_study_qc_clustering.h5ad \
       --name my_study \
       --species mouse \
       --ortholog_hcop_directory hcop \
       --outdir results/downstream

Without it, -entry downstream stops with an error for a non-human --species. Human data does not need it. See Reference data for details.

Quick start

Note

Install the requirements listed under Prerequisites first. If you are unsure about the filtered / unfiltered distinction in the samplesheet, see Filtered and unfiltered matrices.

Prepare a samplesheet describing your per-sample matrices:

sample,unfiltered
sample1,/absolute/path/to/sample1.h5ad
sample2,/absolute/path/to/sample2.h5
sample3,relative/path/to/sample3.rds
sample4,/absolute/path/to/sample4.csv

Each entry is an h5ad, 10x h5, RDS or CSV file. RDS files may contain any object convertible to a SingleCellExperiment via Seurat's as.SingleCellExperiment. CSV files should hold a matrix with genes as columns and cells as rows, the first column being cell barcodes.

Stage 1 — QC, integration and clustering:

nextflow run UKDRI/scdownstream -r dev_ukdri -entry qc_clustering \
   -profile apptainer \
   --input samplesheet.csv \
   --metadata sample_metadata.tsv \
   --umap_color_by diagnosis,sex \
   --name my_study \
   --species human \
   --outdir results/qc_clustering

--metadata is optional: a tab-separated file with one row per sample, whose columns (donor, diagnosis, sex, ...) are added to every cell's obs. See Per-sample metadata. --umap_color_by plots such columns on the PCA and scVI UMAPs in the reports, i.e. before and after integration. See UMAP colourings.

Stage 2 — downstream analysis:

nextflow run UKDRI/scdownstream -r dev_ukdri -entry downstream \
   -profile apptainer \
   --base_adata results/qc_clustering/my_study_qc_clustering.h5ad \
   --name my_study \
   --species human \
   --outdir results/downstream

Stage 3 — differential expression:

nextflow run UKDRI/scdownstream -r dev_ukdri -entry differential_genes \
   -profile apptainer \
   --base_adata results/downstream/my_study_downstream.h5ad \
   --diffgenes_contrasts contrasts.tsv \
   --diffgenes_group_col cell_type \
   --diffgenes_sample_col sample \
   --name my_study \
   --outdir results/differential_genes

Warning

Provide pipeline parameters on the command line or via -params-file. Custom config files passed with -c can supply any Nextflow configuration except parameters.

Container images

Every process runs from a public container image; nothing has to be built by hand. Nextflow pulls each image on first use, converting it to a .sif in $NXF_SINGULARITY_CACHEDIR under the singularity or apptainer profile. Two images are specific to this fork; they are hosted on Docker Hub and pulled the same way. Their Dockerfiles are kept in the repository only as the recipe they were built from:

Image (Docker Hub) Required by Recipe
docker.io/nhecker/scanpy-report:1.11.4-coreinf0.4 SCANPY_GENERATE_REPORT, SCANPY_GENERATE_REPORT_QC, PYDESEQ2_GENERATE_REPORT, SCANPY_ENRICH modules/local/scanpy/report/Dockerfile
docker.io/nhecker/pydeseq2:0.1 DIFFERENTIAL_GENES_PER_CONTRAST modules/local/pydeseq2/differential_genes/Dockerfile

Both extend the public gcfntnu/scanpy:1.11.4 image, which DECOUPLER_PSEUDOBULK, FILTER_PSEUDOBULK and SCANPY_EXPORT_MARKERS use directly. The images are published for linux/amd64 only.

Important

We recommend running with -profile apptainer (or singularity), not -profile docker. These images are hosted on Docker Hub, which limits anonymous pulls on free accounts. Apptainer pulls each image once and reuses the .sif from the cache for every later run. Set NXF_SINGULARITY_CACHEDIR to a shared directory so every node and every run uses the same copy. Docker pulls through each machine's own daemon, so a multi-node or cloud run can repeat the same pull many times and hit the limit mid-run. If you must use Docker, run docker login first, or pre-pull the images on each host.

Changes and known limitations

This pipeline is work in development. The following are known and, for now, expected behaviours. Several of them silently affect results, so please read before interpreting output.

  1. Doublets are flagged, not removed, by default. scrublet writes its scores and predictions into the object. Set --doublet_removal to remove them right after detection, with --doublet_detection_threshold methods having to agree (default 1).
  2. Keep scvi in --integration_methods. Everything after the merge is built on the scVI latent space; a selection that omits scvi produces no output and no error.
  3. Per-sample QC overrides in the samplesheet are ignored. The per-sample min_genes, max_mito_percentage, automatic_cell_filtering and related columns are read into the sample metadata but the global --min_* / --max_* parameters are what actually reach the filtering step. Set filters globally.
  4. --prep_cellxgene is no longer supported. Leave it at its default.
  5. -profile test_offline is no longer supported. Use -profile test instead.
  6. Set --species explicitly. It defaults to human, and mouse data analysed under the human default produces wrong enrichment and cell–cell communication results without any error. See Running the pipeline on non-human species.
  7. The upstream single-pass workflow has been removed. Always pass -entry; without it the pipeline runs qc_clustering (see the note above).
  8. --unify_gene_symbols is no longer supported. HUGO-based gene symbol unification only applies to human data and is not reliable enough to recommend. Gene symbols are still harmonised across samples without it.
  9. Several other inherited parameters are also not currently supported: --skip_enrichment, --skip_liana, --skip_rankgenesgroups, --pseudobulk*, --cluster_per_label, --cluster_global, and the exclude_samples_col / exclude_samples_values columns of the contrasts file.
  10. MultiQC coverage is partial — the Quarto reports are the more complete view of a run.

Documentation

  • Usage — samplesheet format, all parameters, per-entry-point guidance
  • Output — the files each stage produces and how to read them
  • UK DRI Informatics wiki — running the pipeline on the UK DRI cluster

Parameters are defined in nextflow_schema.json; nextflow run … --help lists them.

Credits

This pipeline is derived from nf-core/scdownstream, originally written by Nico Trummer, and would not exist without the work of its authors and contributors. Full attribution — original authors, upstream contributors, the nf-core framework and the pipelines this work builds on — is in ACKNOWLEDGEMENTS.md.

The fork is maintained by UK DRI Informatics.

Contributions and support

Contributions are welcome via pull request against the dev_ukdri branch of UKDRI/scdownstream. For help, open an issue in this repository or contact UK DRI Informatics.

Citations

An extensive list of references for the tools used by the pipeline can be found in CITATIONS.md.

This pipeline is no longer part of nf-core, but it was built on the nf-core framework and template, which you may wish to cite:

The nf-core framework for community-curated bioinformatics pipelines.

Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.

Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x.

About

A single cell transcriptomics pipeline for QC, integration and making the data presentable

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages