Tutorials

This section provides step-by-step walkthroughs for common single-cell RNA sequencing (scRNA-seq) analysis workflows using SiCell.jl.


1. Basic Analysis Workflow

This tutorial demonstrates the standard analysis pipeline from raw count matrices to clustered cell populations and biological interpretation.

๐Ÿ“‚ Loading Data

SiCell.jl supports commonly used single-cell data formats.

using SiCell

# Load 10x Genomics output
obj = read_10x("data/filtered_feature_bc_matrix/")

# Load an AnnData (.h5ad) file
# obj = read_h5ad("data/dataset.h5ad")

๐Ÿ”ฌ Quality Control

Calculate quality-control metrics and remove low-quality cells.

calculate_qc_metrics!(obj)

filter_cells!(
    obj;
    min_genes=200,
    max_genes=5000,
    max_mito=10.0
)

โš–๏ธ Normalization and Feature Selection

Normalize expression values and identify highly variable genes.

# Library-size normalization followed by log transformation
normalize_data!(obj)

# Identify highly variable genes
find_variable_features!(
    obj;
    n_features=2000
)

# Optional: scale gene expression values
scale_data!(obj)

๐Ÿ“Š Dimensionality Reduction

Reduce the dimensionality of the dataset and construct the neighborhood graph.

# Principal component analysis
run_pca!(obj)

# Construct K-nearest-neighbor graph
find_neighbors!(
    obj;
    k=20,
    dims=30
)

# Compute UMAP embedding
run_umap!(obj)

๐Ÿงฉ Clustering

Identify cellular populations using graph-based community detection.

run_clustering!(obj, k=12)
run_graph_clustering!(obj, method="label_propagation", key="graph_cluster")
run_graph_clustering!(obj, method="louvain", key="graph_cluster")

๐ŸŽจ Visualization

Visualize cell clusters and gene expression patterns.

# Visualize clusters
dim_plot(
    obj,
    reduction="umap",
    group="graph_cluster"
)

# Visualize expression of a marker gene
feature_plot(
    obj,
    "CD3D"
)

2. ๐Ÿ”— Batch Integration with Harmony

For datasets containing multiple samples, patients, or experimental batches, SiCell.jl provides Harmony-based batch correction and integration.

# Merge datasets
merged_obj = merge_batches!(
    obj1,
    obj2;
    batch1_name="Patient1",
    batch2_name="Patient2"
)

# Standard preprocessing
calculate_qc_metrics!(merged_obj)
normalize_data!(merged_obj)
find_variable_features!(merged_obj)
scale_data!(merged_obj)
run_pca!(merged_obj)

# Harmony correction
run_harmony!(
    merged_obj,
    batch_key="batch"
)

# Construct graph using Harmony embedding
find_neighbors!(
    merged_obj,
    reduction="harmony"
)

# Visualization and  etc
run_umap!(
    merged_obj,
    reduction="harmony"
)


3. ๐ŸŒŠ Trajectory Inference

SiCell.jl provides diffusion-based pseudotime analysis for studying continuous biological processes such as differentiation, cellular activation, and disease progression.

Diffusion Maps

Compute a low-dimensional manifold representation suitable for trajectory analysis.

run_diffusion_map!(
    obj;
    n_components=10
)

โฑ๏ธ Pseudotime

Select a biologically meaningful root cell or population and compute pseudotime.

# Example using a root cell index
root_cell = 1

run_pseudotime!(
    obj,
    root_cell;
    method="graph"
)

Pseudotime values are stored in:

obj.meta_data.pseudotime

๐ŸŽจ Visualize Pseudotime

dim_plot(
    obj,
    reduction="umap",
    group="pseudotime"
)

4. ๐ŸŒฑ Trajectory Uncertainty Framework (TUF)

Traditional pseudotime methods assign cells a position along a trajectory but do not quantify the local uncertainty associated with that trajectory.

The Trajectory Uncertainty Framework (TUF) introduces two complementary metrics:

  • Temporal Entropy Score (TES): quantifies temporal inconsistency among neighboring cells.
  • Trajectory Divergence Score (TDS): quantifies disagreement among forward developmental directions.

Together, TES and TDS provide complementary views of local trajectory uncertainty.

๐Ÿงฎ Compute TES and TDS

After constructing a neighborhood graph and computing pseudotime:

trajectory_uncertainty!(
    obj;
    pseudotime_key="pseudotime",
    reduction="diffusion"
)

The computed scores are stored in:

obj.meta_data.traj_unc_tes
obj.meta_data.traj_unc_tds

๐ŸŽจ Visualize TES and TDS

    trajectory_uncertainty_plot(obj; reduction="umap", key_prefix="traj_unc")

Interpreting TUF

TESTDSInterpretation
LowLowStable and well-defined cellular trajectory
HighLowTemporal mixing or heterogeneous developmental states
LowHighDirectional divergence and potential branching
HighHighComplex transitions involving both temporal and directional ambiguity

5. ๐Ÿงฌ Differential Expression Analysis

Identify genes associated with specific cellular populations or states.

find_markers!(
    obj;
    group="graph_cluster"
)

The output includes:

  • Differentially expressed genes
  • Effect sizes
  • Statistical significance values
  • Multiple-testing corrected p-values

6. ๐Ÿท๏ธ Automated Cell Type Annotation

SiCell.jl supports marker-based cell type annotation using Cell marker 2.

db = load_cellmarker(species="Hs")

annotate_clusters!(obj, markers, db,
    species="Hs",
    min_score=0.03)

๐Ÿš€ What's Next?

After completing the basic workflow, you can:

  • Perform differential expression analysis to identify marker genes.
  • Annotate cellular populations using PanglaoDB.
  • Integrate multi-sample datasets with Harmony.
  • Study developmental and disease trajectories using pseudotime.
  • Identify transitional and branching states using TES and TDS.
  • Generate high quality visualizations.

For more advanced examples, consult the full API documentation and case-study notebooks.