HPC operation

RustScenic computation is scheduler-neutral. Allocate one host, set one Rayon pool for Rust work, and pin every BLAS/OpenMP pool to one thread. LSF launchers under validation/hpc/minerva/ provide Minerva-specific wrapping and evidence collection; the Python and Rust APIs do not depend on LSF or personal paths.

Thread contract

export RAYON_NUM_THREADS="${LSB_DJOB_NUMPROC:-8}"
export OMP_NUM_THREADS=1
export OPENBLAS_NUM_THREADS=1
export MKL_NUM_THREADS=1
export PYTHONNOUSERSITE=1

Request span[hosts=1]. RustScenic is a shared-memory process, not MPI, so a multi-host allocation wastes hosts and does not increase throughput. In GRN thread-scaling jobs, child points may use fewer Rayon threads than the job's allocated ceiling. Nested BLAS/OpenMP pools must remain at one.

For Gibbs topics, pass topics_n_threads no larger than the allocation and record it with the seed. Parallel AD-LDA is reproducible at a fixed thread count, but changing the count changes the partition and can change the fitted posterior mode. Hold it fixed for biological comparisons.

On the audited 10x human-brain ATAC matrix (5,000 cells x 130,127 peaks, K=30, 100 sweeps), eight threads were 1.81x faster than four; sixteen threads did not improve on eight and reached a different posterior. Start Gibbs topics at eight threads on comparable data and scale only with a dataset-specific check. GRN on real PBMC3k retained 88% parallel efficiency at sixteen threads, so sixteen is the measured starting point for GRN-heavy RNA runs.

Rebuild immediately before a preflight when testing local source. Package and extension version strings alone cannot distinguish a stale same-version shared library:

python -m maturin develop --release
python -m rustscenic doctor --pretty

Memory model

  • Sparse RNA and ATAC matrices stay sparse across the production kernels.
  • GRN memory is bounded by the binned expression matrix and per-Rayon-worker scratch buffers. More Rayon threads increase scratch memory.
  • Parallel Gibbs adds approximately topics_n_threads × n_topics × n_peaks × 4 bytes for thread-local topic-word deltas, in addition to the sparse matrix, token assignments and outputs.
  • Correlation dichotomisation operates on requested TF-target pairs without a dense edge-by-cell matrix. Sparse CSC inputs use column intersections.
  • Motif ranking parquet/feather inputs are projected to the required features; do not eagerly load atlas-scale ranking databases.
  • Fragment preprocessing should receive cell-called barcodes. Carrying raw observed/empty droplets can turn a nominal 10k experiment into hundreds of thousands of matrix rows.
  • The integrated rna_with_regulons.h5ad write can require another materialised output copy. SKIP_INTEGRATED_ADATA=1 is appropriate for compute profiling, but not for a complete end-to-end production artefact.

Start with measured pilot memory and leave headroom for the input matrices, external ranking projection and final writes. The audited 44,222-cell cortex report used 16.8 GB peak RSS, but its input shapes and stages are not a universal memory formula.

The 0.5.0 IFB tests reached 1.2 million cells for a fixed 300-gene/30-TF GRN at 4.2 GB process peak RSS, and the synthetic seven-stage 200k-cell run used a 42.31 GB in-process high-water mark. Neither number predicts a real 1.2-million-cell full-gene multiome. For a new atlas, run the exact input at 100k, 200k, 400k and 800k cells and attempt all cells only after the measured memory curve leaves safe headroom for motif data and final writes.

Scientific configuration

  • early_stop_mode="arboreto" is the default. It uses the strict mean of the exact trailing OOB-improvement window and retains the stopping tree.
  • early_stop_mode="legacy_inbag" preserves the historical RustScenic two-point in-bag rule. Use it only for an explicitly labelled legacy rerun.
  • early_stop_window=0 disables stopping. This can fit all 5,000 trees for every target and materially increase runtime.
  • subsample=1.0 leaves no OOB rows, so arboreto mode also fits to the estimator ceiling. Keep the arboreto default subsample=0.9 unless that is an intentional full-sample experiment.
  • grn_regulon_polarities="both" is the production pipeline default. It emits separate <TF>_activator and <TF>_repressor programmes using Pearson rho_threshold=0.03; neutral edges are excluded.
  • "activating" keeps only activators. "unsigned" is the documented legacy migration path and deliberately skips dichotomisation.

The Minerva launchers default to GRN_N_ESTIMATORS=100 for smoke/scaling turnaround. Those runs must not be presented as the 5,000-estimator production configuration. Use the following exact production-validation sequence after the branch has been committed on Minerva, because --require-clean correctly rejects an uncommitted validation tree:

cd /sc/arion/projects/DiseaseGeneCell/Huang_lab_projects/rustscenic/repo
. /sc/arion/work/kahrae01/rustscenic/envs/rustscenic-v047/bin/activate
python -m maturin develop --release
export GRN_N_ESTIMATORS=5000
export SKIP_INTEGRATED_ADATA=0
bsub < validation/hpc/minerva/run_real_pbmc3k_full_pipeline.lsf
bsub < validation/hpc/minerva/run_real_pbmc3k_grn_scaling.lsf

For a cheaper harness exercise, leave GRN_N_ESTIMATORS unset. Every generated benchmark JSON declares path_policy="portable"; repo paths are relative and external inputs are represented by basenames plus hashes, so archived artefacts contain no workstation or cluster absolute paths.