HPC operation¶
RustScenic computation is scheduler-neutral. Allocate one host, set one Rayon
pool for Rust work, and pin every BLAS/OpenMP pool to one thread. LSF launchers
under validation/hpc/minerva/ provide Minerva-specific wrapping and evidence
collection; the Python and Rust APIs do not depend on LSF or personal paths.
Thread contract¶
export RAYON_NUM_THREADS="${LSB_DJOB_NUMPROC:-8}"
export OMP_NUM_THREADS=1
export OPENBLAS_NUM_THREADS=1
export MKL_NUM_THREADS=1
export PYTHONNOUSERSITE=1
Request span[hosts=1]. RustScenic is a shared-memory process, not MPI, so a
multi-host allocation wastes hosts and does not increase throughput. In GRN
thread-scaling jobs, child points may use fewer Rayon threads than the job's
allocated ceiling. Nested BLAS/OpenMP pools must remain at one.
For Gibbs topics, pass topics_n_threads no larger than the allocation and
record it with the seed. Parallel AD-LDA is reproducible at a fixed thread
count, but changing the count changes the partition and can change the fitted
posterior mode. Hold it fixed for biological comparisons.
On the audited 10x human-brain ATAC matrix (5,000 cells x 130,127 peaks, K=30, 100 sweeps), eight threads were 1.81x faster than four; sixteen threads did not improve on eight and reached a different posterior. Start Gibbs topics at eight threads on comparable data and scale only with a dataset-specific check. GRN on real PBMC3k retained 88% parallel efficiency at sixteen threads, so sixteen is the measured starting point for GRN-heavy RNA runs.
Rebuild immediately before a preflight when testing local source. Package and extension version strings alone cannot distinguish a stale same-version shared library:
python -m maturin develop --release
python -m rustscenic doctor --pretty
Memory model¶
- Sparse RNA and ATAC matrices stay sparse across the production kernels.
- GRN memory is bounded by the binned expression matrix and per-Rayon-worker scratch buffers. More Rayon threads increase scratch memory.
- Parallel Gibbs adds approximately
topics_n_threads × n_topics × n_peaks × 4 bytesfor thread-local topic-word deltas, in addition to the sparse matrix, token assignments and outputs. - Correlation dichotomisation operates on requested TF-target pairs without a dense edge-by-cell matrix. Sparse CSC inputs use column intersections.
- Motif ranking parquet/feather inputs are projected to the required features; do not eagerly load atlas-scale ranking databases.
- Fragment preprocessing should receive cell-called barcodes. Carrying raw observed/empty droplets can turn a nominal 10k experiment into hundreds of thousands of matrix rows.
- The integrated
rna_with_regulons.h5adwrite can require another materialised output copy.SKIP_INTEGRATED_ADATA=1is appropriate for compute profiling, but not for a complete end-to-end production artefact.
Start with measured pilot memory and leave headroom for the input matrices, external ranking projection and final writes. The audited 44,222-cell cortex report used 16.8 GB peak RSS, but its input shapes and stages are not a universal memory formula.
The 0.5.0 IFB tests reached 1.2 million cells for a fixed 300-gene/30-TF GRN at 4.2 GB process peak RSS, and the synthetic seven-stage 200k-cell run used a 42.31 GB in-process high-water mark. Neither number predicts a real 1.2-million-cell full-gene multiome. For a new atlas, run the exact input at 100k, 200k, 400k and 800k cells and attempt all cells only after the measured memory curve leaves safe headroom for motif data and final writes.
Scientific configuration¶
early_stop_mode="arboreto"is the default. It uses the strict mean of the exact trailing OOB-improvement window and retains the stopping tree.early_stop_mode="legacy_inbag"preserves the historical RustScenic two-point in-bag rule. Use it only for an explicitly labelled legacy rerun.early_stop_window=0disables stopping. This can fit all 5,000 trees for every target and materially increase runtime.subsample=1.0leaves no OOB rows, soarboretomode also fits to the estimator ceiling. Keep the arboreto defaultsubsample=0.9unless that is an intentional full-sample experiment.grn_regulon_polarities="both"is the production pipeline default. It emits separate<TF>_activatorand<TF>_repressorprogrammes using Pearsonrho_threshold=0.03; neutral edges are excluded."activating"keeps only activators."unsigned"is the documented legacy migration path and deliberately skips dichotomisation.
The Minerva launchers default to GRN_N_ESTIMATORS=100 for smoke/scaling
turnaround. Those runs must not be presented as the 5,000-estimator production
configuration. Use the following exact production-validation sequence after
the branch has been committed on Minerva, because --require-clean correctly
rejects an uncommitted validation tree:
cd /sc/arion/projects/DiseaseGeneCell/Huang_lab_projects/rustscenic/repo
. /sc/arion/work/kahrae01/rustscenic/envs/rustscenic-v047/bin/activate
python -m maturin develop --release
export GRN_N_ESTIMATORS=5000
export SKIP_INTEGRATED_ADATA=0
bsub < validation/hpc/minerva/run_real_pbmc3k_full_pipeline.lsf
bsub < validation/hpc/minerva/run_real_pbmc3k_grn_scaling.lsf
For a cheaper harness exercise, leave GRN_N_ESTIMATORS unset. Every generated
benchmark JSON declares path_policy="portable"; repo paths are relative and
external inputs are represented by basenames plus hashes, so archived artefacts
contain no workstation or cluster absolute paths.