PlasRisk: A Multi-Dimension Data-Driven Weighted Risk Assessment Framework for Bacterial Plasmids (A) Data-driven weight determination · (B) Command-line risk scoring · (C) Dimensionality analysis & lite mode A. Data-driven weight determination PIPdb database 792,964 PSCs Train 634,359 / Test 158,605 Component extraction 10 risk dimensions (0–1) ARG · VF · MOB · BM · SIZE … Training labels (4 outcomes) ● High-risk ARG carriage (3.1%) ● MDR–virulence fusion (1.8%) ● Conjugative mobility (5.1%) ● Biocide/metal resistance (34.0%) Weight inference methods ① Random Forest MDG (4 outcomes) ② LASSO logistic regression ③ Entropy weight method ④ Grid-search AUC optimization + LORO cross-validation (40 replicons) Consensus final weights RF-MDG + LASSO + Entropy + Grid Σw = 1.000 · top-4 = 84.1% Final weights: S_ARG 0.237 S_BM 0.222 S_MOB 0.189 S_SIZE 0.164 S_VF 0.126 S_HOST 0.030 S_HAB 0.019 S_GEO 0.005 S_REP 0.004 S_GROW 0.002 5-fold CV validation (4 outcomes) ● High-risk ARG: AUC = 0.958 ● MDR–VF fusion: AUC = 0.964 ● Conjugative: AUC = 0.856 ● Biocide/metal: AUC = 0.903 ● Mean AUC = 0.919 · LORO-CV = 0.962 ● ±30% perturbation: Spearman ρ = 0.995 Composite risk score S = Σ wᵢ · Sᵢ / Σwᵢ S ∈ [0, 1] · Σwᵢ = 1.000 Consensus of RF-MDG + LASSO + Entropy + Grid Two modes: --mode full (10-dim) | --mode lite (5-dim) weights feed into tool B. PlasRisk command-line tool (plasrisk *.fasta) Input Plasmid FASTA *.fasta / *.fa / *.fna single / batch / directory Annotation (abricate) CARD → ARGs (n, WHO, high-risk) VFDB → VFs (exotoxin, T3SS/T4SS) PlasmidFinder → replicon type BacMet → biocide/metal genes Sequence analysis Length (sigmoid S_SIZE) T4CP · relaxase · oriT · MPF (S_MOB) Component scores (each 0–1) S_ARGARG burden + WHO/high-risk S_BMBiocide/metal co-selection S_MOBConjugation potential S_SIZECargo capacity (sigmoid) S_VFVirulence burden S_HOSTHost range S_GROWTemporal growth S_REPReplicon prior S_HABHabitat breadth S_GEOGeographic spread Lite mode (--mode lite): S_ARG + S_BM + S_MOB + S_SIZE + S_VF Renormalized weights: 0.253 / 0.237 / 0.201 / 0.175 / 0.135 Equivalent mean AUC to full model (0.919 vs 0.919) Weighted sum S = Σ wᵢ·Sᵢ Risk grade A ≥ 0.60 Very High B ≥ 0.45 High C ≥ 0.30 Moderate ■ D ≥ 0.15 ■ E < 0.15 Output TSV / JSON component scores S_total, S_norm grade (A–E) replicon, mobility high-risk genes model_mode (full / lite) $ plasrisk *.fasta -o results/ $ plasrisk --mode lite *.fasta # --no-abricate · --json · --mode {full,lite} Replicon prior lookup: 66 typed replicons with pre-computed S_REP / S_GEO / S_HAB / S_GROW / S_HOST Unknown replicons → neutral defaults (0.30); prefix matching handles subtypes (IncFII(K) → IncFII). dimensionality optimization C. Dimensionality analysis & PlasRisk-lite (all-subsets, 1,023 models, 5-fold CV) All-subsets enumeration 2¹⁰ − 1 = 1,023 subsets Weights renormalized per subset AUC on 4 outcomes, 5-fold CV Stepwise selection Forward: ARG→SIZE→MOB→BM→VF… Backward: drop HOST→GEO→HAB… Removing BM causes largest drop Key results ● Plateau at k=5: mean AUC = 0.920 (same as full) ● 5-dim lite (ARG+VF+MOB+SIZE+BM): AUC = 0.920 ● No overfitting: train-test gap <0.003 DeLong tests (5-dim lite vs 10-dim full) ● HR-ARG: ΔAUC = 0.000 (p = NS) ● Fusion: ΔAUC = −0.001 (p = NS) ● All outcomes equivalent → 5-dim sufficient Mean CV AUC vs. number of dimensions (plateau at k=5): 0.93 0.91 0.89 0.87 0.85 0.77 0.768 0.862 0.912 0.918 0.920 0.920 0.920 0.920 0.920 0.920 1 2 3 4 5 6 7 8 9 10 Plateau (k=5) Number of dimensions Full (10-dim) vs. Lite (5-dim) comparison Metric Full (10-dim) Lite (5-dim) Δ High-risk ARG AUC 0.958 0.958 0.000 MDR–VF fusion AUC 0.964 0.967 −0.001 Conjugative AUC 0.856 0.856 0.000 Biocide/metal AUC 0.903 0.904 +0.001 Mean AUC (4 outcomes) 0.919 0.919 equivalent Required annotations ARG+VF+MOB+REP+BM+meta ARG+VF+MOB+SIZE+BM Why retain 10-dim if 5-dim is equivalent? → Pareto optimality: 10-dim is non-dominated across all 4 outcomes; the 5 context dimensions (HOST, REP, GEO, HAB, GROW) add epidemiological value for One Health surveillance even though they contribute <0.1% to mean AUC. → No overfitting: train-test gap <0.003, bootstrap optimism <0.005. → Both modes offered: full for surveillance, lite for rapid screening. PlasRisk-lite: rapid screening mode (5-dim core) ● 5 dimensions: S_ARG (0.253) + S_BM (0.237) + S_MOB (0.201) + S_SIZE (0.175) + S_VF (0.135) ● Requires ARG + VFDB + mobility genes + length + BacMet (no metadata priors) ● Mean AUC = 0.919, equivalent to full 10-dim model; S_VF raises fusion AUC to 0.967 ● CLI: plasrisk --mode lite *.fasta ● Python API: PlasRiskScorer(mode="lite") ● Ideal for resource-limited settings, large-scale screening, One Health surveillance