PlasRisk: A Multi-Dimension Data-Driven Weighted Risk Assessment Framework for Bacterial Plasmids
(A) Data-driven weight determination · (B) Command-line risk scoring · (C) Dimensionality analysis & lite mode
A. Data-driven weight determination
PIPdb database
792,964 PSCs
Train 634,359 / Test 158,605
Component extraction
10 risk dimensions (0–1)
ARG · VF · MOB · BM · SIZE …
Training labels (4 outcomes)
● High-risk ARG carriage (3.1%)
● MDR–virulence fusion (1.8%)
● Conjugative mobility (5.1%)
● Biocide/metal resistance (34.0%)
Weight inference methods
① Random Forest MDG (4 outcomes)
② LASSO logistic regression
③ Entropy weight method
④ Grid-search AUC optimization
+ LORO cross-validation (40 replicons)
Consensus final weights
RF-MDG + LASSO + Entropy + Grid
Σw = 1.000 · top-4 = 84.1%
Final weights:
S_ARG 0.237
S_BM 0.222
S_MOB 0.189
S_SIZE 0.164
S_VF 0.126
S_HOST 0.030
S_HAB 0.019
S_GEO 0.005
S_REP 0.004
S_GROW 0.002
5-fold CV validation (4 outcomes)
● High-risk ARG: AUC = 0.958
● MDR–VF fusion: AUC = 0.964
● Conjugative: AUC = 0.856
● Biocide/metal: AUC = 0.903
● Mean AUC = 0.919 · LORO-CV = 0.962
● ±30% perturbation: Spearman ρ = 0.995
Composite risk score
S = Σ wᵢ · Sᵢ / Σwᵢ
S ∈ [0, 1] · Σwᵢ = 1.000
Consensus of RF-MDG + LASSO + Entropy + Grid
Two modes: --mode full (10-dim) | --mode lite (5-dim)
weights feed into tool
B. PlasRisk command-line tool (plasrisk *.fasta)
Input
Plasmid FASTA
*.fasta / *.fa / *.fna
single / batch / directory
Annotation (abricate)
CARD → ARGs (n, WHO, high-risk)
VFDB → VFs (exotoxin, T3SS/T4SS)
PlasmidFinder → replicon type
BacMet → biocide/metal genes
Sequence analysis
Length (sigmoid S_SIZE)
T4CP · relaxase · oriT · MPF (S_MOB)
Component scores (each 0–1)
S_ARGARG burden + WHO/high-risk
S_BMBiocide/metal co-selection
S_MOBConjugation potential
S_SIZECargo capacity (sigmoid)
S_VFVirulence burden
S_HOSTHost range
S_GROWTemporal growth
S_REPReplicon prior
S_HABHabitat breadth
S_GEOGeographic spread
Lite mode (--mode lite): S_ARG + S_BM + S_MOB + S_SIZE + S_VF
Renormalized weights: 0.253 / 0.237 / 0.201 / 0.175 / 0.135
Equivalent mean AUC to full model (0.919 vs 0.919)
Weighted sum
S = Σ wᵢ·Sᵢ
Risk grade
A ≥ 0.60 Very High
B ≥ 0.45 High
C ≥ 0.30 Moderate
■ D ≥ 0.15 ■ E < 0.15
Output
TSV / JSON
component scores
S_total, S_norm
grade (A–E)
replicon, mobility
high-risk genes
model_mode
(full / lite)
$ plasrisk *.fasta -o results/
$ plasrisk --mode lite *.fasta
# --no-abricate · --json · --mode {full,lite}
Replicon prior lookup: 66 typed replicons with pre-computed S_REP / S_GEO / S_HAB / S_GROW / S_HOST
Unknown replicons → neutral defaults (0.30); prefix matching handles subtypes (IncFII(K) → IncFII).
dimensionality optimization
C. Dimensionality analysis & PlasRisk-lite (all-subsets, 1,023 models, 5-fold CV)
All-subsets enumeration
2¹⁰ − 1 = 1,023 subsets
Weights renormalized per subset
AUC on 4 outcomes, 5-fold CV
Stepwise selection
Forward: ARG→SIZE→MOB→BM→VF…
Backward: drop HOST→GEO→HAB…
Removing BM causes largest drop
Key results
● Plateau at k=5: mean AUC = 0.920 (same as full)
● 5-dim lite (ARG+VF+MOB+SIZE+BM): AUC = 0.920
● No overfitting: train-test gap <0.003
DeLong tests (5-dim lite vs 10-dim full)
● HR-ARG: ΔAUC = 0.000 (p = NS)
● Fusion: ΔAUC = −0.001 (p = NS)
● All outcomes equivalent → 5-dim sufficient
Mean CV AUC vs. number of dimensions (plateau at k=5):
0.93
0.91
0.89
0.87
0.85
0.77
0.768
0.862
0.912
0.918
0.920
0.920
0.920
0.920
0.920
0.920
1
2
3
4
5
6
7
8
9
10
Plateau (k=5)
Number of dimensions
Full (10-dim) vs. Lite (5-dim) comparison
Metric
Full (10-dim)
Lite (5-dim)
Δ
High-risk ARG AUC
0.958
0.958
0.000
MDR–VF fusion AUC
0.964
0.967
−0.001
Conjugative AUC
0.856
0.856
0.000
Biocide/metal AUC
0.903
0.904
+0.001
Mean AUC (4 outcomes)
0.919
0.919
equivalent
Required annotations
ARG+VF+MOB+REP+BM+meta
ARG+VF+MOB+SIZE+BM
Why retain 10-dim if 5-dim is equivalent?
→ Pareto optimality: 10-dim is non-dominated across all 4 outcomes;
the 5 context dimensions (HOST, REP, GEO, HAB, GROW) add epidemiological value
for One Health surveillance even though they contribute <0.1% to mean AUC.
→ No overfitting: train-test gap <0.003, bootstrap optimism <0.005.
→ Both modes offered: full for surveillance, lite for rapid screening.
PlasRisk-lite: rapid screening mode (5-dim core)
● 5 dimensions: S_ARG (0.253) + S_BM (0.237) + S_MOB (0.201) + S_SIZE (0.175) + S_VF (0.135)
● Requires ARG + VFDB + mobility genes + length + BacMet (no metadata priors)
● Mean AUC = 0.919, equivalent to full 10-dim model; S_VF raises fusion AUC to 0.967
● CLI: plasrisk --mode lite *.fasta
● Python API: PlasRiskScorer(mode="lite")
● Ideal for resource-limited settings, large-scale screening, One Health surveillance