Metadata-Version: 2.4
Name: vantami
Version: 0.4.4
Summary: A collection of ML/AI tools for chemistry applications.
Home-page: https://github.com/MateuszIwan/VantaMI
Author: Mateusz Iwan
Author-email: mateusz.iwan@hotmail.com
Project-URL: Source, https://github.com/M-Iwan/VantaMI
Project-URL: Bug Tracker, https://github.com/M-Iwan/VantaMI/issues
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Chemistry
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.13
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=2.4
Requires-Dist: pandas>=3.0
Requires-Dist: polars>=1.40
Requires-Dist: matplotlib>=3.10
Requires-Dist: seaborn>=0.13
Requires-Dist: scipy>=1.17
Provides-Extra: full
Requires-Dist: scikit-learn>=1.8; extra == "full"
Requires-Dist: rdkit>=2026.03; extra == "full"
Requires-Dist: transformers>=5.5; extra == "full"
Dynamic: author
Dynamic: author-email
Dynamic: classifier
Dynamic: description
Dynamic: description-content-type
Dynamic: home-page
Dynamic: license-file
Dynamic: project-url
Dynamic: provides-extra
Dynamic: requires-dist
Dynamic: requires-python
Dynamic: summary

![](files/Vantami.png)

### What is this repository?
A collection of ML/AI code for chemistry applications developed during my PhD.

If you find the repo or its parts useful, it can be installed using:

`pip install vantami`

Full install:

`pip install "vantami[full]"`

The usual versioning conventions are followed loosely, with minor version bumps usually
corresponding to a substantial update to a specific module.

### Repository Structure
Last updated on version: 0.4.4

```
vantami/
├── deprecated/                    Old code
├── dev/                           New stuff / helper files
├── dist/                          PyPI distribution files
├── environments/                  Environments for special-need code (CDDD, Mordred) 
├── vantami/
│   ├── cache.py                   Functions for caching required data
│   ├── api/
│   │   ├── convert.py             Conversion from IUPAC names to SMILES using OPSIN
│   │   └── resolve.py             Get SMILES from name using PubChem, CIR, or (WIP) CAS
│   ├── chemistry/
│   │   └── molecule.py            Standardize molecule, embed it in 3D, and dock using SMINA
│   ├── cli/
│   │   ├── CDDD.py                CLI wrapper for CDDD descriptors
│   │   ├── CIR.py                 CLI wrapper of CIR name resolver
│   │   ├── Dock.py                Wrapper for molecule.py
│   │   ├── MolStandardizer.py     Molecule standardization using RDKit
│   │   ├── Mordred.py             CLI wrapper for Mordred descriptors
│   │   └── OptunaOptimization.py  Old code for Optuna optimization.
│   ├── data/
│   │   ├── cluster.py             Clustering using Butina/Murcko/Connected Components algorithms
│   │   ├── descriptors.py         Descriptor calculations: ECFP, MACCS, Klek, CDDD, RDKit, Mordred, ChemBERTa, MAPC
│   │   ├── distance.py            Parallel distance matrix / k-neighbors / group k-neighbors calculations
│   │   ├── filter.py              Filter outliers based on molecular parameters
│   │   ├── manager.py             Main class for managing data during training and inference
│   │   ├── partition.py           Partitioning algoriths; convenience wrappers around scikit-learn and cluster.py
│   │   └── transform.py           Main class for normalizing/processing data before training
│   ├── deep/
│   │   ├── dataset.py             MMDataset / MMBatch for multi-modal Polars-backed samples
│   │   ├── loader.py              MMLoader (collate → MMBatch)
│   │   ├── models.py              MMTUnit base class and concrete units (e.g. TestModel)
│   │   ├── modules.py             GNN/CNN/RNN/linear builders used by units
│   │   ├── utils.py               Activations and small helpers
│   │   └── vectorizer.py          GraphVectorizer, StringVectorizer (legacy MMGV → deprecated.deep.mmgv)
│   ├── io/
│   │   ├── database.py            Preprocessing of ChEMBL and BindingDB files (largely useless)
│   │   └── file.py                IO functions for several formats I'm using; works with Pandas/Polars DFs
│   ├── metrics/
│   │   └── modellability.py       MODI index
│   ├── ml/
│   │   ├── augood.py              AU-GOOD framework for model's performance evaluation
│   │   ├── evaluate.py            Code for simple unit/ensemble evaluations
│   │   ├── models.py              Self-contained, sklearn-compatibile models and ensembles
│   │   ├── optimize.py            Hyperparameter optimization; functions at the top of the file are outdated
│   │   ├── params.py              Pre-defined parameters for Optuna
│   │   ├── score.py               Functions for scoring models; Outdated, now included with Unit and Ensemble classes
│   │   ├── select.py              Sequential feature selection; Outdated
│   │   └── utils.p                Helper functions for building Units from just names of models and descriptors
│   ├── nlp/  
│   │   ├── article.py             Article class for retrieving metadata based on DOI/Names
│   │   ├── cluster.py             Latent Dirichlet Allocation for abstract-based clustering
│   │   └── tokenize.py            Word and document tokenizers
│   ├── standardize/   
│   │   ├── clean.py               Wrappers around RDKit functions for standaradizing SMILES
│   │   ├── duplicates.py          Duplicate processing based on Median Absolute Deviation
│   │   ├── filter.py              Filter based on selected features
│   │   └── validate.py            Structure validation and error finding
│   ├── stats/
│   │   ├── lmm.py                 Linear Mixed Models-realted code
│   │   ├── tests.py               Statistical tests
│   │   └── utils.py               Helper functions
│   └── visualize/
│       ├── ecdf.py                Emprical Cumulative Distribution Function of molecular inter-distance 
│       ├── embedding.py           t-SNE and UMAP
│       ├── performance.py         WIP: AU-GOOD framework-related plots
│       ├── predictions.py         Bunch of plots for assessing model performance; currently only Regression
│       ├── properties.py          Plot and compare molecular properties between datasets
│       └── utils.py               Helper functions and my custom palette
├── vantami-r/
│   └── R/
│       └── lmm.R                  Linear Mixed Models code
├── projects/
│   ├── cddd_setup                 Files for setting up CDDD environment anywhere
│   ├── osmordred_setup            WIP: Corrected Mordred descriptors
│   └── qcg_template               Template for QCG PilotJob training on bigger scale (one node)
├── temp/           
├── tests/          
├── .gitignore   
├── CHANGELOG.md 
├── LICENSE   
├── pyproject.toml  
├── README.md       
└── setup.py        
```
