Metadata-Version: 2.4
Name: dcatoolkit
Version: 0.4.0
Summary: Collection of useful modules and representations for managing DCA output data.
Author-email: Raheel Syed Ahmed <raheelsyedahmed@gmail.com>
Maintainer-email: Raheel Syed Ahmed <raheelsyedahmed@gmail.com>
License: MIT License
        
        Copyright (c) 2024 Raheel Syed Ahmed
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Homepage, https://github.com/RaheelSyedAhmed/dcatoolkit
Project-URL: Changelog, https://github.com/RaheelSyedAhmed/dcatoolkit/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/RaheelSyedAhmed/dcatoolkit/issues
Keywords: dca,toolkit,DI,coevolution
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: biotite>=1.7
Requires-Dist: numpy>=1.26.0
Requires-Dist: pandas>=2.1.0
Requires-Dist: scipy>=1.11.0
Provides-Extra: tests
Requires-Dist: pytest; extra == "tests"
Provides-Extra: docs
Requires-Dist: sphinx; extra == "docs"
Requires-Dist: pdoc; extra == "docs"
Requires-Dist: numpydoc; extra == "docs"
Provides-Extra: lint
Requires-Dist: ruff; extra == "lint"
Provides-Extra: plot
Requires-Dist: matplotlib>=3.10.9; extra == "plot"
Dynamic: license-file

# dcatoolkit
 Collection of useful modules and representations for managing DCA output data.

## Installation

```bash
pip install dcatoolkit

# optional: adds matplotlib for the plotting example"
pip install "dcatoolkit[plot]"  
```

Requires Python 3.12+.
Upgrading from 0.2.x? See the [changelog](https://github.com/RaheelSyedAhmed/dcatoolkit/blob/main/CHANGELOG.md).

## Major Sections
### Representations
  * Use Pairs to load lists, tuples, sets, and ndarrays with the correct orientation of elements. This will allow you to store integer pairs, in the form of structured arrays with fields `residue1` and `residue2` that can be mirrored (where y becomes x and vice versa) and subset.
  * Use DirectInformationData to create structured ndarrays with `residue1`, `residue2`, and `DI` fields that can be sorted by `DI`, mapped to a protein with a ResidueAlignment, and used to generate output for other programs (including UCSF Chimera)
  * Use ResidueAlignment to generate a reference map. Indices of one sequence of characters can be linked to their corresponding indices of the other sequence of characters. The dictionaries produced, domain-to-protein and protein-to-domain, allow for forward mapping and backmapping.
  * Use StructureInformation to find contacts in a protein structure and find atomic information related to specific pairs of interest. It can read in PDBx/mmCIF and PDB files or fetch them from RCSB and find contacts between residues.
### Analytics
  * Use MSATools to load in Multiple Sequence Alignment (MSA) data and provide functionality including generating frequency statistics on "gappiness" in the MSA and filtering and cleaning MSAs.


## Quick Start
```python
from dcatoolkit import DirectInformationData, ResidueAlignment, StructureInformation

# Rank DI pairs and map them from MSA (domain) numbering onto the protein sequence.
di = DirectInformationData.load_from_DI_file("my_dca_output.DI")
alignment = ResidueAlignment.load_from_align_file("my_domain.align")
top_pairs = di.get_ranked_mapped_pairs(alignment, alignment, number=50)

# Compare the top pairs with residue contacts in a structure.
structure = StructureInformation.fetch_pdb("2KLL")
contacts = structure.get_contacts(ca_only=False, threshold=8, chain1="A", chain2="A")
hits = [pair for pair in top_pairs.tolist() if pair in contacts]
```

## Diagram of Hidden Markov Model & Direct Coupling Analysis Pipeline
<p align="center">
  <img src="https://github.com/user-attachments/assets/4768e08f-d513-4dbf-abc5-c80c1b3d42aa"/>
</p>

## Development

This project uses [uv](https://docs.astral.sh/uv/) for dependency management.

```bash
git clone https://github.com/RaheelSyedAhmed/dcatoolkit.git
cd dcatoolkit
uv sync --all-extras
```

Run the test suite (requires internet to fetch from RCSB):

```bash
uv run pytest

uv sync --all-extras # Installs packages from tests, docs, lint, and plot.
uv run ruff check # linting
```
