Metadata-Version: 2.4
Name: m_strat
Version: 0.0.2
Summary: A Directed Acyclic Graph Automated Machine Learning Tool using Entity-Component Systems concepts.
Author-email: Pharez Vitalis <orcid.conceded130@passinbox.com>
Classifier: Programming Language :: Python :: 3
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: datasets
Requires-Dist: dearpygui
Requires-Dist: loguru
Requires-Dist: numpy
Requires-Dist: omegaconf
Requires-Dist: pandas
Requires-Dist: rapidfuzz
Requires-Dist: scikit-learn
Requires-Dist: torch
Dynamic: license-file

# m-strat

This project was done by Pharez Vitalis for a Master's of AI at Queen Mary's University of London.

Please [Give me some feedback!](https://forms.gle/sw6MmfX8j3cNiwi9A), the docs are available [here]( https://m-strat-e0776b.gitlab.io/m_strat.html)

# Project Description
The project represents the implementation of an AutoML GUI system with DAGs fusing with the Entity-Component System framework implemented in python.  The project inspirations were drawn from the idea of "What would a game engine to game developer look like for an ML developer?". A system where a developer can extend and alter components in order to produce more complex ML algorithms but also, able to lean into existing systems to fill in gaps in their knowledge.

# Benefits and Architechture
The project is in pre-alpha concept of what is being dubbed a Machine Intelligence Learning Environment (MILE). The system is designed for rapid ML prototyping and works under the followoing core ethos':
1. **State agnostice processing**: There is a main data pipeline called "context" which keeps the current running variables, and node_params which only show the parameterised running of the current node. By using this operational system the nodes remains independent and very reuasable across different projects.
2. **Fully Moddable**: The project is designed under a tenant of glassboxing, so that you the user can create your own Nodes within the AutoML graph simply and easily. And because you are using open source code with open source blueprints, there is a glass sandbox of alterations that can be done to your project, you could even go into the code and write your own execution system and your own custom templates!
3. **Rapid Prototyping**: Create Models quickly and on the Fly using Pipeline features from both PyTorch and SkLearn's libraries using yaml files only! Plug in a url and a local project location and download datasets from HF and anywhere easily. Don't like the way something works? Copy it and change it your way!

#### Note: A full manual is supplied in **user_manual.md**

## Setup and Installation
1. Download and extract the project if it is in a zip folder.
2. Open a command window in the root project folder (in `m_strat` **NOT** 'm_strat/m_strat'!)
3. `pip install .` does a standard installation. `pip install -e .` Is recommended for developers
   The `-e` flag means editable, this allows the project to customised at source without needing to reinstall the package
4. to open the gui now, type in a command window anywhere `m_strat gui`  your GUI should now open!
5. you can also do `m_strat init` or any standard command using the command line interface yourself
 
# Usage
Once the project is installed using `pip install .`, you can launch the app in command terminal by simply typing `m_strat` anywhere.



## Libraries
- **Loguru**: Was used because it is a simple, well-designed logging system. You can use the observer class to help you make simple attachers in api/logger, but you don't have to and you can attach your own stuff.Additionally, when using blueprints it is recommended that you use loguru, which will log to the observers created by m-strat
- **DearPyGui**: Rescued my project when I lost faith in coding GUIs in python! The OOTB Node Space really helped me somewhat land this thing, their use of GPU based rending using C++ bridging is most well implemented.
- **OmegaConf**: A comprehensive but simple to use library that does indeed elegantly do yaml. 
- **RapidFuzz**: Had a query text problem, remembered  a library I used and, to be honest very straightforward.
- **PyTorch, Pandas, NumPy & Scikit-learn**: This is ML, of course they are here!

### Package Citations
```bibtex
@Misc{Yadan2019Hydra,
  author =       {Omry Yadan},
  title =        {Hydra - A framework for elegantly configuring complex applications},
  howpublished = {Github},
  year =         {2019},
  url =          {https://github.com/facebookresearch/hydra}
}

@misc{DearPyGui,
  author = {Hoffstadt, Jonathan and Cothren, Preston},
  title = {Dear PyGui: A fast and powerful Graphical User Interface Toolkit for Python},
  year = {2026},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/hoffstadt/DearPyGui}},
}

@inproceedings{paszke2019pytorch,
  title={PyTorch: An Imperative Style, High-Performance Deep Learning Library},
  author={Paszke, Adam and Gross, Sam and Massa, Francisco and Lerer, Adam and Bradbury, James and Chanan, Gregory and Killeen, Trevor and Lin, Zeming and Gimelshein, Natalia and Antiga, Luca and Desmaison, Alban and K{\"o}pf, Andreas and Yang, Edward and DeVito, Zachary and Raison, Martin and Tejani, Alykhan and Chilamkurthy, Sasank and Steiner, Benoit and Fang, Lu and Bai, Junjie and Chintala, Soumith},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  pages={8024--8035},
  year={2019}
}

@article{scikit-learn,
 title={Scikit-learn: Machine Learning in {P}ython},
 author={Pedregosa, F. and Varoquaux, G. and Gramfort, A. and Michel, V. and Thirion, B. and Grisel, O. and Blondel, M. and Prettenhofer, P. and Weiss, R. and Dubourg, V. and Vanderplas, J. and Passos, A. and Cournapeau, D. and Brucher, M. and Perrot, M. and Duchesnay, E.},
 journal={Journal of Machine Learning Research},
 volume={12},
 pages={2825--2830},
 year={2011}
}

@inproceedings{lhoest-etal-2021-datasets,
  title = {Datasets: A Community Library for Natural Language Processing},
  author = {Lhoest, Quentin and Villanova del Moral, Albert and Jernite, Yacine and Thakur, Abhishek and von Platen, Patrick and Patil, Suraj and Chaumond, Julien and Moro, Mariama and Sutawika, Julien},
  year = {2021},
  publisher = {Association for Computational Linguistics},
  booktitle = {Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations},
  pages = {175--184},
  url = {https://aclanthology.org/2021.emnlp-demo.21}
}

@software{max_bachmann_rapidfuzz,
  author       = {Max Bachmann},
  title        = {rapidfuzz/RapidFuzz},
  publisher    = {Zenodo},
  doi          = {10.5281/zenodo.15133267},
  url          = {https://doi.org/10.5281/zenodo.15133267}
}

@Article{harris2020array,
  title = {Array programming with {NumPy}},
  author = {Charles R. Harris and K. Jarrod Millman and St{\'{e}}fan J. van der Walt and Ralf Gommers and Pauli Virtanen and David Cournapeau and Eric Wieser and Julian Taylor and Sebastian Berg and Nathaniel J. Smith and Robert Kern and Matti Picus and Stephan Hoyer and Marten H. van Kerkwijk and Matthew Brett and Allan Haldane and Jaime Fern{\'{a}}ndez del R{\'{i}}o and Mark Wiebe and Pearu Peterson and Pierre G{\'{e}}rard-Marchant and Kevin Sheppard and Tyler Reddy and Warren Weckesser and Hameer Abbasi and Christoph Gohlke and Travis E. Oliphant},
  year = {2020},
  month = sep,
  journal = {Nature},
  volume = {585},
  number = {7825},
  pages = {357--362},
  doi = {10.1038/s41586-020-2649-2},
  publisher = {Springer Science and Business Media {LLC}},
  url = {https://doi.org/10.1038/s41586-020-2649-2}
}

@software{reback2020pandas,
  author       = {The pandas development team},
  title        = {pandas-dev/pandas: Pandas},
  month        = feb,
  year         = 2020,
  publisher    = {Zenodo},
  version      = {latest},
  doi          = {10.5281/zenodo.3509134},
  url          = {https://doi.org/10.5281/zenodo.3509134}
}
```
