Metadata-Version: 2.4
Name: slimcuda
Version: 1.6.0.0
Summary: CUDA, MLX, and NumPy backends for large-scale hologram generation and SLM wavefront synthesis.
Author-email: jwangXTS <junlei.wang@me.com>
License: MIT
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: numpy>=2.2
Provides-Extra: cuda
Requires-Dist: cupy-cuda12x<14.0,>=13.0; extra == "cuda"
Requires-Dist: cuda-python<13.0,>=12.4; extra == "cuda"
Requires-Dist: PyOpenGL; extra == "cuda"
Requires-Dist: opencv-python; extra == "cuda"
Requires-Dist: glfw; extra == "cuda"
Requires-Dist: matplotlib; extra == "cuda"
Provides-Extra: mlx
Requires-Dist: mlx>=0.26; (platform_system == "Darwin" and platform_machine == "arm64") and extra == "mlx"
Requires-Dist: PyOpenGL; extra == "mlx"
Requires-Dist: opencv-python; extra == "mlx"
Requires-Dist: glfw; extra == "mlx"
Requires-Dist: matplotlib; extra == "mlx"
Provides-Extra: cpu
Requires-Dist: matplotlib; extra == "cpu"
Provides-Extra: cpu-gl
Requires-Dist: PyOpenGL; extra == "cpu-gl"
Requires-Dist: opencv-python; extra == "cpu-gl"
Requires-Dist: glfw; extra == "cpu-gl"
Requires-Dist: matplotlib; extra == "cpu-gl"
Provides-Extra: auto
Requires-Dist: cupy-cuda12x<14.0,>=13.0; platform_system != "Darwin" and extra == "auto"
Requires-Dist: cuda-python<13.0,>=12.4; platform_system != "Darwin" and extra == "auto"
Requires-Dist: mlx>=0.26; (platform_system == "Darwin" and platform_machine == "arm64") and extra == "auto"
Requires-Dist: PyOpenGL; extra == "auto"
Requires-Dist: opencv-python; extra == "auto"
Requires-Dist: glfw; extra == "auto"
Requires-Dist: matplotlib; extra == "auto"

# SLiM-CUDA
[![PyPI Downloads](https://static.pepy.tech/personalized-badge/slimcuda?period=total&units=INTERNATIONAL_SYSTEM&left_color=BLACK&right_color=RED&left_text=downloads)](https://pepy.tech/projects/slimcuda)
---
**SLiM-CUDA** is a backend collection for large-scale hologram generation and wavefront synthesis, designed for high-performance spatial light modulator (SLM) workflows.

The package contains CUDA, Apple MLX/Metal, and NumPy CPU fallback backends under one `slimcuda` namespace. The CUDA and MLX backends are intentionally independent because they target machines that normally cannot share the same accelerator stack.

---

## Features

- CUDA-accelerated weighted Gerchberg–Saxton (WGS)–style solvers
- MLX/Metal backend for Apple Silicon using the original `_og` math path
- NumPy CPU backend for RS-only fallback and oracle testing
- Designed for large multi-focus hologram synthesis
- Backend-specific optional dependencies
- Drop-in CUDA upgrade path for optimized kernels (no API changes)

---

## Installation

```bash
pip install slimcuda
```

The base install is CPU-safe and only installs shared Python code plus NumPy. Select the accelerator backend explicitly:

```bash
# CUDA machines
pip install "slimcuda[cuda]"

# Apple Silicon / Metal machines
pip install "slimcuda[mlx]"

# Let packaging markers choose CUDA on non-macOS and MLX on Apple Silicon
pip install "slimcuda[auto]"

# CPU oracle plus optional simulation plotting
pip install "slimcuda[cpu]"
```

Canonical dispatching import:

```python
from slimcuda import SlimCuda

slm = SlimCuda()                            # CUDA + OpenGL render, legacy default
slm = SlimCuda(backend="cuda", render=False)
slm = SlimCuda(backend="mlx", render=False)
slm = SlimCuda(backend="cpu", render=False)
```

Direct backend imports remain available:

```python
from slimcuda import SlimCudaGl, SlimCuda_base        # CUDA backend
from slimcuda import SlimCudaMlx, SlimCudaMlx_base    # MLX backend
from slimcuda import SlimCudaCpu, SlimCudaCpu_base    # NumPy fallback
```

Backend and render-environment diagnosis is available through the tester:

```bash
python slimcuda_tester.py --backend auto --diagnose-only
python slimcuda_tester.py --backend auto --render --diagnose-only
```

To inspect kernel paths, file presence and versions without loading CUDA or opening a window:

```python
from pprint import pprint
from slimcuda import diagnose_kernels

pprint(diagnose_kernels())
```

The report includes `SLIMCUDA_KERNELS_DIR`, packaged PTX, and default, override,
and selected fatbin paths. The override applies only to the fatbin; a missing
override file falls back to the packaged location. File presence does not check
GPU compatibility or whether a kernel can load.

The report also includes `python_version`, `version_check_enabled`, and
`version` / `version_matches` for the PTX and selected fatbin. Versions are read
directly from the embedded marker; missing, unreadable, unmarked or conflicting
versions return `None`. Compressed fatbins may hide this marker and also return
`None`; actual loading still performs the version check. Disabling that check
does not suppress version information in the diagnostic report.

Python, PTX and fatbin must have exactly the same release version. The loader
checks the embedded kernel version before use. A mismatched or unversioned
fatbin raises an error instead of falling back to PTX; if no fatbin is present,
the packaged PTX is used and checked. Pin the Python package to the version of
your fatbin when needed, for example `pip install slimcuda==1.6.0.0`.

To explicitly bypass this check, set `SLIMCUDA_IGNORE_KERNEL_VERSION=1` before
creating the CUDA backend. This skips only version checks, including missing
version metadata; it does not change kernel loading or make incompatible
interfaces compatible. All other values retain the check.

`build_kernels.ps1` reads the release version from `pyproject.toml`, synchronizes
Python's `__version__`, and embeds it in both kernel builds. For manual builds,
pass the same version to nvcc, for example `-DSLIMCUDA_VERSION=1.6.0.0`.

`SlimCudaCpu_base` implements RS methods only. WGS is intentionally not implemented on CPU because it is not practically useful at full SLM scale.

Beam mode IDs in the NumPy, CUDA and MLX implementations are:

| ID | Mode |
|---|---|
| 0 | Gaussian |
| 1 | Vortex (legacy `rs_lg` / `wgs_lg`) |
| 2 | 2D Airy |
| 3 | 1D Airy |
| 4 | Phase-only HG |
| 5 | Annular Bessel |
| 6 | Tilted vortex |
| 7 | Full LG |
| 8 | Full tilted LG |
| 9 | Perfect vortex (PV) |

IDs 0–2 retain their original meanings. The packaged PTX uses this ordering;
separately distributed fatbins must be rebuilt to match. CUDA and MLX RS/WGS support
all ten modes, including mixtures of phase-only and complex modes. If any beam
uses mode 7, 8 or 9, the entire sum uses complex-amplitude encoding with an
eight-pixel x carrier. These modes assume Gaussian input with waist equal to
half the SLM height; WGS feedback removes the carrier and uses the same
normalized pupils. The existing simulation retains its uniform circular pupil.

RS/WGS methods accept `initial_phases=None` as their last argument: one phase
per beam in radians. `None` retains random initialization; explicit values are
converted to backend FP32 arrays without shape, length, or finite-value checks
and do not consume random numbers. WGS copies the initial phases and continues
updating them during iteration, without modifying the caller's array.

For two-stage coherent composition, save the first encoded phase before
generating the holding light (NumPy/CUDA example):

```python
slm.rs_all(x, y, z, beammode=3, th=theta, ir=ir, initial_phases=phases_a)
phi_a = slm.phase_gpu.copy()  # NumPy/CuPy; use mx.array(...) for MLX
slm.rs(x_hold, y_hold, z_hold, initial_phases=phases_h)
phi_final = slm.compose_phases(phi_a, slm.phase_gpu, g=0.5, delta=0.0)
slm.phase_gpu = phi_final
slm.gl_draw()
```

`compose_phases` returns `arg(exp(1j*phi_a) + g*exp(1j*(phi_h + delta)))`
wrapped to `[0, 2*pi)` as a new backend FP32 array. `g` is relative **amplitude**,
and `delta` is in radians relative to the supplied phase maps, including their
existing global phase offsets. The method does not change its inputs, stored
phase, or display. Keeping the first encoding discards its amplitude before
the second superposition; summing all original beams at once is different.
To display only the final result, generate both stages with `render=False` and
pass the final phase to the existing renderer.

## Kernel Architecture & Performance Model
### Public CUDA kernel set

The PyPI wheel ships with CUDA source and public PTX assets, but CUDA runtime dependencies are installed only through the `cuda` or `auto` extras.

The public CUDA kernel set includes:
- slimcuda_og.ptx 
- Corresponding CUDA source (.cu, .cuh) files

This build prioritizes:
- Broad GPU compatibility
- Reproducibility
- Ease of installation

**⚠️ Performance note**

The PTX kernels are not **performance-optimized** for modern GPUs. They exist to ensure correctness and portability.

### Optimized builds (collaborators)

Highly optimized, GPU-specific kernels are distributed as **fatbin / cubin** binaries and are **not included** in the public wheel.

If an optimized kernel is present locally, SLiM-CUDA will automatically detect and load it.

Benefits:
- Substantially higher throughput
- Reduced launch overhead
- Architecture-specific tuning

If you are a collaborator or have a supported GPU and need optimized kernels, please contact the author.

## Runtime Banner

SLiM-CUDA runtime banners have two layers:
- a loaded-backend indicator, for example `[SLiM-CUDA] Loaded legacy PTX kernels ...`
- an optional explanatory note for non-fatbin backends such as public PTX, MLX/Metal, or NumPy CPU

The optimized fatbin path only prints the loaded-backend indicator. Non-fatbin
backends also print a short performance note by default; hide that note with:
```bash
# Linux / macOS
export SLIMCUDA_BANNER=0

# Windows (PowerShell)
setx SLIMCUDA_BANNER 0
```

To silence both layers programmatically, pass `show_banner=False` to the backend
constructor/factory.

## GPU Compatibility

- Public PTX kernels: should run on most CUDA-capable GPUs
- Optimized kernels: GPU- and build-specific

If you have an optimized kernel but encounter issues on your GPU, please contact the author for a tailored build.

## License

- **Python code**: MIT License
- **Public CUDA source (PTX / .cu)**: MIT License
- **Optimized CUDA binaries**: distributed separately under collaborator-specific terms

## Citation

If you use **SLiM-CUDA** in academic work, please cite the following:

### Primary citation (recommended)
SLiM-CUDA was originally developed to support the methodology described in:

> **Z. Qu et al.**,
> *Deep-learning-aided multi-focal hologram generation*, 
> **Optics & Laser Technology**, 2025.
> DOI: 10.1016/j.optlastec.2024.112056

```bibtex
@article{jwangSlimCuda,
  title   = {Deep-learning-aided multi-focal hologram generation},
  author  = {Qu, Z. and others},
  journal = {Optics & Laser Technology},
  year    = {2025},
  doi     = {10.1016/j.optlastec.2024.112056}
}
```

If your work builds upon or uses the algorithms and concepts enabled by SLiM-CUDA,  
**please cite this publication**.

### Software citation
If you prefer to cite the software directly (e.g. for tooling or infrastructure use), you may cite:

> SLiM-CUDA: GPU-accelerated hologram generation backend.  
> https://pypi.org/project/slimcuda/

A formal software citation entry (BibTeX) will be provided in a future release.


## Disclaimer

This software is intended for research and advanced technical use.

API stability is maintained, but internal kernel implementations may evolve.
