Metadata-Version: 2.4
Name: socket-lm
Version: 0.1.0
Summary: A training-speed-focused LoRA/PEFT library offering four parameter-efficient fine-tuning methods (LoRA, ReFT, LiteLadder, GaLore)
Author: Omur Bera Isik
License: MIT License
        
        Copyright (c) 2026 Ömür Bera Işık
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Keywords: lora,peft,fine-tuning,pytorch,llm,deep-learning
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: torch>=2.0
Provides-Extra: quantization
Requires-Dist: bitsandbytes>=0.41.0; extra == "quantization"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: bitsandbytes>=0.41.0; extra == "dev"
Dynamic: license-file

# Socket

A LoRA/PEFT library focused on training speed. Offers four different parameter-efficient
fine-tuning methods — all written from scratch with no dependency on external PEFT libraries
(`peft` etc.).

```bash
pip install socket
```

> **Important:** PyPI package name is `socket`, but to avoid clashing with Python's built-in `socket`
> (networking) module, **the import name is `socketimport`**:
> ```python
> import socketimport as sk
> ```

## Methods

| Method | What it does | Status |
|---|---|---|
| **LoRA** (+ rsLoRA, DoRA) | Low-rank addition to weights: `W + (α/r)·BA` | Mature, tested |
| **ReFT** (LoReFT) | Intervenes in hidden representations, not weights | Mature, tested |
| **LiteLadder** | Trains a separate "side network" with no backprop to backbone | **Experimental** — validated at small scale, not yet tested with real language data |
| **GaLore** | Full-parameter training with low-rank gradient projection to reduce optimizer memory | Mature, tested |

When to use each method differs: LoRA/ReFT/LiteLadder reduce parameter count (adapter-based),
while GaLore allows **full-parameter** training but only reduces optimizer memory — they are not
interchangeable, but complementary tools.

## Quick Start

### LoRA

```python
import torch
from socketimport import LoRAAdapter, LoRAConfig

model = ...  # any nn.Module (e.g., a Llama model)
config = LoRAConfig(r=16, alpha=32, use_rslora=True, dropout=0.05)
adapter = LoRAAdapter(model, config)

optimizer = torch.optim.AdamW(adapter.trainable_parameters(), lr=1e-4)
# ... normal training loop, forward with adapter(x) ...

adapter.merge()              # zero-overhead embedding for inference
saved = adapter.adapter_state_dict()   # only LoRA weights (KB not MB)
```

Config parameters accept alternative names
(`rank`, `lora_r`, `lora_alpha`, `dora`, `rslora` etc.) — if conflicting values
are provided, an error is raised, not silently chosen.

### ReFT

```python
from socketimport import ReFTAdapter, ReFTConfig

# `layers` must be provided explicitly - Socket does not attempt to guess model architecture
adapter = ReFTAdapter(
    model, layers=model.model.layers, embed_dim=4096,
    config=ReFTConfig(r=4, layers=(8, 16, 24)),
)
optimizer = torch.optim.AdamW(adapter.trainable_parameters(), lr=1e-3)
```

### LiteLadder (experimental)

```python
from socketimport import LiteLadderAdapter, LiteLadderConfig

adapter = LiteLadderAdapter(
    model, layers=model.model.layers, embed_dim=4096, output_dim=32000,
    config=LiteLadderConfig(side_width=256, rank=32, n_taps=4),
)
```

Backward never enters the backbone (showed much lower overhead than LoRA as depth increased in
small-scale tests) — but this has only been validated on synthetic tasks with a single CPU core.
**Not recommended** as production default; should be considered an opt-in experimental option.

### GaLore

```python
from socketimport import GaLoreAdamW, GaLoreConfig, create_galore_param_groups

groups = create_galore_param_groups(model, GaLoreConfig(rank=128, update_proj_gap=200))
optimizer = GaLoreAdamW(groups, lr=1e-4)
# Model is trained FULL-PARAMETER - no adapter, no merge, GaLore
# only reduces optimizer memory usage
```

`GaLoreAdamW` does **not** implement `SocketAdapterBase` — it is not an adapter,
but a standard `torch.optim.Optimizer` subclass.

## Architecture

```
socketimport/
├── core/
│   ├── base.py          # SocketAdapterBase - shared interface for LoRA/ReFT/LiteLadder
│   ├── lora.py
│   ├── reft.py
│   ├── lite_ladder.py
│   └── galore.py
```

All adapter classes (`LoRAAdapter`, `ReFTAdapter`, `LiteLadderAdapter`)
share the same `SocketAdapterBase` interface: `trainable_parameters()`,
`merge()`/`unmerge()`, `adapter_state_dict()`/`load_adapter_state_dict()`.
This allows the training infrastructure (trainer, distributed, profiler) to be
written independently of which method is selected.

## Development

```bash
pip install -e ".[dev]"
pytest tests/ -v
```

444 tests covering all modules: mathematical correctness (e.g., zero-initialization as no-op,
merge/unmerge being inverses), model freezing behavior, end-to-end training actually reducing loss,
and save/load round-trips.

> **Note:** `quantize_linear`/`quantize_model` and their tests require `bitsandbytes`
> (see [Dependencies](#dependencies)). The `dev` extra installs this automatically; if not installed,
> quantization tests will fail with `DependencyError` (not a library error).

## Dependencies

| Package | Required? | For |
|---|---|---|
| `torch>=2.0` | Yes | Entire library |
| `bitsandbytes>=0.41.0` | No (`pip install socket[quantization]`) | Only `quantize_linear`/`quantize_model` (4-bit/8-bit weight quantization) |

All other features (LoRA, ReFT, LiteLadder, GaLore, checkpointing, OOM recovery,
VRAM guard, distributed backend selection, etc.) work with `torch` alone.

## Limitations (honestly)

- All tests run on 1 CPU core with small synthetic tasks — no scale testing on GPU
  or with real language data yet.
- LiteLadder is not a published method; Socket-specific architecture tested at small scale
  (combination of LST + ReFT-style lightweight interventions).
- No multi-seed statistical validation; results are single-seed.

-Note: PyPI does not normally allow this package name; please use -pip install socket-lm- to install it.

## License

MIT — see [LICENSE](LICENSE).
