Metadata-Version: 2.4
Name: leeroo-kapso
Version: 0.4.5
Summary: A Self-Improving AI Software Factory (for Measurable Objectives)
Author-email: Leeroo Team <team@leeroo.com>
License: MIT
Project-URL: Homepage, https://leeroo.com
Project-URL: Documentation, https://docs.leeroo.com/docs
Project-URL: Quickstart, https://docs.leeroo.com/docs/quickstart
Project-URL: Benchmarks, https://docs.leeroo.com/docs/benchmarks/mle-bench
Project-URL: Docs in your coding agent, https://docs.leeroo.com/docs/agent-access
Project-URL: Changelog, https://github.com/Leeroo-AI/kapso/releases
Project-URL: Repository, https://github.com/leeroo-ai/kapso
Project-URL: Issues, https://github.com/leeroo-ai/kapso/issues
Keywords: ai,ml,optimization,code-generation,llm,agents
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: GitPython
Requires-Dist: litellm==1.75.0
Requires-Dist: PyYAML
Requires-Dist: python-dotenv
Requires-Dist: mcp<2,>=1.9
Requires-Dist: neo4j
Requires-Dist: weaviate-client>=4.0.0
Requires-Dist: tiktoken
Requires-Dist: openai>=1.0.0
Provides-Extra: aider
Requires-Dist: aider-chat>=0.35.0; extra == "aider"
Requires-Dist: cryptography<46.0.0; extra == "aider"
Provides-Extra: mle
Requires-Dist: neo4j; extra == "mle"
Requires-Dist: openai; extra == "mle"
Provides-Extra: ale
Provides-Extra: dev
Requires-Dist: pytest>=7.0.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0.0; extra == "dev"
Requires-Dist: black>=23.0.0; extra == "dev"
Requires-Dist: flake8>=6.0.0; extra == "dev"
Dynamic: license-file

<h1 align="center">Kapso</h1>

<h4 align="center">A Self-Improving AI Software Factory (for Measurable Objectives)</h4>

<p align="center">
  <a href="https://docs.leeroo.com/docs">Learn more</a> ·
  <a href="https://discord.gg/hqVbPNNEZM">Join Discord</a> ·
  <a href="https://leeroo.com">Website</a>
</p>

<p align="center">
  <a href="https://pypi.org/project/leeroo-kapso/"><img src="https://img.shields.io/pypi/v/leeroo-kapso?color=blue" alt="PyPI"></a>
  <a href="https://discord.gg/hqVbPNNEZM"><img src="https://dcbadge.limes.pink/api/server/hqVbPNNEZM?style=flat" alt="Discord"></a>
  <a href="https://github.com/leeroo-ai/kapso"><img src="https://img.shields.io/github/commit-activity/m/leeroo-ai/kapso" alt="GitHub commit activity"></a>
  <a href="https://www.ycombinator.com/companies/leeroo"><img src="https://img.shields.io/badge/Y%20Combinator-X25-orange?logo=ycombinator&logoColor=white" alt="Y Combinator X25"></a>
</p>

<p align="center">
  If you like this project, please support us by giving it a star ⭐
</p>

<p align="center">
  <a href="https://github.com/Leeroo-AI/kapso/raw/main/docs/images/kapso-hero.mp4">
    <img src="https://raw.githubusercontent.com/Leeroo-AI/kapso/main/docs/images/kapso-hero.gif" alt="Kapso running: a campaign searches a tree of experiments while the leaderboard climbs; every trajectory is harvested into an expertise wiki of scored insights and procedures, which grounds the next campaign" width="880">
  </a>
</p>

---

## News

- **IOAI 2026 · the AI Model Track.** The [International Olympiad in AI](https://ioai-official.org) is the IMO of the AI era — 471 contestants from 108 countries and territories, six expert-designed tasks under a single-GPU budget. In 2026 it opened [IOAI²](https://ioai-official.org/ai-model-track/), where AI systems sit the *same* exam in two fully autonomous 6-hour sessions. Once the clock starts, no human may solve, correct, or improve anything. Kapso entered as one of 14 Founding AI Participants.
  - 🎓 **Outscored all 471 humans** — 536.07 total, above every contestant in the hall.
  - 🏆 **IOAI² Grand Master Trophy** — top 3 among all AI systems entered. [Full results →](benchmarks/ioai2026/README.md)
- **Beats the best foundation model on [RelBench](benchmarks/relbench/README.md)**: on Stanford's benchmark for predictive ML over enterprise data, Kapso passes [KumoRFM-v2](https://docs.nvidia.com/sdgm/rfm/overview) in outcome prediction (81.2 vs 79.6 AUROC over 12 tasks) and forecasting (0.2476 vs 0.2912 NMAE over 9 tasks, 15% less error), and the best reported results in recommendations (18.4 vs 15.7 MAP over 10 tasks). Published results live on the [official RelBench leaderboard](https://huggingface.co/spaces/relbench/leaderboard).

  <img src="https://raw.githubusercontent.com/Leeroo-AI/kapso/main/docs/images/relbench.png" alt="RelBench, three panels: outcome prediction, Kapso 81.2 AUROC against KumoRFM-v2's 79.6; forecasting, Kapso 0.2476 NMAE against 0.2912, lower being better; recommendations, Kapso 18.4 MAP against the best reported 15.7. Each panel is drawn from its own truncated axis" width="820">

- **[Leeroopedia MCP Integration](https://leeroopedia.com)**: Kapso now connects to **Leeroopedia MCP** — your ML & Data Knowledge Wiki. Learnt by AI, built by AI, for AI. A centralized playbook of best practices and expert-level knowledge for Machine Learning and Data domains. Kapso agents use it during ideation and implementation to search knowledge, build plans, diagnose failures, and more.
- **[Moltbook Agents 🦞](https://www.moltbook.com/)**: Build AI agents that optimize other agents and debate on Moltbook! [Get started →](moltbook_bot/README.md)
- **Technical Report**: Our technical report is now available! [Read the paper](https://arxiv.org/abs/2601.21526)
- **#1 on [MLE-Bench](benchmarks/mle/README.md)**: KAPSO achieved top ranking among open-source systems on Kaggle ML competitions (MLE Benchmark).

  <img src="https://raw.githubusercontent.com/Leeroo-AI/kapso/main/docs/images/mle-bench.png" alt="MLE-Bench: medal rate by difficulty split against R&amp;D-Agent, AIRA-dojo, ML-Master and AIDE" width="820">

- **#1 on [ALE-Bench](benchmarks/ale/README.md)**: KAPSO achieved top ranking on long-horizon algorithmic discovery problems (ALE Benchmark).

  <img src="https://raw.githubusercontent.com/Leeroo-AI/kapso/main/docs/images/ale-bench.png" alt="ALE-Bench: Kapso reaches a final rating of 1909 Elo against ALE Agent's 1879, above a breakdown across ten AtCoder Heuristic Contest problems; every row is drawn from its own truncated Elo origin, and the per-problem bar gaps share one scale" width="820">

## What is KAPSO?

Kapso is a self-improving software factory. State an objective and it runs a campaign: it designs candidate solutions, has coding agents implement them, measures how far each one is from the objective, and keeps refining the closest until the objective is met. The result ships to your infrastructure.

The factory improves with use. When a campaign ends, Kapso studies its own work: which ideas closed the distance to the objective, which did not, and under what conditions. Each finding is kept as a lesson with the evidence that earned it, and a lesson stays trusted only as long as it keeps holding up. Kapso also reads outside your walls, repositories and papers, and folds what it finds into the same knowledge hub. Every new campaign begins from that hub, so it starts with what earlier work already established, about the problem and about your systems.

Kapso is an open-source Python framework by [Leeroo](https://leeroo.com), published on PyPI as [`leeroo-kapso`](https://pypi.org/project/leeroo-kapso/).

### The Four Pillars

| Pillar | Method | Description |
|--------|--------|-------------|
| [**Evolve**](https://docs.leeroo.com/docs/evolve/overview) | `.evolve()` | Run iterative experiments to build software for a goal. Uses tree search, coding agents, and KG context to generate and refine solutions. |
| **Learn** | [`.learn()`](https://docs.leeroo.com/docs/learning/overview) / [`.learn_knowledge()`](https://docs.leeroo.com/docs/knowledge/overview) | Two memories: `learn()` mines your own finished campaigns into evidence-priced knowledge cards (experience); `learn_knowledge()` ingests repositories and research into the Knowledge Graph (imported knowledge). |
| [**Research**](https://docs.leeroo.com/docs/research/overview) | `.research()` | Run deep web research to gather ideas and implementation references. Returns structured findings you can feed into the knowledge base or use as context for evolving solutions. |
| [**Deploy**](https://docs.leeroo.com/docs/deployment/overview) | `.deploy()` | Turn a solution into running software. Supports local execution, Docker containers, or cloud platforms like Modal. |

## 🚀 Quickstart

### Installation

**1. Prerequisites.** Kapso runs its inference through coding-agent CLIs
(there is no direct-API fallback), so you need Node.js and both agent
CLIs logged in before anything works:

```bash
# Node.js 18+ (https://nodejs.org), then:
npm install -g @openai/codex            # research, judging, utilities
codex login

npm install -g @anthropic-ai/claude-code  # ideation + implementation (default mode)
claude auth login
```

Add an OpenAI key for embeddings (memory and knowledge-search indexing):

```bash
echo 'OPENAI_API_KEY=sk-...' >> .env
```

**2. Install the package** (Python 3.10+):

```bash
pip install leeroo-kapso
```

The package is `leeroo-kapso`. The PyPI package named `kapso` is an unrelated
WhatsApp tool that installs a `kapso` command of its own and shadows this one;
if it is present, `pip uninstall kapso` first.

**3. Verify the setup:**

```bash
kapso doctor
```

`doctor` reports what **your config** actually needs, and names the exact
fix for anything missing. Requirements follow the config, so an all-codex
setup is never asked for `claude`. Narrow it to one verb to see just that
verb's requirements:

```bash
kapso doctor evolve            # or research / learn_knowledge / learn / deploy
```

Those are the same checks the verb itself runs before it does any work —
`kapso.evolve(...)` fails in seconds on a missing CLI rather than deep
inside a session. Items marked `[-- ]` are optional; they limit features
you may not need (docker, Weaviate, Neo4j, the deploy targets).

**Knowledge-graph backends (optional)** — `learn_knowledge()` and
`kg_index` store into local Weaviate + Neo4j. From a source checkout:

```bash
bash scripts/start_infra.sh   # starts both via docker
```

**From source (for development)**

```bash
git clone https://github.com/leeroo-ai/kapso.git
cd kapso

conda create -n kapso python=3.12 && conda activate kapso
pip install -e .
```

The legacy aider adapter is an extra (`pip install "leeroo-kapso[aider]"`,
Python <3.13); the default claude/codex agents need no extras.

**Leeroopedia MCP (optional)** — connect Kapso to [Leeroopedia](https://leeroopedia.com), a curated ML/AI knowledge base. Get an API key from the [Leeroopedia dashboard](https://app.leeroopedia.com/dashboard), then:

```bash
pip install leeroopedia-mcp
echo 'LEEROOPEDIA_API_KEY=kpsk_your_key_here' >> .env
```

### Basic Usage

The core loop needs nothing beyond the prerequisites above:

```python
from kapso import Kapso

kapso = Kapso()   # no knowledge graph needed to start

# Evolve: build a solution through experimentation. The campaign prints
# `status: <path>` at launch — watch it live from another terminal with
#     kapso watch ./campaign
# If a session needs something only you can provide (an API key, a file),
# the campaign pauses and `kapso inbox reply` resumes it:
#     https://docs.leeroo.com/docs/evolve/inbox
solution = kapso.evolve(
    goal="Optimize the model in train.py; target accuracy > 0.80 on evaluate.py",
    initial_repo="./my_project",         # or omit to start from scratch
    output_path="./campaign",
    time_budget_minutes=120,
)
print(solution.explain())

# Learn from the campaign you just ran: mine the trajectory, grade the
# lessons, and bank evidence-priced knowledge cards. The bank (a local
# git repo) is created automatically on first use — lessons stay on your
# machine until you share them:
#     kapso bank connect <git-url>   # or: kapso bank create org/name
# after which every learn() pushes the bank there.
lesson = kapso.learn(solution)
print(lesson.explain())

# Evolve again — with `learning.serving.enabled: true` in your config,
# the next campaign is served the cards it just earned.
solution2 = kapso.evolve(goal="...", output_path="./campaign2")
```

With the knowledge-graph backends running, you can also import outside
knowledge and serve it to campaigns:

```python
from kapso import Kapso, Source

kapso = Kapso()

# Research the web, then ingest findings + a repository into the KG
findings = kapso.research(
    "RLHF and DPO fine-tuning for legal contract analysis",
    mode=["idea", "implementation"],
)
kapso.learn_knowledge(
    Source.Repo("https://github.com/huggingface/trl"),
    findings.ideas,
    findings.implementations,
    wiki_dir="data/wikis",
)

# Campaigns on this Kapso now consult the knowledge graph automatically
solution = kapso.evolve(goal="Fine-tune Llama-3.1-8B for clause risk classification")
```

A note on budgeting: `depth="light"` bounds the *research* stage only.
`learn_knowledge()` extracts everything the material supports — a small
findings set can still become dozens of linked wiki pages and an
hours-long ingest. Ingest time scales with extractable substance, not
with the depth flag; pass fewer sources when you want a faster ingest.

And to turn a solution into running software:

```python
from kapso import DeployStrategy

deployed = kapso.deploy(solution, strategy=DeployStrategy.LOCAL)
result = deployed.run({"input": "data"})
deployed.stop()
```

### Choosing models

Every model Kapso uses is named in one config file. The packaged default
runs evolve sessions on `claude-opus-5`, the learning crews on
`claude-fable-5`, and codex roles on `gpt-5.6-sol` — but model access is
subscription-dependent (a plan can cap one model while serving another).
To run on different models, copy the packaged config, edit, and point
Kapso at yours:

```python
from pathlib import Path
import yaml
from kapso import Kapso
from kapso.kapso import DEFAULT_CONFIG_PATH

config = yaml.safe_load(Path(DEFAULT_CONFIG_PATH).read_text())
# e.g. run the learning crews on opus instead of fable:
crews = yaml.safe_dump(config).replace("claude-fable-5", "claude-opus-5")
Path("kapso-config.yaml").write_text(crews)

kapso = Kapso(config_path="kapso-config.yaml")
```

Before a long run, probe every model your config names against your
actual subscriptions — a model your login cannot serve fails here in
seconds instead of hours into a run. A usage cap on a model you can
serve is not visible to a one-token probe:

```bash
kapso doctor --models                            # packaged config
kapso doctor --models --config kapso-config.yaml # yours
kapso doctor learn --models                      # just the learning crews
```

To make that probe part of every call rather than a thing you remember to
run, turn it on in your config. It costs one throwaway token per distinct
`{cli, model}` pair, which is worth it before an unattended multi-hour
`learn()`:

```yaml
preflight:
  enabled: true            # static checks before every verb (default)
  live_model_probe: true   # + one-token model probes (default: false)
```

Model swaps change *pacing* too: the crew `timeout_minutes` caps in the
config were calibrated on the default models, and a swapped model that
reasons longer may need them raised.

For detailed integration steps, see the [Quickstart](https://docs.leeroo.com/docs/quickstart) and [Installation](https://docs.leeroo.com/docs/installation) guides.

## Examples

| Example | Description |
|---------|-------------|
| [**CUDA Optimization**](examples/cuda_optimization/README.md) | Optimize CUDA kernels for GPU performance |
| [**PyTorch Optimization**](examples/pytorch_optimization/README.md) | Cut wall-clock and memory — fuse ops, kill sync points and host-device chatter, saturate the GPU without changing numerics |
| [**ML Model Development**](examples/ml_model_development/README.md) | End-to-end delivery of prediction models — data prep, features, training, and validation evolved into a deployable artifact |
| [**Harness Optimization**](examples/prompt_engineering/README.md) | Evolve the harness around a model — prompts, decoding, parsing, and scoring tuned against a measurable target |
| [**Agent Optimization**](examples/agentic_scaffold/README.md) | Agents improving agents — workflows, tools, and prompts evolved until the metric climbs |

## Supported Benchmarks

| Benchmark | Description |
|-----------|-------------|
| [**MLE-Bench**](benchmarks/mle/README.md) | OpenAI's ML-engineering benchmark — full competitions across tabular, vision, text, and audio, from raw data to graded submission |
| [**ALE-Bench**](benchmarks/ale/README.md) | Sakana AI's algorithmic-optimization benchmark — design, implement, and iterate contest heuristics over hours-long searches |
| [**RelBench**](benchmarks/relbench/README.md) | Stanford's benchmark for predictive ML over enterprise data — forecasting, classification, and recommendation straight from the multi-table databases of SAP, Amazon, H&M, and more |
| [**IOAI 2026**](benchmarks/ioai2026/README.md) | Timed olympiad ML across vision, language, and optimization — expert-set tasks, contest hardware, zero human help |

Each benchmark also has a documentation page covering how to run it, its CLI options and what the output holds:
[IOAI 2026](https://docs.leeroo.com/docs/benchmarks/ioai-2026) ·
[MLE-Bench](https://docs.leeroo.com/docs/benchmarks/mle-bench) ·
[ALE-Bench](https://docs.leeroo.com/docs/benchmarks/ale-bench) ·
[RelBench](https://docs.leeroo.com/docs/benchmarks/relbench)

## 📚 Documentation & Support

- **Full Documentation**: [docs.leeroo.com](https://docs.leeroo.com/docs)
  - [Installation](https://docs.leeroo.com/docs/installation) — the coding-agent CLIs, the package, and `kapso doctor`
  - [Quickstart](https://docs.leeroo.com/docs/quickstart) — your first campaign
  - [CLI reference](https://docs.leeroo.com/docs/reference/cli) — every command, flag and default
  - [Python API](https://docs.leeroo.com/docs/reference/kapso-api) — every public method on `Kapso`
  - [Configuration](https://docs.leeroo.com/docs/reference/configuration) — every key in `config.yaml`
  - [Trajectory learning](https://docs.leeroo.com/docs/learning/overview) — the lesson bank, grading and serving
  - [Evolve](https://docs.leeroo.com/docs/evolve/overview) — how a campaign runs, and [what happens inside one experiment](https://docs.leeroo.com/docs/evolve/execution-flow)
  - [Knowledge graph](https://docs.leeroo.com/docs/knowledge/overview) — repositories and papers turned into searchable knowledge
  - [Deployment](https://docs.leeroo.com/docs/deployment/overview) — local, Docker, Modal and the other strategies
  - [Kapso skill for coding agents](https://docs.leeroo.com/docs/coding-agent-skills) — launch and resume campaigns from Claude Code, Codex or OpenCode
- **From your coding agent**: the docs run an MCP server at `https://docs.leeroo.com/mcp`, and every page is served as Markdown — see [llms.txt](https://docs.leeroo.com/llms.txt) and [Docs in your coding agent](https://docs.leeroo.com/docs/agent-access).
  ```bash
  claude mcp add --transport http kapso-docs https://docs.leeroo.com/mcp   # Claude Code
  codex mcp add kapso-docs --url https://docs.leeroo.com/mcp              # Codex CLI
  ```
- **Let your coding agent run Kapso**: one `SKILL.md` per agent for Claude Code, Codex and OpenCode lives in [`skills/`](https://github.com/Leeroo-AI/kapso/tree/main/skills) — copy the folder for your agent into the project; see [Kapso skill for coding agents](https://docs.leeroo.com/docs/coding-agent-skills).
- **Community**: [Discord](https://discord.gg/hqVbPNNEZM)
- **Website**: [leeroo.com](https://leeroo.com)


## Kapso for Enterprise

Kapso gets better at your company the longer it works: every task feeds a living knowledge bank of your systems, your data, and your hard-won lessons. To onboard Kapso for your challenging enterprise tasks and build that live company context, [talk to us](https://leeroo.com/contact-us).

## Contributing

We welcome contributions! Please see our [Contributing Guide](CONTRIBUTING.md) for details on how to get started.
