Metadata-Version: 2.4
Name: asi-evolve
Version: 0.1.4
Summary: Pythonic wrapper around the ASI_Evolve (arXiv:2603.29640) autonomous evolutionary-search framework. Give it a problem and an evaluator; it runs until it finds a solution. CLI + --serve dashboard, cross-provider.
Author-email: lordxmen2k <lordxmen2k@users.noreply.github.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/lordxmen2k/asi-evolve
Project-URL: Repository, https://github.com/lordxmen2k/asi-evolve
Project-URL: Paper, https://arxiv.org/abs/2603.29640
Project-URL: Upstream, https://github.com/GAIR-NLP/ASI_Evolve
Keywords: ai,llm,evolution,agent,asi-evolve,evolutionary-search,automl,research,optimization
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: LICENSE-VENDORED
License-File: NOTICE
Requires-Dist: openai>=1.0.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: jinja2>=3.0
Requires-Dist: numpy>=1.20.0
Requires-Dist: faiss-cpu>=1.7.0
Requires-Dist: sentence-transformers>=2.2.0
Requires-Dist: rich>=13.0.0
Requires-Dist: flask>=2.3.0
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.20.0; extra == "anthropic"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: twine>=4.0; extra == "dev"
Dynamic: license-file

# asi-evolve

Pythonic wrapper around [ASI-Evolve](https://github.com/GAIR-NLP/ASI-Evolve) (arXiv:2603.29640) — the autonomous evolutionary-search framework that already produces state-of-the-art linear-attention architectures, pretraining-data curation pipelines, RL algorithms, and drug-target interaction models.

**Give it a problem and an evaluator. It runs until it finds a better solution.**

```
$ export MINIMAX_API_KEY=sk-...
$ asi-run --problem ./problem.md --initial ./baseline.py --evaluator ./eval.py
[asi-evolve] starting run 'asi-evolve-2026-09-24-205333'
[asi-evolve] provider=minimax model=MiniMax-M3
[2026-09-24 20:53:33] [INFO] === Step 1 ===
[2026-09-24 20:53:33] [INFO] [Engineer] Eval score: 0.9598
[2026-09-24 20:53:50] [INFO] [Researcher] Generated: lp_optimized_with_multistart_search
[2026-09-24 20:53:53] [INFO] [Engineer] Eval score: 2.2301
[2026-09-24 20:54:00] [INFO] Updated best snapshot: lp_optimized_with_multistart_search (score=2.2301)
```

In one round the loop improved the circle-packing score from `0.9598` to `2.2301` (a 132% improvement over the naive ring baseline). ASI-Evolve upstream is the engine; this package is the ergonomic Python surface.

---

## Table of contents

1. [Step-by-step: your first run](#step-by-step-your-first-run) ← **start here**
2. [What you get](#what-you-get)
3. [Install](#install)
4. [Detailed walkthrough](#detailed-walkthrough) — CLI flags, dashboard, loop, stop conditions, provider matrix, v0.1.4 changes
5. [Examples](#examples) — 7 fully-embedded example files (copy-paste and run)
6. [FAQ](#faq)
7. [License](#license)

---

## Step-by-step: your first run

This walkthrough is for someone who has never used asi-evolve. Follow each step in order. By the end you'll have run a real evolutionary search against a real LLM, gotten a better solution than you started with, and known exactly which files to look at when things go wrong.

### Step 1 — Install

asi-evolve ships on PyPI as a single Python package. You don't need git, you don't need to clone anything. Just `pip install` from any venv:

```bash
# Create a fresh venv (one-time per project)
python -m venv .venv

# Activate it
source .venv/bin/activate                # Git Bash / Linux / macOS
.venv\Scripts\activate                   # cmd.exe (Windows PowerShell: .venv\Scripts\Activate.ps1)

# Install asi-evolve (pulls ~5 deps: openai, httpx, faiss-cpu, sentence-transformers, flask)
pip install --upgrade asi-evolve

# Verify — this prints "asi-evolve 0.1.4" and exits 0
asi-run --version
```

If `asi-run` is not found after install, your `pip install` and `python` are pointing to different Pythons (common on systems with multiple Python installs). Run `which pip` and `which python` — they should print paths inside `.venv/`. If not, see the FAQ entry "Why a venv?" at the end of this README.

Requires Python 3.10+. 3.11 recommended.

### Step 2 — Pick a problem and write three files

asi-evolve doesn't care what problem you give it, as long as you provide these files in the same directory:

**`problem.md`** — a markdown description of what you want the code to do. Be specific about inputs, outputs, and constraints. The LLM reads this every round.

```bash
cat > problem.md <<'EOF'
# Pack 26 circles in a unit square to maximize the sum of radii.

Inputs:
  - n_circles: int = 26
  - unit_square: a 1.0 x 1.0 box at origin

Output:
  - centers: list of (x, y) tuples, each in [0, 1]
  - radii: list of n_circles floats, all > 0

Constraints:
  - Every circle must fit entirely inside the unit square
  - No two circles may overlap (distance >= r_i + r_j)
  - Maximize the sum of radii
EOF
```

**`initial_program.py`** — a baseline implementation. Doesn't have to be good — the loop improves it.

```bash
cat > initial_program.py <<'EOF'
"""Naive baseline: pack 26 equal-radius circles on a hex grid."""
import math

def pack(n_circles=26):
    side = math.ceil(math.sqrt(n_circles))
    radius = 0.5 / side
    centers, radii = [], []
    for i in range(side):
        for j in range(side):
            x = (j + 0.5 + (i % 2) * 0.5) / side
            y = (i + 0.5) / side
            if 0 < x < 1 and 0 < y < 1 and len(centers) < n_circles:
                centers.append((x, y))
                radii.append(radius)
    return centers, radii
EOF
```

**`evaluator.py`** — a Python function that scores a candidate. Takes the candidate file path, returns a dict with at least `eval_score`:

```bash
cat > evaluator.py <<'EOF'
import math
from pathlib import Path
import importlib.util

def evaluate(code_path):
    spec = importlib.util.spec_from_file_location("candidate", code_path)
    candidate = importlib.util.module_from_spec(spec)
    spec.loader.exec_module(candidate)
    centers, radii = candidate.pack(26)

    score = 0.0
    if not centers or len(radii) != 26:
        return {"success": False, "eval_score": 0.0, "error": "wrong shape"}

    for (x, y), r in zip(centers, radii):
        if x - r < 0 or x + r > 1 or y - r < 0 or y + r > 1:
            return {"success": False, "eval_score": 0.0, "error": "out of bounds"}

    for i in range(26):
        for j in range(i + 1, 26):
            (x1, y1), r1 = centers[i], radii[i]
            (x2, y2), r2 = centers[j], radii[j]
            if math.hypot(x1 - x2, y1 - y2) < r1 + r2 - 1e-9:
                return {"success": False, "eval_score": 0.0, "error": "overlap"}

    return {
        "success": True,
        "eval_score": sum(radii),
        "sum_radii": sum(radii),
    }

if __name__ == "__main__":
    import sys, json
    result = evaluate(sys.argv[1])
    Path("results.json").write_text(json.dumps(result, indent=2))
    print(f"eval_score: {result.get('eval_score', 0.0):.4f}")
EOF
```

**`eval.sh`** — a thin bash wrapper. The framework runs `bash eval.sh` with no arguments from a per-step cwd; the candidate code is at `$PWD/code` and `results.json` must be written to `$PWD`:

```bash
cat > eval.sh <<'EOF'
#!/usr/bin/env bash
# The framework invokes `bash eval.sh` (no args) from a per-step cwd.
# The candidate code is at $PWD/code and results.json must be in $PWD.
set -uo pipefail
STEP_DIR="$(pwd)"
python "$(dirname "$0")/evaluator.py" "$STEP_DIR/code"
EOF
chmod +x eval.sh
```

**Optional: `cognition.md`** — domain knowledge the LLM reads each round. Highly recommended for any non-trivial problem:

```bash
cat > cognition.md <<'EOF'
# Circle packing heuristics

1. **Hexagonal close packing** is the densest infinite-plane arrangement.
   Edge effects reduce achievable density in a unit square.

2. **Variable radii help a LOT.** Equal radii waste space at corners.
   Larger circles belong in the interior; smaller ones along edges.

3. **scipy.optimize.linprog** with HiGHS is fast for the LP relaxation
   of the contact graph.

4. **scipy.optimize.differential_evolution** searches over center
   positions to escape local minima.

5. **Validate first.** If any circle is out of bounds or overlapping,
   score is 0. Don't waste compute on invalid configurations.
EOF
```

That's it. You now have 4-5 files in a directory and you're ready to run.

### Step 3 — Set your API key

The recommended default is **MiniMax-M3 via MiniMax**. Set the key as an env var (do NOT put it on the command line — it leaks into shell history):

```bash
export MINIMAX_API_KEY=sk-cp-...
```

Other supported providers — set the matching env var:

| Provider         | Env var               | Recommended model |
|------------------|----------------------|---------------------|
| MiniMax          | `MINIMAX_API_KEY`    | `MiniMax-M3` or `MiniMax-M2.7-highspeed` |
| OpenAI           | `OPENAI_API_KEY`     | `gpt-4o` |
| Anthropic        | `ANTHROPIC_API_KEY`  | `claude-3-5-sonnet-latest` |
| Google           | `GOOGLE_API_KEY`     | `gemini-1.5-pro` |
| OpenRouter       | `OPENROUTER_API_KEY` | any model on openrouter.ai |
| xAI              | `XAI_API_KEY`        | `grok-2` |
| DeepSeek         | `DEEPSEEK_API_KEY`   | `deepseek-chat` |
| Mistral          | `MISTRAL_API_KEY`    | `mistral-large-latest` |
| Together         | `TOGETHER_API_KEY`   | `meta-llama/Llama-3-70b-chat-hf` |
| Groq             | `GROQ_API_KEY`       | `llama-3.1-70b-versatile` |
| Fireworks        | `FIREWORKS_API_KEY`  | `accounts/fireworks/models/llama-v3p1-70b-instruct` |
| Ollama (local)   | (none)               | `qwen2.5-coder:7b` |
| LM Studio (local) | (none)              | any loaded model |
| llama.cpp (local) | (none)               | any GGUF |
| vLLM (local)     | (none)               | any model |

For local servers, no API key is needed — the wrapper falls back to `"EMPTY"`.

### Step 4 — Run

The minimum command is:

```bash
asi-run \
  --problem problem.md \
  --initial initial_program.py \
  --evaluator evaluator.py
```

If you made `eval.sh` and `cognition.md`, add those too:

```bash
asi-run \
  --problem problem.md \
  --initial initial_program.py \
  --evaluator evaluator.py \
  --eval-script eval.sh \
  --cognition cognition.md \
  --name my-first-run \
  --output runs/my-first-run
```

**What you'll see in the console:**

```
[asi-evolve] starting run 'my-first-run'
[asi-evolve] provider=minimax model=MiniMax-M3
[asi-evolve] cognition: populated 7 items from cognition.md
[2026-09-25 14:23:01] [INFO] Starting new experiment: my-first-run
[2026-09-25 14:23:01] [INFO] [Engineer] Eval score: 0.9598   ← baseline score
[2026-09-25 14:23:25] [INFO] [Researcher] Generated: lp_optimized_with_multistart_search
[2026-09-25 14:23:31] [INFO] [Engineer] Eval score: 2.2301   ← better!
[2026-09-25 14:23:32] [INFO] Updated best snapshot: lp_optimized_with_multistart_search (score=2.2301)
[2026-09-25 14:24:00] [INFO] [Researcher] Generated: lp_with_polar_init
[2026-09-25 14:24:08] [INFO] [Engineer] Eval score: 2.4189   ← even better
```

The wrapper saves the best candidate at every improvement and the entire history under `runs/my-first-run/experiments/my-first-run/`.

### Step 5 — Stop the run

Three options:

1. **Wait for it to finish** — default `--max-rounds 50`. Override with `--max-rounds 8` for a smoke test.
2. **Stop when stuck** — `--plateau 5` aborts after 5 consecutive rounds with no score improvement.
3. **Press Ctrl-C** — soft-stops cleanly at the next round boundary. Press twice to force-kill.

```bash
# Long job with sensible stop conditions
asi-run \
  --problem problem.md --initial initial_program.py --evaluator evaluator.py \
  --eval-script eval.sh --cognition cognition.md \
  --max-rounds 50 \
  --plateau 5 \
  --max-hours 4 \
  --name my-run --output runs/my-run
```

Stops at the **first** of: 50 rounds, 5-round plateau, or 4 hours.

### Step 6 — Inspect the results

The whole run lives in `runs/<name>/`:

```
runs/my-first-run/
└── experiments/
    └── my-first-run/
        ├── input.md              ← problem.md + auto-extracted program interface
        ├── initial_program       ← your baseline (verbatim copy)
        ├── evaluator.py          ← your evaluator (verbatim copy)
        ├── eval.sh               ← your eval script (verbatim copy)
        ├── config.yaml           ← merged config the wrapper built
        ├── cognition_data/       ← FAISS-indexed cognition chunks
        │   ├── cognition.json
        │   └── faiss/
        ├── database_data/        ← the candidate database
        ├── steps/
        │   ├── step_0_initial/   ← your baseline
        │   ├── step_1/           ← first improvement
        │   ├── step_2/           ← second improvement
        │   └── best/             ← symlink to the current best step
        └── logs/
            ├── evolve.log        ← human-readable run log
            └── errors.log        ← full Python tracebacks (only on failure)
```

**Your best candidate**: `runs/<name>/experiments/<name>/steps/best/code`. Score in `runs/<name>/experiments/<name>/steps/best/results.json`.

**Full history**: `runs/<name>/experiments/<name>/database_data/nodes.json`. Every candidate, its score, its parent, the Analyzer's lesson.

### Step 7 — Run with the dashboard (optional)

For a long run, the web dashboard is helpful:

```bash
asi-run \
  --serve --port 5000 \
  --problem problem.md --initial initial_program.py --evaluator evaluator.py \
  --eval-script eval.sh --cognition cognition.md \
  --name my-run --output runs/my-run
```

Open `http://localhost:5000` in a browser. You see:
- Current best score and step number
- Per-round timeline (which step improved the score)
- Live cognition store contents
- A Stop button that soft-stops the run

The dashboard polls every 2 seconds. When you close the browser tab, the run continues.

---

## What you get

| Surface | What it does |
|---|---|
| `asi-run` CLI | One command, no boilerplate, runs the loop. |
| `--serve` flag | Spins up a Flask dashboard on `localhost:5000` for live monitoring. |
| Library import | `from asi_evolve import Runner; Runner(opts).run()` for embedding. |
| 7 working examples | Each is a complete, runnable file with a 6-section docstring header (inlined below). |
| 12+ provider integrations | OpenAI, Anthropic, MiniMax, Google, OpenRouter, xAI, DeepSeek, Mistral, Together, Groq, Fireworks, plus any local OpenAI-compat server. |

## Install

The quick install is at the top of this README. Summary:

```bash
# Create a fresh venv
python -m venv .venv
source .venv/bin/activate                # Git Bash / Linux / macOS
.venv\Scripts\activate                   # cmd.exe

# Install asi-evolve from PyPI
pip install --upgrade asi-evolve

# Verify
asi-run --version
# asi-evolve 0.1.4
```

Requires Python 3.10+.

### Building from source (contributors only)

```bash
git clone https://github.com/lordxmen2k/asi-evolve.git
cd asi-evolve

python -m venv .venv
source .venv/bin/activate                # Git Bash / Linux / macOS
.venv\Scripts\activate                   # cmd.exe

pip install --upgrade build twine
python -m build
pip install --force-reinstall dist/asi_evolve-*.whl
```

---

## Detailed walkthrough

### A. Seven working examples

Each example is a single Python file you can copy, save as `example.py`, and run end-to-end. The full source is below.

| # | File | Demonstrates | LLM calls |
|---|---|---|---|
| 1 | `01_quickstart.py` | Library import — drop-in `Runner` instead of CLI | yes |
| 2 | `02_cli_run.py` | CLI from your own files (`problem.md`, `initial.py`, `evaluator.py`, `eval.sh`) | yes |
| 3 | `03_stop_conditions.py` | Plateau / max-hours / threshold / max-rounds | yes |
| 4 | `04_cross_provider.py` | Same loop on MiniMax / OpenAI / Anthropic / DeepSeek | yes |
| 5 | `05_local_model.py` | Ollama / LM Studio / llama.cpp / vLLM | yes |
| 6 | `06_resilient_run.py` | `--inter-call-delay` + circuit breaker + `--request-timeout` | yes |
| 7 | `07_serve_dashboard.py` | Web dashboard (`--serve`) for live monitoring | yes |

### B. CLI anatomy (`asi-run --help`)

```
usage: asi-run [-h] --problem PROBLEM --initial INITIAL --evaluator EVALUATOR
               [--eval-script EVAL_SCRIPT] [--cognition COGNITION]
               [--name NAME] [--output OUTPUT]
               [--provider {openai,anthropic,minimax,google,openrouter,...}]
               [--base-url BASE_URL] [--api-key API_KEY] [--model MODEL]
               [--temperature TEMPERATURE] [--top-p TOP_P] [--max-tokens MAX_TOKENS]
               [--thinking-effort {low,high}]
               [--thinking-effort-low | --thinking-effort-high]
               [--max-rounds MAX_ROUNDS] [--sample-n SAMPLE_N]
               [--max-hours MAX_HOURS] [--plateau PLATEAU] [--threshold THRESHOLD]
               [--inter-call-delay INTER_CALL_DELAY]
               [--max-consecutive-connection-failures MAX_CONSECUTIVE_CONNECTION_FAILURES]
               [--serve] [--port PORT] [--host HOST]
```

**Required (kronos-finance fail-fast rule):** `--problem`, `--initial`, `--evaluator`. The CLI validates all three files exist before launching the loop.

**Provider (defaults to MiniMax):** `--provider`, `--base-url`, `--api-key`, `--model`. The `--api-key` defaults to the matching env var (`MINIMAX_API_KEY`, `OPENAI_API_KEY`, etc.). The `--model` defaults to the per-provider canonical model if omitted.

**MiniMax thinking effort:** `--thinking-effort high` uses careful planning (recommended for `Researcher` and `Manager` calls); `--thinking-effort low` is the cheap, fast default. If neither is set, the wrapper uses a per-agent default: `researcher=high, manager=high, engineer=low, analyzer=low`.

**Loop control:** `--max-rounds` (default 50), `--max-hours`, `--plateau K`, `--threshold T`. Composite semantics: stop on **first-of** `{threshold, plateau, max-hours, max-rounds}`.

**Resilience (v0.1.4):** `--inter-call-delay SECONDS` sleeps N seconds AFTER each successful LLM call (helps throttle provider rate limits); `--max-consecutive-connection-failures N` is the circuit-breaker that stops the run after N consecutive `APIConnectionError` steps instead of burning the rest of the round budget on dead network. See section E.2 below.

**Surface:** `--serve` boots the Flask dashboard on `--host:--port` (default `127.0.0.1:5000`).

### C. The web dashboard (`--serve`)

Seven routes, all returning 200:

| Endpoint | Purpose |
|---|---|
| `GET /` | Single-page HTML dashboard |
| `GET /api/run/<id>/status` | Current state, best score, elapsed time |
| `GET /api/run/<id>/best` | Top candidate's full code + motivation |
| `GET /api/run/<id>/history` | Per-round timeline (round, score, code, motivation) |
| `GET /api/run/<id>/cognition` | Domain-knowledge store contents |
| `GET /api/run/<id>/stop` | Soft-stop the run (returns immediately) |
| `GET /api/runs` | List all runs in `runs/` |

The JS polls `/api/run/<id>/status` every 2 seconds and updates the status card, best-code panel, history table, and cognition preview. The Stop button hits `/api/run/<id>/stop` which sets a flag the runner checks at the next round boundary.

### D. How the loop actually runs

1. **Setup** — The CLI builds the upstream-compatible folder under your `--output` directory: `runs/<name>/experiments/<name>/` gets `input.md`, `initial_program`, `evaluator.py`, `eval.sh`, `config.yaml`. The vendored tree is never written to.
2. **Cognition population** — If `--cognition` is a markdown file, the wrapper reads it, chunks it (markdown headings + numbered lists), and writes the items + FAISS vectors into `exp_dir/cognition_data/`. If `--cognition` is `.py`, it's run as a custom script. Otherwise skipped.
3. **Pipeline init** — The vendored `Pipeline` reads your `config.yaml`, builds its DB + cognition store + LLM client. We hook `LLMClient.chat()` to strip `<think>` blocks (MiniMax-specific) and inject per-agent thinking effort.
4. **Loop** — Each round samples N context nodes, asks the `Researcher` to write a new candidate (diff or full rewrite), runs it via `Engineer` and your `eval.sh`, scores it, records it in the database, and asks the `Analyzer` to write 1-2 lessons learned.
5. **Stop check** — At each round boundary we check the stop conditions you specified. First-of-composite wins.
6. **Output** — Best snapshot is updated on every improvement. Final summary prints the best score and the winning candidate's path.

### E. Stop conditions reference

| Condition | Fires when | CLI flag |
|---|---|---|
| Threshold | any candidate's score >= T | `--threshold 2.0` |
| Plateau | best score hasn't improved in K rounds | `--plateau 5` |
| Wall-clock | elapsed >= H hours | `--max-hours 6` |
| Round budget | ran M rounds | `--max-rounds 50` |

**Composite semantics:** threshold is checked first (most decisive), then plateau, then wall-clock, then round budget. `first-of {threshold, plateau, max-hours, max-rounds}`.

To stop on `"first of plateau or 8 rounds or 2 hours"`: `--plateau 3 --max-rounds 8 --max-hours 2`.

### E.2 Resilience knobs (v0.1.4)

When your provider is flaky or rate-limited, two extra flags keep a run from
burning hours on dead networks:

| Knob | Fires when | CLI flag | Default |
|---|---|---|---|
| Inter-call delay | sleep N seconds AFTER each successful LLM call | `--inter-call-delay 2.0` | 0 (off) |
| Connection circuit-breaker | N consecutive `APIConnectionError` steps | `--max-consecutive-connection-failures 3` | 3 |

**Why inter-call-delay:** providers like MiniMax-M3 enforce a per-minute
quota during peak hours. The default retry policy already waits 5s between
retries of the *same* call; this knob adds a sleep *between successful calls*
so you don't trip the quota on long runs.

**Why circuit-breaker:** without it, a single TCP-reset storm keeps the
runner spinning through 3 retries × 9 minutes = 27 minutes per failed step.
With `--max-consecutive-connection-failures 3`, the run aborts with
`StopReason("connection_failures", ...)` after 3 dead steps in a row
(~3 × 9 minutes = ~27 minutes) instead of running the full budget.

**Note on plateau (v0.1.4 behavior change):** a round where the LLM call
drops or the eval script crashes now increments the plateau counter the same
way a non-improving successful step does. Previously, `--plateau 3` would
silently ignore failed rounds and never fire during connection storms.

### F. Provider matrix (cross-provider)

| Provider | `--base-url` | Default model | Thinking-effort |
|---|---|---|---|
| `minimax` (recommended) | `https://api.minimax.io/v1` | `MiniMax-M3` | yes (low/high) |
| `openai` | `https://api.openai.com/v1` | `gpt-4o` | no |
| `anthropic` | `https://api.anthropic.com/v1` | `claude-sonnet-4-5` | no |
| `google` | `https://generativelanguage.googleapis.com/v1beta/openai` | `gemini-2.5-pro` | no |
| `openrouter` | `https://openrouter.ai/api/v1` | `anthropic/claude-sonnet-4-5` | no |
| `xai` | `https://api.x.ai/v1` | `grok-3` | no |
| `deepseek` | `https://api.deepseek.com/v1` | `deepseek-chat` | no |
| `mistral` | `https://api.mistral.ai/v1` | `mistral-large-latest` | no |
| `together` | `https://api.together.xyz/v1` | `meta-llama/Llama-3.1-70B-Instruct` | no |
| `groq` | `https://api.groq.com/openai/v1` | `llama-3.1-70b-versatile` | no |
| `fireworks` | `https://api.fireworks.ai/inference/v1` | `accounts/fireworks/models/llama-v3p1-70b-instruct` | no |
| `ollama` (local) | `http://localhost:11434/v1` | `llama3.1:8b` | depends on the model |
| `lmstudio` (local) | `http://localhost:1234/v1` | `local-model` | depends on the model |
| `llamacpp` (local) | `http://localhost:8080` | `gpt-3.5-turbo` | no |
| `vllm` (local) | `http://localhost:8000/v1` | `meta-llama/Llama-3.1-8B-Instruct` | no |
| `koboldcpp` (local) | `http://localhost:5001/v1` | `koboldcpp` | no |
| `textgen` (local) | `http://localhost:7860` | `text-generation-webui` | no |
| `local` | `http://localhost:<port>/v1` | provider-specific | catch-all for any localhost URL |

`detect_provider()` infers the family from a substring match against the base URL — most popular local servers have unique default ports that get auto-detected. Unknown hosts fall through to `openai-compat`. Anthropic requires `--provider=anthropic` because the wire format is structurally different (different message format, `max_tokens` required). Local servers accept `"EMPTY"` as an API key — the wrapper falls back to that automatically when no env var is set, so you don't need a real key for Ollama/LM Studio/etc.

### G. Citation

If you publish results from this package, cite both the upstream paper and (if useful) this wrapper:

```
@misc{asi-evolve-paper,
  title   = {ASI-Evolve: An Agentic Framework for Autonomous Evolutionary Search},
  author  = {{GAIR-NLP}},
  year    = {2026},
  journal = {arXiv preprint arXiv:2603.29640}
}
```

### H. License

* **Wrapper code** (everything outside `src/asi_evolve/_vendor/`): MIT. See `LICENSE`.
* **Vendored upstream**: Apache License 2.0. See `LICENSE-VENDORED`.
* **Citation** requirements: see `NOTICE`.

### I. New in v0.1.4 — runtime behavior changes you should know about

These are user-visible behavior changes in v0.1.4 (vs v0.1.3). If you've been
running v0.1.3 and the run looks different now, this section is why.

#### I.1 — `--version` / `-V` no longer errors

```bash
$ asi-run --version
asi-evolve 0.1.4

$ asi-run -V
asi-evolve 0.1.4
```

Both flags exit 0 with the formatted version string and short-circuit
**before** required-arg validation. (Previously `asi-run --version`
dumped full help + `error: the following arguments are required: ...`,
which was a UX bug.)

#### I.2 — Ctrl-C actually stops the run

The previous version set a soft-stop flag at SIGINT but only checked it
between rounds, so Ctrl-C during a 90-second LLM call would sit and
wait for the call to drain. v0.1.4 does both:

* **First Ctrl-C:** sets the soft-stop flag + prints
  `[asi-evolve] Ctrl-C received — finishing current step and stopping
  cleanly. Press Ctrl-C again to kill immediately.`
* **Second Ctrl-C:** restores the default SIGINT handler and re-raises
  `KeyboardInterrupt` — process exits with code **130** within ~1s,
  even mid-LLM-call.

Tested end-to-end via subprocess: broken API key + SIGINT after 15s
yields exit code 130 within 30s.

#### I.3 — Circuit breaker raises `CircuitBreakerTripped` mid-run

When `--max-consecutive-connection-failures N` is hit (default 3), the
runner stops with `StopReason("connection_failures", ...)` **without
waiting for the current step to finish**. Previously the breaker only
fired at loop boundaries; now it interrupts the current `pipeline.run()`
call directly. Worst case before abort: ~3 × 9min retry budget = ~27min
of dead network, vs infinite in v0.1.3.

#### I.4 — Tracebacks go to `errors.log`, not stdout

The framework's `Researcher failed:`, `Engineer failed:`, etc. errors
used to dump a 30-line Python stack trace to the console every time
something failed. v0.1.4 keeps one line on stdout:

```
[2026-09-25 09:21:48] [ERROR] Researcher failed: APIConnectionError: Connection error
(full traceback: runs/my-run/experiments/my_run/logs/errors.log)
```

The full traceback is in `runs/my-run/experiments/my_run/logs/errors.log`.
Look there when you need the stack.

#### I.5 — Local-model servers auto-detect by URL

asi-evolve now recognizes 6 popular local-model servers by their
default ports — no `--provider` needed:

| Family | Auto-detect URL | Default model |
|---|---|---|
| `ollama` | `localhost:11434` | `llama3.1:8b` |
| `lmstudio` | `localhost:1234` | `local-model` |
| `llamacpp` | `localhost:8080` | `gpt-3.5-turbo` |
| `vllm` | `localhost:8000` | `meta-llama/Llama-3.1-8B-Instruct` |
| `koboldcpp` | `localhost:5001` | `koboldcpp` |
| `textgen` | `localhost:7860` | `text-generation-webui` |

No API key needed — wrapper falls back to `"EMPTY"` automatically.
See section F (provider matrix) and the FAQ entry below for full
per-server recipes.

#### I.6 — Run output dir is no longer inside the vendored package

v0.1.0–v0.1.3 created a symlink inside the installed package's
`experiments/` directory and made log messages say `Code written to
...site-packages/asi_evolve/_vendor/...`. v0.1.4 uses a class-level
override (`_Pipeline._asi_evolve_override_base_dir`) so writes go
directly to your `--output` dir. The vendored tree is never touched.

If you have stale `experiments/<name>` folders inside
`site-packages/asi_evolve/_vendor/ASI_Evolve/experiments/` from old
v0.1.0–v0.1.3 runs, you can delete them — they were old symlinks that
are no longer needed.

#### I.7 — `--researcher-mode {diff,full}` for small local models

The upstream Researcher's default diff mode (emit `<<<<<<< SEARCH /
>>>>>>> REPLACE` blocks applied to the parent code) saves tokens on
large files, but small local models like `qwen2.5-coder:7b` and
`llama-3.1-8b` produce diffs whose search-text rarely matches the
source — every round logs a `Search text not found` warning and forces
a wasteful full-rewrite retry. Use full mode for local models:

```bash
asi-run --provider ollama --model qwen2.5-coder:7b \
        --researcher-mode full ...   # ← skip diff mode for small models
```

Strong cloud models (MiniMax-M3) still default to `diff` and get
better results from it.

#### I.8 — Auto-injected program-interface block in `input.md`

The wrapper now extracts the function signature and every dict-key
access from `initial_program.py` using `ast` and regex, and prepends
a `## Program Interface` section to the problem description. The LLM
sees exactly what fields are valid on the function arguments, so it
stops inventing nonexistent ones (e.g. `gpu_state["latency_budget_ms"]`)
that crash at runtime with `KeyError`.

Before the fix, qwen2.5-coder:7b on the `tiny_ml_v014` scheduler
problem kept producing candidates that referenced
`gpu_state["latency_budget_ms"]` (which doesn't exist), scoring 0.0
every round. With the auto-injected interface block, the LLM knows
exactly what keys are available.

You don't have to do anything — it's automatic. To preview the
injected block, check `runs/<name>/experiments/<name>/input.md`
after a run starts.

#### I.9 — Duplicate-candidate warning

When the LLM emits byte-identical code across consecutive rounds
(small local models at `temperature=0` get stuck in a local minimum
and copy themselves), the wrapper prints:

```
[asi-evolve] round 4: LLM produced a byte-identical candidate
(3 prior copies in DB) (same as prior rounds: Foo V1, Foo V2, Foo V3).
Try --temperature=0.7 or a stronger model.
```

This surfaces the underlying stagnation — from the score alone you
can't tell whether the LLM is iterating or stuck. If you see this,
the LLM is stuck; raise `--temperature` or switch to a stronger model.

#### I.10 — Default temperature 0.3, fixed seed removed

Combined with `temperature=0`, the hard-coded `seed=42` in the config
made small local models produce byte-identical output round after
round. v0.1.4 raises the default temperature to `0.3` (LLM breaks out
of the local min) and removes the fixed seed (API uses fresh
per-call sampling).

Override for full determinism with strong cloud models:

```bash
asi-run --provider minimax --temperature 0.0 ...   # reproducible
```

#### I.11 — HTTP keepalive disabled (kills wedged-socket hangs)

The default `httpx.Client` used by the OpenAI SDK pools TCP
connections and reuses them across calls. On Windows +
long-running sessions a connection can go half-open (firewall idle
timeout, VPN reconnect, sleep/wake) without the kernel noticing —
the next read blocks forever. v0.1.4 passes a custom `httpx.Client`
with `max_keepalive_connections=0, keepalive_expiry=0` so every chat
completion opens a fresh TCP connection. ~30ms handshake per call,
zero wedged-socket hangs.

#### I.12 — Console: zero eval-error noise

Eval crashes (`KeyError`, `NameError`, `SyntaxError`, etc.) used to
dump the full Python traceback to the user's console. v0.1.4
shortens them to one line on stdout and writes the full traceback
to `experiment_dir/logs/errors.log` for grep:

```
[Engineer] Eval failed: KeyError: 'latency_budget_ms'
  → step_1/code:17
```

The full traceback (file/line/frame stack) is in `errors.log`.

#### I.13 — Markdown code-fence stripping

Some LLMs (qwen2.5-coder in full-rewrite mode, in particular) wrap
their entire output in ```` ```python ... ````` fences. The Engineer
wrote that verbatim to the candidate `code` file, and the eval
script crashed on line 1 with `SyntaxError: invalid syntax`. v0.1.4
strips leading/trailing fences (with or without language tag)
before writing.

#### I.14 — `FutureWarning: get_sentence_embedding_dimension`

Modern sentence-transformers renamed the method to
`get_embedding_dimension`. v0.1.4 calls the new name first and
falls back to the old one, so the deprecation warning no longer
pollutes every run console.

#### I.15 — Local-provider silent exit 2 fixed

When `--provider ollama` (or any other local provider) was used
without the matching env var, the CLI exited 2 silently. Root cause:
a stray indent meant `return 2` was inside the outer
`if not args.api_key:` block rather than the inner `else:` (cloud)
branch. Fixed: `return 2` now only fires for cloud providers
without a key. Local providers fall through to `args.api_key =
"EMPTY"`.

#### I.16 — Daemon-thread Ctrl-C handler

The Ctrl-C handler from v0.1.4 was unreliable on Windows + Git Bash:
CPython can't deliver SIGINT while the main thread is blocked in a
C-level `httpx` socket read. v0.1.4's `runner.run_with_interrupt()`
runs the entire runner loop in a **daemon worker thread** so the main
thread stays free to receive SIGINT. Second Ctrl-C calls
`os._exit(130)` (the conventional exit code for SIGINT) so the
interpreter exits immediately, killing the daemon worker with it.

#### I.17 — Plateau doesn't tick on eval crashes

When the LLM produced code that crashed at runtime (e.g. `KeyError`),
the Engineer recorded `score=0.0` and the broken node was added
to the DB. The wrapper then ticked the plateau counter because the
score didn't improve — but the LLM didn't produce a worse solution,
it produced broken code. v0.1.4 detects this case (`score=0.0 +
results.temp.error` non-empty) and skips the plateau tick so the
run budget isn't wasted on a single eval-script bug. You'll see:

```
[asi-evolve] round 2 eval crashed (score=0.0 + error in results);
not ticking plateau
```

#### I.18 — `--cognition` populates a real Cognition store, no stub script

The wrapper reads your `--cognition` file, chunks it on markdown
headings (## / ###) and numbered lists, instantiates the framework's
`Cognition` class against `exp_dir/cognition_data/`, and saves. No
`init_cognition.py` is generated, no `runpy` is invoked, no subprocess
is spawned — it's a direct in-process call.

You see this on every run:

```
[asi-evolve] cognition: populated 7 items from cognition.md
```

Per-round retrieval pulls the 1-3 chunks most relevant to the
analysis text instead of matching-the-whole-doc-or-nothing:

```
[2026-09-25 14:23:11] [INFO] Retrieved 2 cognition items
```

If `cognition.md` has only ONE `#` heading and N numbered rules
(common shape: "1. **Memory is the binding constraint.** ... 2. **Group
by ...** ..."), the numbered-list split kicks in and each rule becomes
its own retrievable chunk.

If `--cognition` is omitted entirely, the wrapper skips cognition
population and the round still proceeds (just no per-round context
beyond what the analyzer wrote).

---

# Examples

The seven examples below are also in the `examples/` directory of the source repository. Each one is a single file you can copy-paste and run. The first five use the bundled `circle_packing_demo` (ships inside the wheel, so no extra setup); the last two work with any user-provided files.

All examples need an API key set as an env var (e.g. `export MINIMAX_API_KEY=sk-...`) unless you're using a local server (Ollama, LM Studio, etc.).

---

## Example 1 — Library import: smallest possible run

Drop the vendored `Pipeline` and call our `Runner` directly. No shell-out, no subprocess.

```python
"""Example 1 — Quickstart: smallest possible run, library import."""
import os, sys
from pathlib import Path
from asi_evolve.runner import Runner, RunOptions


def main() -> int:
    api_key = os.environ.get("MINIMAX_API_KEY", "")
    if not api_key:
        print("Set MINIMAX_API_KEY first.", file=sys.stderr)
        return 2

    # The bundled circle_packing_demo ships inside the installed wheel
    import asi_evolve
    demo = Path(asi_evolve.__file__).parent / "_vendor" / "ASI_Evolve" / "experiments" / "circle_packing_demo"

    options = RunOptions(
        name="example-01-quickstart",
        problem=demo / "input.md",
        initial_program=demo / "initial_program",
        evaluator=demo / "evaluator.py",
        eval_script=demo / "eval.sh",
        cognition=demo / "init_cognition.py",  # optional; the wrapper chunks it
        provider="minimax",
        base_url="https://api.minimax.io/v1",
        api_key=api_key,
        model="MiniMax-M3",
        max_rounds=1,                                # 1 round = quick smoke test
        max_consecutive_connection_failures=3,
        output_dir=Path("./runs/example-01-quickstart"),
    )

    reason = Runner(options).run()
    print(f"[example-01] stopped: {reason}")
    print(f"[example-01] results in: {options.output_dir}")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
```

**Expected console output**:

```
[asi-evolve] starting run 'example-01-quickstart'
[asi-evolve] provider=minimax model=MiniMax-M3
[asi-evolve] cognition: populated 4 items from init_cognition.py
[2026-09-25 14:23:01] [INFO] Starting new experiment: example-01-quickstart
[2026-09-25 14:23:01] [INFO] [Engineer] Eval score: 0.9598
[2026-09-25 14:23:25] [INFO] [Researcher] Generated: lp_optimized_with_multistart_search
[2026-09-25 14:23:31] [INFO] [Engineer] Eval score: 2.2301
[2026-09-25 14:23:32] [INFO] Updated best snapshot: lp_optimized_with_multistart_search (score=2.2301)
[asi-evolve] stopped: max_rounds: completed 1 of 1 rounds
```

---

## Example 2 — CLI run from your own files

The CLI surface for users who already have `problem.md`, `initial_program.py`, and `evaluator.py` on disk. This is the path you'll use 99% of the time.

```python
"""Example 2 — CLI run from your own files (no library import)."""
import os, subprocess, sys, tempfile
from pathlib import Path


PROBLEM_MD = '''# Reverse a string

Given an input string `s`, return a new string with the characters in
reverse order. The function must be named `reverse_string` and accept
one positional argument. Return value is a string.
'''

INITIAL_PY = '''def reverse_string(s):
    """Naive baseline: return the input unchanged."""
    return s
'''

EVALUATOR_PY = '''import json, random, string, sys
from pathlib import Path

def evaluate(code_path):
    namespace = {}
    with open(code_path, "r", encoding="utf-8") as f:
        exec(f.read(), namespace)
    fn = namespace.get("reverse_string")
    if fn is None:
        return {"success": False, "eval_score": 0.0, "error": "no reverse_string"}

    cases = [("hello", "olleh"), ("", ""), ("a", "a"), ("abc", "cba"), ("racecar", "racecar")]
    random.seed(42)
    for _ in range(20):
        n = random.randint(5, 30)
        s = "".join(random.choices(string.ascii_lowercase, k=n))
        cases.append((s, s[::-1]))

    correct = 0
    for s, exp in cases:
        try:
            if fn(s) == exp:
                correct += 1
        except Exception:
            pass
    return {"success": True, "eval_score": correct / len(cases),
            "score": correct / len(cases), "correct": correct, "total": len(cases)}

if __name__ == "__main__":
    result = evaluate(sys.argv[1])
    Path("results.json").write_text(json.dumps(result, indent=2))
    print(f"eval_score: {result.get('eval_score', 0.0):.4f}")
'''

EVAL_SH = '''#!/usr/bin/env bash
# Framework invokes `bash eval.sh` (no args) from a per-step cwd;
# the candidate code is at $PWD/code.
set -uo pipefail
STEP_DIR="$(pwd)"
python "$(dirname "$0")/evaluator.py" "$STEP_DIR/code"
exit 0
'''


def main() -> int:
    if not os.environ.get("MINIMAX_API_KEY"):
        print("Set MINIMAX_API_KEY first.", file=sys.stderr)
        return 2

    workdir = Path(tempfile.mkdtemp(prefix="asi-evolve-ex02-"))
    (workdir / "problem.md").write_text(PROBLEM_MD)
    (workdir / "initial_program.py").write_text(INITIAL_PY)
    (workdir / "evaluator.py").write_text(EVALUATOR_PY)
    (workdir / "eval.sh").write_text(EVAL_SH)
    (workdir / "eval.sh").chmod(0o755)

    cmd = [
        sys.executable, "-m", "asi_evolve",
        "--problem",     str(workdir / "problem.md"),
        "--initial",     str(workdir / "initial_program.py"),
        "--evaluator",   str(workdir / "evaluator.py"),
        "--eval-script", str(workdir / "eval.sh"),
        "--provider",    "minimax",
        "--base-url",    "https://api.minimax.io/v1",
        "--model",       "MiniMax-M3",
        "--max-rounds",  "2",
        "--name",        "example-02-cli",
        "--output",      str(workdir / "runs"),
    ]
    print(f"[example-02] files written to: {workdir}")
    return subprocess.call(cmd)


if __name__ == "__main__":
    raise SystemExit(main())
```

Or directly without Python — once you've written the four files:

```bash
asi-run \
    --problem     problem.md \
    --initial     initial_program.py \
    --evaluator   evaluator.py \
    --eval-script eval.sh \
    --provider    minimax \
    --base-url    https://api.minimax.io/v1 \
    --model       MiniMax-M3 \
    --max-rounds  2 \
    --name        my-run \
    --output      runs/my-run
```

---

## Example 3 — Stop conditions: plateau, max-hours, threshold

asi-evolve has FOUR stop conditions that all fire together (whichever comes first wins):

- `--plateau N` — abort after N consecutive rounds without improvement
- `--max-hours H` — abort after H wall-clock hours
- `--threshold T` — abort as soon as a candidate reaches score T
- `--max-rounds N` — hard cap on rounds (default 50)

```python
"""Example 3 — Stop conditions."""
import os, sys
from pathlib import Path
from asi_evolve.runner import Runner, RunOptions


def main() -> int:
    api_key = os.environ.get("MINIMAX_API_KEY", "")
    if not api_key:
        print("Set MINIMAX_API_KEY first.", file=sys.stderr)
        return 2

    import asi_evolve
    demo = Path(asi_evolve.__file__).parent / "_vendor" / "ASI_Evolve" / "experiments" / "circle_packing_demo"

    options = RunOptions(
        name="example-03-stops",
        problem=demo / "input.md",
        initial_program=demo / "initial_program",
        evaluator=demo / "evaluator.py",
        eval_script=demo / "eval.sh",
        provider="minimax",
        base_url="https://api.minimax.io/v1",
        api_key=api_key,
        model="MiniMax-M3",
        # Stop conditions — first of these to fire wins:
        plateau=2,            # 2 consecutive no-improvement rounds → stop
        max_hours=1.0,        # 1 hour wall clock → stop
        threshold=None,       # (or e.g. 2.5 — stop when score reaches this)
        max_rounds=20,        # hard cap
        max_consecutive_connection_failures=3,
        output_dir=Path("./runs/example-03-stops"),
    )

    reason = Runner(options).run()
    print(f"[example-03] stopped: {reason}")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
```

CLI equivalent:

```bash
asi-run \
    --problem problem.md --initial initial.py --evaluator evaluator.py \
    --eval-script eval.sh --cognition cognition.md \
    --plateau 2 --max-hours 1 --max-rounds 20 \
    --name my-run --output runs/my-run
```

---

## Example 4 — Same loop, four different providers

The provider flag is the only thing that changes between runs. Pick whichever matches your account:

```python
"""Example 4 — Cross-provider."""
import os, sys
from pathlib import Path
from asi_evolve.runner import Runner, RunOptions


# (env-var-name, provider-flag, base-url, model)
PROVIDERS = [
    ("MINIMAX_API_KEY",   "minimax",   "https://api.minimax.io/v1",   "MiniMax-M3"),
    ("OPENAI_API_KEY",    "openai",    "https://api.openai.com/v1",   "gpt-4o"),
    ("ANTHROPIC_API_KEY", "anthropic", "https://api.anthropic.com",  "claude-3-5-sonnet-latest"),
    ("DEEPSEEK_API_KEY",  "deepseek",  "https://api.deepseek.com/v1", "deepseek-chat"),
]


def main() -> int:
    chosen = next(
        ((env, prov, url, mdl) for env, prov, url, mdl in PROVIDERS
         if os.environ.get(env)),
        None,
    )
    if chosen is None:
        print("Set one of:", ", ".join(env for env, *_ in PROVIDERS), file=sys.stderr)
        return 2
    env_name, provider, base_url, model = chosen
    api_key = os.environ[env_name]

    import asi_evolve
    demo = Path(asi_evolve.__file__).parent / "_vendor" / "ASI_Evolve" / "experiments" / "circle_packing_demo"

    options = RunOptions(
        name=f"example-04-{provider}",
        problem=demo / "input.md",
        initial_program=demo / "initial_program",
        evaluator=demo / "evaluator.py",
        eval_script=demo / "eval.sh",
        provider=provider,
        base_url=base_url,
        api_key=api_key,
        model=model,
        max_rounds=1,
        max_consecutive_connection_failures=3,
        output_dir=Path(f"./runs/example-04-{provider}"),
    )

    reason = Runner(options).run()
    print(f"[example-04] {provider}: {reason}")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
```

---

## Example 5 — Local model (Ollama, LM Studio, llama.cpp, vLLM)

asi-evolve auto-detects 6 popular local servers by URL/port. No API key needed — the wrapper falls back to `"EMPTY"`.

- **Ollama** → http://localhost:11434/v1
- **LM Studio** → http://localhost:1234/v1
- **llama.cpp** → http://localhost:8080
- **vLLM** → http://localhost:8000/v1
- **KoboldCpp** → http://localhost:5001/v1
- **text-gen-webui** → http://localhost:7860

For local models, ALWAYS pass:
- `--researcher-mode full` (small models struggle with diff SEARCH/REPLACE)
- `--temperature 0.3` (escapes local-minimum loops)
- `--request-timeout 600` (cold-cache first call can be slow)

```python
"""Example 5 — Local model server."""
import os
from pathlib import Path
from asi_evolve.runner import Runner, RunOptions


def main() -> int:
    import asi_evolve
    demo = Path(asi_evolve.__file__).parent / "_vendor" / "ASI_Evolve" / "experiments" / "circle_packing_demo"

    base_url = os.environ.get("ASI_LOCAL_BASE_URL", "http://localhost:11434/v1")
    model = os.environ.get("ASI_LOCAL_MODEL", "qwen2.5-coder:7b")

    options = RunOptions(
        name="example-05-local",
        problem=demo / "input.md",
        initial_program=demo / "initial_program",
        evaluator=demo / "evaluator.py",
        eval_script=demo / "eval.sh",
        provider="ollama",                       # auto-detected from URL too
        base_url=base_url,
        api_key=os.environ.get("ASI_LOCAL_API_KEY", ""),  # usually empty
        model=model,
        max_rounds=2,
        researcher_mode="full",                  # important for small models
        temperature=0.3,                         # breaks out of local-min loops
        request_timeout=600,                     # first call can be slow
        max_consecutive_connection_failures=5,
        output_dir=Path("./runs/example-05-local"),
    )

    reason = Runner(options).run()
    print(f"[example-05] {model} @ {base_url}: {reason}")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
```

CLI:

```bash
# Terminal 1: ollama serve
ollama pull qwen2.5-coder:7b
ollama serve

# Terminal 2:
asi-run \
    --provider ollama \
    --base-url http://localhost:11434/v1 \
    --model qwen2.5-coder:7b \
    --researcher-mode full \
    --temperature 0.3 \
    --request-timeout 600 \
    --problem problem.md --initial initial.py --evaluator evaluator.py \
    --eval-script eval.sh --cognition cognition.md \
    --max-rounds 4 --max-consecutive-connection-failures 5 \
    --name my-run --output runs/my-run
```

---

## Example 6 — Resilient run on a flaky provider

When your provider is flaky or rate-limited (MiniMax-M3 at peak hours, a local server that OOMs, etc.), three knobs prevent wasting hours on dead network:

- `--inter-call-delay SECONDS` — sleep N seconds AFTER each successful LLM call (separate from per-call retry-delay, which is between RETRIES of the same call). Helps throttle rate limits.
- `--max-consecutive-connection-failures N` — circuit breaker. Stops the run after N consecutive `APIConnectionError` steps. Default 3.
- `--request-timeout SECONDS` — per-call read timeout. Each provider has a default (MiniMax=240s, Anthropic=600s, local=60s) but you can override.

```python
"""Example 6 — Resilient run."""
import os
from pathlib import Path
from asi_evolve.runner import Runner, RunOptions


def main() -> int:
    api_key = os.environ.get("MINIMAX_API_KEY", "")
    if not api_key:
        import sys
        print("Set MINIMAX_API_KEY first.", file=sys.stderr)
        return 2

    import asi_evolve
    demo = Path(asi_evolve.__file__).parent / "_vendor" / "ASI_Evolve" / "experiments" / "circle_packing_demo"

    options = RunOptions(
        name="example-06-resilient",
        problem=demo / "input.md",
        initial_program=demo / "initial_program",
        evaluator=demo / "evaluator.py",
        eval_script=demo / "eval.sh",
        provider="minimax",
        base_url="https://api.minimax.io/v1",
        api_key=api_key,
        model="MiniMax-M3",
        # v0.1.4 resilience knobs:
        inter_call_delay_seconds=3.0,
        max_consecutive_connection_failures=2,
        request_timeout=240,
        # Normal stop conditions:
        max_rounds=30,
        plateau=5,
        max_hours=8.0,
        output_dir=Path("./runs/example-06-resilient"),
    )

    reason = Runner(options).run()
    print(f"[example-06] stopped: {reason}")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
```

---

## Example 7 — Web dashboard for live monitoring

`--serve` spins up a Flask dashboard on `localhost:5000` that polls every 2 seconds and shows the current best score, per-round timeline, and live cognition store contents. The Stop button soft-stops the run.

```python
"""Example 7 — Web dashboard."""
import os
from pathlib import Path
from asi_evolve.runner import Runner, RunOptions


def main() -> int:
    api_key = os.environ.get("MINIMAX_API_KEY", "")
    if not api_key:
        import sys
        print("Set MINIMAX_API_KEY first.", file=sys.stderr)
        return 2

    import asi_evolve
    demo = Path(asi_evolve.__file__).parent / "_vendor" / "ASI_Evolve" / "experiments" / "circle_packing_demo"

    options = RunOptions(
        name="example-07-dashboard",
        problem=demo / "input.md",
        initial_program=demo / "initial_program",
        evaluator=demo / "evaluator.py",
        eval_script=demo / "eval.sh",
        provider="minimax",
        base_url="https://api.minimax.io/v1",
        api_key=api_key,
        model="MiniMax-M3",
        serve=True,                            # enables the dashboard
        port=5000,
        max_rounds=10,
        max_consecutive_connection_failures=3,
        output_dir=Path("./runs/example-07-dashboard"),
    )

    print("[example-07] Dashboard at http://localhost:5000")
    reason = Runner(options).run()
    print(f"[example-07] stopped: {reason}")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
```

CLI:

```bash
asi-run --serve --port 5000 \
    --problem problem.md --initial initial.py --evaluator evaluator.py \
    --eval-script eval.sh --cognition cognition.md \
    --name my-run --output runs/my-run

# Open http://localhost:5000 in a browser
```

---
# FAQ

**Why is `asi_evolve._vendor.ASI_Evolve` private (`_vendor`)?**
Because Python doesn't allow dashes in package names — the upstream's `ASI-Evolve` directory had to become `ASI_Evolve` to be importable, and we wanted to make it loud that users shouldn't import internals.

**Why isn't Anthropic in the choices for `--base-url`?**
Anthropic requires `--provider=anthropic` because its wire format (system prompt as a separate field, `max_tokens` required) is structurally different from OpenAI's. We have an adapter — the flag just tells the wrapper to use it.

**What happens if I delete a previous version on PyPI?**
**Don't.** Once deleted, the filename is blacklisted forever. If you've already shipped a broken version, use the `Yank` button in the PyPI web UI to mark it as withdrawn without breaking the filename namespace.

**The dashboard says "stop_requested: false" forever.**
The polling is 2-second intervals. Click Stop again — if the run is mid-round (a 60-second LLM call), the soft-stop flag only takes effect at the next round boundary.

**Can I run multiple experiments in parallel from the same machine?**
Yes — but each one needs its own `--name` and `--output` directory to avoid colliding. ASI-Evolve's database and cognition store will overwrite each other if they share a directory.

**How do I run a long job against a flaky provider without burning 8 hours on dead network?**
Use the resilience knobs together:

```bash
asi-run \
  --problem problem.md \
  --initial initial_program.py \
  --evaluator evaluator.py \
  --eval-script eval.sh \
  --name my-job \
  --inter-call-delay 3 \
  --max-consecutive-connection-failures 2 \
  --request-timeout 240 \
  --plateau 5 \
  --max-hours 8 \
  --max-rounds 30 \
  --output runs/my-job
```

This will: throttle every successful LLM call by 3 seconds, abort after 2 consecutive connection drops (each call is bounded by `--request-timeout 240` seconds so a single call can't burn more than ~4 minutes), and stop on whichever of {5-round plateau, 8-hour wall clock, 30-round budget} fires first.

**How do I run against a local model (Ollama, LM Studio, llama.cpp, vLLM)?**
asi-evolve auto-detects 6 popular local servers by URL/port. Just point `--base-url` at the running server:

```bash
# Ollama — first `ollama pull qwen2.5-coder:32b` then `ollama serve`
asi-run \
  --problem problem.md --initial initial.py --evaluator eval.py \
  --eval-script eval.sh --cognition cognition.md \
  --name my_run --provider ollama \
  --base-url http://localhost:11434/v1 \
  --model qwen2.5-coder:32b \
  --max-rounds 8 --max-consecutive-connection-failures 5 \
  --output runs/my_run

# LM Studio — start server in GUI on port 1234
asi-run \
  --problem problem.md --initial initial.py --evaluator eval.py \
  --eval-script eval.sh --cognition cognition.md \
  --name my_run --provider lmstudio \
  --base-url http://localhost:1234/v1 \
  --model "" \
  --output runs/my_run

# llama.cpp server — `./server -m model.gguf --port 8080`
asi-run \
  --problem problem.md --initial initial.py --evaluator eval.py \
  --eval-script eval.sh --cognition cognition.md \
  --name my_run --provider llamacpp \
  --base-url http://localhost:8080 \
  --output runs/my_run

# vLLM — `vllm serve meta-llama/Llama-3.1-70B-Instruct --port 8000`
asi-run \
  --problem problem.md --initial initial.py --evaluator eval.py \
  --eval-script eval.sh --cognition cognition.md \
  --name my_run --provider vllm \
  --base-url http://localhost:8000/v1 \
  --model meta-llama/Llama-3.1-70B-Instruct \
  --output runs/my_run
```

No API key needed for any local server — the wrapper falls back to `"EMPTY"` automatically. For local inference, bump `--max-consecutive-connection-failures` to 5+ (local 70B models can take 30-60s per call, longer than cloud).

**Why a venv? Why not `pip install asi-evolve` straight into my system Python?**
If you have one Python on your machine and your `pip install` lands scripts on PATH cleanly, you don't need a venv — just `pip install --upgrade asi-evolve` and go.

But if you have multiple Python installs (system 3.14 + a 3.11 venv + MSYS Python + …), `pip install` tends to land in whichever Python's `pip` happens to be first on PATH. The venv gives you a private Python whose `pip`, `python`, and `Scripts\` are all owned by that venv — `pip install` always lands in the venv's site-packages, the `asi-run.exe` always lands in the venv's `Scripts\` (which is on PATH while the venv is active), and nothing leaks into system Python.

```bash
# One-time per project:
python -m venv .venv
source .venv/bin/activate    # Git Bash / Linux / macOS
.venv\Scripts\activate       # cmd.exe
pip install --upgrade asi-evolve

# Now `asi-run` works:
asi-run --version
# asi-evolve 0.1.4
```

Deactivate with `deactivate` when you're done.

# License

* `LICENSE` — MIT for the wrapper code (everything outside `src/asi_evolve/_vendor/`).
* `LICENSE-VENDORED` — Apache 2.0 for the upstream ASI-Evolve code.
* `NOTICE` — Citation requirements + third-party dependency credits.

Both licenses ship in the wheel. Default install respects both.

# Acknowledgements

* **GAIR-NLP** for the ASI-Evolve framework and the paper.
* **MiniMax** for the underlying model that powers the recommended default provider.
* The original open-source community of evolutionary-search researchers who validated the technique on problems we're only beginning to apply it to.
