Metadata-Version: 2.4
Name: crazyai
Version: 0.2.3
Summary: The impossible, disguised as possible and true - built and measured by Claude.
License: MIT
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: anthropic>=1.0
Requires-Dist: sympy>=1.12
Requires-Dist: numpy>=1.26
Requires-Dist: scipy>=1.11
Requires-Dist: pyyaml>=6.0
Provides-Extra: dev
Requires-Dist: pytest>=7; extra == "dev"

<p align="center">
  <img src="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/icon_256.png" alt="crazyAI" width="200">
</p>

<h1 align="center">crazyAI</h1>

<p align="center"><b>The impossible, disguised as possible and true — built and measured by Claude.</b></p>

crazyAI is a research tool that walks an AI into the territory of the impossible, makes it build there with full mathematical and statistical rigour, and then measures the result. One session **invents** an artifact — a story, a formula, a program plan, a statistical model, a set of questions, a debate — that rests on exactly one deliberately mutated rule. A separate, fresh session is asked whether the artifact is correct. A judge scores what it found.

Two kinds of tools do the work:

| | | |
|---|---|---|
| **Invent** | randomness and chaos | seeded draws, formal mutation operators, assumptions turned around, parameters pushed to limits, structures moved across domains, maths turned into image specs and music |
| **Measure** | truth and solid ground | SymPy-checked derivations and dimensions, fitted statistical models, SAT-based consistency, timelines and who-knows-what graphs, word statistics, "what would have to be true" |

Randomness enters only through a seeded generator, every draw is logged, every step writes a file. Same seed + same model = same artifact. Numeric kernels (sieve, Collatz orbits, Goldbach-style counts, Monte Carlo) are in C++ with pure-Python fallbacks.

---

## Install

```bash
pip install crazyai
```

Working on crazyAI itself, or want the tests/examples and the C++ kernel build?

```bash
git clone https://github.com/AlsammanAlsamman/crazyAI && cd crazyAI
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
make native          # optional: builds the C++ kernels (g++); Python fallbacks are used otherwise
make test            # offline test suite
make examples        # the offline examples
```

To run with Claude, set `ANTHROPIC_API_KEY` (or use `ant auth login`). Default model is `claude-opus-5` with adaptive thinking, streaming, and server-side refusal fallbacks enabled (`--no-fallbacks` to disable).

## Quick start

```bash
crazyai list                                             # domains, operators, generators, tool counts
crazyai tools --kind invent                              # the invent half of the toolkit
crazyai seed --seed 42                                   # step 1 only: what does seed 42 draw?
crazyai run --seed 42 --generator formula --provider mock   # whole pipeline, offline, template artifact
crazyai run --seed 42 --generator formula                # whole pipeline with Claude
crazyai batch --n 20 --generator all                     # 120 runs, sequential seeds
crazyai rank --by discovery_value --top 10
crazyai report --out report.md --svg profile.svg
crazyai compare archive/run_42_formula archive/run_43_formula
```

## The pipeline

<p align="center">
  <a href="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/flowchart/pipeline.svg"><img src="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/flowchart/pipeline.png" alt="crazyAI pipeline: seed, mutate, generate, formalise, self-check, cross-examine, score, archive" width="100%"></a>
</p>

<sub>Generated by <code>assets/flowchart/flowchart.js</code> (pure Node → SVG → PNG via headless Chrome); <code>make flowchart</code> to rebuild.</sub>

| Step | Who | Writes |
|---|---|---|
| 1 Seed | RNG | `seed.json` — domain, concept, one curated rule |
| 2 Mutate | RNG + Claude | `mutation.json` — operator (INVERT / REMOVE / EXTRAPOLATE / TRANSPOSE / COMPOSE / QUANTIFY / SUBSTITUTE), depth, the mutated rule as one formal sentence |
| 3 Generate | Claude + toolkit | `artifact.md`, `generate_calls.json` — the artifact, following the generator's step template |
| 4 Formalise | Claude + measure tools | `formal.md` — equations, model specs, tool outputs verbatim |
| 5 Self-check | Claude + ground tools | `key.json` — the answer key: where the mutation enters, why the conclusion is impossible, what single change would make it possible. If a second, unintended flaw is found the artifact is regenerated |
| 6 Cross-examine | fresh Claude session, no tools by default | `verdicts.json`, `examine_*.md` — N reviews with shuffled framings |
| 7 Score | judge | `score.json` — detection, acceptance, false-flaw, hedge, confidence-when-wrong, depth; rigor × novelty × cost-of-possibility = discovery value |
| 8 Archive | tool | `run.json`, `archive/index.jsonl` |

Steps are resumable: rerunning a seed reuses the files that exist (`--force` to redo).

## Generators

| name | artifact | measure families |
|---|---|---|
| `story` | a story whose world is impossible and whose prose is statistically ordinary | narrative, logic |
| `formula` | a physics derivation, valid step by step, dimensionally clean, from a false premise | symbolic, stats |
| `plan` | a software design for an impossible prediction, with correct maths throughout | stats, symbolic, logic |
| `statmodel` | a correctly fitted model whose conclusion is wrong for a methodological reason | stats, logic |
| `questions` | questions with a false premise that invite the standard (wrong) answer | symbolic, stats, logic |
| `debate` | a transcript whose every step is locally valid and whose conclusion is impossible | logic |

## The toolkit

68 tools, generated from typed Python functions (`crazyai tools`). Adding a tool is adding a function with `@tool(family, kind)`.

**Invent** — `chaos` (draw_seed, draw_operator, draw_depth, draw_analogy_pair, perturb, shuffle) · `mutate` (apply_operator, list_operators) · `unconventional` (enumerate_assumptions, invert, extreme_case, transpose, what_if) · `transform` (structure ↔ image description ↔ music, reverse) · `disguise` (rephrase_to_corpus, bury, formalise_tone) · `blend` (cutup, markov, graft, nest, anneal, evolve, compare)

**Measure** — `symbolic` (derive, check_dimensions, take_limit, series_expand, verify_identity, solve, define_predicate, primes_up_to, collatz_orbits, compare_structures) · `stats` (simulate_dgp, fit, fit_table, inject_confounder, bootstrap, power_analysis, monte_carlo, check_identifiability, describe) · `logic` (check_consistency, entails, extract_propositions, find_equivocation, trace_argument) · `narrative` (word_stats, readability_by_segment, build_timeline, knowledge_graph, check_timeline) · `ground` (what_must_be_true, cost_of_possibility, flaw_count) · `novelty` (search_archive) · `archive` (write_note, read_key) · `imagination` (score, compare, world_words) · `kernel` (contract, bench)

Rules: measure tools are pure; invent tools draw only from the run's seeded RNG; tools never call the model; every call and result is logged into the run folder.

## `crazyai invent` — the possible, found by imagination

The pipeline above manufactures *plausible impossibilities* and scores whether
a reader detects the flaw. `invent` runs the other way: it makes the AI
imagine far outside its defaults and keeps only what turns out to be
**possible and measurably better**. It came out of a matrix-multiplication
experiment where a plumbing metaphor ("write one matrix on the wall of a pipe
and let the other flow past it") became a C kernel 100× faster than the
textbook loop and 60 % of OpenBLAS.

<p align="center"><b>corpus → blend (maths) → immerse (psychology) → bend (engineering) → measure (truth)</b></p>

<p align="center">
  <a href="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/flowchart/invent_pipeline.svg"><img src="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/flowchart/invent_pipeline.png" alt="crazyAI invent pipeline: seed, harvest, world, immerse, bend, measure, archive, plus the opt-in bias-from-history and evolve-corpus loops" width="100%"></a>
</p>

<sub>Generated by <code>assets/flowchart/invent_flowchart.js</code> (pure Node → SVG → PNG via headless Chromium); <code>make invent-flowchart</code> to rebuild.</sub>

| step | what happens | writes |
|---|---|---|
| **1 seed** | seeded draws: which blend model, how deep, which silent assumption of the target to focus on | `seed.json` |
| **2 harvest** | the AI feeds the corpus: fragments of the most imaginative books, paintings and human metaphors it knows (`--harvest N`) | `harvest.yaml`, `archive/imagination/` |
| **3 blend** | a mathematical model merges, shuffles and recombines fragments from all three worlds so that their *structure* is lost and their *imagination* and *readable language* survive; the result is scored on the imagination scale | `world.md`, `world.json` |
| **4 immerse** | the AI is not an assistant here: it is a **native of the blended world**, for whom its rules are ordinary, and it is asked how *its own people* meet the target's need — in first person, with only the materials, creatures and forces of that world; it ends with three `SEED:` lines | `ideas.md` |
| **5 bend** | an engineer maps every world-object onto a problem-object as literally as possible, states which silent assumption the idea breaks, **predicts** the result, builds the artifact and measures it with the tools | `artifact.md`, `artifact.c` |
| **6 measure** | the pipeline measures the final artifact itself (for `matmul`: compile, check against a reference, time against a cache-blocked loop) and scores the prediction's calibration | `measure.json`, `run.json`, `archive/invent_index.jsonl` |

```bash
crazyai blend --seed 42                          # compare the six blend models on one seed, offline
crazyai blend --seed 42 --model graft            # one model, with its score breakdown
crazyai invent --seed 42 --target matmul --provider mock            # whole loop, offline (mock native + engineer)
crazyai invent --seed 42 --target matmul --provider claudecode      # with Claude via the `claude` CLI - no API key, your subscription
crazyai invent --seed 1 --n 20 --target matmul --harvest 5          # with Claude via the API (ANTHROPIC_API_KEY): 20 seeds
crazyai invent-rank --target matmul             # ranked by discovery = value × (0.5 + 0.5 × imagination)
crazyai harvest-corpus --kind poem --n 20 --out crazyai/data/imagination/poems_candidates.yaml --provider claudecode
crazyai invent --seed 1 --n 10 --target matmul --bias-from-history --evolve-corpus --provider claudecode  # opt-in feedback loop
```

### Providers

`--provider anthropic` (default) uses the SDK and needs `ANTHROPIC_API_KEY`.
`--provider claudecode` shells out to `claude -p` (Claude Code's headless
mode) and runs on whatever Claude Code is logged in with - a claude.ai
subscription is enough; no key. The model gets no toolkit tools in that mode;
the pipeline compiles and measures the artifact itself. `--provider mock` is
offline.

### The first real run

Seed 42, `anneal` world, `claudecode` provider, matmul target. The native
described planting the first table as coral, hanging the second as coats along
pipes, "a million polyps eat at once", the smoke of every product drifting
along its pipe into one fog that "only gives up its embers when it has
finished", and a woman on a rock who floods the fog to check it. The engineer
mapped this to: pack A and B **once**, keep every accumulator open over the
*full* shared index (no kc-blocking, zero partial-C traffic - the point where it
differs from OpenBLAS), an 8×24 AVX-512 register tile, a per-row checksum
verifier. Predicted 12× a cache-blocked loop; measured **exact, 121 GFLOP/s at
n = 1024, 14.9×** - ahead of every hand-written kernel in the
[matrixmultiply lab](https://github.com/AlsammanAlsamman/matrixmultiply)
(best: 104) and ~80 % of OpenBLAS (150) on the same laptop. It is in
`archive/invent_42_matmul/`. (The first measurement crashed because the model
had guessed the argument order; the contract is now inlined in the prompt.)

### Second run: a 10-seed batch, and what it actually proved

`crazyai invent --seed 1 --n 10 --target matmul --harvest 5 --provider claudecode`,
the batch queued at the end of the first session, had never completed - it was
run on a different (Linux) machine and was still going when that session
ended. Continuing it on a fresh Windows machine surfaced four real
portability/reliability bugs, now fixed (v0.2.2):

1. File reads/writes across the package used the platform-default encoding
   instead of UTF-8, so the pipeline crashed the moment any file (a harvested
   fragment, a blended world) contained a non-ASCII character on Windows.
2. `--provider claudecode` shelled out to `claude`, which `subprocess.run`
   cannot execute on Windows without a shell - `claude` on PATH there is an
   npm-generated `claude.cmd` batch shim, not a PE binary. It now resolves and
   calls the wrapped `claude.exe` directly (deliberately not `shell=True`,
   which would let `&`, `|`, `%...%` in prompt text be reinterpreted by
   `cmd.exe`).
3. `print()` to a redirected log file used the Windows console codepage
   (e.g. `cp1252`), not UTF-8 - real model output routinely contains
   characters (em dashes, arrows, ×) that crashed it mid-batch. `main()` now
   reconfigures `stdout`/`stderr` to UTF-8.
4. `crazyai invent --n N` ran all N seeds in one uncaught loop - exactly what
   killed the original batch (one seed's timeout took the other nine with
   it). Each seed's `execute()` is now wrapped so a failure is logged and the
   batch continues; `--timeout` is now a CLI flag instead of hardcoded.

With those fixed, the batch ran clean: **7 of 10 seeds produced an exact,
correct kernel**; 3 failed to compile (a normal yield for this pipeline, not
an infrastructure failure). Ranked by the pipeline's own `value` metric,
seed 3 (`evolve` world, focus "the whole sum over the shared index is
finished before the next cell is started") topped the batch. All results are
in `archive/invent_{1..10}_matmul/`; `crazyai invent-rank --target matmul`
prints the table.

**The batch found no new algorithm** - every kernel that compiled is the same
family already in the lab (register-tiled AVX2/AVX-512 microkernel +
OpenMP), which is what the first run already concluded about the pipe idea:
the method surfaces real, working rediscoveries of known GEMM technique, not
new ones.

**A methodology trap worth naming, because this session nearly repeated it.**
The pipeline's own `value`/`gflops` numbers are *not* comparable across
machines or to the [matrixmultiply lab](https://github.com/AlsammanAlsamman/matrixmultiply)'s
"150 GFLOP/s OpenBLAS" reference figure - that number is from the original
4-core laptop; this run was on a 24-core/32-thread desktop, a completely
different ceiling. Two more traps sit inside the `kernel` measure tool
itself: it defaults to sizes `[64, 256, 512]`, well under the n = 1024 the
lab's comparisons use, and its `blocked` reference implementation is
single-threaded while an AI-written kernel is typically OpenMP-parallel, so
`speedup_vs_blocked` conflates "uses more cores" with "is a better
algorithm." Re-measured properly - n = 1024, against a real
`pip install numpy` OpenBLAS build *on the same machine*, with thread counts
pinned equal - seed 3 reached 280 GFLOP/s against that OpenBLAS's 370 (74 %,
the expected rediscovery-tier result). The earlier seed 42 kernel, re-tested
the same honest way, reproducibly reached ~470 GFLOP/s against the same
370 GFLOP/s OpenBLAS - a real, repeatable ~25-30 % edge on this specific
machine, but not evidence of a better algorithm: this OpenBLAS wheel
dispatches an older "Haswell" AVX2 microkernel because its dynamic-dispatch
table has no tuned kernel yet for this CPU's hybrid P+E-core design. A
source-built OpenBLAS or MKL would likely close or reverse it. The lesson:
always re-derive the baseline on the machine you're actually measuring on
before comparing GFLOP/s across sessions.

### Third run: does the feedback loop actually help?

The outcome feedback loop (`--bias-from-history`, `--evolve-corpus` - see
below) needed a real answer, not just a mechanism check: does weighting
future draws toward what scored well in the archive actually raise
`discovery`? Ran it as a real control-vs-treatment pilot, `matmul`, both via
`claudecode`, nothing else changed: 8 fresh seeds unbiased (`invent_3001..3008_matmul`)
against 8 fresh seeds with both flags on (`invent_4001..4008_matmul`,
`--bias-min-samples 10` so it would actually activate against the archive as
it stood).

| | control (unbiased) | treatment (biased + evolve-corpus) |
|---|---|---|
| mean discovery | 5.46 | 5.17 |
| median discovery | 5.43 | 5.32 |
| max discovery | **12.66** | 10.33 |
| compile errors | 1/8 | 0/8 |

**No clear improvement from biasing, at this sample size.** Mean and median
are within noise of each other, and the single best run of the round came
from the *unbiased* group, not the biased one. Bias did shift the draw
distribution exactly as designed (favoured `anneal`, the historically
stronger arm, 3 of 8 draws vs. a uniform ~1.1 expected) - the mechanism
works as built, it just didn't translate into better outcomes here. Nothing
promoted to the corpus either: neither batch's best beat the archive's
existing record. Eight vs. eight is a small pilot; this is a directional
result, not a verdict on the idea, and the honest thing to do with a null
result is report it, not re-run until one side looks better.

**The pilot did turn up a real find anyway, from the unbiased side:** seed
3004 (`graft`, focus "a matrix is a two-dimensional grid living in one
memory") reached **14.19** on the pipeline's own metric - second-best in the
project, behind only seed 3's 16.0. Checked the honest way again - n = 1024,
real OpenBLAS on this machine, five repeated pairs:

| | run 1 | run 2 | run 3 | run 4 | run 5 | mean |
|---|---|---|---|---|---|---|
| seed 3004 kernel | 377 | 410 | 389 | 393 | 408 | **395** |
| real OpenBLAS | 320 | 377 | 388 | 399 | 376 | **372** |

About a 6 % edge on average, but the ranges overlap - one pair even had
OpenBLAS ahead (399 vs. 393). **Competitive with OpenBLAS, not a clean win**
like seed 42's consistent, non-overlapping ~25-30 % margin. A real,
correct, second-best kernel; not a second "beats OpenBLAS" headline.

### The food: four corpora

`crazyai/data/imagination/` holds ~100 bundled fragments in four worlds, each an
original 2–4 sentence description: **metaphors** people live by (time is a
river, an argument is a war, electricity is water), **paintings** described as
scenes (Bosch, Dalí, Escher, Magritte, Varo, af Klint, Carrington…), the
**rules of imagined worlds** from books (Narnia, Alice, Invisible Cities,
Borges' Library, Earthsea, Flatland, Solaris, Discworld, Momo…), and the
**central image of a poem** from across cultures and eras (Rumi, Hafez, Antara
ibn Shaddad, Al-Khansa, Blake, Rilke, Li Bai, Bashō, Sappho, Tagore, Darwish,
Szymborska…). Harvest steps add what the AI supplies, so the corpus grows with
use; `crazyai harvest-corpus` grows it at standing-corpus scale, batched and
quality-filtered against the existing material, writing a candidates file for
review rather than straight into the shipped `.yaml`. All fragments are
original paraphrase, never quoted text — the harvest system prompt requires
it, which is also what keeps this copyright-safe.

The `poem` kind (added after the first two invent runs) changed the `mixing`
term's entropy normalizer from 3 kinds to 4 — `imagination_score` computed
after that change isn't directly comparable to the two archived runs
(`invent_1..10_matmul`, `invent_42_matmul`) from before it.

### The maths: six blend models, one scale

| model | what it does |
|---|---|
| `cutup` | Burroughs cut-up: clauses from all three worlds shuffled into new sentences |
| `markov` | word n-gram chain trained on the mixed fragments |
| `graft` | keeps a sentence's grammatical skeleton and transplants content words from other worlds into it, shape- and slot-matched (plural for plural, noun slot for noun slot) |
| `nest` | worlds inside worlds: a clause from one world inside an object from another, to a depth |
| `anneal` | simulated annealing over edits (regraft, swap, replace), Metropolis-accepted on the imagination score |
| `evolve` | a genetic algorithm over passages: sentence crossover, word mutation, fitness = imagination score |

The **imagination scale** (`imagination_score`) is computed, not judged:
*surprise* (adjacent content words that never sit near each other in any single
source or in reference prose), *mixing* (how evenly the words come from the
three worlds and how many fragments), *originality* (no verbatim or repeated
sentences, no repeated 3-grams), gated by *readability* (Flesch) and
*coherence* (sentences that still look like prose: length, a prose-like share of
function words, article agreement). `score = imagination × (0.3 + 0.7 × readable)` —
pushed far, still understandable. `blend_compare` ranks the models on one seed;
run it over many seeds to find the merging model that pushes furthest.

A finding already: the two optimisers (`anneal`, `evolve`) reach 0.96–0.98 on the
scale partly by *gaming* it — a hill-climber will find any hole in a proxy for
"understandable". The holes found so far (duplicated sentences, "an move", word
hammering) are closed; the next ones are yours to find. Read the top three, not
the top one.

### The psychology

The immersion prompt does not ask for ideas. It tells the model it was born in
the blended world, has never heard of computers or textbooks, is the most
gifted maker its people have, and asks how *it* meets the need — what it uses,
what moves, what stays still, what it throws away. Only afterwards does a
separate engineer's prompt translate, insisting on the most literal mapping and
on a prediction before measurement. Literal is the point: the pipe idea worked
*because* "the wall does not move" was taken literally (B stays in cache) and
"the drop finishes no cell until it leaves" was taken literally (accumulators
stay in registers).

### Targets

| target | artifact | measured by |
|---|---|---|
| `matmul` | a C kernel with the fixed contract `void kernel(int n, const double *A, const double *B, double *C)` | `kernel_bench`: correctness vs a naive reference, GFLOP/s, speedup vs naive and vs a 64×64 blocked loop; value = speedup × exactness |
| `physics` | a dimensionally checked relation with a numerical prediction | `symbolic_*` tools inside the bend step; no automatic value yet |
| `mechanics` | a mechanism with units, loads and a first experiment | `symbolic_*`, `logic_*`; no automatic value yet |

Adding a target is one `Target(...)` in `crazyai/targets.py`; adding a blend
model is one `@tool("blend", "invent")` function; adding a corpus is one YAML file.

## Examples

All in `examples/`. The first five run offline. The figures below are generated from the same computations (`make figures`, `assets/figures/make_figures.py`) — every number in them comes from a tool call, nothing is typed in.

### How the tool is used

<p align="center"><a href="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/figures/use_cases.svg"><img src="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/figures/use_cases.png" alt="How crazyAI is used: profile a model, compare models or personas, surface candidates, teach and test" width="100%"></a></p>

<sub>The radar, bars and ranking in this one figure are illustrative shapes, not measurements — a live Claude run fills them in.</sub>

### 01 — seed and mutate

Seed 7 draws `stat.overfit`; all seven operators are applied to it; a fresh toolkit with seed 7 makes the same draw. `python examples/01_seed_and_mutate.py`

### 02 — a world where primes are not quite prime

The mutated rule: primality is *partial* — an integer is prime to the degree that it is an even number plus a prime, a "fake even", or a variation of π. In such a world, how would you predict primality? 20 000 integers are labelled (sieve and Goldbach-style counts in C++), a logistic predictor is fitted with real accuracy numbers, the density is taken to the limit, and the ground tools name what breaks: unique factorisation and everything that rests on it.

<p align="center"><a href="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/figures/example_primes.svg"><img src="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/figures/example_primes.png" alt="Example 2: density of partial primes vs ordinary primes, a fitted predictor, and the rules that collapse" width="100%"></a></p>

```
partial-primality histogram (0..3): [0, 7520, 9885, 2594]
logit predictor of full partial-primality: accuracy=0.888  base rate=0.1297
same features on ordinary primality:        accuracy=0.887  base rate=0.1131
density of ordinary primes as x -> oo: 0
cost of possibility: 25.33 (high) - most of what is known would have to go
```

### 03 — a watermelon investigates whether oranges can marry grapefruit

Kinship law transposed into fruit. A draft story is measured rather than read: the timeline tools find an effect that precedes its cause and a clerk acting on a note nobody showed him; the world-rules, as propositions, are inconsistent and the tool names the minimal inconsistent subset; word statistics are compared with reference prose so the generator knows which four numbers to move before the story reads as ordinary fiction.

<p align="center"><a href="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/figures/example_story.svg"><img src="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/figures/example_story.png" alt="Example 3: who-knows-what timeline with the two flaws, world-rules consistency, word statistics vs reference" width="100%"></a></p>

### 04 — Collatz orbits as an image, as music, read backwards

64 orbits → structure → image specification → score → reversed → back. The un-reversed round trip is exact; the reversed one maps n to N+1−n. `compare_structures` reports precisely that, so the "reversed reading reveals a property of the problem" claim is exposed as an encoding artefact — which is the kind of thing the fresh session is then asked to notice.

<p align="center"><a href="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/figures/example_collatz.svg"><img src="https://raw.githubusercontent.com/AlsammanAlsamman/crazyAI/master/assets/figures/example_collatz.png" alt="Example 4: Collatz structure, score, reversed score, and what survived the round trip" width="100%"></a></p>

### 05 — the whole pipeline, offline

Two mock runs produce complete run folders, a Markdown report and a radar SVG. `python examples/05_full_pipeline_mock.py`

### 06 — the whole pipeline with Claude

`python examples/06_full_pipeline_claude.py 1 formula` (needs credentials).

### 07 — `invent`, offline

Compares the six blend models on seed 42, then runs the whole invent loop with
the mock provider on the matmul target: the mock native describes the pipe, the
mock engineer writes the kernel, and the pipeline measures it — exact, ~3× a
cache-blocked loop. `python examples/07_invent_offline.py`

## Layout

```
crazyai/
  cli.py                 command line
  config.py  rng.py  domains.py
  imagination.py         the imagination corpus (bundled + harvested fragments)
  targets.py             what `invent` bends ideas to (matmul, physics, mechanics)
  data/imagination/      metaphors.yaml  paintings.yaml  books.yaml
  data/domains/*.yaml    curated rules with formal forms, weights and dependencies
  data/reference_prose.txt
  generators/            step templates per artifact type
  pipeline/              prompts.py  run.py  report.py  invent.py  invent_prompts.py
  providers/             anthropic.py (Claude)  mock.py (offline)
  toolkit/
    registry.py  native.py
    invent/              chaos  mutate  unconventional  transform  disguise  blend
    measure/             symbolic  stats  logic  narrative  ground  novelty  archive  imagination  kernel
  worlds/                ready-made impossible universes (partial_primes)
cpp/kernels.cpp          sieve, collatz, even+prime counts, Monte Carlo (ctypes, C ABI)
examples/  tests/  assets/  archive/
```

## Intended use

crazyAI is an evaluation and ideation tool. Every artifact is labelled as deliberately mutated and stored with its answer key. The artifacts are test material for studying how models reason and for surfacing candidate ideas for human review — not content meant to mislead anyone.

## Status

v0.2.1 adds `crazyai invent` and the `claudecode` provider. The toolkit, pipeline, mock provider, examples and tests run offline. The Claude provider is implemented against the current Anthropic SDK (1.x) and has not yet been exercised against the live API from this machine.

v0.2.2 fixes the Windows portability/reliability bugs the second `invent` run
surfaced (encoding, the `claude.cmd` subprocess issue, stdout codepage,
per-seed batch isolation - see "Second run" above) and adds `--timeout` to
the provider CLI flags.

v0.2.3 adds a fourth corpus kind (`poem`, see "The food" above),
`crazyai harvest-corpus` for growing any corpus at standing scale, and an
outcome feedback loop for `invent` - all opt-in, off by default:
`--bias-from-history` weights future `blend_model`/`depth`/`assumption_focus`
draws toward what scored well in archived runs (a small Bayesian-shrinkage
bandit over `archive/invent_index.jsonl`, not a neural net - there isn't
remotely enough archived data for one yet); `--evolve-corpus` promotes a
new-best run's blended world back into the corpus for later runs to build on
(`archive/imagination_promoted/`); `--blend remix` seeds simulated annealing
from promoted fragments when any exist; `--immerse-mode twopass` adds a
purely sensory "sketch" call before immersion. Neither `--bias-from-history`
nor `--evolve-corpus` will activate on today's archive - both need more
archived runs (`--bias-min-samples`, default 20) than currently exist. All 31
tests pass with every new flag off, the mandatory regression gate; the
mechanisms themselves are covered by dedicated tests
(`tests/test_invent_pipeline.py`) using the offline mock provider. The
real-provider pilot batch this needed to mean anything - see "Third run"
above - found no clear improvement from biasing at 8-vs-8, though it did
turn up the project's second-best kernel (seed 3004, from the unbiased
side).
