Metadata-Version: 2.4
Name: isynth-ai-data-config
Version: 0.8.1
Summary: Generate iSynth data_config test-data definitions from a plain-language request, using an LLM that reads the environment's own metadata.
Author-email: Markus Herrmann <markus.herrmann@itopia.ch>
License-Expression: LicenseRef-Proprietary
Project-URL: Homepage, https://isynth.io
Keywords: isynth,mcp,llm,data-config,test-data,synthetic-data
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.13
Description-Content-Type: text/markdown
Requires-Dist: litellm>=1.0
Requires-Dist: fastmcp>=4.0
Requires-Dist: openpyxl>=3.0
Requires-Dist: python-box>=7.0
Requires-Dist: python-dotenv>=1.0
Requires-Dist: httpx>=0.28
Requires-Dist: PyYAML>=6.0
Requires-Dist: starlette>=1.0
Requires-Dist: pytest>=7.0
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: build; extra == "dev"
Requires-Dist: twine; extra == "dev"

<!-- Generated from public_docs/ by `make readme`. Edit those pages, not this file. -->

# isynth-ai-data-config

Generate iSynth `data_config` files from a request written in plain language, instead
of hand-writing YAML against a data model you have to memorise first.

You describe the test data you want — *"a private client with one account, no
portfolio, no phone number"* — and you get back a ready `data_config` that iSynth can
execute:

```yaml
simple_client:
  based_on: banking.individual_client
  number_of_copies: 1
  parameters:
    number_of_accounts: 1
    number_of_portfolios: 0
    number_of_phone_numbers: 0
    number_of_email_addresses: 0
  composers: []
```

An LLM writes that file, but it does not guess the vocabulary. `banking.individual_client`,
`number_of_accounts` and every other name in the output come from **your** environment's
own metadata, which the model reads through a set of tools while it works. The draft is
then validated against that same metadata, and anything that does not check out goes
back to the model for correction before you ever see it.

## Where to start

- **[Getting started](#getting-started)** — the iSynth vocabulary, how it works,
  requirements, installing, and a first call from Python.
- **[Configuration](#configuration)** — choosing the model, the LLM server
  settings, API keys, and GitHub Copilot.
- **[Integrating into an env](#integrating-into-an-env)** — the plugin wrappers,
  the composer order, the optional MCP server service, and checking that it all works.
- **[Using the package in several envs](#using-the-package-in-several-envs)** — how the package finds its
  env, what is per env and what is shared, and what each env needs.
- **[Installing by copying the folder](#installing-by-copying-the-folder)** — the alternative to
  `pip install`: copying the package folder into an env.
- **[Troubleshooting](#troubleshooting)** — symptoms and their causes.

If you are setting the package up for the first time, read them in the order above.

These pages also ship in the source distribution
(`isynth_ai_data_config-<version>.tar.gz`), in `public_docs/`. The package is published on pypi.org; its licence is
proprietary.

## Getting started

This guide takes you from an iSynth env to a first generated `data_config`.

### Some iSynth vocabulary first

If you have not worked with iSynth: it is a platform for generating synthetic test
data. A few terms show up throughout this documentation.

| Term | What it means |
|---|---|
| **env** (environment) | One iSynth project — its own data model, its own directory on disk. This package is installed per env. |
| **object type** | An entity in that model: `NaturalPerson`, `Account`, `PostalAddress`. |
| **constellation template** | A reusable bundle of objects that belong together — a client with their accounts and addresses. It is the `based_on` value of a case. |
| **composer** | A building block that adds objects to a constellation, in a defined order. |
| **data_config** | The YAML file this package produces: one or more named cases, each saying what to build. iSynth executes it to create the actual data. |

Everything above already exists in your env before this package is installed. The
package only reads it.

### How it works

The model is not given a dump of your data model. It gets 14 tools and fetches only
what it needs:

| Tool group | Examples |
|---|---|
| Object types | `list_object_types`, `get_object_type` |
| Composers | `list_composers`, `get_composer`, `get_composer_order` |
| Constellation templates | `list_constellation_templates`, `get_constellation_template` |
| Enums and functions | `list_enums`, `get_enum`, `list_functions`, `get_function` |
| Curated examples | `list_ai_training_examples`, `get_ai_training_example` |
| Validation | `validate_data_config` |

Those tools are plain Python functions in one registry. Generation calls them
**in-process** — no HTTP, no running server. The same registry is also served over MCP
(streamable-http) for external clients such as IDE agents; that is optional, see
[*MCP server service*](#3-mcp-server-service-optional).

LLM access goes through [LiteLLM](https://www.litellm.ai/), so a local model in LM
Studio and a hosted one are configured the same way. The model must support **native
tool calling**.

### Requirements

- **Python 3.13** or newer.
- **An iSynth env**, with these already in place:

  | Path | Where it comes from |
  |---|---|
  | `__generated__/*.json` | the iSynth action **Initialize environment** (object types, composers, enums, functions) |
  | `constellations/__generated__/**/*.json` | generated from the env's constellation modules |
  | `data_configs/ai_training/` | optional — without it there are simply no training examples |
  | `__generated__/composer_sequence.txt` | written once by the restart plugin, see [*Composer order*](#2-composer-order-once) |

  Generation works without training examples, but noticeably worse: they are what the
  model orients itself on.

- **A tool-calling LLM** — local (LM Studio) or hosted. It must handle *native*
  function calling over several turns: choose a tool, read the result, choose the next,
  produce valid JSON arguments, and finally answer with a single JSON object. Prompt-
  engineered tool use is not enough. See [*LLM settings*](#llm-settings).

### Install

```
pip install isynth-ai-data-config
```

Or add `isynth-ai-data-config` to your env's `requirements.txt` and install
that as usual.

Installing gives you the library. It does **not** yet put anything in the iSynth UI —
that needs the plugin wrappers from [Integrating into an env](#integrating-into-an-env).
If you only want the MCP tool server, or want to call `generate_data_order()` from your
own code, you can stop after [Configuration](#configuration).

#### Alternative: copy the folder

The package can also be copied wholesale into an env instead of installed — the way it
was distributed before it was published, still supported for an env that cannot reach
the private index. The steps (copying the folder, including its `requirements.txt` from
the env's own list, the Docker build context and the `.dockerignore` that keeps it
small) are in [Installing by copying the folder](#installing-by-copying-the-folder).

### Using it from Python

```python
from ai_data_config.llm_helper import generate_data_order

data_config = generate_data_order("a private client with one account, no portfolio")
```

`generate_data_order(prompt, content=None)` builds the messages, runs the tool-calling
loop, then validates and orders the result. `content` is an existing `data_config` to
modify rather than start from scratch. It returns a dict, ready to be dumped as YAML.

This needs no running MCP server and no iSynth runtime — only the env's generated
metadata on disk, found via `ENV_HOME`. iSynth sets that variable for its own
processes; elsewhere (a `docker exec` shell, a script) set it yourself, or run from the
env's directory, which is the fallback.

### Next steps

- [Configuration](#configuration) — pick the model, set API keys.
- [Integrating into an env](#integrating-into-an-env) — put the package into the iSynth UI.

## Configuration

### Model

Defaults ship in the package's own `settings.yml`. To deviate, repeat the same key in
your **env's** `settings.yml` — the env wins:

```yaml
LLM_MODEL: "lm_studio/qwen3.8-27b-gguf"
LLM_API_BASE: "http://192.168.1.100:1234/v1"
```

`${VARIABLE_NAME}` substitution works as in any iSynth `settings.yml`. You never edit a
file inside the package, which is what keeps it upgradable.

Only keys the package declares are read from the env; an `LLM_*`/`AI_*` key with a typo
is reported as a warning naming the key. The ones you are most likely to touch:

| Key | Default | What it does |
|---|---|---|
| `LLM_MODEL` | `openai/gpt-5.6-sol` | any LiteLLM model identifier |
| `LLM_API_BASE` | — | for local or self-hosted endpoints |
| `LLM_REQUEST_TIMEOUT` | `1800` | seconds; local models can be slow |
| `LLM_MAX_OUTPUT_TOKENS` | `10000` | |
| `AI_MAX_TOOL_TURNS` | `20` | budget for the tool-calling loop |
| `AI_MAX_VALIDATION_ROUNDS` | `2` | correction rounds; `0` disables validation |
| `AI_MAX_TRAINING_EXAMPLES` | `3` | how many curated examples may be fetched |

The full list, each with a comment explaining it, is in the package's `settings.yml`.

### LLM settings

The package sends **no sampling parameters** — no temperature, no top_p. Those come
from the defaults of whatever serves the model (LM Studio, vLLM, the hosted API). The
one exception is `LLM_CHAT_TEMPLATE_KWARGS`, which travels with each request. So these
are settings you make on the server, not here. **A model vendor's own recommendation
always takes precedence over the table below.**

| Setting | Suggested | Why |
|---|---|---|
| Reasoning / thinking mode | **off** | Decode dominates the runtime, and thinking tokens are decode. Measured on a local 27B: 85% of all generated tokens went into reasoning. Risk of thinking blocks or empty answers instead of JSON. |
| Tool / function calling | **on, native** | Mandatory — the tools are passed over the API as `tools=[...]`. |
| Temperature | 0.1–0.2 to start | Qwen warns against greedy decoding (`0.0` causes repetition loops, it recommends `0.7`); gpt-oss is specified for `1.0`. Start low, raise if you see loops. |
| Top P | 0.8–0.9 | Keep conservative for structured output. |
| Top K | 20–40 | Qwen recommends 20. |
| Repeat penalty | 1.0–1.05 | Higher values corrupt recurring field names. |
| Context length | 64K | A run including two correction rounds needs ~15–25K tokens. Use 128K only if you need it. |
| Max response tokens | 2048–4096 | Per assistant turn, tool-call arguments included. The final JSON and the largest tool argument stay under 1K. |
| Prompt truncation | **disabled** | Set the overflow policy to stop at the limit. The transcript grows over the correction rounds; "truncate middle" would silently cut away rules or examples. |
| Structured output / JSON schema | use with care | Forcing a schema on every answer can suppress tool calls. `validate_data_config` checks the final JSON anyway. |
| Quantization | 6–8 bit where memory allows | 4-bit produces measurably less precise tool arguments. Drop to 4-bit only as a memory fallback. |
| Seed | fixed | Makes a validation run reproducible. |

Running locally, pick a model by class rather than by name — a specific ranking dates
quickly. With 16 GB expect to be limited to a small MoE or an aggressively quantized
model, and to lose precision in tool arguments; 24–32 GB comfortably runs a mid-size
MoE with few active parameters at 4-bit; from 64 GB a dense model in the 27–35B range
at 6–8 bit becomes practical. Mixture-of-experts models with roughly 3–4B active
parameters give the best interactive latency, because decode speed, not model size, is
what a run waits on.

### API keys

Keys do **not** belong in `settings.yml`. They go in `<env>/.keys`, in dotenv format.
Anything ending in `_API_KEY` or `_API_BASE` is picked up:

```
OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...
```

A missing file is not an error — an env that sets its keys as real environment
variables does not need one.

### GitHub Copilot

`LLM_MODEL: "github_copilot/<model>"` runs on a Copilot subscription instead of a
provider API key. There is no key to put in `.keys`: litellm authenticates with a
GitHub OAuth token that it obtains once through a device flow and then keeps in a
token directory, as two files.

| File | What it is |
|---|---|
| `access-token` | the long-lived GitHub token, written by the device flow |
| `api-key.json` | the short-lived Copilot key, refreshed automatically |

The default directory is `~/.config/litellm/github_copilot`, which inside a container
is lost on the next rebuild. So set the directory **first**, then log in.

**1. Point the token directory at the env.** In `<env>/.keys`, next to the other
credentials:

```
GITHUB_COPILOT_TOKEN_DIR=/isynth_envs/<env>/.copilot
```

Every `GITHUB_COPILOT_*` variable in `.keys` is exported, the same way as an
`_API_KEY`. Add `.copilot/` to the env's `.gitignore` — it holds a credential.

**2. Log in once, interactively.** The device flow prints a code and waits for
someone at a browser, so do it by hand rather than letting it happen inside a plugin
call. Importing `ai_data_config.llm_helper` is what applies step 1 — with plain
`import litellm` the token would land in the container's home directory instead.

Run this in a terminal on the **host** — `docker exec` executes it inside the appserver
container. It is one line on purpose, so it survives copy and paste. `-e ENV_HOME=…`
tells the package which env it is in: a `docker exec` shell, unlike iSynth's own
processes, does not have it set, and without it an installed package finds no `.keys` —
the token then lands in the container's home directory instead of `.copilot/`:

```bash
docker exec -it -e ENV_HOME=/isynth_envs/<env> <appserver> sh -c 'cd /isynth_envs/<env> && python3 -c "import ai_data_config.llm_helper, litellm; r = litellm.completion(model=\"github_copilot/<model>\", messages=[{\"role\": \"user\", \"content\": \"hi\"}]); print(r.choices[0].message.content)"'
```

Keep `-it`: without a terminal attached you do not see the code the device flow prints.

It prints `Please visit https://github.com/login/device and enter code XXXX-XXXX`.
Open that page, enter the code, approve. The call then answers normally.

**3. Check that both files are there.**

```bash
docker exec <appserver> ls /isynth_envs/<env>/.copilot
# access-token  api-key.json
```

**4. Set the model.** In the env's `settings.yml`:

```yaml
LLM_MODEL: "github_copilot/<model>"
```

Leave `LLM_API_BASE` empty — litellm resolves `https://api.githubcopilot.com` itself.
Which models you can name is Copilot's list, not litellm's.

**5. Verify**, with the check under [*Check that it works*](#4-check-that-it-works). From
here on nothing is Copilot-specific; generation, the MCP tools and validation runs
behave as with any other model.

Steps 1 and 2 are needed once per env. The `api-key.json` expires regularly and is
refreshed without asking; the `access-token` survives until it is revoked on GitHub.

Two notes on the integration. It is litellm's **provider** integration
(`docs.litellm.ai/docs/providers/github_copilot`) — not the *tutorial* page of a
similar name, which points the Copilot IDE extension at a LiteLLM proxy and is the
opposite direction. And tool calling, which this package depends on, works even
though that page does not mention it.

For a non-default setup the remaining variables also go in `.keys`:
`_ACCESS_TOKEN_FILE` and `_API_KEY_FILE` rename the two files; `_API_BASE`,
`_DEVICE_CODE_URL`, `_ACCESS_TOKEN_URL` and `_API_KEY_URL` point at a GitHub
Enterprise endpoint.

## Integrating into an env

The steps that put the package into the iSynth UI, after
[installing](#install) and [configuring](#configuration) it. Skip
any you do not need.

### 1. Plugin wrappers

The four plugin wrappers are only the iSynth UI; the functionality lives in the
package. They are not part of the installed package: take them from the `plugins/`
folder of the isynth-ai-data-config repository and copy them into your env's `plugins/`
folder.

Each wrapper needs `base.plugin` (the iSynth plugin API), most also `base.logging`.
Beyond that:

| Wrapper | Calls | Additionally needs |
|---|---|---|
| `plugins/ai_data_config.py` | `llm_helper.generate_data_order()` | — |
| `plugins/validate_ai_prompts.py` | `validation.config_validator.prepare_and_execute_validation_run()` | — |
| `plugins/restart_mcp_server.py` | `mcp_server.admin.request_restart()` | `base.composers.Composer`, `settings.BASE_DIR` |
| `plugins/bill_of_materials.py` | `bill_of_materials.calculate_bill_of_materials()` | `base.plugin_pro.ParameterType` |

Afterwards run **"Reload plugins"** from the context menu of the `plugins/` folder in
the UI — the REST API serves new definitions immediately, the UI only after that.

### 2. Composer order (once)

`get_composer_order()` reads `__generated__/composer_sequence.txt`. Without that file it
returns an **empty list**, and the system prompt then fails to tell the model which
order composer entries have to be in.

The restart plugin writes it: run **Restart MCP Server** with the checkbox ticked. It
fetches the order via `Composer.execution_sequence()` and writes the file *before*
restarting the server, so it appears even when no MCP server is running yet (the
restart call then fails, which is fine).

Repeat after any model change that adds, removes or reorders composers. The file is
read at import time, so a running process only picks up a change after a restart.

### 3. MCP server service (optional)

Only needed for external MCP clients (IDE agents) and for the **Restart MCP Server**
plugin. Generation and validation runs do **not** need it — they call the same tools
in-process and keep working with the service stopped.

```yaml
  i-mcpserver:
    build:
      context: ../
      dockerfile: deployment/Dockerfile
    restart: unless-stopped          # safety net, should os.execv fail on restart
    environment:
      - ENV_HOME=/isynth_envs/<env>  # must match the mount target below
      - MCP_TRANSPORT=streamable-http
      - MCP_HOST=0.0.0.0
      - MCP_PORT=8010
      - MCP_PATH=/mcp
      - MCP_WAIT_FOR_GENERATED_TIMEOUT=60
    volumes:
      - "../:/isynth_envs/<env>"
    working_dir: /isynth_envs/<env>  # so "-m ai_data_config.mcp_server" resolves
    networks: [<the appserver's network>]
    ports:
      - "8010:8010"
    command: "python3 -m ai_data_config.mcp_server"
```

**The service name is not free.** `mcp_server/admin.py` has
`DEFAULT_BASE_URL = "http://i-mcpserver:8010"`, which is where the restart plugin looks.
Either the service is called `i-mcpserver`, or the appserver service sets
`MCP_SERVER_URL` to the right address.

One server serves one env. For several envs, see
[*The MCP server with several envs*](#the-mcp-server-with-several-envs).

Installed from the index, the package also puts an `ai-data-config-mcp` command on
`PATH`, equivalent to `python3 -m ai_data_config.mcp_server`.

### 4. Check that it works

```bash
# package loads, settings arrive, tools answer (one line, so it survives copy and paste)
docker exec -e ENV_HOME=/isynth_envs/<env> <appserver> sh -c 'cd /isynth_envs/<env> && python3 -c "from ai_data_config import config, llm_helper; from ai_data_config.mcp_server.tools import TOOLS, call_tool; print(\"ENV_HOME :\", config.ENV_HOME); print(\"LLM_MODEL:\", llm_helper.DEFAULT_MODEL); print(\"Tools    :\", len(TOOLS)); print(\"Composer :\", call_tool(\"get_composer_order\", {})[:3], \"...\"); print(\"Examples :\", len(call_tool(\"list_ai_training_examples\", {})))"'

# MCP server, if you set one up
curl -s http://localhost:8010/health
```

Expected: `ENV_HOME` is your env (not `site-packages`), 14 tools, a **non-empty** composer order, the number of training examples, and
`{"status":"ok",...}`. An empty composer list means step 2 is missing.

### What needs iSynth and what does not

`llm_helper` and `mcp_server` import nothing from `base` — generation and the MCP tools
run in any directory where the dependencies are installed, iSynth or not.

Only `bill_of_materials.py` needs the iSynth runtime (`base.constellation_utils`,
`base.data_config_utils`), and through it the validation run, which compares the bill of
materials of a generated `data_config` against the case's reference. `base` is part of
the iSynth image and cannot be installed from an index, so importing that module
elsewhere raises a `ModuleNotFoundError` saying so.

The MCP server container is needed by none of this except `restart_mcp_server`.

## Using the package in several envs

One iSynth instance usually hosts several envs, each in its own directory under
`/isynth_envs/`. The package is built for that: a single installation serves all of
them, and everything it reads or writes belongs to the env it is running in. This page
explains how it tells the envs apart, what each env needs, and the few things that are
shared.

### How the package finds its env

Everything env-specific is resolved from one directory, `ENV_HOME`, and read **once**,
when the package is first imported in a process:

- **Inside iSynth you do nothing.** Every env process loads its env's `settings.py`, and
  that sets `ENV_HOME` to the env's directory. Plugins, data orders and validation runs
  therefore always see the env they were started in.
- **Outside iSynth, set it yourself.** A `docker exec` shell, a script or the MCP server
  container is not an iSynth env process, so nothing sets `ENV_HOME` there:

  ```bash
  docker exec -e ENV_HOME=/isynth_envs/<env> <appserver> ...
  ```

  Without it, the package uses the directory above the package folder if that is an env
  (the package was [copied into the env](#installing-by-copying-the-folder)), otherwise
  the current working directory. Up to version 0.8.0 the working-directory fallback
  does not exist and an installed package ends up looking in `site-packages` — find no
  metadata, no `.keys`. With those versions, always set `ENV_HOME`.

To see which env a process uses, run the check under
[*Check that it works*](#4-check-that-it-works): its first line is `ENV_HOME`.

### What is per env and what is shared

| What | Where | Per env |
|---|---|---|
| Generated metadata, constellation templates, training examples | `<env>/__generated__/`, `<env>/constellations/**/__generated__/`, `<env>/data_configs/ai_training/` | yes |
| Composer order | `<env>/__generated__/composer_sequence.txt` | yes — write it in each env |
| Setting overrides (model, endpoint, budgets) | `<env>/settings.yml` | yes |
| API keys, GitHub Copilot token | `<env>/.keys`, and the token directory it names | yes |
| Customisation hooks | `<env>/ai_config/` | yes, optional |
| Validation cases and results | `<env>/ai_config/validation_cases/`, `<env>/logs/validation_runs.xlsx` | yes |
| Log files | `<env>/logs/*.log` | yes |
| Plugin wrappers | `<env>/plugins/` | yes — copy them into each env |
| MCP server | one service per env | yes — see [below](#the-mcp-server-with-several-envs) |
| Package code and version | the image's `site-packages` | **shared** by all envs of the image |
| Package defaults | `settings.yml` inside the package | **shared** |

So two envs can use different models and keys, and each generates from its own data
model. What they share is the code: installed into the image, every env runs the same
version with the same defaults.

### What each env needs

Every env that uses the package needs, on its own:

1. **Generated metadata** — run **Initialize environment** in that env.
2. **The composer order** — run **Restart MCP Server** once in that env (see
   [*Composer order*](#2-composer-order-once)); it writes the file into the env it runs
   in.
3. **The plugin wrappers** in its `plugins/` folder, then **Reload plugins**.
4. **Its settings and keys**, if it deviates from the package defaults: `settings.yml`
   and `.keys` in the env directory.

Nothing has to be registered centrally, and an env without the plugin wrappers is
simply not affected.

**A different package version for one env.** Installing into the image gives every env
the same version. An env that needs another one gets its own copy of the package folder
(see [*Installing by copying the folder*](#installing-by-copying-the-folder)): iSynth puts
the env directory first on the module search path of its env processes, so that env
imports its copy while the others keep using the installed package.

### One process, one env

iSynth runs an env's code in env processes that belong to that env and are reused for
later requests. The package relies on this: the env's directory, settings, API keys,
constellation templates and composer order are taken over when the package is first
imported in a process, and kept for the life of that process.

What that means in practice:

- **Changes to object types, composers, enums, functions and training examples** are
  picked up by running processes by themselves: the package re-reads those files when
  they change.
- **A new or changed constellation template, a new composer order, a changed
  `settings.yml` or `.keys`** reach an env only through new processes. iSynth ends all
  env processes after **Sync definitions** and **Generate models**; a long-running MCP
  server needs **Restart MCP Server**.
- **Never serve two envs from one process.** A script that loops over several envs
  keeps the first one — start one process per env, each with its own `ENV_HOME`.

### The MCP server with several envs

A server serves exactly one env: the one in its `ENV_HOME` when it starts. Its `/health`
answer names it (`"server": "<env>-mcp"`). For several envs, run one service per env,
each with its own `ENV_HOME`, mount, port and service name — the service template under
[*MCP server service*](#3-mcp-server-service-optional) is written for one.

The **Restart MCP Server** plugin restarts the server at `MCP_SERVER_URL`, or
`http://i-mcpserver:8010` when that is not set. The variable comes from the appserver's
environment, so it is the same for every env in that appserver: all envs restart the
same server. With one MCP server per env, restart the others by restarting their
service. Writing the composer order is not affected — the plugin always writes it into
its own env.

## Installing by copying the folder

`pip install isynth-ai-data-config` is the normal way in, and
[Getting started](#getting-started) describes it. This page covers the
alternative: copying the package folder `ai_data_config/` wholesale into an env. You
find it under `src/` in the source distribution (`isynth_ai_data_config-<version>.tar.gz`)
or in the repository.
That is how the package was distributed before it was published, and it stays
supported — for an env that cannot reach the private index, or one where the package
is edited in place rather than upgraded.

The difference is confined to getting the code and its dependencies into the env.
Everything after that — plugin wrappers, composer order, the MCP service, model and API
keys — is identical either way; see [Configuration](#configuration) and
[Integrating into an env](#integrating-into-an-env).

### 1. Copy the folder

```bash
cp -R src/ai_data_config/ <target-env>/
```

The folder name is not arbitrary: imports are `ai_data_config.*`, and
`plugins/restart_mcp_server.py` filters on `path_starts_with='ai_data_config'`.

### 2. Wire up the dependencies

Nothing installed them for you, so the env has to. The package keeps its own list;
include it rather than copying the names out of it:

```
# <target-env>/deployment/requirements.txt
-r ../ai_data_config/requirements.txt
```

pip resolves a `-r` path relative to the file it stands in. One list, so a new
dependency of the package needs no change in the env's own file.

### 3. Docker build context

For the build to reach **both** files, the build context has to be the repository root,
not `deployment/`:

```yaml
# deployment/compose.yml, on every service that builds the image
build:
  context: ../
  dockerfile: deployment/Dockerfile
```

```dockerfile
# deployment/Dockerfile
COPY  deployment/requirements.txt      ./requirements/deployment/requirements.txt
COPY  ai_data_config/requirements.txt  ./requirements/ai_data_config/requirements.txt
RUN   pip install -r ./requirements/deployment/requirements.txt --no-cache-dir
```

### 4. Keep the context small

A repository-root context means the whole repository would be sent to the Docker daemon
unless you exclude it — in one env 630 MB instead of 670 bytes. Add a `.dockerignore`
at the repository root, written as an allow-list so folders added later are not shipped
silently:

```
*
!deployment/requirements.txt
!ai_data_config/requirements.txt
```

### Then continue with the configuration

From [Configuration](#configuration) onwards, nothing about the copied folder is
special.

## Troubleshooting

| Symptom | Cause |
|---|---|
| `FileNotFoundError: ... __generated__/...` | **Initialize environment** has not run in the env yet |
| `ModuleNotFoundError: No module named 'box'` or similar | dependencies not installed — see [*Install*](#install) |
| `ModuleNotFoundError: ... needs the iSynth runtime ('base')` | `bill_of_materials` or the validation run is running outside an iSynth env |
| `MCP server at ... did not answer` | service is not called `i-mcpserver` and `MCP_SERVER_URL` is not set |
| Plugin does not appear in the UI | run "Reload plugins" from the context menu of the `plugins/` folder |
| Env override is ignored | key not declared in the package's `settings.yml` — the log warning names it |
| `get_composer_order()` is empty | `__generated__/composer_sequence.txt` missing — see [*Composer order*](#2-composer-order-once) |
| `ENV_HOME` shows `site-packages`, or keys and metadata are not found from a shell | `ENV_HOME` not set outside iSynth — see [*How the package finds its env*](#how-the-package-finds-its-env) |
| An env still uses an old constellation template or composer order | the env process predates the change — see [*One process, one env*](#one-process-one-env) |
| Logs are empty | not on STDOUT: the package logs with `propagate=False` into `<env>/logs/*.log` |
