Metadata-Version: 2.4
Name: bipixie-mcp
Version: 0.6.0
Summary: MCP server that gives AI agents read-only access to BI Pixie usage and engagement data via the Power BI REST executeQueries endpoint, and management tools over the BI Pixie API as the signed-in person
Author-email: DataChant <support@bipixie.com>
Maintainer-email: DataChant <support@bipixie.com>
License: MIT
Project-URL: Homepage, https://bipixie.com
Project-URL: Documentation, https://bipixie.com/docs/cloud/mcp-server/
Project-URL: Support, https://bipixie.com/docs/cloud/contact-support
Keywords: mcp,power-bi,bi-pixie,analytics,model-context-protocol
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: fastmcp<3,>=2.11
Requires-Dist: msal>=1.28
Requires-Dist: azure-identity>=1.17
Requires-Dist: httpx>=0.27
Requires-Dist: pydantic>=2.7
Requires-Dist: pydantic-settings>=2.3
Provides-Extra: test
Requires-Dist: pytest>=8; extra == "test"
Requires-Dist: pytest-asyncio>=0.23; extra == "test"
Requires-Dist: respx>=0.21; extra == "test"
Requires-Dist: PyYAML>=6; extra == "test"
Provides-Extra: dev
Requires-Dist: bipixie-mcp[test]; extra == "dev"
Requires-Dist: mypy>=1.10; extra == "dev"
Requires-Dist: ruff>=0.4; extra == "dev"
Dynamic: license-file

# BI Pixie MCP Server

A [Model Context Protocol](https://modelcontextprotocol.io/) server that lets AI agents (Claude Code, Codex, Microsoft Fabric data agents, and Azure AI Foundry) query your BI Pixie usage and engagement data, read-only, and, on your own machine, act on BI Pixie as you through the BI Pixie API (see [Management tools](#management-tools-bi-pixie-as-you-over-the-bi-pixie-api-local-stdio-only)).

> **Note:** Requires an active BI Pixie license and a deployed BI Pixie semantic model. "BI Pixie" is a trademark of DataChant. This package is a client and grants no rights to the BI Pixie service or data.

**Phase 1 data source:** the existing "BI Pixie" semantic model, queried via the Power BI REST `executeQueries` endpoint. No new data infrastructure is required. The server is a pure-additive Python 3.12 package (`bipixie_mcp`) that lives in `mcp_server/` and does not touch any existing Function App, portal, or workload code.

**Validated against the live model.** All 8 core data tools (`describe_model`, `list_reports`, `top_reports_by_usage`, `report_health`, `user_adoption`, `page_engagement`, `feedback_summary`, `run_dax`) were executed successfully against the BI Pixie SaaS sample dataset (plus `suggest_questions` for discovery and four stdio-only navigation tools — see the tool catalog). Every measure resolved; ranked output is deterministically ordered. Results match the customer's own Power BI report numbers (e.g. top report "Executive Dashboard" = 484 report sessions; CSAT 0.59, NPS -58.3; data through 2026-05-07).

---

## How it works

```
AI agent (Claude Code / Codex / Foundry)
        |
        | MCP JSON-RPC  (stdio  OR  POST /mcp)
        v
  bipixie_mcp.server  (FastMCP, Python 3.12)
        |
        | Power BI REST  executeQueries  (read-only DAX)
        | POST https://api.powerbi.com/v1.0/myorg/groups/{workspaceId}/datasets/{datasetId}/executeQueries
        v
  "BI Pixie" semantic model  (IMPORT-mode, ~60 tables, validated DAX measures)
        ^
        | data already loaded at refresh time
  ADLS Gen2  bipixielake-{license_key}/events/
```

The server wraps the model's curated measures — the same numbers you see in your BI Pixie dashboard — in typed, filterable MCP tools. Because the measures are already validated by the model, agent results match your Power BI reports exactly.

**Two runtime targets, one codebase:**

| Target | Transport | `BIPIXIE_MCP_TRANSPORT` | Typical consumer |
|--------|-----------|------------------------|-----------------|
| Local developer machine | `stdio` | `stdio` (default) | Claude Code, Codex |
| Customer Azure environment | `streamable-http` | `streamable-http` | Fabric data agents, Azure AI Foundry, remote Claude Code |

---

## Prerequisites

### Python

Python 3.12 or later.

```
python --version   # must be 3.12+
```

On Windows with multiple Python versions installed, use `py -3.12` instead of `python`.

### Power BI tenant settings (customer admin action — required before first use)

The following settings must be enabled by a Power BI admin in your tenant. These are customer-side prerequisites; the server cannot self-provision them. They are the same for cloud and self-hosted customers: the only difference is which tenant's Power BI admin turns them on, the tenant that holds your BI Pixie semantic model.

| Setting | Location in Power BI Admin portal | Why |
|---------|-----------------------------------|-----|
| **Dataset Execute Queries REST API** | Tenant settings > Integration settings | Hard gate. Disabled = 403 with no helpful body. |
| **Allow service principals to use Power BI APIs** | Tenant settings > Developer settings | Required for `service_principal` and `managed_identity` auth modes. |
| Workspace **Member** (or Admin) role for your SP or MI | Workspace settings > Manage access | Required so the app identity has Read + Build on the dataset. |

For `device_code` / `interactive` / `azure_cli` auth (local use), the customer's own user identity is used and workspace membership follows normal Power BI access — no service-principal enrollment needed.

### Entra app registration

For `azure_cli` (simplest local validation): **no app registration needed.** Reuses an existing `az login` session via `AzureCliCredential`. Run `az login` once in your shell, then set `BIPIXIE_MCP_AUTH_MODE=azure_cli` — no `BIPIXIE_MCP_CLIENT_ID` required.

For `device_code` / `interactive` (local, no existing az session): create a **public-client** app registration in your Entra tenant with:
- Redirect URI: `https://login.microsoftonline.com/common/oauth2/nativeclient` (for device code)
- API permission: `Power BI Service > Dataset.Read.All` (delegated)
- Token version: `requestedAccessTokenVersion = 2` (v2 tokens)

For `service_principal` (hosted): create a **confidential-client** app registration and add:
- API permission: `Power BI Service > Dataset.Read.All` (application)
- Grant admin consent
- Token version: `requestedAccessTokenVersion = 2`

For `managed_identity` (hosted, preferred): no app registration needed on the server side. The managed identity must be added as workspace Member.

For sign-in on the hosted endpoint itself, create one more app registration that callers get tokens for: Application ID URI `api://<clientId>`, `requestedAccessTokenVersion = 2`, no secret and no API permissions. The hosted template (`deploy/mcp-container.bicep`) requires its application ID as `clientId` and configures the Container App's built-in authentication to accept only v2.0 tokens from your tenant whose audience is `<clientId>` or `api://<clientId>`. Step-by-step commands are in [deploy/README.md](deploy/README.md), step 2.

---

## Install

The server runs locally (stdio) in VS Code, Cursor, Claude Code / Desktop, and Codex. Install it from **PyPI** — no clone required:

```bash
# Zero-install runner (recommended) — always fetches the latest published version:
uvx bipixie-mcp

# Or install into the current environment:
pip install bipixie-mcp
```

[![Install in VS Code](https://img.shields.io/badge/VS_Code-Install_BI_Pixie_MCP-0098FF?logo=visualstudiocode&logoColor=white)](https://vscode.dev/redirect/mcp/install?name=bipixie&config=%7B%22command%22%3A%22uvx%22%2C%22args%22%3A%5B%22bipixie-mcp%22%5D%7D)

### First-run setup wizard

Run `uvx bipixie-mcp` (or `bipixie-mcp setup`) **in a terminal** with no config yet and it launches an interactive guided setup instead of erroring out: it signs you in (reusing `az login` when present, else device code), lists the workspaces your identity can see, flags your `BI Pixie` models, lets you pick one from a numbered menu, runs a one-row test query to confirm auth + the *Dataset Execute Queries* tenant setting + workspace access, and writes a `.env` plus a ready-to-paste `.mcp.json` block. Re-run any time with `bipixie-mcp setup`.

The wizard is gated on an interactive TTY (`stdin` **and** `stderr` are terminals): when an MCP client launches the server over pipes it never prompts — it just reads the env below — so the JSON-RPC stdout stream is never touched and a misconfigured non-interactive run still fails fast with the original `ConfigError`. Set `BIPIXIE_MCP_NO_SETUP=1` to disable the auto-trigger.

Check which build you have at any time with `bipixie-mcp --version` (or `-V`) — it prints `bipixie-mcp <version>` to stdout and exits before reading any config, so it works on any install. Note `uvx` caches by version: when testing a locally built wheel that reuses a version number, pass `uvx --refresh` so it picks up the new code.

After installing (or to configure by hand), set the required `BIPIXIE_MCP_*` variables (see the **Configuration reference** section below) — at minimum a tenant, workspace, and dataset, plus `BIPIXIE_MCP_AUTH_MODE=azure_cli` (then `az login`) for the simplest local auth.

<details>
<summary><strong>From source (maintainers / contributors)</strong></summary>

```bash
pip install -e "mcp_server/[test]"   # from the repo root
pip install -e ".[test]"             # from the mcp_server/ directory
```

</details>

**Dependencies installed automatically:**

| Package | Purpose |
|---------|---------|
| `fastmcp>=2.11` | MCP server framework (stdio + streamable-HTTP) |
| `msal>=1.28` | MSAL Python — device-code / interactive token cache |
| `azure-identity>=1.17` | `DefaultAzureCredential`, `ClientSecretCredential` |
| `httpx>=0.27` | Async HTTP for Power BI REST calls |
| `pydantic>=2.7` | Data validation |
| `pydantic-settings>=2.3` | `BIPIXIE_MCP_*` env-var config |

Test extras (`pytest`, `pytest-asyncio`, `respx`) are included with `[test]`.

---

## Configuration reference

All configuration is via environment variables (or a `.env` file in the working directory). Copy `.env.example` to `.env` and fill in your values. **Never commit `.env` to source control.**

| Variable | Required | Default | Description |
|----------|----------|---------|-------------|
| `BIPIXIE_MCP_TENANT_ID` | Yes | — | Entra (AAD) tenant GUID. Cloud customers: DataChant production tenant. Self-hosted: your own tenant GUID. Never hardcoded — wrong value breaks auth. |
| `BIPIXIE_MCP_AUTH_MODE` | No | `device_code` | Token-acquisition flow: `device_code` \| `interactive` \| `azure_cli` \| `service_principal` \| `managed_identity`. `azure_cli` reuses an existing `az login` session (no app registration, no client secret — simplest local validation). |
| `BIPIXIE_MCP_CLIENT_ID` | Yes (except azure_cli / managed_identity) | — | Entra app (client) ID. Public client for device-code/interactive; confidential for service_principal. Not needed for `azure_cli` or `managed_identity` (both obtain credentials without an explicit client registration). |
| `BIPIXIE_MCP_CLIENT_SECRET` | For SP only | — | Client secret for `service_principal` mode. In Azure, use a Key Vault reference (`@Microsoft.KeyVault(...)`). Never logged. |
| `BIPIXIE_MCP_WORKSPACE_ID` | One of ID/name | — | GUID of the Fabric/Power BI workspace containing the "BI Pixie" dataset. If unset, resolved from `BIPIXIE_MCP_WORKSPACE_NAME`. |
| `BIPIXIE_MCP_WORKSPACE_NAME` | One of ID/name | — | Workspace display name, used to resolve the workspace GUID at startup. Ignored when `WORKSPACE_ID` is set. |
| `BIPIXIE_MCP_DATASET_ID` | One of ID/name | — | GUID of the "BI Pixie" semantic model. If unset, resolved from `BIPIXIE_MCP_DATASET_NAME`. Fails fast on name collision — use an explicit ID when multiple "BI Pixie" datasets exist in the workspace. |
| `BIPIXIE_MCP_DATASET_NAME` | No | `BI Pixie` | Dataset display name for ID resolution. Override only if the customer renamed the model. |
| `BIPIXIE_MCP_POWERBI_API_BASE` | No | `https://api.powerbi.com/v1.0/myorg` | Power BI REST base URL. Change only for Sovereign Cloud (GovCloud / China). |
| `BIPIXIE_MCP_POWERBI_SCOPE` | No | `https://analysis.windows.net/powerbi/api/.default` | OAuth scope for executeQueries. Confirmed in `SemanticModelClient.ts:122-127` — Fabric-audience tokens are rejected by api.powerbi.com. |
| `BIPIXIE_MCP_TRANSPORT` | No | `stdio` | Runtime transport: `stdio` (local) or `streamable-http` (hosted). |
| `BIPIXIE_MCP_HTTP_HOST` | No | `0.0.0.0` | Bind host for streamable-http mode. Ignored in stdio mode. |
| `BIPIXIE_MCP_HTTP_PORT` | No | `8000` | Bind port for streamable-http mode. Azure injects `PORT` / `WEBSITES_PORT` automatically. |
| `BIPIXIE_MCP_DEFAULT_DAYS` | No | `30` | Default lookback window (days) when a tool's `days` argument is omitted. Cannot exceed the model's loaded history. |
| `BIPIXIE_MCP_DEFAULT_TOP_N` | No | `100` | Default row cap for ranking/list tools when `top_n` is omitted. |
| `BIPIXIE_MCP_MAX_ROWS` | No | `1000` | Hard server-side cap on rows returned by any tool. Enforced before serialization. |
| `BIPIXIE_MCP_PII_COLUMNS_ALLOWED` | No | `true` | When `true` (default), results pass through unfiltered, so `Username` and `Client IP` reach the caller if the model exposes them. Set `false` to have `ResponsePiiFilter` strip those columns. The BI Pixie model ships no RLS, so this flag is the only PII control across all auth modes. |
| `BIPIXIE_MCP_QUERY_RATE_LIMIT` | No | `60` | Client-side cap on executeQueries calls per minute per identity, to stay under the Power BI ~120/min quota. |
| `BIPIXIE_MCP_TOKEN_CACHE_PATH` | No | (ignored) | **Deprecated, no effect.** The device_code / interactive sign-in is saved by the Azure identity library under the name `bipixie-mcp` in `.IdentityService` in your user profile; see the security checklist. Setting it logs a warning. |
| `BIPIXIE_MCP_LOG_LEVEL` | No | `INFO` | Logging level (`DEBUG` / `INFO` / `WARNING` / `ERROR`). All logs go to **stderr** — stdout is reserved for the MCP JSON-RPC stream. |
| `BIPIXIE_MCP_ALLOW_TARGET_SWITCHING` | No | `true` | On the local stdio transport, registers the four navigation tools (`list_workspaces`, `list_datasets`, `set_active_dataset`, `get_active_dataset`). Set `false` to lock the server to the configured dataset. The hosted streamable-http transport never registers them. |
| `BIPIXIE_MCP_DEPLOYMENT_MODE` | No | `self_hosted` | `cloud` or `self_hosted`. Informational only: it labels logs and diagnostics and does not change how the server queries Power BI. |
| `BIPIXIE_MCP_DATA_LAKE_URL` | No | — | Your per-tenant ADLS Gen2 container URL. Not used when querying the semantic model; kept for diagnostics and the documented direct-ADLS fallback. |
| `BIPIXIE_MCP_API_BASE` | No | `https://api.bipixie.com/v1` | Base URL of the BI Pixie API the management tools call as you (`https://api-dev.bipixie.com/v1` on dev). Always https. |
| `BIPIXIE_MCP_API_APP_ID` | For the management tools | — | The BI Pixie API's Entra application ID, which names the token the management tools present (`api://<id>/...`). Unset, the management tools are not registered and the usage tools are unchanged. Cloud production (`api.bipixie.com`): `114a7e20-3796-4ddd-ab99-568a78b7759b`. Cloud development (`api-dev.bipixie.com`): `2e8a58e5-12c0-4a8f-a754-576086c505af`. Self-hosted: your own BI Pixie API's application ID. |
| `BIPIXIE_MCP_WRITE_TOOLS` | No | `true` | On the local stdio transport, registers the management write tools beside the reads; each one asks you before it spends or changes anything. Set `false` for a reads-only server: no write tool is registered, the previews and the coverage check included, and a browser or device-code sign-in asks for the API's read scope alone (the Azure CLI asks for the scopes it is pre-authorized for either way). The hosted streamable-http transport ignores it and registers no management tool. |
| `BIPIXIE_MCP_WAIT_SECONDS` | No | `60` | How long `wait_for_operation` follows a BI Pixie operation when the caller names no budget, in seconds. |

**Cloud vs. self-hosted:** the config keys are identical in both deployments. Only the values differ — cloud customers supply DataChant's tenant GUID and their own workspace/dataset IDs; self-hosted enterprise customers supply their own tenant GUID, their own Entra app registration, and their own workspace/dataset IDs. Nothing is hardcoded in the server.

---

## Management tools: BI Pixie as you, over the BI Pixie API (local stdio only)

Beside the usage tools, the local server carries a second family that acts through the [BI Pixie API](https://bipixie.com/docs/api) as the signed-in person, with that person's own permissions and plan. Every plan, Free included, may read and change things, within the same limits the portal applies. Each tool is one API call; the API's refusals, allowances and audit lines are yours. Nothing here is computed by the server, and no token of yours is stored by BI Pixie.

| Ask your agent | Tool |
|---|---|
| What is my account, my plan, what do I have left today and this month | `get_account` |
| Which reports and semantic models count as tracked items | `list_tracked_items` |
| Which semantic models BI Pixie knows, their AI Readiness score, their assessments | `list_semantic_models`, `get_semantic_model_readiness`, `list_assessments`, `get_assessment` |
| Which data agents BI Pixie knows, their readiness, their assessments | `list_data_agents`, `get_data_agent_readiness`, `list_data_agent_assessments`, `get_data_agent_assessment` |
| The saved benchmark questions, the benchmark runs, one scored run | `get_benchmark_questions`, `list_benchmark_runs`, `get_benchmark_run` |
| The A/B tests, one test, one run, a run's comparison, the tests an item is on | `list_ab_tests`, `get_ab_test`, `get_ab_test_run`, `get_ab_test_result`, `list_ab_tests_of_semantic_model`, `list_ab_tests_of_data_agent` |
| The assessment schedules | `list_schedules`, `get_schedule` |
| Whether a report has Pixies | `get_report_pixies` |
| The AIs my account has saved and the credentials they use (an account admin's; never a key) | `list_saved_ais` |
| What BI Pixie is doing for me, one operation, its result, wait for it to end | `list_operations`, `get_operation`, `get_operation_result`, `wait_for_operation` |

**Switching it on.** Set `BIPIXIE_MCP_API_APP_ID` to the BI Pixie API's application ID (the cloud value is in the table above) and, on dev, `BIPIXIE_MCP_API_BASE=https://api-dev.bipixie.com/v1`. Without the application ID the server carries the usage tools alone, as before.

**Sign-in.** One sign-in, two tokens: the same identity yields a token for Power BI (the usage tools) and a token for the BI Pixie API (the management tools). With `BIPIXIE_MCP_AUTH_MODE=azure_cli`, `az login` is enough for both: the Azure CLI is a pre-authorized client of the API, so nothing is registered and nothing is consented to. With `device_code` or `interactive`, the management tools ask the API's scope through your public client, and you consent to it yourself on first use. A permission a call needs and you have not granted is answered with the API's own consent link; open it, select Connect Power BI, and call again. Tenant admin consent is never needed and never asked for.

**Where.** The family is registered on the `stdio` transport only, and only for a person's sign-in (`device_code`, `interactive`, `azure_cli`): the BI Pixie API refuses an application identity, so a `service_principal` or `managed_identity` build, and the hosted `streamable-http` build, carry the usage tools alone. The write tools sit behind `BIPIXIE_MCP_WRITE_TOOLS`, which is on by default; set it to `false` for a reads-only server.

**The write tools.** Each one is a single BI Pixie API write made as you. Every tool that spends an allowance, claims a tracked item, changes an item or deletes a result **answers a preview on its first call and does nothing**: what it will do, to which item, what it would spend in the numbers your portal shows, and what it would change. Where your AI client can put a question to you on the server's behalf, it asks you there and one call does both. If you decline there, nothing is done and no agreement is handed back, so the agent cannot act on it without you; asking again asks you again. Where your client cannot ask, or asks and nobody answers, the preview carries an `agreement` that the second call brings back; it is good once, for that call alone, for ten minutes, and it lives in the server's memory only. Nothing waives it.

**Running without anyone present.** Claude Code run from a script or a pipeline (`claude -p`) has nobody to answer the question, so it dismisses it. BI Pixie treats that as unanswered, never as a refusal: the agent gets the preview, its `agreement`, and a note saying that nobody answered. What the agent does next is what you told it to do. If the prompt you gave it already asks for exactly that change, it can call the tool again with the agreement; otherwise it reports the preview back to you. An agreement the agent invents is refused, and the refusal says how to get a real one.

| Ask your agent | Tool |
|---|---|
| Assess a semantic model, several of them, or a data agent | `assess_semantic_model`, `assess_semantic_models`, `assess_data_agent` |
| Show me what BI Pixie would change on a semantic model, then write it, then undo it | `preview_optimization`, `apply_optimization`, `revert_optimization` |
| The same for a data agent's AI instructions and grounding | `preview_data_agent_optimization`, `apply_data_agent_optimization`, `revert_data_agent_optimization` |
| Add Pixies to a report or to several; remove them | `add_pixies`, `remove_pixies` |
| Write benchmark questions and show them to me, save them, check they can all be scored | `generate_benchmark_questions`, `save_benchmark_questions`, `check_benchmark_coverage` |
| Run the benchmark | `run_benchmark` |
| Create an A/B test and run it, run it again, rename it, delete it or one run | `create_ab_test`, `run_ab_test_again`, `update_ab_test`, `delete_ab_test`, `delete_ab_test_run` |
| Schedule assessments; stop a schedule | `put_schedule`, `delete_schedule` |
| Stop a run | `cancel_operation` |

The two optimization previews and the coverage check spend nothing and change nothing, so they answer at once and ask you nothing. Three things are never offered to an agent: proposing and writing an optimization in one unattended step, applying or reverting a whole bundle of optimizations, and a raw call to any BI Pixie address. Saved AIs are read and never changed through an agent: adding one or replacing its key would put your AI provider key into the agent's conversation, so adding, changing and removing a saved AI stay on the portal's AI setup page. An account that is closed, or a sign-in whose role or permission cannot write through the API, is told so before you are asked anything.

**Your own wording.** `apply_optimization` takes `values`, your own text for some of the preview's changes in place of what BI Pixie proposed: AI instructions, a description, or a display folder. The question put to you quotes each text in full, and the agreement stands for that exact text, so different wording needs a fresh preview. What the text may be is the API's rule, and a change it cannot take is refused before anything is written.

**Who answered.** A benchmark run and an A/B test run name the AI provider that answered, beside the data agent when there was one.

**Operations.** A write answers at once with the operation it started; `wait_for_operation` follows it for up to `BIPIXIE_MCP_WAIT_SECONDS` (or the seconds you pass), honouring the API's `Retry-After`, and answers the state and, once it succeeded, the result. An operation whose next step starts on your read (an A/B run's side B) goes on only while the server waits on it; the tool says so when it stops waiting.

---

## Local stdio quickstart — Claude Code

### 1. Set up your `.env`

```bash
# mcp_server/.env
BIPIXIE_MCP_TENANT_ID=72f988bf-86f1-41af-91ab-2d7cd011db47   # your Entra tenant
BIPIXIE_MCP_CLIENT_ID=<your-public-client-app-id>
BIPIXIE_MCP_WORKSPACE_ID=<your-fabric-workspace-guid>          # or use WORKSPACE_NAME
BIPIXIE_MCP_DATASET_NAME=BI Pixie
BIPIXIE_MCP_AUTH_MODE=device_code
```

### 2. Register with Claude Code (one-liner)

```bash
claude mcp add --transport stdio --scope project bipixie -- python -m bipixie_mcp.server
```

Pass required env vars inline or let Claude Code inherit them from your shell.

### 3. Register via `.mcp.json` (project-scoped, checked in)

Create or add to `.mcp.json` in your project root:

```json
{
  "mcpServers": {
    "bipixie": {
      "type": "stdio",
      "command": "python",
      "args": ["-m", "bipixie_mcp.server"],
      "env": {
        "BIPIXIE_MCP_TENANT_ID": "${BIPIXIE_MCP_TENANT_ID}",
        "BIPIXIE_MCP_CLIENT_ID": "${BIPIXIE_MCP_CLIENT_ID}",
        "BIPIXIE_MCP_WORKSPACE_ID": "${BIPIXIE_MCP_WORKSPACE_ID}",
        "BIPIXIE_MCP_DATASET_NAME": "BI Pixie",
        "BIPIXIE_MCP_AUTH_MODE": "device_code"
      }
    }
  }
}
```

On first use, the server performs a device-code flow: it prints a URL and code to stderr, and the user authenticates in a browser. The Azure identity library saves the sign-in under the name `bipixie-mcp` (in `.IdentityService` in your user profile) so subsequent restarts are silent.

**Important:** add `~/.bipixie_mcp/` to your global `.gitignore`. The token cache contains long-lived credentials.

---

## Local stdio quickstart — Codex

Add to `~/.codex/config.toml`:

```toml
[mcp_servers.bipixie]
command = "python"
args = ["-m", "bipixie_mcp.server"]
env_vars = [
  "BIPIXIE_MCP_TENANT_ID",
  "BIPIXIE_MCP_CLIENT_ID",
  "BIPIXIE_MCP_WORKSPACE_ID",
  "BIPIXIE_MCP_DATASET_NAME",
  "BIPIXIE_MCP_AUTH_MODE",
  "BIPIXIE_MCP_CLIENT_SECRET",    # only for service_principal mode
]
```

Set the referenced variables in your shell environment before starting Codex. Use `BIPIXIE_MCP_AUTH_MODE=device_code` for interactive local use, `service_principal` for CI/automation.

---

## Hosted streamable-HTTP quickstart

### Overview

Set `BIPIXIE_MCP_TRANSPORT=streamable-http`. The server exposes a single `POST /mcp` endpoint (MCP Streamable-HTTP protocol, not the deprecated SSE transport). The server runs with `stateless_http=True` (passed to `mcp.run()`, since current FastMCP no longer accepts it on the constructor), so it scales horizontally with no server-side session state.

### Recommended host: Azure Container App

Deploy with the template in [`deploy/`](deploy/README.md). It creates the Container App, its managed identity, and the sign-in gate in one step:

```bash
az deployment group create -g rg-bipixie-mcp \
  --template-file mcp_server/deploy/main.bicep \
  --parameters mcp_server/deploy/main.parameters.json \
  --parameters tenantId=<tenant-id> clientId=<sign-in-app-client-id> \
               workspaceId=<workspace-guid> image=ghcr.io/bi-pixie/bipixie-mcp:<tag>
```

Do not create the Container App by hand with `az containerapp create --ingress external` and no authentication: the server checks no incoming token, so that endpoint would answer anyone on the internet.

For `service_principal` mode, inject `BIPIXIE_MCP_CLIENT_SECRET` as a Key Vault reference rather than a plaintext env var.

The managed identity assigned to the Container App must be added as **Member** (or Admin) on the target Power BI workspace.

### Sign-in at the edge (Container App built-in authentication)

The server does not validate incoming tokens itself. The template turns on the Container App's built-in authentication (`Microsoft.App/containerApps/authConfigs`), which checks every request before it reaches the server:

- Requests without a valid token get **401**; no tool runs.
- Identity provider: Microsoft Entra ID, `clientId` = the sign-in app registration.
- `allowedAudiences`: both `<clientId>` and `api://<clientId>`.
- `openIdIssuer`: `https://login.microsoftonline.com/<tenantId>/v2.0` (v2.0 tokens only; the login host follows the Azure cloud you deploy to).
- `BIPIXIE_MCP_PII_COLUMNS_ALLOWED` defaults to `false` on the hosted template.

Callers request a token with the scope `api://<clientId>/.default` and send it as `Authorization: Bearer <token>`.

### Register in Claude Code (remote)

```bash
claude mcp add \
  --transport http \
  --header "Authorization: Bearer <token>" \
  bipixie \
  https://<your-app>.azurecontainerapps.io/mcp
```

### Register in Azure AI Foundry

1. Open your Azure AI Foundry project.
2. Navigate to **Tools > Add tool > Custom > Model Context Protocol**.
3. Enter:
   - **Endpoint URL**: `https://<your-app>/mcp`
   - **Authentication**: Microsoft Entra
   - **Type**: Project Managed Identity (machine-to-machine) or OAuth Identity Passthrough (per-user delegated tokens). Because the BI Pixie semantic model ships no RLS, both types see the same full usage data — choose based on your operational preference, not for data-access reasons.
   - **Audience**: the Application ID URI of the sign-in app registration, `api://<clientId>` (the template's `tokenAudience` output)
4. Test connectivity. The tool list should populate with the eight BI Pixie tools.

Note the 100-second non-streaming timeout that Azure AI Foundry enforces on MCP tool calls. All BI Pixie tools complete well within this limit for typical query sizes.

### Register as a Fabric data agent tool

In the Fabric data agent builder, add an MCP tool pointing at the `/mcp` endpoint with Microsoft Entra auth. The data agent will use the eight typed tools alongside (or instead of) the model's built-in VerifiedAnswers and CopilotTooling conversational path. The two are complementary: the custom MCP server provides structured JSON output with filter parameters and pagination; the model's native Fabric Copilot integration provides natural-language Q&A over the existing 39 VerifiedAnswers.

---

## Tool catalog

Call `describe_model` first — it returns the full queryable surface so the agent knows which table names, measures, and dimension columns are available before constructing a `run_dax` query.

| Tool | Required args | Optional args | What it returns |
|------|--------------|---------------|----------------|
| `describe_model` | — | `include_measure_descriptions` | Dataset ID/name, table list, measure catalog by domain, dimension columns, data freshness |
| `list_reports` | — | `days`, `top_n`, `workspace_name` | Tracked reports with sessions, total hours, unique users, last activity |
| `top_reports_by_usage` | — | `metric`, `days`, `top_n`, `ascending` | Reports ranked by `report_sessions` / `total_hours` / `users` / `interactive_sessions` / `avg_session_duration` / `interactions` (use `interactions` for "most/least clicks") |
| `report_health` | `report_name` | `days` | Full health card: sessions, hours, avg duration, interactive vs passive, users, CSAT (last + multiple), NPS |
| `user_adoption` | — | `days`, `granularity`, `report_name` | MAU, DAU, DAU/MAU, WAU, Engaged Users, New Users, Returning Users — summary or time series |
| `page_engagement` | — | `report_name`, `days`, `top_n` | Per-page: page sessions, avg duration, total interactions, avg interactions/session, slicer clicks, visual interactions, tooltip opens |
| `feedback_summary` | — | `days`, `report_name`, `group_by`, `granularity`, `survey_type`, `top_n` | CSAT (both `csat_multiple` — the report's headline, click-weighted — and `csat_last`), the sentiment split (positive/negative clicks, neutral users, response rate), NPS (score, rating, promoters/detractors/passives), and the survey story (respondents, responses, time savings, self-reported **financial gains $**). `group_by` (`report`/`icon`/`workspace`/`question`/`answer`/`survey_type`) breaks it out (domain-aware — only the measures that vary across the dimension); `granularity` (`daily`/`weekly`/`monthly`) returns a trend; an empty window self-heals with a `data_coverage` hint. `group_by_report` kept as a deprecated alias for `group_by="report"`. |
| `field_usage` | — | `report_name`, `page_name`, `field_kind`, `field_name`, `field_table`, `days`, `ascending`, `top_n` | Semantic model columns and measures used by report visuals: visuals using each field and its clicks (a column's data-auditing selections, a measure's hosting-visual clicks). With `field_name`, also the report / page / visual / field-well role hosting it |
| `unused_fields` | — | `report_name`, `page_name`, `field_kind`, `days`, `top_n` | Fields on report visuals with zero clicks in the window, ranked by how many visuals use them, plus unused column and measure counts |
| `audit_overview` | None | `group_by`, `interaction`, `days`, `report_name`, `workspace_name`, `top_n` | Data auditing: which tables, columns or values of the audited reports' semantic models users view or click most, ranked by `Views or Clicks`, with the last selection date |
| `audit_column_values` | `filtered_table`, `filtered_column` | `interaction`, `days`, `report_name`, `workspace_name`, `top_n` | The values of one audited column, each with its view or click count, largest distinct count and `interaction_class` (`click` or `view`) |
| `audit_value_lookup` | `values` | `match`, `interaction`, `days`, `report_name`, `top_n` | Who viewed or clicked a value, and when and where: user key, date, report, page, table, column, value. `match="contains"` is flagged `approximate` |
| `audit_resolve_entity` | `entity` | `days`, `top_n` | Audited table and column names that contain a word such as "product", ranked by activity, with a hint naming the next tool to call |
| `audit_user_activity` | `user_key` | `interaction`, `days`, `top_n` | One user's data selections by date, report and page |
| `ai_readiness_scores` | — | `band`, `workspace_name`, `top_n` | Each semantic model's newest AI Readiness assessment, ranked by score: band, when it ran, how many assessments, the change since the previous and the first, open and high findings. Answers `available: false`, with the reason, when the semantic model holds no AI Readiness results |
| `run_dax` | `dax` | `row_limit` | Escape hatch: execute any read-only `EVALUATE` DAX query and get rows as JSON |

**Data auditing returns your business data.** The `audit_*` tools return `Filtered Values`, the actual values people selected in your reports (a customer name, a region, a price). Treat those results as potentially sensitive business data, just as you would the reports themselves. `BIPIXIE_MCP_PII_COLUMNS_ALLOWED` governs user identity columns only; it does not hide `Filtered Values`. Two rules apply to every audit result, and `describe_model` returns both under `auditing_heuristics`: a row whose distinct count is 1 is a click on one value, and a row whose distinct count is more than 1 is a view whose `Filtered Values` is a span (for example `CityA - CityZ`), so whether one value inside that span was seen is approximate.

**AI Readiness results.** An updated BI Pixie Dashboard carries 17 `AI` tables: semantic model and data agent assessments, their findings, benchmarks and A/B tests. The tables are there even when they are empty, so `describe_model` checks whether they hold any results. When they do, `describe_model` returns an `ai_readiness` block naming the tables present, their measures and how to read them, so an assistant can answer "which of our semantic models are ready for AI?" with `ai_readiness_scores` and questions about history with `run_dax`. A semantic model without the tables, or with the tables but no results, gets no such block. Three things hold for these results:

- **Data residency.** They are read through your semantic model, which reads your own Lakehouse. With data residency on, the results live in that Lakehouse, and this server never reads them from BI Pixie's storage.
- **No answers or DAX.** No table holds an AI's answer to a benchmark question, a correct value or a DAX query, so no tool can return one.
- **Read only.** The usage tools cannot start an assessment or a benchmark, whatever is asked of them. The management tools can, through the BI Pixie API on your own machine, and each of those asks you first.

### Local navigation tools (stdio only)

On the local **stdio** transport, the server also registers four navigation tools so you can point it at a different workspace/dataset without editing config or restarting — handy when one identity can see several `BI Pixie` models. They are enabled by default (`BIPIXIE_MCP_ALLOW_TARGET_SWITCHING=true`); set it to `false` to lock the server to the configured dataset. They are **never** exposed on the hosted (streamable-http) transport, which keeps its single-configured-dataset contract.

| Tool | Required args | What it does |
|------|--------------|--------------|
| `list_workspaces` | — | List the Power BI workspaces the server identity can see (read-only, one `GET /groups`). |
| `list_datasets` | `workspace_id` | List datasets in a workspace, flagging likely `BI Pixie` models via `is_bi_pixie`. |
| `set_active_dataset` | `workspace_id`, `dataset_id` | Re-point every tool at a different workspace+dataset for the rest of the session (resets on restart). Not read-only — it changes session state. |
| `get_active_dataset` | — | Show the workspace+dataset the tools are currently querying, with its freshness footer. |

### Freshness footer

Every tool response includes a freshness footer:

```json
{
  "data_through": "2026-05-29",
  "last_refresh_utc": "2026-05-30T04:12:00Z"
}
```

`data_through` is the `[Last Activity]` date in the model. `last_refresh_utc` is the most recent successful refresh from the Power BI refresh history API. Because the model is IMPORT mode, data is current as of the last scheduled refresh — not real-time.

### `run_dax` safety

`run_dax` passes every query through `DaxGuard` before dispatch:

- Query must begin with `EVALUATE` (a `DEFINE ... EVALUATE ...` preamble is allowed).
- The following write verbs hard-reject the query and return `{"error": "dax_rejected", "reason": "..."}`: `CREATE`, `ALTER`, `DELETE`, `DROP`, `MERGE`, `INSERT`, `UPDATE`, `PROCESS`, `REFRESH`, and related DDL keywords.
- The agent cannot change the workspace or dataset ID — the server is scoped to the single configured tenant context.
- Results pass `ResponsePiiFilter`, which is a pass-through by default and strips `Username` / `Client IP` columns when `BIPIXIE_MCP_PII_COLUMNS_ALLOWED=false`.
- Row count is capped at `BIPIXIE_MCP_MAX_ROWS`.

---

## Security checklist

- **Never commit** `.env` or any file containing `BIPIXIE_MCP_CLIENT_SECRET`. Add `mcp_server/.env` to `.gitignore`.
- **The BI Pixie semantic model ships no row-level security (RLS).** Every auth mode therefore sees the same full usage data — this is by design. Each team or Enterprise tenant installs the BI Pixie Dashboard under its own license key and container and runs its own MCP server against its own model, so anyone who can run the server is already entitled to that model's usage data. The PII filter (`BIPIXIE_MCP_PII_COLUMNS_ALLOWED`, default `true`), not RLS, is the control for user-identifying columns (`Username` / `Client IP`); it passes them through unless you set the flag `false`.
- **`run_dax` is read-only** — `DaxGuard` enforces this before every dispatch. No write, DDL, or refresh operation can reach the dataset.
- **Client secrets never appear in logs.** A redacting filter masks UUIDs and 40-plus-character hex strings from all log output. All logs go to stderr; stdout is reserved for the JSON-RPC stream.
- **Saved sign-in** (`device_code` / `interactive`) contains a refresh token. Protect it like a password. On Windows it is `%LOCALAPPDATA%\.IdentityService\bipixie-mcp.nocae`, encrypted for your Windows account. On macOS the secret is in your login Keychain, with a marker file at `~/.IdentityService/bipixie-mcp.nocae`. On Linux it is `~/.IdentityService/bipixie-mcp.nocae`, encrypted with the desktop keyring when one is available and saved as plain text when it is not (for example over SSH or in a container). On Linux without a keyring, protect the file with `chmod 600`. Delete it to sign out. `BIPIXIE_MCP_TOKEN_CACHE_PATH` is deprecated and does not move it. Nothing is saved in `managed_identity` mode.
- **PII in the model:** the `User and IP Addresses` table contains plaintext UPNs and client IPs, and by default they reach the caller. This is deliberate: the model is the customer's own usage data, and stripping columns the caller asked for made the server disagree with the dashboard and the Fabric data agent. Set `BIPIXIE_MCP_PII_COLUMNS_ALLOWED=false` to re-enable `ResponsePiiFilter` on a shared host, or where a data-governance review calls for it.
- **Multi-tenant cloud hosting:** if the hosted server runs in the vendor Azure subscription and serves multiple customers, each customer should have its own service principal (separate `BIPIXIE_MCP_CLIENT_ID` / `BIPIXIE_MCP_CLIENT_SECRET`) so the ~120 req/min per-identity Power BI quota is not shared.

---

## Tenant settings checklist

Complete this checklist before connecting the server. All items require a Power BI admin.

```
[ ] Power BI Admin portal > Tenant settings > Integration settings:
    "Dataset Execute Queries REST API" = ENABLED
    (disabled = silent 403; this is the most common setup failure)

[ ] Power BI Admin portal > Tenant settings > Developer settings:
    "Allow service principals to use Power BI APIs" = ENABLED
    (required for service_principal and managed_identity auth modes)

[ ] Fabric workspace > Manage access:
    Service principal / managed identity added as Member or Admin
    (grants Read + Build on the dataset)

[ ] Entra app registration:
    API permission "Power BI Service > Dataset.Read.All" granted and admin-consented
    requestedAccessTokenVersion = 2  (v2 tokens required)

[ ] Hosted only: sign-in app registration for the endpoint
    Application ID URI = api://<clientId>, requestedAccessTokenVersion = 2
    Its application ID passed to the template as clientId (required)
```

---

## Architecture notes

### Why the existing semantic model (Phase 1)

The "BI Pixie" semantic model ships ~60 tables, a curated DAX measure library covering adoption, engagement, feedback, survey, and security, 39 VerifiedAnswers, and CopilotTooling annotations. The MCP server's typed tools wrap these confirmed measures — `[Report Sessions]`, `[Total Hours]`, `[Avg Session Duration (sec)]`, `[Interactive Sessions]`, `[MAU]`, `[DAU]`, `[DAU / MAU]`, `[Engaged Users]`, `[New Users]`, `[Returning Users]`, `[CSAT (Last Response)]`, `[CSAT (Multiple Responses)]`, `[Feedback Clicks]`, `[Positive Clicks]`, `[Negative Clicks]`, `[Average Satisfaction]`, `[Neutral Users]`, `[Respondents]`, `[% Feedback Responses]`, `[NPS]`, `[NPS Rating]`, `[NPS Promoters]`, `[NPS Detractors]`, `[NPS Passives]`, `[Survey Respondents]`, `[Survey Responses]`, `[Survey Time Savings (Hours)]`, `[Survey Financial Gains]`, `[Survey Avg Financial Gains]`, `[Page Sessions]`, `[Total Interactions Within Page]`, `[Slicer Clicks]`, `[Visual Interactions]`, `[Tooltip Opens]`, `[Avg Duration in Page (Sec)]` — so agent numbers match the Power BI report exactly.

The `executeQueries` endpoint is a public REST endpoint available on all Power BI tiers (Pro, PPU, Premium, Fabric). Aggregated measure queries return tens of rows, comfortably inside the 100K-row / 1M-value / 15 MB / 120-rpm caps.

### Package layout

```
mcp_server/
  bipixie_mcp/
    __init__.py       # package marker, lazy public re-exports: __version__, Settings, get_settings, build_server, main
    config.py         # Settings (pydantic-settings, BIPIXIE_MCP_* prefix), get_settings(), ConfigError
    auth.py           # TokenProvider, build_credential(), AuthError
    powerbi_client.py # DaxGuard, ResponsePiiFilter, PowerBIClient, QueryResult, ModelMetadata, McpSecurityError
    tools.py          # register_tools(), build_*_dax() helpers, METRIC_MEASURE_MAP, MEASURE_CATALOG
    onboarding.py     # first-run setup wizard (TTY-gated): az/device-code bootstrap, workspace/dataset pickers, .env writer
    server.py         # build_server(), main() — composition root and entrypoint
  tests/
    test_tools.py     # pytest suite (no live Azure/Power BI — uses respx + AsyncMock)
  examples/
    agent-config.md   # copy-paste registration recipes for Claude Code, Codex, Foundry
  pyproject.toml      # PEP 621 metadata, console-script bipixie-mcp = bipixie_mcp.server:main
  .env.example        # full BIPIXIE_MCP_* env template
  README.md           # this file
```

### Module contracts

- `config.Settings` — loaded once via `get_settings()` (lru_cache singleton). No network calls, no credentials built at import time. `validate_runtime()` fails fast with a human-readable `ConfigError` (no secrets in the message) before any network I/O.
- `auth.TokenProvider` — caches the bearer token process-wide, transparently refreshing ~5 minutes before the ~1-hour expiry. Never calls Power BI per-request.
- `powerbi_client.PowerBIClient` — async; `execute_dax()` enforces DaxGuard, rate limit, `max_rows`, error-envelope parsing, and `ResponsePiiFilter`. `resolve_ids()` resolves workspace/dataset names to GUIDs once and caches the result.
- `tools.register_tools()` — decorates the usage tools on the `FastMCP` instance; `management.register_management_tools()` adds the management family on `stdio` for a person's sign-in, each tool one `BiPixieApiClient` call, with the API's refusals passed through as the tool's error. DAX builders (`build_*_dax`) are pure functions with no I/O, independently unit-tested. `report_name` and other user-supplied filter values are passed via `TREATAS({value}, table[column])` — never string-concatenated into a measure name — preventing DAX injection.
- `onboarding.run_onboarding()` — interactive first-run setup. All terminal I/O is injectable (`_Console`) and all network goes through the existing `PowerBIClient` discovery surface, so the flow is unit-tested with canned input + a fake client (no TTY, no Azure, no Power BI). Gated on an interactive TTY so an agent-launched server never prompts.
- `server.build_server()` — composition root: loads config, builds credential/client, registers tools, returns a configured `FastMCP` instance. `main()` is the console-script entry point and the `python -m bipixie_mcp.server` target; it runs the setup wizard first when invoked as `setup` or when config is incomplete on an interactive terminal.

### Running tests

```bash
cd mcp_server
py -3.12 -m pytest tests/ -v --tb=short
```

Tests run without any Azure or Power BI connection (no live credentials required). Fixtures use `AsyncMock` for `PowerBIClient` methods and `respx` for HTTP-layer tests.

---

## Phase 2 — Fabric RTI scale-out

When Phase 1 limits bite — `executeQueries` 429 throttling under agent load, raw row-level drill-down beyond ~50K rows, time-series anomaly detection (`series_decompose_anomalies`), sub-2-second historical scans over 30+ days, or near-real-time "who is viewing this report right now" — the preferred upgrade path is **Microsoft Fabric Real-Time Intelligence (RTI)**.

The design:
1. Tee the existing BI Pixie Event Hub into a Fabric Eventstream -> Eventhouse `Events` table (new consumer group only — `event_consumer_app` -> ADLS stays untouched).
2. Backfill history from `bipixielake-{license_key}/events/` via the Fabric "Get data from Azure Storage" TSV wizard.
3. Replace (or supplement) the Phase-1 server with the **open-source `microsoft-fabric-rti-mcp` server** (`pip install microsoft-fabric-rti-mcp`) pointing at the Eventhouse KQL endpoint — or use the zero-install `api.fabric.microsoft.com/v1/mcp/dataPlane/.../kqlEndpoint` remote MCP endpoint.

Phase 2 is future work and nothing in it is built. It also changes what agents see: `microsoft-fabric-rti-mcp` has its own tools, which answer KQL questions and are not the tools listed above, so an agent set up against today's tools would need new instructions. Keeping today's tool names would mean wrapping the KQL queries behind them in this server, which is a decision for when one of the triggers above fires.

See `project/ai-next-generation-proposal.md` (Pillar 2, "Engagement-Enriched Fabric Ontology") for strategic context.

---

## License

MIT — the full text is in the bundled `LICENSE` file. Published on PyPI as [`bipixie-mcp`](https://pypi.org/project/bipixie-mcp/); customer setup guide at [bipixie.com/docs/cloud/mcp-server](https://bipixie.com/docs/cloud/mcp-server/).

> **NOTICE:** Requires an active BI Pixie license and a deployed BI Pixie semantic model. "BI Pixie" is a trademark of DataChant. This package is a client connector and grants no rights to the BI Pixie service or its data.
