Metadata-Version: 2.4
Name: sharp-sparc
Version: 1.0.1
Summary: Hosted statistical capability for identifying a compact predictive feature core in high-dimensional biological data.
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: httpx>=0.25.0
Requires-Dist: keyring>=24.0
Provides-Extra: dev
Requires-Dist: pytest>=8.0.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.23.0; extra == "dev"

# sharp-sparc

Hosted statistical capability for identifying a compact predictive feature core in high-dimensional biological data.

`sharp-sparc` is the thin, official Python client for SPARC. It provides reproducible statistical discovery and decision boundaries for high-dimensional biological datasets ($P \gg N$).

---

## Installation

```bash
pip install sharp-sparc
```

---

## OAuth-First Quickstart

Sign in once. `sharp-sparc` discovers the hosted Native SDK OAuth client itself;
you do not set Auth0 IDs, browser callbacks, or environment variables. The
session is stored in your operating-system credential vault; it is never written
to your project, notebook, or shell profile.

```bash
pip install sharp-sparc
sharp-sparc login
```

If the service has not enabled SDK sign-in yet, the command stops before opening
the identity provider and explains that hosted configuration is unavailable.

Then use the client without putting a credential in code:

```python
from sharp_sparc import SPARC

# 1. Initialize client; it uses the local OAuth session
sparc = SPARC()

# 2. Discover predictive core signals from tabular data
result = sparc.discover(X, y, top_k=20, approve_upload=True)

# 3. Inspect discovered core signals
for rank, signal in enumerate(result.core_signals, start=1):
    print(f"#{rank}: {signal.name} (relevance={signal.predictive_relevance:.4f})")
```

### Service-key fallback

Use `SPARC_API_KEY` or `SPARC(api_key="...")` only for CI, servers, or a client
that cannot complete OAuth. Treat it as a secret: keep it in a secret manager or
the operating-system environment, never in source control or shared configuration.

---

## Privacy Boundary and Local Preflight

SPARC enforces a strict privacy boundary:
1. **Local Metadata Inspection:** Your dataset's dimensions, non-finite cell counts, and SHA256 digest are computed **locally on your machine**. Raw cell values and feature names never leave your environment during preflight.
2. **Explicit Caller Approval:** Raw data uploads only after explicit caller approval using a server-issued signed URL (`approve_upload=True`). All remote uploads enforce TLS (`https://`) transport.

### Discovering from a Local CSV/TSV File

```python
result = sparc.discover_file(
    path="patient_cohort_rnaseq.csv",
    target_column="treatment_response",
    top_k=15,
    approve_upload=True,
)
```

---

## Granular REST Lifecycle

For production pipelines and workflow orchestrators, `sharp-sparc` exposes the full granular lifecycle:

```python
from sharp_sparc import SPARC, inspect_dataset

client = SPARC()

# 1. Check account quota
account = client.get_account_status()
print(f"Tier: {account.tier}, Runs Remaining: {account.runs_remaining}")

# 2. Inspect aggregate metadata locally
metadata = inspect_dataset("data.csv", target_index=10)

# 3. Validate against tier limits (zero raw data sent)
preflight = client.preflight(metadata=metadata)

# 4. Stream upload with signed authorization
upload = client.upload_dataset(
    file_path_or_data="data.csv",
    preflight_token=preflight.upload_token,
    metadata=metadata,
)

# 5. Submit analysis with idempotency protection
receipt = client.submit_analysis(
    dataset_ref=upload.dataset_ref,
    target_column="target",
    top_k=10,
    idempotency_key="pipeline-run-2026-08-batch-1",
)

# 6. Poll status
status = client.poll_analysis(receipt.job_id)

# 7. Export standalone Python and C++ decision trees (Pro tier)
coretree = client.export_coretree(receipt.job_id)
print(coretree.python_source)
```

---

## Async Client (`AsyncSPARC`)

```python
import asyncio
from sharp_sparc import AsyncSPARC

async def main():
    async with AsyncSPARC() as client:
        account = await client.get_account_status()
        print(f"Account Tier: {account.tier}")

asyncio.run(main())
```

---

## Canonical API Route Map

| Operation | Canonical REST Endpoint |
|---|---|
| Account Status | `GET /v1/account` |
| Preflight Validation | `POST /v1/preflight` |
| Upload Authorization | `POST /v1/datasets/uploads` |
| Direct Data Upload | `PUT /v1/datasets/uploads/{dataset_id}` |
| Upload Completion | `POST /v1/datasets/uploads/{dataset_id}/complete` |
| Submit Analysis | `POST /v1/analyses` (supports `Idempotency-Key`) |
| Poll Job Status | `GET /v1/analyses/{job_id}` |
| Job Results | `GET /v1/analyses/{job_id}/result` |
| Cancel Job | `POST /v1/analyses/{job_id}/cancel` |
| Export CoreTree Code | `GET /v1/analyses/{job_id}/coretree` |

---

## Hosted MCP Interface & Activation

SPARC is also available natively as a hosted Model Context Protocol (MCP) service over Streamable HTTP:
- **Endpoint**: `https://mcp.sharpmachine.ai/mcp`
- **Preview Activation**: Sign in at `https://mcp.sharpmachine.ai/activate` to activate a Preview account. OAuth-capable clients reconnect without a copied key; service keys are an explicit fallback for non-OAuth clients.

### Enforced Pricing & Usage Contract

- **Preview Tier (Free)**: 3 private analyses per UTC day; up to 616,100 matrix elements per analysis.
- **SPARC Pro Tier ($20/month)**: 20 private analyses per billing period; up to 2,000,000 matrix elements per analysis.
- **Zero Metering on Non-Compute Calls**: Discovery, account status, preflight validations, and cached demo retrievals are always free. Only completed private analyses decrement quota.
