Metadata-Version: 2.5
Name: clousight-bench
Version: 0.7.0
Summary: Clousight Bench: reproducible, reproducibility-classed benchmarking for cloud products.
Project-URL: Homepage, https://clousight.com
Project-URL: Repository, https://github.com/clousight/clousight-bench
Project-URL: Issues, https://github.com/clousight/clousight-bench/issues
Project-URL: Changelog, https://github.com/clousight/clousight-bench/blob/main/CHANGELOG.md
Author: Clousight Bench contributors
License: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: agent-runtime,benchmark,cloud,clousight,reproducibility
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: System :: Benchmark
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: opentelemetry-api<2,>=1.44
Requires-Dist: opentelemetry-sdk<2,>=1.44
Requires-Dist: pyyaml>=6.0
Provides-Extra: agent
Requires-Dist: langchain-core<2,>=1.2.22; extra == 'agent'
Requires-Dist: langchain<2,>=1; extra == 'agent'
Requires-Dist: openinference-instrumentation-langchain>=0.1.71; extra == 'agent'
Requires-Dist: opentelemetry-api; extra == 'agent'
Requires-Dist: opentelemetry-exporter-otlp-proto-http; extra == 'agent'
Requires-Dist: opentelemetry-sdk; extra == 'agent'
Provides-Extra: aliyun
Requires-Dist: alibabacloud-agentrun20250910; extra == 'aliyun'
Requires-Dist: alibabacloud-arms20190808; extra == 'aliyun'
Requires-Dist: alibabacloud-credentials; extra == 'aliyun'
Requires-Dist: alibabacloud-eci20180808; extra == 'aliyun'
Requires-Dist: alibabacloud-ecs20140526; extra == 'aliyun'
Requires-Dist: alibabacloud-tea-openapi; extra == 'aliyun'
Requires-Dist: alibabacloud-vpc20160428; extra == 'aliyun'
Requires-Dist: oss2; extra == 'aliyun'
Requires-Dist: requests>=2.28; extra == 'aliyun'
Provides-Extra: aws
Requires-Dist: boto3>=1.34; extra == 'aws'
Requires-Dist: requests>=2.28; extra == 'aws'
Provides-Extra: dev
Requires-Dist: mypy>=1.10; extra == 'dev'
Requires-Dist: pre-commit>=3.7; extra == 'dev'
Requires-Dist: pytest-cov>=5.0; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff==0.16.6; extra == 'dev'
Requires-Dist: types-pyyaml; extra == 'dev'
Provides-Extra: llm
Requires-Dist: requests>=2.28; extra == 'llm'
Provides-Extra: otlp
Requires-Dist: opentelemetry-exporter-otlp-proto-http<2,>=1.44; extra == 'otlp'
Provides-Extra: probe
Requires-Dist: oss2>=2.18; extra == 'probe'
Requires-Dist: requests>=2.28; extra == 'probe'
Provides-Extra: store
Requires-Dist: duckdb>=1.0; extra == 'store'
Requires-Dist: pyarrow>=16; extra == 'store'
Provides-Extra: swebench
Requires-Dist: swebench>=3.0; extra == 'swebench'
Provides-Extra: tpcds
Requires-Dist: duckdb>=1.0; extra == 'tpcds'
Provides-Extra: tpch
Requires-Dist: duckdb>=1.0; extra == 'tpch'
Provides-Extra: validate
Requires-Dist: jsonschema>=4.20; extra == 'validate'
Description-Content-Type: text/markdown

# Clousight Bench

[![CI](https://github.com/clousight/clousight-bench/actions/workflows/ci.yml/badge.svg)](https://github.com/clousight/clousight-bench/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/clousight-bench.svg)](https://pypi.org/project/clousight-bench/)
[![Python](https://img.shields.io/pypi/pyversions/clousight-bench.svg)](https://pypi.org/project/clousight-bench/)
[![License: Apache-2.0](https://img.shields.io/badge/License-Apache_2.0-blue.svg)](LICENSE)
[![Docs](https://img.shields.io/badge/docs-docs.clousight.com-blue.svg)](https://docs.clousight.com)
[![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/clousight/clousight-bench)

**[Clousight](https://clousight.com) Bench is an eval tool for cloud computing** —
it measures **cloud products and services** with **recognized benchmark suites, unmodified**,
across five product categories today: agent runtimes (SWE-bench Verified/Lite/Multimodal), data warehouses
(TPC-H/TPC-DS), transactional databases (TPC-C), key-value stores (YCSB), and LLM
endpoints (MMLU/GSM8K/HumanEval). Every published number stands on three legs:

1. **The verdict comes from the recognized suite's own harness** — never re-scored,
   never blended.
2. **The cloud dimensions ride alongside** — latency, cost, provisioning/teardown
   behavior, trajectory — as separate measurements that can never change the verdict.
3. **A verifiable provenance chain** — which suite, which pinned dataset revision,
   which evaluator, which scaffold — folded into a content fingerprint you can diff.

> **0.7.0 Developer Preview.** The whole pipeline runs locally with no cloud account
> (`mode: mock`), and every data/LLM suite also runs against a real local engine or a
> live endpoint you already operate (config-connect). On the provisioning side, the
> Aliyun AgentRun adapter is `experimental` and live-validated (`cn-hangzhou`
> real-cloud campaigns); the docker-capable ECS driver host, the AgentRun SUT agent
> (oracle/llm modes) and the SWE-bench Verified suite are code complete — the first
> live SWE-bench smoke is gated only on account preconditions (see the
> [live runbook](docs/swe-bench-live-runbook.mdx)). Other clouds are skeletons.

**Repository status.** This repository is public and Apache-2.0 licensed.
`main` is protected: every change lands through a pull request that passes
ruff, pytest and the no-cloud smoke on Python 3.10–3.13 plus a separate
installed-wheel smoke, and CodeQL + dependency-review workflows gate security.
No approving review is required, force pushes and branch
deletion are blocked, and the rules bind administrators too. Commercial plugins
are developed in a separate private repository and are not required to run
anything in this one.

## Quick start (no cloud account, no docker)

```bash
git clone https://github.com/clousight/clousight-bench.git && cd clousight-bench
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"

# what is installed? (registered suites, evaluators, adapters)
.venv/bin/csbench list

# run SWE-bench Verified through the full stage machine in mock mode:
cat > mock.yaml <<'EOF'
target:
  mode: mock
EOF
.venv/bin/csbench run --domain agent-runtime --benchmark swe-bench \
    --platform local-sim --config mock.yaml

# open the web viewer: watch a run while it happens, then read the results (EN | 中文)
.venv/bin/csbench serve
```

The mock run exercises the real code path — the four lifecycle phases *prepare →
connect → measure → conclude* (twelve stages; see
[Architecture](docs/architecture.mdx)) — over the suite's bundled fixture
artifacts, and persists a schema-0.4
record with `swe-bench.resolved` and a full provenance block. `csbench serve` renders
it at `http://127.0.0.1:8787` (React UI, bilingual, dark/light, strict CSP). Leave it
open while a longer run goes and it shows that run's stages, waterfall and log as they
happen; every run — not just the ones that emit a trajectory artifact — gets a trace.
See [docs/viewer.mdx](docs/viewer.mdx).

One run is not a measurement — repeat and pool:

```bash
csbench run --domain agent-runtime --benchmark swe-bench --platform local-sim \
    --config mock.yaml --repeat 5 --warmup 1
```

Only runs sharing a `benchmark` **and** `environment` fingerprint are ever pooled.

## The reproducibility contract (read this first)

Every schema-0.4 record is attributable on independent axes, so you can tell whether
two numbers are even comparable:

| Field | Answers |
|---|---|
| `provenance` | *which suite* — suite id + pinned dataset revision, evaluator id, `unmodified` flag, scaffold |
| `fingerprints.benchmark` | *what* was measured — task, suite revision, dataset digest, controlled params |
| `fingerprints.environment` | *where* — region, mode (`mock`/`real`), environment facts |
| `fingerprints.implementation` | *which code* — core, domain pack, adapter, installed plugins |
| `fingerprints.record_digest` | the content digest of the record itself |

Each measurement carries `value`, `unit`, a `reproducibility_class`
(`deterministic` / `environmental` / `judge-based`) and an `official` flag —
official measurements are the upstream suite's own verdict under the suite's
namespace (`swe-bench.resolved`); custom evaluators are structurally confined to
their own namespace (`csbench conformance --suite` enforces it). A run ends in
exactly one `status`: `completed`, `failed`, `invalid` or `unsupported` (plus
`interrupted`, which a Ctrl-C or a cancel requested from the viewer writes) — there is
no boolean `ok`, because "the platform does not support this" and "the run crashed"
are different results. We publish **per-dimension results, never a single blended
score**.

## Architecture

```
BenchmarkSuite.resolve → prepare → run   (the suite's OWN upstream harness, unmodified)
        │                        │
        │                 SUT invocation (cloud agent runtime; real trajectory + tokens)
        ▼                        ▼
Evaluator.evaluate(RawArtifacts) — pure, offline, namespaced Measurements
        ▼
schema-0.4 record  →  csbench serve (web viewer: live + results)  /  csbench query (SQL)
```

The core only orchestrates the lifecycle. Everything product- or suite-specific is a plugin:

| Plugin | One per | Examples |
|---|---|---|
| **BenchmarkSuite** | recognized suite | `swe-bench` (SWE-bench Verified, pinned HF revision) |
| **Evaluator** | scoring view | `official-swe-evaluator` (pure passthrough of the upstream verdict) |
| **DomainPack** | product category | `agent-runtime` (the cloud-infra shell: adapters, probes, reaper) |
| **ProviderAdapter** | (domain, cloud) | `local-sim`, `aliyun-agentrun`, `aws-agentcore`, … |
| **WorkloadEngine** | load generator | any language: `manifest.yaml` + executable + JSONL on stdout |

All of them register via entry points (`clousight_bench.benchmark_suites`,
`.evaluators`, `.domains`, `.runtime_providers`; plus `.metrics`, `.judges`,
`.enrichers`, `.resource_reapers`, `.span_exporters` for the finer seams) —
third-party packs install like any Python package and appear in `csbench list`.

| Adapter | Status | Runnable |
|---|---|---|
| `local-sim` | reference | yes |
| `aliyun-agentrun` | experimental | preview (live-validated) |
| `aws-agentcore` | skeleton (provider in-tree) | mock |
| `huawei-agentarts` | skeleton | mock |
| `volcengine-agentkit` | skeleton | mock |

`skeleton` clouds run end-to-end in `mode: mock` with no account and become live-runnable
when a **runtime provider** registers via `clousight_bench.runtime_providers`.

## Benchmarking a real platform

Credentials are **never** stored in configs — the cloud's own default credential
chain (env vars / CLI profile / attached role) is reused. The real-cloud SWE-bench
path runs on a **docker-capable ECS driver host** provisioned by `csbench submit`
(OSS-only control plane, self-destructing controller, terraform backstop):

```bash
csbench init aliyun                 # scaffold a private config (auto-gitignored)
csbench doctor --config agent-runtime-aliyun.local.yaml
csbench submit configs/swe-bench-smoke.plan.yaml --config agent-runtime-aliyun.local.yaml
csbench status <campaign-id> --config ...   # then: logs / fetch / teardown
```

The step-by-step live runbook — preconditions, cn-region gotchas (docker registry
mirror, HF mirror), the live-verification checklist and expected outcomes — is
[docs/swe-bench-live-runbook.mdx](docs/swe-bench-live-runbook.mdx) (EN | 中文).
Adapters surface the runtime's own behavior and must **never** touch suites or
scoring. You pay your own cloud bill; you get numbers for your own account,
network and region. That is the point.

## Analysis & viewing

```bash
csbench serve                 # web viewer: watch a run live; board → suite comparison → record → trace
csbench query "SELECT platform, avg(value_num) FROM measurements WHERE name='swe-bench.resolved' GROUP BY platform"
csbench export measurements --out m.parquet   # optional [store] extra: Parquet + DuckDB
csbench trace list|show|import                # per-run traces; import external OTLP/JSONL
csbench verify <record>                       # record-digest integrity check
```

Cost is presented as **list → discount → net** (public price feed via
`CLOUSIGHT_PRICING_DATA`, private discounts via `CLOUSIGHT_PRICING_DISCOUNTS`;
see [docs/querying.mdx](docs/querying.mdx)). Operator-supplied `system_prices`
entries in the same feed yield price/performance composites
(`extensions.pricing.price_performance`, e.g. price per QphH) — never invented,
additive-only.

## Status

- [x] Core: lifecycle orchestrator, `RunSpec`/`ResultRecord` schema 0.4 with provenance-folded fingerprints, entry-point plugin registry, cross-language workload protocol, DuckDB-backed `csbench query`, cost budget + live-run gate + resource reaper (`csbench sweep`)
- [x] **Suite contract (Sub-project B)**: `BenchmarkSuite`/`Evaluator` ABCs, `suite:<id>` runs, SWE-bench Verified at a pinned HF revision with real gold-patch fixtures, official evaluator + namespace conformance, real SUT invocation on Aliyun AgentRun (oracle/llm agent modes) with real trajectory + token capture
- [x] **Driver host (Sub-project A)**: docker-capable ECS controller (`csbench submit`), suite-aware LaunchSpec, OSS-only control plane, self-destruct reaper
- [x] **Web viewer (Sub-project C)**: `csbench serve` — React UI (prebuilt, shipped in the wheel). Watch a run **while it happens** (SSE-fed stages, live waterfall, log tail, stop button) off a disposable progress plane under `results/.progress/`; read finished runs through a domain board → suite comparison → record → trace, with every run getting a trace (run-trace fallback) and internal vocabulary rendered in plain language. EN | 中文, dark/light, strict CSP, offline-first. See [docs/viewer.mdx](docs/viewer.mdx)
- [x] **OLAP suites (`data-warehouse` domain)**: TPC-DS **and** TPC-H on a `duckdb-local` reference platform. Both run offline (`suite:tpc-ds` / `suite:tpc-h`, mock + real DuckDB); correctness vs SF-keyed verified references (SF 0.01/0.1/1 for TPC-H, verified against DuckDB `tpch_answers()`), honest per-query latency, plus a `mode: official` official-formula mode computing QphH@Size / QphDS@SF via a Load/Power/Throughput/ACID phase machine — unaudited, no TPC audit claimed. See [docs/tpch-suite.mdx](docs/tpch-suite.mdx) / [docs/tpcds-suite.mdx](docs/tpcds-suite.mdx)
- [x] **Key-value domain + config-connect abstraction**: **YCSB** on a `key-value` domain — the SUT-connection abstraction generalized so a suite runs against a local reference (`ycsb-local`, binding=basic) or an **already-running service via config** (`ycsb-endpoint`, binding+endpoint). Wraps the recognized upstream YCSB tool; offline mock path in CI, honest throughput + tail-latency (environmental). See [docs/ycsb-suite.mdx](docs/ycsb-suite.mdx)
- [x] **OLTP domain**: **TPC-C via BenchBase** on a `transactional-db` domain — `benchbase-local` (dbtype=sqlite reference) or `jdbc-endpoint` (config-connect to an already-running database). Wraps the recognized upstream BenchBase tool (Apache-2.0); offline mock path in CI, honest throughput/goodput/latency plus a labeled tpmC-style estimate (`tpc-c.tpmc_estimate`, derived from goodput × NewOrder mix) and `tpc-c.goodput_ratio` (environmental; audited tpmC not claimed). See [docs/tpcc-suite.mdx](docs/tpcc-suite.mdx). Data-systems coverage is now OLAP + KV + OLTP.
- [x] **LLM domain (test the managed model itself)**: **MMLU** on an `llm` domain — the SUT is a managed LLM endpoint (Bedrock/DashScope/Vertex/any OpenAI-compatible), config-connected via `llm-endpoint` (base_url + model + credentials) or the offline `llm-mock` reference. Runs recognized benchmarks unmodified (**MMLU** + **GSM8K** + **HumanEval**) → objective accuracy (deterministic) + serving dimensions latency/tokens/cost (environmental). See [docs/mmlu-suite.mdx](docs/mmlu-suite.mdx) / [docs/gsm8k-suite.mdx](docs/gsm8k-suite.mdx) / [docs/human-eval-suite.mdx](docs/human-eval-suite.mdx)
- [x] **pytest & CI gating**: any suite runs as a native pytest test (`assert_run` / the `clousight` fixture, auto-loaded via the `pytest11` entry point) or a CI exit-code gate (`csbench run --assert`), with min/max thresholds per measurement — so a benchmark becomes a red/green check in an enterprise CI pipeline. See [docs/pytest-ci.mdx](docs/pytest-ci.mdx)
- [x] **OTel-native tracing**: OTel SDK in core, span schema v3 (gen_ai/db semconv, v2 accepted), per-run trace with stage spans, `csbench trace list/show/import` (external OTLP/JSONL ingest), measurements-as-gauges + findings-as-logs via the `[otlp]` extra / `CLOUSIGHT_OTLP_ENDPOINT`, plugin API 3.0. See [docs/tracing.mdx](docs/tracing.mdx)
- [x] **Official TPC modes**: `mode: official` QphH@Size (TPC-H) and QphDS@SF (TPC-DS) on a shared engine-agnostic phase machine (Load / Power+refresh / multi-stream Throughput / ACID); official-formula, explicitly unaudited
- [x] **Reliability (R5)**: driver-side DisruptionProxy (TCP relay, `reset`/`stall` at a planned offset) for `ycsb-endpoint`, tool-evidence metrics `ycsb.error_rate` / `tpc-c.goodput_ratio` / `ycsb.completed_under_disruption`; fail-loud when the plan can't be honored
- [x] **`csbench doctor` connectivity probes**: TCP reach, Redis RESP PING, java-version gates (BenchBase ≥17 / YCSB ≥11), SSRF-guarded; mock targets skip
- [x] **Cloud-connect runbooks**: [docs/cloud-connect-kv.mdx](docs/cloud-connect-kv.mdx) + [docs/cloud-connect-rdbms.mdx](docs/cloud-connect-rdbms.mdx) with committed example profiles in `examples/cloud-connect/` (managed Redis / RDS-class endpoints)
- [ ] First **live** SWE-bench smoke on Aliyun (code complete; gated on account preconditions — see the runbook)
- [ ] Wire the remaining clouds (`huawei-agentarts` / `volcengine-agentkit` / `aws-agentcore` live paths); cloud-*provisioned* backends for the data domains (EMR/Spark, cloud DWH) at big-data scale — connecting to existing managed endpoints already ships (see the cloud-connect runbooks)
- [ ] More suites (τ-bench, Nexmark/streaming) + domain packs (streaming / graph / ml-systems)

## Contributing

Sign your commits (`git commit -s`, [DCO](https://developercertificate.org/)).
Adding a suite = one `BenchmarkSuite` + one `Evaluator` (the GSM8K suite in
`src/clousight_bench/suites/gsm8k/` is the simplest template; SWE-bench is the most
complete example; see docs/adding-a-suite); adding a platform = one adapter file + one
example config; adding a product category = one DomainPack. PRs that change suite
wiring or scoring for a shipped suite require a version bump and a changelog entry —
published numbers must stay attributable.

This checkout has no `origin` remote — commit/push/PR/merge go through `scripts/gitsync.sh` (requires the `clousight-dev` `gh` account and forces commit identity to that account's noreply email; `push` refuses `main` — land via a feature-branch PR with squash merge; run `cp .gitsync.env.example .gitsync.env` once to set the target repo).

## License

[Apache-2.0](LICENSE)
