Metadata-Version: 2.4
Name: jev2semopt
Version: 0.1.0
Summary: Bounded semantic dataframe operators powered by Jev-style decisions
Project-URL: Repository, https://github.com/Qingbolan/Jev2SemOpt
Project-URL: Documentation, https://github.com/Qingbolan/Jev2SemOpt/tree/main/docs
Author-email: "Silan.Hu" <silan.hu@u.nus.edu>
License-Expression: MIT
License-File: LICENSE
License-File: NOTICE
Keywords: dataframe,jev,llm,semantic-operators
Classifier: Development Status :: 3 - Alpha
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Typing :: Typed
Requires-Python: >=3.8
Requires-Dist: llm2jev<0.3,>=0.2.1
Requires-Dist: pandas<3,>=1.5
Provides-Extra: jev
Requires-Dist: httpx<1,>=0.27; extra == 'jev'
Provides-Extra: ollama
Requires-Dist: llm2jev[ollama]<0.3,>=0.2.1; extra == 'ollama'
Provides-Extra: transformers
Requires-Dist: llm2jev[transformers]<0.3,>=0.2.1; extra == 'transformers'
Description-Content-Type: text/markdown

# Jev2SemOpt — Typed decisions over tables

**Filter tickets. Assign queues. Rank records. Match pairs.**

A support pipeline needs to keep refund requests, assign a fixed queue, and count
the resulting tickets. Jev2SemOpt makes those dataset operations explicit: select
the fields the model may see, bind a typed question once, evaluate each record,
and apply the result with ordinary Pandas operations.

Jev2SemOpt is an independent, early-stage Python library inspired by
[LOTUS semantic operators](https://github.com/lotus-data/lotus), using
[LLM2Jev](https://github.com/Qingbolan/llm2jev-releases) for Jev-style decisions.
The Jev interface concepts and `Choice`, `Score`, `Noul` vocabulary originate with
[TypeSafe](https://typesafe.ai/blog/introducing-system-one-models-and-jev).
Use a local LLM2Jev runtime or the official TypeSafe HTTP API through `JevBackend`.
The HTTP adapter is contract-tested; hosted Jev accuracy has not been measured.

## Install and run

Python **3.8+**. The core requires Pandas and LLM2Jev; model providers are opt-in.

```bash
pip install jev2semopt
# Official hosted Jev (requires a TypeSafe API key):
pip install 'jev2semopt[jev]'
# Local providers:
pip install 'jev2semopt[transformers]'
# or: pip install 'jev2semopt[ollama]'
```

For a model-free example or development:

```bash
git clone https://github.com/Qingbolan/Jev2SemOpt.git
cd Jev2SemOpt
uv sync --group dev
uv run python examples/refund_pipeline.py
uv run python examples/decorated_decision.py
```

These examples use synthetic scores and require no model or API key. Production
and development dependencies resolve from PyPI. See [release verification](docs/releasing.md).

## A refund pipeline

Configure execution explicitly. The application owns the runtime and closes it;
Jev2SemOpt only borrows it. The following requires compatible local model weights
and a model-appropriate LLM2Jev encoder/label configuration; it is an API example,
not a validated inference configuration.

```python
import pandas as pd
from llm2jev import LLM2Jev, TransformersRuntime
from jev2semopt import LLM2JevBackend, SemEngine, Choice

tickets = pd.DataFrame({
    "id": [101, 102],
    "text": ["Please refund my order.", "Where is my parcel?"],
})

with TransformersRuntime("/path/to/local/model", device="cpu") as runtime:
    engine = SemEngine(LLM2JevBackend(
        LLM2Jev(runtime=runtime, model_identity=runtime.identity),
        model=runtime.identity.name,
    ))
    refunds = engine.sem_filter(
        tickets, "Does text explicitly request a refund?",
        columns=["text"], threshold=0.7, probability_column="refund_support",
    )
    routed = engine.sem_map(
        refunds,
        Choice(instructions="Which team should handle text?", criteria={
            "billing": "Payments, invoices, or refunds",
            "delivery": "Shipping, tracking, or missing parcels",
            "other": "Requests outside billing and delivery",
        }),
        columns=["text"], output="queue",
    )
    counts = routed.groupby("queue").size()
```

Only `text` enters the decision state. IDs remain available in the output. Column
names in instructions refer to keys in that state; there is **no `{column}` string
interpolation**. Thresholds require validation on your own labeled workload.

## Operators and boundaries

Official hosted Jev is also supported through `JevBackend`, using a caller-owned
HTTP client and TypeSafe API key. See the [official API integration](docs/official-jev.md)
for setup and verification limits. An OpenRouter key does not authenticate TypeSafe.

| Operator | Jev decision | Output and scope |
| --- | --- | --- |
| `sem_filter(frame, instructions, ...)` | `Noul` | Rows with support **≥ threshold**, original order and index |
| `sem_map(frame, Choice(...), output=...)` | `Choice` | All rows plus a chosen label from supplied alternatives |
| `sem_score(frame, Score(...), ...)` | `Score` | All rows plus expected ordinal rubric level |
| `sem_topk(frame, Score(...), k=...)` | `Score` | Largest expected levels; stable source-order ties |
| `sem_join(left, right, instructions, ...)` | `Noul` | Inner join over candidate pairs; namespaced source fields and positions |

LOTUS provides a broader semantic operator model, including generated projections,
extraction, aggregation, and comparator-based ranking. Jev2SemOpt deliberately
restricts mappings to finite alternatives and ranks by an explicit ordinal rubric.
It is not a drop-in LOTUS replacement. Free-text extraction, summaries, vector
search, learned cascades, SQL planning, asynchronous execution, and distributed
execution are not implemented. Count/group/sum the typed outputs with Pandas;
there is no misleading `sem_agg` alias for a generative summary.

```python
from jev2semopt import Score

ranked = engine.sem_topk(
    tickets,
    Score(instructions="How urgent is text?", criteria=[
        "Routine request", "Time-sensitive issue", "Immediate safety or service emergency",
    ]),
    columns=["text"], k=10, output="urgency",
)

matches = engine.sem_join(
    tickets, policies,
    "Does left.text describe a case covered by right.policy?",
    left_columns=["text"], right_columns=["policy"],
    candidates=[(0, 1), (1, 0)],  # row POSITIONS, not index labels
    threshold=0.8,
)
```

The join output includes `left.<column>`, `right.<column>`, `_left_position`,
`_right_position`, and `_probability`. Without candidates it evaluates every pair,
subject to a default 100,000-pair limit. A shortlist can reduce work but can also
exclude true matches; the caller owns its recall. See [API contracts](docs/api.md).

## Python-native integration

The Pandas accessor is opt-in and uses Pandas' registration decorator. It forwards
to the same engine implementation and does not install global model settings:

```python
import jev2semopt.pandas

refunds = tickets.jev.sem_filter(
    engine, "Does text request a refund?", columns=["text"], threshold=0.7,
)
```

Use `@decision` when application code already builds the state for one decision:

```python
from jev2semopt import Noul, decision

@decision(engine, question=Noul(instructions="Does text explicitly request a refund?"))
def refund_requested(text):
    return {"text": text}

answer = refund_requested("Please refund my order.")
print(answer.noul)
```

The decorator binds once, preserves function metadata with `functools.wraps`, and
evaluates new state on each call. It does not cache results across calls. Decorated
functions must be synchronous JSON-state builders and now return typed answers.

## Measured results

On **50 balanced SciFact records**, the same GPT-4o-mini scored **31/50 with LOTUS**
and **32/50 with Jev-style / LLM2Jev**. The one-record difference does not establish
an accuracy improvement: the paired 95% interval spans −8 to +12 percentage points.
Jev-style improves precision but lowers recall and F1 in this run.

![Small-sample accuracy, precision, recall and F1 for LOTUS and Jev-style](https://raw.githubusercontent.com/Qingbolan/Jev2SemOpt/v0.1.0/docs/assets/scifact-small-sample.png)

On **400 local Qwen2.5-0.5B candidate pairs**, Jev-style filtering was 1.97× faster,
but both filters had roughly 4% precision. Three-level Jev-style scoring was 2.70×
slower than LOTUS scoring and reduced ranking quality. BM25 had the highest nDCG.

![Local ranking quality and operator time, including the BM25 baseline](https://raw.githubusercontent.com/Qingbolan/Jev2SemOpt/v0.1.0/docs/assets/scifact-local.png)

These are **official SciFact dataset subsets with adapted protocols**, not full
LOTUS paper reproduction or measurements of TypeSafe's hosted Jev model. Unjudged
documents count as negatives under the qrel convention. Prompts differ between
operators; these experiments do not isolate probability assembly as the cause.
[Sampling, confusion matrices, costs, raw observations, and reproduction](https://github.com/Qingbolan/Jev2SemOpt/blob/v0.1.0/benchmarks/README.md).

## Execution cost and score meaning

`SemEngine(backend, deduplicate=True)` reuses identical selected JSON state **within
one operation**. This is opt-in and assumes the backend is deterministic and has no
per-call side effects. There is no global cache and no stale reuse across operations.
Question binding avoids recompilation; it is not a KV cache or a batched model call.

For `N` rows, filtering requires `N` one-candidate evaluations; mapping with `C`
choices requires `N × C` binary candidates; scoring with `R` rubric levels requires
`N × R`. An exhaustive join needs `L × R` pair evaluations. Duplicate reuse reduces
these counts to unique selected states. Calls are sequential. Join candidates and
results are materialized in memory; this version targets bounded in-memory tables.

No end-to-end speedup, calibrated correctness, or equivalence to LOTUS quality is
claimed. `Noul` is label-conditioned support; `Score` is an expected equally spaced
rubric index. These are decision signals, not probabilities that the answer is right.
LLM2Jev's Transformers adapter reads next-token logits; its Ollama adapter requests
one token to obtain exact binary logprobs. Jev2SemOpt never parses generated prose.
See [evaluation protocol](docs/evaluation.md).

## Architecture and development

```text
DataFrame API / @decision
          ↓
SemEngine — positional relational semantics
          ↓
EvaluationSession — one rule, detached state, operation-local reuse
          ↓
DecisionBackend.bind → BoundDecision.evaluate → typed Answer
          ↓
├─ LLM2JevBackend → public LLM2Jev API → caller-owned runtime
└─ JevBackend → TypeSafe HTTP API → caller-owned HTTP client
```

The engine depends on a backend protocol, not Ollama or Transformers. Models,
prompts, device configuration, binary labels, and runtime lifecycle stay in
LLM2Jev for local execution; hosted execution uses the supplied HTTP client. There is no new resource lifecycle to duplicate in the table layer.

```text
src/jev2semopt/
├── contracts.py       Backend and bound-decision protocols
├── execution.py       State isolation, answer checks, safe failure boundary
├── engine.py          Filter/map/score/top-k/join semantics
├── table.py           Dataframe schema and state projection
├── decorators.py      Function-to-decision binding
├── pandas.py          Opt-in accessor registration
└── adapters/          LLM2Jev service and official Jev HTTP integration
```

Run the checks in [CONTRIBUTING.md](CONTRIBUTING.md). Architectural decisions,
privacy constraints, and delivery status are documented in
[architecture](docs/architecture.md), [privacy](docs/privacy.md), and
[implementation plan](docs/implementation-plan.md). Attribution is recorded in
[NOTICE](NOTICE). Licensed under [MIT](LICENSE); [dependency and benchmark attribution](docs/licenses.md).

[Local verification record](docs/verification.md): 50 deterministic tests passed
on Python 3.8, 3.12, and 3.14. Real-model observations are recorded separately in
the [benchmark report](benchmarks/README.md).
