Metadata-Version: 2.4
Name: rdfsolve
Version: 0.1.0
Summary: Wraps several RDF schema solver tools
Author-email: Javier Millán Acosta <javier.millanacosta@maastrichtuniversity.nl>
Maintainer-email: Javier Millán Acosta <javier.millanacosta@maastrichtuniversity.nl>
License: MIT License
        
        Copyright (c) 2024 Javier Millán Acosta
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Bug Tracker, https://github.com/jmillanacosta/rdfsolve/issues
Project-URL: Homepage, https://github.com/jmillanacosta/rdfsolve
Project-URL: Repository, https://github.com/jmillanacosta/rdfsolve.git
Project-URL: Documentation, https://rdfsolve.readthedocs.io
Keywords: snekpack,cookiecutter,rdf,sparql,schema,linked-data
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Framework :: Pytest
Classifier: Framework :: tox
Classifier: Framework :: Sphinx
Classifier: Natural Language :: English
Classifier: Programming Language :: Python
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pydantic>=2.0
Requires-Dist: requests>=2.0
Requires-Dist: httpx>=0.24
Requires-Dist: sssom>=0.4
Requires-Dist: typing-extensions>=4.0
Requires-Dist: rdflib>=6.0.0
Requires-Dist: pyoxigraph<0.6,>=0.5.11
Requires-Dist: ujson>=5.0.0
Requires-Dist: pandas<3,>=2.0.0
Requires-Dist: numpy
Requires-Dist: bioregistry>=0.13.25
Requires-Dist: linkml>=1.10.0
Requires-Dist: linkml-runtime>=1.10.0
Requires-Dist: deprecated>=1.2.13
Requires-Dist: pyld
Requires-Dist: more_itertools
Requires-Dist: tqdm
Requires-Dist: pyyaml>=6.0
Requires-Dist: pyarrow>=23.0.1
Requires-Dist: click
Requires-Dist: more_click
Requires-Dist: flask
Provides-Extra: validation
Requires-Dist: pyshacl<1,>=0.27; extra == "validation"
Provides-Extra: agents
Requires-Dist: pydantic-ai-slim[anthropic,openai]==2.0.0; extra == "agents"
Requires-Dist: python-dotenv==1.2.3; extra == "agents"
Provides-Extra: mcp
Requires-Dist: mcp==2.2.0; extra == "mcp"
Provides-Extra: tests
Requires-Dist: pyshacl<1,>=0.27; extra == "tests"
Requires-Dist: psutil>=5.9; extra == "tests"
Requires-Dist: pydantic-ai-slim[openai]==2.0.0; extra == "tests"
Requires-Dist: mcp==2.2.0; extra == "tests"
Requires-Dist: pytest; extra == "tests"
Requires-Dist: coverage; extra == "tests"
Requires-Dist: networkx>=2.8; extra == "tests"
Requires-Dist: httpx>=0.24; extra == "tests"
Provides-Extra: docs
Requires-Dist: sphinx>=8; extra == "docs"
Requires-Dist: sphinx-rtd-theme>=3.0; extra == "docs"
Requires-Dist: sphinx-click; extra == "docs"
Requires-Dist: sphinx-automodapi; extra == "docs"
Provides-Extra: dev
Requires-Dist: ruff; extra == "dev"
Requires-Dist: mypy; extra == "dev"
Requires-Dist: pandas-stubs; extra == "dev"
Requires-Dist: types-networkx; extra == "dev"
Requires-Dist: types-PyYAML; extra == "dev"
Requires-Dist: types-requests; extra == "dev"
Requires-Dist: types-ujson; extra == "dev"
Provides-Extra: web
Requires-Dist: flask>=3.0; extra == "web"
Requires-Dist: flask-cors>=4.0; extra == "web"
Requires-Dist: gunicorn>=22.0; extra == "web"
Provides-Extra: notebooks
Requires-Dist: psutil>=5.9; extra == "notebooks"
Requires-Dist: matplotlib>=3.10.7; extra == "notebooks"
Requires-Dist: seaborn>=0.11.0; extra == "notebooks"
Requires-Dist: plotly>=5.0.0; extra == "notebooks"
Requires-Dist: networkx>=2.8; extra == "notebooks"
Requires-Dist: jupyterlab>=4.0.0; extra == "notebooks"
Requires-Dist: ipywidgets>=8.0.0; extra == "notebooks"
Requires-Dist: narwhals==2.11.0; extra == "notebooks"
Requires-Dist: plotly==6.4.0; extra == "notebooks"
Requires-Dist: scipy>=1.15.3; extra == "notebooks"
Requires-Dist: kgx>=2.7.0; extra == "notebooks"
Dynamic: license-file

# rdfsolve

<p align="center">
    <a href="https://github.com/jmillanacosta/rdfsolve/actions/workflows/tests.yml">
        <img
        alt="Tests"
        src="https://github.com/jmillanacosta/rdfsolve/actions/workflows/tests.yml/badge.svg"
    /></a>
    <a href="https://pypi.org/project/rdfsolve">
        <img
        alt="PyPI"
        src="https://img.shields.io/pypi/v/rdfsolve"
    /></a>
    <a href="https://pypi.org/project/rdfsolve">
        <img
        alt="PyPI - Python Version"
        src="https://img.shields.io/pypi/pyversions/rdfsolve"
    /></a>
    <a href="https://github.com/jmillanacosta/rdfsolve/blob/main/LICENSE">
        <img
        alt="PyPI - License"
        src="https://img.shields.io/pypi/l/rdfsolve"
    /></a>
    <a href='https://rdfsolve.readthedocs.io/en/latest/?badge=latest'>
        <img
        src='https://readthedocs.org/projects/rdfsolve/badge/?version=latest'
        alt='Documentation Status'
    /></a>
</p>

Tools and MCP to retrieve RDF metadata, test endpoint availability, run batched
SPARQL queries and maintain source registries. Extract and convert schemas,
generate typed Python clients, follow links between records and derive mappings
across datasets. Keep the queries and results behind each exploration, and
export back into RDF, allowing to generate RDF subsets.

## Installation

```bash
pip install rdfsolve
```

## Quick Start

### Mine a local RDF snapshot

```python
from rdflib import Graph
from rdfsolve import SchemaMiner

graph = Graph().parse("data.ttl", format="turtle")
with SchemaMiner.from_graph(graph, strategy="one-shot", delay=0) as miner:
    schema = miner.mine(dataset_name="example")
    report = miner.last_report
print(len(schema.patterns), len(schema.collections or []))
```

Local mining uses RDFLib queries. `one-shot` sends four unpaged discovery
queries; `two-phase` batches classes. Runtime depends on graph structure;
compare report phases before choosing a strategy. Counting is enabled by
default. Set `counts=False` if the task needs model structure without
population/count enrichment; structural coverage checks still run.

### Mine around chosen resources

A large endpoint does not have to be mined completely when a client reads only
some resources. `ScopeStrategy` reads every statement of seed subjects, of a
sample of the members of chosen classes, and of the resources that chosen
predicates reach from them in one hop. The schema is small, and a new run shows
when the statements of these resources change. On Wikidata, items have no
rdf:type: they are classified by P31, and by the class that an IRI prefix gives.
`WikibaseScopeStrategy` also names each predicate after its property label
(`wdt:P50` is `author`, `p:P50` is `author statement`).

```python
from rdfsolve import SchemaMiner
from rdfsolve.mining.wikibase_strategy import WikibaseScopeStrategy

WD = "http://www.wikidata.org/entity/"
strategy = WikibaseScopeStrategy(
    [WD + "Q42"],
    follow=["http://www.wikidata.org/prop/P50"],
    membership="http://www.wikidata.org/prop/direct/P31",
    prefix_classes={WD + "Q": "http://wikiba.se/ontology#Item"},
)
with SchemaMiner("https://query.wikidata.org/sparql", strategy=strategy, counts=False) as miner:
    schema = miner.mine(dataset_name="wikidata-around-q42")
```

Local SPARQL uses Oxigraph by default. Choose `local_backend="rdflib"` on
`SchemaMiner.from_graph(...)`, `Client(...)`, `Client.open(...)`, or
`MinedSchema.from_void(...)` to use RDFLib. Graph Store downloads use the
same option on `SchemaMiner(...)`. HTTP endpoints use their own query engine.

Oxigraph reads a snapshot of the supplied graph or dataset. Reopen the miner
or client after changing the input. Named graphs, blank-node identifiers and
literal forms are retained. If Oxigraph storage would change terms, or named graphs share triples, RDFLib
executes the queries and a warning explains why. The overlap fallback preserves
set-based RDF merge counts across multiple `FROM` clauses. Mining reports and client
session metadata record the requested backend, actual engine, version and
fallback reason under `local_backend`. RDFLib remains the graph/term API for
parsing, authoring, exports and SHACL validation. No graph digest is computed.

In-memory graphs use bulk structural mining automatically. One SPARQL result
supplies uncovered edges, exact counts, distinct nodes and witness bindings.
Recount and witness queries remain available for independent verification.
A typed graph with no uncovered edges skips structural profile discovery.
This path retains uncovered bindings in memory; SPARQL endpoints keep
aggregate queries. Local reports separate `graph-census` (triple and type-presence
counts), `structural-discovery` (uncovered edges and property sets), and
`structural-patterns` (profile assembly). Query timings remain available under
`query_stats`.

Collection profiles are available as `schema.collections` and under
`document["schema"]["collections"]` in canonical JSON.

### Mine an endpoint

```python
from rdfsolve import SchemaMiner

# Query a SPARQL endpoint
miner = SchemaMiner(endpoint_url="https://sparql.example.org/sparql")
schema = miner.mine(dataset_name="example")

# Enrich with labels and definitions, and examples
schema.enrichment = miner.query_enrichment(schema)

# Export formats
schema.to_void_graph()  # VoID RDF graph
schema.to_linkml_yaml()  # LinkML schema YAML string
schema.to_shacl()  # Deactivated observed SHACL templates
schema.to_pydantic()  # Python source for dataset-specific Pydantic classes
schema.to_dict()  # Versioned canonical document (dict)

# Inspect
schema.get_classes()  # List with classes
schema.get_properties()  # List with properties

print(schema.get_metadata())  # Print a summary view of found metadata across graphs
print(schema.get_metadata().to_trig())  # Print metadata as trig
print(schema.get_metadata().to_turtle())  # Print metadata as turtle

import json

with open("example_schema.json", "w", encoding="utf-8") as f:
    json.dump(schema.to_dict(), f, indent=2)
```

`schema.model_dump_json()` returns the `MinedSchema` data as a JSON string.

Use `schema.to_dict()` with `json.dump()` for saved rdfsolve files.
`MinedSchema.model_json_schema()` describes the internal model;
`schema.to_pydantic()` instead generates Python classes for the mined RDF types.

#### Probe for SHACL NodeShapes
To measure joined support for longer paths, use `navigation_hops`,
`navigation_limit` and `navigation_probes` in the mining API:

```python
from rdfsolve.api import mine_schema

schema = mine_schema(endpoint, navigation_hops=3, navigation_limit=50, navigation_probes=20)
```

The `MinedSchema` JSON retains the queries and observations and its SHACL
exports include deactivated, nested query profiles for observed routes. See
[the mining API](docs/source/miner.rst).

### Discover existing VoID descriptions

VoID is a published dataset description. It may contain metadata without any
relationship patterns:

```python
from rdfsolve.api import discover_void_source

void = discover_void_source(
    endpoint="https://aopwiki.rdf.bigcat-bioinformatics.org/sparql",
    name="aopwikirdf",
)
print(void.has_void, void.has_partitions, void.has_patterns)
print(void.get_metadata())  # Show retrieved metadata; no new request.

schema = void.to_mined_schema()
print(schema.get_metadata())  # Show metadata from the stored schema.
```

`has_void` means any VoID description was found; `has_patterns` means it
provides relationship patterns.

Discovery checks all graphs, one request at a time. Use `void.graph_uris` to
list matches and `void.for_graph(uri).get_metadata()` to inspect one. Set
`output_dir` only when you want files.

### Query metadata without mining

Read the available metadata (queried with the package):

```python
from rdfsolve.api import query_metadata

metadata = query_metadata(
    "https://aopwiki.rdf.bigcat-bioinformatics.org/sparql",
    graph_uris=["http://aopwiki.org/"],
)
print(metadata)  # Readable view; notebooks also display it automatically.
```

The view shows what was retrieved, not everything the endpoint may contain. Use
`metadata.to_turtle()`/`metadata.to_trig()` for all retained details, or
`metadata.to_markdown()` to save the readable view.

### Read schema files

```python
from pathlib import Path
from rdfsolve import MinedSchema

schema = MinedSchema.from_json("example_schema.json")
schema = MinedSchema.from_void(Path("example_void.ttl").read_text(encoding="utf-8"))
schema = MinedSchema.from_shacl(Path("example_shacl.ttl").read_text(encoding="utf-8"))
```

Allows (partial) interconversion through `MinedSchema`.

### Load and convert existing schemas

```python
from rdfsolve import VoidParser

# Load VoID Turtle or JSON-LD
parser = VoidParser(void_source="schema.ttl")
schema = parser.to_mined_schema()

# Convert between formats
schema.to_jsonld()  # To JSON-LD
schema.to_linkml_yaml()  # To LinkML
schema.to_shacl()  # To SHACL
```

### Typed client generation

Use `Client` to find names or identifiers, then follow the selected records.
Start from a saved schema. Opening it makes no endpoint requests and does not
mine the data:

```python
from rdfsolve.api import Client

data = Client.open("aopwikirdf.schema.json")
matches = data.find("thyroid")
matches.types()
```

The schema supplies the types, fields and endpoint. Searches read the data when
you ask for it. Use `Client.open(schema)` or `schema.client()` if you already
have a `MinedSchema` in Python.

Choose another endpoint, local data, or a supported RDF schema export:

```python
data = Client.open("schema.json", source="https://example.org/sparql")
data = Client.open("schema.json", data_file="subset.ttl")
data = Client.open("shapes.ttl", format="shacl", source=endpoint)
data = Client.open("void.ttl", format="void", source=endpoint)
```

Prefer canonical schema JSON: it keeps all stored fields. SHACL and VoID use
their supported import fields. Metadata-only VoID cannot supply a typed client.
No missing fields are filled by automatic mining.

Pick the records you want and see where they lead:

```python
pathways = matches.of_type("Adverse outcome pathway")
pathways.paths()

stressors = pathways.related("Stressor")
chemicals = stressors.related("Chemical entity")
chemicals.show("identifier")
```

Use `values("title")` to list values across your selected records. Use
`target_value` when you know a name but not its class:

```python
from IPython.display import Markdown, display

paths = data.paths_between("Adverse outcome pathway", target_value="Phenobarbital", max_hops=3)
display(paths)
display(Markdown(data.diagram(paths=paths)))
```

This verifies mined class routes against matching names or identifiers, ignoring
case. It does not link records just because they share a type. Passing a second
class instead lists possible class routes without querying the data. Use
`diagram(paths=paths, path=1)` for the first complete path. Add
`instances=False` to show its classes instead of its records.

Pass the whole selection to evaluate connections for every matching resource:

```python
pathways = data.find("lung", kind="Adverse Outcome Pathway")
chemical_paths = data.paths_between(pathways, "Chemical entity", max_hops=2)
pathways.summary()
data.trace()
```

Record and `Results` inputs retain their exact identities in the path queries.
The path table contains observed bindings, source query IDs and per-route
outcomes in its `attrs`. Its coverage distinguishes partial search from no
match. A single record selects that record's connections. Two class names
describe schema routes.

Retrieve a discovered route and its available fields:

```python
route = chemical_paths.iloc[0]["Reference"]
result = data.retrieve(route, source="pathway", target="chemical", fields={"chemical": ["title"]})
result.table()
data.query_log()
```

Missing optional fields preserve the linked records. Exact RDF terms remain in
`result.rows`; `data.query_log()` shows the queries and their results.

Read the generated types and field paths without endpoint requests:

```python
data.describe(owners=["Key Events"], targets=["cellular organisms"])
data.describe("measurement", owners=["Key Events"], source=False)
```

Start with a label or literal when the relevant class is unknown:

```python
matches = data.describe("donepezil")
matches[["Kind", "Resource", "Types", "Predicate", "Literal", "Graph"]]
matches.attrs["coverage"]
```

The source lookup tries the supplied spelling, lowercase, uppercase and title
case as exact plain RDF literals. It uses direct object lookups, with no automatic
substring scan. Pass an RDFLib `Literal` to match a specific language or datatype,
for example `data.describe(Literal("Donepezil", lang="en"))`. These finite variants
do not cover every mixed-case spelling or language. Coverage records the literals
searched and marks budget-limited results as partial.

Disambiguate a name with an identifier supplied by the user:

```python
chemical = data.describe("donepezil", identifier="CHEBI:53289")
chemical.attrs["resolution"]
```

The identifier restricts the lookup before its row limit. Registered CURIEs
use Bioregistry namespace candidates; full IRIs use exact identity. A candidate
must also have matching literal evidence in the selected source. Multiple
observed namespace forms remain ambiguous. Registry version and candidates
are retained in the resolution metadata.

`connections` accepts description tables directly when they contain one
verified source identity. It rejects ambiguous, partial and external-only
descriptions before querying; it never selects the first row.

Resource matches retain reusable references for `prepare_network(values=...)`,
including resources whose types have no generated model. Schema matches and
external ontology candidates appear separately in `Kind`; ontology candidates
require `ontology_grounding` and do not establish presence in the data source.
Use `source=False` for schema-only inspection. A `targets` filter selects schema
fields only. For broader text searches, use `search` with a selected class and
fields.

Names resolve within their class. Exact generated names remain usable when human
labels collide. Missing metadata produces an empty description search; the
client does not invent vocabulary for an unexplained field.

Press Tab after `pathways.fields.` to discover fields while typing. `show()`
retrieves only the fields you ask for; displaying results does not send
requests.

To create a new schema, use `SchemaMiner` separately and save its output.
`explore(endpoint, graph=graph_iri)` is a quick mining shortcut, not required to
open a client. Show each query and its returned data with `data.query_log()`.
Save the queries, results, and steps with `data.save_session("session.json")`,
then close the connection with `data.close()`.

[Example](notebooks/pydantic_clients/01_mine_explore.ipynb).

Client operations are in `rdfsolve.client`; the model tools are in
`rdfsolve.mcp`. `rdfsolve.api` gives both for notebook use.

### SparqlHelper

Large SPARQL queries can time out, and endpoints can fail intermittently.
`SparqlHelper` retries temporary failures and fetches large results in smaller
batches. It reduces page sizes after timeouts and spaces requests to ease the
load on endpoints.

```python
from rdfsolve.sparql_helper import SparqlHelper

endpoint = "https://aopwiki.rdf.bigcat-bioinformatics.org/sparql"
query = "SELECT DISTINCT ?class WHERE { ?s a ?class }"

with SparqlHelper(endpoint, timeout=30) as helper:
    pages = helper.prepare_paginated_query(query)
    for rows in helper.select_chunked(pages, chunk_size=100, pagination="cursor"):
        print(rows)
```

For a single request, use `helper.select(query)`. Use `helper.ask(query)` for
yes/no questions or `helper.construct_graph(query)` to retrieve RDF. No mining
pipeline or registry is required.

Cursor paging continues after the last returned value, avoiding server limits on
large offsets. For `SELECT DISTINCT`, it uses the returned columns as keys;
`cursor_keys=["class"]` selects keys explicitly. Keys must identify each row.
Use `pagination="offset"` for offset paging. Paging cannot make every costly
query finish.

Mining accepts the same option:
`SchemaMiner(endpoint_url=endpoint, pagination="cursor")` or
`mine_schema(endpoint, pagination="cursor")`. It applies to paginated phases and
their fallbacks; the mining report records the choice.

### Keep and share useful queries

Give queries names, run them again, and share them as Turtle:

```python
with SparqlHelper(endpoint, timeout=30) as helper:
    helper.add_query("classes", query)
    results = helper.run_query("classes")
    helper.export_queries_as_ttl("queries.ttl")
    print(helper.history)  # Named runs: time, duration, and success or error
```

Loading a SHACL with Sparql Examples:

```python
with SparqlHelper(endpoint, timeout=30) as helper:
    names = helper.load_shacl("queries.ttl")
    print(names)
    helper.queries.rename(names[0], "my query")
    results = helper.run_query("my query")
```

SHACL paths can also become queries for a particular entity:

```python
helper.load_shacl("shapes.ttl")
print(helper.queries.paths)  # Choose a property shape
query = helper.queries.path_query(property_shape_id, entity_iri, limit=20)
```

### Retrieve source metadata

Start with a name and endpoint. Add `sources_file` to read existing settings.
The function returns observations and leaves the registry unchanged:

```python
from rdfsolve import enrich_source

source = enrich_source(
    "aopwikirdf",
    "https://aopwiki.rdf.bigcat-bioinformatics.org/sparql",
    sources_file="sources.yaml",
)
print(source["dataset_metadata"])
print(source["enrichment"])  # Completed, failed, or skipped retrieval steps
```

This tests availability and retrieves dataset descriptions, not instance
patterns. It fills supported metadata such as the description, license, and
version when one dataset can be identified. Missing or ambiguous metadata stays
blank; failed retrieval leaves previous values intact. Existing query settings
and unrelated fields are kept. Save the returned observations in a separate
report. Edit the human-curated source specification directly.

Metadata comes from the default graph unless `metadata_graph_uris=[...]` is
supplied or stored. Add `discover_void=True` to find published VoID descriptions
and record their graph locations and pattern availability. That scan can take
many requests; it does not run by default. Metadata locations do not become
instance-query scopes.

Reuse the registry for your own queries:

```python
from rdfsolve import load_sources

source = next(s for s in load_sources("sources.yaml") if s["name"] == "aopwikirdf")
with SparqlHelper.from_source_entry(source) as helper:
    results = helper.select(query + " LIMIT 20")
```

### Batch mining

Mine multiple endpoints from a YAML file:

**Create `sources.yaml`:**

```yaml
- name: uniprot
  endpoint: https://sparql.uniprot.org/sparql

- name: rhea
  endpoint: https://sparql.rhea-db.org/sparql
```

To retrieve metadata for an entry, use
`enrich_source(name, endpoint, sources_file="sources.yaml")` as above. The registry
is read-only. Keep retrieved observations in separate reports; choose settings
such as `chunk_size`, `class_batch_size`, and `timeout` for the workload.

**Run batch mining:**

```bash
python scripts/pipeline.py --sources-file sources.yaml --remote-only
```

**Output:**

```text
output/
├── uniprot/
│   ├── uniprot_schema.jsonld
│   ├── uniprot_void.ttl
│   └── uniprot_report.json
└── rhea/
    ├── rhea_schema.jsonld
    ├── rhea_void.ttl
    └── rhea_report.json
```

### Local RDF Files (with QLever)

Mine local RDF dumps using QLever:

```yaml
- name: drugbank
  download_nt:
    - https://example.org/drugbank.nt.gz
```

```bash
# Download, index, and mine
python scripts/pipeline.py --sources-file sources.yaml --local-only
```

### Check endpoint health

Test endpoint availability and response times:

```python
from rdfsolve.endpoint_health import check_endpoint_health

check_endpoint_health("https://aopwiki.rdf.bigcat-bioinformatics.org/sparql")
# EndpointHealthCheck(
#     endpoint_url='https://aopwiki.rdf.bigcat-bioinformatics.org/sparql',
#     status='up', response_time=0.1596362590789795, error_message='',
#     timestamp='2026-09-08T08:16:36.126659+00:00'
# )
```

### Build connectivity graphs

Create graphs showing dataset relationships via shared classes and mappings:

```bash
python scripts/build_graphs.py output/schemas/ --mappings output/mappings/
```

### Questions through MCP

Install `rdfsolve[agents,mcp]`. A saved schema and one source (an endpoint or a
local RDF file) are sufficient:

```python
from rdfsolve.api import ask_rdf

answer = await ask_rdf(
    "Which Key Events of AOPs have NCBI gene identifiers?",
    schema="notebooks/mcp/schemas/aopwikirdf.schema.json",
    base_url="http://127.0.0.1:8080/v1",
    model_name="qwen36-35b-a3b",
)
print(answer.text)
if answer.state == "complete":
    display(answer.table())
```

The model writes SPARQL SELECT queries. Five tools help it:

| Tool     | Result                                                                                                                                                                    |
| -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `schema` | The properties of classes, with value types, counts, example values and the links to each class. Search words find classes and properties.                                |
| `find`   | Resources by name, words in their text, IRI or registered identifier.                                                                                                     |
| `paths`  | The shortest chains of properties between two classes, as triple patterns.                                                                                                |
| `run`    | The first rows of a query. Known prefixes are added. Notes tell about terms that are not in the schema, columns with no value, and the triple patterns that give no rows. |
| `answer` | The final query is run on all data (in pages on an endpoint), and the task ends. The query must select the `output_variables`.                                            |

All rows of the answer go to the caller (`answer.bindings`); the model sees
only some rows. `answer.files` names the files of the tool calls, the answer
and the source queries. When one tool call gives the same result three times,
the task stops.

`model=` accepts another PydanticAI model. `endpoint=`, `data_file=`,
`graph_uris=` and `output_dir=` set the source and the files.
`ontology_grounding=True` lets `find` use ontology names (OLS, or Ontobee with
`ontology_provider="ontobee"`) when no name in the source matches.
`ontology_cache=` and `ontology_offline=True` give reproducible runs.

The client can use the same ontology lookup directly:

```python
from rdfsolve.api import Client, Ontologies

lookup = Ontologies("ols", cache="ontology-cache.json")
with Client.open("schema.json", ontology_grounding=lookup) as data:
    display(data.describe("measurement method", owners=["Key Event"]))
    print(data.trace()["ontology"])
```

External definitions, synonyms and direct parents are kept as separate
evidence, with the provider, the basis of the IRI match and the fetch time. The
mined schema does not change. An ontology alias is used only when a record of
the source has the same IRI or registered identifier.

For Claude or another MCP host, start a stdio server:

```bash
python -m rdfsolve.mcp --schema /absolute/path/schema.json --endpoint https://example.org/sparql
```

The resource `rdfsolve://overview` gives a summary of the source for the
instructions of the host. `rdfsolve://diagnostics` gives the source queries.

## Documentation

Full docs: [rdfsolve.readthedocs.io](https://rdfsolve.readthedocs.io)

## License

MIT — see [LICENSE](LICENSE).

### Query catalogues

The [WikiPathways catalogue](notebooks/sparql_helper/01_sparql_examples.ipynb)
and [AOPWiki queries to SHACL](notebooks/sparql_helper/02_aopwiki_catalog.ipynb)
examples to load, name and save queries through `QueryCollection`.

### Retrieve a connected network

Use references returned by `describe()` and `paths_between()` with named roles.
Reusing a role joins the same resource. Available child fields remain optional.

```python
from rdfsolve.api import QueryPattern

query = data.prepare_network(
    [
        QueryPattern(reference=route, bindings=["pathway", "chemical"]),
        QueryPattern(reference=name_field, bindings=["chemical", "name"], optional=True),
    ],
    outputs=["pathway", "chemical", "name"],
)
result = data.select(query, exhaustive=True)
result.table()
```

`route` and `name_field` are selected discovery references. An exact restriction
uses `values={role: term_reference}`. MCP uses these client operations to
construct queries and returns summaries; custom SPARQL remains available through
Python.

Schema connectivity and mapping analysis: [workflow and evidence](docs/source/analysis.rst).


### External labels during source discovery

`client.describe("IC50", ontology_fallback=True)` checks an ontology service
when no matching source-used class is found. Labelled measurement instances
do not count as a class match. Each external candidate retains its provider,
a scoped literal check on the candidate IRI and a class-use witness.
Use `ontology_grounding=Ontologies(...)` to configure caching, offline use
and the request budget.

An external label does not become a source label. A negative literal check
applies only to the recorded literal forms and graph scope. Source errors
propagate; incomplete source results remain incomplete. Candidates remain
separate from source identities and require review before query composition.
`save_session()` records the lookup strategy, candidates and query evidence.
No schema patterns are added by this lookup.


### Resolve names when composing a query

`client.resolve(name, kind="class")` returns a `Resolution`: `status`
(`resolved`, `ambiguous` or `unresolved`), every `candidate` with its origin
(supplied IRI, registered identifier, schema label, source label or external
ontology), whether it is used in the selected graphs and one witness. Nothing is
chosen by precedence; two used candidates are ambiguous. `kind="class"` matches
nodes typed with the class (not members of its subclasses); `kind="resource"`
matches one exact RDF term, including a class IRI used as a value. CURIEs such as
`CHEBI:53289` resolve through registered namespaces; a supplied IRI needs a
statement in scope. `notes` state what the constraint means, `warnings` what the
evidence does not cover (truncated searches, related-synonym matches).
`external_names=True` adds external ontology class names to the candidates.

`client.prepare_network(patterns, outputs=..., resolve=True)` accepts class
names, IRIs or CURIEs and field names that match exactly (case, width and
punctuation aside). Unresolved or ambiguous classes raise `ResolutionError`
with the full resolution. A field declared for a class other than the bound one
is allowed but adds its owner type; the prepared query's `warnings` say so.
Shared bindings specify the joins. Resolutions are recorded in the prepared query
and saved session. Dataset-specific roles and graph topologies belong in the
caller's patterns.
