Search architecture — every retrieval path
The three retrieval stacks — the ARIEL Postgres stack with its enhancement pipeline and query modes, the qmd sidecar sharing one hybrid index between ARIEL and facility knowledge, and the embedding-free channel finder — each feeding MCP tools the OSPREY agent calls.
ARIEL · LOGBOOK · POSTGRES STACK
QMD SIDECAR · SHARED HYBRID INDEX
CHANNEL FINDER · NO EMBEDDINGS ANYWHERE
INGEST
EMBED
ROWS
VECTORS
WRITES .MD
provider reachable at query time
collection "ariel"
read-only mount · collection "okf"
RANKED HITS
sidecar down → substring scan
facility logbook
ALS · JLab · ORNL
embedding provider
Ollama · nomic-embed
or OpenAI
enhancement pipeline
runs at ingest
semantic_processor
LLM keywords + summary
text_embedding
vectors via provider
qmd_export
writes markdown mirror
Postgres
logbook database
tsvector index
full-text · for keyword
pgvector tables
one table per model
enhanced_entries
canonical rows
keyword
tsquery · ts_rank · trigram
semantic
pgvector cosine
sql_query
read-only SQL · exact
hybrid
answered by qmd sidecar
markdown mirror
one .md per entry
OKF bundle
facility knowledge .md
qmd sidecar
self-contained container
BM25 keyword index
own embedder · gemma-300M
LLM reranker · Qwen3-0.6B
OKF search
MCP tool + web panel
channel database
names + metadata
DuckDB FTS
BM25 · middle layer
hierarchical tree
agent navigates · no index
CF MCP tools
run_sql · list_* · navigate
OSPREY agent
via MCP tools
LEGEND
embedding vectors
model inference in the query path
fallback
Figure 1 — every retrieval path, left to right: sources → ingest-time
processing → indexes → query paths → the agent. The hybrid module and OKF
search share the one qmd sidecar — the only cross-stack connection.
query text
+ optional intent
lex
BM25 keyword match
vec
gemma-300M embed · cosine
hyde
1.7B writes a fake answer,
embeds it — no OSPREY caller
merge
first type ×2 weight
≤ candidateLimit (40)
LLM rerank
Qwen3-0.6B · ~4× latency
top hits
scores + line snippets
rerank: true
rerank: false
re-scored order
Figure 2 — one qmd query: two sub-queries fan out, merge into a bounded
candidate pool, and the LLM reranker optionally re-scores the pool. The
greyed hyde branch ships in the image but has no caller.