{% extends "base.html" %} {% block title %}Model comparison · Evalground{% endblock %} {% block head %}{% endblock %} {% block content %}
{% from '_ui.html' import page_header, evalground_tabs %} {% call page_header('New comparison', 'Evaluate quality, latency and cost on a shared set of questions.', 'Evaluate', nav=evalground_tabs('compare'), title_id='comparison-page-title', description_id='comparison-page-description') %}{% endcall %} {#- Live summary: the planned workload while setting up, the saved run's figures on a result. Filled by the script below. -#}
{% for label in ['Dataset', 'Questions', 'Answer models', 'Rerankers', 'Planned answers', 'Spend threshold'] %}
{{ label }}—
{% endfor %}
01

Define the comparison

Choose what changes and the questions you will use to measure it.

Dataset format & example

Provide expected answers and relevant memory IDs for each question. Use conversations that were not used to tune your prompts. Paths refer to files accessible by the Evalground server.

[{"case_id":"q1","corpus_id":"customer-1","category":"preference",
  "question":"Which drink does Asha prefer?","answers":["tea"],
  "scorer":"exact_match","relevant_source_ids":["m1"],
  "documents":[{"source_id":"m1","content":"Asha prefers tea."}]}]
Use a saved candidate pool

Optional: reuse previously retrieved evidence to isolate model and reranker performance. Retrieval and embedding costs are excluded from this replay.

Start from a sample configuration

These examples use six synthetic questions and saved Oracle/Voyage candidates. New runs make live provider calls.

Examples & developer guides

Run sample comparisons here or explore the educational appbook. The appbook provides standalone Python lessons and Oracle-backed decision experiments.

Other comparisons
Loading…
View all evaluations →
{% endblock %} {% block scripts %}{% endblock %}