{% extends "base.html" %} {% block title %}Evalground - Memorizz{% endblock %} {% block head %}{% endblock %} {% block content %} {% from '_ui.html' import page_header, evalground_tabs %} {% call page_header('Evalground', 'Choose models and retrieval strategies with evidence from your workloads.', 'Evaluate', nav=evalground_tabs('evaluations'), actions_class='evaluation-actions') %}New comparison Evaluate an agent Test a harness{% endcall %} {#- Filled from the run library by evalground-runs.js; placeholders keep the layout steady while it loads. -#}
{% for label in ['Evaluations', 'In progress', 'Completed', 'Needs attention', 'Known API cost', 'Latest run'] %}
{{ label }}—
{% endfor %}
Compare answer models

Same evidence. Compare answer quality, latency and API cost.

Compare rerankers

Same candidates. Measure which method finds the right evidence.

Evaluate an agent

Run a memory benchmark against an agent or retrieval pipeline.

Test a harness

Run Codex, Claude Code or a MemAgent on Terminal-Bench 4.0 tasks in Docker.

Recent evaluations

Search, compare and revisit your evaluation results.

j k move↵ open result

Saved model comparisons and agent evaluations. Select a row and press Enter to open its result.
Models Result

Loading saved runs…

About these measurements

{% if warnings %}
Heads up
{% endif %} {% if error %}
! {{ error }}
{% endif %}
Suggested experiments

Accepted Trace Insights become versioned drafts here. A draft is evidence for an evaluation plan; it does not mutate the production agent.

{% if trace_experiments %}
{% for experiment in trace_experiments %} {% endfor %}
Experiment Version Status Agent Component Target Created
{{ experiment.experiment_id|truncate(24) }} v{{ experiment.experiment_version }} {{ experiment.status|default('draft')|title }} {{ experiment.agent_id|default('—')|truncate(16) }} {{ experiment.component or '—' }} {{ experiment.target or '—' }} {{ experiment.created_at or '—' }}
{% else %}

No accepted trace recommendations yet.

{% endif %}
Benchmark datasets

Official datasets remain external. Configure each environment variable or enter a path for the run.

{% for item in benchmark_catalog %}
{{ item.variants|join(', ') }} · protocol {{ item.protocol_version }} · {{ item.dataset_env }}
{% if item.source_sync_supported %}
memorizz eval dataset sync {{ item.benchmark_id }}
{% endif %}
{% if item.dataset_ready %} Ready {% else %} Path needed {% endif %} Diagnostic
{% endfor %}
{% if download_message %}
{{ download_message }}
{% endif %} {% if download_error %}
! {{ download_error }}
{% endif %}
Legacy LongMemEval dataset helper
{% for item in dataset_status %}
LongMemEval {{ item.variant|upper }}
{{ item.filename }} {% if item.exists %}Ready{% else %}Missing{% endif %}
{% endfor %}
{% if missing_variants %}
{% endif %}
Configure an agent evaluation
Diagnostic mode isolates retrieval and reader quality. Full mode runs the selected agent through MemAgent.run() on isolated benchmark memory.
Required for Full MemAgent execution; optional attribution in diagnostic mode.
Separates retrieval failures from answer-synthesis failures.

Every run is fail-closed as Diagnostic unless its versioned manifest proves the official runner, scorer, full split, source revision, models, prompts, and runtime metadata all match. Full MemAgent mode disables side-effect tools and uses isolated benchmark memory while preserving the selected agent's memory configuration. Ollama readers cost $0 externally; OpenAI usage is metered. Paper profiles ignore the sample field.

Results saved to {{ eval_results_dir }}.

Test a harness on Terminal-Bench 4.0

Harbor runs each task in its own Docker container and grades it with the task’s own tests. Codex and Claude Code are installed inside the container; a MemAgent works it from here. With MemoRizz memory on, a harness starts with what earlier runs learned, but never with lessons from the task it is solving.

Tasks
{% for group in terminal_bench.groups %}

{{ group.category }} {{ group.tasks|length }}

{% for task in group.tasks %} {% endfor %}
{% endfor %}

{{ terminal_bench.task_count }} of the dataset’s tasks run without a GPU; {{ terminal_bench.gpu_tasks|join(', ') }} are left out.

Agent evaluation history {% if eval_charts %}{% from '_analytics_charts.html' import bars %}
{{ bars(eval_charts.cost, 'Recent evaluation charges', 'USD · benchmark-reported estimates') }}

Charges are benchmark estimates, not provider invoices. Runs may differ in samples, model and configuration; compare like for like. Local hardware cost is not measured.

{% endif %} {% if runs_history %}
{% for run in runs_history %} {% set run_status = run.status or 'unknown' %} {% set badge_class = 'badge-warning' %} {% if run_status == 'completed' %} {% set badge_class = 'badge-success' %} {% elif run_status in ['failed', 'canceled'] %} {% set badge_class = 'badge-danger' %} {% endif %} {% endfor %}
Run ID Status Benchmark Mode Agent Dataset Samples Accuracy API cost estimate / time Created Finished Actions
{{ run.run_id[:8] }}... {{ run_status }} {{ run.benchmark or 'longmemeval' }} {{ 'Full MemAgent' if run.evaluation_mode == 'memagent' else 'Retrieval' }} {{ run.agent_name }} {{ run.dataset_variant or 'oracle' }} {{ run.num_samples }} {% if run.overall_accuracy is not none %} {{ run.overall_accuracy }}% {% else %} - {% endif %} {{ '$%.6f'|format(run.cost_usd) if run.cost_usd is not none else 'Unknown' }} / {{ '%.2fs'|format(run.processing_seconds) if run.processing_seconds is not none else 'Unknown' }} {{ run.created_at or '-' }} {{ run.finished_at or '-' }} View {% if run_status in ['queued', 'running', 'canceling'] %} {% endif %}
{% else %}

No agent evaluations in this server session. Saved model comparisons are listed in Recent evaluations above.

{% endif %}

Examples & developer guides · Explore sample datasets and runnable memory experiments.

{% if selected_agent %}
Agent {% if selected_agent.persona and selected_agent.persona.name %} {{ selected_agent.persona.name }} {% else %} Agent {% endif %}
Agent ID {{ selected_agent.agent_id }}
Mode {{ selected_agent.application_mode or 'assistant' }}
{% if selected_agent.memory_ids %}
Memory IDs {{ selected_agent.memory_ids|length }}
{% endif %}
{% endif %} {% if eval_results and eval_results.metadata and eval_results.metadata.benchmark == 'terminal-bench' %} {% set tb = eval_results.metadata %}

Terminal-Bench 4.0 · {{ tb.harness }}{% if tb.model %} · {{ tb.model }}{% endif %}

Tasks passed
{{ (eval_results.overall_accuracy * 100)|round(1) }}%
{{ tb.passed }} of {{ tb.num_trials }} trial{{ '' if tb.num_trials == 1 else 's' }}
API cost{% if eval_results.cost_is_estimate %} · estimate{% endif %}
{% if eval_results.external_api_cost_usd is not none %}${{ '%.2f'|format(eval_results.external_api_cost_usd) }}{% else %}Unknown{% endif %}
Memory
{% if tb.memory_id %}{{ tb.memory_id }}{% else %}Off{% endif %}
{% if tb.lessons_saved is defined %}{{ tb.lessons_saved }} lesson{{ '' if tb.lessons_saved == 1 else 's' }} saved{% endif %}
Run time
{% if tb.total_processing_time %}{{ (tb.total_processing_time / 60)|round(1) }} min{% else %}—{% endif %}
{% if eval_charts %}{% from '_analytics_charts.html' import bars %}
{{ bars(eval_charts.categories, 'Pass rate by category', 'share of trials whose tests passed') }}
{% endif %}
{% for case in eval_results.cases %} {% endfor %}
Each trial: its task, result, time, tokens and cost.
TaskResultTimeTokens in / outCostMemory
{{ case.task }}{{ case.category }}{% if case.subcategory %} · {{ case.subcategory }}{% endif %} {% if case.passed %}Passed{% elif case.reward is not none %}Failed{% else %}Error{% endif %}{% if case.error %}{{ case.error|truncate(90) }}{% endif %} {% if case.duration_s is not none %}{{ (case.duration_s / 60)|round(1) }} min{% else %}—{% endif %} {% if case.input_tokens is not none %}{{ '{:,}'.format(case.input_tokens) }} / {{ '{:,}'.format(case.output_tokens or 0) }}{% else %}—{% endif %} {% if case.cost_usd is not none %}${{ '%.2f'|format(case.cost_usd) }}{% else %}—{% endif %} {% if tb.memory_id %}{{ case.memory_sources }} record{{ '' if case.memory_sources == 1 else 's' }}{% else %}—{% endif %}

Dataset {{ tb.dataset }} · Harbor job folder {{ tb.job_dir }}{% if tb.harbor_exit_code %} · Harbor exited with code {{ tb.harbor_exit_code }}{% endif %}. Runs with a shortened time limit or memory from earlier runs are not leaderboard-comparable.

{% elif eval_results %}

Results

{% if eval_charts %}{% from '_analytics_charts.html' import bars %}
{{ bars(eval_charts.categories, 'Quality by category', 'reported accuracy') }}{{ bars(eval_charts.latency, 'Evaluation latency', 'mean seconds; stages may overlap') }}
{% endif %}
Overall accuracy
{{ (eval_results.overall_accuracy * 100)|round(2) }}%
Overall score
{{ eval_results.overall_score|round(3) }}
Samples
{{ eval_results.metadata.num_samples }}
Processing time
{{ eval_results.metadata.total_processing_time|round(2) }}s
{% if eval_results.retrieval %}
Retrieval recall@k
{% if eval_results.retrieval.recall_at_k is not none %}{{ eval_results.retrieval.recall_at_k|round(3) }}{% else %}—{% endif %}
Gold-evidence ceiling
{% if eval_results.answer_quality.gold_evidence_oracle_score is not none %}{{ eval_results.answer_quality.gold_evidence_oracle_score|round(3) }}{% else %}—{% endif %}
Grounded answers
{{ (eval_results.answer_quality.grounded_rate * 100)|round(1) }}%
Corpus cache
{{ eval_results.efficiency.embedding_cache_hits }} hit / {{ eval_results.efficiency.embedding_cache_misses }} miss
{% endif %}
{% if eval_results.benchmark_name %} {% endif %} {% if eval_results.schema_version %} {% endif %} {% if eval_output_path %} {% endif %}
Benchmark {{ eval_results.benchmark_name }}
Dataset {{ eval_results.metadata.dataset_variant }}
Execution {% if eval_results.metadata.evaluation_mode == 'memagent' %} Full MemAgent · {{ eval_results.metadata.application_mode|default('assistant') }} {% else %} Memory retrieval diagnostic {% endif %}
Timestamp {{ eval_results.metadata.timestamp }}
Retrieval basis {{ eval_results.retrieval.basis or 'diagnostic_fusion' }} · top {{ eval_results.retrieval.top_k }} of {{ eval_results.retrieval.candidate_pool_size }}
Models {{ eval_results.metadata.model_provider }}:{{ eval_results.metadata.reader_model }} reader · {{ eval_results.metadata.judge_model }} judge · {{ eval_results.metadata.embedding_model }} embeddings
External API cost estimate {% if eval_results.metadata.external_api_cost_usd is not none %}${{ eval_results.metadata.external_api_cost_usd }}{% else %}Unknown{% endif %}
Comparison status {{ eval_results.comparison_label or 'Diagnostic' }}
Output {{ eval_output_path }}
{% include "_eval_usage.html" %}

Ability breakdown

{% for cat_key, metrics in eval_results.category_results.items() %}
{{ cat_key|replace('-', ' ')|replace('_', ' ')|title }}
Accuracy {% if eval_results.schema_version %}{{ (metrics.accuracy * 100)|round(1) }}%{% else %}{{ metrics.accuracy|round(2) }}{% endif %}
Avg Score {{ metrics.average_score|round(3) }}
Samples {{ metrics.num_samples }}
{% if metrics.retrieval_recall_at_k is defined and metrics.retrieval_recall_at_k is not none %}
Recall@k {{ metrics.retrieval_recall_at_k|round(3) }}
{% endif %} {% if metrics.faithfulness is defined and metrics.faithfulness is not none %}
Faithfulness {{ metrics.faithfulness|round(3) }}
{% endif %} {% if metrics.oracle_reader_score is defined and metrics.oracle_reader_score is not none %}
Oracle reader {{ metrics.oracle_reader_score|round(3) }}
{% endif %} {% if metrics.grounded_rate is defined %}
Grounded {{ (metrics.grounded_rate * 100)|round(1) }}%
{% endif %}
{% endfor %}
{% if eval_results.schema_version %}
ⓘ
{{ eval_results.comparison_label or 'Diagnostic' }} — {% if eval_results.paper_comparable %} Every required paper-protocol field matched the versioned manifest. {% else %} This run is useful for engineering diagnosis, not as an official paper or leaderboard score. {% if eval_results.non_comparability_reasons %}
    {% for reason in eval_results.non_comparability_reasons %}
  • {{ reason }}
  • {% endfor %}
{% endif %} {% endif %}
{% endif %}
{% elif selected_run_id %}
EVAL

No results yet

Run a benchmark to see results.

{% endif %} {% if run_output %}

Run Output

{{ run_output }}
{% endif %} {% endblock %} {% block scripts %} {% endblock %}