{% extends "base.html" %} {% block title %}Evalground - Memorizz{% endblock %} {% block head %}{% endblock %} {% block content %} {% from '_ui.html' import page_header, evalground_tabs %} {% call page_header('Evalground', 'Choose models and retrieval strategies with evidence from your workloads.', 'Evaluate', nav=evalground_tabs('evaluations'), actions_class='evaluation-actions') %}New comparison Evaluate an agent Test a harness{% endcall %} {#- Filled from the run library by evalground-runs.js; placeholders keep the layout steady while it loads. -#}
Same evidence. Compare answer quality, latency and API cost.
Compare rerankersSame candidates. Measure which method finds the right evidence.
Evaluate an agentRun a memory benchmark against an agent or retrieval pipeline.
Test a harnessRun Codex, Claude Code or a MemAgent on Terminal-Bench 4.0 tasks in Docker.
Search, compare and revisit your evaluation results.
| Models | Result |
|---|
Loading saved runs…
Accepted Trace Insights become versioned drafts here. A draft is evidence for an evaluation plan; it does not mutate the production agent.
{% if trace_experiments %}| Experiment | Version | Status | Agent | Component | Target | Created |
|---|---|---|---|---|---|---|
{{ experiment.experiment_id|truncate(24) }} |
v{{ experiment.experiment_version }} | {{ experiment.status|default('draft')|title }} | {{ experiment.agent_id|default('—')|truncate(16) }} |
{{ experiment.component or '—' }} | {{ experiment.target or '—' }} | {{ experiment.created_at or '—' }} |
No accepted trace recommendations yet.
{% endif %}Official datasets remain external. Configure each environment variable or enter a path for the run.
{{ item.dataset_env }}memorizz eval dataset sync {{ item.benchmark_id }}Harbor runs each task in its own Docker container and grades it with the task’s own tests. Codex and Claude Code are installed inside the container; a MemAgent works it from here. With MemoRizz memory on, a harness starts with what earlier runs learned, but never with lessons from the task it is solving.
Charges are benchmark estimates, not provider invoices. Runs may differ in samples, model and configuration; compare like for like. Local hardware cost is not measured.
| Run ID | Status | Benchmark | Mode | Agent | Dataset | Samples | Accuracy | API cost estimate / time | Created | Finished | Actions |
|---|---|---|---|---|---|---|---|---|---|---|---|
{{ run.run_id[:8] }}... |
{{ run_status }} | {{ run.benchmark or 'longmemeval' }} | {{ 'Full MemAgent' if run.evaluation_mode == 'memagent' else 'Retrieval' }} | {{ run.agent_name }} | {{ run.dataset_variant or 'oracle' }} | {{ run.num_samples }} | {% if run.overall_accuracy is not none %} {{ run.overall_accuracy }}% {% else %} - {% endif %} | {{ '$%.6f'|format(run.cost_usd) if run.cost_usd is not none else 'Unknown' }} / {{ '%.2fs'|format(run.processing_seconds) if run.processing_seconds is not none else 'Unknown' }} | {{ run.created_at or '-' }} | {{ run.finished_at or '-' }} | View {% if run_status in ['queued', 'running', 'canceling'] %} {% endif %} |
No agent evaluations in this server session. Saved model comparisons are listed in Recent evaluations above.
{% endif %}Examples & developer guides · Explore sample datasets and runnable memory experiments.
{% if selected_agent %}{{ selected_agent.agent_id }}
| Task | Result | Time | Tokens in / out | Cost | Memory |
|---|---|---|---|---|---|
| {{ case.task }}{{ case.category }}{% if case.subcategory %} · {{ case.subcategory }}{% endif %} | {% if case.passed %}Passed{% elif case.reward is not none %}Failed{% else %}Error{% endif %}{% if case.error %}{{ case.error|truncate(90) }}{% endif %} | {% if case.duration_s is not none %}{{ (case.duration_s / 60)|round(1) }} min{% else %}—{% endif %} | {% if case.input_tokens is not none %}{{ '{:,}'.format(case.input_tokens) }} / {{ '{:,}'.format(case.output_tokens or 0) }}{% else %}—{% endif %} | {% if case.cost_usd is not none %}${{ '%.2f'|format(case.cost_usd) }}{% else %}—{% endif %} | {% if tb.memory_id %}{{ case.memory_sources }} record{{ '' if case.memory_sources == 1 else 's' }}{% else %}—{% endif %} |
Dataset {{ tb.dataset }} · Harbor job folder {{ tb.job_dir }}{% if tb.harbor_exit_code %} · Harbor exited with code {{ tb.harbor_exit_code }}{% endif %}. Runs with a shortened time limit or memory from earlier runs are not leaderboard-comparable.
| Benchmark | {{ eval_results.benchmark_name }} |
| Dataset | {{ eval_results.metadata.dataset_variant }} |
| Execution | {% if eval_results.metadata.evaluation_mode == 'memagent' %} Full MemAgent · {{ eval_results.metadata.application_mode|default('assistant') }} {% else %} Memory retrieval diagnostic {% endif %} |
| Timestamp | {{ eval_results.metadata.timestamp }} |
| Retrieval basis | {{ eval_results.retrieval.basis or 'diagnostic_fusion' }} · top {{ eval_results.retrieval.top_k }} of {{ eval_results.retrieval.candidate_pool_size }} |
| Models | {{ eval_results.metadata.model_provider }}:{{ eval_results.metadata.reader_model }} reader · {{ eval_results.metadata.judge_model }} judge · {{ eval_results.metadata.embedding_model }} embeddings |
| External API cost estimate | {% if eval_results.metadata.external_api_cost_usd is not none %}${{ eval_results.metadata.external_api_cost_usd }}{% else %}Unknown{% endif %} |
| Comparison status | {{ eval_results.comparison_label or 'Diagnostic' }} |
| Output | {{ eval_output_path }} |
Run a benchmark to see results.
{{ run_output }}