{% extends "base.html" %} {% block title %}Judges — SimpleAudit{% endblock %} {% block content %}

Judges

A judge is how conversations are graded: the criteria (what to evaluate), the output format (severity, score, yes/no or checklist) and the probe prompt the auditor follows. Start from one of SimpleAudit's judges, or write your own criteria. You pick the model that grades on New Experiment. Editing a judge saves a new version; runs keep the version they used.

{% if judges %} {% endif %} + New judge
{% if judges and missing_bases %}
{% csrf_token %} SimpleAudit has {{ missing_bases|length }} judge{{ missing_bases|pluralize }} this workspace doesn't have yet: {% for r in missing_bases %}{{ r.name }}{% if not forloop.last %}, {% endif %}{% endfor %}
{% endif %} {% if judges %}
{% for j in judges %}
{{ j.name }} {% if j.current %}{% include "partials/version_badge.html" with number=j.current.version count=j.version_count %}{% endif %}
{% if j.description %}

{{ j.description }}

{% endif %}
{% if j.current %}
Grades as
{{ j.current.output_label }}{% if j.current.output == "score" and j.current.options.dimensions %} · {{ j.current.options.dimensions|join:", " }}{% endif %}
Criteria
{% if not j.current.base %}own{% else %}{% if j.current.resolved.custom_criteria %}edited{% else %}SimpleAudit's{% endif %}{% if j.current.base_name != j.name %} · from {{ j.current.base_name }}{% endif %}{% endif %}{% if j.current.resolved.custom_probe_prompt %} · own probe{% endif %}
Used by
{% if j.usage %}{{ j.usage }} run{{ j.usage|pluralize }}{% else %}no runs{% endif %}{% if j.monitor_count %} · {{ j.monitor_count }} monitor{{ j.monitor_count|pluralize }}{% endif %}
{% endif %}
{% endfor %}
{% else %}

Set up your judges

Start with SimpleAudit's judges: safety (its default), harm, helpfulness, factuality, abstention, checklist and more. You can edit their criteria or clone them later.

{% csrf_token %}

Or write your own judge.

{% endif %} {% endblock %}