Dataset report
{{ report.source.filename }}
Analysis complete in {{ report.processing_seconds | seconds }} s
- Rows
- {{ report.summary.row_count | count }} {{ 'No empty rows' if report.summary.empty_row_count == 0 else (report.summary.empty_row_count | count) ~ (' empty row' if report.summary.empty_row_count == 1 else ' empty rows') }}
- Columns
- {{ report.summary.column_count | count }} {{ report.summary.cell_count | count }} cells analyzed
- Missing cells
- {{ report.summary.missing_percent }}% {{ report.summary.missing_count | count }} cells
- Duplicate rows
- {{ report.summary.duplicate_row_count | count }} Beyond the first occurrence
Quality information ({{ report.issues | length }})
{% if report.issues %}
-
{% for issue in report.issues %}
- {{ issue.severity }}: {{ issue.message }} {{ issue.count | count }} {% endfor %}
No issues detected by the current checks.
{% endif %}Column types
-
{% for kind, count in types.items() %}
- {{ kind | capitalize }}{{ count }} {% endfor %}
Columns ({{ report.summary.column_count }})
| Column | Missing (%) | Distinct | Inferred type | Semantic type | Error (%) | Examples |
|---|---|---|---|---|---|---|
| {{ column.position }}{{ column.name or '(unnamed)' }} | {% if column.missing_percent != 0 %} {{ percent_label(column.missing_percent) }}{{ column.missing_count | count }} {% endif %} |
{{ column.distinct_count | count }} | {{ column.inferred_type }} | {% if column.semantic_type %}{{ column.semantic_type }}{% endif %} | {% if column.type_error_percent is not none and column.type_error_percent != 0 %} {{ percent_label(column.type_error_percent) }}{{ column.type_error_count | count }} {% endif %} |
{% set inline_count = report.config.value_examples.inline_display_size %}
{% for item in column.value_profile.values[:inline_count] %}{{ item.value }}{% else %}No present values{% endfor %}
{% set remaining = column.value_profile.values[inline_count:] %}
{% if remaining %}{% endif %}
{% if column.value_profile.values %}
Values {{ column.value_profile.values | length }}{% if column.value_profile.selection != 'complete' %} / {{ column.distinct_count | count }}{% endif %}
{% for item in column.value_profile.values %}
{{ item.value }}({{ item.count | count }}) {% endfor %}
|
Transformations ({{ report.columns | length }})
| Column | Missing (%) | Trimmed (%) | Whitespace collapsed (%) | Distinct |
|---|---|---|---|---|
| {{ column.position }}{{ column.name or '(unnamed)' }} | {% for percent, count, track in [(column.missing_percent, column.missing_count, 'danger'), (column.normalization.trim_percent, column.normalization.trim_count, 'concept'), (column.normalization.collapse_internal_whitespace_percent, column.normalization.collapse_internal_whitespace_count, 'concept')] %}{% if percent != 0 %} {{ percent_label(percent) }}{{ count | count }} {% endif %} |
{% endfor %}
{{ column.distinct_count | count }} |
Numeric analysis ({{ numeric_columns | length }})
| Column | Minimum | Maximum | Mean | Median |
|---|---|---|---|---|
| {{ column.position }}{{ column.name or '(unnamed)' }} | {% for value in [column.numeric.minimum, column.numeric.maximum, column.numeric.mean, column.numeric.median] %}{{ value | number }} | {% endfor %}
Date analysis ({{ date_columns | length }})
| Column | Status | Valid (%) | Ambiguous (%) | Invalid (%) | Other (%) | Formats |
|---|---|---|---|---|---|---|
| {{ column.position }}{{ column.name or '(unnamed)' }} | {{ column.date_profile.status | replace('_', ' ') }} | {% set present_count = column.date_profile.valid_count + column.date_profile.ambiguous_count + column.date_profile.invalid_date_count + column.date_profile.not_date_count %} {% for value, track in [(column.date_profile.valid_count, 'success'), (column.date_profile.ambiguous_count, 'accent'), (column.date_profile.invalid_date_count, 'danger'), (column.date_profile.not_date_count, 'muted')] %}{% set percentage = ((100 * value / present_count) | round(2)) if present_count else 0 %}{% if percentage != 0 %} {{ percent_label(percentage) }}{{ value | count }} {% endif %} | {% endfor %}
{% set variant_label = column.date_profile.format_count ~ (' variant' if column.date_profile.format_count == 1 else ' variants') %}
{{ variant_label }}
{{ variant_label }}
{% for item in column.date_profile.breakdown %}
{{ item.label }}{{ item.count | count }} ({{ item.percent }}%) {% endfor %}
|
String analysis ({{ string_columns | length }})
| Column | Class | Fixed | Min | Max | Mean | Median | Distinct | Examples |
|---|---|---|---|---|---|---|---|---|
| {{ column.position }}{{ column.name or '(unnamed)' }} | {{ column.string_profile.status | replace('_', ' ') }} | {% if fixed %}{{ column.string_profile.fixed_length | count }}{% endif %} | {% if not fixed %}{{ column.string_profile.minimum_length | count }}{% endif %} | {% if not fixed %}{{ column.string_profile.maximum_length | count }}{% endif %} | {% if not fixed %}{{ column.string_profile.mean_length | number }}{% endif %} | {% if not fixed %}{{ column.string_profile.median_length | number }}{% endif %} | {{ column.distinct_count | count }} |
{% set representative_lengths = column.string_profile.length_distribution | sort(attribute='count', reverse=true) %}
{% for item in representative_lengths[:3] %}{% if item.examples %}{{ item.examples[0].value }}{% endif %}{% endfor %}
{% if representative_lengths | length > 3 %}{% endif %}
{% if column.string_profile.length_distribution %}
Lengths {{ column.string_profile.distinct_length_count }}
{% for item in representative_lengths %}
{{ item.length }} character{{ '' if item.length == 1 else 's' }}{% if item.examples %}{% for example in item.examples[:3] %}{{ example.value }}{% if not loop.last %} / {% endif %}{% endfor %}{% endif %} {{ item.percent }}%{{ item.count | count }} |
Data sample
{% if report.preview %}{% endif %}
{% if report.preview %}
| Row | {% for column in report.columns %}{{ column.name or '(unnamed)' }} | {% endfor %}
|---|---|
| {{ row.row_number }} | {% for value in row.values %}{{ value }} | {% endfor %}
{{ 'No data records.' if report.summary.row_count == 0 else 'Data preview disabled.' }}
{% endif %}Analysis settings
- Scope
- All {{ report.summary.row_count | count }} records
- Missing-value markers
- {% for marker in report.config.missing_values %}
{{ marker | tojson }}{% else %}None{% endfor %} - Missing-value comparison
- Whitespace trimmed; case-sensitive
- Value normalization
- Trim: {{ 'enabled' if report.config.normalization.trim else 'disabled' }}; collapse internal horizontal whitespace: {{ 'enabled' if report.config.normalization.collapse_internal_whitespace else 'disabled' }}. Raw preview values are preserved.
- Duplicate comparison
- Exact raw values across every column
- Type inference
- At least {{ (report.config.type_inference.minimum_confidence * 100) | round(1) }}% agreement across present values; configured strict date formats, dot-decimal numbers and true/false. Leading-zero identifiers remain text.
- Date detection
- Orders: {{ report.config.date_detection.orders | join(', ') }}; separators: {% for separator in report.config.date_detection.separators %}
{{ separator | tojson }}{% endfor %}; ambiguous order: {{ report.config.date_detection.ambiguous_order or 'inferred only from unambiguous column evidence' }} - Enum candidates
- Text columns with at most {{ report.config.enum_detection.maximum_distinct_values }} values when the dataset has at least {{ report.config.enum_detection.minimum_row_count | count }} rows
- String lengths
- Very short through {{ report.config.string_analysis.very_short_max_length }}; short through {{ report.config.string_analysis.short_max_length }}; medium through {{ report.config.string_analysis.medium_max_length }}; long through {{ report.config.string_analysis.long_max_length }}; otherwise very long. Up to {{ report.config.string_analysis.examples_per_length }} examples per length are retained when the maximum is {{ report.config.string_analysis.length_distribution_max_length }} or less.
- Value representation
- Complete through {{ report.config.value_examples.full_distribution_max_distinct }} distinct values; otherwise a reproducible sample of {{ report.config.value_examples.candidate_sample_size }}
- Row numbering
- Data records start at 1, excluding the header. Quoted multiline values count as one record.
- Preview selection
- First {{ report.config.preview_rows }} records; raw values preserved
- JSON format
- {{ report.format_version }} - revision {{ report.format_revision }} (experimental)