Dataset report
{{ source_stem }}{{ source_suffix }}
- {{ report.source.format | upper }} {% if report.source.format == 'json' %}
- Dataset
{{ ds.id }} - {{ ds.summary.row_count | count }} records
- {{ ds.summary.column_count | count }} fields {% else %}
- {{ ds.summary.row_count | count }} rows
- {{ ds.summary.column_count | count }} columns {% endif %}
- {{ report.source.size_bytes | size }}
- {{ report.source.encoding }} {% if report.source.delimiter is not none %}
- Delimiter
{{ report.source.delimiter }}{% endif %}
01 / Overview
Dataset overview
Size, completeness and the main signal to check first.
- Rows
- {{ ds.summary.row_count | count }}{% if ds.summary.empty_row_count %}{{ ds.summary.empty_row_count | count }} empty rows{% else %}No empty rows{% endif %}
- Columns
- {{ ds.summary.column_count | count }}{{ ds.summary.cell_count | count }} cells analyzed
- Missing cells
- {{ ds.summary.missing_percent }}%{{ ds.summary.missing_count | count }} cells
- Duplicate rows
- {% if ds.summary.duplicate_row_status == 'disabled' %}–Not checked{% elif ds.summary.duplicate_row_status == 'limited' %}≥ {{ ds.summary.duplicate_row_count | count }}At least, beyond the first occurrence{% else %}{{ ds.summary.duplicate_row_count | count }}Beyond the first occurrence{% endif %}
ambiguous dates, kept unresolved.
Values that match more than one configured date order. That is {{ ds.date_summary.ambiguous_percent }}% of the {{ ds.date_summary.present_count | count }} present date values. Tabalyst does not guess the order.
{% if ds.date_summary.columns %}{% endif %}See date analysis {% else %}Date analysisdate columns detected.
No date-shaped values were identified with the configured rules.
{% endif %}Quality observations
- {% for issue in ds.issues %}
- {{ issue.severity }}: {% if issue.code in ['limited_measures','excluded_records'] and has_limits %}{{ issue.message }}{% elif issue.code in ['missing_values','empty_columns','constant_columns','ambiguous_headers','mixed_types','limited_measures'] %}{{ issue.message }}{% elif issue.code in ['trimmed_cells','collapsed_whitespace','variant_groups'] %}{{ issue.message }}{% elif issue.code == 'ambiguous_dates' %}{{ issue.message }}{% else %}{{ issue.message }}{% endif %}{{ issue.count | count }} {% endfor %}
Column types
- {% for name, count in ds.summary.inferred_type_counts.items() %}
- {{ name | replace('_',' ') | title }} {{ count }} {% endfor %}
Semantic types
- {% for name, count in ds.summary.semantic_type_counts.items() %}
- {{ 'No semantic type' if name == 'none' else name | replace('_',' ') | title }} {{ count }} {% endfor %}
Mixed: fewer than {{ (report.config.scan.types.minimum_confidence * 100) | round(1) }}% of present values agree on one type.
02 / Structure
Columns({{ ds.summary.column_count }})
Inferred and semantic types per column. With issues: missing values or mixed type.
| {{ column.inferred_type }} | {% if column.semantic_type %}{{ column.semantic_type | replace('_',' ') }}{% else %}–{% endif %} | {{ percent_cell(column.type_error_percent,column.type_error_count or 0,'var(--attn)',column.type_error_count != 0) }}{{ value_examples(column, px=px) }}
03 / Shape
JSON structure({{ st.fields | length }})
Every path of the records, containers included, with presence per parent and array lengths.
- Records
- {{ ds.summary.row_count | count }}{% for name, count in st.record_types.items() %}{{ count | count }} {{ name }}{% if not loop.last %}, {% endif %}{% else %}No records{% endfor %}
- Paths
- {% if st.path_count is none %}limited{% else %}{{ st.path_count | count }}{% endif %}{{ ds.summary.column_count | count }} with values
- Maximum depth
- {{ st.max_depth_seen }}Limit {{ report.config.scan.limits.max_depth }}
{{ f.path }} | {{ f.depth }} | {% for name, count in f.native_types.items() %}{{ name }}{% endfor %} | {% if f.present_percent is none %}items {{ f.present_count | count }} | {% else %}{{ f.present_percent }}%{{ f.present_count | count }} | {% endif %}{% if f.absent_count is none %}– | {% else %}{% if f.absent_count %}{{ f.absent_count | count }}{% else %}0{% endif %} | {% endif %}{% if f.arrays %}{% if f.arrays.minimum_length == f.arrays.maximum_length %}{{ f.arrays.minimum_length | count }}{% else %}{{ f.arrays.minimum_length | count }}–{{ f.arrays.maximum_length | count }}mean {{ f.arrays.mean_length | number }}{% endif %} | {{ f.arrays.empty_count | count }} | {% else %}– | – | {% endif %}{% set role = 'collection' if f.collection else ('column' if f.column else 'container') %}{% if f.collection %}collection {{ f.collection }}{% elif f.column %}column{% else %}container{% endif %} |
|---|
No paths: the dataset has no records.
{% endif %}04 / Cleaning
Transformations({{ ds.summary.column_count }})
Occurrences changed by each normalization stage, and the spellings it groups. Raw preview values remain unchanged.
Variant groups({{ variant_rows | length }})
Raw spellings that every normalization stage compares as equal: an analytical equivalence, not proof of identity. Whitespace is shown as written.
{{ group.key }} | {{ group.count | count }} | {{ group.distinct_count | count }} | {% for item in group.variants %}{{ item.value }}{{ item.count | count }}{% endfor %}{% if group.truncated %}+{{ group.distinct_count - group.variants | length }}{% endif %} |
Only the largest groups of each column are listed (scan.limits.max_variant_groups).
05 / Numeric
Numeric analysis({{ ds.summary.numeric_column_count }})
Range and distribution statistics for accepted numeric values.
| {{ column.numeric.range | number }} | {{ column.numeric.minimum | number }} | {{ column.numeric.maximum | number }} | {{ column.numeric.mean | number }} | {% if column.numeric.median is none %}limited | {% else %}{{ column.numeric.median | number }} | {% endif %}{{ distinct_cell(column) }}{{ value_examples(column,'num',px) }}
06 / Dates
Date analysis({{ ds.summary.date_column_count }})
Strict date parsing keeps ambiguous, invalid and non-date values separate. Ambiguous values are never resolved from the other values of the column.
| {{ column.date_profile.status | replace('_',' ') }} | {{ column.date_profile.format_count }} variant{{ '' if column.date_profile.format_count == 1 else 's' }} {{ column.date_profile.format_count }} variant{{ '' if column.date_profile.format_count == 1 else 's' }} {% set evidence = column.date_profile.ambiguity_evidence %}{% if column.date_profile.ambiguous_count and evidence %}Unambiguous values: {% for order, count in evidence.items() %}{{ order }} {{ count | count }}{% if not loop.last %} · {% endif %}{% endfor %}.{% set attested = evidence | dictsort | selectattr(1) | list %}{% if attested | length == 1 %} Only {{ attested[0][0] }} is attested: set {% for item in column.date_profile.breakdown %} {{ item.label }}{{ item.count | count }} ({{ item.percent }}%) {% endfor %} |
07 / Text
String analysis({{ ds.summary.string_column_count }})
Length classes, fixed widths and representative values.
| {{ column.string_profile.status | replace('_',' ') }} | {% if column.string_profile.fixed_length is not none %}{{ column.string_profile.fixed_length }}{% else %}–{% endif %} | {% if column.string_profile.fixed_length is not none %}– | – | – | – | {% else %}{{ column.string_profile.minimum_length }} | {{ column.string_profile.maximum_length }} | {{ column.string_profile.mean_length }} | {{ column.string_profile.median_length }} | {% endif %}{% for item in column.string_profile.length_distribution %}{% endfor %} | {{ distinct_cell(column) }}{% for value in column.string_profile.representative_examples %}{{ value }}{% endfor %} Lengths {{ column.string_profile.distinct_length_count }} {% for item in column.string_profile.length_distribution | sort(attribute='count', reverse=true) %} {{ item.length }} character{{ '' if item.length == 1 else 's' }}{% for example in item.examples[:3] %}{{ example.value }}{% if not loop.last %} / {% endif %}{% endfor %} {{ item.percent }}%{{ item.count | count }} |
08 / Detectors
Detectors and formats({{ detector_rows | length }})
What each detector recognized per column, with the formats it found. Primary: the interpretation shown as semantic type.
| {{ detector.id }}{% if detector.primary %} primary{% endif %} | {% if detector.status == 'failed' %}failed | – | – | – | {% else %}{{ percent_cell(detector.matched_percent, detector.matched_count, 'var(--ok)') }}{{ detector.ambiguous_count | count }} | {{ detector.invalid_count | count }} | {% for item in detector.formats %}{{ item.format }}{% else %}–{% endfor %} | {% endif %}
09 / Limits
Limits and diagnostics
What the scan could not measure completely, and the technical events it recorded. Counters and statistics continue past a limit.
- Untracked observations
- {{ lim.untracked_observations | count }}Paths beyond
scan.limits.max_fields - Truncated by depth
- {{ lim.depth_truncated_observations | count }}Deeper than
scan.limits.max_depth
Limited measures({{ lim.measures | length }})
{% if m.path is none %}dataset{% else %}{{ m.path }}{% endif %} | {{ measure_labels.get(m.measure, m.measure) }}{{ m.measure }} | {{ m.reason }}{% if reason_settings.get(m.reason) %}{{ reason_settings[m.reason] }}{% endif %} | {{ m.limit | count }} | {% if m.lower_bound is none %}– | {% else %}{{ m.lower_bound | count }} | {% endif %}
|---|
Diagnostics({{ lim.diagnostics | length }})
- {% for d in lim.diagnostics %}
- {{ d.level }}
{{ d.code }}{{ d.message }}{% if d.dataset is none %} whole scan{% endif %}{% if d.locations %}Records {% for loc in d.locations %}{{ loc.record }}{% if loc.line is not none %} (line {{ loc.line }}){% endif %}{% if not loop.last %}, {% endif %}{% endfor %}{% if d.count > d.locations | length %}…{% endif %}{% endif %}{{ d.count | count }} {% endfor %}
10 / Raw values
Data sample({{ ds.preview | length }})
First {{ ds.preview | length }} records with original row numbers and raw values.
Row | {% for column in ds.columns %}{{ column.name or '(unnamed)' }}{{ column.inferred_type }} | {% endfor %}||
|---|---|---|---|
| {{ row.row_number }} | {% for value in row.values %}{% if loop.index0 in row.absent %}absent | {% elif value is none %}hidden | {% else %}{% if value == '' %}empty{% else %}{{ value }}{% endif %} | {% endif %}{% endfor %}
No data records.
{% endif %}11 / Method
Analysis settings
Rules used for this analysis, so the result can be reproduced.
- Scope
- All {{ ds.summary.row_count | count }} records {% set scan = report.config.scan %}
- Missing values
- {{ scan.values.missing | join(', ') }}; markers: {% for marker in scan.values.null_markers %}
{{ marker | tojson }}{% else %}none{% endfor %} - Marker comparison
- Whitespace trimmed; {{ 'case-sensitive' if scan.values.null_markers_case_sensitive else 'case-insensitive' }}
- Value normalization
- Unicode composition (NFC): {{ 'enabled' if scan.normalization.nfc else 'disabled' }}; trim: {{ 'enabled' if scan.normalization.trim else 'disabled' }}; collapse internal whitespace: {{ 'enabled' if scan.normalization.collapse_whitespace else 'disabled' }}; comparison also with case folding: {{ 'enabled' if scan.normalization.casefold else 'disabled' }}, accents removed: {{ 'enabled' if scan.normalization.strip_accents else 'disabled' }}. Raw preview values are preserved.
- Duplicate comparison
- {% if scan.records.duplicates %}Exact raw values across every column; up to {{ scan.limits.max_tracked_records | count }} distinct records compared{% else %}Disabled{% endif %}
- Type inference
- At least {{ (scan.types.minimum_confidence * 100) | round(1) }}% agreement across present values; numbers with {{ scan.detectors.number.conventions | join(' or ') }} decimals, configured dates and true/false. Leading-zero identifiers remain text.
- Date detection
- {% if scan.detectors.date.enabled %}Orders: {{ scan.detectors.date.orders | join(', ') }}; separators: {% for separator in scan.detectors.date.separators %}
{{ separator | tojson }}{% endfor %}; month names: {{ scan.detectors.date.month_languages | join(', ') or 'none' }}; ambiguous order: {{ scan.detectors.date.ambiguous_order or 'unresolved, column evidence shown but never applied' }}{% else %}Disabled{% endif %} - Semantic types
- Dates, or the only detector matching at least {{ (scan.detection.minimum_share * 100) | round(1) }}% of present values
- Enumerations
- {% if scan.detectors.enumeration.enabled %}At most {{ scan.detectors.enumeration.maximum_distinct }} distinct values among at least {{ scan.detectors.enumeration.minimum_values | count }} present values{% else %}Disabled{% endif %}
- Sensitive values
- {{ {'mask': 'Masked', 'hide': 'Hidden', 'show': 'Shown'}[scan.exposure.sensitive_values] }} in examples and the data sample
- String lengths
- Very short through {{ report.config.string_analysis.very_short_max_length }}; short through {{ report.config.string_analysis.short_max_length }}; medium through {{ report.config.string_analysis.medium_max_length }}; long through {{ report.config.string_analysis.long_max_length }}; otherwise very long. Up to {{ report.config.string_analysis.examples_per_length }} examples per length are retained when the maximum is {{ report.config.string_analysis.length_distribution_max_length }} or less.
- Value representation
- Complete through {{ report.config.value_examples.full_distribution_max_distinct }} distinct values; otherwise a reproducible sample of {{ scan.limits.max_samples }} distinct values
- Row numbering
- Data records start at 1, excluding the header. Quoted multiline values count as one record.
- Preview selection
- First {{ scan.records.preview }} records; raw values preserved
- JSON format
- {{ report.format_version }} - revision {{ report.format_revision }} (experimental)
{{ report.source.sha256 }}