Audit before training

Find image dataset leakage and shortcut risks

Inspect up to 200 train, validation and test images for exact leakage, near-duplicate derivatives, label conflicts, filename families and resolution shortcuts.

Audit a dataset
Completely free No account or payment Image stays on this device
USE THIS WHENUse this before model training or evaluation to find cross-split copies, label conflicts and source shortcuts that may inflate results.SUPPORTED INPUTUp to 200 dataset images
What you will getA clear result before technical detail
01A severity-ranked cross-split evidence queue
02Dataset split, label and shortcut summaries
03CSV and JSON reports for cleanup and reproducibility
Private analysis labFiles stay inside this browser
DATASET / 200 FILES
Choose a train / validation / test folderUp to 200 JPEG, PNG or WebP files · folder names are used only to infer split and label
Selection is processed locally
Session diagnosticsNo processing errors

Stored only in this browser tab. Image bytes, filenames and metadata are never included.

No tool error has been recorded in this tab.
No upload or account Originals remain unchangedReview methodology
Useful and careful

What is built into the result.

Important findings come first. Technical fields remain available without taking over the page.

01

Reads common split and label folders without uploading the dataset

02

Combines byte-exact and perceptual duplicate evidence

03

Separates confirmed byte identity from near-match review candidates

Understand the method

A useful result with its reasoning attached.

This page explains what the lab measures, how to interpret it and where human review remains essential.

01

Why visual dataset leakage is different from duplicate cleanup

A computer vision test set should represent unseen evaluation data. If the same image, a resized copy or a crop appears in training and testing, reported accuracy can describe memory of the dataset rather than generalization. Conflicting labels on identical bytes create a separate data-quality problem.

The auditor reads train, validation and test folder names when available. It computes exact SHA-256 identifiers and three compact perceptual fingerprints for each image, then reviews only cross-split pairs above the selected evidence threshold.

  • Exact cross-split copies
  • Near-duplicate derivatives
  • Identical bytes with conflicting labels
  • Related filename families
02

Shortcut signals the model may learn

Even when no duplicate exists, a model can exploit a source artifact that correlates with the label. Different image widths, aspect ratios, compression pipelines, corner text or acquisition devices can become easier predictors than the intended subject.

This browser release reports class-level differences in resolution and aspect-ratio profiles and surfaces filename families that cross splits. These are screening signals. A domain expert must decide whether the correlation is legitimate, spurious or caused by how the dataset was assembled.

  • Class-specific resolution profiles
  • Aspect-ratio separation
  • Source-key review
  • Transparent evidence values
03

A cleanup report designed for reproducibility

Every finding records severity, type, affected files and the evidence that triggered review. CSV is convenient for filtering and assignment; JSON preserves fingerprint distances, split names and the audit threshold for a future run.

The tool does not delete, move or rename source images. Cleanup remains an explicit action outside the browser so a researcher can preserve the original dataset and document every exclusion.

  • No automatic deletion
  • Versioned report structure
  • User-controlled threshold
  • Explicit methodology limits
Three simple steps

Know what happens before you start.

Use this before model training or evaluation to find cross-split copies, label conflicts and source shortcuts that may inflate results.

  1. SelectChoose an extracted dataset folder; train, validation, test and label names are inferred locally.
  2. AuditCompare exact and perceptual fingerprints, filename families and class-level image profiles.
  3. TriageReview severity-ranked evidence and export CSV or JSON without modifying a source file.
RESULT ORDER
01 · A severity-ranked cross-split evidence queue02 · Dataset split, label and shortcut summaries03 · CSV and JSON reports for cleanup and reproducibility04 · Limits and next step
Clear before you rely on it

Questions this tool should answer upfront.

Short answers keep important privacy, evidence and professional-use limits visible.

01Does a near-duplicate finding prove data leakage?

No. It identifies strong visual evidence across dataset splits. Review the pair, subject identity, capture sequence and experimental protocol before deciding whether it is leakage.

02Can the auditor open a ZIP dataset?

This release reads a selected folder or image set directly. It intentionally avoids unpacking untrusted ZIP paths in the browser; extract the dataset locally first.

03Why is the browser run limited to 200 images?

Pairwise visual comparison grows quickly and mobile memory is limited. The cap keeps the page responsive and prevents a public browser tool from becoming an uncontrolled compute service.

Continue when useful

Your next step, without starting over.

Move to another page only when its outcome matches what you need.

Important limitations

The browser edition audits up to 200 images per run · Perceptual hashes do not replace embedding or domain-specific leakage analysis · A clean report cannot prove that subjects, acquisition sites or labels are independent

Full limitations