AeternaData
Dataset QA

Dataset Validation & QA
for Annotated Image Datasets.

Find the annotation errors before your model learns from them. IAA measurement, label consistency audit, and structured quality reporting on existing datasets.

IAA MeasurementCore Method
κ ≥ 0.80Quality Threshold
48 HoursScoping Response
Flat RatePilot Entry

What Is Dataset Validation?

Dataset validation is the systematic review of an annotated dataset before it is used for model training. It audits the dataset for label consistency errors, class distribution issues, systematic annotation mistakes, and guideline violations — and produces a structured report with specific findings and reannotation recommendations.

Annotation errors are significantly cheaper to fix before training than after. A model trained on a corrupted dataset must be retrained from scratch once the errors are discovered — which may only happen after a failed evaluation or a production incident. Dataset validation catches these problems at the source.

Aeterna Data's validation service measures inter-annotator agreement on existing datasets, identifies error patterns by class and by annotator, and delivers a structured quality report with specific reannotation recommendations. The report tells you exactly what is wrong, where it is, and how to fix it.

When to Validate

Before First Training Run

Validate a new dataset before the first training run to catch systematic annotation errors before they corrupt the baseline model.

After Crowd Annotation

Crowd-sourced annotation frequently produces inconsistent labels across annotators. Validate before using crowd-labeled data in a production training pipeline.

Before Model Iteration

When adding new data to an existing training set, validate the new data for consistency with the existing annotation standards before merging.

After Model Underperformance

If a model performs below expectations on a specific class or scenario, dataset validation can identify whether annotation quality is a contributing cause.

What We Audit

Label Consistency

Whether the same class label is applied consistently to the same type of object or image across the full dataset. Inconsistency at the class level is the most common source of annotation quality problems.

Boundary Accuracy (for detection and segmentation datasets)

Whether bounding box or polygon boundaries are drawn consistently around objects of the same class. Loose boundaries, tight boundaries, and misaligned boundaries are each measured and reported by class.

Class Distribution

Whether the class distribution in the annotated dataset reflects the actual distribution in the source data — or whether certain classes are systematically over- or under-labeled.

Systematic Annotator Errors

Whether specific annotators are producing errors at a higher rate than others, and whether those errors follow a pattern — for example, consistently mislabeling a specific class or consistently drawing boundaries too loosely.

Guideline Violations

Whether annotations violate the stated annotation guidelines — for example, annotating objects that should be skipped, or missing attributes that are required by the guidelines.

Edge Case Coverage

Whether the dataset adequately represents the edge cases the model will encounter in deployment — and whether those edge cases are correctly annotated.

IAA Measurement on Existing Datasets

Inter-annotator agreement measurement on an existing dataset works by having a second annotator label a sample of the data independently, then comparing the two sets of labels using Cohen's Kappa (for pairwise comparison) or Fleiss' Kappa (for multi-annotator comparison).

The IAA score tells you how consistent the original annotation is — not just overall, but by class and by annotator. A dataset might have high overall IAA but low IAA on a specific class that is critical to model performance. That class-level measurement is what drives the reannotation recommendation.

κ ScoreInterpretation
κ < 0.40Poor — significant errors
0.40 – 0.59Moderate — review needed
0.60 – 0.79Substantial — minor issues
κ ≥ 0.80Strong — production ready

Quality Report

Every dataset validation engagement delivers a structured quality report. The report is designed to give your engineering team a clear, actionable picture of the dataset's quality and exactly what needs to be fixed.

Executive Summary

Overall IAA score, total errors found, classes with quality issues, and overall reannotation recommendation — ready for technical and non-technical stakeholders.

IAA Analysis by Class

Per-class Cohen's Kappa scores with interpretation, highlighting which classes are production-ready and which require reannotation.

Annotator Performance Analysis

Per-annotator error rates and patterns, identifying whether quality issues are systematic across the dataset or concentrated in specific annotators.

Error Pattern Documentation

Specific error types found, with annotated examples — showing what the error looks like and what the correct annotation should be.

Reannotation Recommendation

Specific recommendation on which images, classes, or annotator batches require reannotation — with priority ranking by impact on model performance.

Guideline Revision Proposal

Where annotation errors trace back to ambiguous guidelines, we propose specific guideline revisions to prevent the same errors in future annotation rounds.

Our Workflow

01

Project Brief & Scoping

Share your annotated dataset, original annotation guidelines, and specific quality concerns if any. We scope the validation sample size and timeline.

02

Pilot Phase

We run IAA measurement on a validation sample, audit label consistency, and identify error patterns. Findings are reviewed before the full report is compiled.

03

Production Annotation

Full dataset annotation following the validated workflow. IAA measured on every batch. Rework at no cost for any batch below threshold.

04

Delivery & Documentation

Annotated dataset delivered in your specified format with a complete quality report — IAA scores, class distribution, and annotator notes.

Annotation Tools

Aeterna Data works with the annotation platform your team already uses. We do not require you to adopt a specific tool. If you do not have a platform, we can advise on setup based on your task type and scale.

CVATLabel StudioRoboflowV7 DarwinCSV/JSON/COCO FormatCustom Pipeline

How to Start

Step 01

Send a Project Brief

Share your annotated dataset or a sample, your annotation guidelines, and the annotation format. Describe any specific quality concerns or classes you want prioritised in the audit.

Send Brief
Step 02

Receive a Pilot Proposal

We send a scoped pilot proposal within 48 hours — sample size, timeline, flat-rate pilot fee, and deliverables.

Step 03

NDA and DPA Signed

Before any data is shared, NDA and Data Processing Agreement are signed. Your dataset stays confidential.

Step 04

Pilot Begins

Annotation starts inside your platform. IAA report delivered with the pilot dataset. Production follows on your confirmation.

Ready to Audit Your Dataset?

IAA measurement on your existing annotations. Error patterns identified by class and annotator. Structured quality report with reannotation recommendations.