Dataset Validation & QA
for Annotated Image Datasets.
Find the annotation errors before your model learns from them. IAA measurement, label consistency audit, and structured quality reporting on existing datasets.
What Is Dataset Validation?
Dataset validation is the systematic review of an annotated dataset before it is used for model training. It audits the dataset for label consistency errors, class distribution issues, systematic annotation mistakes, and guideline violations — and produces a structured report with specific findings and reannotation recommendations.
Annotation errors are significantly cheaper to fix before training than after. A model trained on a corrupted dataset must be retrained from scratch once the errors are discovered — which may only happen after a failed evaluation or a production incident. Dataset validation catches these problems at the source.
Aeterna Data's validation service measures inter-annotator agreement on existing datasets, identifies error patterns by class and by annotator, and delivers a structured quality report with specific reannotation recommendations. The report tells you exactly what is wrong, where it is, and how to fix it.
When to Validate
Before First Training Run
Validate a new dataset before the first training run to catch systematic annotation errors before they corrupt the baseline model.
After Crowd Annotation
Crowd-sourced annotation frequently produces inconsistent labels across annotators. Validate before using crowd-labeled data in a production training pipeline.
Before Model Iteration
When adding new data to an existing training set, validate the new data for consistency with the existing annotation standards before merging.
After Model Underperformance
If a model performs below expectations on a specific class or scenario, dataset validation can identify whether annotation quality is a contributing cause.
What We Audit
Label Consistency
Whether the same class label is applied consistently to the same type of object or image across the full dataset. Inconsistency at the class level is the most common source of annotation quality problems.
Boundary Accuracy (for detection and segmentation datasets)
Whether bounding box or polygon boundaries are drawn consistently around objects of the same class. Loose boundaries, tight boundaries, and misaligned boundaries are each measured and reported by class.
Class Distribution
Whether the class distribution in the annotated dataset reflects the actual distribution in the source data — or whether certain classes are systematically over- or under-labeled.
Systematic Annotator Errors
Whether specific annotators are producing errors at a higher rate than others, and whether those errors follow a pattern — for example, consistently mislabeling a specific class or consistently drawing boundaries too loosely.
Guideline Violations
Whether annotations violate the stated annotation guidelines — for example, annotating objects that should be skipped, or missing attributes that are required by the guidelines.
Edge Case Coverage
Whether the dataset adequately represents the edge cases the model will encounter in deployment — and whether those edge cases are correctly annotated.
IAA Measurement on Existing Datasets
Inter-annotator agreement measurement on an existing dataset works by having a second annotator label a sample of the data independently, then comparing the two sets of labels using Cohen's Kappa (for pairwise comparison) or Fleiss' Kappa (for multi-annotator comparison).
The IAA score tells you how consistent the original annotation is — not just overall, but by class and by annotator. A dataset might have high overall IAA but low IAA on a specific class that is critical to model performance. That class-level measurement is what drives the reannotation recommendation.
| κ Score | Interpretation |
|---|---|
| κ < 0.40 | Poor — significant errors |
| 0.40 – 0.59 | Moderate — review needed |
| 0.60 – 0.79 | Substantial — minor issues |
| κ ≥ 0.80 | Strong — production ready |
Quality Report
Every dataset validation engagement delivers a structured quality report. The report is designed to give your engineering team a clear, actionable picture of the dataset's quality and exactly what needs to be fixed.
Executive Summary
Overall IAA score, total errors found, classes with quality issues, and overall reannotation recommendation — ready for technical and non-technical stakeholders.
IAA Analysis by Class
Per-class Cohen's Kappa scores with interpretation, highlighting which classes are production-ready and which require reannotation.
Annotator Performance Analysis
Per-annotator error rates and patterns, identifying whether quality issues are systematic across the dataset or concentrated in specific annotators.
Error Pattern Documentation
Specific error types found, with annotated examples — showing what the error looks like and what the correct annotation should be.
Reannotation Recommendation
Specific recommendation on which images, classes, or annotator batches require reannotation — with priority ranking by impact on model performance.
Guideline Revision Proposal
Where annotation errors trace back to ambiguous guidelines, we propose specific guideline revisions to prevent the same errors in future annotation rounds.
Our Workflow
Project Brief & Scoping
Share your annotated dataset, original annotation guidelines, and specific quality concerns if any. We scope the validation sample size and timeline.
Pilot Phase
We run IAA measurement on a validation sample, audit label consistency, and identify error patterns. Findings are reviewed before the full report is compiled.
Production Annotation
Full dataset annotation following the validated workflow. IAA measured on every batch. Rework at no cost for any batch below threshold.
Delivery & Documentation
Annotated dataset delivered in your specified format with a complete quality report — IAA scores, class distribution, and annotator notes.
Annotation Tools
Aeterna Data works with the annotation platform your team already uses. We do not require you to adopt a specific tool. If you do not have a platform, we can advise on setup based on your task type and scale.
How to Start
Send a Project Brief
Share your annotated dataset or a sample, your annotation guidelines, and the annotation format. Describe any specific quality concerns or classes you want prioritised in the audit.
Send BriefReceive a Pilot Proposal
We send a scoped pilot proposal within 48 hours — sample size, timeline, flat-rate pilot fee, and deliverables.
NDA and DPA Signed
Before any data is shared, NDA and Data Processing Agreement are signed. Your dataset stays confidential.
Pilot Begins
Annotation starts inside your platform. IAA report delivered with the pilot dataset. Production follows on your confirmation.
Ready to Audit Your Dataset?
IAA measurement on your existing annotations. Error patterns identified by class and annotator. Structured quality report with reannotation recommendations.