Data Annotation and Training Data: What It Actually Involves
Annotation looks like the simple part of a machine-learning project and is where most of them are quietly ruined. What actually has to happen first.
Where it goes wrong
Two annotators shown the same ambiguous photograph will label it differently unless someone has already decided, in writing, what to do with a partially hidden object, a reflection, a blurred edge, or a case that falls between two classes. Annotation looks like the simple part of a machine-learning project and is exactly where inconsistency quietly enters it.
Guidelines before volume
A specimen set is labelled first, the disagreements are found, the guideline is amended to settle them, and only then does volume work begin. It is slower for the first week and considerably cheaper across the project.
What gets annotated
- Bounding boxes and 3D-oriented boxes for detection and counting
- Polygon and semantic segmentation where the shape matters, not just the location
- Keypoint and landmark annotation for pose and alignment tasks
- Video annotation with object tracking across frames
- Text annotation — entity extraction, intent and sentiment labelling
How quality is actually measured
Inter-annotator agreement is calculated and reported, not assumed. A sampled percentage of every batch is independently re-checked, and a failed sample means the batch is redone. You keep a held-back gold set so our quality claims are checkable independently.
How it is priced
Per annotation unit — per box, per polygon, per frame, per document — with the rate set by how fiddly the task is and how tight the required agreement level is. Segmentation costs several times what a bounding box does, because it takes several times as long.