Model customization

Data Labeling

Data annotation · data labelling

Attaching the correct answer to each example in a dataset, such as a category, a tag, a box around an object or a quality score, so that a model can learn from it or be tested against it.

“Money was taken from my card”a customer messageA: cardB: fraudAgreement 78%on the same 500 messagesRewrite the guidelinewith borderline examplesAgreement 94%on the same 500 messagesLabel 6,000 messagesevery tenth one twiceMost quality problems are guideline problems. Illustrative: three weeks of labelling, an afternoon of training.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/data-labeling

In plain terms

A model learns from examples the way a trainee learns from a marked answer sheet. Someone has to do the marking: this email is a complaint, this X-ray shows a fracture, this answer is acceptable and that one is not. Labelling is that marking, done thousands of times and to one consistent standard. The unglamorous truth of many AI projects is that most of the human effort goes here.

Why it matters

Labels set the ceiling for the model: it cannot learn a distinction that the labellers did not make consistently. They are also where cost and time go, since careful labelling needs people who understand the subject, and expert hours are expensive. Language models have changed the economics: they can propose labels for people to check, and many tasks now need a few hundred labelled examples for testing where they once needed tens of thousands for training. The need has not gone away. Without a trusted labelled set there is no way of knowing whether any system, bought or built, works.

Example

A bank wants to route 30,000 customer messages a month into twelve categories. Two staff label the same 500 messages and agree on only 78%, mostly where “card” and “fraud” overlap. The guideline is rewritten with examples of the borderline, agreement rises to 94%, and 6,000 messages are then labelled, every tenth one twice as a check. The labelling takes three weeks. Training the model takes an afternoon.

Most often confused with

Data Labeling vs. Ground truth

Data LabelingThe activity: people or models assigning answers
Ground truthThe result you trust: answers accepted as correct

Labelling produces labels; only some labels deserve to be called ground truth. A single annotator's quick judgement is a label. An answer that two experts agreed on, or that was confirmed by what later happened, is ground truth. The difference matters when measuring a model: scoring it against careless labels measures the carelessness as well as the model.

Under the hood

Label types follow the task: classes, entity spans in text, bounding boxes and segmentation masks in images, transcripts for audio, rankings, and for language-model work ratings or pairwise preferences between answers. The core document is the labelling guideline, with definitions, borderline cases and examples; most quality problems are guideline problems. Quality controls: overlapping assignments, inter-annotator agreement (for example Cohen's kappa), seeded items with known answers, and adjudication of disagreements by an expert. Ways to cut cost: pre-labelling by a model with human review, active learning, which sends people the examples the model is least sure of, and weak supervision from rules. Tools include Label Studio and Amazon SageMaker Ground Truth; much of the work is outsourced to specialist providers, which raises questions of confidentiality and of working conditions. Watch for label bias, class imbalance and labels that go stale when policies or products change.

Written by Mehmet Erkek · Last updated: