In plain terms
A bathroom scale that always shows two kilos too much is biased: its error is the same every time, so weighing yourself again does not help. A model can be skewed in the same way, and the skew can fall unevenly on people. A system trained on a company's past hiring decisions learns the pattern in those decisions, including preferences nobody would defend today, and then applies it to every new candidate, consistently and at scale.
Why it matters
A biased model repeats the same unfair decision thousands of times, which turns an old habit into legal, financial and reputational exposure in hiring, lending, insurance and pricing. It also loses money directly: good candidates and creditworthy customers are turned away. The limit should be stated plainly: bias cannot be removed completely. It can be measured, reduced and monitored, and that requires deciding which groups to compare and which definition of fairness applies. Those are management and legal decisions; a data science team cannot make them alone.
Example
A company trains a screening model on ten years of its own hiring decisions. An audit compares candidates with the same skills-test score: 34% of the graduates of the three universities the company has always hired from are invited to interview, and 18% of everyone else. The model had learned the old habit. After the university name and related signals are removed and the training data is rebalanced, the figures are 31% and 28%.
Most often confused with
Bias vs. Fairness
Bias can be measured: approval rates, error rates and scores can be compared across groups. Fairness is the standard those numbers are held against, and it has several definitions. Equal approval rates, equal error rates and equally reliable scores usually cannot all be met at once. Measuring bias is a technical task. Choosing the fairness goal is a decision for management and legal counsel, taken before the model is built.
Under the hood
Two senses. In statistics, bias is systematic error, the counterpart of variance, which is random scatter. In AI ethics it means systematically different outcomes for groups. Sources along the pipeline: historical bias in the data, unrepresentative samples, labels that record past human judgements, proxy features that stand in for a protected attribute (a postcode can stand in for income or origin), the choice of target, and use on a population the model was not built for. Language models add stereotyped associations absorbed from web text. Metrics: demographic parity, equal opportunity, equalised odds and calibration within groups; several of these are mathematically incompatible when base rates differ. Mitigation: better data collection, reweighting or resampling, constraints in training, group-specific thresholds where the law permits, and human review. Removing the protected attribute alone rarely works, because proxies remain. Tooling includes Fairlearn and AI Fairness 360; NIST Special Publication 1270 groups AI bias into systemic, statistical and human.