In plain terms
In ordinary software a person writes the rule: “flag the payment if it exceeds 10,000”. In machine learning you show the system thousands of past payments, each marked as fraud or genuine, and it works out the rule itself. The result is called a model, and it can judge payments it has never seen.
Why it matters
It is the right tool where the rules are too many, too subtle or too changeable to write down, and examples are plentiful: fraud, demand forecasting, customer churn, credit risk, quality inspection, recommendations. Its two requirements are also its limits: enough good historical data, and a future that resembles the past. A model is only as good, and as fair, as the examples it learned from.
Example
A retailer wants to know how many units of each product each store will sell next week. The planners' rule of thumb is off by 28% on average. A machine-learning model trained on three years of sales, prices, promotions, weather and holidays brings the error down to 14%. Nobody told it that ice cream sells in a heatwave; it found that in the data.
Most often confused with
ML vs. Deep learning
Deep learning is a part of machine learning. Classical methods such as decision trees and regression work well on tables of numbers, need less data and are easier to explain. Deep learning leads where the input is unstructured: images, audio and language. For a business problem shaped like a spreadsheet, the classical methods are often still the better choice.
Origin: The term was popularised in 1959 by Arthur Samuel of IBM, in his work on a program that learned to play checkers.
Under the hood
Three main settings: supervised learning (labelled examples, for classification and regression), unsupervised learning (finding structure without labels, as in clustering) and reinforcement learning (learning from reward). The workflow: define the target, collect and clean the data, prepare features, split into training, validation and test sets, train, evaluate, deploy and monitor. Common algorithms: linear and logistic regression, decision trees, random forests, gradient boosting (XGBoost, LightGBM), support vector machines and neural networks. The central risks: overfitting (memorising the training data), data leakage (the answer slipping into the inputs), bias inherited from historical data, and drift when the world changes after deployment.