Foundations

Neural Network

Artificial neural network

A computing system made of layers of simple connected units, loosely modelled on neurons, that learns by adjusting the strength of the connections between them.

INPUT LAYERHIDDEN LAYEROUTPUT LAYERNeuronsums its inputs using weightsLayersbuild abstraction step by stepInspired by the brain but not a working model of it: in essence a very large mathematical function.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/neural-network

In plain terms

Picture a very large panel of dials. Data goes in on one side, passes through layers of small calculating units, and a result comes out on the other. Every connection has a dial, called a weight, that sets how much one unit influences the next. Learning means turning those dials a little at a time until the outputs are right. Nobody sets them by hand; the training process does.

Why it matters

It is the building block under every modern AI system; a language model is a neural network with billions of dials. Knowing this explains two things managers meet in practice. First, what the model knows is spread across its weights as numbers, so there is no line of code to inspect or correct when it is wrong. Second, its behaviour is changed by training or by what you place in its context, never by editing a rule.

Example

A network learns to read handwritten digits. An image of 28 by 28 pixels goes in as 784 numbers; two hidden layers transform them; ten outputs give the probability of each digit. At first its guesses are random. After it has seen 60,000 labelled examples, adjusting its weights after each small batch, it reads new handwriting with about 98% accuracy.

Most often confused with

Neural Network vs. The human brain

Neural NetworkA mathematical function with adjustable numbers
The human brainA living organ that is still poorly understood

The name suggests more than is there. Artificial neurons are a simplified idea borrowed from biology in the 1940s; they do not work as real neurons do, and a network does not learn the way a child does. The metaphor helps intuition and misleads as an explanation. The name is no evidence that these systems think as people do.

Origin: The idea dates from the 1940s; Frank Rosenblatt built the perceptron in 1958, and backpropagation made training multi-layer networks practical in 1986.

Under the hood

Each unit computes a weighted sum of its inputs, adds a bias and applies a non-linear activation function such as ReLU. Units are arranged in layers: input, hidden, output. Training repeats four steps over many batches of data: run a forward pass; measure the error with a loss function; use backpropagation to work out how much each weight contributed to the error; and shift the weights slightly with gradient descent. Architectures differ in how the units are wired: fully connected, convolutional, recurrent, transformer. Size is counted in parameters, from thousands to hundreds of billions. Known properties: with enough units a network can approximate almost any function, it needs a great deal of data, and its internal representations are hard to interpret.

Written by Mehmet Erkek · Last updated: