Infrastructure & economics

Edge AI

On-device AI

AI that runs on the device where the data arises, such as a phone, a laptop, a car, a camera or a machine on the factory floor, with no need to send each request to a data centre.

ON THE DEVICE · beside the production lineCamera600 bottles/minuteSmall vision modelon the deviceDecision15 msWHEN NEEDEDCloudrecords, retraining0.4% of imagesWhy at the edge?latency · privacy · offline · no per-request feeLimitsmodel size · battery and heat · fleet updatesMost products combine the two: routine work on the device, hard cases in the cloud.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/edge-ai

In plain terms

A pocket calculator works in a tunnel, answers instantly and tells nobody what you typed, because the arithmetic happens in your hand. Edge AI puts a model in the same position: it sits inside the phone, the car or the camera and does its work there. The price is size. Only a model small enough for the device's chip, memory and battery can live on it.

Why it matters

Four things change when the model sits on the device. Answers arrive in milliseconds, with no network round trip. Raw data, whether a face, a voice or a production line, stays where it is. The system works without a connection. And there is no fee per request, because hardware already paid for does the work. The limits are firm: small models are less capable, heavy use drains the battery, and every update must reach thousands of devices of different ages. Most products therefore combine the two: routine work on the device, hard cases in the cloud.

Example

A bottling plant checks 600 bottles a minute for cap defects. Sending every image to the cloud would take about 300 milliseconds per bottle, too slow for the line, and inspection would stop whenever the connection dropped. A small vision model on a device beside the camera decides in 15 milliseconds. Only the 0.4% of images flagged as defective are uploaded, for quality records and retraining.

Most often confused with

Edge AI vs. On-premise

Edge AIRuns on the end device: one small model per device
On-premiseRuns on servers in your own data centre, shared by many users

Both keep data away from an outside provider, and the similarity ends there. On-premise means a server room with GPUs that serves the whole organisation and can run large models. Edge means the model is inside the product or the worker's device, limited by its chip and battery, and updated across a fleet. A factory can use both: cameras decide at the edge; an on-premise server analyses the day's production.

Under the hood

Hardware: neural processing units (NPUs) in phones and laptops, GPUs in vehicles and industrial computers, microcontrollers in the smallest sensors. Models are made to fit through distillation, pruning and quantisation, commonly down to 4-bit weights for language models. Runtimes include Core ML on Apple devices, LiteRT (the successor to TensorFlow Lite), ONNX Runtime, ExecuTorch and llama.cpp. Operating systems and browsers now ship built-in models that applications can call, such as Gemini Nano on Android and in Chrome, and the on-device models of Apple Intelligence. Metrics that matter: memory footprint, tokens or frames per second, energy per inference and load time. The common architecture is hybrid: a small on-device model for routine or private work, with escalation to a cloud model. Operational issues: fragmentation across chips and operating-system versions, staged model updates, monitoring quality without collecting user data, and protecting model weights that sit on hardware the customer holds. The term also covers local gateways and base stations placed close to the devices.

Written by Mehmet Erkek · Last updated: