In plain terms
Think of a cook working without a recipe. Taste, decide what is missing, add it, taste again. Nobody told the cook in advance how many rounds it would take; the dish decides. An agent works the same way. Each round is one call to the model: it looks at everything so far, picks the next move and sees what comes back. That loop is what turns a model that answers once into a system that keeps working.
Why it matters
The loop is where an agent's flexibility, cost and risk all come from. Because the model decides in each round whether to carry on, nobody knows in advance how many rounds a task will take, so cost and duration vary by design. Errors also carry forward: a wrong reading in round three shapes rounds four to twenty. This is why every serious deployment sets limits on the loop itself: a maximum number of steps, a spending cap, a time limit and a way for a person to stop it.
Example
A logistics company asks its operations agent why shipment 4471 is late. Round one: it queries the tracking system. Round two: it reads the carrier's status page. Round three: it checks the customs record and finds a missing document. Round four: it writes the explanation and requests no tool, so the loop ends. Four model calls, three tool calls, 40 seconds; the step limit of 25 was never reached.
Most often confused with
Agent Loop vs. Agentic Workflow
Both call a model several times. In a workflow the developer wrote the sequence, so every run has the same steps and roughly the same cost. In an agent loop the model chooses the number and order of steps while it works. That is why a loop needs stopping rules and a budget, and why two runs of one task can differ in length and price.
Under the hood
In code the loop is short: send the system prompt, tool definitions and message history to the model; if the reply contains tool calls, execute them, append the results and call the model again; if it contains none, stop. One pass is usually called a turn or a step, and a model may request several tool calls in parallel within one. Stopping conditions: the model ends its turn, a maximum number of turns, a token or cost budget, a timeout, an error, or a human interrupt. The whole history is resent on every turn, so input cost grows with the length of the run; prompt caching and compaction keep it in check. Failure modes: repeating the same failing call, declaring success too early, and compounding error, since twenty consecutive steps that are each 95% reliable all succeed only about 36% of the time. Hooks before and after each tool call are where permissions, guardrails and tracing attach.