// Explainer · Sep 27, 2026
AI Inference
The phase where a trained AI model processes an input and generates an answer.

// How it works
Training is when an AI model learns, running massive batches on expensive hardware. Inference happens billions of times a day when you ask a question. The model does not get smarter; it simply runs math set during training to produce an output.
// For example
Browser Use demonstrated an agent completing a flight booking in 7 seconds for $0.0039 in inference cost.
// The fuller picture
Training is when an AI model learns, but inference is when it answers 1. Training happens occasionally, in massive batches, on the most expensive hardware on Earth 1. Inference happens billions of times a day on whatever hardware sits closest to the user 1.
Every time you ask an AI model a question, watch an AI-generated video, or receive a transcript from a meeting tool, you trigger an inference 1. During this step, the model does not learn or get smarter 1. It simply runs the math that was set in stone during its prior training run to produce an output 1.
Most of the AI economy now hinges on inference 1. In real-world tasks, this compute translates directly to operational cost 12. For example, Browser Use founder Gregor Zunic demonstrated a browser agent completing an end-to-end flight booking in 7 seconds for $0.0039 in inference cost 24. Meanwhile, running perception models locally offers a way to avoid per-inference bills or datacentre round trips 3.
// Sources
- [1] Inference, explained
- [2] Episode 2026-W39 · Four frontier models shipped in seven days, the price of a token fell by half, and the people who built them went to the UN Security Council to explain themselves
- [3] Devices & Robotics — W39: the reflex-plus-planner robot brain shows up in Minecraft, and the voice tier gets its own model
- [4] LLM Weekly — W39: GPT-6 halves the price of a token, four frontier models ship in seven days
// Where we covered it
- episode · 2026-09-25Episode 2026-W39 · Four frontier models shipped in seven days, the price of a token fell by half, and the people who built them went to the UN Security Council to explain themselves
- article · 2026-09-25Devices & Robotics — W39: the reflex-plus-planner robot brain shows up in Minecraft, and the voice tier gets its own model
- article · 2026-09-25LLM Weekly — W39: GPT-6 halves the price of a token, four frontier models ship in seven days
- article · 2026-09-25OpenAI Cut Token Prices 50%. Every H1 Business Case Is Now Mispriced.