It's easy to assume that "using AI" means building an agent — a system that plans, reasons, calls tools, and pursues a goal across multiple steps. That's a legitimate and powerful pattern, and much of our work is exactly that. But it's not the only way LLMs create value, and it's often not even the highest-leverage one.

The other pattern

A fast, low-cost model — like Google's Gemini Flash line — can be called as a single step inside a conventional, deterministic application: classify this, score this, summarise this, decide this — and return control immediately. No planning loop, no autonomy, no multi-step reasoning. Just intelligence injected at one precise point in a process that otherwise runs exactly as it always has.

Why speed and cost change what's possible

The largest, most capable models are relatively slow and relatively expensive per call — reasonable for a complex, high-value agentic task, but uneconomical for a decision that has to run continuously across an entire operation. Fast, lightweight models like Gemini Flash are priced and built for a different job: extremely high call volumes, tight latency budgets, and simple, bounded decisions — the kind of judgment call that used to require either a hand-tuned rules engine or a human, and now doesn't need either.

That combination — good enough judgment, low latency, low cost per call — is what makes it possible to inject a model into places nobody would seriously consider putting a large, slow, expensive one: a single API call inside a request path that has to respond in milliseconds, at a volume of millions of calls a day.

The strategic unlock isn't "smarter AI." It's AI cheap and fast enough to sit inside processes that were never designed to wait for one.

A real example: traffic lights

Google's own Project Green Light is a clean illustration of this pattern operating at genuine civic scale. Rather than building a conversational agent, Google combined AI modelling with Google Maps driving-trend data to analyse traffic patterns at city intersections and recommend better signal-timing configurations to city traffic engineers. It's now running in dozens of cities worldwide, and reported reductions in stop-and-go traffic — and the resulting emissions — of up to 30% at participating intersections, using the existing traffic light hardware already installed at each junction.

Nothing about that system needs to "decide" anything autonomously in real time, hold a multi-turn conversation, or call external tools. It's intelligence applied to a narrow, well-defined judgment — how should this specific intersection's signal timing change — running at a scale (thousands of intersections) that would be impractical to solve by hand, and delivered as a recommendation back into an existing municipal process. That is the essential shape of this second pattern: not an autonomous agent, a smarter decision embedded at a critical junction of ordinary infrastructure.

Where this fits in an enterprise's AI portfolio

01

Look for high-volume, bounded decisions first

Routing a support ticket, flagging a transaction for review, scoring a lead, tagging a document — decisions made thousands of times a day, each one simple enough for a fast model, too varied for a fixed rules engine.

02

Keep it inside your existing process, not a new one

The highest-leverage version of this pattern doesn't ask users to adopt a new chat interface — it improves a process they already run, invisibly, the way Green Light improves traffic lights the driver never notices changed.

03

Match the model to the job, not the other way around

Reserve larger, slower, more expensive models for the genuinely complex, high-value reasoning tasks. Route the high-volume, lower-stakes decisions to fast, cheap models — often through a unified interface so the choice stays flexible as models improve.

04

Apply the same governance discipline regardless of pattern

A single fast model call embedded in a critical process still needs the same rigor as an agent — guardrails on its output and a clear escalation path when its confidence is low.

The executive takeaway

When your team asks "where should we use AI?", resist the instinct to only answer with agent projects. Some of the best opportunities are quieter: a single well-placed, fast, inexpensive model call embedded at a critical decision point inside a process you already run — at a scale no team of humans, and no fixed rules engine, could match.

Key takeaways

  • Not all valuable AI is agentic — a single fast, cheap model call embedded in an existing process is a distinct, high-leverage pattern.
  • Fast, affordable models like Gemini Flash make it economical to inject intelligence into high-volume, latency-sensitive decisions.
  • Google's Project Green Light shows this pattern at civic scale — improving traffic light timing across thousands of intersections without any conversational agent.
  • Match the model to the job: save large models for complex reasoning, route high-volume bounded decisions to fast, cheap ones.

Wondering where a fast model could quietly upgrade a process you already run?

We help organisations find the highest-leverage places to inject AI — agentic or not.

Talk to us →