Author: Piyush Choubey - Practice Director
Key takeaways:
- Artificial intelligence (AI) budgets are rising fast, but the question attached to them has shifted; It is no longer how much AI a company uses, it is how much value that spending returns.
- Snowflake's dynamic model routing, part of the Cortex AI Gateway, automatically sends each step of an agent's work to the least expensive model that can still meet the quality bar, inside your approved-model list and data-residency rules, with every decision logged. It is pre-preview today, which makes this a moment to prepare rather than wait.
- The lasting advantage is not the router itself; It is the operating discipline around it. KPI Partners helps organizations set the approved-model policy, quality bars, and spend governance that make efficient AI a repeatable practice.
Every leader I talk with is spending more on AI this year than last, and most plan to spend more again next year. Industry forecasts put 2026 enterprise AI spend well into the trillions, a sharp rise on the year before. The budgets are there. What has changed in the last few months is the question attached to them. It is no longer how much AI a company is using; It is whether that spending is turning into value.
For the last two years, enterprise AI strategy has largely been framed as a model-selection problem: which large language model (LLM) should we standardize on? That may be the wrong question. As AI workloads become multi-step and agentic, the architecture itself needs to decide when a task requires frontier-level reasoning and when a faster, cheaper model is enough. The competitive advantage will increasingly come not from choosing one “best” model, but from orchestrating intelligence economically. In that framing, model choice stops being a procurement decision and becomes a runtime decision the platform makes for you.
The common instinct, when the invoice grows, is to slow down, or to pick one cheaper model and use it everywhere. My view is that both moves miss the point. The waste in enterprise AI is rarely that a company uses too much AI. It is that it uses the most powerful and most expensive model for tasks that never needed it. The opportunity is to pay only for the intelligence each task actually requires. Snowflake's dynamic model routing, announced in August 2026 for the Cortex AI Gateway, is a direct move in that direction, and because it is still early, now is the moment to build the discipline that will make the most of it.
Spending More on AI is Not the Same as Getting More from It
The stakes are easy to see in the numbers, and so is the shift in mood. Alongside the growth in overall spending, industry observers have noted that as investment in AI platforms and models climbs, enterprise budgets are coming under closer scrutiny, with a sharper focus on usage efficiency, cost control, and measurable outcomes, and that spending is moving toward providers who can show clear value across cost, latency, performance, and reliability. That is not a retreat from AI. It is a maturing market asking a fair question.
Here is where the value quietly leaks. Inside almost any AI workflow, the tasks are not equally hard. Summarizing a support ticket, drafting a product description, or tidying a routine data-pipeline step is not the same kind of work as forecasting demand across a supply chain or reasoning through a complex root-cause analysis.
Yet when every request is sent to the same frontier model by default, the easy tasks are billed at the price of the hard ones. Multiply that across dozens of agents and thousands of requests a day, and the gap between what you spend and what you needed to spend becomes the number a Chief Financial Officer (CFO) and a Chief Information Officer (CIO) now want to close. The encouraging part is that this is largely an efficiency problem, and one that can often be addressed without treating every task as if it requires maximum intelligence.
The Real Question is Which Model Fits the Task, Not Which Model is Best
Dynamic model routing answers that question automatically. Within the Cortex AI Gateway, it selects the most affordable model that can confidently complete each step of an agent's work. Routine and repetitive steps go to efficient models; Steps that need deeper reasoning go to frontier models. Snowflake describes the mechanism as an advisor pattern: A smaller model attempts the task first, and if it cannot finish, it calls a larger model as a tool and hands off the rest, while a classifier trained on your own history learns which requests should skip the small model entirely. The point is that your teams no longer have to build and maintain model-selection logic by hand, and as new models arrive or prices change, the routing adapts without anyone rewriting the application.
What makes this safe for an enterprise is where it sits. The router only chooses among models an administrator has approved, and it checks your existing data-residency settings before it ever weighs cost or latency, so requests stay inside the compliance boundaries you have already set. Every routing decision is logged, which gives the people accountable for both spend and risk a clear record of which model handled which request. The capability is built into Snowflake's own AI products, including CoCo and CoWork, and extends to third-party agents that connect through the gateway. In Snowflake’s own internal testing, routing delivered up to 3x greater token efficiency on a data build tool (dbt) pipeline workload and roughly 25% fewer tokens on a coding workload, both at comparable quality. Those are useful illustrations of what efficient routing can do with Snowflake’s own figures as the source, not independent proof.
One honest note on timing. As of late August 2026, dynamic model routing is described as entering private preview soon, and the newest open models it can draw on are in private preview or coming shortly, even though the Cortex AI Gateway that houses it has been available since July. This is not a reason to wait. It is the reason to start now on the part that is entirely in your hands, which is the policy and the discipline around model choice, so you are ready to switch it on with intent the moment it is generally available.
If You Have Standardized on One Model, You Are Leaving Value on the Table
Some teams have already reacted to cost by choosing a single default model, or by negotiating a good rate and standardizing on it. That is a reasonable first step, and this is not a reason to undo it. Where this still falls short is granularity: a single default cannot match the shape of the work, because complexity varies from step to step inside the very same workflow. One agent might use the default for triage, for a planning call, for a code edit, and for a verification check. Four steps, four different levels of required intelligence. Routing optimizes at that finer grain, which is where most of the savings actually live.
The other common hesitation is the worry about handing model choice to a black box. It is a fair concern, and the design answers it directly. The router never reaches for a model you have not approved, it respects your data-residency rules, it logs every decision, and you set the tradeoffs it optimizes for, whether that is cost, latency, or quality for a given workload. You are not giving up control; You are giving a repeatable set of rules something to enforce.
Turn Cost Control into a Discipline, and Where KPI Partners Enter
A router is a capability. Getting durable value from it is an operating practice. Some teams already have a name for this kind of work: Financial Operations, or FinOps, a cross-industry discipline for governing technology spend that began in cloud and is now extending to AI. The label matters less than the habit. The habit is deciding which models are approved and for what, defining the quality bar each workload has to clear, governing spend with visibility and limits by team and by agent, and reviewing the whole thing on a cadence as models and prices keep changing.
That practice is where my team focuses. As a Snowflake partner, KPI Partners helps organizations design the approved-model policy, set the quality bars that protect outcomes, put spend governance and review loops in place, and connect all of it to the way the business actually measures value. We approach it business first, starting from the workflows that cost the most and matter the most, and translating an infrastructure capability into a model a CFO and a CIO can both stand behind.
This is how we would deliver it:
- Identify the workflows where model spend is concentrated, and the steps inside them whose complexity varies the most.
- Define model policies and quality thresholds for each class of work, aligned to the outcomes the business actually measures.
- Enable routing within approved governance boundaries, approved-model list, data-residency rules, audit logging, and spend limits by team and by agent.
- Continuously review cost, quality, and routing behavior as models, prices, and workloads evolve, so the rules stay current instead of going stale.
The aim is efficiency you can trust, not a one-time cleanup that quietly erodes.
The Advantage Goes to the Disciplined, Not the Biggest Spender
If there is one stance I would leave you with, it is this. As AI budgets grow and every model keeps improving, the winners will not be the organizations that spend the most, or the ones that chase the single best model. They will be the ones with the discipline to match each task to the right level of intelligence, govern that choice, and keep refining it as the landscape shifts. Dynamic model routing gives that discipline a natural home inside Snowflake. The companies that build the practice now, while the capability is still early, are the ones that will lower the cost of every AI outcome without ever slowing down.
Related reading
- From Analytics to Action: Building Enterprise-Ready Data Agents on Snowflake
- Getting Your Data AI-Ready: Turning Snowflake into a Trusted AI Data Foundation
- Meet Coco: Snowflake's Native AI Coding Agent for Data
Ready to Transform Your Data Strategy? Talk to our experts and discover how KPI Partners can accelerate your data and analytics initiatives.
Talk to Our Experts | Book a Demo