Motivation — The Modern Productivity Paradox: LLMflation vs AI bill inflation
The token, that is the discrete representation of text characters processed by an LLM, is considered as the new economic unit of Artificial Intelligence. Its dual role is unusual: the same object that determines how a model parses and generates information also determines how users are charged and how organizations reason about capacity, utilization and ROI [1].
Over the past five years, the economics of foundation model inference have followed an aggressive deflationary trajectory of the order of 99.7%, for instance on March 2023 GPT-4: 30$/1M tokens in contrast to 0.1$/1M tokens for frontier open-weight ones by mid-2026. This nominal pricing for LLMs collapsed primarily due to silicon advancements, architectural innovations (Mixture-of-Experts, speculative decoding) and model distillation. This collapse fueled an assumption that machine intelligence was rapidly converging to a zero-marginal-cost commodity.
Yet, as enterprises transitioned from isolated proof-of-concept experiments to production deployments, a trend emerged: aggregate AI expenditures surged. Instead of observing deflationary operating bills, enterprises usually face double-digit year-over-year spend increases such as the case of Uber that exhausted the annual budget by April 2026 and Microsoft that cancelled Claude code licences when token bills spiraled out of control ahead of schedule.
The root cause for this apparent paradox of collapsing pricing and spiking enterprise invoices is not an accounting anomaly. It reflects a shift in how software engineering and deployment consumes machine intelligence. Utilization volume and task complexity surge, shifting consumption from small exploratory queries (such as the static interactive prompts from the early chatGPT days) to continuous background workloads such as automated reasoning chains (such as the ‘thinking’ part while waiting for a chatbot response) or autonomous pipelines. Consumption exploded precisely because unit prices have fallen (Jevon’s paradox) unmasking an underlying economic reality that enterprises must address to remain profitable and competitive.
Traditional enterprise software is treated either as: i) capital expenditure or ii) as operational spend: per-seat SaaS subscriptions. In both regimes the cost remains detached from the ‘cognitive throughput it delivers’. In a paper [2] by Microsoft researchers, the authors argue that advanced AI tools cannot be evaluated as passive digital tools as they represent ‘digital labor’ associating autonomous cognitive capabilities that function as independent factors of production alongside human labor. In other words, tokens are not mere bandwidth or storage bytes but a meter of digital labor. Because this metered cognition is embedded across multiple layers of enterprise architecture (from customer-facing applications and internal classification routines to autonomous agentic loops) every single LLM invocation represents an implicit expenditure decision. When enterprises fail to govern this expenditure dynamically, treating every transaction as an unmetered call to the most capable model, the result is runaway operational spend that undermines the very efficiencies generative models promised in the first place.
The solution: allocating smarter or more economical models based on the background task (or workflow step) complexity, thus optimizing cost instead of always asking for frontier model reasoning at one or two orders of magnitude above what is required.
Problem Definition — Cost-Aware Request Routing & Agentic Amplification
To understand why enterprise token consumption defies linear projections, one must examine the architectural evolution of how LLMs are invoked. In traditional human-in-the-loop deployments (e.g., conversational chatbots, inline code completions), token generation is strictly bounded by human pacing. A user submits a query, reviews the output, and formulates a follow-up. In this regime, latency is measured in seconds, sessions terminate when a human pauses, and token consumption scales linearly with user activity. Cost variance is naturally capped by the throughput of the human operator.
With the introduction of autonomous (multi-step) agentic workflows, inference consumption takes a steep increase. Modern architectures operate largely decoupled from human latency: examples include ReAct, plan-and-solve decompositions and automated reflection loops. A high-level objective such as “reconcile cross-system data discrepancies and generate an audit summary” expands to a recursive loop of many model invocations. In addition, agentic execution introduces context accumulation on every downstream invocation carrying previous responses that grows monotonically. In effect, what once was a series of discrete/bounded queries is now a compound workload where token volume scales super-linearly across a task. Even worse, multiple tool (code function call) retries, dead-end search paths, chain-of-thought reasoning, memory management (summarization of context before storing to short term memory and transformation of the latter to long term memory [7]) can cause token consumption to surge, requiring the need for budget allocation per business task.
Simultaneously, the supply side of foundation models has fragmented into specialized capability tiers, ranging from distilled, high-throughput utility ones (often termed -flash or -mini) to massive frontier ones with high test-time reasoning (e.g. Fable 5). Pricing is very differentiated across tiers as well due to performance vs cost trade-off and empirical benchmarks across foundation models demonstrate that it spans from 15x to 100x [6], due to extensive parameter scale, high-memory bandwidth and energy consumption [1]. We suggest that interested reader should consult [1] for a very rigorous analysis on tokenomics, based on model complexity, resource allocation and economic value creation.
For the scope of this article, we assume that the total inference expenditure into to variables:
Total cost = Tokens Consumed x price per Token
(In this case we assume that the input and output token prices are the same, though in general output tokens are more expensive and input token caching is an additional cost). Cost-aware dynamic model routing operates on the unit price multiplier. It would be tempting for developers to route all application traffic to a pinned frontier reasoning model in order to insure against errors or hallucinations. This would lead to burning capital on ‘surplus’ intelligence. On the contrary, naive under-provisioning routing complex tasks involving multi-step reasoning to a budget model induces multi-turn repair cascades due to limited context, hallucinations etc. So the engineering challenge is to decide how to allocate cognition at the request-level and route traffic to the minimum-cost model capable of satisfying the quality constraints. This is the decision layer that evaluates incoming task complexity at real time assisted by tracking telemetry of past decisions (cost, latency, retries).
Related work on Cost-Aware LLM Routing
As already described, an LLM router balances performance and cost. Early approaches such as FrugalGPT [6] and AutoMix [8] employ a cascade method which sequentially queries different LLMs until a reliable response is obtained, though this strategy usually needs to query multiple times and leads to high latency. Other cost-aware approaches are: RouteLLM, HybridLLM where queries are directly routed to the most suitable LLM (predictive routing).
MetaRouter [5] adapts on individual user cost-performance preference feedback by training a policy neural network based on the encoded feedback during an adaptation phase.
CoDyn [4] is a code specific LLM router exploring how classifier-based routers specifically evaluate software engineering workloads, proving that a tuned, low-cost router model can match frontier coding capabilities while securing over 43% in code generation savings.
While tokenomics clarifies why cost-aware allocation is necessary, operationalizing it requires routing algorithms that respect commercial performance commitments. Recent breakthroughs address two core real-world challenges: adapting to sparse feedback and enforcing strict accuracy targets (eg 95% accuracy threshold on critical queries for premium frontier models in contrast to 75% for the corresponding economy tier model). The latter is also referred to as Service Level Agreement (SLA) and it applies on latency besides answer relevance/quality as well.
PROTEUS [9] is a multi-LLM router that accepts accuracy targets τ as runtime input, in contrast to conventional routers requiring offline hyper-parameter tuning and guessing τ through trial and error, and uses Lagrangian Reinforcement Learning. It employs a learned dual variable λ that tracks constraint violations during training (once) to condition the underlying policy network hence a single trained model can serve the entire accuracy spectrum (τ in [75%, 95%]) dynamically.
In production systems, where sparse, one-sided feedback is available (meaning that for the training set samples for which a request is dispatched to a given model, the system observes user satisfaction only for that chosen model; counterfactual outcomes for unselected models remain unobserved), SLARouter [3] provides theoretical guarantees for meeting SLA requirements under this relaxed requirement on training data. In addition, SLARouter allows online adaptation as new data arrives so that it can account for distributional drifts, by updating routing boundaries from production data without retraining. Both approaches involve training a small neural network.
CARROT [11] introduces rate-optimal guarantees but does not adapt online (during inference). Relies on full-feedback datasets and assumes access to data where each query has been evaluated by all available LLMs, that is not always available in real data scenarios.
Various evaluation benchmarks have been designed to assess the efficiency and accuracy of LLM routers. For instance, RouterBench [10] comprises approximately 405K inference outcomes and includes responses and metrics from 11 distinct LLMs, combining open-source options (like Llama-70B-chat, Mixtral-8x7B-chat, and Mistral-7B) and proprietary systems (like GPT-4 and Claude) as an effort towards a standardized benchmark.
SPROUT [11] complements RouterBench with 45K queries across 14 models. Together, these benchmarks test complementary aspects: RouterBench provides scale and task diversity (reasoning, factual recall, dialogue, mathematics) with established models, while SPROUT tests generalization to modern model pools with extreme cost variation.
The progression of request routing reflects a steady shift from rigid developer intuition toward adaptive, SLA-enforceable gateway orchestration. Table 1 summarizes the approaches mentioned so far. This evolutionary trajectory reveals a clear design consensus: high-performance enterprise routing cannot rely solely on static heuristics or opaque third-party black boxes.
Instead, bridging research and production demands a layered gateway architecture capable of executing sub-millisecond filtering on straightforward traffic while enforcing strict quality and SLA constraints on complex cognitive tasks.
| Routing Paradigm | Core Decision Mechanism | Routing Overhead | Adaptivity | Quality & SLA Safeguards |
| Static / Manual | Hardcoded model endpoints, prompt prefixes | 0 ms | None (Rigid) | Unbounded failure risk on complex queries |
| Semantic Routers | Vector embedding distance against exemplars | 50–200 ms(local or hosted encoder, encoder size dependent) | Static reference sets | Heuristic similarity thresholds without performance bounds |
| Cloud Aggregator Routing (e.g., OpenRouter openrouter/auto) | Managed proxy heuristics, provider load balancing, and price-tier fallback | ~25–50 ms (additional proxy latency, Cloudflare edge [12]) + peak hour latency + credit balance checks | Dynamic provider availability & load | Best-effort provider failover; no formal accuracy floor guarantees |
| Predictive Classifiers (e.g., FrugalGPT, RouteLLM) | Supervised scoring / fast LLM classification rubric | >1000 ms (Encoder size dependent) | Offline batch retrained | Confidence scoring; requires per-dataset parameter tuning |
| Online SLA Routing (SLARouter) | Dual-primal contextual bandit optimization | ModernBert Encoder + <1 ms for the MLP classifier | Online (sparse one-sided feedback) | Provable probabilistic SLA satisfaction guarantees |
| Lagrangian RL Routing (PROTEUS) | Policy network conditioned on learned dual variable λ (for online RL training) | 2.6-8.7ms (based on batching on an A100) | Direct runtime input of accuracy target τ | Strict floor compliance (Accuracy≥τ) across variable targets |
| Hierarchical Hybrid Gateway (LiteLLM Auto Routing) | Cascaded Heuristics → Keywords → Fast LLM → Adaptive Pool | <1 ms (fast path) to ~50 ms | Thompson Sampling & health metrics | Multi-tier failover, fallback chains, and session affinity rails |
LiteLLM – OpenSource, self-hosted, versatile Dynamic Router
Having examined the microeconomic foundations of tokenomics and the algorithmic guarantees of SLA-constrained routing, we now outline the architectural blueprint for operationalizing these concepts. Rather than prescribing a fixed code recipe, this section articulates the architectural intent of a dynamic, cost-aware routing layer built upon the open-source LiteLLM Auto Routing framework (Figure 1). LiteLLM is the open source alternative to OpenRouter [12].

Figure 1: LiteLLM Gateway Proxy Architecture
The first architectural decision is decoupling application logic from routing intelligence. In standard enterprise codebases, specific models are often hardcoded into agent prompts. This tight coupling complicates failover and prevents cost management. LiteLLM is a central (user managed) AI Gateway providing telemetry and spend attribution to agentic applications by evaluating incoming request complexity and steering execution to the optimal model tier (that is a user-defined pool of comparable models such as ‘cheap’, ‘reasoning’, ‘coding’ etc.).
The decision (routing) on incoming requests comes with a classification tax: either spending hundreds of milliseconds and tokens on an external LLM to decide or implement a hierarchical classification cascade to resolve requests via low-latency, low-cost mechanisms:
Stage 1: Heuristic scorer (Fast)
The gateway inspects structural features of the latest prompt (token length, code syntax tokens, multiple question marks etc) so that zero API calls are made and queries end up immediately to the corresponding model tier in sub-millisecond speed.
Stage 2: Deterministic & semantic keyword rules (Fast)
Words in a prompt are matched or approximately matched (fast semantic vector search) to the corresponding tier, for example “hi” → low-cost model, “kubernetes” → reasoning model. In this case the developer can pin a specific model to an agentic tool call.
Stage 3: (small) LLM classification
A lightweight, dedicated classifier model such as claude-Haiku or a finetuned and deployed LLM can handle ambiguous requests in order to evaluate their complexity against a structured rubric. Latency can be controlled by a strict timeout so that routing falls under one of the other stages instead.
Stage 4: Custom classifier plugin
Custom plugins evaluate organizational metadata such as caller tenant tier (free vs. enterprise SLA), remaining departmental token budgets, or time-of-day cost constraints thus overriding or refining the assigned capability tier before dispatch. This is not about the prompt complexity but rather about a python based instruction allowing more controlled decisions.
A subtle yet severe failure mode in agentic and multi-turn routing is cache thrashing. Modern model providers offer significant cost discounts (typically 50% to 90%) and latency reductions for input tokens that hit the provider’s Key-Value (KV) prompt cache. If a router evaluates every conversational turn in isolation, Turn 1 might route to Provider A, Turn 2 to Provider B, and Turn 3 back to Provider A. This naive per-turn oscillation invalidates the KV cache on every exchange, forcing full context re-computation, driving up time-to-first-token (TTFT), and inflating aggregate input token costs. LiteLLM incorporates session affinity to track past conversations and pass subsequent turns to the same provider to maximize KV cache hits, when necessary.
LiteLLM pros:
- LiteLLM is an opensource (SOC-2 type 2, ISO 27001 certified), proxy server that runs as a container with a db (production features: per key budget limits, cost tracking). Useful for regulated industries in contrast to OpenRouter (3rd party cloud service) where it stores metadata and delegates privacy (training/retention) discretion to the corresponding model provider it internally routes data to. Zero data retention (ZDR) using OpenRouter can be enforced [13] though it might defeat the purpose of (not) choosing a smaller/cheaper model especially considering that this can be achieved with a self-hosted open weight model accessible by LiteLLM.
- Provides Thompson sampling within models in each tier in order to learn latency distributions between providers and dynamically adapt.
- Although LiteLLM is another stateful service in the stack in contrast to the fully managed OpenRouter [12] and other variants, we are in control of this important component instead of relying on potential downtimes of a fully managed one.
- With respect to latency, LiteLLM adds <1ms (correctly-sized container orchestration, fast caching, logging etc) in contrast to OpenRouter ~25-40ms through its edge network.
Conclusion & Next Steps — From Pinned Gateways to Dynamic Allocation
As foundation models evolve from isolated conversational tools into autonomous digital labor, treating machine cognition as an unmetered, monolithic utility is no longer economically sustainable. While macro-level token prices will continue to trend downward, the superlinear token demand driven by multi-step agentic workflows and the growing cost spread between utility and frontier models make dynamic request routing the single most decisive lever for enterprise cost governance.
The academic literature—from the structural tokenomics of Zhu (2026) to the SLA-constrained algorithms of SLARouter and PROTEUS—proves that intelligent routing can cut inference expenditure by 50% to nearly 90% without compromising task satisfaction. Yet, the ultimate test of these principles lies in their production adoption.
In our own infrastructure, we have already established the necessary architectural foundation by deploying LiteLLM as a centralized proxy gateway to manage traffic control and cost monitoring across our MultiAgent community [14]. Today, this gateway successfully aggregates credentials, enforces rate limits, and provides unified spend visibility across our active agents. However, our current deployment remains fundamentally predetermined: models are statically pinned to specific agent roles or pipeline tasks. While this ensures predictable execution, it leaves substantial economic efficiency on the table. High-frequency intermediate steps (such as agent handoffs, structured argument extraction, and routine status checks) routinely consume expensive frontier tokens simply because an agent’s assigned model was chosen for its peak cognitive capability rather than its median task requirement.
Our immediate roadmap focuses on evolving this infrastructure from static model pinning to dynamic, cost-aware cognitive allocation. By transitioning our existing LiteLLM gateway into an active routing intelligence plane, we aim to operationalize the insights of tokenomics research—scaling our multi-agent community with rigorous cost efficiency while preserving frontier reasoning exactly where it matters most.
Bibliography
- [1] (2026). AI Tokenomics: The Economics of Tokens, Computation, and Pricing in Foundation Models. https://arxiv.org/abs/2606.24616
- [2] (2025). Evolving the Productivity Equation: Should Digital Labor Be Considered a New Factor of Production? https://arxiv.org/abs/2505.09408
- [3] (2026). Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees. https://arxiv.org/abs/2606.19376
- [4] (2025). CoDyn: Dynamic LLM Routing for Coding Tasks. NeurIPS 2025 Fourth Workshop on Deep Learning for Code. https://openreview.net/forum?id=0ox03jE6jb
- [5] (2026). Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning. https://arxiv.org/abs/2606.06178
- [6] (2023). FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. https://arxiv.org/abs/2305.05176
- [7] Medium-Tokenomics: FinOps for the Agentic Harness. https://aiadvances.org/tokenomics-for-the-agentic-harness-41dbaa822f7f
- [8] (2025). AutoMix: Automatically Mixing Language Models. https://arxiv.org/abs/2310.12963
- [9] (2026). PROTEUS: SLA-Aware Routing via Lagrangian RL for Multi-LLM Serving Systems. https://arxiv.org/abs/2601.19402
- [10] (2024). RouterBench: A Benchmark for Multi-LLM Routing System. https://arxiv.org/abs/2403.12031
- [11] (2025). CARROT: A Cost Aware Rate Optimal Router. https://arxiv.org/abs/2502.03261
- [12] (2026). OpenRouter vs LiteLLM: Which LLM Gateway Fits Your Stack? https://openrouter.ai/blog/insights/openrouter-vs-litellm/
- [13] (2026). Data Collection. https://openrouter.ai/docs/guides/privacy/data-collection
- [14] A research agenda for agent expert communities. https://research.wpp.com/blog/a-research-agenda-for-expert-agent-communities