Silvia Labs
Adaptive Model Routing for Agentic Workloads: Knowing the Work Before the Agent Does It
Every chat turn in our product currently runs on the most capable and most expensive frontier model available. That is a safe default and a poor economic one. We show that agent effort (how many times an agent iterates its loop) is predictable before the turn runs, and that a sub-millisecond classifier can route the shallow tail to a smaller, cheaper model while the frontier handles the deep work.
Jacob Weiss, Head of AIAugust 20260.34 ms p50 inference358 req/s sustained
0.34 ms
Median classifier inference. Four to five orders of magnitude below the frontier model call it gates.
358 req/s
Sustained throughput under a live load test against the deployed endpoint, with zero errors at 64× concurrency.
≈ 29%
Upper bound on token-cost reduction at a 5× per-token price ratio between the small and frontier models.
Abstract
Production chat assistants are increasingly powered by agents: a language model wrapped in a harness of tools and a control loop that the model iterates until the task is done. An instant factual lookup and a multi-minute deep-research report are produced by the same model and harness, and differ only in how many times the agent loops.
Tracing the production turns of a frontier financial assistant, we find that the work is strongly bimodal in loop depth, and that this depth is predictable from the question text before the agent runs a single step. We formalize an agent as a model–harness pair, autonomous work as the agent composed with itself L times, and routing as choosing the cheapest model expected to clear a quality bar at the predicted depth. A sub-millisecond effort classifier (a linear softmax over hashed n-grams) then maps each turn to a model.
Short-routing precision reaches 0.74 over half of all traffic and ~0.90 on its most confident slice, and the confidence is well calibrated, so the threshold can be read directly as a precision dial. Inference costs 0.34 ms at the median, four to five orders of magnitude below the model call it gates. Because cost scales with loop depth and the routed turns are the shallow ones, the token-cost reduction is bounded above by ~29% at a 5× price ratio. The question-only signal saturates near 0.72 ROC-AUC; adding leakage-safe conversation-state features lifts short-detector ROC-AUC by ~0.06 and accuracy by 3–4 points under a strict user-grouped split. Trained on the frontier agent's own behavior, the router tracks the frontier as models and harnesses evolve rather than fighting it.
The framework, in three definitions
Agent
A model and a harness: the tools, the control loop, and the context management that wrap it. One step of the agent turns the current state into an action and a successor state.
Work
The agent composed with itself L times. A one-shot answer, a few-step lookup, and a deep-research report differ only in L, not in the model, not in the harness.
Routing
Choosing the cheapest model expected to clear the quality bar at the predicted depth. Since L is unknown before the turn runs, a learned classifier predicts it from the question, in microseconds.
Both wall-clock time and token cost scale approximately linearly with L, at roughly 18–20 seconds per loop under the frontier model. The two quantities a router controls are the model's per-loop speed and price.
Headline findings
Agent effort is bimodal, and predictable from the question alone
The same model and harness produce a one-line answer and a deep-research report. The only thing that differs is how many times the agent loops. Tracing production turns of a frontier financial assistant, we find that loop depth is strongly bimodal and can be predicted before the agent runs a single step. A linear softmax over hashed n-grams of the question text separates short from deep work well enough to route.
Confidence is well calibrated, and the threshold reads as a precision dial
Raw accuracy is modest (0.68 over a 0.58 base rate), but a router only needs to be right on the turns it acts on. Short-routing precision reaches 0.74 at 50% coverage and ~0.90 on the most confident slice. The predicted probability tracks the observed short rate almost diagonally, so the confidence threshold τ can be turned directly into a precision knob without post-hoc calibration.
The text-only signal saturates, but conversation state breaks the ceiling
Hyperparameter sweeps, semantic embeddings, and different label cuts all plateau near 0.72 ROC-AUC. The residual signal is not in the question. Adding leakage-safe session and user "effort-momentum" features (prior-turn loop counts, running mean, session opener flag) lifts short-detector ROC-AUC by +0.06 and accuracy by 3–4 points under a user-grouped split, rescuing terse follow-ups like "yes, run it" that carry no lexical signal.
Routing is effectively free relative to the model call it gates
Inference is a sparse dot product costing 0.34 ms at the median, four to five orders of magnitude below the ~18 s frontier model call. Under load, the deployed endpoint sustains 358 req/s with zero errors, and even at 64× concurrency the tail stays sub-second. Compute is never the bottleneck; the network is, and the classifier itself adds no measurable serving cost.
Read the full write-up
The framework, the deployed model card, calibration and error analysis, the effort-momentum extension, latency measurements, and the operating curve. Eighteen pages.
Data and materials
The training data derives from private customer conversations traced in production and is not released; neither are the model artifact or the internal experiment code. This paper shares the framework, the measurement methodology, and the findings so that others can apply the same approach to their own traces. Absolute traffic volumes are withheld and reported as shares and rates instead.
How to cite
Jacob Weiss. Adaptive Model Routing for Agentic Workloads: Knowing the Work Before the Agent Does It. Silvia, Inc. / Silvia Labs, August 2026. https://www.cfosilvia.com/ai-lab/adaptive-model-routing
Related
From Silvia Labs
A New Frontier in AI Tax Intelligence
Does retrieval tooling make AI tax guidance more accurate, and easier to check? 200 expert-level questions, 1,511 answers, 2,949 blind judgments, run twice.