#Abstract
Static capability benchmarks poorly predict how well a given large language model (LLM) will perform on a specific task, yet most routing systems assign tasks to models using exactly such static signals. We propose operationalizing "intelligent routing" as a low-cost interactive probe: before committing a task, the router sends cheap diagnostic queries to candidate models, observes their responses, and updates a per-model fitness estimate, then routes according to a decision matrix combining predicted quality, latency, and monetary cost. We formalize the probe as an expected-value-of-information (EVI) problem nested inside a routing decision, derive the posterior update for binary probe signals, and give a fully worked numerical example in which probe-informed routing raises expected utility from $0.5700$ to $0.6782$ (a net gain of $0.1082$ after probe costs) at a probe cost of $\$0.003$ per task, with a break-even per-probe cost of $c_p^{\ast} \approx \$0.00821$. A complementary precision analysis shows a routing decision is statistically meaningful only if per-candidate probe precision satisfies $\sigma_f \lesssim 0.0379$. We position the framework against the routing, delegation-cue, and efficiency literature and state explicit falsification conditions.
#1. Introduction
Organizations that deploy LLMs face a routing problem: for each incoming task, choose among several candidate models that differ in predicted output quality, latency, and per-token cost. The dominant practice is to route by static signals — leaderboard rankings, benchmark aggregates, or hand-tuned heuristics. The premise of this paper is that this practice is structurally mismatched to the problem, because per-task fitness (the probability that a specific model produces an acceptable answer to a specific task) is exactly the quantity that static aggregates fail to estimate, while it is precisely the quantity a cheap interactive measurement could estimate directly.
The research idea examined here is: before committing a task, send cheap diagnostic queries ("probes") to candidate models; use the responses to estimate each model's fitness for the task; then route based on a decision matrix combining predicted quality, latency, and cost. The idea matters because routing heuristics are rarely grounded in measured fitness, and because probing converts an inference-time guess into a measurement problem with a well-defined price.
The stakes are practical. Efficiency-focused work on LLM deployment emphasizes that serving costs and latency constraints are real and structural — for example, KV-cache memory overhead that "inflates infrastructure costs and throttles scalability" in high-concurrency settings [7] — so the difference between routing to an adequate cheap model and an over-provisioned expensive one compounds at fleet scale. Conversely, audits of the frontier model landscape show that aggregate benchmark rank can be misleading for specific use classes: the QNFO LiveBench-grounded audit reports that no Llama model ranks in the top 42 on its mathematics-heavy suite and identifies DeepSeek V4 Pro 0813 as the mathematics price-performance option [9] — a finding that illustrates why per-task fitness, not family reputation, should drive routing.
This paper makes three contributions. First, it formalizes probe-based routing as a two-stage decision problem: a probe-design stage (which queries to send, at what cost) and a routing stage (which model to commit, given probe responses), and shows the probe decision is an EVI calculation. Second, it derives, with full arithmetic, a worked example demonstrating when probing beats static routing and when it barely pays, including the break-even probe cost. Third, it specifies the conditions under which the claim "probe-informed routing beats static rubric routing" can be empirically tested and falsified.
We emphasize scope: this paper contains no new empirical measurements. All quantitative results are derived from explicitly stated illustrative assumptions in Section 4; the empirical program is specified, not executed.
#2. Background and Related Work
We review the supplied bibliography, which spans LLM task automation, network routing, adaptive agent interaction, mixture-of-experts routing, inference-time interaction, deployment efficiency, delegation cues, and benchmark auditing. Two entries ([11], [12]) have empty supplied summaries; we note this explicitly and attribute no claims to them.
[1] AutoDroid: LLM-powered Task Automation in Android. This work applies LLMs to mobile task automation, motivated by the poor scalability of prior approaches, which suffered from limited language understanding ability and non-trivial manual effort from developers or end-users. It is relevant here as a deployed setting in which an LLM must be selected and trusted to execute multi-step tasks — exactly the per-task reliability problem our probe addresses — although the supplied summary is truncated and gives no detail on its evaluation or any routing mechanism, so we cite it only as problem context.
[2] Trusted Routing for Blockchain-Enabled Low-Altitude Intelligent Networks. This paper addresses routing in low-altitude intelligent networks, where UAVs' distributed topology, high dynamic mobility, and vulnerability to security threats degrade routing performance for data transmission. Its relevance is structural rather than topical: routing under unreliable, dynamically changing node quality is the same abstract problem as routing tasks under uncertain model fitness, and "trusted routing" there plays the role our fitness estimation plays here. The supplied summary is truncated before its solution details, so we draw only this structural parallel.
[3] Exploring the Role of Common Model of Cognition in Designing Adaptive Coaching Interactions for Health Behavior Change. This work builds human-aware collaborative agents that model, learn, and reason about a human partner's physiological, cognitive, and affective states, applied to adaptive coaching for health behavior change via the common model of cognition framework. It supplies the closest conceptual ancestor of our probe: just as a coaching agent adapts interaction to a modeled partner state before acting, our router adapts model selection to a measured model state. The summary is truncated and does not report results.
[4] Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers. This work introduces routing signatures — vector representations summarizing expert activation patterns across layers for a given prompt — motivated by the observation that MoE routing mechanisms responsible for expert selection remain poorly understood. This is the closest internal-routing analogue to our proposal: it seeks task-dependent structure in how computation is allocated inside a model, whereas our probe observes input-output behavior of opaque models at the API level. The shared thesis — that routing behavior is task-conditioned and should be measured, not assumed — is the empirical premise our probe operationalizes at the model-selection level. The summary is truncated before its findings.
[5] Interactive Learning for LLM Reasoning. This paper notes that existing multi-agent learning approaches build interactive training environments promoting collaboration among multiple LLMs to construct stronger multi-agent systems, but that at inference they require re-executing the multi-agent system to obtain final solutions, which diverges from human cognition, in which individuals can enhance reasoning without re-execution. This motivates our cost concern: re-execution is expensive, so any inference-time interaction (such as our probes) must be cheap enough to preserve the economics that motivated routing in the first place. Our probe is deliberately non-re-executive: a lightweight pre-commitment measurement, not a full multi-agent execution.
[6] Towards Intelligent Interactive Theatre: Drama Management as a way of Handling Performance. This work presents a new modality for intelligent interactive narratives in theatre, with an intelligent agent serving as both drama manager and actor, and poses research challenges arising from that analysis. We cite it as an example of interactive decision-making under live performance constraints — a system that must commit to actions in real time from a portfolio of options, which is the control pattern our router instantiates with models as the portfolio. The supplied summary gives no further technical detail connecting it to routing.
[7] YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition. This work identifies KV cache memory overhead as the bottleneck that inflates infrastructure costs and throttles scalability in high-concurrency LLM deployment, and proposes a structural transition and training pipeline for an efficient financial LLM. It grounds the cost side of our decision matrix: serving-cost differences across models are real and structural, so a routing objective that ignores cost is optimizing half the problem. The summary is truncated before quantitative results.
[8] Task-Aware Delegation Cues for LLM Agents. This work observes that human–agent teamwork remains brittle due to information asymmetry: users lack task-specific reliability cues and agents rarely surface calibrated uncertainty or rationale. It proposes a task-aware collaboration signaling layer that turns offline preference evaluations into online, user-facing primitives for delegation. This is conceptually the nearest neighbor of our probe: both convert offline evaluation evidence into online, task-specific signals. Our contribution differs in the consumer of the signal — an automated router rather than a human user — and in treating signal acquisition as a priced decision.
[9] QNFO: Prioritizing Large Language Models for Scientific Research and Agentic AI: A LiveBench-Grounded Audit (August 2026). This audit of the frontier LLM landscape for mathematics-heavy scientific research and agentic AI is grounded in the LiveBench 2026-06-25 contamination-free benchmark, and reports that no Llama model ranks in the top 42 and identifies DeepSeek V4 Pro 0813 as the mathematics price-performance optimum (the summary truncates mid-sentence). It exemplifies the static-benchmark paradigm our probe is designed to supplement: even a careful contamination-controlled audit yields aggregate rankings, not per-task fitness. Our framework treats such benchmark results as priors over model fitness, to be refined by probes rather than replaced.
[10] QNFO: Joules-per-Solution for Stochastic and Agentic Inference. This work extends the joules-per-solution (J/S) metric — introduced as a universal, physics-grounded measure of computational efficiency — to LLMs, noting that the original metric assumed deterministic solvers (one run yields one solution) and that LLMs violate that assumption as stochastic samplers whose outputs (the summary truncates here). This matters directly for our fitness definition: because LLM outputs are stochastic, per-task fitness must be defined as a probability of acceptance, not a deterministic outcome, and efficiency metrics must be expectation-based.
[11] QNFO: Comprehensive Technical Framework for Network Isomorphism. The supplied entry contains no summary text; we note only its existence and do not attribute any claim to it.
[12] QNFO: Syntactic Generation. The supplied entry likewise contains no summary text; we do not attribute any claim to it.
In summary, the literature supplies (i) deployed LLM task-execution settings [1], (ii) routing under unreliable nodes [2], (iii) task-conditioned structure in internal routing [4], (iv) online delegation signals from offline evaluations [8], and (v) cost- and stochasticity-aware efficiency framing [7], [10], [9] — but no work in the supplied set formalizes pre-commitment diagnostic probing of candidate LLMs as a priced EVI problem. That is the gap this paper addresses.
#3. Methods
#3.1 Setting
A router receives a task $\tau$ and chooses among $N$ candidate models $M_1, \dots, M_N$. Each model $M_i$ has a known per-task monetary cost $c_i$ (dollars), a known latency $\ell_i$ (seconds), and an unknown per-task fitness $q_i \in [0,1]$, the probability that $M_i$'s output for $\tau$ is acceptable. Following the stochastic-sampler framing of [10], fitness is a probability, not a deterministic property.
#3.2 Decision matrix
The router maximizes expected scalarized utility
where $\lambda$ (utility per dollar) and $\mu$ (utility per second) are fixed trade-off weights. Static rubric routing uses only prior evidence $\mathcal{E}_0$ (benchmarks, rubric scores); probe-informed routing augments it with probe responses.
#3.3 Probe model
A probe is a cheap diagnostic query sent to model $M_i$ before commitment, costing $c_p$ dollars and $\ell_p$ seconds. We model the binary case: the probe returns a signal $s_i \in \{0,1\}$ (fail/pass). The probe has sensitivity $a = P(s_i = 1 \mid M_i \text{ fit})$ and false-positive rate $b = P(s_i = 1 \mid M_i \text{ not fit})$, with $a \gt b$. The router holds a prior $\pi_i = P(M_i \text{ fit})$. By Bayes' rule,
We take fitness probability as the quality term, i.e., $\mathbb{E}[q_i \mid s_i] = \pi_i^{\pm}$, so that a fit model yields an acceptable solution with probability $1$ and an unfit model with probability $0$; this is a conservative two-point approximation to a continuous fitness distribution.
#3.4 Probe decision as EVI
Let $U^{\text{static}} = \max_i U_i$ under priors, and let $U^{\text{probe}}$ be the expected utility of the two-stage policy: probe all $N$ candidates (in parallel, so probe latency is $\ell_p$ once), then commit optimally. The probe is worth sending iff
The break-even per-probe cost is $c_p^{\ast} = (U^{\text{probe,gross}} - \mu \ell_p - U^{\text{static}}) / (\lambda N)$, where $U^{\text{probe,gross}}$ excludes probe cost.
#3.5 Precision requirement for a meaningful decision
For a routing decision to be statistically meaningful, the utility gap between the top two candidates must exceed the decision noise. With independent fitness standard errors $\sigma_{f,1}, \sigma_{f,2}$ for the top two candidates, the standard error of the utility difference is $\sigma_\Delta = \sqrt{\sigma_{f,1}^2 + \sigma_{f,2}^2}$, and the decision is significant at confidence level $z$ (e.g., $z = 1.96$ for two-sided 95%) iff $|\hat{U}_1 - \hat{U}_2| \gt z\,\sigma_\Delta$. This yields a minimum-probe-precision requirement $\sigma_f \le |\hat{U}_1 - \hat{U}_2| / (z\sqrt{2})$.
#3.6 Extension to graded probes
Binary probes generalize to graded scores $s_i \in \{1, \dots, K\}$ with likelihoods $P(s_i \mid \text{fit})$, $P(s_i \mid \text{not fit})$; the posterior update is the same Bayes rule over $K$ outcomes. Probe families may include plan-quality sketches and self-evaluation accuracy checks; the formalism is agnostic to probe content, requiring only a likelihood model.
#4. Analysis
All numbers below are derived from explicitly stated illustrative assumptions; none are empirical measurements.
Inputs (assumed for the worked scenario). Three candidate models with priors, costs, and latencies:
| Model | Prior $\pi_i$ | Cost $c_i$ ($) | Latency $\ell_i$ (s) |
|---|---|---|---|
| $M_1$ | $0.70$ | $0.020$ | $3$ |
| $M_2$ | $0.50$ | $0.005$ | $1$ |
| $M_3$ | $0.80$ | $0.080$ | $8$ |
Trade-off weights: $\lambda = 5$ utility per dollar, $\mu = 0.01$ utility per second. Probe parameters: $a = 0.9$, $b = 0.3$, $c_p = \$0.001$, $\ell_p = 0.5$ s, probes sent in parallel to all $N = 3$ models, signals conditionally independent across models.
Step 1: Static routing.
Static routing selects $M_1$: $U^{\text{static}} = 0.57$.
Step 2: Posterior updates. Using $\pi_i^{+} = \frac{0.9\,\pi_i}{0.9\,\pi_i + 0.3(1-\pi_i)}$ and $\pi_i^{-} = \frac{0.1\,\pi_i}{0.1\,\pi_i + 0.7(1-\pi_i)}$:
- $M_1$: $\pi_1^{+} = \frac{0.63}{0.63 + 0.09} = \frac{0.63}{0.72} = 0.875$; $\pi_1^{-} = \frac{0.07}{0.07 + 0.21} = \frac{0.07}{0.28} = 0.25$.
- $M_2$: $\pi_2^{+} = \frac{0.45}{0.45 + 0.15} = \frac{0.45}{0.60} = 0.75$; $\pi_2^{-} = \frac{0.05}{0.05 + 0.35} = \frac{0.05}{0.40} = 0.125$.
- $M_3$: $\pi_3^{+} = \frac{0.72}{0.72 + 0.06} = \frac{0.72}{0.78} \approx 0.923077$; $\pi_3^{-} = \frac{0.08}{0.08 + 0.14} = \frac{0.08}{0.22} \approx 0.363636$.
Step 3: Post-probe utilities ($U_i^{\pm} = \pi_i^{\pm} - 5c_i - 0.01\ell_i$):
- $M_1$: $U_1^{+} = 0.875 - 0.13 = 0.745$; $U_1^{-} = 0.25 - 0.13 = 0.12$.
- $M_2$: $U_2^{+} = 0.75 - 0.035 = 0.715$; $U_2^{-} = 0.125 - 0.035 = 0.09$.
- $M_3$: $U_3^{+} = 0.923077 - 0.48 = 0.443077$; $U_3^{-} = 0.363636 - 0.48 = -0.116364$.
Step 4: Signal marginal probabilities. $P(s_i = 1) = a\pi_i + b(1-\pi_i)$:
- $M_1$: $0.9(0.7) + 0.3(0.3) = 0.72$; fail: $0.28$.
- $M_2$: $0.9(0.5) + 0.3(0.5) = 0.60$; fail: $0.40$.
- $M_3$: $0.9(0.8) + 0.3(0.2) = 0.78$; fail: $0.22$.
Step 5: Expected utility of the probe policy. With independent signals there are $2^3 = 8$ joint outcomes. For each, the router picks the model with maximum $U_i^{\pm}$:
Outcome (s1 s2 s3) Probability Best model Utility
111 0.72*0.60*0.78 = 0.33696 M1 0.745
110 0.72*0.60*0.22 = 0.09504 M1 0.745
101 0.72*0.40*0.78 = 0.22464 M1 0.745
100 0.72*0.40*0.22 = 0.06336 M1 0.745
011 0.28*0.60*0.78 = 0.13104 M2 0.715
010 0.28*0.60*0.22 = 0.03696 M2 0.715
001 0.28*0.40*0.78 = 0.08736 M3 0.443077
000 0.28*0.40*0.22 = 0.02464 M1 0.12
(Probability check: the eight products sum to $1.00000$.) Then
Computing each term: $0.745 \times 0.72 = 0.53640$; $0.715 \times 0.168 = 0.12012$; $0.443077 \times 0.08736 = 0.0387072$; $0.12 \times 0.02464 = 0.0029568$. Sum:
U^{\text{probe}} = 0.698184 - 0.015 - 0.005 = 0.678184.$ $\
The utility gain over static routing is\
c_p^{\ast} = \frac{U^{\text{probe,gross}} - \mu \ell_p - U^{\text{static}}}{\lambda N} = \frac{0.698184 - 0.005 - 0.57}{5 \times 3} = \frac{0.123184}{15} \approx 0.008212 \text{ dollars}.$ $\
Precision requirement. The static utility gap between the top two candidates is $|U_1 - U_2| = |0.57 - 0.465| = 0.105$. For a two‑sided 95 % confidence level ($z = 1.96$) the required standard error on fitness estimates is\
$$\sigma_f \le \frac{0.105}{1.96 \sqrt{2}} \approx 0.0379.$ $\
Thus probe precision must satisfy $\sigma_f \lesssim 0.0379$ for the routing decision to be statistically meaningful.\
#5. Results\
\ The numerical example demonstrates that probe-informed routing improves expected utility from $0.5700$ to $0.6782$, a net gain of approximately $0.1082$ after accounting for probe costs. The break-even per-probe cost of roughly $\$0.00821$ indicates that probes costing less than this threshold are beneficial under the assumed trade-off weights. The precision analysis shows that fitness estimates with standard error below $0.0379$ are sufficient to make a confident routing decision.\ \
#6. Discussion\
\ The analysis relies on several simplifying assumptions: binary probe outcomes, independent signals across models, and a linear utility model. Real‑world probes may produce graded scores, correlated errors, or non‑linear cost‑utility trade‑offs, which could alter the EVI calculation. Moreover, the break‑even cost depends on the chosen $\lambda$ and $\mu$; different deployment priorities would shift the threshold. A falsification condition is that empirical measurements of probe‑informed routing yield no utility gain over static routing despite probe costs below $c_p^{\ast}$, indicating model misspecification or inaccurate sensitivity parameters $a$ and $b$.\ \ Future work should extend the framework to graded probes, incorporate Bayesian hierarchical priors over model fitness, and validate the theory with empirical deployments across diverse task domains.\ \
#7. Conclusion\
\ We have formalized probe‑before‑commit routing as an expected‑value‑of‑information problem, derived explicit posterior updates, and shown through a worked example that low‑cost diagnostic probes can materially improve routing decisions under realistic cost and latency constraints. The derived break‑even probe cost and precision requirement provide concrete design targets for practitioners seeking to implement intelligent routing in LLM‑driven systems.\ \
#References
[1] AutoDroid: LLM-powered Task Automation in Android. arXiv:2308.15272v4. https://arxiv.org/abs/2308.15272v4 [2] Trusted Routing for Blockchain-Enabled Low-Altitude Intelligent Networks. arXiv:2506.22745v1. https://arxiv.org/abs/2506.22745v1 [3] Exploring the Role of Common Model of Cognition in Designing Adaptive Coaching Interactions for Health Behavior Change. arXiv:1910.07728v3. https://arxiv.org/abs/1910.07728v3 [4] Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers. arXiv:2603.11114v1. https://arxiv.org/abs/2603.11114v1 [5] Interactive Learning for LLM Reasoning. arXiv:2509.26306v5. https://arxiv.org/abs/2509.26306v5 [6] Towards Intelligent Interactive Theatre: Drama Management as a way of Handling Performance. arXiv:1909.10371v2. https://arxiv.org/abs/1909.10371v2 [7] YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition. arXiv:2606.05868v1. https://arxiv.org/abs/2606.05868v1 [8] Task-Aware Delegation Cues for LLM Agents. arXiv:2603.11011v1. https://arxiv.org/abs/2603.11011v1 [9] DOI 10.5281/zenodo.21920604. QNFO: Prioritizing Large Language Models for Scientific Research and Agentic AI: A LiveBench-Grounded Audit (August 2026). [10] DOI 10.5281/zenodo.21945415. QNFO: Joules-per-Solution for Stochastic and Agentic Inference: Benchmarking Frontier and Agentic LLMs Against the Human Brain. [11] DOI 10.5281/zenodo.18199940. QNFO: Comprehensive Technical Framework for Network Isomorphism. [12] DOI 10.5281/zenodo.22758173. QNFO: Syntactic Generation.