stage 1: cheap MLP gate
Reject requests that are clearly not worth probing.
- Inputs
- operation, text, device state, session context
- Output
- probe or skip to cloud
OnDevice determines how much of your cloud-based LLM traffic can move on-device while preserving application quality, latency, reliability, and power requirements.
Your application runs inference entirely in the cloud, drawing on open- or closed-source models.
Now you want to move some of that intelligence onto the device.
But you do not know which operations belong there. You do not know which models will survive the move without degrading the application. And you do not know what optimal means once accuracy, latency, memory, power, hardware, and the shape of a user session all begin pulling in different directions.
Should you run a quantized 3B parameter model or a smaller 1.5B parameter model? Should the model remain resident in memory to minimize time to first token? Can it? What happens when the user opens another memory-intensive application, the battery begins to fall, or the session lasts longer than expected?
Even once you discover a viable on-device policy, another set of problems appears. You need to route eligible traffic onto the device, fall back to the cloud when local execution is no longer safe, and preserve the behavior of the application across both paths.
Then the ground moves again.
Traffic patterns change. New features are built. New models are released. New quantizations become available. Devices become more capable. A policy that was optimal six months ago may no longer be optimal today.
An offline inference-policy search engine. It observes representative application traffic and continuously searches across models, quantizations, execution strategies, memory configurations, and device classes to determine which operations can run locally, under what conditions, and when they should fall back to the cloud.
We model discovery as a constrained contextual bandit because each request arrives with observable context \(x\), each feasible inference pathway is an action \(a \in \mathcal{A}\), and the system must learn which action should handle the request while respecting production constraints.
The key quantity is the conditional success model:
The optimizer is not simply asking which model is best. It is solving a constrained assignment problem:
subject to:
\(L\) is latency, \(M\) is memory, \(s\) is device/runtime state, and \(q_{\min}\) is the required quality threshold.
What makes this different from a textbook bandit is that actions are compositional. A pathway includes model family, quantization, prompt, context size, decoding policy, model residency, prefix/cache state, checkpoint strategy, and fallback behavior. Some knobs change answer quality; others change feasibility, latency, memory, or future runtime state. Discovery has to reason over both.
In the larger system, discovery is the offline evidence-acquisition and policy-search stage. It consumes captured traffic, local/cloud evaluator outcomes, synthetic expansions, model/runtime measurements, and device constraints. It selectively measures request-path pairs, estimates where local pathways succeed, compares cheaper challengers against stronger incumbents, and searches for the policy with maximum expected cloud reduction under quality and systems constraints.
Its output is a frozen candidate policy: local-serving regions, the inference paths assigned to those regions, uncertainty estimates, expected coverage, fallback behavior, and the runtime assumptions needed to execute safely.
A discovered router policy can look like this. The score checks are cheap router gates; the local model is only invoked after a rule accepts.
{
"serving_policy": "ordered_local_rules_with_cloud_fallback",
"quality_target": 0.95,
"rules": [
{
"when": "score(qwen1.5b-q4km) >= 0.795",
"then": "run qwen1.5b-q4km locally",
"runtime": {
"context_tokens": 512,
"residency": "HOT",
"shared_prefix_state": "shared_prefix_hot",
"estimated_latency_ms": 30,
"estimated_memory_gb": 1.50
},
"fallback": "next_rule_then_cloud"
},
{
"when": "score(qwen3b-q4km) >= 0.905",
"then": "run qwen3b-q4km locally",
"runtime": {
"context_tokens": 512,
"residency": "HOT",
"shared_prefix_state": "shared_prefix_hot",
"estimated_latency_ms": 234,
"estimated_memory_gb": 2.55
},
"fallback": "next_rule_then_cloud"
}
],
"fallback": "cloud"
}
A production policy layer. For each request, it considers the operation, device state, session context, and application requirements, then chooses the appropriate on-device or cloud execution path.
Router compilation turns the discovered offline policy into an executable on-device decision program. Discovery may find that a region \(R\) can be handled by pathway \(a\), but production needs a legal recognizer \(g(x,z)\) that can decide membership at request time using only available signals.
Those signals can include request features, runtime state, and optionally model-state features \(z\) acquired during a lightweight prefill or probe pass.
The compiler objective is closer to:
subject to:
So the compiler is not just training a classifier. It is lowering a discovered policy into an efficient cascade: grouping compatible rules, sharing signal-acquisition passes, calibrating conservative thresholds, and emitting MLP gates that can run locally.
Reject requests that are clearly not worth probing.
Use the model's own representation to decide local versus cloud.
In production, multiple logical rules should not necessarily mean multiple model passes. If several rules depend on the same model, prompt, or prefix, the compiler should acquire the shared probe once, extract all required features, and evaluate the gates in memory.
The key innovation is using hidden-state information as router training signal. The router is trying to predict the local model's competence boundary, and that boundary is partly encoded inside the model's own representation of the request.
Surface text features can say what the request looks like to us; hidden states say what the request looks like to the model. By training MLP gates over these model-state features, the router learns a boundary in the model's latent geometry instead of relying only on brittle lexical features. This is why prefill-derived signals can improve segmentation: the local model helps expose the shape of its own reliable region before generation.
Examples of hidden-state features we consider include:
The on-device runtime. It carries out the policy efficiently, managing model residency, memory, scheduling, and execution across supported hardware.