auto, that you declare in a per-model policy document. When a request names an alias, Hivenet Router picks one concrete model from the request’s content, rewrites the request’s model field, and then applies quota, admission control, and replica routing to that model exactly as if the client had named it.
Requests that name a concrete model never enter this path. They cost one map lookup and are forwarded byte-identical.
The decision uses structural signals only: tool activity, conversation size, images, output constraints, and lexical markers for code, stack traces, and maths, plus keyword lists you write. It does not call a model or an embedding service. Topics without such markers (for example, “why is the sky blue?”) fall through to the alias’s default route.
POST /v1/chat/completions, POST /v1/messages, and POST /v1/messages/count_tokens. On any other path an alias name is treated as an ordinary, unknown model name.
Prerequisites
- A policy directory set with
--policy-model-dirorHIVENET_ROUTER_POLICY_MODEL_DIR. See Policy YAML reference. - The concrete model IDs exactly as your agents register them. Check them with
GET /v1/models.
Declare model profiles
Aprofile: block describes the models its document claims. The resolver uses it to skip models that cannot serve a request. Every field is optional: an unknown value never excludes a model.
A document that contains only
models: and profile: may omit routing_policy; its models keep the global routing policy. A model can be claimed by only one policy document, so if a model already has its own routing document, put the profile: block in that document.
Declare an alias
Analias: block turns the names in its document’s models: list into aliases. An alias document must not contain routing_policy or profile:; the resolved model’s own policy applies after resolution.
Alias keys
Route keys
Signals
Each signal sets exactly one offeature or keywords, and a weight of 0 or more. The weights of a route’s signals must sum to 1 (the loader rejects the document otherwise), so a route’s score is the sum of the weights of its matching signals and always lies between 0 and 1.
featurewithgteand/orltematches when the value is within the bounds.featurewithout bounds matches when the value is greater than0.keywordsmatches when any keyword appears, case-insensitively and at word boundaries, in the lastrecent_turnshuman turns.
How a model is chosen
1
Pin lookup
If any route of the alias sets
pin_task, the router looks up the task’s pin. A pin is used only if its route still exists and still pins, its model is still one of that route’s candidates, and the model passes the eligibility checks below. A pin that fails only the health check, or a capability only this request needs, is kept for later requests; any other failure drops it. See Task pinning.2
Score routes
Each route scores
Σ weight × match over its own signals. Among routes that reach their own min_score, the highest-scoring route wins if its score is above 0 and at least abstain_below; ties go to the route declared first. Otherwise the decision abstains to default_route.3
Conversation affinity
If
affinity_ttl is set and the conversation was resolved recently, the router keeps its previous model, unless the route chosen in the previous step scores at least affinity_margin more than the conversation’s current route, or that model is no longer eligible.4
Pick a candidate
The router takes the first eligible candidate of the chosen route. If there is none, it tries the default route, then the remaining routes in descending score order.
A route with a single signal scores
1.0 whenever that signal matches, which passes any min_score and abstain_below. If one weak signal should not be enough on its own, give the route a second signal and set min_score above the weak signal’s weight. For example, two signals weighted 0.5 with min_score: 1 require both to match.- It is not itself an alias.
- The caller’s API key may use it: its allowlist, or its per-model quota entries when the key uses per-model quotas. Quota is then charged to the resolved model, so resolution never ends in a quota rejection for an unlisted model.
- Its profile does not rule out a capability the request needs.
- The estimated prompt tokens plus the request’s
max_tokens(ormax_completion_tokens) fit withincontext_window × context_buffer. - At least one healthy agent serves it.
Task pinning
Routes withpin_task: true keep a multi-call task, such as an agent session, on one model. The task key is:
- the
X-Hivenet-Task-IDrequest header, when present; or - a fingerprint of the API key, the end user (OpenAI
useror Anthropicmetadata.user_id), the system prompt, and the first human message.
pin_ttl, up to pin_max_age.
When the pinned model cannot serve one request, that request goes to another candidate:
- If the pinned model has no healthy agent, or lacks a capability only this request needs (for example, one turn with an image on a
vision: falsemodel), the pin is kept. The next request that the model can serve returns to it. - If the key may no longer use the model, or the prompt no longer fits its context window, the pin is dropped. If the route that serves the request sets
pin_task, the task is pinned to the model that served it.
X-Hivenet-Task-ID values longer than 128 bytes are rejected with 400.
Pins live in router memory and are lost on restart. The store holds at most --semantic-pin-max pins in total and --semantic-pin-max-per-key per API key; beyond either limit the least recently used pin is evicted.
POST /v1/messages/count_tokens on an alias always uses the default route and never reads or writes a pin or affinity entry.
Conversation affinity
Routes withoutpin_task are decided again on every request, so without affinity a conversation can move between models when one turn happens to contain code and the next does not. Each move discards the prompt prefix the previous backend had cached. With affinity_ttl set, the router records which route and model answered a conversation, keyed exactly like task pins, and keeps using them while the conversation continues within affinity_ttl, up to pin_max_age.
affinity_margin more than its current route. It also moves on when a reload removes that model from the route, or when the key may no longer use the model or the prompt outgrows its context window. A turn whose model has no healthy agent, or lacks a capability only that turn needs, is served elsewhere, and the conversation returns to its model on the next turn. Affinity entries share the pin store and its limits. A route with pin_task uses its task pin instead.
Responses
A resolved request carries these response headers:
The response body’s
model field is the resolved model, not the alias.
The router does not rewrite the response body, so
model is the name the backend reports for the resolved model. If a client or SDK checks that the response echoes the model it requested, relax that check for aliases, or read X-Hivenet-Routed-Model instead.
The error message names only models the API key may use.
Listing aliases
GET /v1/models lists an alias for an API key when at least one of its candidates is allowed for that key and has a healthy agent. GET /v1/models/{alias} is not supported and returns 404.
Observability
- Audit log: alias requests add
aliasandsemantic_routefields;modelholds the resolved model. See Audit logging. - Metrics:
hivenet_semantic_decisions_total,hivenet_semantic_decision_seconds,hivenet_semantic_decision_log_dropped_total, andhivenet_semantic_decision_log_write_errors_total. See Prometheus metrics. - Decision log: see Decision log.
Decision log
Set--semantic-decision-log (or HIVENET_ROUTER_SEMANTIC_DECISION_LOG) to a file path to append one JSON line per decision. Each line holds the requested alias, resolved model, route, source, per-route scores, matched signals, features, rejected candidates with reasons, task key hash, and decision time in microseconds.
A single background writer appends to the file, so logging never blocks a request. When more than --semantic-decision-log-buffer records are queued, new records are dropped and counted in hivenet_semantic_decision_log_dropped_total. Encode and write failures, such as a full disk, are counted in hivenet_semantic_decision_log_write_errors_total.
The router opens the file once, in append mode, and never rotates or reopens it. Rotate it with copy-and-truncate (for example, copytruncate in logrotate). With rename-based rotation the router keeps writing to the renamed file until it restarts.
Prompt text is not written, but the log is still sensitive.
route_matches lists which of your configured keywords matched each request, which reveals that the prompt contained those words. key_id holds the caller’s key ID, or the tenant and masked key preview for static keys. filtered_out names every rejected candidate, including models the caller’s key may not use. The router creates the file with mode 0600. Keep it at least as protected as your audit log.Reload
Alias and profile blocks reload with the rest of the policy directory onSIGHUP and through PUT /admin/policy/models/{name}. The global policy (--policy-file, _default.yaml, PUT /admin/policy) rejects both blocks. Existing pins are revalidated against the new configuration on their next use.

