Skip to main content
A semantic model alias is a virtual model name, such as auto, that you declare in a per-model policy document. When a request names an alias, Hivenet Router picks one concrete model from the request’s content, rewrites the request’s model field, and then applies quota, admission control, and replica routing to that model exactly as if the client had named it. Requests that name a concrete model never enter this path. They cost one map lookup and are forwarded byte-identical.
The decision uses structural signals only: tool activity, conversation size, images, output constraints, and lexical markers for code, stack traces, and maths, plus keyword lists you write. It does not call a model or an embedding service. Topics without such markers (for example, “why is the sky blue?”) fall through to the alias’s default route.
Aliases apply to POST /v1/chat/completions, POST /v1/messages, and POST /v1/messages/count_tokens. On any other path an alias name is treated as an ordinary, unknown model name.

Prerequisites

  • A policy directory set with --policy-model-dir or HIVENET_ROUTER_POLICY_MODEL_DIR. See Policy YAML reference.
  • The concrete model IDs exactly as your agents register them. Check them with GET /v1/models.

Declare model profiles

A profile: block describes the models its document claims. The resolver uses it to skip models that cannot serve a request. Every field is optional: an unknown value never excludes a model.
An unset capability counts as supported. A model whose profile omits vision, for example, or a model with no profile: at all, can be chosen for a request with images, tools, or JSON output that it cannot serve. Set every capability a model lacks to false.
A document that contains only models: and profile: may omit routing_policy; its models keep the global routing policy. A model can be claimed by only one policy document, so if a model already has its own routing document, put the profile: block in that document.

Declare an alias

An alias: block turns the names in its document’s models: list into aliases. An alias document must not contain routing_policy or profile:; the resolved model’s own policy applies after resolution.
The weights and candidate order above are an example, not a recommendation.

Alias keys

Route keys

Signals

Each signal sets exactly one of feature or keywords, and a weight of 0 or more. The weights of a route’s signals must sum to 1 (the loader rejects the document otherwise), so a route’s score is the sum of the weights of its matching signals and always lies between 0 and 1.
  • feature with gte and/or lte matches when the value is within the bounds.
  • feature without bounds matches when the value is greater than 0.
  • keywords matches when any keyword appears, case-insensitively and at word boundaries, in the last recent_turns human turns.

How a model is chosen

1

Pin lookup

If any route of the alias sets pin_task, the router looks up the task’s pin. A pin is used only if its route still exists and still pins, its model is still one of that route’s candidates, and the model passes the eligibility checks below. A pin that fails only the health check, or a capability only this request needs, is kept for later requests; any other failure drops it. See Task pinning.
2

Score routes

Each route scores Σ weight × match over its own signals. Among routes that reach their own min_score, the highest-scoring route wins if its score is above 0 and at least abstain_below; ties go to the route declared first. Otherwise the decision abstains to default_route.
3

Conversation affinity

If affinity_ttl is set and the conversation was resolved recently, the router keeps its previous model, unless the route chosen in the previous step scores at least affinity_margin more than the conversation’s current route, or that model is no longer eligible.
4

Pick a candidate

The router takes the first eligible candidate of the chosen route. If there is none, it tries the default route, then the remaining routes in descending score order.
A route with a single signal scores 1.0 whenever that signal matches, which passes any min_score and abstain_below. If one weak signal should not be enough on its own, give the route a second signal and set min_score above the weak signal’s weight. For example, two signals weighted 0.5 with min_score: 1 require both to match.
A candidate is eligible when all of these hold, checked in this order:
  1. It is not itself an alias.
  2. The caller’s API key may use it: its allowlist, or its per-model quota entries when the key uses per-model quotas. Quota is then charged to the resolved model, so resolution never ends in a quota rejection for an unlisted model.
  3. Its profile does not rule out a capability the request needs.
  4. The estimated prompt tokens plus the request’s max_tokens (or max_completion_tokens) fit within context_window × context_buffer.
  5. At least one healthy agent serves it.
Grant API keys the concrete candidate models, not the alias name. Model access is checked on the candidates, through the key’s allowlist or its per-model quota entries. Listing auto in allowed_models is neither needed nor enough. A key that may use none of an alias’s candidates gets 403 model_forbidden. See Model restrictions.

Task pinning

Routes with pin_task: true keep a multi-call task, such as an agent session, on one model. The task key is:
  • the X-Hivenet-Task-ID request header, when present; or
  • a fingerprint of the API key, the end user (OpenAI user or Anthropic metadata.user_id), the system prompt, and the first human message.
Keys are scoped to the alias and the API key, and stored as hashes. A pin lasts while the task keeps calling within pin_ttl, up to pin_max_age. When the pinned model cannot serve one request, that request goes to another candidate:
  • If the pinned model has no healthy agent, or lacks a capability only this request needs (for example, one turn with an image on a vision: false model), the pin is kept. The next request that the model can serve returns to it.
  • If the key may no longer use the model, or the prompt no longer fits its context window, the pin is dropped. If the route that serves the request sets pin_task, the task is pinned to the model that served it.
Without X-Hivenet-Task-ID or an end-user field, conversations that share an API key, a system prompt, and a first message share one pin. Send X-Hivenet-Task-ID when many users share one key.
X-Hivenet-Task-ID values longer than 128 bytes are rejected with 400. Pins live in router memory and are lost on restart. The store holds at most --semantic-pin-max pins in total and --semantic-pin-max-per-key per API key; beyond either limit the least recently used pin is evicted. POST /v1/messages/count_tokens on an alias always uses the default route and never reads or writes a pin or affinity entry.

Conversation affinity

Routes without pin_task are decided again on every request, so without affinity a conversation can move between models when one turn happens to contain code and the next does not. Each move discards the prompt prefix the previous backend had cached. With affinity_ttl set, the router records which route and model answered a conversation, keyed exactly like task pins, and keeps using them while the conversation continues within affinity_ttl, up to pin_max_age.
The conversation switches when the newly chosen route scores at least affinity_margin more than its current route. It also moves on when a reload removes that model from the route, or when the key may no longer use the model or the prompt outgrows its context window. A turn whose model has no healthy agent, or lacks a capability only that turn needs, is served elsewhere, and the conversation returns to its model on the next turn. Affinity entries share the pin store and its limits. A route with pin_task uses its task pin instead.

Responses

A resolved request carries these response headers: The response body’s model field is the resolved model, not the alias.
The router does not rewrite the response body, so model is the name the backend reports for the resolved model. If a client or SDK checks that the response echoes the model it requested, relax that check for aliases, or read X-Hivenet-Routed-Model instead.
When no candidate is eligible, the router answers before quota or routing runs: The error message names only models the API key may use.

Listing aliases

GET /v1/models lists an alias for an API key when at least one of its candidates is allowed for that key and has a healthy agent. GET /v1/models/{alias} is not supported and returns 404.

Observability

  • Audit log: alias requests add alias and semantic_route fields; model holds the resolved model. See Audit logging.
  • Metrics: hivenet_semantic_decisions_total, hivenet_semantic_decision_seconds, hivenet_semantic_decision_log_dropped_total, and hivenet_semantic_decision_log_write_errors_total. See Prometheus metrics.
  • Decision log: see Decision log.

Decision log

Set --semantic-decision-log (or HIVENET_ROUTER_SEMANTIC_DECISION_LOG) to a file path to append one JSON line per decision. Each line holds the requested alias, resolved model, route, source, per-route scores, matched signals, features, rejected candidates with reasons, task key hash, and decision time in microseconds. A single background writer appends to the file, so logging never blocks a request. When more than --semantic-decision-log-buffer records are queued, new records are dropped and counted in hivenet_semantic_decision_log_dropped_total. Encode and write failures, such as a full disk, are counted in hivenet_semantic_decision_log_write_errors_total. The router opens the file once, in append mode, and never rotates or reopens it. Rotate it with copy-and-truncate (for example, copytruncate in logrotate). With rename-based rotation the router keeps writing to the renamed file until it restarts.
Prompt text is not written, but the log is still sensitive. route_matches lists which of your configured keywords matched each request, which reveals that the prompt contained those words. key_id holds the caller’s key ID, or the tenant and masked key preview for static keys. filtered_out names every rejected candidate, including models the caller’s key may not use. The router creates the file with mode 0600. Keep it at least as protected as your audit log.

Reload

Alias and profile blocks reload with the rest of the policy directory on SIGHUP and through PUT /admin/policy/models/{name}. The global policy (--policy-file, _default.yaml, PUT /admin/policy) rejects both blocks. Existing pins are revalidated against the new configuration on their next use.