Connect the Pi coding agent to Hivenet Router through OpenAI Chat Completions and configure models, authentication, tools, security, and compatibility.
Pi can use Hivenet Router as a custom OpenAI-compatible model provider.Pi sends requests to:
POST /v1/chat/completions
Hivenet Router authenticates the request, checks model access and quotas, selects an eligible llm agent, and forwards the request to its inference backend.
Pi’s built-in tools can read, create, edit, and delete files and run shell commands with the permissions of the operating-system user that started Pi.Pi does not provide a built-in approval system or security sandbox. Use a dedicated account, container, virtual machine, or another execution boundary when the project or environment requires stronger isolation.
Pi can answer text prompts through any compatible chat model.For coding work, the selected model and inference engine must also support structured tool calls for operations such as:
reading files
editing and creating files
running shell commands
searching the project
invoking extension-provided tools
For vLLM, automatic tool calling normally requires:
with the parser supported by that model family.Some models also require a tool-compatible chat template.
Tool-call parsers are model-specific.Do not copy a parser from another model family without checking the inference-engine documentation and testing the complete tool loop.
Register the same public model ID with the Hivenet Router agent:
Verify that the variable is present without printing it:
test -n "$HIVENET_ROUTER_API_KEY" \ && echo "HIVENET_ROUTER_API_KEY is set"
The configuration value:
"apiKey": "$HIVENET_ROUTER_API_KEY"
resolves the environment variable when Pi sends a request.The provider setting:
"authHeader": true
makes Pi send the resolved credential as:
Authorization: Bearer <hivenet-router-api-key>
Do not put:
the SHA-256 key hash
an administrator key
the agent JWT secret
a provider fallback key
in this variable.
HIVENET_ROUTER_API_KEY is a client-side environment variable chosen for this guide.It does not need to follow the case-sensitive HIVENET_ROUTER_* convention used by the Hivenet Router process.
Pi can resolve a credential by running a shell command:
{ "apiKey": "!op read 'op://Infrastructure/Hivenet Router API key/credential'"}
or:
{ "apiKey": "!pass show services/hivenet-router/pi"}
The command must print only the raw API key to standard output.Pi resolves command-backed values when a request is sent. It does not add its own caching, stale-value reuse, or retry policy for arbitrary credential commands.Use a wrapper script when the secret command needs:
Pi then sends the resolved apiKey value in the standard bearer header Hivenet Router accepts. Without this setting, a custom provider can resolve the key without attaching it as Authorization: Bearer.
This is the human-readable label shown as secondary model information.The model picker and status display still identify the model primarily by its configured id.
This tells Pi how much conversation context the model can accept.Use the effective context limit of the deployed backend, not only the theoretical limit of the underlying model.For example, a model repository may support 128,000 tokens while the backend is started with:
--max-model-len 32768
In that deployment, configure:
"contextWindow": 32768
An accurate value helps Pi compact the conversation before the backend rejects it.
This is the largest output Pi should request.The prompt and requested output must fit inside the backend context limit.Use a value that leaves enough room for:
Pi uses these values for its local cost display.They do not affect Hivenet Router billing, quotas, or routing.For a self-hosted deployment, zero values are reasonable unless you maintain your own internal per-token cost model.
A model may remain configured in Pi while temporarily unavailable in Hivenet Router. Selecting it then produces the relevant model or availability error.
Do not put the raw Hivenet Router API key in a project settings file.A project-level .pi/settings.json may be committed to source control. Keep the credential in an environment variable, protected secret file, or credential command.
Current vLLM versions support reasoning_effort for several reasoning-model integrations, so disabling it unconditionally can remove a feature that the deployment supports.
Pi also supports model-specific thinking-level mappings and compatibility formats.The appropriate configuration depends on:
model family
reasoning parser
chat template
vLLM version
whether the server expects reasoning_effort
whether it expects enable_thinking
how reasoning is represented in streamed responses
Establish basic chat and tool compatibility before enabling reasoning controls.
A model being described as a reasoning model does not prove that its tool calls work reliably while reasoning is enabled.Test reasoning and tool use together.
Long coding requests may exceed it because they can include:
large system instructions
many tool definitions
long project context
slow model startup
large requested outputs
backend queueing
Increase it when representative requests need more time:
./bin/hivenet-router \ --request-timeout 5m \ ...
The client cannot extend a shorter deadline enforced by the router.Avoid increasing the timeout before checking whether the backend is healthy and making progress.
Pi performs some startup network operations independently of model inference.Disable install and update telemetry:
export PI_TELEMETRY=0
Disable only the version check:
export PI_SKIP_VERSION_CHECK=1
Disable Pi’s startup network operations:
export PI_OFFLINE=1
You can also start one session with:
pi --offline
Offline mode disables Pi’s startup update, package-update, and telemetry requests. It does not prevent the selected model provider from calling the configured Hivenet Router router.Network policy remains the authoritative control for:
command uploads the current session as a private GitHub gist.Do not use it when a session may contain:
private source code
credentials
customer data
internal prompts or documents
security findings
regulated information
The command is explicit rather than automatic, but Pi does not provide an internal policy control that prevents a user or extension from invoking external network operations.Use network restrictions and organizational controls where sharing must be prohibited.
Pi can load extensions, skills, prompt templates, and packages.These can change:
available tools
model behavior
project instructions
network access
command execution
session handling
Third-party extensions execute with the same system access as Pi.Review their source and provenance before installation.For a minimal Hivenet Router integration test, begin without third-party extensions and add them only after the model, tools, and router path work correctly.
the chat template supports tool and tool-result messages
the backend returns structured tool_calls
the model follows the schemas reliably
Test the backend directly with a Chat Completions request containing a tools array.Hivenet Router forwards the response. It does not convert plain model text into a structured tool call.
Tool calls appear as text
The backend parser did not recognize the model’s tool syntax.Review:
model family
tool-call parser
chat template
backend version
streaming tool-call support
Inspect the backend logs first.
The backend rejects the `developer` role
Add:
"compat": { "supportsDeveloperRole": false}
to the provider or affected model.Open /model again to reload the configuration.
The backend rejects `reasoning_effort`
Add:
"compat": { "supportsReasoningEffort": false}
Do this only for models or backends that reject the field.
The backend rejects `max_completion_tokens`
Add:
"compat": { "maxTokensField": "max_tokens"}
Current vLLM supports both fields, but other OpenAI-compatible servers may not.
Streaming usage causes a validation error
Add:
"compat": { "supportsUsageInStreaming": false}
This prevents Pi from requesting the final streaming usage block.
Pi compacts too early
Increase the configured:
"contextWindow"
only when the deployed backend genuinely accepts the larger context.Check the backend’s effective maximum model length.
The backend rejects the prompt as too long
Reduce:
contextWindow
maxTokens
included project context
requested output
An accurate context configuration lets Pi compact before the request reaches the backend limit.
Output arrives only after completion
Test streaming directly with curl.Then check:
backend SSE behavior
reverse-proxy buffering
proxy read and idle timeouts
Hivenet Router request timeout
agent logs
Pi modifies files without asking
This is expected behavior.Pi does not have a built-in permission-prompt system.Restart it with restricted tools:
pi \ --tools read,grep,find,ls
Use a container, virtual machine, or equivalent sandbox when tools need stronger boundaries.
Requests return `504 request_timeout`
The request exceeded the Hivenet Router deadline.Check:
backend readiness
prompt size
output size
engine queue depth
router-side capacity queueing
current request timeout
Increase the deadline only after confirming that the backend is progressing normally.