Inspired by 9router, built with Go for maximum performance and ultra-low footprint.
This tutorial guides you through compiling, running, and configuring your myAiRouter gateway and dashboard.
| Router | Runtime | Memory (idle) |
|---|---|---|
| MyAiRouter | Native Go | ~14 MB - 23 MB |
| 9Router Next.js | Next.js server | ~132 MB |
| 9Router Node process | Node.js | ~58 MB |
| Total 9Router | Node + Next | ~190 MB |
MyAiRouter: 14 MB - 23 MB — 9Router: ~190 MB — ×ばつ less memory
- Pure Pass-Through Gateway: Requests are forwarded byte-for-byte to the selected upstream by default. The gateway never caches responses and never rewrites bodies unless you explicitly opt in — response caching belongs to your client or agent tooling, and provider-native prompt caching keeps working untouched.
-
Error-Classified Fallback Engine: Fallback decisions are status-aware:
429/408/5xxand network errors advance to the next target,401/403try remaining accounts of the same provider only, and terminal errors (400/404/413/422) return the upstream error immediately instead of burning tokens on other routes. Policy is tunable per route (auto/aggressive/conservative) with configurable attempt timeouts and fallback budgets. -
Health-Aware Routing: In-memory per-connection health (failure streaks with exponential cooldown, EWMA latency/TTFB) skips recently-failed upstreams and orders accounts by observed responsiveness, exposed at
GET /api/connections/health. - Model-Centric Routing Policies: Routing, fallback, and compression rules are configured per-model rather than globally.
-
Custom Fallback Models: Fail over seamlessly to alternative model IDs (e.g. falling back from
deepseek/deepseek-v4-flashtoopzen/mimo-v2.5-free) when primary providers return retryable errors. -
Opt-In Context Compression (off by default):
- Protected Prefix (System prompts and tool definitions) — Preserved verbatim.
- Middle History (Older conversation context) — Compressed dynamically using the AST optimizer or RTK fallback.
-
Protected Suffix (Last
$N$ recent chat messages, default 20) — Preserved verbatim.
-
Explicit Compression Triggers:
-
Proactive (
threshold): Compresses when request exceeds user-specified token threshold. -
Reactive (
context_limit): Compresses only when request exceeds model context limits (OpenAI: 128k, Anthropic: 200k, Gemini: 1M).
-
Proactive (
-
Streaming-Safe Concurrency:
race/parallel/ensemblestrategies buffer child responses and replay exactly one winner to the socket; streaming (stream:true) requests always take the sequential path. -
Live Reloading Watcher: Native file watching and recompilation of the Go backend using the integrated
airdev server.
Because myAiRouter embeds all frontend assets directly into the Go executable, you only need to run a simple build step to generate the final standalone binary.
Navigate to the web folder, install dependencies, and build the static production distribution:
cd web npm install npm run build cd ..
This creates the static HTML, JS, and CSS files inside web/dist/.
Compile the Go entry code to produce a standalone executable binary named myAiRouter:
go build -o myAiRouter .This packages the Go web server, the SQLite database migrations, local agent skills, and embedded Vite assets into a single binary.
curl -fsSL https://haslab-dev.github.io/MyAiRouter/website/install.sh | bashInstalls to $HOME/.local/bin/myairouter (or /usr/local/bin/myairouter).
myairouter # start server (foreground) myairouter start # start server (foreground) myairouter start -d # start server (background daemon) myairouter status # show server status, running PIDs & listening ports myairouter stop # stop all running server processes (auto-sweeps duplicates) myairouter restart # restart background daemon myairouter bg # background alias myairouter version # print version
By default, the server runs on port 20128. Set PORT to change:
PORT=8080 myairouter
On startup, myAiRouter will:
- Initialize a SQLite database at
~/.myairouter/db.sqlite. - Apply database migrations and seed default configuration settings.
- Automatically sweep and terminate any duplicate process instances.
- Start the API gateway at
http://localhost:20128/v1/. - Host the space-dark dashboard at
http://localhost:20128/.
To build the client and run the server locally with file watching and live reloading:
make dev-server # Launches Go backend (auto-runs `air` hot reload if installed, falling back to `go run .`) make dev-client # Launches Vite HMR client on port 5173
Open http://localhost:5173 in your browser. All API requests are automatically proxied to the Go backend.
The Request Traces dashboard (http://localhost:20128/traces) displays routing-focused analytics organized into four distinct sections:
- Summary: High-level execution metrics (Latency, TTFB, Input/Output/Cached Tokens, Cost, Prompt Compression %, Streaming, Attempts count, Fallback & Retry counts).
- Cached Tokens: Upstream-reported provider prompt-cache reuse (the gateway itself does not cache responses).
- Route Graph: Visual node tree for Fallback and Load Balance routing strategies showing per-node execution status (✔ Success, ✖ Failed).
- Pipeline: Routing-focused steps (
Resolve Model,Request Preparation(only when compression is opted in),Route,Provider). - Request / Response Preview: Request metadata (
system,user,messages,chars,tokens, plus detailed compression telemetry) and Response metadata (preview,finish_reason).
make patch-version # bump patch (0.1.0 → 0.1.1) make minor-version # bump minor, reset patch (0.1.0 → 0.2.0) make major-version # bump major, reset minor+patch (0.1.0 → 1.0.0) make set-version V=x.y.z # set explicit version
Updates both main.go (backend) and web/package.json (client).
- Open your web browser and navigate to the dashboard at:
http://localhost:20128/. - Go to the Providers section using the sidebar navigation.
- Add or configure your credentials:
- Core Providers: Configure active connections for OpenAI, Anthropic, DeepSeek, Gemini, etc.
- Custom Providers: Click Add OpenAI/Anthropic Compatible to register custom target proxies, configure endpoints, and assign credentials. Easily remove nodes completely via the red Remove button.
- Test connectivity next to the connection card to verify configuration.
By default, API gateway authentication is disabled. You can query completions directly:
Query the completions endpoint using curl:
curl -N http://localhost:20128/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "openai/gpt-4o-mini", "messages": [ {"role": "user", "content": "Hello! What is your name?"} ], "stream": true }'
Your gateway hosts local instructions that autonomous agents (such as Cline, Roo Code, or Claude Code) can load.
- Entry point skill:
http://localhost:20128/skills/myairouter/SKILL.md - Chat skill:
http://localhost:20128/skills/myairouter-chat/SKILL.md - Token Saving details:
http://localhost:20128/skills/myairouter-token-saver/SKILL.md
You can view, read, and copy these skill URLs directly under the Agent Skills section of the web dashboard.