API reference
OpenAI-compatible, routed to independent nodes.
Swap your base URL and keep your code. Requests go through the TezGrid gateway to NVIDIA, Apple Silicon, AMD and CPU nodes run by operators. Usage settles per token in integer micro-units, where $1 is 1,000,000 µ.
https://api.tezgrid.com/v1
Every path below is relative to this base, except /health and /ready. A local gateway runs on http://127.0.0.1:3000.
curl https://api.tezgrid.com/v1/chat/completions \ -H "Authorization: Bearer $TEZGRID_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "Qwen2.5-1.5B-Instruct", "messages": [{"role": "user", "content": "Hello"}]}'
Auth & billing
Chat completions need an API key from the console. Public reads like models, pricing and nodes don't.
- Authorization
Bearer <api_key>. Missing or invalid key returns401.- Charge
- 1 µ per prompt token plus 2 µ per completion token, i.e. $1 and $2 per million tokens.
- Split
- 94.5% to the operator that served the request, 5.5% platform fee.
- Empty wallet
- Returns
402. Nothing is sent to a node. - Idempotency-Key
- Retries with the same key are never billed or executed twice.
Routing headers
All optional. Plain OpenAI clients work without them.
Request
- X-TezGrid-Min-Trust
- Refuse nodes below a tier:
standard,verifiedorconfidential. - X-TezGrid-Developer-Email
- Attribute usage to a developer account.
- X-External-Request-Id
- Your own correlation ID, echoed in usage records.
- X-TezGrid-Debug: latency
- Adds a
tezgrid_debugobject to non-streaming JSON responses.
Response
- X-TezGrid-Request-Id
- Gateway request ID.
- X-TezGrid-Provider-Id
- Node that served the request.
- X-TezGrid-Provider-Region
- Region of that node.
- X-TezGrid-Router-Latency-Ms
- Time spent choosing a node.
- X-TezGrid-Gateway-Overhead-Ms
- Total time added by the gateway.
- X-TezGrid-Estimated-Latency-Ms
- Router's latency estimate for the chosen node.
- X-TezGrid-Failover-Attempts
- How many nodes were tried before this one answered.
- X-TezGrid-TTFT-Ms
- Time to first token. Non-streaming only.
From the CLI, tezgrid ping-api --gateway URL runs the same probe. It needs TEZGRID_API_KEY for the chat check.
Operator setup
Inference runs on the operator's machine. The installer detects hardware and lists the models it can serve before you download any weights.
-
Install the CLI and llama-server
Creates an operator home with
bin/andmodels/, and installs the CLI withllama-serverfrom llama.cpp (Vulkan build on Windows and Linux x64, Metal on macOS). If you already run Ollama, TezGrid can use it instead. To use your own llama.cpp build, setTEZGRID_LLAMA_SERVER.tezgrid doctorshows which runtimes it found. Stop a running node withtezgrid downbefore reinstalling.# Linux (x86_64, arm64) and macOS (Apple Silicon, Intel) curl -fsSL https://get.tezgrid.com/install.sh | bash # Windows (PowerShell) irm https://get.tezgrid.com/install-windows.ps1 | iex
-
See which models fit
The installer prints CPU, RAM, GPU and disk, then splits the catalog into models that fit and models that don't. Run it again any time:
tezgrid doctor tezgrid recommend
-
Sign in
Same email and password as the console. The first login creates a provider account. The session is stored on this machine in
session.json. Each machine is bound to one account and one location.tezgrid login you@example.com '<password>'
-
Install a model
GGUF weights download into the operator
models/directory. With--from ollamathe model is pulled through Ollama instead.tezgrid install-model Qwen2.5-1.5B-Instruct # or: tezgrid install-model Llama-3.2-1B-Instruct # or, with Ollama: tezgrid install-model llama3.2 --from ollama
-
Go online
Starts llama-server with your GGUF (or uses Ollama, or a server already running on the inference port), registers only the models you installed and sends a heartbeat every 15 seconds.
tezgrid startis an alias.tezgrid up
-
Stop
Marks the node offline and stops the processes
tezgrid upstarted. Ctrl+C does the same.tezgrid down
Endpoints
Click an endpoint for its request and response shape.
Inference
POST /v1/chat/completionsChat completions, streaming or not auth
Body
{
"model": "Qwen2.5-1.5B-Instruct",
"messages": [{"role": "user", "content": "Hello from TezGrid."}],
"stream": false,
"temperature": 0.7,
"max_tokens": 512
}
Charged at prompt + 2 × completion tokens, in µ. The operator gets 94.5%. Empty wallet returns 402, missing key returns 401. Routing headers are listed under Routing headers.
GET /v1/modelsModels available on the network public
Response
{
"object": "list",
"data": [
{ "id": "llama3-8b", "object": "model", "owned_by": "tezgrid" }
]
}
Also GET /v1/models/legacy and GET /v1/provider/models?provider_id=…
GET /v1/pricingModeled marketplace pricing for reference models public
Economics from MarketplacePricingEngine for DeepSeek R1 Distill, Llama 8B and 70B, Qwen 72B and DeepSeek V3. These are modeled estimates, not a live per-request quote.
Accounts
POST /v1/auth/signupCreate a developer or provider account public
Body
{
"name": "Ada",
"email": "ada@example.com",
"password": "••••••••",
"role": "developer"
}
POST /v1/auth/login takes email, password and role.
GET /v1/developer/wallet/:developer_idWallet balance in micro-units auth
The developer ID is usually dev_<normalized_email>. The response includes wallet_balance_micro.
Related
GET /v1/developer/usage returns stored usage rows from MySQL. GET /v1/developer/tokens/hourly returns an hourly series for the console chart.
Network & providers
GET /v1/nodesRegistered and online provider nodes public
Returns { nodes: [...], mysql_enabled } with status, hardware, settled earnings and the telemetry fields the console uses.
POST /v1/provider/registerRegister a node and its inference endpoint auth
Body
{
"provider_id": "<uuid>",
"node_name": "studio-mac",
"hardware": "Apple M4 Max",
"device_type": "Apple Silicon",
"inference_endpoint": "http://203.0.113.10:8080",
"models": ["Qwen2.5-1.5B-Instruct"],
"status": "ONLINE"
}
tezgrid up calls this for you. Most operators never call it directly.
GET /v1/network/metricsFleet-wide traffic and latency public
Includes a latency block with avg_router_ms, avg_ttft_ms, avg_decode_tps and success_rate_percent. Pass ?window=30m|24h|7d|30d for the traffic series that powers the network page.
Provider-scoped versions: GET /v1/provider/metrics and GET /v1/provider/dashboard.
GET /v1/diagnostics/latencyRouter and TTFT diagnostics, no inference needed public
Response
{
"gateway_reachable": true,
"online_providers": 12,
"avg_router_latency_ms": 0.9,
"avg_network_ttft_ms": 180,
"avg_network_decode_tps": 42,
"network_success_rate_percent": 99.2
}
Health
GET /healthGateway liveness public
Reports mysql_enabled, pilot readiness and any warnings. GET /ready is the readiness probe. Neither route is under /v1.