API reference

OpenAI-compatible, routed to independent nodes.

Swap your base URL and keep your code. Requests go through the TezGrid gateway to NVIDIA, Apple Silicon, AMD and CPU nodes run by operators. Usage settles per token in integer micro-units, where $1 is 1,000,000 µ.

Base URL
https://api.tezgrid.com/v1

Every path below is relative to this base, except /health and /ready. A local gateway runs on http://127.0.0.1:3000.

quickstart
curl https://api.tezgrid.com/v1/chat/completions \
  -H "Authorization: Bearer $TEZGRID_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "Qwen2.5-1.5B-Instruct", "messages": [{"role": "user", "content": "Hello"}]}'

Auth & billing

Chat completions need an API key from the console. Public reads like models, pricing and nodes don't.

Authorization
Bearer <api_key>. Missing or invalid key returns 401.
Charge
1 µ per prompt token plus 2 µ per completion token, i.e. $1 and $2 per million tokens.
Split
94.5% to the operator that served the request, 5.5% platform fee.
Empty wallet
Returns 402. Nothing is sent to a node.
Idempotency-Key
Retries with the same key are never billed or executed twice.

Routing headers

All optional. Plain OpenAI clients work without them.

Request

X-TezGrid-Min-Trust
Refuse nodes below a tier: standard, verified or confidential.
X-TezGrid-Developer-Email
Attribute usage to a developer account.
X-External-Request-Id
Your own correlation ID, echoed in usage records.
X-TezGrid-Debug: latency
Adds a tezgrid_debug object to non-streaming JSON responses.

Response

X-TezGrid-Request-Id
Gateway request ID.
X-TezGrid-Provider-Id
Node that served the request.
X-TezGrid-Provider-Region
Region of that node.
X-TezGrid-Router-Latency-Ms
Time spent choosing a node.
X-TezGrid-Gateway-Overhead-Ms
Total time added by the gateway.
X-TezGrid-Estimated-Latency-Ms
Router's latency estimate for the chosen node.
X-TezGrid-Failover-Attempts
How many nodes were tried before this one answered.
X-TezGrid-TTFT-Ms
Time to first token. Non-streaming only.

From the CLI, tezgrid ping-api --gateway URL runs the same probe. It needs TEZGRID_API_KEY for the chat check.

Operator setup

Inference runs on the operator's machine. The installer detects hardware and lists the models it can serve before you download any weights.

  1. Install the CLI and llama-server

    Creates an operator home with bin/ and models/, and installs the CLI with llama-server from llama.cpp (Vulkan build on Windows and Linux x64, Metal on macOS). If you already run Ollama, TezGrid can use it instead. To use your own llama.cpp build, set TEZGRID_LLAMA_SERVER. tezgrid doctor shows which runtimes it found. Stop a running node with tezgrid down before reinstalling.

    # Linux (x86_64, arm64) and macOS (Apple Silicon, Intel)
    curl -fsSL https://get.tezgrid.com/install.sh | bash
    
    # Windows (PowerShell)
    irm https://get.tezgrid.com/install-windows.ps1 | iex
  2. See which models fit

    The installer prints CPU, RAM, GPU and disk, then splits the catalog into models that fit and models that don't. Run it again any time:

    tezgrid doctor
    tezgrid recommend
  3. Sign in

    Same email and password as the console. The first login creates a provider account. The session is stored on this machine in session.json. Each machine is bound to one account and one location.

    tezgrid login you@example.com '<password>'
  4. Install a model

    GGUF weights download into the operator models/ directory. With --from ollama the model is pulled through Ollama instead.

    tezgrid install-model Qwen2.5-1.5B-Instruct
    # or: tezgrid install-model Llama-3.2-1B-Instruct
    # or, with Ollama: tezgrid install-model llama3.2 --from ollama
  5. Go online

    Starts llama-server with your GGUF (or uses Ollama, or a server already running on the inference port), registers only the models you installed and sends a heartbeat every 15 seconds. tezgrid start is an alias.

    tezgrid up
  6. Stop

    Marks the node offline and stops the processes tezgrid up started. Ctrl+C does the same.

    tezgrid down

Endpoints

Click an endpoint for its request and response shape.

Inference

POST /v1/chat/completionsChat completions, streaming or not auth

Body

{
  "model": "Qwen2.5-1.5B-Instruct",
  "messages": [{"role": "user", "content": "Hello from TezGrid."}],
  "stream": false,
  "temperature": 0.7,
  "max_tokens": 512
}

Charged at prompt + 2 × completion tokens, in µ. The operator gets 94.5%. Empty wallet returns 402, missing key returns 401. Routing headers are listed under Routing headers.

GET /v1/modelsModels available on the network public

Response

{
  "object": "list",
  "data": [
    { "id": "llama3-8b", "object": "model", "owned_by": "tezgrid" }
  ]
}

Also GET /v1/models/legacy and GET /v1/provider/models?provider_id=…

GET /v1/pricingModeled marketplace pricing for reference models public

Economics from MarketplacePricingEngine for DeepSeek R1 Distill, Llama 8B and 70B, Qwen 72B and DeepSeek V3. These are modeled estimates, not a live per-request quote.

Accounts

POST /v1/auth/signupCreate a developer or provider account public

Body

{
  "name": "Ada",
  "email": "ada@example.com",
  "password": "••••••••",
  "role": "developer"
}

POST /v1/auth/login takes email, password and role.

GET /v1/developer/wallet/:developer_idWallet balance in micro-units auth

The developer ID is usually dev_<normalized_email>. The response includes wallet_balance_micro.

Related

GET /v1/developer/usage returns stored usage rows from MySQL. GET /v1/developer/tokens/hourly returns an hourly series for the console chart.

Network & providers

GET /v1/nodesRegistered and online provider nodes public

Returns { nodes: [...], mysql_enabled } with status, hardware, settled earnings and the telemetry fields the console uses.

POST /v1/provider/registerRegister a node and its inference endpoint auth

Body

{
  "provider_id": "<uuid>",
  "node_name": "studio-mac",
  "hardware": "Apple M4 Max",
  "device_type": "Apple Silicon",
  "inference_endpoint": "http://203.0.113.10:8080",
  "models": ["Qwen2.5-1.5B-Instruct"],
  "status": "ONLINE"
}

tezgrid up calls this for you. Most operators never call it directly.

GET /v1/network/metricsFleet-wide traffic and latency public

Includes a latency block with avg_router_ms, avg_ttft_ms, avg_decode_tps and success_rate_percent. Pass ?window=30m|24h|7d|30d for the traffic series that powers the network page.

Provider-scoped versions: GET /v1/provider/metrics and GET /v1/provider/dashboard.

GET /v1/diagnostics/latencyRouter and TTFT diagnostics, no inference needed public

Response

{
  "gateway_reachable": true,
  "online_providers": 12,
  "avg_router_latency_ms": 0.9,
  "avg_network_ttft_ms": 180,
  "avg_network_decode_tps": 42,
  "network_success_rate_percent": 99.2
}

Health

GET /healthGateway liveness public

Reports mysql_enabled, pilot readiness and any warnings. GET /ready is the readiness probe. Neither route is under /v1.