Public alpha · api.tezgrid.com

Inference on machines that already exist.

TezGrid runs AI models on computers people already own, like gaming PCs, Macs and office servers. If your app uses OpenAI, you switch by changing one line of code. If you own the computer, you keep 94.5% of what it earns.

Alpha software. Pricing and features will change. Read the paper

app.py1 line changed
from openai import OpenAI client = OpenAI(    base_url="https://api.openai.com/v1",    base_url="https://api.tezgrid.com/v1",    api_key=os.environ["TEZGRID_API_KEY"],) client.chat.completions.create(    model="Qwen2.5-1.5B-Instruct",    messages=[{"role": "user", "content": "Hi"}],    stream=True,)
Same SDK, same request shape, same streaming. python · node · curl
Network right now connecting to gateway…
—
Nodes online
—
Models serving
—
Tokens served, 24h
—
Avg. time to first token

01 How it works

One endpoint in front of a fleet you don't have to run.

The gateway handles auth, limits, routing and metering. Operators run the models on their own hardware. Your code only ever sees a standard OpenAI response.

  1. Your app sends a normal chat request

    Point any OpenAI client at api.tezgrid.com/v1. Streaming, models list and errors behave the way your SDK expects.

  2. The gateway picks a node

    Candidates are scored on latency, region, trust tier and price. Unhealthy nodes are probed and skipped, and a stalled request fails over to the next one.

  3. The node runs the model, the ledger settles

    Inference happens on the operator's machine via llama.cpp or Ollama. Every request is metered per token: 94.5% to the operator, 5.5% to the platform.

Your app OpenAI SDK TezGrid gateway auth · limits route · failover meter · settle GPU node NVIDIA · CUDA Mac node Apple Silicon CPU / server x86 · ROCm request out, tokens stream back on the same path

02 Two sides

A marketplace for idle inference.

Developers bring demand through an API they already know. Operators bring supply from hardware that is otherwise sitting idle.

For developers

Keep your SDK. Change the base URL.

Nothing to rewrite. Usage, keys and spend live in the console.

  • OpenAI-compatible chat completions with streaming
  • Per-key spend caps and hourly usage in the console
  • Optional region, latency and minimum-trust headers
  • Idempotency keys, so a retry doesn't charge you twice
Get an API key

For operators

Earn from a machine you already paid for.

Install the CLI, load a model, go online. No public port needed.

  • NVIDIA GPUs, Apple Silicon, AMD and desktop or server CPUs
  • tezgrid doctor tells you which models fit before you download
  • Outbound tunnel, so it works behind home NAT and firewalls
  • 94.5% of every settled charge, tracked per request
Run a node

03 Why it costs less

Fewer hands between the silicon and your request.

Most inference is bought, rented, repackaged and resold before it reaches you, and every layer takes a margin. TezGrid sends requests to machines that are already paid for, where the marginal cost is mostly electricity.

Typical API provider

3 markups stacked on top of the compute

Makes the chipChip vendor
Rents it by the hourCloud provider
Resells it per tokenAPI reseller
Pays all threeYour app

TezGrid

One 5.5% fee. The other 94.5% goes to the operator.

Already paid forIdle GPU, Mac or CPU
Routes and metersTezGrid gateway
Pays per tokenYour app
94.5%
of each settled charge goes to the operator
7–10×
lower modeled cost than typical list prices
4
hardware classes: CUDA, Apple, ROCm, CPU
1 line
to switch an existing OpenAI client

04 Privacy & trust

What we protect, and what we don't.

Prompts carry customer conversations, source code and internal plans. On a network of independent machines, privacy has to come in layers you can pick from, with the limits of each one written down.

  • No inbound port

    Nodes connect out to the gateway over a persistent encrypted channel. Operators never expose an inference port to the internet.

  • Encrypted per request

    Prompts and responses are envelope-encrypted on the tunnel hop between gateway and node, so they don't cross the public internet as plain text.

  • Trust tiers

    Every response tells you which tier served it. Send X-TezGrid-Min-Trust to refuse anything below the tier you need.

  • Confidential tier

    For work the host must not read: only nodes on confidential hardware with an active sealed connection are eligible.

  • One machine, one account

    Each physical host is fingerprinted and bound to one operator account and one location, which keeps duplicate and spoofed nodes out.

  • Baseline on every call

    API keys, rate and spend limits, and request logs that record IDs and timings, not prompt text.

05 API

The API you already use.

Chat completions and models, OpenAI-shaped. TezGrid-specific behaviour is opt-in through headers, so plain OpenAI code keeps working.


          
  • POST /v1/chat/completions
    Streaming or not. Errors come back as 401 for a missing key and 402 for an empty wallet.
  • GET /v1/models
    Models currently served by online nodes.
  • X-TezGrid-Min-Trust
    Refuse nodes below a trust tier, up to Confidential.
  • Idempotency-Key
    Safe retries: the same key is never billed twice.
  • X-TezGrid-Provider-Region
    Response header. Also returned: router latency, TTFT and failover attempts.

Full reference in the API docs.

06 Pricing

Per token, priced by supply.

These are modeled marketplace rates, not a fixed quote. Real prices move with model, region and how much capacity is online.

ModelQuantMemoryTezGrid, modeledTypical listDifference
Llama 3.1 8BQ4_K_M6 GB$0.02 / 1M$0.20 / 1M10× lower
DeepSeek R1 Distill Qwen 7BQ4_K_M5 GB$0.02 / 1M$0.20 / 1M10× lower
Llama 3.1 70BQ4_K_M40 GB$0.12 / 1M$0.90 / 1M7.5× lower
Qwen 2.5 72BQ4_K_M42 GB$0.14 / 1M$0.95 / 1M6.8× lower
DeepSeek V3MoE, 37B active96 GB$0.22 / 1M$1.50 / 1M6.8× lower

"Typical list" means published per-token rates for comparable open models from major hosts. During the alpha the gateway settles a flat 1 µ per prompt token and 2 µ per completion token ($1 and $2 per million) on every model. Wallet top-ups are not open yet; settled charges and the operator split are visible per request in the console.

07 Run a node

From idle to online in four commands.

The installer adds the CLI and llama.cpp's llama-server, then scans the machine. Already use Ollama? That works too. You'll know which models fit before you download any weights.

  1. Install

    One command per platform. Linux, macOS and Windows. Includes llama-server; Ollama works too.

  2. Check the machine

    tezgrid doctor reports CPU, RAM, GPU, disk and runtimes. tezgrid recommend lists models that fit.

  3. Install a model

    GGUF weights from the catalog, Ollama, or your own URL for closed models.

  4. Go online

    tezgrid up registers the node and keeps a heartbeat. tezgrid down takes it offline.

Earnings depend on demand in your region. Please read the installer before piping it to a shell. Full setup guide


          

Earnings estimate

Net monthly after electricity, for the hardware this browser detects.

Tariff preset detecting location…

Using the scenario's modeled probability.

Expected net per monthmodeled
—

Operator revenue minus power cost

Theoretical capacity (100% busy)—
Expected operator revenue—
Power, — kWh—
Ledger settlement—
Network demand
—
Node selection
—
Work received
—
Power draw
—

Estimates come from the gateway's earnings model and your inputs. Actual earnings depend on routing, uptime and real demand, and are not guaranteed.

→ Go deeper

Read how routing, metering and trust actually work.

The research paper explains how a machine is chosen, what happens when one fails mid-answer, how every token is counted, and what we cannot promise yet. The docs cover every endpoint.