Working paper 01Version 1.0Alpha
TezGrid: AI inference on computers people already own
How we route requests, keep count of every token, and decide what to trust in a network of machines that we do not own. Written from inside the alpha, with the open questions left in.
Contents
Abstract
A lot of computing power around the world, and in India, sits unused for most of the day. Gaming PCs, office workstations and small server rooms are already paid for and mostly idle. At the same time, developers here pay in dollars to use AI models running in data centres on the other side of the world. TezGrid tries to connect the two. A developer calls one API that follows the same format as OpenAI's. Behind it, our gateway picks a suitable machine run by an independent operator, sends the request over an encrypted tunnel, counts the tokens and splits the charge: 94.5% to the operator and 5.5% to us.
This paper explains how that works in practice. It covers how a machine is chosen, what happens when one fails halfway through an answer, how money is counted without rounding errors, and what we can and cannot promise about privacy. We also write down, plainly, what we have not proved yet. The biggest open question is not technical. It is whether there will be enough paying demand to keep the operators' machines busy.
01Why we built this
Take a developer in Pune who is building a customer support assistant for a mid-sized business. She has three practical problems.
The first is cost. Every message her users send is billed per token, in dollars, by a company abroad. The bill grows with usage, and it also moves with the rupee. A small team feels both.
The second is distance. A round trip from Mumbai to a data centre on the US east coast commonly takes 200 milliseconds or more, and that is before the model has even started thinking. For a chat window where the user is waiting for the first word to appear, this delay is felt.
The third is choice. If her provider changes a price, retires a model or has a bad day, she has very little room to move.
Now look at the other side. Visit an engineering college hostel, a small animation studio in Hyderabad or an IT office in Noida after 7 pm. There are machines with good graphics cards and plenty of memory that are switched on and doing nothing. A gamer's RTX 3090 sits idle from 9 to 6 while he is at work. An office's workstations sit idle all night. These machines are already bought and already connected. Nobody is asking them to do anything.
People have tried to rent out idle machines before. Broadly, the attempts fall into two groups. Some sell raw machine time by the hour, which means the developer has to set up the software, load the models and handle failures on their own. Others are built around crypto tokens and wallets, which puts off most ordinary developers and most ordinary owners. Neither is what our developer in Pune wants. She wants to change one line in her code. The gamer wants to install something, leave it running and see the earnings add up.
So the question we started with was a simple one. Can a person with a decent PC in Indore serve a developer in Bengaluru, with the developer changing nothing except the address of the API, and with the money adding up correctly on both sides? The rest of this paper is our answer so far.
02The idea in one picture
There are three parties: the developer who sends requests, the operator who runs a machine, and the gateway in the middle. The gateway is the only part we run ourselves. We do not own any of the machines that run models.
- Uses any OpenAI client library
- Changes only the base URL and key
- Pays per token from a wallet
- Checks the key and the wallet
- Picks a machine and watches it
- Counts tokens and settles the charge
- Runs llama.cpp or Ollama
- Opens no ports to the internet
- Keeps 94.5% of what it serves
For the developer, TezGrid looks like any other AI API. For the operator, it looks like a background program that occasionally gets work. The interesting part is the gateway, because it has to make decisions about machines it does not control: which one to trust with a request, how long to wait, and what to do when one goes quiet.
03The life of one request
It is easiest to understand the system by following a single request from start to finish.
-
Your app sends a normal request
The app sends a chat completion request to
api.tezgrid.com/v1with its API key. Code written for OpenAI works as it is, once the base URL and key are changed. -
The key and the wallet are checked first
The gateway checks the key, which we store only as a hash. It then works out the most this request could cost. Prompt length is estimated at roughly one token for every four characters, and the output is taken at whatever
max_tokensthe app asked for (512 if it did not say, and never more than 8,192). If the wallet cannot cover that amount, the request is refused before any machine is touched. Each key can also have its own spending cap. -
A shortlist is made
The gateway looks at every machine that is online and has the requested model. It drops the ones that cannot take this particular request. Section 4 explains how.
-
The shortlist is ranked
Each remaining machine gets a score based on speed, distance, current load and track record. The ranking also decides who gets tried second and third if the first choice fails.
-
The request goes down the tunnel
The operator's machine keeps an outbound connection open to the gateway, so the request travels down that connection. Nothing on the internet can reach the operator's machine directly.
-
The gateway waits for the first word
If no first token arrives within 6 seconds, the gateway gives up on that machine and moves to the next one on the list. For the fast model alias the limit is 4 seconds, and for the large-model alias it is 15.
-
The answer streams back
Tokens flow back to the app as they are generated. Response headers tell the developer which region served the request, how long routing took and how many retries were needed, so nothing about the path is hidden.
-
The charge is settled
When the answer is complete, the gateway counts prompt and completion tokens, works out the charge, deducts it from the wallet and records the operator's share. All of this goes into one row in the database (Section 6).
04How a machine is chosen
Choosing a machine is the heart of the system. A bad choice means a slow answer, a failed request or a node that never gets work. We do it in two passes: first we remove machines that should not be considered at all, then we score the rest.
4.1 First, who is even allowed
A machine is removed from the shortlist if any of these is true:
- It does not have the model, or a model of the right size class when the developer asked for an alias.
- The request will not fit. We check the prompt plus the requested output against the machine's context length and memory.
- It lacks a feature the request needs, such as tool calling or JSON output.
- More than 30% of its recent requests failed.
- The developer set a hard rule it does not meet: a required region, a data residency requirement, a maximum network delay, or a minimum trust level.
4.2 Then, the score
Every machine that survives gets a score. The table below lists what goes into it, in plain words.
| Signal | What we look at | How it moves the score |
|---|---|---|
| Model already loaded | Is the model already in memory, or must it be loaded from disk first? | A large bonus if loaded. For speed-first requests, a machine that would have to load the model is pushed well down. |
| Current load | Requests already running and requests waiting in its queue | Each one pulls the score down. |
| Time to first token | How long, on past requests, the first word took to arrive | Lower is better. Counted more heavily for streaming chats, where a person is waiting. |
| Generation speed | Tokens per second while writing. For long prompts, also how fast it reads the prompt. | Faster is better. Counts for more when a long answer is asked for. |
| Distance | The round trip measured by our own health checks, or a regional estimate if there is no measurement yet | Closer is better. A measured figure earns a small extra bonus over an estimate. |
| Track record | The share of recent requests that failed, once there are at least 5 to go on | Steady machines get a bonus, and each failure costs points. |
| Trust level | Standard, Verified, Attested or Confidential (Section 8) | Higher levels get a modest bonus. |
| Price | The machine's price per token | Matters most when the request asks for the cheapest option. |
| Newness | Fewer than 5 measured requests so far | A small boost, so new machines get a fair chance to prove themselves. |
Not every request wants the same thing. A chat where a user is watching the screen cares most about the first word. A background job summarising a long document cares more about steady writing speed. The router therefore has a few objectives (cheapest, fastest response, most throughput, most reliable, and balanced). They all use the same signals with different weights. Streaming requests are treated as speed-first automatically, so most developers never need to set anything.
4.3 Not always the top scorer
If the router always picked the single best machine, that machine would be overloaded while the second best, almost as good, would sit idle and earn nothing. New machines would never get measured. So when several machines score within 5% (or 5 points) of the best, the router picks one of them, spreading requests across the group. The choice is based on the model, the size of the prompt and the exact time of the request, which in practice works like a fair lottery among near-equals.
This matters for the business as much as for speed. An operator whose machine is nearly as good as the best one should still see some work. Otherwise they leave, and the network gets thinner.
4.4 Asking for "good enough" instead of a model name
Developers often do not care which exact model answers, only roughly how big and fast it is. For them we offer four aliases. The gateway maps each alias to whatever suitable model is loaded on a healthy machine.
| Alias | What you get | Good for |
|---|---|---|
tezgrid-fast | Small models, up to about 3 billion parameters | Quick replies, classification, simple chat |
tezgrid-balanced | Models of about 7 to 8 billion parameters | Most general chat and writing |
tezgrid-quality | 13 billion parameters and above | Harder reasoning, longer answers |
tezgrid-private | Only machines at the Confidential trust level | Requests where the developer wants the strictest available setting |
4.5 Keeping a conversation on one machine
In a multi-turn chat, every new message resends the whole conversation. A machine that served the previous turn often still has that conversation in its prompt cache, so it can skip re-reading it and start answering sooner. The gateway therefore keeps a conversation on the same machine for about 30 minutes, as long as that machine stays healthy. Developers can control this with a session header, or leave it to the API key.
05When a machine fails
Machines at home fail in very ordinary ways. The power goes. The Wi-Fi drops. Someone starts a game. Windows decides to restart for an update. We designed TezGrid on the assumption that any machine can disappear at any moment, and that this must not become the developer's problem.
5.1 Health checks
The gateway checks every machine every 20 seconds. After 3 failed checks in a row, the machine is marked degraded and gets fewer requests. After 6, about two minutes of silence, it is marked offline and gets none. A single successful check brings it back. The same checks measure the round trip between the gateway and the machine, which feeds the distance signal in Section 4.
5.2 Before the first word
If a machine refuses the request, errors out or does not produce a first token in time, the gateway moves to the next machine on the ranked list. By default it tries three machines before telling the developer the request failed (this can be set between one and five). The developer sees one request and one answer. A response header records how many attempts it took.
5.3 In the middle of an answer
This is the harder case. Suppose a machine has streamed half a paragraph and then goes silent. Starting again from scratch would show the user a repeated or jumbled answer. Instead, the gateway sends the next machine the original conversation plus the text written so far, and asks it to carry on from there. The user sees the answer continue.
We should be honest about the limits. The second machine may run a slightly different build of the same model, so the wording can shift a little at the join. For most chat use this is acceptable. For tasks where every word must come from a single run, it is not ideal, and a per-request option to switch it off is on our list.
5.4 Failing loudly on the operator's side
Early on, we learnt that a machine which looks online but is not is worse than an error. An operator would see "online" on their screen while being connected only to a test gateway on their own computer, and would then wonder why no work ever came. Now, if tezgrid up cannot register with the real gateway, it says so clearly and stops, unless the operator has deliberately asked for a local demo.
06Counting money without losing a paisa
A single request costs a fraction of a rupee. Millions of requests add up to real money. If each request is rounded a little carelessly, the totals drift, and then the developer's bill, the operator's earnings and our own fee stop agreeing with each other. In a marketplace, that is a trust problem, not just an accounting one.
6.1 Whole numbers only
Computers store decimal fractions approximately. Ask most programming languages for 0.1 + 0.2 and the answer comes back as 0.30000000000000004. That is harmless once, but it is not acceptable in a ledger. So TezGrid never stores money as a decimal fraction. Every amount is a whole number of micro-dollars, where one micro-dollar (written µ) is one millionth of a US dollar.
Our alpha prices are 1 µ for every prompt token and 2 µ for every completion token. That works out to $1 per million prompt tokens and $2 per million completion tokens. With these prices, the whole calculation needs only addition, one multiplication and one division of whole numbers:
operator = floor(charge × 9,450 ÷ 10,000) // 94.5%, rounded down
platform = charge − operator // the remaining 5.5%
Because the platform fee is worked out as "whatever is left", the two shares always add up exactly to what the developer paid. The rounding remainder, always less than one micro-dollar, falls on the platform's side.
6.2 A worked example
Let us take one ordinary request: a prompt of 1,200 tokens and an answer of 400 tokens.
6.3 Where it is written down
Every request becomes one row in a usage_records table: the token counts, the charge, the operator's share, the platform fee and how long the request took. The developer's usage page and the operator's earnings page both read from these same rows, so the two sides cannot see different numbers.
Credits added to a wallet go through a separate double-entry ledger. Every credit has a matching debit from an "alpha credits" account, so for any balance we can show exactly where it came from.
During the alpha, developers cannot buy credit with a card. We grant credit by invitation. Operator earnings are recorded but not yet paid out. Before we can pay operators, we have to build bank and UPI transfers and handle tax properly (TDS, for example, for operators in India). We will publish the payout terms before any payout happens.
07What an operator can earn, honestly
This is the question every operator asks first, and the one most projects answer too generously. We will show the arithmetic and the assumptions, so you can change any of them and redo the sum.
7.1 Assumptions
- Speed is taken from our hardware catalogue. These are estimates for small models serving one request at a time, not measurements from our network. Real speed depends a great deal on the size of the model and how it is compressed.
- All tokens are counted as completion tokens at $2 per million. Prompt tokens are cheaper but are also read much faster, so leaving them out keeps the sum simple without changing the picture much.
- The operator keeps 94.5%.
- ₹88 to the dollar, and electricity at ₹8 per unit (kWh). Electricity is counted only for the hours the machine is actually serving.
All rupee figures in the table are for one busy hour.
| Machine | Tokens/s | Power | Earned | Electricity | Net |
|---|---|---|---|---|---|
| Ryzen 9 desktop, CPU only | 38 | 120 W | ₹22.8 | ₹0.96 | ₹21.8 |
| Apple M2 / M3 Pro | 45 | 50 W | ₹26.9 | ₹0.40 | ₹26.5 |
| NVIDIA RTX 3090 | 75 | 350 W | ₹44.9 | ₹2.80 | ₹42.1 |
| NVIDIA RTX 4090 | 110 | 350 W | ₹65.9 | ₹2.80 | ₹63.1 |
To check one row: an RTX 4090 at 110 tokens a second writes 396,000 tokens in an hour. At $2 per million that is $0.79, of which the operator keeps $0.75, or about ₹65.9. The card draws 0.35 units of electricity in that hour, which costs ₹2.80.
7.2 The number that really matters: how busy
Those are earnings for a busy hour. No machine is busy every hour. What an operator actually takes home depends almost entirely on utilisation, the share of the month their machine spends serving requests.
Two things stand out. First, a good graphics card earns roughly three times what a strong desktop processor does, so GPUs will carry most of the network's work. Second, and more important, the difference between 5% and 50% busy is ten times the money. Utilisation matters far more than which machine you own.
Today, utilisation on TezGrid is close to zero, because the alpha does not yet have paying developers outside our own team. So treat the earnings calculator on our site as a ceiling, not a promise. The real constraint on operator income is demand, and building demand is our job, not the operator's.
In the current build, the CLI starts the model server with GPU offload switched off unless the operator sets TEZGRID_NGL. Without it, even a machine with a strong graphics card runs on the processor alone and earns at desktop-CPU rates. Making GPU offload automatic is near the top of our list.
08Privacy and trust, without the marketing
This is the section where projects like ours usually oversell. We will try not to. Some risks are covered well today. Some are not covered at all, and we would rather you hear that from us.
8.1 What we do protect
- Traffic in transit. Connections to the gateway use HTTPS.
- The operator's home network. Operators do not open any ports. Their machine connects out to the gateway and keeps that tunnel open, so the home router stays closed and nobody on the internet can reach the model server directly.
- The tunnel itself. A request can be sealed for one specific machine using X25519 and XSalsa20-Poly1305, the same construction as libsodium's sealed boxes. Nothing between the gateway and that machine can read it.
- Fake machines. Machines sign their registration and heartbeat messages with a key and a timestamp, so nobody can pretend to be a machine or replay its old messages.
- One machine, one account. Each machine is identified by a hash of its hardware identifiers. If the same machine tries to join under a different email, it is refused. Its location is taken from its public IP address, not from a setting the operator can edit.
- Accounts. Passwords are hashed with Argon2id. API keys are stored only as hashes.
- Logs. Usage records and logs store token counts and timings. They do not store the text of prompts or answers.
8.2 What we do not protect
To run a model, the operator's machine must decrypt the request. Someone who controls their own hardware can, with effort, look at what passes through it. Our terms forbid this, but software alone cannot prevent it today. Please do not send passwords, payment details, health records or anyone's personal data through TezGrid.
- The Confidential label is a claim, not a proof. Today it is based on what the operator declares, for example that the machine runs inside AMD SEV-SNP or Intel TDX. We do not yet verify the hardware's attestation report. Treat it as a preference, not a guarantee.
- We cannot yet prove which model ran. Our speed and failure measurements catch machines that are broken or slow. They do not catch a machine that quietly runs a smaller model than it claims. We plan to add spot checks with known prompts.
| Risk | Covered today? | How, or what is planned |
|---|---|---|
| Someone on the network reads the traffic | Yes | HTTPS to the gateway, sealed tunnel to the machine |
| Operator's home network is exposed | Yes | Outbound tunnel, no open ports |
| A fake machine joins or impersonates another | Yes | Signed messages and one-machine-one-account binding |
| Operator reads prompts on their own machine | No | Terms only today; verified hardware attestation planned |
| Machine runs a different model than it claims | No | Spot checks with known prompts planned |
| The gateway itself goes down | Partly | Single region today; a second region is planned |
09Getting a machine online
Supply only grows if joining is easy. Our working rule is that an operator should not need to know what a port is. Joining takes five steps:
- Install. One command in PowerShell on Windows, or in a terminal on Linux and macOS. The installer adds the
tezgridprogram and the model server,llama-serverfrom the open-source llama.cpp project. On Windows and most Linux PCs it uses a Vulkan build, which works across NVIDIA, AMD and Intel graphics cards. On Macs it uses Metal. Operators who already use Ollama can use that instead. - Check the machine.
tezgrid doctordetects the processor, memory and graphics card and suggests models that will fit. - Sign in.
tezgrid signuportezgrid logincreates or opens the account and binds this machine to it. - Add a model.
tezgrid install-modelpicks a compressed version of the model that suits the hardware, checks there is enough disk space before downloading, and validates the file before using it. - Go online.
tezgrid upstarts the model server, opens the tunnel and registers with the gateway. The region is worked out from the public IP address and the system time zone, with nothing to configure.
The release pipeline that builds these installers for Windows, Linux and macOS is ready. At the time of writing, the first public release has not yet been published.
10How this compares
TezGrid sits between a few familiar kinds of services. The table is a simplification, and each group contains very different companies, but it shows where we fit.
| Cloud model APIs | API aggregators | GPU rental markets | TezGrid | |
|---|---|---|---|---|
| Who owns the hardware | The provider | Several cloud providers | People and companies renting machines out | Independent operators |
| What the developer buys | Tokens | Tokens, routed across providers | Machine time, usually by the hour | Tokens |
| Developer setup | Their SDK or API | Change the base URL | Set up the software, models and scaling yourself | Change the base URL |
| Supplier setup | Not applicable | Not applicable | List a machine, sometimes with a crypto wallet | Run the installer, then tezgrid up |
| Where prompts are processed | Provider data centres | Provider data centres | The rented machine | The operator's machine |
| Maturity | Mature | Mature | Varies | Alpha |
For developers, TezGrid behaves like an aggregator: one API in front of many sources. On the supply side, it is closer to a homestay platform, where ordinary people rent out something they already own and the platform handles bookings, payments and trust. We are not trying to compete with large cloud providers on the biggest frontier models. We focus on open models from small to mid size, where ordinary hardware does a good job and where distance and price matter.
11What we have not proved yet
A paper like this is only useful if it is clear about its gaps. Here are ours.
- Demand. No developer outside our team is paying for TezGrid yet. Until that changes, everything about operator income is theory.
- Real-world latency. Our response headers measure routing time, first-token time and retries on every request. But we have not yet published a study with real machines spread across different cities. The next paper should contain one.
- Speeds. The speeds in Section 7 are catalogue estimates, not our own measurements. We have left out benchmark figures that we could not reproduce from our own code.
- A single gateway. The gateway runs in one region. If it goes down, the whole network stops, even though every operator machine is fine.
- Quality varies. Two machines serving "the same" model may use different compressed versions of it, and answers can differ slightly between them.
- Payouts, invoices and tax. None of these are built yet.
- Tests are not users. The code is 12 Rust packages, about 37,600 lines, with 266 automated tests. Tests show that the code does what we meant it to do. They do not show that the network holds up under real users with real home internet connections.
12What comes next
In the order we intend to do them:
- Keep the hosted gateway reliably online, and publish the first installer release.
- Turn on GPU offload by default, so graphics cards are actually used.
- Bring on the first few paying developers as design partners, and track requests per day, money spent and utilisation by region.
- Publish a measured latency study from real machines in several Indian cities.
- Build operator payouts over UPI and bank transfer, with proper tax handling.
- Add spot checks that test whether a machine runs the model it claims, and verify hardware attestation for the Confidential level.
- Run a second gateway region, so that one outage does not stop the network.
13Closing
Most of the technology described here already works. Requests are routed, machines that fail are skipped, and the money is counted correctly down to the last micro-dollar. What remains is mostly ordinary business work: finding developers who will pay, and keeping machines busy enough that their owners stay.
We will publish the real numbers as they come in: requests, spending, utilisation and latency. If they are bad, we will publish those too.
Appendix ARequest and response headers
Developers do not need any of these to get started. They exist for people who want control over where a request runs, or visibility into how it was handled.
| Header | Direction | What it does |
|---|---|---|
X-TezGrid-Region | Request | Prefer machines in this region, without ruling others out |
X-TezGrid-Required-Region | Request | Only use machines in this region |
X-TezGrid-Max-Network-Latency-Ms | Request | Skip machines whose network round trip is above this many milliseconds |
X-TezGrid-Data-Residency | Request | Only use machines that match this residency requirement |
X-TezGrid-Min-Trust | Request | Minimum trust level, for example confidential |
X-TezGrid-Sticky-Session | Request | Keep this conversation on the same machine |
X-TezGrid-Provider-Region | Response | Region of the machine that answered |
X-TezGrid-Router-Latency-Ms | Response | Time the gateway spent choosing a machine |
X-TezGrid-Failover-Attempts | Response | How many machines were tried |
X-TezGrid-TTFT-Ms | Response | Time to first token, for non-streaming requests |
The full list, with examples, is in the API reference.
Appendix BGlossary
- Token
- A piece of text, usually a short word or part of a word. Models read and write text in tokens, and pricing is per token. In English, 1,000 tokens are roughly 750 words.
- Prompt and completion
- The prompt is what you send to the model. The completion is what it writes back. Completion tokens cost more because generating text takes more work than reading it.
- Inference
- Running a trained model to get an answer, as opposed to training the model in the first place.
- Time to first token
- How long you wait before the first word of the answer appears. It is the delay a user feels most.
- Quantisation
- Compressing a model so it needs less memory, at a small cost in quality. It is how large models fit on ordinary graphics cards.
- GGUF
- A file format for models used by llama.cpp, the open-source program our operators run.
- Operator
- A person or company that runs a machine on TezGrid and earns from the requests it serves.
- Gateway
- Our server in the middle. It receives requests, picks machines, and keeps the accounts.
- Utilisation
- The share of time a machine spends actually serving requests.
- Micro-dollar (µ)
- One millionth of a US dollar. All TezGrid amounts are stored as whole numbers of micro-dollars.
ReferencesFurther reading
- ggml-org. llama.cpp: LLM inference in C/C++. github.com/ggml-org/llama.cpp
- Ollama. Run large language models locally. ollama.com
- OpenAI. API reference: Chat completions. platform.openai.com/docs/api-reference/chat
- libsodium. Sealed boxes. doc.libsodium.org
- A. Biryukov, D. Dinu, D. Khovratovich, S. Josefsson. RFC 9106: Argon2 memory-hard function for password hashing and proof-of-work applications. IETF, 2021. rfc-editor.org/rfc/rfc9106
- AMD. Secure Encrypted Virtualization (SEV). amd.com/en/developer/sev.html
- Intel. Intel Trust Domain Extensions (Intel TDX). intel.com
- TezGrid. API reference, Alpha terms and Privacy notice. API reference · Terms · Privacy