Next batch: Reservations open, closes Nov 20
Private GLM-5.3 capacity for your agents

Need your agents’ data to stay in the EU? On Tium, it always does.

Every slot runs on GPUs in EU member states, behind a gateway in Germany. Your prompts, code and model outputs are processed only in the EU, and never stored.

Each slot is one GLM-5.3 agent working at full speed, for a flat monthly price instead of a per-token bill.

From $900 per month. Nothing is charged until your batch confirms.

Where your data goes

One route, all of it in the EU.

Your traffic is never sent to another provider, not even to cover an outage. There is no fallback outside the EU to switch on.

  1. 1. Your agent sends a request
    Gateway in Germany

    Checks your key and counts tokens. The request passes through memory and is not written anywhere.

  2. 2. The model runs
    GPUs in EU member states

    GLM-5.3 on servers we rent and control, allocated to a limited number of teams.

  3. 3. The answer comes back
    Same gateway, same route

    Streamed straight to your agent. No copy of the prompt, code or output is kept.

  4. 4. What we keep
    Usage numbers only

    Tokens, cost and timing, for billing and capacity. Never content.

For your security review

Hosted in the EU. Nothing kept.

Your prompts, code and model outputs are processed only in the EU, and never stored.

The model runs on GPUs we rent in EU member states only, behind a gateway in Germany. The gateway records usage (token counts and cost) and nothing else: no prompts, no code, no outputs, no logs of them. A data processing agreement is available on request.

Account, billing and email services sit outside the model path entirely and never see your prompts or outputs. The categories we use are in the privacy policy, and the named list comes with the data processing agreement.

99% monthly uptime, backed by credits

If the service is available less than 99% of a month, you get a credit in proportion to the downtime on your next invoice. Planned maintenance is announced at least 48 hours ahead and kept short. The box restarts itself after a fault, and we are alerted the moment it does.

Your traffic is never sent to another provider to cover an outage. That is what keeps the privacy line above true.

What you get

Capacity that is yours, not a shared queue.

Guaranteed concurrency

A slot is one agent request running at full speed, any time during your hours. Not a rate limit you share with strangers, and not a queue behind someone else's batch job.

A flat monthly price

No per-token bill that grows with every long session. Your agents can read the whole repo, retry, and think for as long as they need. The price does not move.

EU only, never stored

Your prompts, code and model outputs are processed only in the EU, and never stored. No Chinese endpoint in the path, and a person who answers when you write.

The maths

A busy agent is expensive per token.

Coding agents re-read large contexts on every turn. Move the sliders to your own hours.

Per-token, trusted providers (est.)
$2,864 /mo
One 90K business-hours slot
$900 /mo
Difference
$1,964 /mo

Estimate for one coding agent working nonstop during those hours: about 44K tokens of context per turn, mostly cached, roughly 1K tokens out, at GLM-5.3 prices from trusted US providers (about $1.40 in and $4.40 out per million tokens, cached input discounted). Agents that idle part of the time cost less per token. If yours mostly idle, per-token may suit you better, and we will say so.

How it works

Reserve now, pay when your batch confirms.

1

Reserve

Pick a context size, 24/7 or business hours, and how many slots. Nothing is charged, and you can cancel for free until the batch confirms.

2

The batch confirms

The next batch starts on Tuesday, December 1, 2026, or about 7 days after it fills if that comes first. We email you before anything is invoiced. The invoice follows, due in 7 days.

3

You're live

You get an API key for your private GLM-5.3 endpoint. It works with any OpenAI-compatible tool. Invoices come on the 1st of each month after that.

Questions

Before you reserve.

What exactly is a slot?
One request running at full speed at any moment, during the slot's hours, on GLM-5.3. An agent typically runs one request at a time, so one slot is roughly one agent working nonstop. If you need more at once, buy more slots: three slots means three requests in parallel.
What happens outside business hours?
A business-hours slot is off outside its hours, and requests get a clear error saying when it is back. You can turn on after-hours overage in your dashboard: your key keeps working on the same box when there is room, billed per token on your next invoice, up to a monthly cap you set. It is off until you turn it on, and it never goes past your cap.
When does the next batch start?
On its listed start date, or about 7 days after it fills if that comes first. Reservations for this batch close on Nov 20. If it has not filled by then, it extends once, to Dec 15, and you can cancel during that time. If it still has not filled, every reservation is cancelled and nothing is owed.
When and how am I invoiced?
Nothing is charged when you reserve. When the batch confirms we email you first, and the invoice follows at least 2 days later, due in 7 days. Invoices are in US dollars and are paid by bank transfer to the details printed on the invoice; card payment from the invoice link is also accepted. Prices exclude taxes: customers outside the US account for any VAT or GST themselves under the reverse charge. If the batch starts early, the first invoice covers only the days up to the 1st; after that, invoices come on the 1st of each month. Prepaid terms are one invoice for the whole term.
Can I change my slot count or size later?
Yes. Adding slots depends on room in your batch or the next one. Reducing takes effect at the next billing month with 30 days' notice.
How do I cancel?
Before the batch confirms, from the link in your confirmation email, free. After that, monthly plans cancel with 30 days' notice. Prepaid terms run to the end of the term.
What does "never stored" mean?
Your prompts, code and model outputs pass through memory on our gateway and GPU servers while a request runs, and are never written to disk, logged or kept. We store usage numbers (tokens, cost, timing) for billing and capacity planning, and nothing else.
What if the server goes down?
It restarts itself, and we are alerted at once. We commit to 99% uptime each month; below that, you get a credit in proportion to the downtime. We do not send your traffic to another provider while it recovers.
Which model, exactly?
GLM-5.3 from Z.ai, served at NVFP4 precision on current-generation NVIDIA data-centre GPUs, the format that hardware is built for. Context is set by your slot size: 45K, 90K or 128K tokens.
Does it work with my tools?
Yes, it is an OpenAI-compatible endpoint: Cline, Aider, OpenCode, Continue, your own scripts, anything that takes a base URL and a key. Setup guides are in the docs.
Can I test it first?
Yes. We run short paid pilots during your working hours, priced at what the GPU time costs us. Ask us.