Need your agents’ data to stay in the EU? On Tium, it always does.
Every slot runs on GPUs in EU member states, behind a gateway in Germany. Your prompts, code and model outputs are processed only in the EU, and never stored.
Each slot is one GLM-5.3 agent working at full speed, for a flat monthly price instead of a per-token bill.
From $900 per month. Nothing is charged until your batch confirms.
One route, all of it in the EU.
Your traffic is never sent to another provider, not even to cover an outage. There is no fallback outside the EU to switch on.
- 1. Your agent sends a requestGateway in Germany
Checks your key and counts tokens. The request passes through memory and is not written anywhere.
- 2. The model runsGPUs in EU member states
GLM-5.3 on servers we rent and control, allocated to a limited number of teams.
- 3. The answer comes backSame gateway, same route
Streamed straight to your agent. No copy of the prompt, code or output is kept.
- 4. What we keepUsage numbers only
Tokens, cost and timing, for billing and capacity. Never content.
Hosted in the EU. Nothing kept.
Your prompts, code and model outputs are processed only in the EU, and never stored.
The model runs on GPUs we rent in EU member states only, behind a gateway in Germany. The gateway records usage (token counts and cost) and nothing else: no prompts, no code, no outputs, no logs of them. A data processing agreement is available on request.
Account, billing and email services sit outside the model path entirely and never see your prompts or outputs. The categories we use are in the privacy policy, and the named list comes with the data processing agreement.
99% monthly uptime, backed by credits
If the service is available less than 99% of a month, you get a credit in proportion to the downtime on your next invoice. Planned maintenance is announced at least 48 hours ahead and kept short. The box restarts itself after a fault, and we are alerted the moment it does.
Your traffic is never sent to another provider to cover an outage. That is what keeps the privacy line above true.
Capacity that is yours, not a shared queue.
Guaranteed concurrency
A slot is one agent request running at full speed, any time during your hours. Not a rate limit you share with strangers, and not a queue behind someone else's batch job.
A flat monthly price
No per-token bill that grows with every long session. Your agents can read the whole repo, retry, and think for as long as they need. The price does not move.
EU only, never stored
Your prompts, code and model outputs are processed only in the EU, and never stored. No Chinese endpoint in the path, and a person who answers when you write.
A busy agent is expensive per token.
Coding agents re-read large contexts on every turn. Move the sliders to your own hours.
Estimate for one coding agent working nonstop during those hours: about 44K tokens of context per turn, mostly cached, roughly 1K tokens out, at GLM-5.3 prices from trusted US providers (about $1.40 in and $4.40 out per million tokens, cached input discounted). Agents that idle part of the time cost less per token. If yours mostly idle, per-token may suit you better, and we will say so.
Reserve now, pay when your batch confirms.
Reserve
Pick a context size, 24/7 or business hours, and how many slots. Nothing is charged, and you can cancel for free until the batch confirms.
The batch confirms
The next batch starts on Tuesday, December 1, 2026, or about 7 days after it fills if that comes first. We email you before anything is invoiced. The invoice follows, due in 7 days.
You're live
You get an API key for your private GLM-5.3 endpoint. It works with any OpenAI-compatible tool. Invoices come on the 1st of each month after that.