GLM-5.3 slots for agent teams.
A private GLM-5.3 endpoint with guaranteed capacity on GPUs we rent and control. Each server is allocated to a limited number of teams, and its capacity is never resold by the token. Your prompts, code and model outputs are processed only in the EU, and never stored.
Next batch starts Tuesday, December 1, 2026, or earlier if it fills. Nothing is charged until it confirms.
One agent, full speed, guaranteed.
A slot is one request running at full speed at any moment during its hours. Most agents send one request at a time, so a slot is roughly one agent working nonstop. Three slots, three agents in parallel. A slot cannot be over-used, so nobody else’s spike slows you down.
Guaranteed concurrency
A slot is one agent request running at full speed, any time during your hours. Not a rate limit you share with strangers, and not a queue behind someone else's batch job.
A flat monthly price
No per-token bill that grows with every long session. Your agents can read the whole repo, retry, and think for as long as they need. The price does not move.
EU only, never stored
Your prompts, code and model outputs are processed only in the EU, and never stored. No Chinese endpoint in the path, and a person who answers when you write.
Per slot, per month.
Choose the context your agents need, and whether you need the slot around the clock or during your working day.
| Context per slot | 24/7 | Business hours |
|---|---|---|
| 45K | $1,600 /mo | $600 /mo |
| 90Kmost agent teams | $2,400 /mo | $900 /mo |
| 128Klimited availability | $3,400 /mo | $1,300 /mo |
Per slot, per month. 24/7 is guaranteed at all times. Business hours is 08:00 to 20:00 on weekdays, in your own time zone. No minimum: one slot is fine. Monthly plans cancel with 30 days’ notice. Need 12 months or more, many slots, or a whole box? Talk to us.
What the same work costs per token.
Move the sliders to how long your agents actually work.
Estimate for one coding agent working nonstop during those hours: about 44K tokens of context per turn, mostly cached, roughly 1K tokens out, at GLM-5.3 prices from trusted US providers (about $1.40 in and $4.40 out per million tokens, cached input discounted). Agents that idle part of the time cost less per token. If yours mostly idle, per-token may suit you better, and we will say so.
Performance with every slot busy.
Measured 2026-10-06 on the hardware your slot runs on, with a real coding-agent loop over a real codebase: tool calls every turn, sessions compacting as they grow. Speeds are per agent with the box at its full slot count, not a single request on an empty machine.
Hosted in the EU. Nothing kept.
Your prompts, code and model outputs are processed only in the EU, and never stored.
The model runs on GPUs we rent in EU member states only, behind a gateway in Germany. The gateway records usage (token counts and cost) and nothing else: no prompts, no code, no outputs, no logs of them. A data processing agreement is available on request.
Account, billing and email services sit outside the model path entirely and never see your prompts or outputs. The categories we use are in the privacy policy, and the named list comes with the data processing agreement.
99% monthly uptime, backed by credits
If the service is available less than 99% of a month, you get a credit in proportion to the downtime on your next invoice. Planned maintenance is announced at least 48 hours ahead and kept short. The box restarts itself after a fault, and we are alerted the moment it does.
Your traffic is never sent to another provider to cover an outage. That is what keeps the privacy line above true.
Every server runs fully allocated.
Slots are sold in batches. Each batch has a start date and starts early if it fills, so you always know the latest date you will be live.
Reserve
Pick a context size, 24/7 or business hours, and how many slots. Nothing is charged, and you can cancel for free until the batch confirms.
The batch confirms
The next batch starts on Tuesday, December 1, 2026, or about 7 days after it fills if that comes first. We email you before anything is invoiced. The invoice follows, due in 7 days.
You're live
You get an API key for your private GLM-5.3 endpoint. It works with any OpenAI-compatible tool. Invoices come on the 1st of each month after that.
If the batch does not fill
Reservations close on Nov 20. If the batch has not filled by then, it extends once, to Dec 15, and anyone can cancel for free during that time. If it still has not filled, every reservation is cancelled and nothing is owed. There is no open-ended wait.
Before you reserve.
What exactly is a slot?
What happens outside business hours?
When does the next batch start?
When and how am I invoiced?
Can I change my slot count or size later?
How do I cancel?
What does "never stored" mean?
What if the server goes down?
Which model, exactly?
Does it work with my tools?
Can I test it first?
A whole box, many slots, or a longer term.
Dedicated capacity is a full GPU box for your company alone, from about $19,500 a month, with your own context size and engine settings. We also quote prepaid terms of 12 months or more, and large slot counts. A person replies, usually within a working day.
Flex slots, at about 25% less with a lower priority when the box is busy, are coming soon. Mention it in the form if you are interested.