Next batch: Reservations open, closes Nov 20

GLM-5.3 slots for agent teams.

A private GLM-5.3 endpoint with guaranteed capacity on GPUs we rent and control. Each server is allocated to a limited number of teams, and its capacity is never resold by the token. Your prompts, code and model outputs are processed only in the EU, and never stored.

Next batch starts Tuesday, December 1, 2026, or earlier if it fills. Nothing is charged until it confirms.

What a slot is

One agent, full speed, guaranteed.

A slot is one request running at full speed at any moment during its hours. Most agents send one request at a time, so a slot is roughly one agent working nonstop. Three slots, three agents in parallel. A slot cannot be over-used, so nobody else’s spike slows you down.

Guaranteed concurrency

A slot is one agent request running at full speed, any time during your hours. Not a rate limit you share with strangers, and not a queue behind someone else's batch job.

A flat monthly price

No per-token bill that grows with every long session. Your agents can read the whole repo, retry, and think for as long as they need. The price does not move.

EU only, never stored

Your prompts, code and model outputs are processed only in the EU, and never stored. No Chinese endpoint in the path, and a person who answers when you write.

Pricing

Per slot, per month.

Choose the context your agents need, and whether you need the slot around the clock or during your working day.

Context per slot24/7Business hours
45K$1,600 /mo$600 /mo
90Kmost agent teams$2,400 /mo$900 /mo
128Klimited availability$3,400 /mo$1,300 /mo

Per slot, per month. 24/7 is guaranteed at all times. Business hours is 08:00 to 20:00 on weekdays, in your own time zone. No minimum: one slot is fine. Monthly plans cancel with 30 days’ notice. Need 12 months or more, many slots, or a whole box? Talk to us.

Against per-token

What the same work costs per token.

Move the sliders to how long your agents actually work.

Per-token, trusted providers (est.)
$2,864 /mo
One 90K business-hours slot
$900 /mo
Difference
$1,964 /mo

Estimate for one coding agent working nonstop during those hours: about 44K tokens of context per turn, mostly cached, roughly 1K tokens out, at GLM-5.3 prices from trusted US providers (about $1.40 in and $4.40 out per million tokens, cached input discounted). Agents that idle part of the time cost less per token. If yours mostly idle, per-token may suit you better, and we will say so.

Measured

Performance with every slot busy.

1.0 s
median time to first token
p95 1.71 s
50 tok/s
per agent, 90K slots
24 tok/s with every 45K slot busy
96%
of each prompt reused, not reprocessed
so long agent sessions stay fast
100%
tool calls parsed
zero errors at every load level

Measured 2026-10-06 on the hardware your slot runs on, with a real coding-agent loop over a real codebase: tool calls every turn, sessions compacting as they grow. Speeds are per agent with the box at its full slot count, not a single request on an empty machine.

Privacy and uptime

Hosted in the EU. Nothing kept.

Your prompts, code and model outputs are processed only in the EU, and never stored.

The model runs on GPUs we rent in EU member states only, behind a gateway in Germany. The gateway records usage (token counts and cost) and nothing else: no prompts, no code, no outputs, no logs of them. A data processing agreement is available on request.

Account, billing and email services sit outside the model path entirely and never see your prompts or outputs. The categories we use are in the privacy policy, and the named list comes with the data processing agreement.

99% monthly uptime, backed by credits

If the service is available less than 99% of a month, you get a credit in proportion to the downtime on your next invoice. Planned maintenance is announced at least 48 hours ahead and kept short. The box restarts itself after a fault, and we are alerted the moment it does.

Your traffic is never sent to another provider to cover an outage. That is what keeps the privacy line above true.

How batches work

Every server runs fully allocated.

Slots are sold in batches. Each batch has a start date and starts early if it fills, so you always know the latest date you will be live.

1

Reserve

Pick a context size, 24/7 or business hours, and how many slots. Nothing is charged, and you can cancel for free until the batch confirms.

2

The batch confirms

The next batch starts on Tuesday, December 1, 2026, or about 7 days after it fills if that comes first. We email you before anything is invoiced. The invoice follows, due in 7 days.

3

You're live

You get an API key for your private GLM-5.3 endpoint. It works with any OpenAI-compatible tool. Invoices come on the 1st of each month after that.

If the batch does not fill

Reservations close on Nov 20. If the batch has not filled by then, it extends once, to Dec 15, and anyone can cancel for free during that time. If it still has not filled, every reservation is cancelled and nothing is owed. There is no open-ended wait.

Questions

Before you reserve.

What exactly is a slot?
One request running at full speed at any moment, during the slot's hours, on GLM-5.3. An agent typically runs one request at a time, so one slot is roughly one agent working nonstop. If you need more at once, buy more slots: three slots means three requests in parallel.
What happens outside business hours?
A business-hours slot is off outside its hours, and requests get a clear error saying when it is back. You can turn on after-hours overage in your dashboard: your key keeps working on the same box when there is room, billed per token on your next invoice, up to a monthly cap you set. It is off until you turn it on, and it never goes past your cap.
When does the next batch start?
On its listed start date, or about 7 days after it fills if that comes first. Reservations for this batch close on Nov 20. If it has not filled by then, it extends once, to Dec 15, and you can cancel during that time. If it still has not filled, every reservation is cancelled and nothing is owed.
When and how am I invoiced?
Nothing is charged when you reserve. When the batch confirms we email you first, and the invoice follows at least 2 days later, due in 7 days. You can pay by card or bank transfer. If the batch starts early, the first invoice covers only the days up to the 1st; after that, invoices come on the 1st of each month. Prepaid terms are one invoice for the whole term.
Can I change my slot count or size later?
Yes. Adding slots depends on room in your batch or the next one. Reducing takes effect at the next billing month with 30 days' notice.
How do I cancel?
Before the batch confirms, from the link in your confirmation email, free. After that, monthly plans cancel with 30 days' notice. Prepaid terms run to the end of the term.
What does "never stored" mean?
Your prompts, code and model outputs pass through memory on our gateway and GPU servers while a request runs, and are never written to disk, logged or kept. We store usage numbers (tokens, cost, timing) for billing and capacity planning, and nothing else.
What if the server goes down?
It restarts itself, and we are alerted at once. We commit to 99% uptime each month; below that, you get a credit in proportion to the downtime. We do not send your traffic to another provider while it recovers.
Which model, exactly?
GLM-5.3 from Z.ai, served at NVFP4 precision on current-generation NVIDIA data-centre GPUs, the format that hardware is built for. Context is set by your slot size: 45K, 90K or 128K tokens.
Does it work with my tools?
Yes, it is an OpenAI-compatible endpoint: Cline, Aider, OpenCode, Continue, your own scripts, anything that takes a base URL and a key. Setup guides are in the docs.
Can I test it first?
Yes. We run short paid pilots during your working hours, priced at what the GPU time costs us. Ask us.
Talk to us

A whole box, many slots, or a longer term.

Dedicated capacity is a full GPU box for your company alone, from about $19,500 a month, with your own context size and engine settings. We also quote prepaid terms of 12 months or more, and large slot counts. A person replies, usually within a working day.

Flex slots, at about 25% less with a lower priority when the box is busy, are coming soon. Mention it in the form if you are interested.