Status and availability
Measured values, not promises. We deliberately state no SLA percentage: a documented history says more than a pledge we might not be able to keep.
| Component | Right now |
|---|---|
|
spark-1 + 2 (dgx-cluster)
vLLM · Qwen 3.5 397B
|
up |
|
spark-3
vLLM · Qwen 3.6 35B
|
up |
|
spark-4
Ollama · models on demand
|
up |
|
spark3_nemotron
|
up |
Last 30 days
Per compute node, sampled every five minutes. Planned maintenance counts as downtime here; we do not subtract it.
| Node | Available | Samples |
|---|---|---|
|
spark-1 + 2 (dgx-cluster)
vLLM · Qwen 3.5 397B, TP2 across both nodes
|
96.48 % | 8.818 |
|
spark-3
vLLM · Qwen 3.6 35B and Nemotron 3.5 Lightning
|
99.95 % | 8.850 |
|
spark-4
Ollama · all models loaded on demand
|
99.86 % | 8.850 |
|
Platform overall
At least one node ready to serve
|
100.00 % | 8.850 |
The last row is the number that matters to you: if a single node goes down, the fallback chain takes over and your request is still answered, just by a different model.
Requests
All customer requests of the last 30 days by outcome.
| Requests total | 14.401 |
| Successful | 14.263 · 99.04 % |
| Errors on our side (5xx) | 129 · 0.90 % |
| Rejected as malformed (4xx) | 9 |
We report the 5xx as well. They stem mostly from maintenance windows in which a model was being redeployed.
Collected by our own agent, which queries the services directly on the nodes every five minutes, and from the gateway request log. Both have been running without a gap since 02.04.2026.