Status and availability

Measured values, not promises. We deliberately state no SLA percentage: a documented history says more than a pledge we might not be able to keep.

Component Right now
spark-1 + 2 (dgx-cluster)
vLLM · Qwen 3.5 397B
up
spark-3
vLLM · Qwen 3.6 35B
up
spark-4
Ollama · models on demand
up
spark3_nemotron
up

Last 30 days

Per compute node, sampled every five minutes. Planned maintenance counts as downtime here; we do not subtract it.

Node Available Samples
spark-1 + 2 (dgx-cluster)
vLLM · Qwen 3.5 397B, TP2 across both nodes
96.48 % 8.818
spark-3
vLLM · Qwen 3.6 35B and Nemotron 3.5 Lightning
99.95 % 8.850
spark-4
Ollama · all models loaded on demand
99.86 % 8.850
Platform overall
At least one node ready to serve
100.00 % 8.850
The last row is the number that matters to you: if a single node goes down, the fallback chain takes over and your request is still answered, just by a different model.

Requests

All customer requests of the last 30 days by outcome.

Requests total 14.401
Successful 14.263 · 99.04 %
Errors on our side (5xx) 129 · 0.90 %
Rejected as malformed (4xx) 9
We report the 5xx as well. They stem mostly from maintenance windows in which a model was being redeployed.

Collected by our own agent, which queries the services directly on the nodes every five minutes, and from the gateway request log. Both have been running without a gap since 02.04.2026.