An OpenAI-compatible API to open-weights models on our own hardware in Austria. Aliases that carry policy. Privacy in stages, up to an LLM-backed check. And every number on this page is measured nightly — not promised.
The smart part was never the model. It's what surrounds it: routing, policy, privacy, proof.
An alias is more than a model name — it bundles model, provider and privacy policy. Swap the model centrally; your application never changes.
Mask, pseudonymize, or go strict: an LLM on our own hardware double-checks for indirect identifiers before anything leaves the EU.
Our own hardware in Austria. No sub-processors, no third countries. External frontier models are reachable — but only through the masking layer.
Nightly benchmarks with published methodology. Detection recall for the privacy layer as a number, not an adjective.
If your software speaks OpenAI, it already speaks SmartLLM. Change one URL.
10M tokens free to test and set up. No card, no subscription.
One key, all models, all aliases. Usage-based, nothing else.
Call an alias, not a model version. Privacy is one parameter.
curl https://api.smartllm.at/v1/chat/completions \ -H "Authorization: Bearer $KEY" \ -d '{ "model": "frontier", "privacy": {"mode": "pseudonymize"}, "messages": [...] }'
Later, host it yourself: same API on your own hardware. When your box is running, you change one URL — done.
Vetted, benchmarked, and kept current. Aliases follow our recommendation — or pin a version and stay.
| Model | Alias | Context | Measured | Status |
|---|---|---|---|---|
| Qwen 3.5 397B | frontier | 256K | 26.5 t/s | loaded |
| Nemotron 3 Omni | — | 128K | 39.9 t/s | on-demand |
| Nemotron Cascade 2 | — | 256K | 40.2 t/s | on-demand |
| Phi 4 | — | 16K | — | on-demand |
| Qwen 3.6 35B | mid vision coder | 32K | 53.5 t/s | loaded |
All models, one price. 10M tokens free. No base fee, no subscription — usage only.