
SovInfra runs open-weight AI models on GPUs in Europe and sells access by the token. You point any OpenAI-compatible client at its API, pick a model or let its router pick, and your data stays in the EU, with no retention and no training on it.
It's a small French company answering a question I hear more and more from European customers: where does my data go when I call an AI model, and who can reach it?

| Company | FPLC SAS, Paris, France. No US parent company |
| Product | SovAPI, an OpenAI-compatible inference API, plus a web console |
| Models | Qwen 3.8-27B and Gemma 4 31B on NVIDIA H200, Whisper Large V3, BGE-M3, Kokoro |
| Price | $0.12 per million tokens in, $0.38 out. 1 billion free tokens, no card needed |
| Hosting | GPUs in the Czech Republic, orchestration in France, backup in Germany |
| Data | No logging of content, no retention, no training on customer data |
What it offers
- 2 main language models. Qwen 3.8-27B in FP8 for code, reasoning, vision and tool calls. Gemma 4 31B in BF16 for text, vision and structured extraction.
- SovAPI routing. Set the model to
sovapiand the router picks between Qwen and Gemma per request. The site calls this "modeless" inference. - Audio and search services. Whisper Large V3 for transcription, Kokoro for French and English speech, BGE-M3 for embeddings.
- Specialised models. SovInfra deploys other open-weight models on reserved capacity for a given workload, with up to 10 billion tokens for a production-scale test.
- Arena. A public page to run your own prompt and see time to first token and generation speed, without an account.
- Transparency. A status page with incident history, a sovereignty page that names every subcontractor, and configuration published on Codeberg rather than GitHub.
For business people
The problem: most AI APIs run on US clouds, which the US CLOUD Act can reach, even when the servers sit in Europe. For regulated sectors, public bodies and anyone careful with customer data, that's a blocker.
What SovInfra offers: a European operator, European data centres and a written data policy. Requests are processed in memory and not stored. The company names who does what:
| Provider | Role | Country |
|---|---|---|
| VS Hosting | GPU hosting | Czech Republic |
| RINF, Selectiv T & C | Inference operations, network monitoring | Romania |
| Hetzner | Redundancy and backup | Germany |
| SovInfra | Gateway and billing | France |
Cost: cheap. $0.12 per million input tokens and $0.38 per million output tokens for both language models, $0.04 for cached input on Qwen. Whisper costs $0.00048 per audio minute, so 1,000 hours of audio cost about $29. The 1 billion free tokens cover a real pilot.
The limits:
- A very young company, with no public customers, funding or team page.
- Few models. 2 mid-size language models, not frontier ones. For the hardest tasks, Claude or GPT still lead, and the site links to their published benchmarks next to its own models.
- Availability claims need reading closely. The data centres are "designed for 99.9999%", but the site itself says that isn't a measured or contractual figure for the API. SLAs come with a contract.
- Payments and email go through Stripe and Brevo, which the privacy policy covers. Stripe is a US company, although it handles no inference data.
For technical people
The API follows the OpenAI format, so migration is a base URL and a key:
from openai import OpenAI
client = OpenAI(base_url="https://api.sovinfra.ai/v1", api_key="YOUR_SOVINFRA_KEY")
response = client.chat.completions.create(
model="qwen3.8-27b", # or "gemma-4-31b", or "sovapi" to let the router choose
messages=[{"role": "user", "content": "Summarise this contract in 5 bullet points."}],
)
print(response.choices[0].message.content)
- Precision is stated per model: FP8 for Qwen, BF16 for Gemma. Many providers don't say, and quantisation changes output quality.
- Latency: first token from about 50 ms in Europe, per the site. Test it on the Arena with your own prompt.
- Prefix caching keeps context in volatile memory, isolated per API key. That's what makes the $0.04 cached price work for agent loops that resend the same long context.
- Security: TLS 1.3 in transit, AES-256 for stored account data, revocable API keys, no third-party connector on by default.
- Continuity: primary orchestration in France with a standby in Falkenstein, Germany, and a standby gateway in a second country.
- Codeberg: the published repository holds a GPUStack configuration, which suggests how the GPUs are managed.
An earlier launch post on the SovInfra blog describes a different offer: Llama models billed per GPU second. The service moved to open-weight models billed per token since then.
How it compares
| Service | Approach |
|---|---|
| SovInfra | French operator, EU GPUs, open-weight models, per-token pricing |
| Mistral | French, serves its own Mistral models |
| Scaleway Generative APIs | French cloud, open-weight models on its own data centres |
| OpenRouter | 1 API to 100s of models from many providers, mostly US (see OpenRouter ) |
Value
The model caught my eye first. My own NicAI model is a LoRA fine-tune of Qwen 3.8-27B that I serve on my Mac Studio (see Fine-tuning my own model (LLM): NicAI ). SovInfra serves the same base model on H200s, and it offers to deploy specialised models for a workload. That makes it a candidate host if I ever need my model reachable outside my home network, without sending the data to a US cloud.
For work, it's a useful reference in the sovereignty conversation. European customers ask where their data goes, and a provider that names each subcontractor and country is a good example of what a clear answer looks like. The 1 billion free tokens make it cheap to test on a real workload before trusting it.
Further Reading


