Skip to main content
The Infercom Inference Service provides access to a broad selection of AI models through the Global Model Catalog. Models are available in two categories: EU-hosted models running on Infercom’s sovereign infrastructure in Germany, and additional models available via global infrastructure.

EU-hosted models

The following models run on Infercom’s EU infrastructure in Germany. For these models, all data processing happens entirely within the EU — no data leaves EU jurisdiction, with full GDPR compliance and no US CLOUD Act exposure.

Text generation models

Embedding model

Audio model

Global Model Catalog

In addition to EU-hosted models, the Infercom Inference Service provides access to a broader selection of models through the Global Model Catalog. Models not hosted in our EU datacenters are served via global infrastructure. You can always check a model’s hosting region via the API or the Playground.
Every model is clearly labeled with its hosting region — both in the API response and in the Playground. You always know where your data is being processed.
See Identifying model regions below for how to check where each model runs. Sampling parameters are model-specific. Use the values the model’s publisher recommends rather than a single setting across all models. The table below reproduces each publisher’s own recommendation, with a link to the source. A dash means the publisher does not publish a value for that parameter. Where no recommendation exists, start from the API defaults and tune for your own workload. DeepSeek-V3.1 publishes no prose guidance, but the generation config it ships defaults to temperature 0.6 and top_p 0.95. The Infercom API accepts a temperature between 0 and 2. temperature=1.0 does not mean “maximum randomness” - it means the model’s probability distribution is used exactly as trained. Values below 1.0 sharpen it, values above 1.0 flatten it. Adjust temperature, top_p, or top_k, but generally not more than one of them at a time.
Very low temperatures can cause non-terminating reasoning. Reasoning models can enter a repetition loop, restating the same point until max_tokens runs out, returning finish_reason: "length" and no usable answer. Raising max_tokens does not help. A low temperature is still a legitimate choice for reproducible or structured output - just move it back towards the recommended value if you see truncated responses. See Reasoning.

Default system prompt for MiniMax-M2.7

MiniMax publish a default system prompt for MiniMax-M2.7. Use it verbatim unless you have a reason to replace it:

Identifying model regions

You can identify where a model is hosted through the API or the Playground.

Via the API

The /v1/models endpoint includes an sn_metadata object for each model. Use the region field to determine where a model is hosted: "EU" for sovereign models on Infercom’s EU infrastructure, or a non-EU region (e.g. "US", "JP") for models on global infrastructure. Use the ?verbose=true query parameter to retrieve detailed model metadata including sovereignty information:
Example response for an EU-hosted model:
Example response for a globally-routed model:

Via the Playground

In the Infercom Playground, region flags are displayed next to each model name, letting you see at a glance where each model runs.

Data sovereignty

EU sovereignty applies to EU-hosted models only. When using models from the Global Model Catalog that are not hosted on EU infrastructure, requests are processed on global infrastructure outside the EU. Always check the model’s region before processing sensitive or regulated data.
For EU-hosted models, Infercom provides:
  • EU data residency — inference runs in our EU datacenters
  • GDPR compliance — full compliance with EU data protection regulations
  • No US CLOUD Act exposure — your inference data is not subject to US jurisdiction
  • AI Act readiness — designed for compliance with the EU AI Act