Features
Implement OpenAI-Compatible Features - Developer Guide
Use OpenAI client libraries with the Infercom API. Drop-in compatible — just change the base URL and API key to switch to EU sovereign inference.
Infercom inference APIs are designed to be compliant with OpenAI client libraries to simplify the adoption of our inference technologies to enhance your AI applications.
Download the library
Section titled “Download the library”Run the command below to download the library.
pip install openaiUse Infercom APIs with OpenAI client libraries
Section titled “Use Infercom APIs with OpenAI client libraries”Configuring your OpenAI client libraries to use Infercom inference APIs is as simple as setting two values: the base_url and your api_key, as shown below.
from openai import OpenAI
client = OpenAI( base_url="https://api.infercom.ai/v1", api_key="your-infercom-api-key")Now you can make an API request to a model and choose how to receive your output.
Non-streaming example
Section titled “Non-streaming example”The following code demonstrates using the OpenAI Python client for non-streaming completions.
completion = client.chat.completions.create( model="Meta-Llama-3.1-8B-Instruct", messages = [ {"role": "system", "content": "Answer the question in a couple sentences."}, {"role": "user", "content": "Share a happy story with me"} ])
print(completion.choices[0].message)Streaming example
Section titled “Streaming example”The following code demonstrates using the OpenAI Python client for streaming completions.
completion = client.chat.completions.create( model="Meta-Llama-3.1-8B-Instruct", messages = [ {"role": "system", "content": "Answer the question in a couple sentences."}, {"role": "user", "content": "Share a happy story with me"} ], stream= True)
for chunk in completion: print(chunk.choices[0].delta.content)Accepted and ignored
Section titled “Accepted and ignored”Two OpenAI parameters are accepted for API compatibility and have no effect on any model:
presence_penaltyfrequency_penalty
To control repetition, use repetition_penalty instead.
Supported, but not on every model
Section titled “Supported, but not on every model”These parameters work. Which models honour them differs, so check the table before you rely on one.
| Parameter | Honoured on | No effect on | Request fails on |
|---|---|---|---|
n |
Every model | - | - |
logprobs, top_logprobs |
Every model except MiniMax-M3 |
MiniMax-M3 |
- |
logit_bias |
gpt-oss-120b |
gemma-4-31B-it |
DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct, MiniMax-M3 |
seed |
gpt-oss-120b, DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct |
gemma-4-31B-it, MiniMax-M3 |
- |
repetition_penalty |
gpt-oss-120b, gemma-4-31B-it |
DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct |
MiniMax-M3, at any value other than 1.0 |
Measured on /v1/chat/completions. A failing request returns HTTP 500, except logit_bias on MiniMax-M3,
which returns HTTP 400. The OpenAI SDKs retry a 500 twice, so do not send these parameters to these models.
Feature differences
Section titled “Feature differences”temperature: The Infercom API accepts values between 0 and 2, the same range as OpenAI. Values above 2 are rejected with a 400. Unlike OpenAI, the Infercom API applies no default. Omit temperature and most models decode greedily, which behaves like temperature 0. Set it on every request. See If you omit temperature.
Infercom API features not supported by OpenAI clients
Section titled “Infercom API features not supported by OpenAI clients”The Infercom API supports two parameters that the OpenAI client libraries do not:
top_k- limits sampling to the K most probable tokens. It has no effect unlesstemperatureis 0.001 or above. See How top_k interacts with temperature.repetition_penalty- the only repetition control this API applies. It also penalises tokens that appear in your prompt, so a long system prompt can suppress words the prompt relies on. It is not a drop-in replacement forfrequency_penalty, which is accepted and ignored. Accepted range is 1 to 2. Set a correcttemperaturefirst, then keeprepetition_penaltyat or below 1.2. Applied on/v1/chat/completionsonly.
The OpenAI SDKs reject unknown top-level arguments, so send both through extra_body.
response = client.chat.completions.create( model="gpt-oss-120b", messages=[{"role": "user", "content": "..."}], temperature=1.0, extra_body={"top_k": 40, "repetition_penalty": 1.05},)