Skip to content
InfercomInfercomInfercom Documentation

Features

Implement OpenAI-Compatible Features - Developer Guide

Use OpenAI client libraries with the Infercom API. Drop-in compatible — just change the base URL and API key to switch to EU sovereign inference.

Infercom inference APIs are designed to be compliant with OpenAI client libraries to simplify the adoption of our inference technologies to enhance your AI applications.

Run the command below to download the library.

pip install openai

Use Infercom APIs with OpenAI client libraries

Section titled “Use Infercom APIs with OpenAI client libraries”

Configuring your OpenAI client libraries to use Infercom inference APIs is as simple as setting two values: the base_url and your api_key, as shown below.

from openai import OpenAI
client = OpenAI(
base_url="https://api.infercom.ai/v1",
api_key="your-infercom-api-key"
)

Now you can make an API request to a model and choose how to receive your output.

The following code demonstrates using the OpenAI Python client for non-streaming completions.

completion = client.chat.completions.create(
model="Meta-Llama-3.1-8B-Instruct",
messages = [
{"role": "system", "content": "Answer the question in a couple sentences."},
{"role": "user", "content": "Share a happy story with me"}
]
)
print(completion.choices[0].message)

The following code demonstrates using the OpenAI Python client for streaming completions.

completion = client.chat.completions.create(
model="Meta-Llama-3.1-8B-Instruct",
messages = [
{"role": "system", "content": "Answer the question in a couple sentences."},
{"role": "user", "content": "Share a happy story with me"}
],
stream= True
)
for chunk in completion:
print(chunk.choices[0].delta.content)

Two OpenAI parameters are accepted for API compatibility and have no effect on any model:

  • presence_penalty
  • frequency_penalty

To control repetition, use repetition_penalty instead.

These parameters work. Which models honour them differs, so check the table before you rely on one.

Parameter Honoured on No effect on Request fails on
n Every model - -
logprobs, top_logprobs Every model except MiniMax-M3 MiniMax-M3 -
logit_bias gpt-oss-120b gemma-4-31B-it DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct, MiniMax-M3
seed gpt-oss-120b, DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct gemma-4-31B-it, MiniMax-M3 -
repetition_penalty gpt-oss-120b, gemma-4-31B-it DeepSeek-V3.1, DeepSeek-V3.2, Meta-Llama-3.3-70B-Instruct MiniMax-M3, at any value other than 1.0

Measured on /v1/chat/completions. A failing request returns HTTP 500, except logit_bias on MiniMax-M3, which returns HTTP 400. The OpenAI SDKs retry a 500 twice, so do not send these parameters to these models.

temperature: The Infercom API accepts values between 0 and 2, the same range as OpenAI. Values above 2 are rejected with a 400. Unlike OpenAI, the Infercom API applies no default. Omit temperature and most models decode greedily, which behaves like temperature 0. Set it on every request. See If you omit temperature.

Infercom API features not supported by OpenAI clients

Section titled “Infercom API features not supported by OpenAI clients”

The Infercom API supports two parameters that the OpenAI client libraries do not:

  • top_k - limits sampling to the K most probable tokens. It has no effect unless temperature is 0.001 or above. See How top_k interacts with temperature.
  • repetition_penalty - the only repetition control this API applies. It also penalises tokens that appear in your prompt, so a long system prompt can suppress words the prompt relies on. It is not a drop-in replacement for frequency_penalty, which is accepted and ignored. Accepted range is 1 to 2. Set a correct temperature first, then keep repetition_penalty at or below 1.2. Applied on /v1/chat/completions only.

The OpenAI SDKs reject unknown top-level arguments, so send both through extra_body.

response = client.chat.completions.create(
model="gpt-oss-120b",
messages=[{"role": "user", "content": "..."}],
temperature=1.0,
extra_body={"top_k": 40, "repetition_penalty": 1.05},
)