Chat Completions

Generate a model response for a conversation. This is the core endpoint — it works identically to OpenAI's chat completions API.

Endpoint

POST https://dev-backend.sovereigneg.com/v1/chat/completions

Request body

ParameterTypeRequiredDescription
modelstringYesModel ID from GET /v1/models or the Model Library
messagesarrayYesConversation messages (see below)
temperaturefloatNoSampling temperature (0.0–2.0). Default: 1.0
max_tokensintegerNoMax tokens to generate. Default: model-dependent
top_pfloatNoNucleus sampling (0.0–1.0). Default: 1.0
streambooleanNoStream response via SSE. Default: false
stopstring/arrayNoStop sequences
frequency_penaltyfloatNoFrequency penalty (-2.0 to 2.0). Default: 0.0
presence_penaltyfloatNoPresence penalty (-2.0 to 2.0). Default: 0.0
reasoning_effortstringNoReasoning depth on models that reason: none, minimal, low, medium, high, xhigh. Dropped on models that do not reason and mapped to the nearest value the selected model accepts; see Parameter adjustments below

Parameter adjustments

Agent frameworks send reasoning_effort on every request, including to models that do not reason. Instead of failing the request with the provider's 400, the gateway adjusts the parameter for the selected model and names the adjustment in a response header:

HeaderMeaning
X-Dropped-Params: reasoning_effortThe model does not reason (for example gpt-4o-mini), or its provider has no equivalent of the field (Claude models served directly by Anthropic use thinking budgets instead), so the field was removed before the request reached the provider.
X-Mapped-Params: reasoning_effort=<value>The model reasons but does not accept the value you sent; <value> is what it received (for example none becomes minimal on gpt-5-mini).

Neither header appears when the value went through unchanged. Models served by OpenAI-compatible providers other than OpenAI itself (OpenRouter, DeepInfra, Together and the like) receive the value as sent and never trigger an adjustment. On streaming and non-streaming responses alike, the headers describe the backing that served the request, including after a failover to another backing.

Message roles

RolePurpose
systemSets the behavior of the assistant
userThe human's input
assistantThe model's previous response (for multi-turn)

Example request

from openai import OpenAI
 
client = OpenAI(
    api_key="sk-...",
    base_url="https://dev-backend.sovereigneg.com/v1"
)
 
response = client.chat.completions.create(
    model="gpt-oss-20b",
    messages=[
        {"role": "system", "content": "You are a concise technical writer."},
        {"role": "user", "content": "Explain how transformers work in 3 sentences."}
    ],
    temperature=0.7,
    max_tokens=256
)
 
print(response.choices[0].message.content)

Response

{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1711000000,
  "model": "gpt-oss-20b",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Transformers use self-attention mechanisms..."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 28,
    "completion_tokens": 65,
    "total_tokens": 93
  }
}

Streaming

Set stream: true to receive tokens as server-sent events. See Streaming for details.

Multi-turn conversation

Pass the full conversation history in the messages array:

response = client.chat.completions.create(
    model="gpt-oss-20b",
    messages=[
        {"role": "system", "content": "You are a helpful math tutor."},
        {"role": "user", "content": "What is 15% of 200?"},
        {"role": "assistant", "content": "15% of 200 is 30."},
        {"role": "user", "content": "And 20% of the same number?"}
    ]
)