POST
OpenAI-Compatible Chat Completions

Overview

The Geoff API also exposes an OpenAI-compatible chat completions endpoint. Configure an OpenAI SDK client with https://geoff.ai/api/v1.

Configuration

Supported features and limits

The standard chat path accepts text/image messages, conversation history, system/developer instructions, and function tools. It forwards temperature, top_p, stop, frequency_penalty, presence_penalty, seed, response_format, and parallel_tool_calls to the worker. Model capabilities determine whether each control is available. max_completion_tokens takes precedence over max_tokens; the cap is constrained by the request budget. Only one completion (n: 1) is supported. Logprobs, logit bias, audio output, predictions, storage, hosted web search, and service-tier selection are unsupported and return validation errors. Omit stream (or set it to false) for JSON. For SSE, set stream: true. Set stream_options: { include_usage: true } to receive a final usage chunk with an empty choices array. Other chunks then contain usage: null. reasoning_effort stays on the selected inference path. Stacknet-specific sequential-thinking, multi-model, and media workflows have separate execution contracts; they are not certified against this standard chat parameter subset. Paid traffic through the current Rust settlement gate is buffered until billing commits, even when SSE was requested. Other paths can stream progressively. Tool arguments may arrive as a complete block after the worker finishes. Disconnecting a client does not guarantee cancellation of inference. The legacy /api/v1/text/chat endpoint is an alias of /api/v1/chat/completions. Responses include HTTP status codes and request identifiers.

Authorizations

Authorization
string
header
required

Bearer API_key, can be found in Account Management > API Keys.

Headers

X-Stack-Id
string
default:stk_jmqog66ha0mugmro
required

Supplied automatically by the documentation playground.

Body

application/json
model
enum<string>
default:magma
required

Select a supported model layer.

Available options:
magma,
pyro,
pyro:max
messages
object[]
required

Array of message objects.

max_tokens
integer

Maximum number of tokens to generate.

temperature
number
default:1

Sampling temperature between 0 and 2.

top_p
number
default:1

Nucleus sampling parameter.

stream
boolean
default:false

Whether to stream the response via SSE.

tools
object[]

Tool definitions for function calling.

Response

200 - application/json

Successful response in OpenAI chat completion format.

id
string
object
string
created
integer
model
string
choices
object[]
usage
object