Skip to Content
API๐Ÿ”„ OpenAI-Compatible API

OpenAI-Compatible API

DocsGPT exposes /v1/chat/completions following the standard chat completions protocol. Point any compatible client โ€” opencode, Aider, LibreChat or the OpenAI SDKs โ€” at your DocsGPT Agent by changing only the base URL and API key.

Quick Start

from openai import OpenAI client = OpenAI( base_url="http://localhost:7091/v1", # or https://gptcloud.arc53.com/v1 api_key="your_agent_api_key", ) response = client.chat.completions.create( model="docsgpt-agent", messages=[{"role": "user", "content": "Summarize our refund policy"}], ) print(response.choices[0].message.content)

The model field is accepted but ignored โ€” the agent bound to your API key determines the model. The agentโ€™s prompt, sources, tools, and default model are loaded automatically.

Base URL & Auth

EnvironmentBase URL
Localhttp://localhost:7091/v1
Cloudhttps://gptcloud.arc53.com/v1

Authenticate with Authorization: Bearer <agent_api_key>. Every /v1 request is an agent-key request, even from the agentโ€™s owner.

Nobody can approve a tool action through /v1. A tool that needs approval comes back to you as an ordinary entry in tool_calls, as if it were one of your own client-side tools: whatever you post back as its role: "tool" message is used as its result, and the server-side tool never runs. To approve server-side tools, use the native /api/answer or /stream with tool_actions.

Write actions that use the agent ownerโ€™s connected accounts or saved credentials are refused unless the owner allows them in the agentโ€™s Access Details > Changes others can make as you. See Letting API callers make changes.

Endpoints

MethodPathDescription
POST/v1/chat/completionsChat request (streaming or non-streaming)
GET/v1/modelsReturns the one agent bound to your key, with the agent id as the model id

Streaming

Set "stream": true. Youโ€™ll receive SSE chunks with choices[0].delta.content. DocsGPT-specific events (sources, tool calls) arrive as extra frames that carry a top-level docsgpt key on an otherwise-empty chunk โ€” standard clients ignore them.

stream = client.chat.completions.create( model="docsgpt-agent", stream=True, messages=[{"role": "user", "content": "Explain vector search"}], ) for chunk in stream: print(chunk.choices[0].delta.content or "", end="", flush=True)

Set "stream_options": {"include_usage": true} to get a chunk with usage (prompt, completion and total tokens, summed over every model call of the turn). It arrives just before the chunk that carries finish_reason, not after it as in OpenAIโ€™s API, so read it wherever it appears rather than only from the last chunk. Non-streaming responses always include usage.

Sampling Parameters

Standard OpenAI sampling parameters are forwarded to the model. When omitted, the agentโ€™s configured defaults apply. Supported: temperature, max_tokens (or max_completion_tokens), top_p, frequency_penalty, presence_penalty, stop, seed. When the request has tools, tool_choice and parallel_tool_calls are forwarded too.

Options DocsGPT canโ€™t honor are rejected with HTTP 400 (invalid_request_error) rather than ignored: n other than 1, logprobs set to anything but false, and a stream_options that is not an object.

{ "model": "docsgpt-agent", "messages": [{"role": "user", "content": "Write a haiku about search"}], "temperature": 0.2, "max_tokens": 256, "seed": 42 }

Structured Output

You can force the model to return JSON matching a schema, using either the OpenAI response_format field or the response_schema convenience field.

{ "model": "docsgpt-agent", "messages": [{"role": "user", "content": "Extract the order id and total"}], "response_format": { "type": "json_schema", "json_schema": { "name": "order", "strict": true, "schema": { "type": "object", "properties": { "order_id": {"type": "string"}, "total": {"type": "number"} }, "required": ["order_id", "total"] } } } }
  • response_format follows OpenAI Structured Outputs. strict defaults to true; set strict: false to relax enforcement.
  • response_format: {"type": "json_object"} requests JSON without a fixed schema (the model is steered by the prompt).
  • response_schema is a DocsGPT convenience: pass a raw JSON Schema object (or a {"schema": {...}} wrapper) directly.

Multimodal Input (text + images)

User messages may use OpenAI typed-content arrays with image_url parts. Images are forwarded to vision-capable models.

{ "model": "docsgpt-agent", "messages": [ { "role": "user", "content": [ {"type": "text", "text": "What's in this screenshot?"}, {"type": "image_url", "image_url": {"url": "https://example.com/shot.png"}} ] } ] }

Tool Calling (client-side, stateless)

You can register your own tools and execute them on the client. The flow is stateless โ€” OpenAI clients that donโ€™t carry a conversation_id re-send the full message history each turn, and DocsGPT rebuilds the agent from it.

  1. Send a request with a tools array.
  2. If the agent decides to call a tool, the response comes back with finish_reason: "tool_calls" and a tool_calls array (and content: null).
  3. Execute the tool(s) on your side, then re-POST the full message history with the assistantโ€™s tool_calls message followed by role: "tool" result messages.
  4. DocsGPT continues the run and returns the final answer.
{ "model": "docsgpt-agent", "messages": [ {"role": "user", "content": "What's the weather in Paris?"}, {"role": "assistant", "tool_calls": [ {"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": "{\"city\":\"Paris\"}"}} ]}, {"role": "tool", "tool_call_id": "call_1", "content": "18ยฐC, clear"} ], "tools": [ { "type": "function", "function": { "name": "get_weather", "...": "..." } } ] }

Reasoning

For models that emit reasoning (โ€œthinkingโ€) tokens, the response surfaces them in a non-standard reasoning_content field (a reasoning_content delta when streaming). Standard clients ignore it; clients that understand it can display the modelโ€™s thinking separately from the answer.

Idempotent Retries

Add an Idempotency-Key header so a retried request returns the stored first response instead of re-running the agent (which would duplicate the answer and double-bill tokens).

curl -X POST http://localhost:7091/v1/chat/completions \ -H "Authorization: Bearer your_agent_api_key" \ -H "Idempotency-Key: 8f1c...unique-per-request" \ -H "Content-Type: application/json" \ -d '{"model":"docsgpt-agent","messages":[{"role":"user","content":"hi"}]}'
  • Opt-in โ€” no header means todayโ€™s behavior (every request runs).
  • Non-streaming only โ€” streaming replay is not supported.
  • A completed key replays the cached body (and status) for 24 hours.
  • A request with a key whose first attempt is still in flight returns HTTP 409.
  • Keys are scoped per agent and capped at 256 characters (oversized keys are rejected).

System Prompt Override

System messages are dropped by default โ€” the agentโ€™s configured prompt is used. To allow callers to override it, enable Allow prompt override in the agentโ€™s Advanced settings.

When an override is active, the agentโ€™s prompt template is replaced wholesale โ€” template variables like {summaries} are not substituted.

Conversation Persistence

Conversations are always persisted server-side, and the response includes docsgpt.conversation_id. They never appear in the agent ownerโ€™s sidebar โ€” /v1 traffic is stored hidden, so external clients canโ€™t clutter the ownerโ€™s conversation list.

A tool-result round that carries no conversation (no conversation_id and no linked session, as with a client that resends the whole transcript) is not stored, so it doesnโ€™t leave orphan rows. docsgpt.persist and the legacy docsgpt.save_conversation flag from older releases have no effect.

Continuing a conversation

By default each request is a new conversation built from the messages you send. To continue a stored one, pass its id in any of these, in order of precedence:

  1. the X-DocsGPT-Conversation-ID header;
  2. a top-level conversation_id field;
  3. docsgpt.conversation_id.

The conversation must belong to the agent bound to your key; otherwise the request fails with HTTP 400 and "code": "conversation_not_found".

Session headers

Chat-completions requests have no conversation field, so coding clients such as opencode send a stable session header instead. When a request carries no conversation id, DocsGPT reads the first of these it finds and links the session to a conversation:

  • the X-DocsGPT-Session-ID header, or docsgpt.session_id in the body;
  • the X-Session-ID or X-Session-Affinity header.

Requests with the same session value, the same agent and the same system and developer messages continue the same conversation. The link is kept in Redis, only as a hash of the value, for V1_SESSION_TTL_SECONDS (24 hours by default) after the last request that used it: each request restarts the time limit. If the linked conversation is gone, a new one starts. Without Redis, session headers have no effect.

DocsGPT Extension Fields

DocsGPT adds an optional docsgpt object to both requests and responses for features outside the OpenAI schema.

Request (docsgpt.*):

FieldDescription
attachmentsList of attachment IDs to include as context for this turn. Upload them with the native /api/store_attachment.
conversation_idContinue a stored conversation (see Continuing a conversation). A top-level conversation_id works too.
session_idLink requests into one conversation (see Session headers).

Response (docsgpt.*):

FieldDescription
conversation_idServer-side conversation ID for this exchange.
sourcesRAG sources used to answer.
tool_callsCompleted tool-call results from the run.

When streaming, these arrive on otherwise-empty chunks that carry a top-level docsgpt key, so strict OpenAI clients still validate each frame.

When to Use Native Endpoints Instead

Use /api/answer or /stream if you need passthrough template variables or sidebar visibility control via visibility. Uploading an attachment also uses the native /api/store_attachment; the resulting id then works in docsgpt.attachments.