OpenAI-Compatible API
DocsGPT exposes /v1/chat/completions following the standard chat completions protocol. Point any compatible client โ opencode, Aider, LibreChat or the OpenAI SDKs โ at your DocsGPT Agent by changing only the base URL and API key.
Quick Start
Python
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:7091/v1", # or https://gptcloud.arc53.com/v1
api_key="your_agent_api_key",
)
response = client.chat.completions.create(
model="docsgpt-agent",
messages=[{"role": "user", "content": "Summarize our refund policy"}],
)
print(response.choices[0].message.content)The model field is accepted but ignored โ the agent bound to your API key determines the model. The agentโs prompt, sources, tools, and default model are loaded automatically.
Base URL & Auth
| Environment | Base URL |
|---|---|
| Local | http://localhost:7091/v1 |
| Cloud | https://gptcloud.arc53.com/v1 |
Authenticate with Authorization: Bearer <agent_api_key>. Every /v1 request is an agent-key request, even from the agentโs owner.
Nobody can approve a tool action through /v1. A tool that needs approval comes back to you as an ordinary entry in tool_calls, as if it were one of your own client-side tools: whatever you post back as its role: "tool" message is used as its result, and the server-side tool never runs. To approve server-side tools, use the native /api/answer or /stream with tool_actions.
Write actions that use the agent ownerโs connected accounts or saved credentials are refused unless the owner allows them in the agentโs Access Details > Changes others can make as you. See Letting API callers make changes.
Endpoints
| Method | Path | Description |
|---|---|---|
POST | /v1/chat/completions | Chat request (streaming or non-streaming) |
GET | /v1/models | Returns the one agent bound to your key, with the agent id as the model id |
Streaming
Set "stream": true. Youโll receive SSE chunks with choices[0].delta.content. DocsGPT-specific events (sources, tool calls) arrive as extra frames that carry a top-level docsgpt key on an otherwise-empty chunk โ standard clients ignore them.
stream = client.chat.completions.create(
model="docsgpt-agent",
stream=True,
messages=[{"role": "user", "content": "Explain vector search"}],
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)Set "stream_options": {"include_usage": true} to get a chunk with usage (prompt, completion and total tokens, summed over every model call of the turn). It arrives just before the chunk that carries finish_reason, not after it as in OpenAIโs API, so read it wherever it appears rather than only from the last chunk. Non-streaming responses always include usage.
Sampling Parameters
Standard OpenAI sampling parameters are forwarded to the model. When omitted, the agentโs configured defaults apply. Supported: temperature, max_tokens (or max_completion_tokens), top_p, frequency_penalty, presence_penalty, stop, seed. When the request has tools, tool_choice and parallel_tool_calls are forwarded too.
Options DocsGPT canโt honor are rejected with HTTP 400 (invalid_request_error) rather than ignored: n other than 1, logprobs set to anything but false, and a stream_options that is not an object.
{
"model": "docsgpt-agent",
"messages": [{"role": "user", "content": "Write a haiku about search"}],
"temperature": 0.2,
"max_tokens": 256,
"seed": 42
}Structured Output
You can force the model to return JSON matching a schema, using either the OpenAI response_format field or the response_schema convenience field.
response_format
{
"model": "docsgpt-agent",
"messages": [{"role": "user", "content": "Extract the order id and total"}],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "order",
"strict": true,
"schema": {
"type": "object",
"properties": {
"order_id": {"type": "string"},
"total": {"type": "number"}
},
"required": ["order_id", "total"]
}
}
}
}response_formatfollows OpenAI Structured Outputs.strictdefaults totrue; setstrict: falseto relax enforcement.response_format: {"type": "json_object"}requests JSON without a fixed schema (the model is steered by the prompt).response_schemais a DocsGPT convenience: pass a raw JSON Schema object (or a{"schema": {...}}wrapper) directly.
Multimodal Input (text + images)
User messages may use OpenAI typed-content arrays with image_url parts. Images are forwarded to vision-capable models.
{
"model": "docsgpt-agent",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "What's in this screenshot?"},
{"type": "image_url", "image_url": {"url": "https://example.com/shot.png"}}
]
}
]
}Tool Calling (client-side, stateless)
You can register your own tools and execute them on the client. The flow is stateless โ OpenAI clients that donโt carry a conversation_id re-send the full message history each turn, and DocsGPT rebuilds the agent from it.
- Send a request with a
toolsarray. - If the agent decides to call a tool, the response comes back with
finish_reason: "tool_calls"and atool_callsarray (andcontent: null). - Execute the tool(s) on your side, then re-POST the full message history with the assistantโs
tool_callsmessage followed byrole: "tool"result messages. - DocsGPT continues the run and returns the final answer.
{
"model": "docsgpt-agent",
"messages": [
{"role": "user", "content": "What's the weather in Paris?"},
{"role": "assistant", "tool_calls": [
{"id": "call_1", "type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\":\"Paris\"}"}}
]},
{"role": "tool", "tool_call_id": "call_1", "content": "18ยฐC, clear"}
],
"tools": [ { "type": "function", "function": { "name": "get_weather", "...": "..." } } ]
}Reasoning
For models that emit reasoning (โthinkingโ) tokens, the response surfaces them in a non-standard reasoning_content field (a reasoning_content delta when streaming). Standard clients ignore it; clients that understand it can display the modelโs thinking separately from the answer.
Idempotent Retries
Add an Idempotency-Key header so a retried request returns the stored first response instead of re-running the agent (which would duplicate the answer and double-bill tokens).
curl -X POST http://localhost:7091/v1/chat/completions \
-H "Authorization: Bearer your_agent_api_key" \
-H "Idempotency-Key: 8f1c...unique-per-request" \
-H "Content-Type: application/json" \
-d '{"model":"docsgpt-agent","messages":[{"role":"user","content":"hi"}]}'- Opt-in โ no header means todayโs behavior (every request runs).
- Non-streaming only โ streaming replay is not supported.
- A completed key replays the cached body (and status) for 24 hours.
- A request with a key whose first attempt is still in flight returns HTTP 409.
- Keys are scoped per agent and capped at 256 characters (oversized keys are rejected).
System Prompt Override
System messages are dropped by default โ the agentโs configured prompt is used. To allow callers to override it, enable Allow prompt override in the agentโs Advanced settings.
When an override is active, the agentโs prompt template is replaced wholesale โ template variables like {summaries} are not substituted.
Conversation Persistence
Conversations are always persisted server-side, and the response includes docsgpt.conversation_id. They never appear in the agent ownerโs sidebar โ /v1 traffic is stored hidden, so external clients canโt clutter the ownerโs conversation list.
A tool-result round that carries no conversation (no conversation_id and no linked session, as with a client that resends the whole transcript) is not stored, so it doesnโt leave orphan rows. docsgpt.persist and the legacy docsgpt.save_conversation flag from older releases have no effect.
Continuing a conversation
By default each request is a new conversation built from the messages you send. To continue a stored one, pass its id in any of these, in order of precedence:
- the
X-DocsGPT-Conversation-IDheader; - a top-level
conversation_idfield; docsgpt.conversation_id.
The conversation must belong to the agent bound to your key; otherwise the request fails with HTTP 400 and "code": "conversation_not_found".
Session headers
Chat-completions requests have no conversation field, so coding clients such as opencode send a stable session header instead. When a request carries no conversation id, DocsGPT reads the first of these it finds and links the session to a conversation:
- the
X-DocsGPT-Session-IDheader, ordocsgpt.session_idin the body; - the
X-Session-IDorX-Session-Affinityheader.
Requests with the same session value, the same agent and the same system and developer messages continue the same conversation. The link is kept in Redis, only as a hash of the value, for V1_SESSION_TTL_SECONDS (24 hours by default) after the last request that used it: each request restarts the time limit. If the linked conversation is gone, a new one starts. Without Redis, session headers have no effect.
DocsGPT Extension Fields
DocsGPT adds an optional docsgpt object to both requests and responses for features outside the OpenAI schema.
Request (docsgpt.*):
| Field | Description |
|---|---|
attachments | List of attachment IDs to include as context for this turn. Upload them with the native /api/store_attachment. |
conversation_id | Continue a stored conversation (see Continuing a conversation). A top-level conversation_id works too. |
session_id | Link requests into one conversation (see Session headers). |
Response (docsgpt.*):
| Field | Description |
|---|---|
conversation_id | Server-side conversation ID for this exchange. |
sources | RAG sources used to answer. |
tool_calls | Completed tool-call results from the run. |
When streaming, these arrive on otherwise-empty chunks that carry a top-level docsgpt key, so strict OpenAI clients still validate each frame.
When to Use Native Endpoints Instead
Use /api/answer or /stream if you need passthrough template variables or sidebar visibility control via visibility. Uploading an attachment also uses the native /api/store_attachment; the resulting id then works in docsgpt.attachments.