Settings Reference
Every setting DocsGPT reads, generated from docsgpt/core/settings/. Each one is
an environment variable of the same name, set in .env or the process
environment; see App Configuration for how the
file is found and for worked examples. <DOCSGPT_HOME> below is the data home
described there.
Authentication
How users authenticate: none, a shared token, per-session JWTs, or OIDC SSO.
AUTH_TYPE
Type "simple_jwt" | "session_jwt" | "oidc", default unset.
Authentication mode: simple_jwt, session_jwt, oidc, or unset (None) for no authentication.
JWT_SECRET_KEY
Type str, default "".
Signing key for session tokens and other signed capabilities. Required on every replica in production; local development may fall back to a key generated on disk.
ENCRYPTION_SECRET_KEY
Type str, default default-docsgpt-encryption-key.
Key used to encrypt stored credentials such as tool and connector secrets. Set your own value before connecting services on a multi-user install; the default is public.
ENCRYPTION_SECRET_KEY_PREVIOUS
Type str, default unset.
Previous ENCRYPTION_SECRET_KEY, tried when a stored credential (a connection, a tool or custom-model secret) was encrypted with it. Set it while rotating the key, run docsgpt connectors reencrypt, and remove it once the command reports nothing unreadable.
INTERNAL_KEY
Type str, default unset.
Required: shared secret the worker uses to hand finished indexes and files to the API. Set the same value on the API and the worker; without it the API rejects the worker’s uploads and every ingest fails.
OIDC_ISSUER
Type str, default unset.
OIDC issuer URL with discovery, e.g. https://auth.example.com/application/o/docsgpt/ .
OIDC_CLIENT_ID
Type str, default unset.
OIDC client id.
OIDC_CLIENT_SECRET
Type str, default unset.
OIDC client secret. Optional; PKCE is always used.
OIDC_SCOPES
Type str, default openid profile email.
Scopes requested from the IdP.
OIDC_USER_ID_CLAIM
Type str, default sub.
ID-token claim mapped to the DocsGPT user id.
OIDC_FRONTEND_URL
Type str, default unset.
Browser-facing app origin, e.g. http://localhost:5173 .
OIDC_REDIRECT_URI
Type str, default unset.
Override for the callback URL; default is <request host>/api/auth/oidc/callback.
OIDC_SESSION_LIFETIME_SECONDS
Type int, default 28800, must be > 0.
Lifetime of the minted session JWT in seconds (8h).
OIDC_PROVIDER_NAME
Type str, default unset.
Sign-in button label, e.g. “Acme SSO”.
OIDC_ALLOWED_GROUPS
Type str, default unset.
Comma-separated group allowlist; unset admits any authenticated user.
OIDC_GROUPS_CLAIM
Type str, default groups.
ID-token/userinfo claim carrying group membership.
OIDC_ADMIN_GROUPS
Type str, default unset.
Comma-separated groups granted admin; unset means no OIDC admin mapping.
LOCAL_MODE_ADMIN
Type bool, default false.
Grant admin without a database role. Persisted admin grants live in user_roles (AUTH_TYPE=oidc only); this is the only non-DB admin path, for AUTH_TYPE=None self-host. MUST stay False if networked.
SCIM_ENABLED
Type bool, default false.
Enable SCIM 2.0 provisioning at /scim/v2.
SCIM_TOKEN
Type str, default unset.
Bearer token for IdP SCIM clients (required when SCIM is enabled).
PAT_ENABLED
Type bool, default true.
Master switch for personal access tokens. When false, no token can be created AND every existing token stops authenticating immediately (pipelines using them get 401); tokens are kept and work again when re-enabled. Tokens are only available under AUTH_TYPE=oidc or unset (None); switching to simple_jwt or session_jwt disables them the same way.
PAT_DEFAULT_LIFETIME_DAYS
Type int, default 90, must be > 0.
Lifetime of a personal access token created without an explicit expiry.
PAT_MAX_LIFETIME_DAYS
Type int, default 365, must be > 0.
Longest lifetime a user may request for a personal access token.
PAT_ALLOW_NON_EXPIRING
Type bool, default false.
Let users create personal access tokens that never expire. Off by default.
PAT_MAX_PER_USER
Type int, default 25, must be > 0.
Maximum number of live personal access tokens per user.
LLM providers
Which model answers, how it is reached, and provider-specific behaviour.
LLM_PROVIDER
Type str, default docsgpt.
Provider whose first model is the default when LLM_NAME names none: docsgpt, openai, anthropic, google, groq, openrouter, novita or openai_compatible. For your own OpenAI-compatible server use openai with OPENAI_BASE_URL.
LLM_NAME
Type str, default unset.
Default model id. For a cloud provider it must be an id from docsgpt/core/models/*.yaml or a MODELS_CONFIG_DIR YAML, e.g. gpt-5.5; any other name is ignored with a warning. With OPENAI_BASE_URL it is required and names the model(s) the server serves, comma-separated.
API_KEY
Type str, default unset.
LLM API key used by LLM_PROVIDER.
OPENAI_API_KEY
Type str, default unset.
OpenAI API key.
ANTHROPIC_API_KEY
Type str, default unset.
Anthropic API key.
GOOGLE_API_KEY
Type str, default unset.
Google AI API key.
GROQ_API_KEY
Type str, default unset.
Groq API key.
OPEN_ROUTER_API_KEY
Type str, default unset.
OpenRouter API key.
NOVITA_API_KEY
Type str, default unset.
Novita API key.
OPENAI_API_BASE
Type str, default unset.
Azure OpenAI API base URL.
OPENAI_API_VERSION
Type str, default unset.
Azure OpenAI API version.
AZURE_DEPLOYMENT_NAME
Type str, default unset.
Azure deployment name for answering.
AZURE_EMBEDDINGS_DEPLOYMENT_NAME
Type str, default unset.
Azure deployment name for embeddings.
OPENAI_BASE_URL
Type str, default unset.
Base URL for OpenAI-compatible model servers.
LLM_ALLOW_PLAINTEXT_ENDPOINTS
Type bool, default false.
Allow a configured LLM API key to be sent over plain http to any endpoint already in that key’s configured scope. Off, plain http is accepted only for loopback, private and link-local addresses, single-label hosts (Docker service names), host.docker.internal and names ending in .local, .svc, .internal or .localhost; every other endpoint needs https.
FALLBACK_LLM_PROVIDER
Type str, default unset.
Provider for the fallback LLM.
FALLBACK_LLM_NAME
Type str, default unset.
Model name for the fallback LLM.
FALLBACK_LLM_API_KEY
Type str, default unset.
API key for the fallback LLM. Unset uses the fallback provider’s own key (e.g. ANTHROPIC_API_KEY), or API_KEY when FALLBACK_LLM_PROVIDER is LLM_PROVIDER; the primary provider’s key never goes to another provider.
TITLE_MODEL_ID
Type str, default unset.
Optional cheaper model for conversation titles; unset reuses the answer model.
MODELS_CONFIG_DIR
Type str, default unset.
Directory of operator-supplied model YAMLs, loaded after the built-in catalog; later wins on duplicate model id. See docsgpt/core/models/README.md.
DEFAULT_LLM_TOKEN_LIMIT
Type int, default 128000.
Context window assumed when the model is not found in the registry.
RESERVED_TOKENS
Type dict[str, int], default {"system_prompt": 500, "current_query": 500, "safety_buffer": 1000}.
Tokens held back from the context window for the system prompt, the query and a safety buffer.
CACHE_REDIS_URL
Type str, default redis://localhost:6379/2.
Redis URL for the LLM cache, the live event and notification streams, SSO state and token denylist, device pairing, MCP OAuth, the API session store, live speech-to-text and the version check.
LLM_CACHE_ENABLED
Type bool, default true.
Cache LLM answers in Redis (CACHE_REDIS_URL) and replay them for an identical request: same model, messages and generation parameters. Calls that pass tools are never cached. False turns it off.
LLM_CACHE_TTL
Type int, default 1800, must be > 0.
Seconds a cached LLM answer is kept in Redis before it expires.
OPENAI_RESPONSES_STORE
Type bool, default false.
True persists Responses API calls server-side so previous_response_id can chain turns. False keeps them stateless, carrying reasoning across the tool loop as encrypted items.
OPENAI_RESPONSES_CHAIN_ACROSS_TURNS
Type bool, default true.
Cross-turn previous_response_id chaining (store mode only). The chained transcript lives on the provider and is invisible to every local guard, so it is bounded: a turn starts from the local history when the previous turn’s reported prompt already reached the budget (default: the model’s context window) or when the conversation was compressed after that turn was produced.
OPENAI_RESPONSES_CHAIN_BUDGET_TOKENS
Type int, default unset.
Prompt-token budget for cross-turn chaining; unset uses the model’s context window.
OPENAI_RESPONSES_TRUNCATION_AUTO
Type bool, default false.
Send truncation: “auto” so the provider drops the oldest input items instead of failing every request once a chain exceeds the model’s window.
OPENAI_PROMPT_CACHE_KEY
Type bool, default true.
Route a user’s Responses API calls to the same prompt-cache shard with an opaque per-user key.
OPENAI_PROMPT_CACHE_RETENTION
Type str, default unset.
Request extended prompt-cache retention where the provider offers it.
OPENAI_REASONING_SUMMARY
Type str, default auto.
Reasoning summary mode requested from the Responses API.
Embeddings
The embedding model, remote or local, and the batching around it.
EMBEDDINGS_NAME
Type str, default huggingface_sentence-transformers/all-mpnet-base-v2.
Embedding model. Leave unset and the first boot pins one in the database: granite for a new install, this legacy default for an install that already has sources, since granite is the same width and a silent swap would degrade retrieval. Setting it overrides the pin; to switch an existing index, set it and run docsgpt.scripts.reembed.
EMBEDDINGS_BASE_URL
Type str, default unset.
Remote embeddings API URL (OpenAI-compatible).
EMBEDDINGS_KEY
Type str, default unset.
API key for remote or OpenAI embeddings. OpenAI embeddings fall back to OPENAI_API_KEY, then to API_KEY when LLM_PROVIDER=openai.
EMBEDDINGS_MAX_INPUT_TOKENS
Type int, default unset.
Truncate each remote embed input to N tokens (overflow is lost).
EMBEDDINGS_MAX_QUERY_TOKENS
Type int, default 512, must be >= 0.
Clip a search query to N tokens before embedding it (0 disables). Embedder memory grows with the square of input length, and a query needs a few hundred tokens: one 9k-token query (a webhook payload, pasted logs) OOM-killed a 12 GB embeddings server. Documents are not affected.
EMBEDDINGS_LOCAL_MAX_TOKENS
Type int, default unset, must be >= 1.
Hard ceiling, in tokens, on every input a local FastEmbed model embeds, documents included (the overflow is dropped). Unset keeps the model’s own maximum, which is 32,768 for granite; its memory grows with the square of input length, so 4096 caps one input at about 3 GB.
EMBEDDINGS_BATCH_SIZE
Type int, default 32, must be >= 1.
Chunks per store transaction and per remote embed request.
EMBEDDINGS_MODEL_BATCH_SIZE
Type int, default 1, must be >= 1.
Documents per local ONNX forward pass. Each pass pads to its longest input, and that waste grows with the square of chunk length: on a 30-document ingest at 1250-token chunks, 32 peaked at 7.7 GB, 1 at 1.5 GB.
EMBEDDINGS_THREADS
Type int, default unset.
Intra-op threads for the local ONNX runner; unset uses every core. It scales sub-linearly, so several single-threaded workers beat one many-threaded process on the same cores.
EMBEDDINGS_CACHE_DIR
Type str, default <DOCSGPT_HOME>/models.
Where embedding models and their tokenizers are cached. Persistent by default: FastEmbed’s own default is the temp dir.
EMBEDDINGS_POOLING
Type "cls" | "mean", default unset.
Pooling strategy (“cls” or “mean”). Read from the model’s own repository; set only for a repository that declares none, or to override what it declares.
EMBEDDINGS_NORMALIZE
Type bool, default unset.
L2-normalise embeddings. Read from the model’s own repository; set only for a repository that declares nothing, or to override what it declares.
EMBEDDINGS_DELEGATE_TO_WORKER
Type bool, default true.
Embed on the worker so the API holds no model (~660 MB down to ~285 MB on a default install), at one broker round trip per query. Ignored when EMBEDDINGS_BASE_URL is set, which is the better answer for production.
EMBEDDINGS_QUEUE
Type str, default embeddings.
Celery queue the embed task is routed to.
EMBEDDINGS_DELEGATE_TIMEOUT
Type int, default 60, must be > 0.
Seconds the API waits for the worker to return an embedding.
Retrieval
Which vector store answers searches and how retrieval fans out across sources.
VECTOR_STORE
Type "faiss" | "elasticsearch" | "mongodb" | "qdrant" | "milvus" | "pgvector", default faiss.
Vector store backend.
RETRIEVAL_MAX_PARALLEL_SOURCES
Type int, default 4, must be >= 1.
Concurrent per-source searches in one retrieval; the query is embedded once and shared.
PER_SOURCE_RETRIEVAL_ENABLED
Type bool, default true.
Kill-switch for per-source retrieval dispatch; False collapses to a single retriever.
GRAPHRAG_ENABLED
Type bool, default false.
Gates graph-aware ingestion and retrieval.
GRAPHRAG_EXTRACTION_MODEL
Type str, default unset.
Model for ingest-time graph extraction; unset reuses LLM_PROVIDER/LLM_NAME.
GRAPHRAG_MAX_CHUNKS_FOR_EXTRACTION
Type int, default 2000, must be >= 0.
Hard cap on chunks extracted per source (cost control); 0 extracts nothing.
GRAPHRAG_EXTRACTION_WORKERS
Type int, default 8, must be >= 1 and <= 32.
Concurrent extraction calls during ingest. Model calls run in parallel while graph writes stay serial, so ordering and idempotency are unchanged; 1 is fully serial.
Vector stores
Per-backend connection details; only the backend named by VECTOR_STORE is read.
MONGO_URI
Type str, default unset.
Only consulted when VECTOR_STORE=mongodb or when running scripts/db/backfill.py; user data lives in Postgres.
ELASTIC_CLOUD_ID
Type str, default unset.
Elastic Cloud id.
ELASTIC_USERNAME
Type str, default unset.
Elasticsearch username.
ELASTIC_PASSWORD
Type str, default unset.
Elasticsearch password.
ELASTIC_URL
Type str, default unset.
Elasticsearch URL.
ELASTIC_INDEX
Type str, default docsgpt.
Elasticsearch index name.
QDRANT_COLLECTION_NAME
Type str, default docsgpt.
Qdrant collection name.
QDRANT_LOCATION
Type str, default unset.
Qdrant location (‘:memory:’ or a URL).
QDRANT_URL
Type str, default unset.
Qdrant server URL.
QDRANT_PORT
Type int, default 6333.
Qdrant REST port.
QDRANT_GRPC_PORT
Type int, default 6334.
Qdrant gRPC port.
QDRANT_PREFER_GRPC
Type bool, default false.
Use gRPC instead of REST where possible.
QDRANT_HTTPS
Type bool, default unset.
Use HTTPS for the Qdrant connection.
QDRANT_API_KEY
Type str, default unset.
Qdrant API key.
QDRANT_PREFIX
Type str, default unset.
URL prefix for a Qdrant behind a proxy.
QDRANT_TIMEOUT
Type float, default unset.
Qdrant request timeout in seconds.
QDRANT_HOST
Type str, default unset.
Qdrant host (alternative to QDRANT_URL).
QDRANT_PATH
Type str, default unset.
Path for an embedded on-disk Qdrant.
QDRANT_DISTANCE_FUNC
Type str, default Cosine.
Qdrant distance function.
PGVECTOR_CONNECTION_STRING
Type str, default unset.
pgvector connection string. postgres://, postgresql:// and postgresql+psycopg:// are all accepted and normalized internally for psycopg.connect(). Unset falls back to POSTGRES_URI.
PGVECTOR_POOL_MAX_SIZE
Type int, default 8, must be >= 0.
Per-process connection pool size; 0 uses one direct connection per store.
PGVECTOR_IVFFLAT_PROBES
Type int, default unset.
IVFFlat probes; unset derives sqrt(lists) from the index. Higher means better recall, more scan.
MILVUS_COLLECTION_NAME
Type str, default docsgpt.
Milvus collection name.
MILVUS_URI
Type str, default <DOCSGPT_HOME>/milvus_local.db.
Milvus server URI. The default is a milvus-lite (embedded) database file under the data home, like the other local stores.
MILVUS_TOKEN
Type str, default "".
Milvus auth token.
User-data database
The Postgres database holding users, conversations and sources, and what startup may do to it.
POSTGRES_URI
Type str, default unset.
User-data Postgres connection URI.
AUTO_MIGRATE
Type bool, default true.
On startup, apply pending Alembic migrations. Disable if you manage schema out-of-band.
AUTO_CREATE_DB
Type bool, default true.
On startup, create the target Postgres database if missing (needs CREATEDB privilege).
AUTO_VECTOR_SCHEMA
Type bool, default true.
On startup, create the pgvector/graph tables and verify the embedding dimension. No Alembic migration covers the vector DB (it may be a separate cluster); set False to manage it yourself.
Workers
How background tasks are queued and how worker processes are recycled.
CELERY_BROKER_URL
Type str, default redis://localhost:6379/0.
Celery broker URL.
CELERY_RESULT_BACKEND
Type str, default redis://localhost:6379/1.
Celery result backend URL.
CELERY_WORKER_PREFETCH_MULTIPLIER
Type int, default 1.
Tasks prefetched per worker process; 1 caps SIGKILL loss to one task.
CELERY_VISIBILITY_TIMEOUT
Type int, default 3600, must be > 0.
Broker visibility timeout in seconds. Must exceed the longest legitimate task runtime but stay short enough that SIGKILLed tasks redeliver promptly.
CELERY_WORKER_MAX_MEMORY_PER_CHILD
Type int, default 4194304, must be >= 0.
Recycle a prefork child past this resident size in KB; backstops docling/torch heap growth. Checked between tasks, so it does not bound the peak within one. 0 disables.
CELERY_WORKER_MAX_TASKS_PER_CHILD
Type int, default 0, must be >= 0.
Recycle a worker child after N tasks; 0 disables.
API_URL
Type str, default http://localhost:7091.
Address of the API. The worker hands finished indexes to it here, and the API builds the agent image, agent webhook, device pairing and MCP OAuth callback links it hands out from it, so on the API set it to the address browsers use. Docker Compose sets the worker’s to http://backend:7091 .
WORKER_API_URL
Type str, default unset.
Address the worker uses for its own calls into the API (handing over finished indexes). Unset falls back to API_URL. Set it when the API and the worker share one settings file and API_URL is a public address, e.g. http://127.0.0.1:7091 ; docsgpt up --native does.
Ingestion and parsing
Upload limits, the parser engine, and per-format byte caps for ingestion and attachments.
UPLOAD_FOLDER
Type str, default inputs.
Directory under the data home for uploaded sources.
UPLOAD_MAX_REQUEST_BYTES
Type int, default 268435456, must be > 0.
Cap on an upload request body; applied by Flask before multipart parsing.
UPLOAD_MAX_FILE_BYTES
Type int, default 104857600, must be > 0.
Cap on a single uploaded file; also enforced while copying.
PARSE_SPEC_MAX_BYTES
Type int, default 10485760, must be > 0.
Cap on an OpenAPI/tool spec file accepted for parsing.
UPLOAD_MAX_ARCHIVE_BYTES
Type int, default 262144000, must be > 0.
Cap on total bytes extracted from one uploaded archive.
UPLOAD_MAX_ARCHIVE_FILES
Type int, default 10000, must be > 0.
Cap on files extracted from one uploaded archive.
UPLOAD_MAX_ARCHIVE_RATIO
Type int, default 1000, must be > 0.
Maximum decompressed-to-compressed ratio before an archive is rejected.
UPLOAD_MAX_ARCHIVE_DEPTH
Type int, default 3, must be >= 0.
Maximum nesting depth of archives inside archives.
ATTACHMENT_ARCHIVE_MAX_MEMBERS
Type int, default 200, must be > 0.
Files unpacked from one zip attachment (nested archives included); the rest are skipped.
ATTACHMENT_ARCHIVE_MAX_ENTRIES
Type int, default 5000, must be > 0.
Entries looked at in one zip attachment, nested archives and skipped members included; the rest are skipped unread. Bounds the work a zip of many tiny or unsupported entries can cause, separately from the file limit.
ATTACHMENT_ARCHIVE_MAX_BYTES
Type int, default 209715200, must be > 0.
Total uncompressed bytes unpacked from one zip attachment; members past it are skipped.
ATTACHMENT_ARCHIVE_MAX_DEPTH
Type int, default 2, must be >= 1.
Archive levels unpacked from a zip attachment (2 = a zip inside the zip); deeper ones are skipped.
ATTACHMENT_ARCHIVE_MAX_RATIO
Type int, default 100, must be > 0.
Uncompressed-to-compressed ratio above which a zip attachment is rejected as a zip bomb.
ATTACHMENT_ARCHIVE_PARALLELISM
Type int, default 4, must be > 0.
Members of one zip attachment parsed at the same time, each as its own worker task; the next member is queued as one finishes, so a large zip never floods the queue ahead of other uploads.
ATTACHMENT_ARCHIVE_MEMBER_TIMEOUT
Type int, default 5400, must be > 0.
Seconds a zip attachment’s member may stay unparsed after it is queued (queue wait included) before the reconciler marks it failed, so a lost task never leaves the zip processing forever. A member whose task is running (its lease heartbeat is live) is never failed. Keep it above CELERY_VISIBILITY_TIMEOUT, after which the broker redelivers a task whose worker died.
PARSE_PDF_AS_IMAGE
Type bool, default false.
Render PDF pages to images before parsing.
PARSE_IMAGE_REMOTE
Type bool, default false.
Send images to a remote parser.
DOC_PARSER_ENGINE
Type "anydoc" | "docling", default anydoc.
Document parser for source ingestion, chat attachments and the read_document tool. “anydoc” (default): firecrawl-anydoc, a Rust converter with no ML models; milliseconds per file, ~100 MB peak RSS. “docling”: the layout/table-model pipeline (optional install; needed for read_document’s structured output and the docling OCR backend). Files anydoc cannot convert (scanned PDFs, malformed input) fall back to docling when it is installed, otherwise to the native OCR parsers (OCR on) or the legacy parsers. Rollback to the previous behaviour is this one variable.
DOCLING_PIPELINE_QUEUE_MAX_SIZE
Type int, default 2.
Pages docling’s threaded pipeline buffers in flight; the library default (100) drives worker RSS to ~3 GB on a mid-size PDF.
DOCLING_COMPILE_TORCH_MODELS
Type bool, default false.
Let docling torch.compile its models (slower start, faster pages).
DOCLING_TABULAR_MAX_BYTES
Type int, default 2000000.
Largest CSV/XLSX docling will parse, in bytes.
DOCLING_MARKUP_MAX_BYTES
Type int, default 8000000.
Largest HTML/XML docling will parse, in bytes.
MARKUP_MAX_BYTES
Type int, default 8000000, must be >= 0.
HTML/XHTML larger than this (bytes) are head-truncated before the markdownify parser runs (the anydoc engine’s HTML path). The tree that path builds costs ~50x the input (30 MB of HTML measured at 1.6 GB RSS) and the upload cap is 100 MB, so the gate is what keeps one upload from taking the ingest worker down. 0 disables it.
PDF_TRUST_CHECK
Type bool, default true.
Trust-check anydoc’s PDF output (docsgpt/parser/file/pdf_trust.py): flag composite (Type0) fonts without a ToUnicode map, and CJK-declaring PDFs whose extracted text has almost no CJK, the two classes where anydoc drops text silently. A flagged file re-parses on the docling fallback when docling is installed; otherwise the anydoc output is kept and the document gets extra_info[“parse_warnings”]. ~30 ms per scanned MB.
ANYDOC_TABLEIZE
Type bool, default false.
Rewrite dot-leader / whitespace-aligned table runs in anydoc’s PDF markdown into GFM tables (docsgpt/parser/file/tableize.py). Off by default: it rewrites content on a heuristic (>=3 uniform label+numbers lines) validated only on a small corpus so far.
ATTACHMENT_PDF_TEXT_FAST_PATH
Type bool, default true.
Read PDF attachments via their embedded text layer (pypdfium2) instead of docling, falling back to docling when there is no text layer. Attachments go into a prompt, so docling’s structural markdown earns far less than the tens of seconds per file it costs; source ingestion is unaffected because chunking and retrieval do depend on that structure.
ATTACHMENT_PDF_TEXT_MIN_MEDIAN_CHARS
Type int, default 32.
Median chars per sampled page below which a PDF attachment is treated as a scan and handed to docling. Measured on real uploads: scans at 0-17 chars/page, text-layer documents at 433-6834.
ATTACHMENT_TEXT_MAX_BYTES
Type int, default 5000000.
Cap on extracted attachment text.
ATTACHMENT_FULL_TEXT_MAX_BYTES
Type int, default 8000000, must be >= 0.
An attachment’s stored text is cut at 100k tokens for the prompt; when it is, the worker also keeps up to this many bytes of the whole extracted text next to the original file, so the attachments tool can search and read past the cut. The tool never loads a larger side copy. 0 keeps none.
AGENT_IMAGE_MAX_BYTES
Type int, default 5000000.
Cap on an image passed to an agent.
AGENT_IMAGE_MAX_PIXELS
Type int, default 16777216.
Cap on the pixel count of an image passed to an agent.
GITHUB_INGEST_MAX_FILE_BYTES
Type int, default 1048576, must be >= 0.
Skip GitHub repo blobs larger than this (0 = no cap).
GITHUB_INGEST_MAX_WORKERS
Type int, default 8, must be >= 1.
Parallel file fetches per GitHub repo ingest.
DOCUMENT_PARSE_QUEUE
Type str, default parsing.
Celery queue the parse_document task is routed to.
DOCUMENT_PARSE_TIMEOUT
Type int, default 120.
Seconds the read_document tool awaits the enqueued parse before degrading.
DOCUMENT_PARSE_TIMEOUT_PER_MB
Type int, default 60.
Extra seconds of parse window per MiB of input. The base timeout is a FLOOR: the window grows with document size because OCR cost scales with pages. Without this a large scan is silently dropped at the base window.
DOCUMENT_PARSE_TIMEOUT_MAX
Type int, default 900.
Absolute ceiling on the size-scaled parse window, in seconds.
DOCUMENT_PARSE_MAX_BYTES
Type int, default 0, must be >= 0.
Cap on a parsed document’s bytes (0 = reuse SANDBOX_MAX_INPUT_BYTES).
DOCUMENT_MAX_DECOMPRESSED_BYTES
Type int, default 314572800.
Cap on bytes decompressed from an archive handed to read_document.
DOCUMENT_MAX_ARCHIVE_ENTRIES
Type int, default 10000.
Cap on entries in an archive handed to read_document.
OCR
Whether OCR runs, which stack performs it, and which engine it uses.
OCR_ENABLED
Type bool, default false, also read from DOCLING_OCR_ENABLED.
OCR scanned PDFs and images during source ingestion.
OCR_ATTACHMENTS_ENABLED
Type bool, default false, also read from DOCLING_OCR_ATTACHMENTS_ENABLED.
OCR scanned PDFs and images attached to a chat.
OCR_BACKEND
Type "auto" | "docling" | "native", default auto.
Which stack runs OCR when it is on (OCR_ENGINE=deepseek always uses native). auto: docling when installed, otherwise native. docling: the layout-model pipeline (hybrid region OCR, reading order, table structure); needs the optional docling extra. native: pypdfium2/Pillow page rendering straight into tesseract or a DeepSeek-OCR endpoint (docsgpt/parser/file/ocr_parser.py); no ML models in the worker, tables come out as text lines under tesseract.
OCR_ENGINE
Type "tesseract" | "deepseek" | "auto" | "ocrmac" | "rapidocr", default tesseract.
OCR engine used when OCR is on. Benched 2026-08 on EN/ZH/table/degraded scans (docs page Sources/ocr has the menu). tesseract (recommended): best classic-engine accuracy (perfect EN word recall, 0.000 bilingual CER, 100% table cells), ~35 MB, CPU-only; needs the system binary and language packs, an optional install like every OCR dependency (build with INSTALL_TESSERACT=true, or apt/brew install tesseract-ocr for a local run); both backends. deepseek: DeepSeek-OCR on a local Ollama/vLLM server or a hosted API (OCR_DEEPSEEK_*); best table/CJK quality, the worker stays light (no layout models) but each page costs seconds on the model server; always runs on the native backend, whatever OCR_BACKEND says; docling never OCRs with it. auto: docling’s pick, ocrmac on macOS (excellent), rapidocr on Linux (silently shreds some long text lines; avoid as a server default). ocrmac | rapidocr: force one of those. auto/ocrmac/rapidocr exist only inside docling; the native backend runs tesseract for them. An engine that is not installed degrades (docling: to auto) with a warning instead of failing the parse.
OCR_LANGS
Type str, default eng.
Tesseract language packs, ”+“-separated (e.g. “eng+chi_sim+deu”). Other engines keep their own defaults; their language codes differ.
OCR_DEEPSEEK_PROVIDER
Type "ollama" | "vllm" | "novita" | "deepinfra" | "custom", default ollama.
Where DeepSeek-OCR runs; each preset supplies the endpoint URL, model and concurrency. ollama: a local Ollama (deepseek-ocr:3b). vllm: a vLLM server on localhost:8000 (deepseek-ai/DeepSeek-OCR). novita: Novita’s hosted API (deepseek/deepseek-ocr-2). deepinfra: DeepInfra’s hosted API (deepseek-ai/DeepSeek-OCR). custom: no defaults; set OCR_DEEPSEEK_URL and OCR_DEEPSEEK_MODEL. The hosted presets need OCR_DEEPSEEK_API_KEY and send every scanned page to that provider.
OCR_DEEPSEEK_URL
Type str, default unset.
Chat-completions URL of the DeepSeek-OCR endpoint. Unset uses the OCR_DEEPSEEK_PROVIDER preset’s URL (Ollama’s http://localhost:11434/v1/chat/completions by default); a value here always wins.
OCR_DEEPSEEK_MODEL
Type str, default unset.
Model name at the DeepSeek-OCR endpoint. Unset uses the OCR_DEEPSEEK_PROVIDER preset’s model.
OCR_DEEPSEEK_API_KEY
Type str, default unset.
Bearer token sent to the DeepSeek-OCR endpoint. Required by the novita and deepinfra presets (novita falls back to NOVITA_API_KEY); local servers need none.
OCR_DEEPSEEK_CONCURRENCY
Type int, default unset, must be >= 1 and <= 32.
Page requests in flight per file. Unset: 4 for the novita, deepinfra and vllm presets, 1 for ollama and custom, since a laptop-hosted model only slows down under parallel requests.
OCR_DEEPSEEK_MAX_RETRIES
Type int, default 3, must be >= 0 and <= 10.
Retries per page after a rate limit (429), a 5xx or a refused connection, with exponential backoff that honours Retry-After. Read timeouts are not retried.
OCR_DEEPSEEK_TIMEOUT
Type float, default 300.0.
Seconds allowed per page request to the DeepSeek endpoint. A 3B model on a laptop needs minutes; a vLLM GPU deployment or a hosted API, seconds.
OCR_DEEPSEEK_PROMPT
Type str, default Free OCR..
Instruction sent with every page image. ‘Free OCR.’ kept 94-100% of the words on real pages in testing and returns tables as Markdown; ‘Convert the document to markdown.’ lost most of a table page on Ollama, whose server strips the HTML cell tags that prompt produces. ’<|grounding|>…’ prompts add bounding boxes, which the output cleanup removes.
OCR_RENDER_DPI
Type int, default 200.
Native backend only: resolution at which pages without a text layer are rendered before OCR. 200 suits tesseract; clamped to 72-600.
OCR_MIN_CHARS_PER_PAGE
Type int, default 20, must be >= 0, also read from DOCLING_OCR_MIN_CHARS_PER_PAGE.
Chars-per-page floor below which an OCR’d PDF/image parse is treated as an OCR dropout rather than as content (long-running docling workers were observed returning zero characters for every scanned page after a long scanned PDF, with no error). docling retries once on a fresh full-page-OCR converter; both backends then fail loudly instead of indexing an empty document. 0 disables the guard.
File storage
Local disk or an S3-compatible bucket, and how download URLs are produced.
STORAGE_TYPE
Type "local" | "s3", default local.
File storage backend.
URL_STRATEGY
Type "backend" | "s3", default backend.
How download links are produced: backend (streamed through the API) or s3 (presigned URLs).
S3_BUCKET_NAME
Type str, default docsgpt-test-bucket.
Bucket name.
S3_ENDPOINT_URL
Type str, default unset.
Custom endpoint for S3-compatible services (MinIO, R2, B2, Spaces); omit for AWS.
S3_ACCESS_KEY_ID
Type str, default unset.
Access key id.
S3_SECRET_ACCESS_KEY
Type str, default unset.
Secret access key.
S3_REGION
Type str, default unset.
AWS region; use “auto” for Cloudflare R2.
S3_PATH_STYLE
Type bool, default false.
Path-style addressing (required by most non-AWS services).
SAGEMAKER_REGION
Type str, default unset.
Deprecated. Set S3_REGION instead; the SAGEMAKER_* fallback will be removed.
Legacy AWS region from the retired SageMaker provider; deprecated fallback for S3_REGION.
SAGEMAKER_ACCESS_KEY
Type str, default unset.
Deprecated. Set S3_ACCESS_KEY_ID instead; the SAGEMAKER_* fallback will be removed.
Legacy AWS access key from the retired SageMaker provider; deprecated fallback for S3_ACCESS_KEY_ID.
SAGEMAKER_SECRET_KEY
Type str, default unset.
Deprecated. Set S3_SECRET_ACCESS_KEY instead; the SAGEMAKER_* fallback will be removed.
Legacy AWS secret key from the retired SageMaker provider; deprecated fallback for S3_SECRET_ACCESS_KEY.
Connectors
Client credentials and callback URLs for Google Drive, Microsoft, Confluence, GitHub and MCP.
GOOGLE_CLIENT_ID
Type str, default unset.
Google OAuth client id.
GOOGLE_CLIENT_SECRET
Type str, default unset.
Google OAuth client secret.
CONNECTOR_REDIRECT_BASE_URI
Type str, default http://127.0.0.1:7091/api/connectors/callback.
OAuth callback URL; register it as-is in your provider’s console (e.g. GCP). Either the API’s /api/connectors/callback, which forwards to the app, or the app’s own <app origin>/connectors/callback.
CONNECTOR_ALLOWED_ORIGINS
Type str, default unset.
Comma-separated frontend origins connector sign-ins may start from and return to, e.g. https://docsgpt.example.com . The origin of CONNECTOR_REDIRECT_BASE_URI and OIDC_FRONTEND_URL are always allowed; a loopback callback also allows localhost:5173. The request’s Host never is.
MICROSOFT_CLIENT_ID
Type str, default unset.
Azure AD application (client) id.
MICROSOFT_CLIENT_SECRET
Type str, default unset.
Azure AD application client secret.
MICROSOFT_TENANT_ID
Type str, default common.
Azure AD tenant id, or ‘common’ for multi-tenant.
MICROSOFT_AUTHORITY
Type str, default unset.
Authority URL override; unset derives https://login.microsoftonline.com/<MICROSOFT_TENANT_ID> .
CONFLUENCE_CLIENT_ID
Type str, default unset.
Confluence Cloud OAuth client id.
CONFLUENCE_CLIENT_SECRET
Type str, default unset.
Confluence Cloud OAuth client secret.
GITHUB_ACCESS_TOKEN
Type str, default unset.
Instance-wide GitHub token for the public-repository upload. It raises GitHub’s rate limit and is never used to read a private repository; users connect their own GitHub account for those.
GITHUB_CLIENT_ID
Type str, default unset.
GitHub App client id. With the secret and slug, offers Sign in with GitHub next to tokens.
GITHUB_CLIENT_SECRET
Type str, default unset.
GitHub App client secret.
GITHUB_APP_SLUG
Type str, default unset.
GitHub App URL name (github.com/apps/<slug>), for the link where users choose repositories.
MCP_OAUTH_REDIRECT_URI
Type str, default unset.
Public callback URL for MCP OAuth; unset derives it from CONNECTOR_REDIRECT_BASE_URI.
Server
Serving the UI, public URLs, and process-level knobs of the API server.
DEPLOYMENT_TYPE
Type str, default unset.
Deployment class: cloud or production makes a missing JWT_SECRET_KEY fatal at startup. Otherwise the API generates .jwt_secret_key in the data home, which suits a single local process only.
SERVE_UI
Type bool, default true.
Serve the web UI shipped in the package (docsgpt/static) from the API process.
FLASK_DEBUG_MODE
Type bool, default false.
Run Flask in debug mode.
VERSION_CHECK
Type bool, default true.
Outbound version check for security advisories: the worker sends its version, a random instance id, the Python version and the platform to gptcloud.arc53.com when it starts and every 7 hours (reusing a recent cached answer), and logs any advisory. Set false to turn it off.
PUBLIC_API_BASE_URL
Type str, default unset.
Public base URL for user-facing endpoint references in prompts.
GRACEFUL_SHUTDOWN_TIMEOUT_SECONDS
Type int, default 30.
Bounds uvicorn’s shutdown drain (uvicorn_worker doesn’t forward —graceful-timeout). Keep below the gunicorn —timeout (180) watchdog. Used by BoundedDrainUvicornWorker.
WSGI_THREADPOOL_WORKERS
Type int, default 96, must be >= 1.
Threads serving the WSGI (Flask) part of the app under the ASGI server.
V1_SESSION_TTL_SECONDS
Type int, default 86400.
Lets OpenAI-compatible clients identify a logical chat by session header, which chat-completions itself has no field for; TTL of that session mapping.
Events and devices
The internal push channel (notifications and durable replay) and the Redis pool behind it.
ENABLE_SSE_PUSH
Type bool, default true.
Internal SSE push channel (notifications and durable replay journal). False makes /api/events emit “push_disabled” and return; clients fall back to polling.
EVENTS_STREAM_MAXLEN
Type int, default 1000, must be >= 1.
Per-user durable backlog cap in entries; ~24h of replay at typical rates.
SSE_KEEPALIVE_SECONDS
Type int, default 15, must be >= 1.
Interval between SSE keepalive comments.
SSE_MAX_CONCURRENT_PER_USER
Type int, default 8, must be >= 0.
Simultaneous SSE connections per user; each holds a pooled async Redis connection for its lifetime. 8 covers multi-tab use without one user starving the pool. 0 disables.
ASYNC_REDIS_MAX_CONNECTIONS
Type int, default 2000, must be >= 1.
Pool size of the async Redis client behind the event-loop routes, per process. Every open notification tab, chat reconnect and device session holds one connection, so this caps concurrent streams per worker (redis-py’s own default is 100). Keep the total across workers below the Redis server’s maxclients (10000 by default).
EVENTS_REPLAY_MAX_PER_REQUEST
Type int, default 200, must be >= 1.
Backlog entries XRANGE returns per /api/events snapshot. Bounds what one replay moves from Redis to the wire: a client looping Last-Event-ID reconnects enumerates at most this many per round-trip.
EVENTS_REPLAY_MAX_AGE_HOURS
Type int, default 48.
Oldest backlog entry a replay will return.
EVENTS_REPLAY_BUDGET_REQUESTS_PER_WINDOW
Type int, default 30.
Sliding-window cap on snapshot replays per user; exhausting it returns 429 with the cursor pinned so the client backs off until the window rolls over.
EVENTS_REPLAY_BUDGET_WINDOW_SECONDS
Type int, default 60.
Length of the replay budget window.
MESSAGE_EVENTS_RETENTION_DAYS
Type int, default 14, must be > 0.
Retention for the message_events journal, enforced by the cleanup_message_events beat task. Replay only needs streams a client could still be tailing.
REMOTE_DEVICE_SESSION_IDLE_SECONDS
Type int, default 60, must be > 0.
Seconds without a heartbeat before a remote-device session is considered idle.
REMOTE_DEVICE_REQUIRE_SIGNATURE
Type bool, default false.
Require signed commands from remote devices.
REMOTE_DEVICE_PAIRING_TTL_SECONDS
Type int, default 600, must be > 0.
Lifetime of a pairing code.
REMOTE_DEVICE_CMD_QUEUE_TTL_SECONDS
Type int, default 900, must be > 605.
Redis TTL of the per-device command queue, routing invocations cross-process so a scheduled run reaches the web-held device session. Must exceed the max drain deadline (605s) so a command for a briefly-offline device isn’t evicted before its own drain gives up.
REMOTE_DEVICE_INVOCATION_TTL_SECONDS
Type int, default 900, must be > 0.
Redis TTL of a pending remote-device invocation.
REMOTE_DEVICE_OUTPUT_STREAM_MAXLEN
Type int, default 10000.
Cap on buffered output entries per remote-device invocation stream.
Notifications
The VAPID key pair Web Push is signed with, and which push services may be sent to.
WEBPUSH_VAPID_PUBLIC_KEY
Type str, default unset.
VAPID public key for Web Push, raw base64url (the uncompressed P-256 point, as npx web-push generate-vapid-keys prints it). Browsers subscribe with it. Derived from the private key when unset; when both are set they must belong together.
WEBPUSH_VAPID_PRIVATE_KEY
Type str, default unset.
VAPID private key for Web Push, raw base64url. Unset, Web Push is off: users with no tab open get only the unread mark. Changing the key pair invalidates every browser’s subscription; browsers subscribe again on their next visit.
WEBPUSH_VAPID_SUBJECT
Type str, default unset.
Contact the push services can reach the operator at: a mailto: or https: URL. Apple rejects pushes without a real one. Unset, the public API URL is used when it is https.
WEBPUSH_EXTRA_ALLOWED_HOSTS
Type list[str], default [].
Push service hosts to accept besides the known ones (Google FCM, Mozilla, Apple, Windows), as a comma list or JSON list. *.example.com matches subdomains. Subscription endpoints must be https on an allowed host, so the server never posts to an arbitrary URL; this is for self-hosted push services and tests.
Agents
What an agent may do per turn and how its context is kept within budget.
AGENT_NAME
Type str, default classic.
Default agent type for agentless chats.
DEFAULT_AGENT_LIMITS
Type dict[str, int], default {"token_limit": 50000, "request_limit": 500}.
Per-agent default quotas: tokens and requests.
DEFAULT_CHAT_TOOLS
Type list[str], default ["memory", "read_webpage", "scheduler", "monitor"].
Config-free tools on by default in agentless chats, as a JSON list of tool names ([“memory”,“scheduler”]) or comma-separated names. none (or []) turns them all off; an empty value keeps the default. scheduler is dual-registered in BUILTIN_AGENT_TOOLS so one synthetic id resolves via defaults or the agent picker. Add code_executor and artifact_generator once a sandbox runner is configured; both execute through it and would fail on every call without one. check_job is never listed or toggled (an entry is ignored): every turn that can hand calls off to a background job gets it. monitor is offered only while MONITORS_ENABLED is on.
WEBHOOK_RUN_TIMEOUT
Type int, default 600, must be > 0.
Wall-clock cap on one agent webhook run, in seconds (at least 30). Past it the run stops and its task ends with a “timeout” result; the worker kills it 60 seconds later if it has not stopped by then.
ENABLE_TOOL_PREFETCH
Type bool, default true.
Pre-fetch retrieval before the agent’s first turn.
TOOL_RESULT_MAX_TOKENS
Type int, default 20000, must be >= 0.
Cap on one tool result entering the LLM context (0 disables); journal and DB keep it whole.
ATTACHMENT_BUDGET_SHARE
Type float, default 0.5, must be > 0 and <= 1.
Largest fraction of the model’s context window a turn’s attached files may take. Files that do not fit are listed in the turn’s manifest and, past the first partial one, left to tools.
ATTACHMENT_MAX_NATIVE_PARTS
Type int, default 40, must be >= 0.
Native file parts (images, PDFs, PDF page images) sent per turn; past the cap files go as extracted text or are left out. Keep it below the provider’s per-request image limit.
ENABLE_CONVERSATION_COMPRESSION
Type bool, default true.
Compress long conversations once they approach the context window.
COMPRESSION_THRESHOLD_PERCENTAGE
Type float, default 0.8, must be > 0 and <= 1.
Fraction of the context window at which compression triggers.
COMPRESSION_MODEL_OVERRIDE
Type str, default unset.
Use a different model for compression; unset reuses the answer model.
COMPRESSION_PROMPT_VERSION
Type str, default v1.0.
Tracks compression prompt iterations.
COMPRESSION_MAX_HISTORY_POINTS
Type int, default 3.
Keep only the last N compression points to prevent DB bloat.
COMPRESSION_RECENT_FIELD_MAX_TOKENS
Type int, default 8000, must be >= 0.
Per-field cap on the verbatim tail kept after a compression point (0 disables).
WORKFLOW_NODE_NATIVE_MAX_FILES
Type int, default 5.
Files per node passed natively to the LLM; past the cap they are extracted to text or dropped, to bound context and cost. Re-uses SANDBOX_MAX_INPUT_BYTES per file.
WORKFLOW_NODE_EXTRACT_MAX_FILES
Type int, default 5.
Documents per node extracted via the parsing worker. Each issues a separate blocking parse; past the cap they are skipped with a truncation note.
WORKFLOW_NODE_EXTRACT_BUDGET_SECONDS
Type int, default 900.
Wall clock one node may spend on blocking parses, shared across all of them. Without it a node could serialize WORKFLOW_NODE_EXTRACT_MAX_FILES full windows on a web threadpool slot.
WORKFLOW_RUN_STALE_SECONDS
Type int, default 3600.
A run row is pre-created as running; a disconnect or crash can strand it there. The beat reaper fails runs still running past this. Generous so a long run is never cut off.
ARTIFACT_MAX_BYTES
Type int, default 52428800.
Cap on a single stored artifact version’s bytes (0 disables).
ARTIFACT_MAX_COUNT_PER_USER
Type int, default 5000.
Cap on artifacts a user may own (0 disables).
ARTIFACT_MAX_TOTAL_BYTES_PER_USER
Type int, default 5368709120.
Cap on a user’s total stored artifact bytes (0 disables).
Guardrails
Input/output checks every agent runs, and the floor no agent may weaken.
GUARDRAILS_ENABLED
Type bool, default true.
Master switch; False disables every stage.
GUARDRAILS_CHECKS_ENABLED
Type list[str], default [].
Allowlist of GuardrailCreator.checks keys, as a JSON list or comma-separated names; unset or empty means every registered check.
GUARDRAILS_FLOOR
Type dict[str, Any], default {}.
A GuardrailsConfig fragment every agent inherits and cannot weaken; agents may add controls or make an action stricter, never looser. “enabled” is required; without it the floor parses but applies to nothing. Example: {“enabled”: true, “mode”: “scan_all”, “controls”: [{“check”: “secrets”, “stage”: “output”, “action”: “redact”}]}
GUARDRAILS_JUDGE_MODEL
Type str, default unset.
Judge model for the topic/policy checks; unset reuses the request’s model.
GUARDRAILS_STORE_SCANNED_TEXT
Type bool, default false.
Persist scanned text alongside guardrail_events. Off by default: pre-redaction text is exactly the material a PII control exists to keep out of storage.
GUARDRAILS_EVENTS_RETENTION_DAYS
Type int, default 30, must be >= 1.
Days guardrail events are kept before the cleanup task removes them.
Execution traces
Recording of agent, LLM, tool and retrieval steps per request.
TRACES_ENABLED
Type bool, default true.
Record an execution trace (agent runs, LLM calls, tool calls, retrieval, embeddings) for every request, store it in request_traces and show it in the Logs UI. False records nothing.
TRACES_CAPTURE_CONTENT
Type bool, default true.
Store short, secret-redacted previews (tool arguments and results, retrieved chunk titles, rephrased queries, answer excerpts) with each stored trace. Full prompts are never stored. OTel export follows OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT instead.
TRACES_PREVIEW_CHARS
Type int, default 2000, must be >= 100.
Maximum characters kept per stored trace preview.
TRACES_MAX_SPANS
Type int, default 500, must be >= 10.
Maximum spans recorded per trace; further spans are counted as dropped, not stored.
TRACES_RETENTION_DAYS
Type int, default 30, must be >= 1.
Days stored traces are kept before the cleanup task removes them.
TRACES_OTEL_EXPORT
Type bool, default true.
Also emit each finished trace as OpenTelemetry GenAI spans (gen_ai.*) and metrics. Has no effect unless an OTel SDK is configured, e.g. by launching under opentelemetry-instrument.
Quotas
Quota window and the treatment of unpriced models.
QUOTA_PERIOD
Type "day" | "week" | "month", default month.
Window every usage quota is measured over. Windows are calendar-aligned in UTC: a day starts at 00:00, a week on Monday, a month on the 1st.
QUOTA_UNPRICED_RATE_PER_MILLION
Type list[float], default unset.
Fallback [input, output] USD rates per 1M tokens for models that declare no price, e.g. [0.5, 1.5] or 0.5,1.5. Unset, such calls are recorded at $0 and only count toward token quotas.
Scheduler
Cadence, quotas and timeouts of scheduled runs.
SCHEDULE_DISPATCHER_INTERVAL
Type int, default 30.
Seconds between dispatcher passes that enqueue due schedules.
SCHEDULE_MIN_INTERVAL
Type int, default 900.
Smallest allowed recurrence interval in seconds.
SCHEDULE_MAX_PER_USER
Type int, default 50.
Cap on schedules a user may own.
SCHEDULE_RUN_TIMEOUT
Type int, default 600.
Wall-clock cap on one scheduled run, in seconds.
SCHEDULE_MISFIRE_GRACE
Type int, default 60.
Seconds past the due time within which a missed run still fires.
SCHEDULE_AUTOPAUSE_FAILURES
Type int, default 3.
Consecutive failures after which a schedule is paused automatically.
SCHEDULE_ONCE_MAX_HORIZON
Type int, default 31536000.
How far ahead a one-off run may be scheduled, in seconds (one year).
SCHEDULE_RUN_OUTPUT_RETENTION_DAYS
Type int, default 90, must be > 0.
Days scheduled-run output is kept.
Background jobs
How long a turn waits on a tool call, what a handed-off job may use, and when the agent is resumed.
BACKGROUND_JOBS_ENABLED
Type bool, default true.
Hand a slow tool call off to a background job instead of keeping the turn waiting on it. Off, every call runs in the foreground until its own timeout, as before.
BACKGROUND_YIELD_SECONDS
Type int, default 30, must be >= 1.
Seconds a chat turn waits on a server-side tool call before handing it off to a background job. A call that finishes sooner behaves exactly as before.
BACKGROUND_JOB_MAX_SECONDS
Type int, default 1000, must be >= 1.
Hard lifetime of one background job in seconds; past it the job is failed. Matches SANDBOX_EXEC_MAX_TIMEOUT by default.
DEVICE_JOB_MAX_SECONDS
Type int, default 3600, must be >= 60.
Hard lifetime of a background job running a command on a paired remote device, in seconds. A command started with background=true may run this long (a foreground one keeps its 600 s cap); a job whose device never reports back before then is marked lost.
BACKGROUND_MAX_JOBS_PER_CONVERSATION
Type int, default 2, must be >= 0.
Running background jobs one conversation may have. Over the cap a slow call keeps running in the foreground; it never fails because of the cap.
BACKGROUND_MAX_JOBS_PER_USER
Type int, default 5, must be >= 0.
Running background jobs one user may have, across conversations; same over-cap rule.
BACKGROUND_RESULT_RETENTION_DAYS
Type int, default 7, must be > 0.
Days a finished background job and its result are kept.
BACKGROUND_LEASE_STALE_SECONDS
Type int, default 60, must be >= 20.
Heartbeat age after which a job whose process stopped reporting is marked lost. Jobs are stamped every 10 s, so this allows several missed beats.
BACKGROUND_RECONCILE_INTERVAL_SECONDS
Type int, default 60, must be >= 10.
Seconds between beat sweeps that mark lost jobs, fail jobs past their deadline and resume conversations whose results were not delivered.
BACKGROUND_POOL_SIZE
Type int, default 16, must be >= 1.
Threads per API or worker process that run tool calls a turn may hand off. A call that finds the pool full runs inline, as before.
AUTO_RESUME_ENABLED
Type bool, default true.
Resume the agent in the same conversation when a background job finishes, so the result reaches the user without them asking. Off, results wait for check_job or the user’s next message.
AUTO_RESUME_MAX_RESULT_CHARS
Type int, default 16000, must be >= 500.
Characters of one job result a continuation turn sees; longer results keep their head and tail.
AUTO_RESUME_MAX_CONSECUTIVE
Type int, default 5, must be >= 1.
Continuation turns a conversation may get in a row without a user message in between. Past it, results wait for check_job or the user’s next message, so a job that starts a job cannot loop.
Monitors and trigger links
What the monitor tool may watch, how often, for how long, and how trigger and approval links behave.
MONITORS_ENABLED
Type bool, default true.
Offer the monitor tool: polled checks of a webpage or any tool, ingest events, webhook trigger links and human approval links that resume the conversation when something happens. Needs AUTO_RESUME_ENABLED, since a monitor reports back by resuming the conversation.
MONITOR_JUDGE_MODEL
Type str, default unset.
Model that judges a monitor’s natural-language condition, only after its deterministic check passed (or on a change when it has none). Unset uses the deployment’s default model.
MONITOR_DEFAULT_INTERVAL_SECONDS
Type int, default 900, must be >= 60.
Seconds between checks of a polled monitor that names no interval.
MONITOR_MIN_INTERVAL_SECONDS
Type int, default 300, must be >= 60.
Shortest interval a polled monitor may use, in seconds; a shorter request is raised to it.
MONITOR_DEFAULT_TTL_DAYS
Type int, default 7, must be >= 1.
Days a monitor (and its trigger or approval link) lives when none is named.
MONITOR_MAX_TTL_DAYS
Type int, default 30, must be >= 1.
Longest lifetime a monitor may ask for, in days.
MONITOR_DEFAULT_MAX_WAKES
Type int, default 1, must be >= 1.
Times a monitor may wake the agent when none is named; at 0 left it finishes.
MONITOR_MAX_WAKES
Type int, default 20, must be >= 1.
Most wakes one monitor may ask for.
MONITOR_MAX_ACTIVE_PER_USER
Type int, default 5, must be >= 0.
Active or paused monitors one user may have; 0 turns creating them off.
MONITOR_MAX_WAKES_PER_HOUR
Type int, default 5, must be >= 1.
Circuit breaker: a monitor that would wake the agent more often than this in an hour is paused, and the conversation is told why.
MONITOR_UNREACHABLE_GRACE_SECONDS
Type int, default 3600, must be >= 60.
How long a monitor’s source may stay unreachable (device offline, server down, timeout, 5xx) before the agent is told once and the monitor pauses. Unreachable checks are skipped, never a change.
MONITOR_JUDGE_TOKEN_BUDGET
Type int, default 100000, must be >= 0.
Tokens one monitor’s condition judge may use over its lifetime; past it the monitor pauses and says so. 0 means no limit.
PUBLIC_APP_URL
Type str, default unset.
Public address of the web UI, used for human approval links (/approve/<token>). Unset uses PUBLIC_API_BASE_URL, then API_URL: the API serves the UI itself when SERVE_UI is on. Set it when the UI runs elsewhere, e.g. http://localhost:5173 for the Vite dev server.
TRIGGER_RATE_PER_MINUTE
Type int, default 30, must be >= 1.
Requests one trigger link accepts per minute; more get 429.
TRIGGER_MAX_PAYLOAD_BYTES
Type int, default 65536, must be >= 1024.
Largest request body a trigger link accepts, in bytes; larger gets 413.
TRIGGER_GET_MAX_HITS
Type int, default 100, must be >= 1.
Calls a webhook link that also accepts GET takes over its lifetime (a POST-only link takes 1000). Lower, since a GET link can be fired by anything that opens it.
TRIGGER_GET_DEFAULT_TTL_HOURS
Type int, default 24, must be >= 1.
Default lifetime, in hours, of a webhook link that also accepts GET, when monitor_create asks for none (a POST-only link defaults to MONITOR_DEFAULT_TTL_DAYS). An explicit expires_in still applies, up to MONITOR_MAX_TTL_DAYS.
TRIGGER_DEDUPE_WINDOW_SECONDS
Type int, default 600, must be >= 1.
Seconds in which a trigger link delivery with the same body as an earlier one (and no Idempotency-Key, webhook-id or X-GitHub-Delivery) counts as a repeat. The same body sent later, such as a nightly job’s, is a new event.
Sandbox
The app is a CLIENT of an always-on runner; defaults are safe so app import never fails unconfigured.
SANDBOX_BACKEND
Type "jupyter" | "daytona", default jupyter.
Sandbox backend: jupyter (self-host) or daytona (Daytona Cloud).
SANDBOX_GATEWAY_URL
Type str, default http://localhost:8888.
URL of the Jupyter Kernel Gateway runner (the docsgpt-sandbox service). Chat offers attached files to the code execution tool only when this is set explicitly; the default alone does not count.
SANDBOX_GATEWAY_AUTH_TOKEN
Type str, default unset.
Gateway auth token, if set.
SANDBOX_KERNEL_NAME
Type str, default docsgpt-python.
Kernelspec per session. The env-scrubbing docsgpt-python spec keeps kernel code from reading the gateway token or operator secrets from os.environ; the stock python3 spec inherits the gateway env verbatim and must not be used with untrusted code.
SANDBOX_MAX_TTL
Type int, default 1200.
Seconds an idle session is kept before it is closed, and the cap on a keep-alive TTL the model asks for. Code Executor sessions stay open between calls until then.
SANDBOX_MAX_SESSIONS
Type int, default 32.
Concurrent live sessions per process, backend-agnostic; at the cap an LRU-idle session is evicted. 0 or negative disables the cap.
SANDBOX_EXEC_TIMEOUT
Type int, default 60.
Default wall-clock cap (s) per exec call.
SANDBOX_EXEC_MAX_TIMEOUT
Type int, default 1000, must be >= 1.
Longest wall-clock cap (s) the model may ask for on one run_code call, for long jobs such as video renders or OCR of many pages; a larger request is clamped to it. Never below SANDBOX_EXEC_TIMEOUT. On the Jupyter runner keep SANDBOX_KERNEL_IDLE_TIMEOUT above SANDBOX_MAX_TTL; a busy kernel is never culled.
SANDBOX_HTTP_TIMEOUT
Type int, default 10.
Fixed cap (s) for REST control calls (create/delete/alive/interrupt).
SANDBOX_MAX_OUTPUT_BYTES
Type int, default 8388608.
Cap on buffered stdout+stderr per exec.
SANDBOX_MAX_FILE_BYTES
Type int, default 10485760.
Cap on get_file size routed through stdout.
SANDBOX_MAX_INPUT_BYTES
Type int, default 26214400.
Cap on an input document staged into a sandbox session.
SANDBOX_MEMORY
Type str, default 4g.
Docker mem_limit for the runner container: the gateway, the warm session kernels and the LibreOffice and Chromium processes they start. Consumed by the docsgpt-sandbox compose service, not the app; part of the untrusted-code security boundary.
SANDBOX_HOME_SIZE
Type str, default 1g.
Size of the runner’s /sandbox-home tmpfs: the kernels’ HOME, where runtime pip installs and caches go. It allows exec so compiled packages load (/tmp stays noexec) and counts against SANDBOX_MEMORY as it fills. Consumed by the docsgpt-sandbox compose service, not the app.
SANDBOX_CPUS
Type str, default 1.0.
Docker CPU quota for the runner container. Consumed by the docsgpt-sandbox compose service, not the app; part of the untrusted-code security boundary.
DAYTONA_API_KEY
Type str, default unset.
Daytona Cloud API key (secret).
DAYTONA_API_URL
Type str, default unset.
Override for the Daytona API base URL, if self-targeting.
DAYTONA_TARGET
Type str, default unset.
Daytona region/target, e.g. “us”.
DAYTONA_SNAPSHOT
Type str, default unset.
Snapshot for new sandboxes; build one with the sandbox’s libraries, tools and fonts via scripts/build_daytona_snapshot.py (default name docsgpt-sandbox-py312-v3).
DAYTONA_LANGUAGE
Type str, default python.
Default runtime language for created sandboxes.
DAYTONA_AUTO_STOP_INTERVAL
Type int, default 15, must be >= 0.
Minutes idle before Daytona auto-stops a sandbox (0 disables).
DAYTONA_AUTO_DELETE_INTERVAL
Type int, default 60, must be >= -1.
Minutes after stop before Daytona auto-deletes a sandbox (-1 disables).
DAYTONA_MAX_SANDBOXES
Type int, default 50.
Cap on concurrent live Daytona sandboxes (cost-DoS guard).
Speech
Voice providers and transcription options.
TTS_PROVIDER
Type "google_tts" | "elevenlabs" | "none", default google_tts.
Text-to-speech provider; none switches it off.
TTS_MAX_CHARS
Type int, default 10000.
Cap on the characters one text-to-speech request may speak, after markdown is stripped.
ELEVENLABS_API_KEY
Type str, default unset.
ElevenLabs API key.
ELEVENLABS_VOICE_ID
Type str, default nPczCjzI2devNBz1zQrb.
ElevenLabs voice to speak with; an ID from your ElevenLabs voice library.
STT_PROVIDER
Type "openai" | "faster_whisper" | "none", default openai.
Speech-to-text provider; none switches it off.
OPENAI_STT_MODEL
Type str, default gpt-4o-mini-transcribe.
OpenAI transcription model.
STT_LANGUAGE
Type str, default unset.
Language hint for transcription; unset auto-detects.
STT_MAX_FILE_SIZE_MB
Type int, default 50.
Cap on an audio file accepted for transcription.
STT_ENABLE_TIMESTAMPS
Type bool, default false.
Return word/segment timestamps.
STT_ENABLE_DIARIZATION
Type bool, default false.
Label speakers in the transcript.