Observability
DocsGPT bundles the OpenTelemetry SDK and auto-instrumentation packages
as core dependencies (pyproject.toml), so they install with the rest of
the backend and ship in the image. OpenTelemetry export is off by default; opt in by
prefixing the launch command with opentelemetry-instrument and setting
OTLP env vars.
Two other things are on by default. Execution traces
are stored locally in Postgres and leave the instance only through an
exporter you configure. The version check is outbound: when the
worker starts and every 7 hours it sends its version, a random
instance_id kept in the database, the Python version, the platform
(sys.platform) and a client name to https://gptcloud.arc53.com/api/check,
reusing a recent answer cached in Redis instead where it has one. Security
advisories in the answer appear in the worker log, and high or critical
ones also print a banner to the workerโs console. Turn it off with
VERSION_CHECK=0.
Auto-instrumentation covers Flask, Starlette, Celery, SQLAlchemy, psycopg, Redis, requests, and Python logging. Agent runs, LLM calls, tool calls and retrieval are recorded by DocsGPT itself and exported as OpenTelemetry GenAI spans โ see Execution traces.
Enabling
Set these env vars in your .env (or compose environment: block):
OTEL_SDK_DISABLED=false
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_EXPORTER_OTLP_ENDPOINT=https://your-collector.example.com
OTEL_EXPORTER_OTLP_HEADERS=Authorization=Bearer%20<token>
OTEL_TRACES_EXPORTER=otlp
OTEL_METRICS_EXPORTER=otlp
OTEL_LOGS_EXPORTER=otlp
OTEL_PYTHON_LOG_CORRELATION=true
OTEL_RESOURCE_ATTRIBUTES=service.name=docsgpt-backend,deployment.environment=prodThen prefix the process command with opentelemetry-instrument. The
simplest way is a Compose override file, with no image rebuild. Save this
as docker-compose.otel.yaml next to your Compose file:
# docker-compose.otel.yaml
services:
backend:
# The image's CMD from docsgpt/Dockerfile behind opentelemetry-instrument.
# Keep the rest in sync with the Dockerfile when you upgrade.
command:
- opentelemetry-instrument
- gunicorn
- -w
- "1"
- -k
- docsgpt.gunicorn_worker.BoundedDrainUvicornWorker
- --bind
- 0.0.0.0:7091
- --timeout
- "180"
- --graceful-timeout
- "120"
- --keep-alive
- "5"
- --worker-tmp-dir
- /dev/shm
- --max-requests
- "5000"
- --max-requests-jitter
- "500"
- --config
- docsgpt/gunicorn_conf.py
- docsgpt.asgi:asgi_app
environment:
- OTEL_SERVICE_NAME=docsgpt-backend
worker:
# The bundled worker command behind opentelemetry-instrument.
command: opentelemetry-instrument celery -A docsgpt.app.celery worker -l INFO -B -Q docsgpt,parsing,embeddings
environment:
- OTEL_SERVICE_NAME=docsgpt-celery-workerCompose loads an override file by itself only when no -f is given, and
the commands in these docs all pass -f, so name it after the main file
on every command, up included. For the checkout Compose files, with the override saved in
deployment/:
docker compose --env-file .env -f deployment/docker-compose-hub.yaml -f deployment/docker-compose.otel.yaml up -dFor the standalone file, run
docker compose -f docker-compose-standalone.yaml -f docker-compose.otel.yaml up -d
in its folder. A docsgpt up stack doesnโt know about the override:
docsgpt up, docsgpt restart and docsgpt upgrade start it without
tracing, so start it with
docker compose -f ~/.docsgpt/server/docker-compose.yaml -f ~/.docsgpt/server/docker-compose.otel.yaml up -d
afterwards.
For local dev, prepend dotenv run -- so the OTEL_* vars from .env
reach opentelemetry-instrument before it boots the SDK:
dotenv run -- opentelemetry-instrument uvicorn docsgpt.asgi:asgi_app --port 7091
dotenv run -- opentelemetry-instrument celery -A docsgpt.app.celery worker -l INFO -B --pool=soloTrace the ASGI app rather than flask run, which serves only the Flask app: the
ASGI-only routes return 404 there and
their spans never appear. -B keeps the beat scheduler running, as in production.
Logs are exported in-process when OTEL_LOGS_EXPORTER=otlp is set โ
docsgpt/core/logging_config.py detects the flag and preserves
the OTEL log handler. Without it, logging writes only to stdout.
Execution traces
Every request records an execution trace: a timed tree of the steps
behind it. Traces are recorded for chat turns (/stream, /api/answer,
/v1/chat/completions, including each round of a tool-approval pause),
scheduled and webhook runs, workflows, the research agent, /api/search,
the MCP search_docs tool, and graph builds.
| Step | Recorded when |
|---|---|
invoke_agent | An agent (or a workflow nodeโs agent) runs |
chat | An LLM call made during the request, including retries, fallbacks, query rephrasing, prescreening, history compression and guardrail judges |
execute_tool | A tool is executed, paused for approval, denied or skipped |
retrieval | A retriever or the multi-source dispatcher searches |
embeddings | The query is embedded |
search | One source is searched |
rerank | Prescreening filters retrieved chunks |
guardrail | A guardrail calls a remote check or fires |
step | A workflow node or research phase runs |
Traces are stored in the request_traces table and shown in the app: open
Settings โ Logs (or an agentโs Logs tab), expand an entry and choose
View trace to see a waterfall of every step with its timing, tokens,
cost and details. See Analytics and Logs for a
userโs guide to those pages.
Stored traces keep short previews โ tool arguments and results, retrieved chunk titles and snippets, rephrased queries, answer excerpts โ truncated and with secret-named fields redacted. Full prompts are never stored. When a guardrail fires during a request, every preview is dropped from its trace.
TRACES_ENABLED=true # record traces at all
TRACES_CAPTURE_CONTENT=true # keep previews in stored traces
TRACES_PREVIEW_CHARS=2000 # characters kept per preview
TRACES_MAX_SPANS=500 # steps kept per trace; the rest are counted
TRACES_RETENTION_DAYS=30 # a daily task deletes older traces
TRACES_OTEL_EXPORT=true # also export traces as OTel GenAI spansGenAI spans and metrics
When DocsGPT runs under opentelemetry-instrument, each finished trace is
also exported as spans that follow the
OpenTelemetry GenAI semantic conventionsย :
invoke_agent {agent}, chat {model}, execute_tool {tool},
embeddings {model} and retrieval, with attributes such as
gen_ai.provider.name, gen_ai.request.model,
gen_ai.usage.input_tokens, gen_ai.usage.output_tokens,
gen_ai.usage.cache_read.input_tokens, gen_ai.conversation.id,
gen_ai.agent.id and gen_ai.tool.name. DocsGPT-specific details use the
docsgpt.* prefix (docsgpt.request_id, docsgpt.token_source,
docsgpt.cache_hit, docsgpt.ttft_ms, and on a Responses API call that did not chain onto the previous response, docsgpt.chain_reset_reason, โฆ). The traceโs root span is a
child of the requestโs HTTP server span, and the stored trace keeps the
OTel trace id so you can move between the two.
Two metrics are recorded for every model call:
gen_ai.client.token.usage and gen_ai.client.operation.duration.
Backends that understand the GenAI conventions โ Langfuse
(/api/public/otel), Arize Phoenix, Datadog LLM Observability, Grafana โ
render these as LLM traces with token and cost views.
Prompt and tool content is not exported by default, because the OTLP backend may be a third party. Opt in with the standard variable:
OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=SPAN_ONLYThis adds the same redacted previews the app stores (for example
gen_ai.tool.call.arguments and gen_ai.tool.call.result).
GenAI spans are exported when the request finishes, with their original timestamps. Consequences: a long research run appears only when it ends; HTTP and database spans made during a step sit beside the step rather than under it; and log records carry the requestโs span ids, not the stepโs.
The GenAI conventions are still in development upstream, so attribute names may change in later releases.
Backend examples
Axiom
OTEL_EXPORTER_OTLP_ENDPOINT=https://api.axiom.co
OTEL_EXPORTER_OTLP_HEADERS=Authorization=Bearer%20xaat-XXXX,X-Axiom-Dataset=docsgpt
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf%20 is the URL-encoded space between Bearer and the token. Create
the dataset in the Axiom UI before sending.
Self-hosted OTLP collector / Jaeger / Tempo
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4317
OTEL_EXPORTER_OTLP_PROTOCOL=grpcHoneycomb / Grafana Cloud / Datadog
Each vendor publishes a single-line OTEL_EXPORTER_OTLP_ENDPOINT plus
OTEL_EXPORTER_OTLP_HEADERS recipe โ drop them in alongside the
service-name override.
Caveats
- The Dockerfile uses
gunicorn -w 1. If you raise worker count, move SDK init into apost_worker_inithook to avoid one-thread-per-process exporter contention. asgi.pymounts the Flask app inside a Starlette app through a2wsgiโsWSGIMiddleware. Both instrumentors are installed, so each request produces a Starlette span enclosing a Flask span. If the duplication is noisy, uninstallopentelemetry-instrumentation-flaskin your image, or setOTEL_PYTHON_DISABLED_INSTRUMENTATIONS=flask. Donโt editdocsgpt/requirements.txt: it is generated fromuv.lock.- OTEL packages add ~50 MB to the image. They install on every build โ
the runtime cost is zero unless you set
opentelemetry-instrumenton the command and set the OTLP env vars.