Skip to Content
Upgrading

Upgrading DocsGPT

Upgrading from 0.16.x? User data moved from MongoDB to Postgres in 0.17.0. Follow the Postgres Migration guide before running docker compose pull or git pull — existing deployments will not start cleanly without it.

To upgrade:

  1. Back up. docsgpt backup for an installer or docsgpt up Docker install whose docsgpt has that command (releases after 0.21.0; on 0.21.0 back up by hand). Otherwise dump Postgres (pg_dump) and copy your uploads and indexes.
  2. Read the upgrade notes for every release after the one you run, and the changelog.
  3. Follow the steps for your install, below.

Check your version

InstallCommand
Installer, docsgpt up, pipdocsgpt --version for the package; docsgpt status for the version the stack runs
Docker Compose (from the stack directory; from a checkout add --env-file .env -f deployment/docker-compose-hub.yaml)docker compose exec backend python -c "from docsgpt.version import get_version; print(get_version())"
Kuberneteskubectl exec deploy/docsgpt-api -- python -m docsgpt.cli --version

Release notes: changelog. Tags: GitHub releases .

Installer or docsgpt up

docsgpt backup # releases after 0.21.0; see below for 0.21.0 docsgpt upgrade # the latest release; --version <version> for another docsgpt status

For a uv tool install, which is what the installer sets up, docsgpt upgrade installs the new package and runs its docsgpt up, which moves the stack to that release’s images and keeps .env and the data. Running the install command again does the same; DOCSGPT_VERSION=<version> picks a release. With pipx or pip, docsgpt upgrade prints the command to run instead (pipx install --force docsgpt or pip install -U docsgpt); run it, then docsgpt up.

A docsgpt up --native install has nothing for docsgpt backup to archive: dump its Postgres with pg_dump and copy indexes/, inputs/ and vectors/ from the data home first.

Back up and restore by hand (0.21.0)

docsgpt backup and docsgpt restore are new after 0.21.0 (docsgpt --help shows whether yours has them). On 0.21.0, take the same backup by hand from the stack directory. It holds a dump of the database and a tar of each data volume (docsgpt_indexes, docsgpt_inputs, docsgpt_vectors). Keep a copy of .env too: it holds the install’s secrets.

cd ~/.docsgpt/server # /opt/docsgpt when installed as root on Linux IMAGE=arc53/docsgpt:0.21.0 # the image the stack runs mkdir -p backups/before-upgrade docker compose stop backend worker # so the dump and the files match docker compose exec -T postgres pg_dump --clean --if-exists -U docsgpt -d docsgpt > backups/before-upgrade/docsgpt.sql for v in indexes inputs vectors; do docker run --rm -v docsgpt_$v:/data:ro "$IMAGE" tar cf - -C /data . > backups/before-upgrade/$v.tar done docker compose start backend worker

To restore it, with the stack on the version the backup came from. Load the dump into an empty database: loaded over a database a newer release migrated, the newer tables’ foreign keys stop it.

First check that every tar reads, since the restore empties each volume before unpacking into it. Go on only if all three print ok:

cd ~/.docsgpt/server # /opt/docsgpt when installed as root on Linux for v in indexes inputs vectors; do tar -tf backups/before-upgrade/$v.tar > /dev/null && echo "$v.tar ok" done

Then restore:

IMAGE=arc53/docsgpt:0.21.0 docker compose down for v in indexes inputs vectors; do docker run --rm -i -u 0 -v docsgpt_$v:/data "$IMAGE" \ sh -c 'find /data -mindepth 1 -delete && tar xf - -C /data' < backups/before-upgrade/$v.tar done docker compose up -d --wait postgres docker compose exec -T postgres psql -U docsgpt -d postgres \ -c 'ALTER DATABASE docsgpt RENAME TO docsgpt_before_restore' -c 'CREATE DATABASE docsgpt OWNER docsgpt' docker compose exec -T postgres psql --set ON_ERROR_STOP=on -U docsgpt -d docsgpt < backups/before-upgrade/docsgpt.sql docker compose up -d

Once DocsGPT works, drop the set-aside copy: docker compose exec postgres psql -U docsgpt -d postgres -c 'DROP DATABASE docsgpt_before_restore'. If the load fails, drop the new docsgpt database and rename docsgpt_before_restore back before starting DocsGPT. If a restore stopped after the rename and DocsGPT has since started, the API will have created an empty docsgpt; drop that empty one (DROP DATABASE docsgpt WITH (FORCE)) before you rename docsgpt_before_restore back. docsgpt restore does all of this itself.

pip install

pip install -U docsgpt # or: pipx install --force docsgpt docsgpt migrate

Then restart docsgpt api and docsgpt worker (docsgpt restart for a docsgpt up --native install). See Install with pip.

Docker Compose — standalone file

Download the new release’s docker-compose-standalone.yaml (attached to every release ) over the old one. If .env pins DOCSGPT_IMAGE_TAG, set it to the new release; without it the file runs latest. Then, in that folder:

docker compose -f docker-compose-standalone.yaml pull docker compose -f docker-compose-standalone.yaml up -d --remove-orphans

Docker Compose — hub images

From the repository root:

git pull # or: git checkout <version> docker compose --env-file .env -f deployment/docker-compose-hub.yaml pull docker compose --env-file .env -f deployment/docker-compose-hub.yaml up -d

--env-file .env makes Compose read settings such as DOCSGPT_BIND and DOCSGPT_IMAGE_TAG from the repository root; without it they are ignored. This file runs the develop images, built from the main branch, unless .env sets DOCSGPT_IMAGE_TAG. To move to a release, set DOCSGPT_IMAGE_TAG=<version> in .env; don’t edit image: in the file, which would leave the frontend image on another tag. With the Ollama overlay, add its -f deployment/optional/... file to both commands.

Docker Compose — from source

cd DocsGPT git pull docker compose --env-file .env -f deployment/docker-compose.yaml build docker compose --env-file .env -f deployment/docker-compose.yaml up -d

Swap git pull for git checkout <tag> if you want to pin a specific release. Build settings such as EXTRAS=docling are read from .env through --env-file .env.

Kubernetes

The Kubernetes pods don’t migrate the database when they start; the postgres-init Job does. Change every arc53/docsgpt:<version> image in deployment/k8s/deployments/docsgpt-deploy.yaml and deployment/k8s/jobs/postgres-init-job.yaml to the new release, and arc53/docsgpt-sandbox:<version> in deployment/k8s/deployments/sandbox-deploy.yaml if you run the opt-in sandbox. Then:

kubectl delete job postgres-init --ignore-not-found kubectl apply -k deployment/k8s/ kubectl wait --for=condition=complete job/postgres-init --timeout=300s kubectl rollout status deployment/docsgpt-api kubectl rollout status deployment/docsgpt-worker

A finished Job can’t be changed, so delete it before apply to let the new image run its migrations. The new API and worker pods wait for them, and the old pods keep serving until then. See Kubernetes: Upgrading.

Migrations

The API and the worker apply pending Alembic migrations when they start (AUTO_MIGRATE, on by default), whichever starts first. On Kubernetes the postgres-init Job applies them (see above). To apply them by hand:

InstallCommand
Installer (--native) or pipdocsgpt migrate
docsgpt up or standalone Compose (from the stack directory)docker compose exec backend python -m docsgpt migrate
Checkout Compose (from the repository root)docker compose --env-file .env -f deployment/docker-compose-hub.yaml exec backend python -m docsgpt migrate

The image has no docsgpt console script, so containers run python -m docsgpt. Images up to 0.21.0 lack that entry point too: use python -m docsgpt.cli migrate there. Migrating is idempotent.

Rollback

If the newer release ran no migrations, set the image tag (or package version) back and start it again. If it did, the older release may not work against the upgraded schema, and the realistic way back is the backup you took before upgrading. Go back to the older version first, then restore it with that version, so the older images start on the restored database. For a docsgpt up Docker install, docsgpt upgrade --version <old> moves back; it may report that DocsGPT does not answer, because the older release can’t start on the newer schema. Then run docsgpt restore <archive>, or on 0.21.0, which has no restore, restore by hand.

Alembic can also undo migrations, but many downgrades drop the data the newer release added. Run the downgrade from the newer image, the one that knows the newer revisions, with the API and the worker stopped: with AUTO_MIGRATE on, a running or restarted container migrates straight back up. From the stack directory (from a checkout, add --env-file .env -f deployment/docker-compose-hub.yaml after docker compose):

docker compose stop backend worker docker compose run --rm backend alembic -c docsgpt/alembic.ini downgrade <revision>

<revision> is the newest file in docsgpt/alembic/versions at the older release’s tag. Then set the older tag and start the stack. For example, 0.21.0 ends at 0031_token_usage_cache_tokens; downgrading to it from the next release drops personal access tokens, usage quotas, request traces, team sharing settings, tool preferences and resource sponsors, so restore the backup instead.

Upgrade notes

Newest first. Each note says what an existing deployment has to do.

Next release (unreleased)

These apply to the release after 0.21.0. The changelog lists the other changes that need action: SCIM_TOKEN with SCIM_ENABLED, six retrieved chunks by default, and the removed providers and settings.

Connectors: set ENCRYPTION_SECRET_KEY first

Service credentials (OAuth tokens for Google Drive, SharePoint, Confluence and MCP servers, and API keys for tools) now live on connections and are encrypted with a key derived from ENCRYPTION_SECRET_KEY. Migration 0040_connections runs on startup, encrypts the stored tokens and removes their plaintext copies.

Multi-user installs (any AUTH_TYPE): set ENCRYPTION_SECRET_KEY to your own value before you upgrade, in the environment of the API, the worker and anything that runs migrations. The migration encrypts with the key it sees; changing it afterwards makes every connection ask its owner to reconnect. While the key is still the public default, DocsGPT refuses to store new credentials and Admin > Connectors shows a warning.

If you already used a key and want to change it, see rotating the key. Single-user local installs keep working with the default and log a warning at startup.

Other changes:

  • Register CONNECTOR_REDIRECT_BASE_URI exactly as set (no ?provider= query) as the redirect URI of each OAuth app. Admin > Connectors shows it.
  • VITE_GOOGLE_CLIENT_ID is optional now; without it, members pick Drive files in DocsGPT’s own picker.
  • Synced sources run as their connection, without a browser session. Keep Celery beat running for scheduled syncs.
  • To undo only this change, for example on a develop build that already had the migrations before it, downgrade to 0039_resource_sponsors, the revision just before 0040_connections, the way Rollback describes. Alembic undoes 0043 to 0041 first, then decrypts the tokens and gives each tool its key again. This does not take you back to 0.21.0, which ends at 0031_token_usage_cache_tokens: going that far drops more data, so restore your pre-upgrade backup instead.

Checkout Compose files: bound to this machine

The checkout Compose files (deployment/docker-compose.yaml, docker-compose-hub.yaml and docker-compose-dev.yaml, which setup.sh and setup.ps1 run) now publish Postgres and Redis on 127.0.0.1 only, and the API (7091) and UI (5173) on DOCSGPT_BIND, 127.0.0.1 by default. A local install needs no change.

If other machines open your DocsGPT (a VPS, a LAN server), add DOCSGPT_BIND=0.0.0.0 to .env in the repository root and start the stack with that file, or it answers only on the server itself:

docker compose --env-file .env -f deployment/docker-compose-hub.yaml up -d

Before you do, set AUTH_TYPE: without it anyone who reaches the port uses DocsGPT as one shared account, with your model key and connected services. See the security checklist.

Anything outside the stack that connected to Postgres on port 5432 or Redis on port 6379 from another machine can no longer reach them; a backend run on the same host still can.

JWT key file: moved to the data home

Without JWT_SECRET_KEY, the API generated .jwt_secret_key in the directory it was started from. It now keeps it in the data home (DOCSGPT_HOME, the checkout, or ~/.docsgpt/server). On its first start the API copies a key it finds in its working directory there, so existing tokens stay valid; start it from the same directory as before once. The worker resolves the key the same way, so start whichever process holds the old file first: one started from another directory would generate a new key before the old one is copied. Docker and multi-process installs should set JWT_SECRET_KEY instead (see Secrets to set before going live).

Default model: LLM_PROVIDER alone picks it

The default model, the one that answers when a request names none, is now the first model of the provider in LLM_PROVIDER whenever that provider registered a model. Before, this step also required the generic API_KEY, so an install that set only a provider-specific key fell through to the hosted DocsGPT model:

  • Only a provider key set, for example LLM_PROVIDER=anthropic with ANTHROPIC_API_KEY: chats used to go to the public DocsGPT API. They now go to that provider, billed to your key.
  • LLM_PROVIDER=openai with the key in API_KEY: the key used to register no OpenAI model, so chats silently used the hosted DocsGPT API. They now go to OpenAI.

To keep using the hosted model, set LLM_NAME=docsgpt-local. To pick a specific model of your provider, set LLM_NAME to its id. The API and the worker log an ERROR at startup when LLM_PROVIDER names a provider but the default is still the hosted model. See Default model selection.

Chatwoot bridge: port 5000 and its own .env

The Chatwoot bridge (extensions/chatwoot/app.py) run with python app.py now listens on port 5000, or on PORT, instead of 80. Change the webhook URL in Chatwoot, or set PORT=80. It also reads .env from its own folder, extensions/chatwoot/, instead of the directory it was started from, so move the file there if you kept it elsewhere. See the Chatwoot guide.

Kubernetes manifests reworked

The manifests in deployment/k8s did not work as shipped: the migration Job ran a script missing from the image, uploads never reached the worker, and both Services were public LoadBalancers. They are rebuilt; see the Kubernetes guide for the full setup. To move an existing cluster onto them:

  1. Back up Postgres. The postgres Deployment switches to the pgvector/pgvector image, which reuses the existing volume:

    kubectl exec deploy/postgres -- pg_dump -U docsgpt docsgpt > docsgpt-backup.sql
  2. Carry your keys into the new docsgpt-secrets.yaml. Its values are plain text now (stringData), and every REPLACE_ME must be replaced or the pods refuse to start. Read the old values with:

    kubectl get secret docsgpt-secrets -o jsonpath='{.data.JWT_SECRET_KEY}' | base64 -d
    • Keep your JWT_SECRET_KEY. Replace INTERNAL_KEY if it is still the shipped internal.
    • The old secret had no ENCRYPTION_SECRET_KEY, so any stored credentials are sealed with the public default. Set a new key from openssl rand -hex 32, and add ENCRYPTION_SECRET_KEY_PREVIOUS: default-docsgpt-encryption-key so they stay readable.
    • Choose a new database password: the old manifests used docsgpt. Put it in POSTGRES_PASSWORD and in POSTGRES_URI, and set it on the existing database before you apply, because the bundled Postgres keeps the password it was created with: kubectl exec deploy/postgres -- psql -U docsgpt -c "ALTER USER docsgpt PASSWORD '<new>'". Keeping docsgpt also works, but anyone who reaches the database inside the cluster can then log in.
    • CACHE_REDIS_URL is new; QDRANT_URL and QDRANT_PORT are gone.
    • Fill in the S3 bucket and credentials: uploads now go to a bucket (STORAGE_TYPE: s3), and vectors to pgvector (VECTOR_STORE: pgvector).
  3. Apply the new manifests and wait for the migration. Delete the old Job first, because a finished Job can’t be changed:

    kubectl delete job postgres-init --ignore-not-found kubectl apply -k deployment/k8s/ kubectl wait --for=condition=complete job/postgres-init --timeout=300s

    Then rebuild the database’s indexes. The old postgres:16-alpine image sorts text with musl, and pgvector/pgvector is built on Debian and sorts with glibc’s en_US.utf8, so the same data files now compare strings in a different order. Postgres records no collation version for the old volume, so it does not warn you: indexes on text columns built under the old order can miss rows in lookups or let a duplicate past a unique constraint until they are rebuilt.

    kubectl exec deploy/postgres -- psql -U docsgpt -d docsgpt -c 'REINDEX DATABASE docsgpt'

    Instead of reindexing, you can put the database on a new volume, where the restore builds every index under the new order. Before the commands above, and only once you have checked that docsgpt-backup.sql from step 1 is complete, delete the old volume, start Postgres alone and load the dump; the migration Job then upgrades it:

    kubectl delete deployment/postgres pvc/postgres-pvc kubectl apply -f deployment/k8s/docsgpt-secrets.yaml -f deployment/k8s/deployments/postgres-deploy.yaml \ -f deployment/k8s/services/postgres-service.yaml kubectl rollout status deployment/postgres kubectl exec -i deploy/postgres -- psql -U docsgpt -d docsgpt < docsgpt-backup.sql
  4. Remove what the stack no longer uses. The API serves the web UI, and Qdrant was never used by the old stack:

    kubectl delete deployment/docsgpt-frontend service/docsgpt-frontend-service --ignore-not-found kubectl delete deployment/qdrant service/qdrant pvc/qdrant-pvc --ignore-not-found
  5. Re-encrypt stored secrets if you set ENCRYPTION_SECRET_KEY_PREVIOUS in step 2. docsgpt connectors reencrypt rewrites every connection, tool secret and custom-model key with the new key. Once it reports nothing unreadable, remove ENCRYPTION_SECRET_KEY_PREVIOUS from the secret, apply again, and restart the pods:

    kubectl exec deploy/docsgpt-worker -- python -m docsgpt connectors reencrypt kubectl apply -k deployment/k8s/ kubectl rollout restart deployment/docsgpt-api deployment/docsgpt-worker
  6. Upload your documents again. The old stack kept uploads and FAISS indexes on the API pod’s disk, which the new stack does not read.

docsgpt-api-service is now ClusterIP, so the old external address stops answering. Use kubectl port-forward service/docsgpt-api-service 7091:80, or set AUTH_TYPE and API_URL (the public https:// address) and publish DocsGPT through an Ingress as described in Publish DocsGPT.

Schedules: review the tools they pre-approve

Schedules created before this release still pre-approve every tool the agent had, including actions that need approval. Open each schedule in the Schedules tab: those tools show ticked under Tools that need approval. Untick any a scheduled run should not use without asking, then save. Saving without unticking keeps them approved.

Swagger UI moved to /api/docs

Bookmarks or scripts that opened the Swagger UI at / should use /api/docs; /swagger.json is unchanged.

Azure and local Compose files removed

deployment/docker-compose-azure.yaml and deployment/docker-compose-local.yaml are gone. Their replacements:

  • -azure (the whole stack): docker-compose-hub.yaml for pre-built images, or docker-compose.yaml to build from the checkout.
  • -local (Postgres, Redis and a frontend, for a backend run on the host): docker-compose-dev.yaml for Postgres and Redis, with the frontend run as in Development Environment.

Uploads and indexes under application/ carry over, but the database does not: the removed files had no project name, so their Postgres volume is deployment_postgres_data, while the remaining files use docsgpt-oss_postgres_data. Before you pull, dump the database and stop the old stack (use docker-compose-local.yaml if that is the file you ran):

docker compose --env-file .env -f deployment/docker-compose-azure.yaml exec -T postgres pg_dump -U docsgpt docsgpt > docsgpt.sql docker compose --env-file .env -f deployment/docker-compose-azure.yaml down

Then, after pulling, load it into the new stack’s database before starting the rest (with docker-compose-dev.yaml in place of the hub file for a former -local setup):

docker compose --env-file .env -f deployment/docker-compose-hub.yaml up -d --wait postgres docker compose --env-file .env -f deployment/docker-compose-hub.yaml exec -T postgres psql -U docsgpt docsgpt < docsgpt.sql docker compose --env-file .env -f deployment/docker-compose-hub.yaml up -d

Upgrading to 0.21.0

pip installs: data home moved

An installed package (pip install docsgpt, pipx, uv tool) used to keep its data home, meaning .env, inputs/, indexes/ and models/, in the directory you ran docsgpt api and docsgpt worker from. It is now ~/.docsgpt/server (/opt/docsgpt for root on Linux). Either move those files there, or set DOCSGPT_HOME to the old directory in the environment of both commands. Both commands print the data home they use, and point out a .env in the working directory that they no longer read. Source checkouts and the Docker images are not affected.

Standalone Compose file: one port

The standalone Compose file (docker-compose-standalone.yaml) no longer runs a frontend container. The backend image serves the web UI on port 7091, and the port is published on 127.0.0.1 unless you set DOCSGPT_BIND. After downloading the new file, start it with --remove-orphans to remove the old frontend container, then open port 7091 instead of 5173. If you opened DocsGPT from other machines, see Upgrading from an earlier standalone file. The checkout Compose files are unchanged.

Upgrading to 0.20.0

Embedding models

DocsGPT now runs embeddings through FastEmbed  (ONNX Runtime) instead of SentenceTransformer. The models are the same and the vectors are identical, so your existing index needs no action — all-mpnet-base-v2 keeps working exactly as before.

Your worker command does need one change. Query embedding now runs on the Celery worker (EMBEDDINGS_DELEGATE_TO_WORKER, on by default), which keeps the API from loading a model of its own. If you start your worker with an explicit -Q, add the embeddings queue:

- celery -A docsgpt.app.celery worker -B -l INFO -Q docsgpt,parsing + celery -A docsgpt.app.celery worker -B -l INFO -Q docsgpt,parsing,embeddings

The bundled Compose and Kubernetes manifests already do this — pull them along with the code. Without it, every search blocks for EMBEDDINGS_DELEGATE_TIMEOUT (60s) and then answers with no retrieved context rather than raising, so the symptom is bad answers, not an error. To keep the model out of the worker too, set EMBEDDINGS_BASE_URL; to run the API on its own, set EMBEDDINGS_DELEGATE_TO_WORKER=false.

New installs default to ibm-granite/granite-embedding-311m-multilingual-r2: multilingual, a 32k-token context, and the same 768 dimensions.

Switching an existing deployment to granite

Changing EMBEDDINGS_NAME on an index that already has vectors breaks retrieval silently. Both models are 768-dimensional, so nothing raises an error — queries are simply compared against vectors that mean something else, and answers quietly get worse. Always re-embed.

Set the model, then rebuild the vectors:

# 1. In your .env EMBEDDINGS_NAME=ibm-granite/granite-embedding-311m-multilingual-r2 # 2. Rebuild the vectors from the chunk text already in your index docker compose exec backend python -m docsgpt.scripts.reembed --dry-run docker compose exec backend python -m docsgpt.scripts.reembed

Run these from the directory of your Compose stack (~/.docsgpt/server for docsgpt up); from a checkout, add --env-file .env -f deployment/docker-compose-hub.yaml after docker compose. A pip install runs docsgpt reembed --dry-run, then docsgpt reembed.

Re-embedding reads the chunk text already stored in your index. It does not re-download, re-parse or re-chunk your documents, so no source files are needed and the run is proportional to index size, not corpus size. Both pgvector and faiss are supported.

Useful flags:

FlagEffect
--dry-runReport how many chunks would change, write nothing
--sources a,bOnly these source ids — also how you retry a failed source
--batch-size NChunks per embed call (default 64)

The script processes sources independently: one failing source is reported and skipped rather than aborting the run, and the exit code is non-zero if any failed. For pgvector it reads a page of chunks at a time and updates rows in place, so memory stays flat on a large index and an interrupted run simply re-does its last batch. For faiss it builds the replacement index in memory, writes each file to a temporary path, and moves it into place — so an interrupt during either the rebuild or the write leaves the existing index intact rather than truncated.

Stop ingest before you run this. It reads each source’s chunks and writes the vectors back; anything ingested while it runs can be overwritten by the rebuild (faiss) or missed by it (pgvector).

Running GraphRAG? The script also rewrites graph_nodes.name_embedding, which seeds every graph traversal. Those vectors are written once at extraction time and share the chunk vectors’ width, so leaving them in the old model’s space degrades graph retrieval just as silently as the chunk vectors would — and needs no LLM re-extraction to fix.

Custom local models

A local model now runs through ONNX Runtime, so its repository must ship an ONNX export (onnx/model.onnx) or be one of FastEmbed’s built-in models. Repositories with PyTorch weights only no longer load; hkunlp/instructor-large, previously supported by name, is one of them. Serve such a model over EMBEDDINGS_BASE_URL instead, or switch to a model with an export.

How to run the model — pooling, and whether outputs are L2-normalised — is read from the repository’s own 1_Pooling/config.json and modules.json. Two cases need attention:

  • Models with a Dense projection layer (sentence-transformers/LaBSE, distiluse-base-multilingual-cased-v1) are now refused at startup. FastEmbed cannot apply the projection, so it would have produced vectors of the wrong width in a different space. If you were running one, its stored vectors were already wrong; move it to EMBEDDINGS_BASE_URL or pick another model.
  • Repositories that declare nothing fall back to mean pooling with normalisation and log a warning. Pin the real values with EMBEDDINGS_POOLING (cls or mean) and EMBEDDINGS_NORMALIZE.

Staying on all-mpnet-base-v2 is a supported choice — it remains in the model registry and in setup.sh. You only need this section if you want to move to granite.

Backend package renamed to docsgpt

The backend’s Python package is docsgpt (it was application), the name it will carry on PyPI. For one release the old name keeps working through an alias, so nothing breaks on upgrade, but update these before the alias goes:

  • Entry points: celery -A docsgpt.app.celery worker, uvicorn docsgpt.asgi:asgi_app, python -m docsgpt.scripts.<name>. The application.… spellings still run and print a FutureWarning. The compose files, Kubernetes manifests and setup scripts in the repository are already updated; only custom copies need editing.
  • Local image builds: the build context is the repository root, so use docker build -f docsgpt/Dockerfile . (or the compose files, which do this).
  • Celery task names changed with the package (docsgpt.api.user.tasks.ingest and so on). A worker on this release also accepts the old names, so tasks queued before the upgrade still run, and beat rewrites the periodic schedule in Redis on start-up. The daily, weekly and monthly source-sync timers restart from the upgrade, so the first sync after it can land later than it would have (a monthly sync by up to a month). Nothing to do.
  • Data directories do not move: the compose files keep your indexes, inputs and vectors under application/ in the checkout, where they already are.
  • The backend is also a package now (pip install docsgpt, see Install with pip). Runtime data lives in a data home: DOCSGPT_HOME, else the checkout, else ~/.docsgpt/server (/opt/docsgpt for root on Linux; see pip installs: data home moved). One consequence for a source checkout: the embedded Milvus (MILVUS_URI) default path now resolves under the checkout instead of the start directory. If you use it at its default path and start DocsGPT from another directory, the old data is at <start directory>/milvus_local.db; point the setting at it, or move it into the checkout. Faiss indexes and uploads were already stored under the checkout and are unaffected.

Upgrading to 0.17.0

User data moved from MongoDB to PostgreSQL. Migrate it before you pull the new images; see Migrating from MongoDB.

Maintenance scripts

The docsgpt command covers the routine tasks; the CLI reference lists them all. A few one-off scripts live only in the repository under scripts/ and are not in the Docker image or the pip package.

ScriptWhen to run it
docsgpt reembedAfter changing EMBEDDINGS_NAME; see Switching an existing deployment to granite. Packaged.
docsgpt grant-adminTo make the first admin under AUTH_TYPE=oidc; see Access Control. Packaged.
scripts/db/migrate_model_ids.py [--map OLD=NEW ...] [--apply]After a provider renames or retires a model id: rewrites the ids that agents and schedules store, which would otherwise fail on their next call. Dry-run unless --apply.
scripts/db/backfill_token_usage_model_id.py [--apply]Once, to attribute usage rows recorded before token usage stored its model, so analytics can group them by model. Dry-run unless --apply.
scripts/db/backfill_tool_attempts_attribution.py [--apply]Once, to attribute tool calls recorded before migration 0018 to their user and agent. Dry-run unless --apply.
scripts/db/backfill.pyMoving from MongoDB (0.16.x); see PostgreSQL for User Data.
scripts/db/init_postgres.pySame as docsgpt migrate, from a checkout.
scripts/migrate_to_v1_vectorstore.py, scripts/migrate_conversation_id_dbref_to_objectid.pyNever on current versions: they change MongoDB data from before the move to Postgres.

Run a script from the repository root of a checkout of the release you run. Directly on the host, POSTGRES_URI in the checkout’s .env (or the environment) must point at the deployment’s database. Against a Docker stack, copy the scripts folder into the backend container and run it there with the app on the path. For the checkout Compose files:

docker compose --env-file .env -f deployment/docker-compose-hub.yaml cp scripts/. backend:/tmp/scripts/ docker compose --env-file .env -f deployment/docker-compose-hub.yaml exec -e PYTHONPATH=/app backend \ python /tmp/scripts/db/migrate_model_ids.py

For a docsgpt up stack, still from the checkout (the stack directory has no scripts/):

docker compose -f ~/.docsgpt/server/docker-compose.yaml cp scripts/. backend:/tmp/scripts/ docker compose -f ~/.docsgpt/server/docker-compose.yaml exec -e PYTHONPATH=/app backend \ python /tmp/scripts/db/migrate_model_ids.py