Upgrading DocsGPT
Upgrading from 0.16.x? User data moved from MongoDB to Postgres in 0.17.0. Follow the Postgres Migration guide before running docker compose pull or git pull — existing deployments will not start cleanly without it.
To upgrade:
- Back up.
docsgpt backupfor an installer ordocsgpt upDocker install whosedocsgpthas that command (releases after 0.21.0; on 0.21.0 back up by hand). Otherwise dump Postgres (pg_dump) and copy your uploads and indexes. - Read the upgrade notes for every release after the one you run, and the changelog.
- Follow the steps for your install, below.
Check your version
| Install | Command |
|---|---|
Installer, docsgpt up, pip | docsgpt --version for the package; docsgpt status for the version the stack runs |
Docker Compose (from the stack directory; from a checkout add --env-file .env -f deployment/docker-compose-hub.yaml) | docker compose exec backend python -c "from docsgpt.version import get_version; print(get_version())" |
| Kubernetes | kubectl exec deploy/docsgpt-api -- python -m docsgpt.cli --version |
Release notes: changelog. Tags: GitHub releases .
Installer or docsgpt up
docsgpt backup # releases after 0.21.0; see below for 0.21.0
docsgpt upgrade # the latest release; --version <version> for another
docsgpt statusFor a uv tool install, which is what the installer sets up, docsgpt upgrade installs the new package and runs its docsgpt up, which moves the stack to that release’s images and keeps .env and the data. Running the install command again does the same; DOCSGPT_VERSION=<version> picks a release. With pipx or pip, docsgpt upgrade prints the command to run instead (pipx install --force docsgpt or pip install -U docsgpt); run it, then docsgpt up.
A docsgpt up --native install has nothing for docsgpt backup to archive: dump its Postgres with pg_dump and copy indexes/, inputs/ and vectors/ from the data home first.
Back up and restore by hand (0.21.0)
docsgpt backup and docsgpt restore are new after 0.21.0 (docsgpt --help shows whether yours has them). On 0.21.0, take the same backup by hand from the stack directory. It holds a dump of the database and a tar of each data volume (docsgpt_indexes, docsgpt_inputs, docsgpt_vectors). Keep a copy of .env too: it holds the install’s secrets.
cd ~/.docsgpt/server # /opt/docsgpt when installed as root on Linux
IMAGE=arc53/docsgpt:0.21.0 # the image the stack runs
mkdir -p backups/before-upgrade
docker compose stop backend worker # so the dump and the files match
docker compose exec -T postgres pg_dump --clean --if-exists -U docsgpt -d docsgpt > backups/before-upgrade/docsgpt.sql
for v in indexes inputs vectors; do
docker run --rm -v docsgpt_$v:/data:ro "$IMAGE" tar cf - -C /data . > backups/before-upgrade/$v.tar
done
docker compose start backend workerTo restore it, with the stack on the version the backup came from. Load the dump into an empty database: loaded over a database a newer release migrated, the newer tables’ foreign keys stop it.
First check that every tar reads, since the restore empties each volume before unpacking into it. Go on only if all three print ok:
cd ~/.docsgpt/server # /opt/docsgpt when installed as root on Linux
for v in indexes inputs vectors; do
tar -tf backups/before-upgrade/$v.tar > /dev/null && echo "$v.tar ok"
doneThen restore:
IMAGE=arc53/docsgpt:0.21.0
docker compose down
for v in indexes inputs vectors; do
docker run --rm -i -u 0 -v docsgpt_$v:/data "$IMAGE" \
sh -c 'find /data -mindepth 1 -delete && tar xf - -C /data' < backups/before-upgrade/$v.tar
done
docker compose up -d --wait postgres
docker compose exec -T postgres psql -U docsgpt -d postgres \
-c 'ALTER DATABASE docsgpt RENAME TO docsgpt_before_restore' -c 'CREATE DATABASE docsgpt OWNER docsgpt'
docker compose exec -T postgres psql --set ON_ERROR_STOP=on -U docsgpt -d docsgpt < backups/before-upgrade/docsgpt.sql
docker compose up -dOnce DocsGPT works, drop the set-aside copy: docker compose exec postgres psql -U docsgpt -d postgres -c 'DROP DATABASE docsgpt_before_restore'. If the load fails, drop the new docsgpt database and rename docsgpt_before_restore back before starting DocsGPT. If a restore stopped after the rename and DocsGPT has since started, the API will have created an empty docsgpt; drop that empty one (DROP DATABASE docsgpt WITH (FORCE)) before you rename docsgpt_before_restore back. docsgpt restore does all of this itself.
pip install
pip install -U docsgpt # or: pipx install --force docsgpt
docsgpt migrateThen restart docsgpt api and docsgpt worker (docsgpt restart for a docsgpt up --native install). See Install with pip.
Docker Compose — standalone file
Download the new release’s docker-compose-standalone.yaml (attached to every release ) over the old one. If .env pins DOCSGPT_IMAGE_TAG, set it to the new release; without it the file runs latest. Then, in that folder:
docker compose -f docker-compose-standalone.yaml pull
docker compose -f docker-compose-standalone.yaml up -d --remove-orphansDocker Compose — hub images
From the repository root:
git pull # or: git checkout <version>
docker compose --env-file .env -f deployment/docker-compose-hub.yaml pull
docker compose --env-file .env -f deployment/docker-compose-hub.yaml up -d--env-file .env makes Compose read settings such as DOCSGPT_BIND and DOCSGPT_IMAGE_TAG from the repository root; without it they are ignored. This file runs the develop images, built from the main branch, unless .env sets DOCSGPT_IMAGE_TAG. To move to a release, set DOCSGPT_IMAGE_TAG=<version> in .env; don’t edit image: in the file, which would leave the frontend image on another tag. With the Ollama overlay, add its -f deployment/optional/... file to both commands.
Docker Compose — from source
cd DocsGPT
git pull
docker compose --env-file .env -f deployment/docker-compose.yaml build
docker compose --env-file .env -f deployment/docker-compose.yaml up -dSwap git pull for git checkout <tag> if you want to pin a specific release. Build settings such as EXTRAS=docling are read from .env through --env-file .env.
Kubernetes
The Kubernetes pods don’t migrate the database when they start; the postgres-init Job does. Change every arc53/docsgpt:<version> image in deployment/k8s/deployments/docsgpt-deploy.yaml and deployment/k8s/jobs/postgres-init-job.yaml to the new release, and arc53/docsgpt-sandbox:<version> in deployment/k8s/deployments/sandbox-deploy.yaml if you run the opt-in sandbox. Then:
kubectl delete job postgres-init --ignore-not-found
kubectl apply -k deployment/k8s/
kubectl wait --for=condition=complete job/postgres-init --timeout=300s
kubectl rollout status deployment/docsgpt-api
kubectl rollout status deployment/docsgpt-workerA finished Job can’t be changed, so delete it before apply to let the new image run its migrations. The new API and worker pods wait for them, and the old pods keep serving until then. See Kubernetes: Upgrading.
Migrations
The API and the worker apply pending Alembic migrations when they start (AUTO_MIGRATE, on by default), whichever starts first. On Kubernetes the postgres-init Job applies them (see above). To apply them by hand:
| Install | Command |
|---|---|
Installer (--native) or pip | docsgpt migrate |
docsgpt up or standalone Compose (from the stack directory) | docker compose exec backend python -m docsgpt migrate |
| Checkout Compose (from the repository root) | docker compose --env-file .env -f deployment/docker-compose-hub.yaml exec backend python -m docsgpt migrate |
The image has no docsgpt console script, so containers run python -m docsgpt. Images up to 0.21.0 lack that entry point too: use python -m docsgpt.cli migrate there. Migrating is idempotent.
Rollback
If the newer release ran no migrations, set the image tag (or package version) back and start it again. If it did, the older release may not work against the upgraded schema, and the realistic way back is the backup you took before upgrading. Go back to the older version first, then restore it with that version, so the older images start on the restored database. For a docsgpt up Docker install, docsgpt upgrade --version <old> moves back; it may report that DocsGPT does not answer, because the older release can’t start on the newer schema. Then run docsgpt restore <archive>, or on 0.21.0, which has no restore, restore by hand.
Alembic can also undo migrations, but many downgrades drop the data the newer release added. Run the downgrade from the newer image, the one that knows the newer revisions, with the API and the worker stopped: with AUTO_MIGRATE on, a running or restarted container migrates straight back up. From the stack directory (from a checkout, add --env-file .env -f deployment/docker-compose-hub.yaml after docker compose):
docker compose stop backend worker
docker compose run --rm backend alembic -c docsgpt/alembic.ini downgrade <revision><revision> is the newest file in docsgpt/alembic/versions at the older release’s tag. Then set the older tag and start the stack. For example, 0.21.0 ends at 0031_token_usage_cache_tokens; downgrading to it from the next release drops personal access tokens, usage quotas, request traces, team sharing settings, tool preferences and resource sponsors, so restore the backup instead.
Upgrade notes
Newest first. Each note says what an existing deployment has to do.
Next release (unreleased)
These apply to the release after 0.21.0. The changelog lists the other changes that need action: SCIM_TOKEN with SCIM_ENABLED, six retrieved chunks by default, and the removed providers and settings.
Connectors: set ENCRYPTION_SECRET_KEY first
Service credentials (OAuth tokens for Google Drive, SharePoint, Confluence and MCP servers, and API keys for tools) now live on connections and are encrypted with a key derived from ENCRYPTION_SECRET_KEY. Migration 0040_connections runs on startup, encrypts the stored tokens and removes their plaintext copies.
Multi-user installs (any AUTH_TYPE): set ENCRYPTION_SECRET_KEY to your own value before you upgrade, in the environment of the API, the worker and anything that runs migrations. The migration encrypts with the key it sees; changing it afterwards makes every connection ask its owner to reconnect. While the key is still the public default, DocsGPT refuses to store new credentials and Admin > Connectors shows a warning.
If you already used a key and want to change it, see rotating the key. Single-user local installs keep working with the default and log a warning at startup.
Other changes:
- Register
CONNECTOR_REDIRECT_BASE_URIexactly as set (no?provider=query) as the redirect URI of each OAuth app. Admin > Connectors shows it. VITE_GOOGLE_CLIENT_IDis optional now; without it, members pick Drive files in DocsGPT’s own picker.- Synced sources run as their connection, without a browser session. Keep Celery beat running for scheduled syncs.
- To undo only this change, for example on a
developbuild that already had the migrations before it, downgrade to0039_resource_sponsors, the revision just before0040_connections, the way Rollback describes. Alembic undoes0043to0041first, then decrypts the tokens and gives each tool its key again. This does not take you back to 0.21.0, which ends at0031_token_usage_cache_tokens: going that far drops more data, so restore your pre-upgrade backup instead.
Checkout Compose files: bound to this machine
The checkout Compose files (deployment/docker-compose.yaml, docker-compose-hub.yaml and docker-compose-dev.yaml, which setup.sh and setup.ps1 run) now publish Postgres and Redis on 127.0.0.1 only, and the API (7091) and UI (5173) on DOCSGPT_BIND, 127.0.0.1 by default. A local install needs no change.
If other machines open your DocsGPT (a VPS, a LAN server), add DOCSGPT_BIND=0.0.0.0 to .env in the repository root and start the stack with that file, or it answers only on the server itself:
docker compose --env-file .env -f deployment/docker-compose-hub.yaml up -dBefore you do, set AUTH_TYPE: without it anyone who reaches the port uses DocsGPT as one shared account, with your model key and connected services. See the security checklist.
Anything outside the stack that connected to Postgres on port 5432 or Redis on port 6379 from another machine can no longer reach them; a backend run on the same host still can.
JWT key file: moved to the data home
Without JWT_SECRET_KEY, the API generated .jwt_secret_key in the directory it was started from. It now keeps it in the data home (DOCSGPT_HOME, the checkout, or ~/.docsgpt/server). On its first start the API copies a key it finds in its working directory there, so existing tokens stay valid; start it from the same directory as before once. The worker resolves the key the same way, so start whichever process holds the old file first: one started from another directory would generate a new key before the old one is copied. Docker and multi-process installs should set JWT_SECRET_KEY instead (see Secrets to set before going live).
Default model: LLM_PROVIDER alone picks it
The default model, the one that answers when a request names none, is now the first model of the provider in LLM_PROVIDER whenever that provider registered a model. Before, this step also required the generic API_KEY, so an install that set only a provider-specific key fell through to the hosted DocsGPT model:
- Only a provider key set, for example
LLM_PROVIDER=anthropicwithANTHROPIC_API_KEY: chats used to go to the public DocsGPT API. They now go to that provider, billed to your key. LLM_PROVIDER=openaiwith the key inAPI_KEY: the key used to register no OpenAI model, so chats silently used the hosted DocsGPT API. They now go to OpenAI.
To keep using the hosted model, set LLM_NAME=docsgpt-local. To pick a specific model of your provider, set LLM_NAME to its id. The API and the worker log an ERROR at startup when LLM_PROVIDER names a provider but the default is still the hosted model. See Default model selection.
Chatwoot bridge: port 5000 and its own .env
The Chatwoot bridge (extensions/chatwoot/app.py) run with python app.py now listens on port 5000, or on PORT, instead of 80. Change the webhook URL in Chatwoot, or set PORT=80. It also reads .env from its own folder, extensions/chatwoot/, instead of the directory it was started from, so move the file there if you kept it elsewhere. See the Chatwoot guide.
Kubernetes manifests reworked
The manifests in deployment/k8s did not work as shipped: the migration Job ran a script missing from the image, uploads never reached the worker, and both Services were public LoadBalancers. They are rebuilt; see the Kubernetes guide for the full setup. To move an existing cluster onto them:
-
Back up Postgres. The
postgresDeployment switches to thepgvector/pgvectorimage, which reuses the existing volume:kubectl exec deploy/postgres -- pg_dump -U docsgpt docsgpt > docsgpt-backup.sql -
Carry your keys into the new
docsgpt-secrets.yaml. Its values are plain text now (stringData), and everyREPLACE_MEmust be replaced or the pods refuse to start. Read the old values with:kubectl get secret docsgpt-secrets -o jsonpath='{.data.JWT_SECRET_KEY}' | base64 -d- Keep your
JWT_SECRET_KEY. ReplaceINTERNAL_KEYif it is still the shippedinternal. - The old secret had no
ENCRYPTION_SECRET_KEY, so any stored credentials are sealed with the public default. Set a new key fromopenssl rand -hex 32, and addENCRYPTION_SECRET_KEY_PREVIOUS: default-docsgpt-encryption-keyso they stay readable. - Choose a new database password: the old manifests used
docsgpt. Put it inPOSTGRES_PASSWORDand inPOSTGRES_URI, and set it on the existing database before you apply, because the bundled Postgres keeps the password it was created with:kubectl exec deploy/postgres -- psql -U docsgpt -c "ALTER USER docsgpt PASSWORD '<new>'". Keepingdocsgptalso works, but anyone who reaches the database inside the cluster can then log in. CACHE_REDIS_URLis new;QDRANT_URLandQDRANT_PORTare gone.- Fill in the S3 bucket and credentials: uploads now go to a bucket (
STORAGE_TYPE: s3), and vectors to pgvector (VECTOR_STORE: pgvector).
- Keep your
-
Apply the new manifests and wait for the migration. Delete the old Job first, because a finished Job can’t be changed:
kubectl delete job postgres-init --ignore-not-found kubectl apply -k deployment/k8s/ kubectl wait --for=condition=complete job/postgres-init --timeout=300sThen rebuild the database’s indexes. The old
postgres:16-alpineimage sorts text with musl, andpgvector/pgvectoris built on Debian and sorts with glibc’sen_US.utf8, so the same data files now compare strings in a different order. Postgres records no collation version for the old volume, so it does not warn you: indexes on text columns built under the old order can miss rows in lookups or let a duplicate past a unique constraint until they are rebuilt.kubectl exec deploy/postgres -- psql -U docsgpt -d docsgpt -c 'REINDEX DATABASE docsgpt'Instead of reindexing, you can put the database on a new volume, where the restore builds every index under the new order. Before the commands above, and only once you have checked that
docsgpt-backup.sqlfrom step 1 is complete, delete the old volume, start Postgres alone and load the dump; the migration Job then upgrades it:kubectl delete deployment/postgres pvc/postgres-pvc kubectl apply -f deployment/k8s/docsgpt-secrets.yaml -f deployment/k8s/deployments/postgres-deploy.yaml \ -f deployment/k8s/services/postgres-service.yaml kubectl rollout status deployment/postgres kubectl exec -i deploy/postgres -- psql -U docsgpt -d docsgpt < docsgpt-backup.sql -
Remove what the stack no longer uses. The API serves the web UI, and Qdrant was never used by the old stack:
kubectl delete deployment/docsgpt-frontend service/docsgpt-frontend-service --ignore-not-found kubectl delete deployment/qdrant service/qdrant pvc/qdrant-pvc --ignore-not-found -
Re-encrypt stored secrets if you set
ENCRYPTION_SECRET_KEY_PREVIOUSin step 2.docsgpt connectors reencryptrewrites every connection, tool secret and custom-model key with the new key. Once it reports nothing unreadable, removeENCRYPTION_SECRET_KEY_PREVIOUSfrom the secret, apply again, and restart the pods:kubectl exec deploy/docsgpt-worker -- python -m docsgpt connectors reencrypt kubectl apply -k deployment/k8s/ kubectl rollout restart deployment/docsgpt-api deployment/docsgpt-worker -
Upload your documents again. The old stack kept uploads and FAISS indexes on the API pod’s disk, which the new stack does not read.
docsgpt-api-service is now ClusterIP, so the old external address stops answering. Use kubectl port-forward service/docsgpt-api-service 7091:80, or set AUTH_TYPE and API_URL (the public https:// address) and publish DocsGPT through an Ingress as described in Publish DocsGPT.
Schedules: review the tools they pre-approve
Schedules created before this release still pre-approve every tool the agent had, including actions that need approval. Open each schedule in the Schedules tab: those tools show ticked under Tools that need approval. Untick any a scheduled run should not use without asking, then save. Saving without unticking keeps them approved.
Swagger UI moved to /api/docs
Bookmarks or scripts that opened the Swagger UI at / should use /api/docs; /swagger.json is unchanged.
Azure and local Compose files removed
deployment/docker-compose-azure.yaml and deployment/docker-compose-local.yaml are gone. Their replacements:
-azure(the whole stack):docker-compose-hub.yamlfor pre-built images, ordocker-compose.yamlto build from the checkout.-local(Postgres, Redis and a frontend, for a backend run on the host):docker-compose-dev.yamlfor Postgres and Redis, with the frontend run as in Development Environment.
Uploads and indexes under application/ carry over, but the database does not: the removed files had no project name, so their Postgres volume is deployment_postgres_data, while the remaining files use docsgpt-oss_postgres_data. Before you pull, dump the database and stop the old stack (use docker-compose-local.yaml if that is the file you ran):
docker compose --env-file .env -f deployment/docker-compose-azure.yaml exec -T postgres pg_dump -U docsgpt docsgpt > docsgpt.sql
docker compose --env-file .env -f deployment/docker-compose-azure.yaml downThen, after pulling, load it into the new stack’s database before starting the rest (with docker-compose-dev.yaml in place of the hub file for a former -local setup):
docker compose --env-file .env -f deployment/docker-compose-hub.yaml up -d --wait postgres
docker compose --env-file .env -f deployment/docker-compose-hub.yaml exec -T postgres psql -U docsgpt docsgpt < docsgpt.sql
docker compose --env-file .env -f deployment/docker-compose-hub.yaml up -dUpgrading to 0.21.0
pip installs: data home moved
An installed package (pip install docsgpt, pipx, uv tool) used to keep its data home, meaning .env, inputs/, indexes/ and models/, in the directory you ran docsgpt api and docsgpt worker from. It is now ~/.docsgpt/server (/opt/docsgpt for root on Linux). Either move those files there, or set DOCSGPT_HOME to the old directory in the environment of both commands. Both commands print the data home they use, and point out a .env in the working directory that they no longer read. Source checkouts and the Docker images are not affected.
Standalone Compose file: one port
The standalone Compose file (docker-compose-standalone.yaml) no longer runs a frontend container. The backend image serves the web UI on port 7091, and the port is published on 127.0.0.1 unless you set DOCSGPT_BIND. After downloading the new file, start it with --remove-orphans to remove the old frontend container, then open port 7091 instead of 5173. If you opened DocsGPT from other machines, see Upgrading from an earlier standalone file. The checkout Compose files are unchanged.
Upgrading to 0.20.0
Embedding models
DocsGPT now runs embeddings through FastEmbed (ONNX Runtime) instead of SentenceTransformer. The models are the same and the vectors are identical, so your existing index needs no action — all-mpnet-base-v2 keeps working exactly as before.
Your worker command does need one change. Query embedding now runs on the Celery worker (EMBEDDINGS_DELEGATE_TO_WORKER, on by default), which keeps the API from loading a model of its own. If you start your worker with an explicit -Q, add the embeddings queue:
- celery -A docsgpt.app.celery worker -B -l INFO -Q docsgpt,parsing
+ celery -A docsgpt.app.celery worker -B -l INFO -Q docsgpt,parsing,embeddingsThe bundled Compose and Kubernetes manifests already do this — pull them along with the code. Without it, every search blocks for EMBEDDINGS_DELEGATE_TIMEOUT (60s) and then answers with no retrieved context rather than raising, so the symptom is bad answers, not an error. To keep the model out of the worker too, set EMBEDDINGS_BASE_URL; to run the API on its own, set EMBEDDINGS_DELEGATE_TO_WORKER=false.
New installs default to ibm-granite/granite-embedding-311m-multilingual-r2: multilingual, a 32k-token context, and the same 768 dimensions.
Switching an existing deployment to granite
Changing EMBEDDINGS_NAME on an index that already has vectors breaks retrieval silently. Both models are 768-dimensional, so nothing raises an error — queries are simply compared against vectors that mean something else, and answers quietly get worse. Always re-embed.
Set the model, then rebuild the vectors:
# 1. In your .env
EMBEDDINGS_NAME=ibm-granite/granite-embedding-311m-multilingual-r2
# 2. Rebuild the vectors from the chunk text already in your index
docker compose exec backend python -m docsgpt.scripts.reembed --dry-run
docker compose exec backend python -m docsgpt.scripts.reembedRun these from the directory of your Compose stack (~/.docsgpt/server for docsgpt up); from a checkout, add --env-file .env -f deployment/docker-compose-hub.yaml after docker compose. A pip install runs docsgpt reembed --dry-run, then docsgpt reembed.
Re-embedding reads the chunk text already stored in your index. It does not re-download, re-parse or re-chunk your documents, so no source files are needed and the run is proportional to index size, not corpus size. Both pgvector and faiss are supported.
Useful flags:
| Flag | Effect |
|---|---|
--dry-run | Report how many chunks would change, write nothing |
--sources a,b | Only these source ids — also how you retry a failed source |
--batch-size N | Chunks per embed call (default 64) |
The script processes sources independently: one failing source is reported and skipped rather than aborting the run, and the exit code is non-zero if any failed. For pgvector it reads a page of chunks at a time and updates rows in place, so memory stays flat on a large index and an interrupted run simply re-does its last batch. For faiss it builds the replacement index in memory, writes each file to a temporary path, and moves it into place — so an interrupt during either the rebuild or the write leaves the existing index intact rather than truncated.
Stop ingest before you run this. It reads each source’s chunks and writes the vectors back; anything ingested while it runs can be overwritten by the rebuild (faiss) or missed by it (pgvector).
Running GraphRAG? The script also rewrites graph_nodes.name_embedding, which seeds every graph traversal. Those vectors are written once at extraction time and share the chunk vectors’ width, so leaving them in the old model’s space degrades graph retrieval just as silently as the chunk vectors would — and needs no LLM re-extraction to fix.
Custom local models
A local model now runs through ONNX Runtime, so its repository must ship an ONNX
export (onnx/model.onnx) or be one of FastEmbed’s built-in models. Repositories
with PyTorch weights only no longer load; hkunlp/instructor-large, previously
supported by name, is one of them. Serve such a model over EMBEDDINGS_BASE_URL
instead, or switch to a model with an export.
How to run the model — pooling, and whether outputs are L2-normalised — is read
from the repository’s own 1_Pooling/config.json and modules.json. Two cases
need attention:
- Models with a Dense projection layer (
sentence-transformers/LaBSE,distiluse-base-multilingual-cased-v1) are now refused at startup. FastEmbed cannot apply the projection, so it would have produced vectors of the wrong width in a different space. If you were running one, its stored vectors were already wrong; move it toEMBEDDINGS_BASE_URLor pick another model. - Repositories that declare nothing fall back to mean pooling with
normalisation and log a warning. Pin the real values with
EMBEDDINGS_POOLING(clsormean) andEMBEDDINGS_NORMALIZE.
Staying on all-mpnet-base-v2 is a supported choice — it remains in the model registry and in setup.sh. You only need this section if you want to move to granite.
Backend package renamed to docsgpt
The backend’s Python package is docsgpt (it was application), the name it
will carry on PyPI. For one release the old name keeps working through an
alias, so nothing breaks on upgrade, but update these before the alias goes:
- Entry points:
celery -A docsgpt.app.celery worker,uvicorn docsgpt.asgi:asgi_app,python -m docsgpt.scripts.<name>. Theapplication.…spellings still run and print aFutureWarning. The compose files, Kubernetes manifests and setup scripts in the repository are already updated; only custom copies need editing. - Local image builds: the build context is the repository root, so use
docker build -f docsgpt/Dockerfile .(or the compose files, which do this). - Celery task names changed with the package (
docsgpt.api.user.tasks.ingestand so on). A worker on this release also accepts the old names, so tasks queued before the upgrade still run, and beat rewrites the periodic schedule in Redis on start-up. The daily, weekly and monthly source-sync timers restart from the upgrade, so the first sync after it can land later than it would have (a monthly sync by up to a month). Nothing to do. - Data directories do not move: the compose files keep your indexes, inputs
and vectors under
application/in the checkout, where they already are. - The backend is also a package now (
pip install docsgpt, see Install with pip). Runtime data lives in a data home:DOCSGPT_HOME, else the checkout, else~/.docsgpt/server(/opt/docsgptfor root on Linux; see pip installs: data home moved). One consequence for a source checkout: the embedded Milvus (MILVUS_URI) default path now resolves under the checkout instead of the start directory. If you use it at its default path and start DocsGPT from another directory, the old data is at<start directory>/milvus_local.db; point the setting at it, or move it into the checkout. Faiss indexes and uploads were already stored under the checkout and are unaffected.
Upgrading to 0.17.0
User data moved from MongoDB to PostgreSQL. Migrate it before you pull the new images; see Migrating from MongoDB.
Maintenance scripts
The docsgpt command covers the routine tasks; the CLI reference lists them all. A few one-off scripts live only in the repository under scripts/ and are not in the Docker image or the pip package.
| Script | When to run it |
|---|---|
docsgpt reembed | After changing EMBEDDINGS_NAME; see Switching an existing deployment to granite. Packaged. |
docsgpt grant-admin | To make the first admin under AUTH_TYPE=oidc; see Access Control. Packaged. |
scripts/db/migrate_model_ids.py [--map OLD=NEW ...] [--apply] | After a provider renames or retires a model id: rewrites the ids that agents and schedules store, which would otherwise fail on their next call. Dry-run unless --apply. |
scripts/db/backfill_token_usage_model_id.py [--apply] | Once, to attribute usage rows recorded before token usage stored its model, so analytics can group them by model. Dry-run unless --apply. |
scripts/db/backfill_tool_attempts_attribution.py [--apply] | Once, to attribute tool calls recorded before migration 0018 to their user and agent. Dry-run unless --apply. |
scripts/db/backfill.py | Moving from MongoDB (0.16.x); see PostgreSQL for User Data. |
scripts/db/init_postgres.py | Same as docsgpt migrate, from a checkout. |
scripts/migrate_to_v1_vectorstore.py, scripts/migrate_conversation_id_dbref_to_objectid.py | Never on current versions: they change MongoDB data from before the move to Postgres. |
Run a script from the repository root of a checkout of the release you run. Directly on the host, POSTGRES_URI in the checkout’s .env (or the environment) must point at the deployment’s database. Against a Docker stack, copy the scripts folder into the backend container and run it there with the app on the path. For the checkout Compose files:
docker compose --env-file .env -f deployment/docker-compose-hub.yaml cp scripts/. backend:/tmp/scripts/
docker compose --env-file .env -f deployment/docker-compose-hub.yaml exec -e PYTHONPATH=/app backend \
python /tmp/scripts/db/migrate_model_ids.pyFor a docsgpt up stack, still from the checkout (the stack directory has no scripts/):
docker compose -f ~/.docsgpt/server/docker-compose.yaml cp scripts/. backend:/tmp/scripts/
docker compose -f ~/.docsgpt/server/docker-compose.yaml exec -e PYTHONPATH=/app backend \
python /tmp/scripts/db/migrate_model_ids.py