Skip to Content
Deploy & Operate🧹 Background Jobs & Retention

Background Jobs and Data Retention

Besides the API and the worker, DocsGPT needs a Celery beat scheduler. Beat puts periodic tasks on the queue, and the worker runs them: source syncs, scheduled agent runs, the reconciliation of stuck work, and the cleanups that keep Postgres from growing without bound. Without beat none of this happens, and nothing tells you so.

Running beat

At least one beat scheduler must run. Beat uses RedBeat, which keeps its schedule and a lock in Redis, so only the instance holding the lock schedules anything. Starting beat in more than one worker is therefore safe: the extra instances wait and take over if the lock holder stops.

InstallHow beat runs
docsgpt up and the one-line installerEmbedded in the worker container (celery ... worker -B). Nothing to do.
docsgpt up --native, docsgpt workerEmbedded (-B) by default. docsgpt worker --no-beat leaves it out; then run docsgpt beat as its own process.
Checkout Compose files (docker-compose.yaml, docker-compose-hub.yaml, docker-compose-standalone.yaml)Embedded in the worker service with -B.
Kubernetes manifestsEmbedded in docsgpt-worker with -B, safe on every replica.
Developmentcelery -A docsgpt.app.celery worker -l INFO -B, or docsgpt worker. See Development Environment.
WindowsThe worker can’t embed beat. Run docsgpt beat (or celery -A docsgpt.app.celery beat -l INFO) in a separate process.

If you write your own worker command, keep -B on it or run celery -A docsgpt.app.celery beat -l INFO next to it. Signs that beat isn’t running: scheduled agent runs and source syncs never fire, and a request stuck after a crash is never marked failed.

Periodic tasks

Beat schedules these tasks. Each runs on the worker, so the worker must also be running and consuming the docsgpt queue. The cleanups skip themselves when POSTGRES_URI is not set.

TaskEveryWhat it doesControlled by
schedule-syncs-dailydayRe-ingests sources whose sync frequency is Daily.The source’s Sync setting in Knowledge.
schedule-syncs-weeklyweekThe same for Weekly sources.The source’s Sync setting.
schedule-syncs-monthly30 daysThe same for Monthly sources.The source’s Sync setting.
dispatch-scheduled-runs30 sQueues scheduled agent runs that are due.SCHEDULE_DISPATCHER_INTERVAL (30, minimum 15 s).
cleanup-schedule-runsdayDeletes scheduled-run records and their output older than the retention window, keeping the 50 most recent runs of each schedule.SCHEDULE_RUN_OUTPUT_RETENTION_DAYS (90).
reconciliation30 sMarks stuck work as failed and logs an alert: answers that stopped streaming, tool calls left mid-way, stalled ingests, stuck idempotency claims and scheduled runs.None.
cleanup-pending-tool-state60 sDeletes paused conversations waiting for a tool approval or a client-side tool result once they expire (30 minutes), and clears their approval prompts.None.
cleanup-idempotency-deduphourDeletes idempotency keys for uploads, tasks and webhooks older than 24 hours.None.
cleanup-message-eventsdayDeletes the per-chunk stream journal that lets a client reconnect to an answer, and the markers left by regenerated answers. Conversations themselves are kept.MESSAGE_EVENTS_RETENTION_DAYS (14).
cleanup-guardrail-eventsdayDeletes guardrail events.GUARDRAILS_EVENTS_RETENTION_DAYS (30).
cleanup-tracesdayDeletes request traces (the trace viewer and per-agent logs).TRACES_RETENTION_DAYS (30).
cleanup-orphan-memoriesdayDeletes memory entries whose tool was deleted.None.
reap-sandbox-sessions60 sCloses sandbox sessions in the worker that have been idle longer than their keep-alive time.SANDBOX_MAX_TTL (1200) caps the keep-alive.
reap-stale-workflow-runs5 minMarks workflow runs still running after a disconnect or crash as failed.WORKFLOW_RUN_STALE_SECONDS (3600).
version-check7 hoursChecks gptcloud.arc53.com for security advisories for your version and logs any it finds.VERSION_CHECK (true); set false to turn it off.

All the settings above are listed in the Settings Reference.

Kept until deleted

No task trims the following. They grow for as long as the instance runs, and they are removed only when a user, an admin or you delete them:

  • Conversations and their messages. Users delete them one by one or all at once in the web app.
  • Token usage (token_usage). Quotas and analytics read it.
  • Request logs (stack_logs, user_logs). The Logs pages and analytics read stack_logs.
  • The audit logs: auth_events (sign-ins, access and configuration changes) and device_audit_log (remote-device commands).
  • Tool call records (tool_call_attempts) and workflow runs (workflow_runs).
  • Sources, attachments and artifacts, including their files in storage and their vectors.
  • Agents, prompts, tools, memories, notes and todos.

If you need a retention limit on any of these, for example for GDPR, delete old rows yourself on a schedule, such as DELETE FROM stack_logs WHERE timestamp < now() - interval '180 days'. Back up the database first, and keep token_usage as long as your quota periods need it.