Background Jobs and Auto-Resume
A chat turn no longer waits on slow work. When a tool call takes longer than a short window, it becomes a background job: the turn ends normally, the work keeps going, and when it finishes the agent comes back to the same conversation with the result, whether or not the user is still looking.
How a call becomes a job
Every server-side tool call in a chat turn runs in a small per-process pool. The turn waits up to
BACKGROUND_YIELD_SECONDS (30 by default):
- A call that finishes in time behaves exactly as before, with the same result.
- A call still running becomes a job. The model gets a result like this instead of the tool’s output:
{
"status": "running",
"job_id": "4f6c…",
"started_as": "background",
"elapsed_s": 30,
"note": "Still running as background job 4f6c…; you are resumed automatically with its result when it finishes. Do NOT run it again…"
}The model is told not to run the call again, not to wait on it and not to guess its result, and usually ends its turn by telling the user in plain words that the work is running (without the job id or how it is resumed). The call’s entry in the conversation shows as running until the job ends.
The model can also start a call as a job at once with background=true. The argument is offered on
code_executor, MCP tools, API tools and read_webpage, in chat turns only. A code_executor run started
that way without a timeout may run for the whole job lifetime (BACKGROUND_JOB_MAX_SECONDS, within
SANDBOX_EXEC_MAX_TIMEOUT) instead of the 60-second default. A background run that hits its timeout is
reported to the user, never re-run unless they ask.
Some calls are never handed off: client-side tools (the /v1 client tool calls), calls inside workflow nodes,
and scheduled and webhook runs. A tool that needs approval starts its clock only after it is approved.
Where a job runs
| Runner | When | Survives an API restart |
|---|---|---|
| Sandbox | Every code_executor run on Daytona; a background=true run on the Jupyter runner | Yes |
| Worker | background=true on other tools: the whole call runs in a Celery worker | Until the worker dies |
| In-process | Any other call that outlived the window (and Jupyter runs, which keep their kernel variables) | No |
A sandbox job is a separate process in the sandbox. A Celery poller follows it, keeps the sandbox from
auto-stopping, captures the files it writes as artifacts and reports the same result the turn would have.
On Daytona every chat run starts this way (about 60 ms more than before for a trivial run), so a hand-off
costs nothing. On the Jupyter runner a normal run stays in the kernel, so variables persist; a background=true
run is a separate Python process that does not share the kernel’s variables, only the workspace files.
A job that a process was running when the process died is marked lost within about two minutes. The agent is resumed with a note that the work may or may not have taken effect and should be verified before retrying.
Auto-resume
When a job finishes in a conversation auto-resume applies to (see below for the poll-only ones), the
conversation is resumed with a continuation turn: a headless turn of the same agent
and model, given the result as an event that is marked as not coming from the user. It answers in the same
conversation, and a conversation.continued event tells open tabs about the new message.
- One delivery. A result reaches the model once. If the model already read it with
check_job, or the user wrote before it was delivered, the continuation does not repeat it. - Defer while busy. If a reply is streaming in the conversation, or a tool call is waiting for approval, the continuation waits and retries.
- Folded into the user’s message. If the user writes first, the result is put in front of their message for the model (their message itself is stored unchanged).
- Batched. Results that land together, up to eight, are answered in one turn.
- Silence is allowed. If nothing is worth telling the user, the model answers
NO_REPLYand nothing is added to the conversation. - A turn like any other. A continuation’s tool calls and files belong to its own message, a slow call in it
becomes a background job too, and it has
check_job(to cancel a failing job a watch pattern reported, for example). Chains stay bounded byAUTO_RESUME_MAX_CONSECUTIVEand the job caps. - Safe by default. A continuation can’t use tools that need approval, and data from tools and outside services is presented to the model as data, never as instructions.
- Guarded like any tool result. The agent’s
tool_resultguardrails (and the instance floor) scan a job’s final result and last output before they are stored, and the data an event carries before it is queued. A block replaces the text with a “withheld by a content policy” note; a redaction masks what it matched. Every later reader (the continuation, the user’s next message,check_job, the job card) gets the scanned text. - Bounded. At most
AUTO_RESUME_MAX_CONSECUTIVEcontinuations follow one another without a user message.
Monitors, trigger links and approval links resume the conversation the same way, with
source monitor, trigger or approval, or monitor_paused and monitor_expired when a monitor reports on
itself.
Continuations are billed to the conversation’s owner like any turn, and only run in conversations the owner
holds outright. Conversations through the OpenAI-compatible /v1 API, an agent API key or the widget, a
shared agent’s public link, or an agent someone else owns get poll-only jobs: the running result says so,
and the result waits for check_job or the next message.
check_job
Every chat turn that can hand calls off has a check_job tool. The server attaches it, so it is not in
DEFAULT_CHAT_TOOLS, the Tools picker, Settings > Tools or the agent tool picker, and no tool setting removes it:
check_job(job_id?, action="get" | "cancel", wait_seconds=0)getreturns the result of a finished job (and takes it, so the conversation is not resumed with it again), or the job’s progress and latest output while it runs.wait_secondswaits up to 30 seconds for it to finish, except for a job the same turn started: that one answers at once, since its result resumes the conversation anyway.- Without a
job_id, it lists this conversation’s jobs. It never sees other conversations’ jobs. cancelasks the job to stop. A sandbox run is stopped; an in-process call can’t be interrupted and endscancelledwhen it returns. A step already under way may still complete.
Statuses follow MCP Tasks: working, completed, failed, cancelled. A lost job reads as failed with the
lost note.
watch
A code_executor call can carry a watch spec for what its output may do while it runs:
"watch": {"patterns": ["Traceback|Error|Killed"], "progress_regex": "PROGRESS (\\d+)%", "heartbeat_s": 0}patterns(up to five) resume the agent while the job still runs, with the matching line: at most one wake every 15 seconds, and after three dropped matches or eight wakes only completion wakes it.progress_regexupdates the job’s progress (sent injob.updated). It never wakes the model.heartbeat_s(0 for off, else at least 60) resumes the agent with the output printed since the last heartbeat, at most 2000 characters, skipped when nothing new was printed.
Watch reads the call’s own output as it is printed, so a run should print progress itself (flush=True); the
output of a child process captured with subprocess.run(..., capture_output=True) only appears when it ends.
In a conversation that auto-resume applies to, completion and failure always resume the agent; in a poll-only
one (see above) the result waits for check_job or the next message, and watch patterns and heartbeats wake
nothing. Output during a run is available for sandbox jobs only; for other jobs progress_regex is applied to
the final result.
Events
The app’s event stream (/api/events) carries:
job.updated:{job_id, conversation_id, status, progress, tool_name, action_name}conversation.continued:{conversation_id, message_id, source}notification.created:{kind, title, body, url, conversation_id}
A continuation’s message carries metadata.wake = {source, ref_id, dedupe_key}. Jobs can be read and cancelled
through GET /api/background_jobs?conversation_id=…, GET /api/background_jobs/<id> and
POST /api/background_jobs/<id>/cancel.
In the chat
- Job card. A handed-off call shows as a card in place of its tool call: running, with the elapsed time,
the progress and the latest output line from
watch, and a Cancel button; then finished, failed (with the error), cancelled or interrupted. It followsjob.updatedlive, and polls the conversation’s job list every 10 seconds while the event stream is down. - Woken turns. A continuation’s prompt is the event that woke the agent, not something the user wrote. It shows as a one-line event row (“Background job finished”, “Monitor matched”, “Webhook received”, “Approval received”, “Monitor paused”, “Monitor expired”, “Background job interrupted”) with the event text behind a toggle, never as a user message. A failed continuation offers no Retry, which would send the event as the user’s own words. Shared conversations show it the same way, without the internal ids.
- Live. The new answer appears in an open conversation without a reload. If the user’s own answer is streaming at that moment, it appears when that stream ends.
Notifications
When a continuation answers (never for NO_REPLY), or a one-time scheduled task (a reminder) answers in its
conversation, the user is notified according to where they are:
- Watching the conversation (a visible tab shows it): nothing; the chat updates.
- A DocsGPT tab open elsewhere: a toast with an Open button that goes to the conversation. A hidden tab also shows a system notification, if the browser allows it.
- No tab open: a Web Push notification to the browsers the user turned it on in, if the operator set it up (below).
Unless the user is watching, the conversation is also marked unread in the sidebar until they open it.
Tabs report what they show to POST /api/presence ({tab_id, conversation_id, visible, closing?}) when their
route or visibility changes and every 20 seconds while visible; a report counts for 45 seconds, and a closing
tab says so. When Redis is unavailable, the user counts as away, so the notification still goes out.
The toast and the push lead with the kind of event:
| Kind | Heading |
|---|---|
job | Background job finished |
lost | Background job interrupted |
monitor | Monitor matched |
monitor_paused | Monitor paused |
monitor_expired | Monitor expired |
trigger | Webhook received |
approval | Approval received |
schedule | Scheduled task finished |
Then comes what the event is about (the conversation’s name for a job, a monitor’s description, a schedule’s name) and the start of the answer as plain text, Markdown removed. Any other kind shows its own title. Toasts are translated; Web Push is sent in English, since the server does not know the browser’s language.
Setting up Web Push
Web Push is off until you give the server a VAPID key pair. Without it, users with no tab open get the unread mark only.
-
Generate a key pair, in the raw base64url format browsers use. Either command works:
npx web-push generate-vapid-keyspython -c "import base64; from cryptography.hazmat.primitives import serialization as s; from cryptography.hazmat.primitives.asymmetric import ec; k = ec.generate_private_key(ec.SECP256R1()); b = lambda r: base64.urlsafe_b64encode(r).rstrip(b'=').decode(); print('WEBPUSH_VAPID_PRIVATE_KEY=' + b(k.private_numbers().private_value.to_bytes(32, 'big'))); print('WEBPUSH_VAPID_PUBLIC_KEY=' + b(k.public_key().public_bytes(s.Encoding.X962, s.PublicFormat.UncompressedPoint)))" -
Set them on the API and the worker (both read them; the worker sends):
WEBPUSH_VAPID_PRIVATE_KEY=<private key> WEBPUSH_VAPID_PUBLIC_KEY=<public key> # optional: derived from the private key when unset WEBPUSH_VAPID_SUBJECT=mailto:ops@example.comWEBPUSH_VAPID_SUBJECTis how the push services reach you. Apple rejects pushes without a realmailto:orhttps:contact; unset, the public API URL is used when it is https. A key that doesn’t parse, or a public key that doesn’t match the private one, turns Web Push off and logs an error. -
Serve the app over HTTPS. Browsers allow service workers and push only in a secure context (
localhostcounts, for development). The service worker is/sw.js, served from the UI’s origin root by Vite, the nginx image and the API’s built-in UI (SERVE_UI) alike. It shows notifications only: it never caches or intercepts requests. A UI served under a sub-path can’t register it. -
Let the worker reach the push services:
fcm.googleapis.com,*.push.services.mozilla.com,*.push.apple.comand*.notify.windows.comover https. The server posts only to these hosts (plusWEBPUSH_EXTRA_ALLOWED_HOSTS, for a self-hosted push service or tests), never to an arbitrary endpoint a browser hands it, and never follows a redirect.
The browser asks for permission only in context: the first time a job goes to the background, a monitor or link is set up, or a notification arrives, a “Get notified when it’s done?” toast offers Enable or Not now. The page never asks on load, and the answer is remembered in that browser. Signing out removes the browser’s subscription.
| Browser | Web Push |
|---|---|
| Chrome, Edge, Opera, Samsung Internet (desktop and Android) | Yes |
| Firefox (desktop and Android) | Yes |
| Safari 16.1+ on macOS 13+ | Yes |
| Safari on iOS and iPadOS 16.4+ | Only once DocsGPT is added to the Home Screen |
Rotating the key pair invalidates every subscription; browsers that allowed notifications subscribe again on their next visit. A subscription the push service reports gone (404 or 410) is deleted, and one that keeps failing is dropped.
Running it
Background jobs need a Celery worker and the beat scheduler. The worker runs continuations, sandbox pollers
and background=true calls; beat runs the sweep that reports lost jobs and redelivers results. See
Development Environment.
| Setting | Default | What it does |
|---|---|---|
BACKGROUND_JOBS_ENABLED | true | Hand slow calls off at all |
BACKGROUND_YIELD_SECONDS | 30 | How long a turn waits before handing a call off |
BACKGROUND_JOB_MAX_SECONDS | 1000 | Hard lifetime of a job |
BACKGROUND_MAX_JOBS_PER_CONVERSATION | 2 | Running jobs per conversation; over it a call stays in the foreground |
BACKGROUND_MAX_JOBS_PER_USER | 5 | Running jobs per user |
BACKGROUND_RESULT_RETENTION_DAYS | 7 | How long finished jobs are kept |
BACKGROUND_POOL_SIZE | 16 | Pool threads per process; a full pool runs calls inline as before |
BACKGROUND_LEASE_STALE_SECONDS | 60 | Heartbeat age after which a job is lost |
BACKGROUND_RECONCILE_INTERVAL_SECONDS | 60 | How often the sweep runs |
AUTO_RESUME_ENABLED | true | Resume conversations with finished results |
AUTO_RESUME_MAX_RESULT_CHARS | 16000 | Result length a continuation sees (head and tail kept) |
AUTO_RESUME_MAX_CONSECUTIVE | 5 | Continuations in a row without a user message |
WEBPUSH_VAPID_PRIVATE_KEY | unset | VAPID private key; unset, Web Push is off |
WEBPUSH_VAPID_PUBLIC_KEY | unset | VAPID public key; derived from the private key when unset |
WEBPUSH_VAPID_SUBJECT | unset | Contact for the push services (mailto: or https:) |
WEBPUSH_EXTRA_ALLOWED_HOSTS | empty | Push service hosts to accept besides the known ones |
Keep BACKGROUND_JOB_MAX_SECONDS below SANDBOX_MAX_TTL, so an idle sandbox session is not closed under a
running job. See the Settings Reference for every setting.
Limits
- The tool-result guardrail stage runs on results inside a turn; a background result reaches the model in a continuation turn, as event data, without that stage.
- Images a tool shows the model (charts) are shown only in the turn the call ran in; a job’s charts are saved as artifacts instead.
- MCP Tasks (an MCP server running a tool as an upstream task) are not used yet; such calls hand off like any other.