Connecting DocsGPT to Local Inference Engines
DocsGPT can be configured to leverage local inference engines, allowing you to run Large Language Models directly on your own infrastructure. This approach offers enhanced privacy and control over your LLM processing.
Currently, DocsGPT primarily supports local inference engines that are compatible with the OpenAI API format. This means you can connect DocsGPT to various local LLM servers that mimic the OpenAI API structure.
Configuration via .env file
Setting up a local inference engine with DocsGPT is configured through environment variables in the .env file. For a detailed explanation of all settings, please consult the DocsGPT Settings Guide.
To connect to a local inference engine, you will generally need to configure these settings in your .env file:
LLM_PROVIDER: Crucially set this toopenai. This tells DocsGPT to use the OpenAI-compatible API format for communication, even though the LLM is local.LLM_NAME: Required. The model name as your inference engine names it (for Ollama, for examplellama3.2:1b). DocsGPT registers exactly the models listed here, so withLLM_NAMEunset orNoneno model is available and every chat fails. To offer several models, separate them with commas:LLM_NAME=llama3.2:1b,qwen2.5:7b. The first one is the default.OPENAI_BASE_URL: This is essential. Set this to the base URL of your local inference engineβs API endpoint. This tells DocsGPT where to find your local LLM server.API_KEY: Generally, for local inference engines, you can setAPI_KEY=Noneas authentication is usually not required in local setups.
Setting OPENAI_BASE_URL hides the hosted DocsGPT model and the OpenAI catalog, so chats go only to your server unless you also set another providerβs key and pick one of its models.
llama.cpp
Run llama.cpp as its OpenAI-compatible server and point DocsGPT at it like any other engine below. There is no separate llama.cpp provider: LLM_PROVIDER=llama.cpp is not a provider name, and the API logs an error at startup if you set it.
llama-server -m ./models/your-model.gguf --port 8000 --alias local-model --jinjaLLM_PROVIDER=openai
OPENAI_BASE_URL=http://localhost:8000/v1
LLM_NAME=local-model
API_KEY=None--jinja lets the server handle tool calls, which agents with tools need; recent llama-server builds turn it on by default.
Supported Local Inference Engines (OpenAI API Compatible)
DocsGPT is also readily configurable to work with the following local inference engines, all communicating via the OpenAI API format. Here are example OPENAI_BASE_URL values for each, based on default setups:
| Inference Engine | LLM_PROVIDER | OPENAI_BASE_URL |
|---|---|---|
| LLaMa.cpp (server mode) | openai | http://localhost:8000/v1 |
| Ollama | openai | http://localhost:11434/v1 |
| Text Generation Inference (TGI) | openai | http://localhost:8080/v1 |
| SGLang | openai | http://localhost:30000/v1 |
| vLLM | openai | http://localhost:8000/v1 |
| Aphrodite | openai | http://localhost:2242/v1 |
| FriendliAI | openai | http://localhost:8997/v1 |
| LMDeploy | openai | http://localhost:23333/v1 |
Important Note on localhost vs host.docker.internal:
The OPENAI_BASE_URL examples above use http://localhost. If you are running DocsGPT within Docker and your local inference engine is running on your host machine (outside of Docker), you will likely need to replace localhost with http://host.docker.internal to ensure Docker can correctly access your hostβs services. For example, http://host.docker.internal:11434/v1 for Ollama.
Plain http and your API key
DocsGPT sends a configured API key (API_KEY, OPENAI_API_KEY, a model YAMLβs key) over plain http only when the server is on your own network: localhost and other loopback addresses, private addresses (10.x, 172.16β31.x, 192.168.x, fc00::/7), link-local addresses, a single-label host such as a Docker Compose service (http://ollama:11434/v1, http://litellm:4000), host.docker.internal, and names ending in .local, .svc, .internal or .localhost (for example http://ollama.llm.svc.cluster.local:11434/v1). Any other host needs https. Otherwise the request fails with Refusing to send a ... request to http://...: the endpoint uses plain http, and nothing is sent.
If your server sits on your own network under another name, such as http://llm.corp.example, set LLM_ALLOW_PLAINTEXT_ENDPOINTS=true. The key still goes only to the endpoint you configured.
How the Model Registry Works
DocsGPT uses a Model Registry to decide which models the model picker offers and which one answers by default.
Automatic Model Detection
At startup the registry loads the model catalog (docsgpt/core/models/*.yaml, plus any YAMLs in MODELS_CONFIG_DIR) and registers the models of every provider that has a key: the providerβs own variable (OPENAI_API_KEY, ANTHROPIC_API_KEY, β¦) or API_KEY for the provider named in LLM_PROVIDER. See Provider API keys for the full list. The hosted DocsGPT model is always registered unless OPENAI_BASE_URL is set.
Custom OpenAI-Compatible Models
When you set OPENAI_BASE_URL along with LLM_PROVIDER=openai and LLM_NAME, the registry creates one model entry per name in LLM_NAME, pointing at your inference server. This is how local engines like Ollama, vLLM, and others get registered.
Default Model Selection
The default model is the one the web app preselects and the one API and widget requests use when they name no model. The registry picks it in this order:
- The first name in
LLM_NAMEthat is a registered model id. A name that is not registered is ignored, with a warning in the log. - Otherwise, the first registered model of the provider named in
LLM_PROVIDER. - Otherwise, the first registered model, which is the hosted DocsGPT model when it is registered.
Step 3 means that if the provider in LLM_PROVIDER registered no model, for example because its key is missing, chats go to the public DocsGPT API. The API and the worker log an ERROR at startup in that case, and docsgpt doctor reports it. Setting LLM_NAME to a catalog id, or to your serverβs model name, makes the default explicit.
Multiple Providers
You can configure multiple API keys simultaneously (e.g., both OPENAI_API_KEY and ANTHROPIC_API_KEY). The registry loads models from all configured providers, so users can switch between them in the UI; LLM_PROVIDER and LLM_NAME still decide the default.
Adding Support for Other Local Engines
While DocsGPT currently focuses on OpenAI API compatible local engines, you can extend it to support other local inference solutions. An integration has two parts:
- An LLM class in
docsgpt/llm/(a subclass ofBaseLLMindocsgpt/llm/base.py) that calls your engine. The existing modules, such asopenai.pyandanthropic.py, are examples. - A provider plugin in
docsgpt/llm/providers/(a subclass ofProviderinbase.py) that names the provider, points at the LLM class and says how to find its API key. Add an instance toALL_PROVIDERSindocsgpt/llm/providers/__init__.py, add its name toModelProviderindocsgpt/core/model_settings.py, and list its models in a YAML underdocsgpt/core/models/or inMODELS_CONFIG_DIR.
For any engine with an OpenAI-compatible API, no code is needed: use OPENAI_BASE_URL as above, or an openai_compatible model YAML.