Skip to Content
ModelsπŸ–₯️ Local Inference

Connecting DocsGPT to Local Inference Engines

DocsGPT can be configured to leverage local inference engines, allowing you to run Large Language Models directly on your own infrastructure. This approach offers enhanced privacy and control over your LLM processing.

Currently, DocsGPT primarily supports local inference engines that are compatible with the OpenAI API format. This means you can connect DocsGPT to various local LLM servers that mimic the OpenAI API structure.

Configuration via .env file

Setting up a local inference engine with DocsGPT is configured through environment variables in the .env file. For a detailed explanation of all settings, please consult the DocsGPT Settings Guide.

To connect to a local inference engine, you will generally need to configure these settings in your .env file:

  • LLM_PROVIDER: Crucially set this to openai. This tells DocsGPT to use the OpenAI-compatible API format for communication, even though the LLM is local.
  • LLM_NAME: Required. The model name as your inference engine names it (for Ollama, for example llama3.2:1b). DocsGPT registers exactly the models listed here, so with LLM_NAME unset or None no model is available and every chat fails. To offer several models, separate them with commas: LLM_NAME=llama3.2:1b,qwen2.5:7b. The first one is the default.
  • OPENAI_BASE_URL: This is essential. Set this to the base URL of your local inference engine’s API endpoint. This tells DocsGPT where to find your local LLM server.
  • API_KEY: Generally, for local inference engines, you can set API_KEY=None as authentication is usually not required in local setups.

Setting OPENAI_BASE_URL hides the hosted DocsGPT model and the OpenAI catalog, so chats go only to your server unless you also set another provider’s key and pick one of its models.

llama.cpp

Run llama.cpp as its OpenAI-compatible server and point DocsGPT at it like any other engine below. There is no separate llama.cpp provider: LLM_PROVIDER=llama.cpp is not a provider name, and the API logs an error at startup if you set it.

llama-server -m ./models/your-model.gguf --port 8000 --alias local-model --jinja
LLM_PROVIDER=openai OPENAI_BASE_URL=http://localhost:8000/v1 LLM_NAME=local-model API_KEY=None

--jinja lets the server handle tool calls, which agents with tools need; recent llama-server builds turn it on by default.

Supported Local Inference Engines (OpenAI API Compatible)

DocsGPT is also readily configurable to work with the following local inference engines, all communicating via the OpenAI API format. Here are example OPENAI_BASE_URL values for each, based on default setups:

Inference EngineLLM_PROVIDEROPENAI_BASE_URL
LLaMa.cpp (server mode)openaihttp://localhost:8000/v1
Ollamaopenaihttp://localhost:11434/v1
Text Generation Inference (TGI)openaihttp://localhost:8080/v1
SGLangopenaihttp://localhost:30000/v1
vLLMopenaihttp://localhost:8000/v1
Aphroditeopenaihttp://localhost:2242/v1
FriendliAIopenaihttp://localhost:8997/v1
LMDeployopenaihttp://localhost:23333/v1

Important Note on localhost vs host.docker.internal:

The OPENAI_BASE_URL examples above use http://localhost. If you are running DocsGPT within Docker and your local inference engine is running on your host machine (outside of Docker), you will likely need to replace localhost with http://host.docker.internal to ensure Docker can correctly access your host’s services. For example, http://host.docker.internal:11434/v1 for Ollama.

Plain http and your API key

DocsGPT sends a configured API key (API_KEY, OPENAI_API_KEY, a model YAML’s key) over plain http only when the server is on your own network: localhost and other loopback addresses, private addresses (10.x, 172.16–31.x, 192.168.x, fc00::/7), link-local addresses, a single-label host such as a Docker Compose service (http://ollama:11434/v1, http://litellm:4000), host.docker.internal, and names ending in .local, .svc, .internal or .localhost (for example http://ollama.llm.svc.cluster.local:11434/v1). Any other host needs https. Otherwise the request fails with Refusing to send a ... request to http://...: the endpoint uses plain http, and nothing is sent.

If your server sits on your own network under another name, such as http://llm.corp.example, set LLM_ALLOW_PLAINTEXT_ENDPOINTS=true. The key still goes only to the endpoint you configured.

How the Model Registry Works

DocsGPT uses a Model Registry to decide which models the model picker offers and which one answers by default.

Automatic Model Detection

At startup the registry loads the model catalog (docsgpt/core/models/*.yaml, plus any YAMLs in MODELS_CONFIG_DIR) and registers the models of every provider that has a key: the provider’s own variable (OPENAI_API_KEY, ANTHROPIC_API_KEY, …) or API_KEY for the provider named in LLM_PROVIDER. See Provider API keys for the full list. The hosted DocsGPT model is always registered unless OPENAI_BASE_URL is set.

Custom OpenAI-Compatible Models

When you set OPENAI_BASE_URL along with LLM_PROVIDER=openai and LLM_NAME, the registry creates one model entry per name in LLM_NAME, pointing at your inference server. This is how local engines like Ollama, vLLM, and others get registered.

Default Model Selection

The default model is the one the web app preselects and the one API and widget requests use when they name no model. The registry picks it in this order:

  1. The first name in LLM_NAME that is a registered model id. A name that is not registered is ignored, with a warning in the log.
  2. Otherwise, the first registered model of the provider named in LLM_PROVIDER.
  3. Otherwise, the first registered model, which is the hosted DocsGPT model when it is registered.

Step 3 means that if the provider in LLM_PROVIDER registered no model, for example because its key is missing, chats go to the public DocsGPT API. The API and the worker log an ERROR at startup in that case, and docsgpt doctor reports it. Setting LLM_NAME to a catalog id, or to your server’s model name, makes the default explicit.

Multiple Providers

You can configure multiple API keys simultaneously (e.g., both OPENAI_API_KEY and ANTHROPIC_API_KEY). The registry loads models from all configured providers, so users can switch between them in the UI; LLM_PROVIDER and LLM_NAME still decide the default.

Adding Support for Other Local Engines

While DocsGPT currently focuses on OpenAI API compatible local engines, you can extend it to support other local inference solutions. An integration has two parts:

  1. An LLM class in docsgpt/llm/ (a subclass of BaseLLM in docsgpt/llm/base.py) that calls your engine. The existing modules, such as openai.py and anthropic.py, are examples.
  2. A provider plugin in docsgpt/llm/providers/ (a subclass of Provider in base.py) that names the provider, points at the LLM class and says how to find its API key. Add an instance to ALL_PROVIDERS in docsgpt/llm/providers/__init__.py, add its name to ModelProvider in docsgpt/core/model_settings.py, and list its models in a YAML under docsgpt/core/models/ or in MODELS_CONFIG_DIR.

For any engine with an OpenAI-compatible API, no code is needed: use OPENAI_BASE_URL as above, or an openai_compatible model YAML.

Last updated on