LLM Models
The LLM Models catalog stores the chat-capable model connections that CoreCube can use. The Orchestrator selects one ready catalog entry for query resolution, orchestration decisions, and final answer generation.
Open Admin Console → Query Pipeline → LLM Models.
Supported connections
| Connection | Configuration |
|---|---|
| Anthropic | Anthropic endpoint, model, and API key. |
| OpenAI | OpenAI endpoint, model, and API key. |
| Google Gemini | Gemini endpoint, model, and API key. |
| Mistral AI | Mistral endpoint, model, and API key. |
| xAI | xAI endpoint, model, and API key. |
| Local | A chat model exposed by the local inference registry. |
| Custom | Any compatible endpoint and model entered manually. |
Custom endpoints are commonly used for vLLM, Ollama, LM Studio, or another OpenAI-compatible
server. Enter the API base URL, normally ending in /v1; CoreCube appends /chat/completions when
it sends a chat request.
Add a model
- Select Add Model.
- Choose a cloud provider, local model, or custom endpoint.
- Enter a display name and select or enter the provider's model identifier.
- Supply credentials when the provider requires them.
- Keep the entry active and use Test to verify the connection.
- Save the model.
Cloud providers support model discovery when valid credentials are available. Local choices come from the configured inference deployment and model catalog.
The first catalog entry becomes the default and remains active. Setting another entry as default promotes it and clears the previous default.
Select the Orchestrator model
Registering or marking a model as default does not update the active answer runtime automatically. Open Configuration → Orchestrator → LLM, select the intended ready model, then Save and Apply the Orchestrator revision.
One selected model is used for:
- standalone-query resolution;
- every Orchestrator decision round;
- final answer generation from frozen evidence.
The Orchestrator also stores model settings such as temperature and verifies that the model's reported context capacity is sufficient for each configured Search mode.
The model binding is copied into the immutable active snapshot. Editing the catalog entry later causes a dependency change but does not silently alter that snapshot.
References and deletion
The catalog shows references from Pipelines, tools, and legacy configuration. Before deactivating or deleting a model, repoint those references and apply an updated Orchestrator snapshot.
CoreCube blocks unsafe deletion when the model is still required. A model already embedded in an active snapshot remains part of that immutable snapshot until another revision is applied.
Provider limits
When CC_PROVIDER_RATE_LIMITS_ENABLED=true (the default), CoreCube checks configured provider
request and token limits before dispatch. A request that exceeds an available limit fails with a
typed 429 instead of knowingly sending work the provider is expected to reject.
The model page exposes connection tests and usage statistics. Runtime readiness is also checked from the Orchestrator because a valid catalog row can still be unavailable—for example, when credentials are missing, a local worker is offline, or the model's context capacity is insufficient.
Local-processing boundary
Selecting a local Orchestrator model keeps query-resolution, Orchestrator, and final-answer calls on that local endpoint. It does not automatically make every other model-backed stage local. Check the separate embedding, reranker, OCR, and tool model bindings before claiming a fully local deployment.
Streaming
Streaming is selected per /v1/chat/completions request, not per LLM Model. Buffered and streaming
requests use the same frozen Orchestrator policy and evidence process. See
Chat Completions.