Chat Completions
POST /v1/chat/completions
CoreCube accepts OpenAI-style chat messages, retrieves authorized organizational evidence, and returns a cited answer. The active Orchestrator snapshot controls the model, prompts, retrieval policy, tools, and budgets for the complete turn.
Request
POST /v1/chat/completions
Authorization: Bearer cc_YOUR_API_KEY
Content-Type: application/json
{
"messages": [
{
"role": "user",
"content": "What do our deployment runbooks say about rollbacks?"
}
],
"searchMode": "balanced",
"stream": false
}
OpenAI-compatible fields
| Field | Type | Default | Behavior |
|---|---|---|---|
messages | array | required | Conversation history. At least one message is required. |
stream | boolean | false | Return one JSON response or Server-Sent Events. |
model | string | omitted | Accepted for client compatibility; the applied Orchestrator model is used. |
temperature | number | omitted | Accepted from 0 to 2; the applied Orchestrator temperature is used. |
max_tokens | integer | omitted | Accepted up to 32768; the active mode's final-answer ceiling is used. |
Message content must be a string. Supported roles are system, user, and assistant.
Unrecognized OpenAI fields such as top_p, n, client-supplied tools, multimodal content arrays,
and role: "tool" are not part of the CoreCube request contract.
The compatibility fields do not override the frozen server policy. To change the answer model, temperature, or maximum output budget, save and apply an Orchestrator revision.
CoreCube fields
| Field | Type | Default | Description |
|---|---|---|---|
searchMode | string | balanced | fast, balanced, or accurate; selects the applied mode policy. |
connectionIds | string[] | all allowed | Narrow retrieval to specific accessible connections. |
compartmentId | string | context-dependent | Select an authorized compartment. Required when the key can access several. |
sensitivityCeiling | string | authorized maximum | Narrow retrieval to public, internal, confidential, or restricted. |
cube_extended | boolean | false | Add structured citations, diagnostics, confidence, and audit metadata. |
connectionIds and sensitivityCeiling can only narrow authorization. They cannot grant access
beyond the API key and effective user's allowed scope.
Use X-Cube-Extended: true as the header equivalent of cube_extended.
Standard response
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1753000000,
"model": "corecube",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Stop the deployment when health checks fail, then restore the previous container image [1].\n\nSources:\n[1] Deployment Runbook (Confluence — Engineering)"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1842,
"completion_tokens": 127,
"total_tokens": 1969
}
}
For ordinary callers, CoreCube appends a rendered Sources section when authorized citations are
available. A service key configured for OpenWebUI receives its compatible sources array instead.
Extended response
Set cube_extended: true or X-Cube-Extended: true to add the cube object:
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1753000000,
"model": "corecube",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Stop the deployment when health checks fail, then restore the previous container image [1].\n\nSources:\n[1] Deployment Runbook (Confluence — Engineering)"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1842,
"completion_tokens": 127,
"total_tokens": 1969
},
"cube": {
"citations": [
{
"citationId": "citation_1",
"citationIndex": 1,
"title": "Deployment Runbook",
"sourceLabel": "Confluence — Engineering",
"sourcePath": "connector",
"snippet": "Stop the deployment if health checks fail.",
"evidenceAsOfAt": "2026-08-30T10:00:00.000Z",
"evidenceAsOfBasis": "connector_modified",
"url": "https://corecube.example/v1/documents/document_1/file"
}
],
"attachments": [],
"searchLatencyMs": 142,
"llmLatencyMs": 1840,
"chunksRetrieved": 3,
"filesRetrieved": 1,
"initialCandidateCount": 12,
"initialAdmittedChunkCount": 3,
"expansionSearchCount": 0,
"expansionAdmittedChunkCount": 0,
"finalEvidenceChunkCount": 3,
"finalEvidenceDocumentCount": 1,
"searchMode": "balanced",
"rerankerOutcome": "reranked",
"confidence": {
"confidencePercent": 92,
"verificationStatus": "verified"
},
"auditReference": "audit_1",
"auditStatus": "written",
"degraded": false,
"emptyReason": null,
"degradations": []
}
}
Extended metadata
| Field | Meaning |
|---|---|
citations | Authorized citation map with snippets and evidence dates. |
attachments | Authorized files explicitly attached by the answer. |
initialCandidateCount | Candidates found by the initial Discovery operation. |
initialAdmittedChunkCount | Discovery chunks admitted to the evidence ledger. |
expansionSearchCount | Later Pipeline or direct-tool searches. |
expansionAdmittedChunkCount | Chunks admitted from later actions. |
finalEvidenceChunkCount | Chunks in the frozen final evidence set. |
finalEvidenceDocumentCount | Distinct documents in the frozen evidence set. |
searchMode | Effective Fast, Balanced, or Accurate policy. |
rerankerOutcome | reranked, disabled, degraded, or not_applicable. |
confidence | Nullable confidence percentage and verification status. |
auditReference / auditStatus | Audit record identity and persistence outcome. |
degraded / degradations | Whether a stage degraded and bounded public details. |
Streaming
Set "stream": true to receive text/event-stream frames:
data: {"id":"chatcmpl-abc","choices":[{"delta":{"role":"assistant"},"index":0}]}
data: {"id":"chatcmpl-abc","choices":[{"delta":{"content":"According"},"index":0}]}
data: [DONE]
Buffered and streaming requests use the same Orchestrator decisions, Pipeline and direct-tool actions, evidence admission, and final frozen evidence. Streaming supports the current Orchestrator tool loop; only final answer delivery is streamed to the client.
In extended mode, CoreCube sends a terminal metadata frame before [DONE]:
data: {"object":"cube.metadata","cube":{"citations":[],"searchMode":"balanced","finalEvidenceChunkCount":3}}
data: [DONE]
Insufficient evidence
After mandatory Discovery, the Orchestrator must obtain answer-bearing evidence from an eligible Pipeline or direct evidence tool before generating a grounded answer. Discovery-only chunks guide action selection but do not support the final answer. If the authorized sources still do not support the requested claim, CoreCube returns an evidence-limited answer rather than filling the gap with ungrounded model knowledge.
The exact response can describe unresolved evidence needs. With extended mode, inspect
confidence.verificationStatus, emptyReason, and degradations for structured diagnostics.
Rate limits and concurrency
CoreCube can return 429 before provider work starts:
| Cause | Response |
|---|---|
| API-key request-per-minute limit exceeded | 429 with the API-key rate-limit error. |
| Chat concurrency limit exceeded | 429, code CONCURRENCY_LIMIT_EXCEEDED, Retry-After: 1. |
| Provider RPM/TPM catalog limit exceeded | 429, code provider_rate_limited, with Retry-After. |
Global chat concurrency defaults to five times the app database pool (100 with the default pool).
The other defaults are 100 per API key, 8 per effective user, and 8 per source IP for public keys. See
Environment Variables.
Examples
cURL
curl https://corecube.your-domain.com/v1/chat/completions \
-H "Authorization: Bearer cc_YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Who owns the authentication service?"}],
"searchMode": "balanced",
"stream": false
}'
Python
The OpenAI SDK requires a model argument. Use a stable placeholder such as CoreCube; the applied
Orchestrator model remains authoritative.
from openai import OpenAI
client = OpenAI(
api_key="cc_YOUR_API_KEY",
base_url="https://corecube.your-domain.com/v1"
)
response = client.chat.completions.create(
model="CoreCube",
messages=[{"role": "user", "content": "What is our data retention policy?"}]
)
print(response.choices[0].message.content)
JavaScript
import OpenAI from 'openai';
const client = new OpenAI({
apiKey: 'cc_YOUR_API_KEY',
baseURL: 'https://corecube.your-domain.com/v1',
});
const response = await client.chat.completions.create({
model: 'CoreCube',
messages: [{ role: 'user', content: 'What are our deployment procedures?' }],
});
console.log(response.choices[0].message.content);
Service key
Service keys can map each request to an effective user so CoreCube enforces that user's permissions:
curl https://corecube.your-domain.com/v1/chat/completions \
-H "Authorization: Bearer cc_SERVICE_API_KEY" \
-H "X-Cube-User: sarah@company.com" \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Show me the Q4 budget projections"}]}'
Public key
Public keys require no user identity. Any identity header is ignored and the key's bound scope always applies.
curl https://corecube.your-domain.com/v1/chat/completions \
-H "Authorization: Bearer cc_PUBLIC_API_KEY" \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "How do I install CoreCube?"}]}'
Configure CORS on the reverse proxy for browser clients. CoreCube expects the proxy to handle preflight rather than writing per-key CORS headers.
Model listing
GET /v1/models
Authorization: Bearer cc_YOUR_API_KEY
The response contains one OpenAI-format model entry representing the CoreCube service. This lets compatible clients discover CoreCube without exposing the internal Orchestrator model binding.
Other /v1 endpoints
| Endpoint | Purpose |
|---|---|
GET /v1/tools | List tools available to the calling key. |
GET /v1/tools/:slug | Read one tool definition. |
POST /v1/tools/:slug/invoke | Invoke an authorized tool. |
GET /v1/documents/:id/file | Fetch a cited source file using its signed citation token. |