Skip to main content

Chat Completions

POST /v1/chat/completions

CoreCube accepts OpenAI-style chat messages, retrieves authorized organizational evidence, and returns a cited answer. The active Orchestrator snapshot controls the model, prompts, retrieval policy, tools, and budgets for the complete turn.

Request​

POST /v1/chat/completions
Authorization: Bearer cc_YOUR_API_KEY
Content-Type: application/json
{
"messages": [
{
"role": "user",
"content": "What do our deployment runbooks say about rollbacks?"
}
],
"searchMode": "balanced",
"stream": false
}

OpenAI-compatible fields​

FieldTypeDefaultBehavior
messagesarrayrequiredConversation history. At least one message is required.
streambooleanfalseReturn one JSON response or Server-Sent Events.
modelstringomittedAccepted for client compatibility; the applied Orchestrator model is used.
temperaturenumberomittedAccepted from 0 to 2; the applied Orchestrator temperature is used.
max_tokensintegeromittedAccepted up to 32768; the active mode's final-answer ceiling is used.

Message content must be a string. Supported roles are system, user, and assistant. Unrecognized OpenAI fields such as top_p, n, client-supplied tools, multimodal content arrays, and role: "tool" are not part of the CoreCube request contract.

The compatibility fields do not override the frozen server policy. To change the answer model, temperature, or maximum output budget, save and apply an Orchestrator revision.

CoreCube fields​

FieldTypeDefaultDescription
searchModestringbalancedfast, balanced, or accurate; selects the applied mode policy.
connectionIdsstring[]all allowedNarrow retrieval to specific accessible connections.
compartmentIdstringcontext-dependentSelect an authorized compartment. Required when the key can access several.
sensitivityCeilingstringauthorized maximumNarrow retrieval to public, internal, confidential, or restricted.
cube_extendedbooleanfalseAdd structured citations, diagnostics, confidence, and audit metadata.

connectionIds and sensitivityCeiling can only narrow authorization. They cannot grant access beyond the API key and effective user's allowed scope.

Use X-Cube-Extended: true as the header equivalent of cube_extended.

Standard response​

{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1753000000,
"model": "corecube",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Stop the deployment when health checks fail, then restore the previous container image [1].\n\nSources:\n[1] Deployment Runbook (Confluence — Engineering)"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1842,
"completion_tokens": 127,
"total_tokens": 1969
}
}

For ordinary callers, CoreCube appends a rendered Sources section when authorized citations are available. A service key configured for OpenWebUI receives its compatible sources array instead.

Extended response​

Set cube_extended: true or X-Cube-Extended: true to add the cube object:

{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1753000000,
"model": "corecube",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Stop the deployment when health checks fail, then restore the previous container image [1].\n\nSources:\n[1] Deployment Runbook (Confluence — Engineering)"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1842,
"completion_tokens": 127,
"total_tokens": 1969
},
"cube": {
"citations": [
{
"citationId": "citation_1",
"citationIndex": 1,
"title": "Deployment Runbook",
"sourceLabel": "Confluence — Engineering",
"sourcePath": "connector",
"snippet": "Stop the deployment if health checks fail.",
"evidenceAsOfAt": "2026-08-30T10:00:00.000Z",
"evidenceAsOfBasis": "connector_modified",
"url": "https://corecube.example/v1/documents/document_1/file"
}
],
"attachments": [],
"searchLatencyMs": 142,
"llmLatencyMs": 1840,
"chunksRetrieved": 3,
"filesRetrieved": 1,
"initialCandidateCount": 12,
"initialAdmittedChunkCount": 3,
"expansionSearchCount": 0,
"expansionAdmittedChunkCount": 0,
"finalEvidenceChunkCount": 3,
"finalEvidenceDocumentCount": 1,
"searchMode": "balanced",
"rerankerOutcome": "reranked",
"confidence": {
"confidencePercent": 92,
"verificationStatus": "verified"
},
"auditReference": "audit_1",
"auditStatus": "written",
"degraded": false,
"emptyReason": null,
"degradations": []
}
}

Extended metadata​

FieldMeaning
citationsAuthorized citation map with snippets and evidence dates.
attachmentsAuthorized files explicitly attached by the answer.
initialCandidateCountCandidates found by the initial Discovery operation.
initialAdmittedChunkCountDiscovery chunks admitted to the evidence ledger.
expansionSearchCountLater Pipeline or direct-tool searches.
expansionAdmittedChunkCountChunks admitted from later actions.
finalEvidenceChunkCountChunks in the frozen final evidence set.
finalEvidenceDocumentCountDistinct documents in the frozen evidence set.
searchModeEffective Fast, Balanced, or Accurate policy.
rerankerOutcomereranked, disabled, degraded, or not_applicable.
confidenceNullable confidence percentage and verification status.
auditReference / auditStatusAudit record identity and persistence outcome.
degraded / degradationsWhether a stage degraded and bounded public details.

Streaming​

Set "stream": true to receive text/event-stream frames:

data: {"id":"chatcmpl-abc","choices":[{"delta":{"role":"assistant"},"index":0}]}

data: {"id":"chatcmpl-abc","choices":[{"delta":{"content":"According"},"index":0}]}

data: [DONE]

Buffered and streaming requests use the same Orchestrator decisions, Pipeline and direct-tool actions, evidence admission, and final frozen evidence. Streaming supports the current Orchestrator tool loop; only final answer delivery is streamed to the client.

In extended mode, CoreCube sends a terminal metadata frame before [DONE]:

data: {"object":"cube.metadata","cube":{"citations":[],"searchMode":"balanced","finalEvidenceChunkCount":3}}

data: [DONE]

Insufficient evidence​

After mandatory Discovery, the Orchestrator must obtain answer-bearing evidence from an eligible Pipeline or direct evidence tool before generating a grounded answer. Discovery-only chunks guide action selection but do not support the final answer. If the authorized sources still do not support the requested claim, CoreCube returns an evidence-limited answer rather than filling the gap with ungrounded model knowledge.

The exact response can describe unresolved evidence needs. With extended mode, inspect confidence.verificationStatus, emptyReason, and degradations for structured diagnostics.

Rate limits and concurrency​

CoreCube can return 429 before provider work starts:

CauseResponse
API-key request-per-minute limit exceeded429 with the API-key rate-limit error.
Chat concurrency limit exceeded429, code CONCURRENCY_LIMIT_EXCEEDED, Retry-After: 1.
Provider RPM/TPM catalog limit exceeded429, code provider_rate_limited, with Retry-After.

Global chat concurrency defaults to five times the app database pool (100 with the default pool). The other defaults are 100 per API key, 8 per effective user, and 8 per source IP for public keys. See Environment Variables.

Examples​

cURL​

curl https://corecube.your-domain.com/v1/chat/completions \
-H "Authorization: Bearer cc_YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Who owns the authentication service?"}],
"searchMode": "balanced",
"stream": false
}'

Python​

The OpenAI SDK requires a model argument. Use a stable placeholder such as CoreCube; the applied Orchestrator model remains authoritative.

from openai import OpenAI

client = OpenAI(
api_key="cc_YOUR_API_KEY",
base_url="https://corecube.your-domain.com/v1"
)

response = client.chat.completions.create(
model="CoreCube",
messages=[{"role": "user", "content": "What is our data retention policy?"}]
)

print(response.choices[0].message.content)

JavaScript​

import OpenAI from 'openai';

const client = new OpenAI({
apiKey: 'cc_YOUR_API_KEY',
baseURL: 'https://corecube.your-domain.com/v1',
});

const response = await client.chat.completions.create({
model: 'CoreCube',
messages: [{ role: 'user', content: 'What are our deployment procedures?' }],
});

console.log(response.choices[0].message.content);

Service key​

Service keys can map each request to an effective user so CoreCube enforces that user's permissions:

curl https://corecube.your-domain.com/v1/chat/completions \
-H "Authorization: Bearer cc_SERVICE_API_KEY" \
-H "X-Cube-User: sarah@company.com" \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Show me the Q4 budget projections"}]}'

Public key​

Public keys require no user identity. Any identity header is ignored and the key's bound scope always applies.

curl https://corecube.your-domain.com/v1/chat/completions \
-H "Authorization: Bearer cc_PUBLIC_API_KEY" \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "How do I install CoreCube?"}]}'

Configure CORS on the reverse proxy for browser clients. CoreCube expects the proxy to handle preflight rather than writing per-key CORS headers.

Model listing​

GET /v1/models
Authorization: Bearer cc_YOUR_API_KEY

The response contains one OpenAI-format model entry representing the CoreCube service. This lets compatible clients discover CoreCube without exposing the internal Orchestrator model binding.

Other /v1 endpoints​

EndpointPurpose
GET /v1/toolsList tools available to the calling key.
GET /v1/tools/:slugRead one tool definition.
POST /v1/tools/:slug/invokeInvoke an authorized tool.
GET /v1/documents/:id/fileFetch a cited source file using its signed citation token.

We use cookies for analytics to improve our website. More information in our Privacy Policy.