Skip to main content

Orchestrator

The Orchestrator is CoreCube's query-time policy authority. It selects how the configured answer model evaluates evidence, which retrieval Pipelines and direct evidence tools it may use, when it must verify a result, and when it may finalize an answer.

The Orchestrator is not a general workflow or action runtime. Its tools are governed, read-only evidence operations, and every accepted result passes through CoreCube's authorization and evidence admission controls.

Where to find it

Open Admin Console → Configuration → Orchestrator.

What the Orchestrator owns​

The Orchestrator configuration contains:

  • one LLM for query resolution, orchestration decisions, and final answer generation;
  • ordered System Prompts for the Orchestrator and Answer phases;
  • separate policies for the Fast, Balanced, and Accurate Search modes;
  • the Pipelines and direct evidence tools available to each mode;
  • stopping rules and hard ceilings for work, evidence, tokens, and latency.

Pipelines remain reusable retrieval definitions. A Pipeline does not choose the answer model or make itself active. Applying an Orchestrator revision embeds the selected Pipeline definitions, tools, prompts, and model binding into one immutable runtime snapshot.

LLM​

The LLM tab selects one ready model from Query Pipeline → LLM Models. That model performs both Orchestrator decisions and final answer generation. Its configured temperature also applies to query resolution.

The model must be active, reachable, and capable of handling the context required by every enabled mode. CoreCube reports insufficient context capacity as a readiness problem before Apply.

Changing which LLM Model is marked as the registry default does not replace the model stored in an already-applied Orchestrator snapshot. Select the intended model here, save the revision, and apply it.

System Prompts​

The System Prompts tab maintains two independent ordered sequences:

PhaseControls
OrchestratorEvidence sufficiency, Pipeline selection, direct-tool use, and stopping.
AnswerThe final response generated from the frozen, admitted evidence.

Prompt content is reusable and shared. Its enabled state, phase, and position belong to the current Orchestrator revision. Save or discard an edited System Prompt before navigating away or saving the Orchestrator configuration.

Saved prompt changes do not mutate the active runtime snapshot. Apply a new Orchestrator revision to put the changed content into service.

See System Prompts for fragment types and other binding surfaces.

Modes​

The Modes tab configures Fast, Balanced, and Accurate independently. Clients select a mode with the searchMode field on /v1/chat/completions; Balanced is used when the field is omitted.

Each mode has four components:

ComponentPurpose
Eligible pipelinesOrdered retrieval strategies the Orchestrator may run.
Direct evidence toolsFocused, governed evidence operations available without running a Pipeline.
Decision rulesVerification, no-progress, source-count, and finalization behavior.
Budget ceilingsMaximum permitted actions, evidence, model tokens, and elapsed time.

Attaching a Pipeline makes the run_pipeline function available to that mode. The Orchestrator chooses a Pipeline from the eligible set using its declared purpose and server-derived capabilities. Adding several Pipelines provides choices; it does not require every Pipeline to run.

Seeded mode policies​

Fresh installations start with these bindings:

ModePipelinePipeline executionStopping behavior
FastExact LookupAdaptiveFirst sufficient
BalancedHybrid CoveragePrefer oneVerify if uncertain
AccurateDeep Document AnalysisRequire oneVerify once

The Pipeline execution setting means:

  • Adaptive — no Pipeline preference is imposed; a Pipeline or direct evidence tool can provide answer-bearing evidence after Discovery.
  • Prefer one Pipeline — normally run one eligible Pipeline before finalizing.
  • Require one Pipeline — do not finalize until one eligible Pipeline has completed.

The requirement applies once across the eligible set, not once per attached Pipeline. Discovery-only chunks are orientation evidence and cannot support a grounded final answer under any of these policies. See Discovery chunks and answer-bearing evidence.

Budget ceilings​

Budgets are maximum permitted work, not quotas. The Orchestrator should stop as soon as the admitted evidence satisfies the request, even when budget remains.

The primary controls cover Orchestrator rounds, total actions, Pipeline attempts, admitted evidence, total model tokens, and total latency. Advanced controls bound individual model phases, retrieval operations, direct tools, source expansion, reranking, and finalization reserves.

Save and Apply​

Authoring state and runtime state are deliberately separate:

  1. Edit the LLM, prompt bindings, or mode policies.
  2. Save validates and persists a new editable revision.
  3. Resolve any validation or readiness issues.
  4. Apply a ready saved revision.
  5. CoreCube resolves every dependency and stores an immutable, hashed active snapshot.

Saving does not change live answer behavior. Applying does.

Each answer turn freezes the active policy envelope when the turn begins. Applying another revision affects subsequent turns only; it cannot change a turn already in progress.

The version selector distinguishes:

  • Saved revision — the latest authoring state;
  • Active revision — the immutable snapshot currently serving requests;
  • older saved snapshots — revisions that can be inspected and, when still compatible, applied.

Editing a Pipeline, tool, LLM Model, or shared System Prompt after Apply does not mutate the active snapshot. The Orchestrator reports the dependency change so an administrator can save and apply a new revision intentionally.

Validation and readiness​

These statuses answer different questions:

StatusQuestion answered
ValidationIs the configuration internally coherent and within its declared constraints?
ReadinessAre the selected models, providers, Pipelines, tools, and workers available?

A configuration can be valid but not ready—for example, when a selected reranker is configured correctly but its inference worker is offline. Follow the remediation link shown in the Admin Console, then recheck readiness before applying.

Admin API​

The Admin Console uses these authenticated endpoints:

EndpointPurpose
GET /api/orchestratorRead the saved and active configuration.
GET /api/orchestrator/catalogRead selectable models, prompts, and tools.
PUT /api/orchestratorValidate and save the editable draft.
POST /api/orchestrator/validateValidate without saving.
POST /api/orchestrator/applyApply a saved revision.
GET /api/orchestrator/snapshotsList immutable snapshot history.
GET /api/orchestrator/snapshots/:revision/readinessCheck a historical snapshot's readiness.

Administrative writes require an administrator role.

We use cookies for analytics to improve our website. More information in our Privacy Policy.