Orchestrator
The Orchestrator is CoreCube's query-time policy authority. It selects how the configured answer model evaluates evidence, which retrieval Pipelines and direct evidence tools it may use, when it must verify a result, and when it may finalize an answer.
The Orchestrator is not a general workflow or action runtime. Its tools are governed, read-only evidence operations, and every accepted result passes through CoreCube's authorization and evidence admission controls.
Open Admin Console → Configuration → Orchestrator.
What the Orchestrator owns
The Orchestrator configuration contains:
- one LLM for query resolution, orchestration decisions, and final answer generation;
- ordered System Prompts for the Orchestrator and Answer phases;
- separate policies for the Fast, Balanced, and Accurate Search modes;
- the Pipelines and direct evidence tools available to each mode;
- stopping rules and hard ceilings for work, evidence, tokens, and latency.
Pipelines remain reusable retrieval definitions. A Pipeline does not choose the answer model or make itself active. Applying an Orchestrator revision embeds the selected Pipeline definitions, tools, prompts, and model binding into one immutable runtime snapshot.
LLM
The LLM tab selects one ready model from Query Pipeline → LLM Models. That model performs both Orchestrator decisions and final answer generation. Its configured temperature also applies to query resolution.
The model must be active, reachable, and capable of handling the context required by every enabled mode. CoreCube reports insufficient context capacity as a readiness problem before Apply.
Changing which LLM Model is marked as the registry default does not replace the model stored in an already-applied Orchestrator snapshot. Select the intended model here, save the revision, and apply it.
System Prompts
The System Prompts tab maintains two independent ordered sequences:
| Phase | Controls |
|---|---|
| Orchestrator | Evidence sufficiency, Pipeline selection, direct-tool use, and stopping. |
| Answer | The final response generated from the frozen, admitted evidence. |
Prompt content is reusable and shared. Its enabled state, phase, and position belong to the current Orchestrator revision. Save or discard an edited System Prompt before navigating away or saving the Orchestrator configuration.
Saved prompt changes do not mutate the active runtime snapshot. Apply a new Orchestrator revision to put the changed content into service.
See System Prompts for fragment types and other binding surfaces.
Modes
The Modes tab configures Fast, Balanced, and Accurate independently. Clients select a mode with
the searchMode field on /v1/chat/completions; Balanced is used when the field is omitted.
Each mode has four components:
| Component | Purpose |
|---|---|
| Eligible pipelines | Ordered retrieval strategies the Orchestrator may run. |
| Direct evidence tools | Focused, governed evidence operations available without running a Pipeline. |
| Decision rules | Verification, no-progress, source-count, and finalization behavior. |
| Budget ceilings | Maximum permitted actions, evidence, model tokens, and elapsed time. |
Attaching a Pipeline makes the run_pipeline function available to that mode. The Orchestrator
chooses a Pipeline from the eligible set using its declared purpose and server-derived capabilities.
Adding several Pipelines provides choices; it does not require every Pipeline to run.
Seeded mode policies
Fresh installations start with these bindings:
| Mode | Pipeline | Pipeline execution | Stopping behavior |
|---|---|---|---|
| Fast | Exact Lookup | Adaptive | First sufficient |
| Balanced | Hybrid Coverage | Prefer one | Verify if uncertain |
| Accurate | Deep Document Analysis | Require one | Verify once |
The Pipeline execution setting means:
- Adaptive — no Pipeline preference is imposed; a Pipeline or direct evidence tool can provide answer-bearing evidence after Discovery.
- Prefer one Pipeline — normally run one eligible Pipeline before finalizing.
- Require one Pipeline — do not finalize until one eligible Pipeline has completed.
The requirement applies once across the eligible set, not once per attached Pipeline. Discovery-only chunks are orientation evidence and cannot support a grounded final answer under any of these policies. See Discovery chunks and answer-bearing evidence.
Budget ceilings
Budgets are maximum permitted work, not quotas. The Orchestrator should stop as soon as the admitted evidence satisfies the request, even when budget remains.
The primary controls cover Orchestrator rounds, total actions, Pipeline attempts, admitted evidence, total model tokens, and total latency. Advanced controls bound individual model phases, retrieval operations, direct tools, source expansion, reranking, and finalization reserves.
Save and Apply
Authoring state and runtime state are deliberately separate:
- Edit the LLM, prompt bindings, or mode policies.
- Save validates and persists a new editable revision.
- Resolve any validation or readiness issues.
- Apply a ready saved revision.
- CoreCube resolves every dependency and stores an immutable, hashed active snapshot.
Saving does not change live answer behavior. Applying does.
Each answer turn freezes the active policy envelope when the turn begins. Applying another revision affects subsequent turns only; it cannot change a turn already in progress.
The version selector distinguishes:
- Saved revision — the latest authoring state;
- Active revision — the immutable snapshot currently serving requests;
- older saved snapshots — revisions that can be inspected and, when still compatible, applied.
Editing a Pipeline, tool, LLM Model, or shared System Prompt after Apply does not mutate the active snapshot. The Orchestrator reports the dependency change so an administrator can save and apply a new revision intentionally.
Validation and readiness
These statuses answer different questions:
| Status | Question answered |
|---|---|
| Validation | Is the configuration internally coherent and within its declared constraints? |
| Readiness | Are the selected models, providers, Pipelines, tools, and workers available? |
A configuration can be valid but not ready—for example, when a selected reranker is configured correctly but its inference worker is offline. Follow the remediation link shown in the Admin Console, then recheck readiness before applying.
Admin API
The Admin Console uses these authenticated endpoints:
| Endpoint | Purpose |
|---|---|
GET /api/orchestrator | Read the saved and active configuration. |
GET /api/orchestrator/catalog | Read selectable models, prompts, and tools. |
PUT /api/orchestrator | Validate and save the editable draft. |
POST /api/orchestrator/validate | Validate without saving. |
POST /api/orchestrator/apply | Apply a saved revision. |
GET /api/orchestrator/snapshots | List immutable snapshot history. |
GET /api/orchestrator/snapshots/:revision/readiness | Check a historical snapshot's readiness. |
Administrative writes require an administrator role.