Skip to main content

Retrieval Pipeline

Every CoreCube answer turn runs against one immutable Orchestrator snapshot. The snapshot fixes the model, prompts, Search mode policy, eligible Pipelines, direct evidence tools, and budgets for the duration of the turn.

The current query-time flow is:

Discovery always runs before the first Orchestrator decision. After that, each accepted action returns to the same Orchestrator loop until the evidence is sufficient or policy prevents more work.

1. Freeze authorization and policy​

When /v1/chat/completions receives a request, CoreCube resolves:

  • the API key and effective user;
  • allowed connections and scopes;
  • the selected compartment;
  • the effective sensitivity ceiling;
  • the requested searchMode, or Balanced when omitted;
  • the currently active Orchestrator snapshot.

CoreCube freezes these values into a turn policy envelope. A later configuration change cannot expand or alter a request already in progress.

2. Resolve the standalone query​

The configured Orchestrator model resolves conversation references into a standalone question. The original conversation remains available for final answer generation, while retrieval works from the resolved query.

Query-resolution input and output are bounded by the selected mode's token budgets. The same authorization scope applies to every derived query.

3. Run mandatory Discovery​

CoreCube starts with one bounded PostgreSQL full-text search. Discovery is FTS-only and uses the active keyword-language configuration. Its purpose is to obtain an inexpensive first evidence set before asking the model what additional retrieval, if any, would help.

Discovery candidates pass through the same authorization, active chunk-version, evidence-budget, and admission checks as later results. Accepted chunks receive opaque evidence references and enter the turn evidence ledger.

Discovery is not a configured Pipeline. It is the server-owned starting step for every Search mode.

Discovery chunks and answer-bearing evidence​

Discovery chunks are orientation evidence for the Orchestrator. They help it decide which Pipeline or direct evidence tool to run, but a chunk whose only producer is Discovery cannot support or be cited in the final answer.

If a later Pipeline retrieves the same chunk, the evidence ledger can retain one copy of the content and add the Pipeline's producer attribution. The chunk is then answer-bearing because it has Pipeline lineage—not because the Discovery evidence was promoted.

Keeping Discovery-only chunks out of final answers preserves retrieval integrity:

  • a required Pipeline cannot become ceremonial while an empty or unrelated run unlocks an FTS-only answer;
  • Discovery's FTS-only results cannot bypass the Pipeline's query expansion, ranking, filtering, and quality policy;
  • sufficiency checks and audits retain a clear distinction between orientation evidence and evidence that supports the answer.

This is not an additional authorization boundary. Discovery is already restricted to the caller's authorized scope and revalidated during admission; the distinction is about how evidence was retrieved and qualified.

Result limits are independent. Discovery permits up to eight chunks, while the seeded Exact Lookup and Hybrid Coverage Pipelines return at most four and seven respectively. A Pipeline may still return fewer chunks, and equal limits would not guarantee the same documents after query expansion, filtering, ranking, and reranking.

4. Let the Orchestrator choose one action​

The Orchestrator receives the standalone question, cumulative admitted evidence, remaining budgets, and only the functions permitted by the frozen mode policy. It must choose exactly one action:

ActionResult
run_pipelineRuns one eligible Pipeline for a specific unresolved evidence need.
Direct evidence toolRuns one focused, governed operation such as document search or neighbor lookup.
finalizeStops retrieval when the evidence and stopping policy allow it.

The server validates every proposed action. The model cannot invent a Pipeline, tool, identifier, scope, or budget. Pipeline and tool output is treated as untrusted data and must pass evidence admission before it returns to the model.

After an accepted Pipeline or direct-tool action, the Orchestrator reevaluates all material evidence needs against the cumulative ledger. Budgets are ceilings, so it may finalize before consuming all available work.

Search mode behavior​

Fresh installations use these policies:

Search modeSeeded PipelinePipeline policyCurrent behavior
FastExact LookupAdaptiveMay use a Pipeline or direct evidence tool after Discovery.
BalancedHybrid CoveragePrefer onceNormally runs one Pipeline before finalizing.
AccurateDeep Document AnalysisRequire onceMust run one eligible Pipeline before finalizing.

Discovery-only chunks cannot support a grounded final answer in any mode. Fast imposes no Pipeline preference, so a focused direct evidence tool can provide the answer-bearing evidence instead.

Adding multiple eligible Pipelines gives the Orchestrator a choice. The policy applies once across the set; it does not require every attached Pipeline to run.

5. Run a Pipeline​

A Pipeline is a frozen retrieval strategy with four stages:

  1. Query — ordered query tools and prompts produce bounded variants.
  2. Retrieval — the default retrieval tool runs FTS, vector, or hybrid candidate search according to its weights and limits.
  3. Chunk — ordered chunk tools select or process server-issued candidate references.
  4. Evidence — the server validates the output, applies final ranking and evidence limits, and returns admitted evidence to the Orchestrator.

When both retrieval weights are greater than zero, the dense and sparse legs run as a hybrid search:

Setting the vector weight to zero produces FTS-only retrieval. Setting the FTS weight to zero produces vector-only retrieval. See Retrieval settings for the candidate pools, fusion weights, score floor, HNSW, freshness, and reranker controls.

Pipeline execution is bounded by both its local tool-call/wall-time limits and the active Search mode's remaining turn budget.

6. Run a direct evidence tool​

Direct evidence tools are attached per Search mode in the Orchestrator. They handle focused gaps without rerunning a complete Pipeline—for example:

  • search for authorized source documents;
  • read a bounded page from a known source reference;
  • fetch neighboring chunks around admitted evidence;
  • run a bounded multi-query context expansion.

These tools are read-only, LLM-callable, and server-executed. They receive only the authority and opaque references granted by the turn. Their results enter the evidence ledger before the Orchestrator can use them.

7. Freeze evidence and generate the answer​

When finalization is accepted, CoreCube freezes the admitted evidence set. The selected Orchestrator model then generates the answer using:

  • the original conversation;
  • the ordered Answer-phase System Prompts;
  • the frozen evidence and citation map;
  • server-owned safety and answerability constraints.

No retrieval or direct tool is available during final generation. Unsupported claims must be reported as limitations rather than filled with model knowledge.

CoreCube validates citation references and renders authorized source details. With extended mode, the response also includes evidence counts, Search mode, reranker outcome, confidence, degradations, audit status, and latency.

Buffered and streaming requests use the same retrieval decisions and frozen evidence process. Streaming changes only delivery: answer text is sent as SSE chunks, followed by optional extended metadata and [DONE].

Ingestion side​

Before content can be retrieved, CoreCube:

  1. extracts and sanitizes source content;
  2. splits it using the global Heading-aware, Paragraph, or Fixed-size chunking strategy;
  3. builds PostgreSQL full-text vectors;
  4. generates embeddings in the background;
  5. stores chunks, provenance, authorization metadata, and vectors in PostgreSQL.

Chunking and embedding are global ingestion configuration, not Pipeline fields. Existing content is not silently re-chunked when those settings change. See Embedding & Chunking.

Inspect a query​

Use Admin Console → Query Pipeline → Query Explorer to inspect retrieval behavior, including:

  • the selected Search mode and effective scope;
  • candidate and ranking diagnostics;
  • which Pipeline and tools ran;
  • admitted evidence and degradations;
  • final context and answerability information.

We use cookies for analytics to improve our website. More information in our Privacy Policy.