Skip to main content

Retrieval Settings

Retrieval settings belong to a retrieval tool attached to a Pipeline. Open the Pipeline, select a Retrieval node, and edit its settings in the node panel.

The settings affect the next Orchestrator snapshot that includes the Pipeline. Saving the tool or Pipeline does not alter an already-active snapshot; save and apply a new Orchestrator revision to put the changes into service.

Search and fusion​

Top K​

Maximum number of candidate chunks returned by the hybrid search stage before reranking. A wider pool can improve recall but increases retrieval and reranking work.

Allowed range: 1–200.

Min Score​

Minimum fused score a candidate must meet. Lower values preserve recall; higher values reject more weak matches before evidence admission.

Allowed range: 0–1.

RRF K​

The Reciprocal Rank Fusion constant. RRF combines the positions from the dense and sparse result lists without assuming their raw scores use the same scale. Lower values give the first few ranks more influence; higher values flatten the difference between nearby ranks.

Allowed range: 1–1000.

Vector Weight​

Weight applied to the dense, embedding-based result list during fusion. Increase it when semantic matches should contribute more strongly.

Allowed range: 0–10.

FTS Weight​

Weight applied to the sparse PostgreSQL full-text result list during fusion. Increase it when exact terms, identifiers, and quoted phrases should contribute more strongly.

Allowed range: 0–10.

Setting either stream's weight to 0 removes that stream's contribution. The seeded Exact Lookup tool uses FTS only; the Hybrid Coverage and High-recall tools use both streams.

Recency Penalty​

Largest ranking demotion that freshness may apply to old evidence. A value below 1 keeps old but otherwise relevant material retrievable. This is a ranking signal, not a deletion or authorization rule.

Allowed range: 0–0.999999.

See Freshness for the scoring model and evidence-date rules.

Number of neighbors the pgvector HNSW index explores at query time. Higher values can improve dense recall at the cost of latency. This setting does not rebuild the index.

Allowed range: 1–2000.

Candidate-pool limits​

Dense Top K​

Maximum candidates requested from vector search before fusion.

Allowed range: 1–500.

Sparse Top K​

Maximum candidates requested from PostgreSQL full-text search before fusion.

Allowed range: 1–500.

Rerank Candidates​

Maximum fused candidates sent to the reranker. Increasing this can improve precision when the useful chunk starts low in the fused order, but it increases reranker work.

Allowed range: 0–500.

Final Top K​

Maximum chunks returned after reranking or, when reranking is disabled or degraded, after fused ordering.

Allowed range: 1–33. Keep this at or below Rerank candidates when reranking is enabled.

Evidence Budget​

Maximum retrieved-evidence tokens across the complete answer turn, including later Pipeline or direct-tool actions. The Orchestrator's mode-level evidence budget can impose a lower effective ceiling.

Allowed range: 512–60000 tokens.

Reranking​

Reranker Model​

Cross-encoder model used to score each query and candidate together. Select a model from the Reranker catalog. Applying an Orchestrator snapshot checks that the selected model and its inference worker are ready.

Reranker Enabled​

When enabled, CoreCube reranks the fused candidate pool before final selection. When disabled, it uses fused retrieval order directly.

If a reranker becomes unavailable during a request, CoreCube can degrade to fused order and records the outcome in extended response metadata.

Reranker Doc Chars​

Maximum characters from each candidate passed to the reranker. Longer candidate text is truncated.

Allowed range: 100–20000 characters.

Reranker Tokens​

Token window scored per candidate. This is an important latency and memory control for local inference.

Allowed range: 64–8192 tokens.

Reranker Timeout​

Maximum time for one reranker call before CoreCube falls back to fused order.

Allowed range: 1000–120000 milliseconds.

Seeded retrieval tools​

The three fresh-install Pipelines use these retrieval tools:

PipelineRetrieval toolShape
Exact Lookupretrieve-lexical-focusedFTS-focused, small pool, reranking disabled.
Hybrid Coverageretrieve-hybrid-coverageBalanced dense and sparse pools with reranking.
Deep Document Analysisretrieve-high-recallWider pools, deeper reranking, and a larger final result set.

These are editable definitions, not global defaults. Duplicating a Pipeline lets you tune a variant without changing the seeded Pipeline used by an existing saved Orchestrator draft.

Setting reference​

Setting keyTypeAllowed range
retrieval_top_kinteger1–200
retrieval_min_scorenumber0–1
retrieval_rrf_kinteger1–1000
retrieval_vector_weightnumber0–10
retrieval_fts_weightnumber0–10
retrieval_recency_max_penaltynumber0–0.999999
retrieval_hnsw_ef_searchinteger1–2000
retrieval_dense_top_kinteger1–500
retrieval_sparse_top_kinteger1–500
retrieval_rerank_candidatesinteger0–500
retrieval_final_top_kinteger1–33
retrieval_evidence_budget_tokensinteger512–60000
reranker_modelmodel identifierready catalog model
reranker_enabled0 or 1—
reranker_max_document_charsinteger100–20000
reranker_max_tokensinteger64–8192
reranker_timeout_msinteger1000–120000 ms

The server rejects invalid retrieval-tool values at the validation boundary. Use PATCH /api/tools/:id to update a tool definition; PATCH /api/pipelines/:id updates Pipeline identity, purpose, execution limits, and Pipeline-managed metadata.

We use cookies for analytics to improve our website. More information in our Privacy Policy.