Skip to main content
Managed reranking evaluates retrieved candidates against a query and orders them by the resulting score before returning the final size documents. It runs after global merge and deduplication, so it can be used with lexical, vector, sparse vector, or hybrid scoring queries. Add a rerank object to a query request. This is a per-query choice, independent of collection settings and managed embedding models. Omitting rerank or setting it to null preserves existing search behavior and response fields. LambdaDB manages the provider credentials; you do not supply a Jev API key. Use a compatible client version when reranking through an SDK, CLI, or MCP server. The REST examples below can be used directly without an SDK.

Request reranking

The examples assume an articles collection with stored scalar string fields title and body, indexed as text or keyword; body must be a text index for the lexical query shown. Reranking evaluates both fields even though the response projection returns only id and title.
LambdaDB Cloud uses region-specific API base URLs. Use your project’s base URL and project name, together with a project API key created in its API Keys tab. Project creation does not automatically issue a key; save the full value when you create it, because it is shown only once. See API key management. Do not assume a global default URL or a fixed project name.
Replace the placeholders with your project’s connection details and export them in your shell:
Save this body as rerank-query.json. Omitting criteria uses the default ten-level relevance criteria; omitting candidateSize uses max(50, size).
Send it using your LambdaDB project credentials:

Request parameters

Each candidate must contain nonblank text in at least one selected field. Missing or null individual values contribute no text. Arrays, objects, non-string selected values, or candidates with no nonblank selected text are validation errors; they are not silently dropped or handled by provider fallback. Only the named fields enter reranker input. Top-level fields controls the public response projection and can exclude fields used for reranking. Rendered candidate text is limited to 16 KiB UTF-8 per candidate. Query text, all rendered candidates, and custom descriptions together are limited to 256 KiB. Inputs are rejected rather than silently truncated; model context limits also apply.

Candidate counts

These three controls serve different purposes: For example, a vector-only query with size: 10, knn.k: 20, and candidateSize: 50 can provide at most 20 vector candidates for reranking. Increase k explicitly if you want deeper vector retrieval; LambdaDB does not rewrite it. In a hybrid query, a vector leg with k: 20 and lexical retrieval can together populate a larger merged pool. Sparse vector and lexical retrieval use the planned candidate depth without a separate k. The pool can contain fewer than candidateSize documents. candidateCount reports the actual pool size, not a guaranteed count or the final response size.

Query restrictions

Reranking requires a scoring retrieval query. Query-less, match-all requests without a query, and filter-only requests cannot use it. Explicit sort, including an empty list, cannot be combined with reranking. Facets keep their existing query restrictions: reranking does not enable facets on vector, sparse vector, or hybrid queries. For supported lexical queries, facet counts cover the full matching population and survive reranking or fallback; they are not recomputed from the final documents. size: 0 remains available for facet-only requests without reranking.

Default criteria and scoring

Omitting criteria or setting it to null selects default-relevance-v1. The model evaluates each candidate’s selected text against queryText using these ten descriptions, ordered from lowest to highest relevance: LambdaDB computes the final score as a weighted mean of the model’s distribution across these levels: score = sum(p[i] * v[i]) / sum(p[i]), where p[i] is the model’s probability mass for level i and v[i] is that level’s fixed scoring value. For example, equal mass on levels 8 and 9 produces (0.5 * 0.80 + 0.5 * 0.90) / (0.5 + 0.5) = 0.85. A candidate can therefore receive an intermediate score rather than one of the ten listed values. The default score range is 0.05–1, within the API’s overall [0,1] domain. LambdaDB does not shift the default minimum to zero, multiply the result by model confidence, or blend it with the retrieval score. This weighted evaluation value is the returned score used for sorting; the original search or fusion score remains in retrievalScore. The model’s level distribution is not included in the query response, and the final score is not a calibrated probability of relevance. These scoring values are fixed for the default criteria version, not caller-configurable weights. Supplying the same descriptions explicitly as custom criteria selects custom scoring values and can produce different scores.

Use custom criteria

Supply descriptions in order from lowest to highest relevance. Each must be nonblank and distinct, with at most 2 KiB UTF-8 per description and 8 KiB total. Text is preserved exactly; choosing meaningful descriptions and their order is your responsibility. This three-level example replaces the default criteria for this request:
For custom criteria with N levels, the server assigns uniformly spaced level values from 0 to 1 (i / (N - 1), with i starting at zero) and applies the same weighted-mean formula described above. Two levels use [0, 1]; three use [0, 0.5, 1], so custom scores can reach zero. The default ten-level criteria use their own fixed values, so explicitly supplying the default descriptions as custom criteria does not reproduce the default scoring contract. There are no caller-supplied weights or threshold options. criteriaVersion: "custom" identifies custom criteria, not a content hash or unique version. Retain the request descriptions when comparing or reproducing evaluations.

Read the response

The following response is illustrative, not a quality or latency guarantee:
On applied results, each document envelope’s score is the actual final sorting value, in [0,1]. It is an evaluation score, not a calibrated relevance probability. The original retrieval or hybrid-fusion score is retrievalScore, alongside score and doc; it is not placed inside the stored document. Results are ordered by descending final score. Exact ties preserve original candidate order. Low scores, including numeric zero, do not remove documents. Preserve score precision when storing or displaying results; do not round and then sort again. Different criteria or models can assign different meanings to the same numeric value, and retrieval scores are not directly comparable to reranking scores. maxScore is the maximum final returned score; it is omitted for empty results. Zero remains a valid numeric score and maximum. total is the number of returned documents, not the corpus match count. Top-level took covers the complete query, including reranking.

Reranking metadata

Without reranking, both top-level rerank and envelope retrievalScore are omitted. rerankScore and rubricVersion are not response fields. For custom criteria, zero is a valid result. The custom example can return:

Empty results

With no candidates, the server skips provider calls and returns skipped/noCandidates with both counts zero. Validation still runs before this shortcut. maxScore, retrievalScore, and criteriaVersion are omitted:

Large results

When isDocsInline is false, download the envelopes from docsUrl as described in Large results. Applied results retain both score and retrievalScore in the downloaded envelopes. Reranking metadata stays in the top-level response; an empty inline docs array does not mean the rerank pool was empty.

Handle failures

The default onFailure: "error" fails the request on provider-stage errors. Eligible timeouts return HTTP 408; other eligible provider-stage failures return HTTP 503. No partial reranking scores are returned. Set onFailure: "returnOriginal" to return the first size documents from the retained pre-rerank pool when the provider times out, rate-limits, becomes unavailable, rejects credentials, or returns malformed/incomplete scores. The server keeps original search scores, omits retrievalScore and criteriaVersion, and reports scoredCount: 0. It never mixes partial reranking scores with search scores. An illustrative timeout fallback:
For hybrid retrieval, fallback preserves the order of the expanded fused pool. It need not match a separate query run without reranking at a smaller candidate depth. These errors do not become fallback results, even with returnOriginal:
  • Invalid input, selected text, criteria, input limits, or disabled model configuration.
  • Retrieval, hydration, read-version routing, authorization, or inference quota errors.
  • Provider-call admission capacity errors, which return HTTP 429.
  • Caller cancellation or an exhausted overall query deadline; fallback does not extend the request deadline.
Fallback reason codes are timeout, rateLimit, unavailable, invalidResponse, or credentials. They expose no provider credentials or request payloads.

Client compatibility and usage

These published releases support optional query-level reranking, final/retrieval score envelopes, stage metadata, and downloaded results: Update the client or tool itself; older serializers can reject or omit the new object and response metadata. Python requires Python 3.10–3.13, Go requires Go 1.22 or later, and CLI/MCP require Node.js 22.14.0 or later. Migration CLI 0.1.7 does not rerank its verification searches. Reranking uses server-managed inference and the existing inference quota. Successful stages with complete valid batch usage record aggregate input/output tokens; provider call count is not a billing unit. Failed, fallback, or incomplete-usage stages do not produce a customer reranking token event. See Managed reranking rates for token-to-LIU rates and a cost example.