size documents. It runs after global merge and deduplication, so it can be used with lexical, vector, sparse vector, or hybrid scoring queries.
Add a rerank object to a query request. This is a per-query choice, independent of collection settings and managed embedding models. Omitting rerank or setting it to null preserves existing search behavior and response fields. LambdaDB manages the provider credentials; you do not supply a Jev API key.
Use a compatible client version when reranking through an SDK, CLI, or MCP server. The REST examples below can be used directly without an SDK.
Request reranking
The examples assume anarticles collection with stored scalar string fields title and body, indexed as text or keyword; body must be a text index for the lexical query shown. Reranking evaluates both fields even though the response projection returns only id and title.
LambdaDB Cloud uses region-specific API base URLs. Use your project’s base URL and project name, together with a project API key created in its API Keys tab. Project creation does not automatically issue a key; save the full value when you create it, because it is shown only once. See API key management. Do not assume a global default URL or a fixed project name.
rerank-query.json. Omitting criteria uses the default ten-level relevance criteria; omitting candidateSize uses max(50, size).
Request parameters
Each candidate must contain nonblank text in at least one selected field. Missing or null individual values contribute no text. Arrays, objects, non-string selected values, or candidates with no nonblank selected text are validation errors; they are not silently dropped or handled by provider fallback.
Only the named fields enter reranker input. Top-level
fields controls the public response projection and can exclude fields used for reranking. Rendered candidate text is limited to 16 KiB UTF-8 per candidate. Query text, all rendered candidates, and custom descriptions together are limited to 256 KiB. Inputs are rejected rather than silently truncated; model context limits also apply.
Candidate counts
These three controls serve different purposes:
For example, a vector-only query with
size: 10, knn.k: 20, and candidateSize: 50 can provide at most 20 vector candidates for reranking. Increase k explicitly if you want deeper vector retrieval; LambdaDB does not rewrite it. In a hybrid query, a vector leg with k: 20 and lexical retrieval can together populate a larger merged pool. Sparse vector and lexical retrieval use the planned candidate depth without a separate k.
The pool can contain fewer than candidateSize documents. candidateCount reports the actual pool size, not a guaranteed count or the final response size.
Query restrictions
Reranking requires a scoring retrieval query. Query-less, match-all requests without a query, and filter-only requests cannot use it. Explicitsort, including an empty list, cannot be combined with reranking.
Facets keep their existing query restrictions: reranking does not enable facets on vector, sparse vector, or hybrid queries. For supported lexical queries, facet counts cover the full matching population and survive reranking or fallback; they are not recomputed from the final documents. size: 0 remains available for facet-only requests without reranking.
Default criteria and scoring
Omittingcriteria or setting it to null selects default-relevance-v1. The model evaluates each candidate’s selected text against queryText using these ten descriptions, ordered from lowest to highest relevance:
LambdaDB computes the final score as a weighted mean of the model’s distribution across these levels:
score = sum(p[i] * v[i]) / sum(p[i]), where p[i] is the model’s probability mass for level i and v[i] is that level’s fixed scoring value. For example, equal mass on levels 8 and 9 produces (0.5 * 0.80 + 0.5 * 0.90) / (0.5 + 0.5) = 0.85. A candidate can therefore receive an intermediate score rather than one of the ten listed values.
The default score range is 0.05–1, within the API’s overall [0,1] domain. LambdaDB does not shift the default minimum to zero, multiply the result by model confidence, or blend it with the retrieval score. This weighted evaluation value is the returned score used for sorting; the original search or fusion score remains in retrievalScore. The model’s level distribution is not included in the query response, and the final score is not a calibrated probability of relevance.
These scoring values are fixed for the default criteria version, not caller-configurable weights. Supplying the same descriptions explicitly as custom criteria selects custom scoring values and can produce different scores.
Use custom criteria
Supply descriptions in order from lowest to highest relevance. Each must be nonblank and distinct, with at most 2 KiB UTF-8 per description and 8 KiB total. Text is preserved exactly; choosing meaningful descriptions and their order is your responsibility. This three-level example replaces the default criteria for this request:i / (N - 1), with i starting at zero) and applies the same weighted-mean formula described above. Two levels use [0, 1]; three use [0, 0.5, 1], so custom scores can reach zero. The default ten-level criteria use their own fixed values, so explicitly supplying the default descriptions as custom criteria does not reproduce the default scoring contract.
There are no caller-supplied weights or threshold options. criteriaVersion: "custom" identifies custom criteria, not a content hash or unique version. Retain the request descriptions when comparing or reproducing evaluations.
Read the response
The following response is illustrative, not a quality or latency guarantee:applied results, each document envelope’s score is the actual final sorting value, in [0,1]. It is an evaluation score, not a calibrated relevance probability. The original retrieval or hybrid-fusion score is retrievalScore, alongside score and doc; it is not placed inside the stored document.
Results are ordered by descending final score. Exact ties preserve original candidate order. Low scores, including numeric zero, do not remove documents. Preserve score precision when storing or displaying results; do not round and then sort again. Different criteria or models can assign different meanings to the same numeric value, and retrieval scores are not directly comparable to reranking scores.
maxScore is the maximum final returned score; it is omitted for empty results. Zero remains a valid numeric score and maximum. total is the number of returned documents, not the corpus match count. Top-level took covers the complete query, including reranking.
Reranking metadata
Without reranking, both top-level
rerank and envelope retrievalScore are omitted. rerankScore and rubricVersion are not response fields.
For custom criteria, zero is a valid result. The custom example can return:
Empty results
With no candidates, the server skips provider calls and returnsskipped/noCandidates with both counts zero. Validation still runs before this shortcut. maxScore, retrievalScore, and criteriaVersion are omitted:
Large results
WhenisDocsInline is false, download the envelopes from docsUrl as described in Large results. Applied results retain both score and retrievalScore in the downloaded envelopes. Reranking metadata stays in the top-level response; an empty inline docs array does not mean the rerank pool was empty.
Handle failures
The defaultonFailure: "error" fails the request on provider-stage errors. Eligible timeouts return HTTP 408; other eligible provider-stage failures return HTTP 503. No partial reranking scores are returned.
Set onFailure: "returnOriginal" to return the first size documents from the retained pre-rerank pool when the provider times out, rate-limits, becomes unavailable, rejects credentials, or returns malformed/incomplete scores. The server keeps original search scores, omits retrievalScore and criteriaVersion, and reports scoredCount: 0. It never mixes partial reranking scores with search scores.
An illustrative timeout fallback:
returnOriginal:
- Invalid input, selected text, criteria, input limits, or disabled model configuration.
- Retrieval, hydration, read-version routing, authorization, or inference quota errors.
- Provider-call admission capacity errors, which return HTTP 429.
- Caller cancellation or an exhausted overall query deadline; fallback does not extend the request deadline.
timeout, rateLimit, unavailable, invalidResponse, or credentials. They expose no provider credentials or request payloads.
Client compatibility and usage
These published releases support optional query-level reranking, final/retrieval score envelopes, stage metadata, and downloaded results:
Update the client or tool itself; older serializers can reject or omit the new object and response metadata. Python requires Python 3.10–3.13, Go requires Go 1.22 or later, and CLI/MCP require Node.js 22.14.0 or later. Migration CLI 0.1.7 does not rerank its verification searches.
Reranking uses server-managed inference and the existing inference quota. Successful stages with complete valid batch usage record aggregate input/output tokens; provider call count is not a billing unit. Failed, fallback, or incomplete-usage stages do not produce a customer reranking token event. See Managed reranking rates for token-to-LIU rates and a cost example.