IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo module

indexer_score

Lightning-indexer scores read straight out of a paged compressed leaf.

DeepSeek-V4's indexer scores every query row t against the compressed entries it may see:

score[t, c] = sum_h relu(q[t, h, :] . k[e(c), :]) * weights[t, h]

The candidate axis has 2 * cap columns in the model's order: columns 0 .. cap - 1 are the request's entries closed before this chunk (entry c, live while c < base[b]), columns cap .. 2 * cap - 1 its fresh windows (entry base[b] + c - cap, live while that entry is below the query's cutoff[t]). Every live entry is already stored in the leaf -- the fresh ones by the store op that runs before this one in the same graph -- so the per-token candidate table is never built. Dead columns are written as 0; callers mask them before ranking.

The sum runs over the heads this op is given. Under tensor parallelism each device passes its share of the heads and the partial scores are all-reduced before the top-k.

GPU: one block per (query row, block_size columns), one thread per column. The row's query and weights sit in shared memory; each thread streams its entry once and keeps one accumulator per head. Blocks with no live column only write zeros. The CPU path is the same arithmetic in the same order.

Functions​

Was this page helpful?