For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo module
indexer_score
Lightning-indexer scores read straight out of a paged compressed leaf.
DeepSeek-V4's indexer scores every query row t against the compressed
entries it may see:
score[t, c] = sum_h relu(q[t, h, :] . k[e(c), :]) * weights[t, h]The candidate axis has 2 * cap columns in the model's order: columns
0 .. cap - 1 are the request's entries closed before this chunk (entry
c, live while c < base[b]), columns cap .. 2 * cap - 1 its fresh
windows (entry base[b] + c - cap, live while that entry is below the
query's cutoff[t]). Every live entry is already stored in the leaf -- the
fresh ones by the store op that runs before this one in the same graph -- so
the per-token candidate table is never built. Dead columns are written as 0;
callers mask them before ranking.
The sum runs over the heads this op is given. Under tensor parallelism each device passes its share of the heads and the partial scores are all-reduced before the top-k.
GPU: one block per (query row, block_size columns), one thread per
column. The row's query and weights sit in shared memory; each thread streams
its entry once and keeps one accumulator per head. Blocks with no live column
only write zeros. The CPU path is the same arithmetic in the same order.
Functions
-
indexer_score_ragged_paged: Scores every query row against its live compressed entries.