For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo module
mla_index_kpool
K-pool compression for the DSA indexer.
A k-pooled indexer stores one candidate key per kpool consecutive tokens
instead of one per token. A pooled key is a weighted average of its members,
where the weights come from a softmax over a gate score plus a learned
within-pool position embedding:
logits[m, c] = gate[member m, c] + ape[m, c]
weights[:, c] = softmax over m # independently per channel c
pooled[p, c] = sum_m weights[m, c] * k[member m, c]The softmax runs per channel, not per member. One weight per member would be a different function.
Pool p covers absolute positions [p * kpool, (p + 1) * kpool) of one
request. Only pools whose members all arrive in the same call are written.
Functions
-
kpool_compress_kernel: Builds one pooled key per block; one thread per channel. -
kpool_expand_topk_kernel: Turns selected pool ids back into the token positions they cover. -
kpool_seed_tail_kernel: Stashes a prefill chunk's trailing tokens into the tail ring. -
kpool_tail_update_kernel: Stashes a request's new tokens, and closes each pool as it fills. -
pool_channel: One pooled channel: softmax overlogits, weighted sum ofvals.