For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo module
latent_sparse_attention
Sparse attention over a shared K=V latent held in two paged leaves.
DeepSeek-V4's compressed sparse attention attends from q[t, h, :] to one
shared latent row per key -- the same row serves as key and value for every
head -- and the keys a query sees come from two caches:
- the last
windowtoken positions, kept in a sliding-window leaf paged by token position, and - a per-query list of compressed entries (the indexer's top-k, or every
closed window on ratio-128 layers), kept in a leaf whose page holds
slots_per_pageentries and is addressed by entry index. Because the leaf'spage_sizeparameter is read off the block shape,loadon that leaf already pages by entry, so the entry index is passed straight through as the token index.
attn_sink[h] enters the softmax denominator only: the running max is
taken over the gathered scores alone, then den += exp(sink - max). There
is no value row for the sink.
The batch is ragged: query row t belongs to the sequence b with
input_row_offsets[b] <= t < input_row_offsets[b + 1] and sits at absolute
position pos = cache_lengths[b] + t - input_row_offsets[b] in the window
leaf. Its window keys are positions max(0, pos - window + 1) .. pos, all
of which must already be stored -- the store ops run before this op in the
same graph, as they do for every other paged attention. Compressed keys are
comp_indices[t, k] >= 0; -1 is skipped.
One block per (query row, group of heads); each warp owns heads_per_warp
heads and reads every key row once as a lane-strided vector, so a key costs
one vector load per warp plus a warp reduction per head. This is the
portable form of the kernel; it does not use tensor cores.
Functions
-
latent_sparse_attention_ragged_paged: Attends every query row to its window keys and its listed compressed entries.