IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo module

latent_sparse_attention

Sparse attention over a shared K=V latent held in two paged leaves.

DeepSeek-V4's compressed sparse attention attends from q[t, h, :] to one shared latent row per key -- the same row serves as key and value for every head -- and the keys a query sees come from two caches:

  • the last window token positions, kept in a sliding-window leaf paged by token position, and
  • a per-query list of compressed entries (the indexer's top-k, or every closed window on ratio-128 layers), kept in a leaf whose page holds slots_per_page entries and is addressed by entry index. Because the leaf's page_size parameter is read off the block shape, load on that leaf already pages by entry, so the entry index is passed straight through as the token index.

attn_sink[h] enters the softmax denominator only: the running max is taken over the gathered scores alone, then den += exp(sink - max). There is no value row for the sink.

The batch is ragged: query row t belongs to the sequence b with input_row_offsets[b] <= t < input_row_offsets[b + 1] and sits at absolute position pos = cache_lengths[b] + t - input_row_offsets[b] in the window leaf. Its window keys are positions max(0, pos - window + 1) .. pos, all of which must already be stored -- the store ops run before this op in the same graph, as they do for every other paged attention. Compressed keys are comp_indices[t, k] >= 0; -1 is skipped.

One block per (query row, group of heads); each warp owns heads_per_warp heads and reads every key row once as a lane-strided vector, so a key costs one vector load per warp plus a warp reduction per head. This is the portable form of the kernel; it does not use tensor cores.

Functions​

Was this page helpful?