IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

Struct_latent_sparse_attention_ragged_paged

struct Struct_latent_sparse_attention_ragged_paged

Registers the mo.latent_sparse_attention.ragged.paged graph op with the graph compiler.

Sparse attention over a shared K=V latent read from two paged leaves: the sliding-window leaf by position range and the compressed leaf by an explicit per-query entry list, with a per-head attention sink in the denominator. See nn.attention.latent_sparse_attention.

Implemented traitsโ€‹

AnyType, Deinitable, Movable

Methodsโ€‹

executeโ€‹

static def execute[q_type: DType, swa_type: DType, comp_type: DType, //, window: Int, target: StringSpan[ImmStaticOrigin]](output: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=output.static_spec], q: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=q.static_spec], input_row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=input_row_offsets.static_spec], comp_indices: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=comp_indices.static_spec], attn_sink: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=attn_sink.static_spec], swa_kv_blocks: ManagedTensorSlice[IOSpec[_, _].MutableInput, static_spec=swa_kv_blocks.static_spec], swa_page_stride: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=swa_page_stride.static_spec], swa_cache_lengths: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=swa_cache_lengths.static_spec], swa_kv_lookup_table: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=swa_kv_lookup_table.static_spec], swa_max_prompt_length: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=swa_max_prompt_length.static_spec], swa_max_cache_length: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=swa_max_cache_length.static_spec], comp_kv_blocks: ManagedTensorSlice[IOSpec[_, _].MutableInput, static_spec=comp_kv_blocks.static_spec], comp_page_stride: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=comp_page_stride.static_spec], comp_cache_lengths: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=comp_cache_lengths.static_spec], comp_kv_lookup_table: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=comp_kv_lookup_table.static_spec], comp_max_prompt_length: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=comp_max_prompt_length.static_spec], comp_max_cache_length: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=comp_max_cache_length.static_spec], layer_swa: UInt32, layer_comp: UInt32, scale: Float32, context: DeviceContext)

Was this page helpful?