IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

DecodeKVConsumer

struct DecodeKVConsumer[dtype: DType, config: MLA_SM100_Decode_Config, num_producer: Int = Int(1), num_consumer: Int = Int(2), num_stages: Int = config.num_kv_stages]

Consumer side of the decode KV pipeline that waits for and releases KV stages.

Parameters

  • dtype (DType): Element type of the KV tiles stored in SMEM.
  • config (MLA_SM100_Decode_Config): Decode config supplying num_kv_stages, BN_QK, and q_depth for pipeline stage sizing.
  • num_producer (Int): Number of producer threads arriving on each producer mbarrier (defaults to 1).
  • num_consumer (Int): Number of consumer threads arriving on each consumer mbarrier (defaults to 2, matching the standard mmaQK+mmaPV dual-consumer KV pipeline).
  • num_stages (Int): Ring depth (defaults to config.num_kv_stages). Set explicitly when a kernel stages KV through two rings of different depths.

Fields

  • pipe (DecodeKVConsumer[dtype, config, num_producer, num_consumer, num_stages].KVPipeType):
  • smem (Pointer[Scalar[dtype], MutAnyOrigin, address_space=AddressSpace.SHARED]):

Implemented traits

AnyType, Copyable, Deinitable, ImplicitlyCopyable, Movable, RegisterPassable, TrivialRegisterPassable

comptime members

kv_stage_elems

comptime kv_stage_elems = (config * config)

KVPipeType

comptime KVPipeType = KVPipelineGeneric[num_stages, Int(1), num_producer, num_consumer]

Methods

__init__

def __init__(pipe: KVPipelineGeneric[num_stages, Int(1), num_producer, num_consumer], smem: Pointer[Scalar[dtype], MutAnyOrigin, address_space=AddressSpace.SHARED]) -> Self

stage_base_ptr

def stage_base_ptr[*, qk_stage: Int = Int(0)](self) -> Pointer[Scalar[dtype], MutAnyOrigin, address_space=AddressSpace.SHARED]

Returns:

Pointer[Scalar[dtype], MutAnyOrigin, address_space=AddressSpace.SHARED]

stage_index

def stage_index[*, qk_stage: Int = Int(0)](self) -> UInt32

Returns:

UInt32

wait

def wait[*, qk_stage: Int = Int(0)](self)

release

def release[*, qk_stage: Int = Int(0)](mut self, e: Int32)

release_all

def release_all(mut self)

Explicit-arrive release for a non-MMA (independent-thread) consumer.

Every thread of the num_consumer-wide consumer role (e.g. a full warpgroup doing a manual SMEM re-swizzle read) calls this once; the num_consumer independent arrive() calls satisfy the mbar's expected count. Use this instead of release[qk_stage](e) (which uses elect_mma_arrive for a single elected thread per warp) when every thread independently participates, not just one MMA-eligible lane per warp.