For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
gated_delta_conv1d_fwd_gpu
def gated_delta_conv1d_fwd_gpu[work_dtype: DType, state_dtype: DType, KERNEL_SIZE: Int, CONV1D_BLOCK_DIM: Int, WRITE_STATE: Bool, qkv_input_ragged_LT: TensorLayout, conv_weight_LT: TensorLayout, conv_state_LT: TensorLayout, slot_idx_LT: TensorLayout, input_row_offsets_LT: TensorLayout, conv_output_ragged_LT: TensorLayout, Engine: TensorEngine](batch_size: Int32, total_seq_len: Int32, conv_dim: Int32, tokens_per_block: Int32, qkv_input_ragged: TileTensor[work_dtype, qkv_input_ragged_LT, MutUntrackedOrigin, Engine=Engine], conv_weight: TileTensor[work_dtype, conv_weight_LT, MutUntrackedOrigin, Engine=Engine], conv_state: TileTensor[state_dtype, conv_state_LT, MutUntrackedOrigin, Engine=Engine], slot_idx: TileTensor[.uint32, slot_idx_LT, MutUntrackedOrigin, Engine=Engine], input_row_offsets: TileTensor[.uint32, input_row_offsets_LT, MutUntrackedOrigin, Engine=Engine], conv_output_ragged: TileTensor[work_dtype, conv_output_ragged_LT, MutUntrackedOrigin, Engine=Engine], qkv_input_seqlen_stride: UInt32, qkv_input_channel_stride: UInt32, conv_weight_channel_stride: UInt32, conv_weight_offset_stride: UInt32, conv_output_seqlen_stride: UInt32, conv_output_channel_stride: UInt32)
Slot-indexed causal depthwise conv1d over a ragged batch.
The thread owning a sequence's first token computes the outputs that
see the old conv window and then rewrites the window; later tokens read
only the ragged input. With WRITE_STATE False the window is left
unchanged.