For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
causal_conv1d_varlen_update_gpu
def causal_conv1d_varlen_update_gpu[x_dtype: DType, weight_dtype: DType, bias_dtype: DType, output_dtype: DType, conv_state_dtype: DType, cache_seqlens_dtype: DType, conv_state_indices_dtype: DType, WIDTH: Int, BLOCK_DIM: Int, x_LT: TensorLayout, weight_LT: TensorLayout, bias_LT: TensorLayout, conv_state_LT: TensorLayout, cache_seqlens_LT: TensorLayout, conv_state_indices_LT: TensorLayout, output_LT: TensorLayout, x_engine: TensorEngine, weight_engine: TensorEngine, bias_engine: TensorEngine, conv_state_engine: TensorEngine, cache_seqlens_engine: TensorEngine, conv_state_indices_engine: TensorEngine, output_engine: TensorEngine](batch: Int32, dim: Int32, seqlen: Int32, state_len: Int32, x: TileTensor[x_dtype, x_LT, MutUntrackedOrigin, Engine=x_engine], weight: TileTensor[weight_dtype, weight_LT, MutUntrackedOrigin, Engine=weight_engine], bias: TileTensor[bias_dtype, bias_LT, MutUntrackedOrigin, Engine=bias_engine], conv_state: TileTensor[conv_state_dtype, conv_state_LT, MutUntrackedOrigin, Engine=conv_state_engine], cache_seqlens: TileTensor[cache_seqlens_dtype, cache_seqlens_LT, MutUntrackedOrigin, Engine=cache_seqlens_engine], conv_state_indices: TileTensor[conv_state_indices_dtype, conv_state_indices_LT, MutUntrackedOrigin, Engine=conv_state_indices_engine], output: TileTensor[output_dtype, output_LT, MutUntrackedOrigin, Engine=output_engine], silu_activation: Int8, pad_slot_id: Int32, has_conv_state_indices: Int8, has_cache_seqlens: Int8, has_bias: Int8)
GPU kernel for causal conv1d update (decode step).
Grid: (batch, ceildiv(dim, BLOCK_DIM)) Block: (BLOCK_DIM,)
Note: silu_activation and flag parameters are Int8 (0 or 1) instead of Bool for DevicePassable compatibility on GPU.