For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
causal_conv1d_varlen_fwd_gpu
def causal_conv1d_varlen_fwd_gpu[x_dtype: DType, weight_dtype: DType, bias_dtype: DType, output_dtype: DType, cu_seqlens_dtype: DType, cache_indices_dtype: DType, has_initial_state_dtype: DType, conv_states_dtype: DType, WIDTH: Int, BLOCK_DIM: Int, BLOCK_SEQ: Int, x_LT: TensorLayout, weight_LT: TensorLayout, bias_LT: TensorLayout, query_start_loc_LT: TensorLayout, cache_indices_LT: TensorLayout, has_initial_state_LT: TensorLayout, conv_states_LT: TensorLayout, output_LT: TensorLayout, x_engine: TensorEngine, weight_engine: TensorEngine, bias_engine: TensorEngine, query_start_loc_engine: TensorEngine, cache_indices_engine: TensorEngine, has_initial_state_engine: TensorEngine, conv_states_engine: TensorEngine, output_engine: TensorEngine, use_residual: Bool = False, channels_last: Bool = False](dim: Int32, total_seqlen: Int32, batch: Int32, x: TileTensor[x_dtype, x_LT, MutUntrackedOrigin, Engine=x_engine], weight: TileTensor[weight_dtype, weight_LT, MutUntrackedOrigin, Engine=weight_engine], bias: TileTensor[bias_dtype, bias_LT, MutUntrackedOrigin, Engine=bias_engine], query_start_loc: TileTensor[cu_seqlens_dtype, query_start_loc_LT, MutUntrackedOrigin, Engine=query_start_loc_engine], cache_indices: TileTensor[cache_indices_dtype, cache_indices_LT, MutUntrackedOrigin, Engine=cache_indices_engine], has_initial_state: TileTensor[has_initial_state_dtype, has_initial_state_LT, MutUntrackedOrigin, Engine=has_initial_state_engine], conv_states: TileTensor[conv_states_dtype, conv_states_LT, MutUntrackedOrigin, Engine=conv_states_engine], output: TileTensor[output_dtype, output_LT, MutUntrackedOrigin, Engine=output_engine], silu_activation: Int8, pad_slot_id: Int32, has_cache_indices: Int8, has_initial_state_flag: Int8, has_conv_states: Int8, has_bias: Int8)
GPU kernel for causal conv1d forward with variable length sequences.
Grid: (batch, ceildiv(dim, BLOCK_DIM)) Block: (BLOCK_DIM, BLOCK_SEQ)
Each block processes BLOCK_DIM channels for one sequence.
use_residual adds x[d, s] to the convolution sum before the activation.
Note: silu_activation and flag parameters are Int8 (0 or 1) instead of Bool for DevicePassable compatibility on GPU.