IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

causal_conv1d_channel_first_fwd_gpu_no_bias

def causal_conv1d_channel_first_fwd_gpu_no_bias[x_dtype: DType, weight_dtype: DType, output_dtype: DType, kNThreads: Int, kWidth: Int, kNElts: Int, x_LT: TensorLayout, weight_LT: TensorLayout, output_LT: TensorLayout, x_engine: TensorEngine, weight_engine: TensorEngine, output_engine: TensorEngine](batch: Int32, dim: Int32, seqlen: Int32, width: Int32, x: TileTensor[x_dtype, x_LT, MutUntrackedOrigin, Engine=x_engine], weight: TileTensor[weight_dtype, weight_LT, MutUntrackedOrigin, Engine=weight_engine], output: TileTensor[output_dtype, output_LT, MutUntrackedOrigin, Engine=output_engine], x_batch_stride: UInt32, x_c_stride: UInt32, x_l_stride: UInt32, weight_c_stride: UInt32, weight_width_stride: UInt32, out_batch_stride: UInt32, out_c_stride: UInt32, out_l_stride: UInt32, silu_activation: Int8)

Optimized causal conv1d implementation for channel first data layout using SIMD operations (no bias).

Key optimizations:

  1. SIMD vectorization for input/output operations
  2. Efficient memory access patterns with coalesced loads
  3. Vectorized weight loading and computation
  4. Optimized activation function with SIMD operations
  5. Better thread utilization and memory bandwidth usage

Grid: (ceildiv(seqlen, kNThreads * kNElts), dim, batch) Block: kNThreads

Parameters:

  • ​x_dtype (DType): Element type of the input tensor x.
  • ​weight_dtype (DType): Element type of the weight tensor weight.
  • ​output_dtype (DType): Element type of the output tensor output.
  • ​kNThreads (Int): Number of threads per block used to process the sequence dimension.
  • ​kWidth (Int): Compile-time convolution kernel width; must match the runtime width argument.
  • ​kNElts (Int): Number of sequence elements each thread processes, used for SIMD vectorization and ILP.
  • ​x_LT (TensorLayout): TensorLayout of the input tensor x.
  • ​weight_LT (TensorLayout): TensorLayout of the weight tensor weight.
  • ​output_LT (TensorLayout): TensorLayout of the output tensor output.
  • ​x_engine (TensorEngine): Engine of the input tensor x.
  • ​weight_engine (TensorEngine): Engine of the weight tensor weight.
  • ​output_engine (TensorEngine): Engine of the output tensor output.

Args:

Was this page helpful?