IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

apple_gemv_kernel

def apple_gemv_kernel[c_type: DType, in_type: DType, c_layout: TensorLayout, a_layout: TensorLayout, w_layout: TensorLayout, c_engine: TensorEngine, a_engine: TensorEngine, w_engine: TensorEngine, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None], tile_m: Int, rows_per_warp: Int, tile_k: Int, unroll: Int](c: TileTensor[c_type, c_layout, MutAnyOrigin, Engine=c_engine], a: TileTensor[in_type, a_layout, ImmutAnyOrigin, Engine=a_engine], weight: TileTensor[in_type, w_layout, ImmutAnyOrigin, Engine=w_engine], m_arg: Int32, n_arg: Int32, k_arg: Int32)

Computes rows_per_warp columns of c[:m] = a[:m] @ weight^T per warp.

c is [M, N], a is [M, K] and weight is [N, K], all row-major, with M <= tile_m. K must be a multiple of tile_k, so every chunk load is aligned to its width. Accumulation is fp32.

Was this page helpful?