For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
apple_gemv_kernel
def apple_gemv_kernel[c_type: DType, in_type: DType, c_layout: TensorLayout, a_layout: TensorLayout, w_layout: TensorLayout, c_engine: TensorEngine, a_engine: TensorEngine, w_engine: TensorEngine, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None], tile_m: Int, rows_per_warp: Int, tile_k: Int, unroll: Int](c: TileTensor[c_type, c_layout, MutAnyOrigin, Engine=c_engine], a: TileTensor[in_type, a_layout, ImmutAnyOrigin, Engine=a_engine], weight: TileTensor[in_type, w_layout, ImmutAnyOrigin, Engine=w_engine], m_arg: Int32, n_arg: Int32, k_arg: Int32)
Computes rows_per_warp columns of c[:m] = a[:m] @ weight^T per warp.
c is [M, N], a is [M, K] and weight is [N, K], all row-major,
with M <= tile_m. K must be a multiple of tile_k, so every chunk load
is aligned to its width. Accumulation is fp32.