For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
enqueue_apple_gemv
def enqueue_apple_gemv[c_type: DType, in_type: DType, //, *, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None](c: TileTensor[c_type, Engine=c.Engine, address_space=c.address_space, linear_idx_type=c.linear_idx_type], a: TileTensor[in_type, Engine=a.Engine, address_space=a.address_space, linear_idx_type=a.linear_idx_type], weight: TileTensor[in_type, Engine=weight.Engine, address_space=weight.address_space, linear_idx_type=weight.linear_idx_type], ctx: DeviceContext)
Enqueues c = a @ weight^T for 1 <= M <= APPLE_GEMV_MAX_M.
Picks the smallest tile_m covering M, one weight row per warp, with the
K unroll measured best for that tile_m.
Parameters:
- c_type (
DType): Output element type. Accumulation is fp32. - in_type (
DType): Activation and weight element type (bf16 or fp16). - elementwise_lambda_fn (
Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None]): Optional epilogue applied to each output.
Args:
- c (
TileTensor[c_type, Engine=c.Engine, address_space=c.address_space, linear_idx_type=c.linear_idx_type]): Output[M, N]. - a (
TileTensor[in_type, Engine=a.Engine, address_space=a.address_space, linear_idx_type=a.linear_idx_type]): Activation[M, K]withK % 8 == 0. - weight (
TileTensor[in_type, Engine=weight.Engine, address_space=weight.address_space, linear_idx_type=weight.linear_idx_type]): Weight[N, K]. - ctx (
DeviceContext): Device context to enqueue on.