IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

apple_gemv_small_batch_supported

def apple_gemv_small_batch_supported(m: Int, k: Int) -> Bool

Returns whether _matmul_gpu routes an [m, k] activation here.

Measured on M5 Max with bench_matmul.mojo (bench_gemv_apple.yaml) against the tiled AppleM5MatMul: up to 3x faster at decode shapes with K >= 1024 and 2 <= M <= 8. At K = 512 it wins for M <= 4 (1.13-1.50x) and loses at M = 8 (0.89-0.96x); at K = 256 it loses from M = 4. At M == 1 it is 0.84-1.10x the split-K GEMV, which already streams weights above 20 MB at 535-580 GB/s, so M == 1 stays there.

Returns:

Bool

Was this page helpful?