For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
apple_gemv_small_batch_supported
def apple_gemv_small_batch_supported(m: Int, k: Int) -> Bool
Returns whether _matmul_gpu routes an [m, k] activation here.
Measured on M5 Max with bench_matmul.mojo (bench_gemv_apple.yaml)
against the tiled AppleM5MatMul: up to 3x faster at decode shapes with
K >= 1024 and 2 <= M <= 8. At K = 512 it wins for M <= 4 (1.13-1.50x) and
loses at M = 8 (0.89-0.96x); at K = 256 it loses from M = 4. At M == 1 it
is 0.84-1.10x the split-K GEMV, which already streams weights above 20 MB
at 535-580 GB/s, so M == 1 stays there.
Returns: