For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
coop_group_size
def coop_group_size(ctx: DeviceContext, batch_size: Int, block_size: Int, n_vec: Int) -> Int
Chooses a resident AMD block group for each row.
Every block must be resident because the barrier spins until all blocks arrive. The launch therefore uses at most one block per compute unit. The timeout catches contention that prevents a group from becoming resident.
Args:
- βctx (
DeviceContext): Device the launch targets. - βbatch_size (
Int): Number of rows in the launch. - βblock_size (
Int): Threads per block. - βn_vec (
Int): Vectors in one row.
Returns:
Int: Blocks per row, or 1 when the row should not be split.