IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo module

topk_fi_cluster

Cluster-launched top-k/top-p kernels.

One block per row uses only as many SMs as there are rows, which leaves most of the GPU idle at decode batch sizes. The kernels here spread each row over the CTAs of one thread-block cluster and combine the per-row reductions over distributed shared memory. The launchers fall back to the single-block kernels in topk_fi when a CTA's slice does not fit in shared memory or when the target has no clusters.

Functions​

Was this page helpful?