For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo module
topk_fi_cluster
Cluster-launched top-k/top-p kernels.
One block per row uses only as many SMs as there are rows, which leaves most
of the GPU idle at decode batch sizes. The kernels here spread each row over
the CTAs of one thread-block cluster and combine the per-row reductions over
distributed shared memory. The launchers fall back to the single-block kernels
in topk_fi when a CTA's slice does not fit in shared memory or when the
target has no clusters.
Functions
-
topk_topp_masked_probs_cluster: Computes per-row top-k/top-p masked softmax on a cluster device. -
topk_topp_sampling_from_prob_cluster: Joint top-k + top-p sampling from probability distribution. -
TopKTopPMaskedProbsClusterKernel:TopKTopPMaskedProbsKernelwith one row spread over a cluster. -
TopKTopPSamplingEmitDistClusterKernel:TopKTopPSamplingFromProbKernelwithfrom_logitsandemit_dist, spread over a cluster.