For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
TopKTopPMaskedProbsClusterKernel
def TopKTopPMaskedProbsClusterKernel[block_size: Int, vec_size: Int, dtype: DType, LogitsLayoutType: TensorLayout, logits_origin: ImmOrigin, cluster_size: Int, LogitsEngine: TensorEngine](logits: TileTensor[dtype, LogitsLayoutType, logits_origin, Engine=LogitsEngine], probs_ptr: Pointer[Float32, MutAnyOrigin], top_k_arr: Optional[Pointer[Int64, ImmutAnyOrigin]], top_k_val: Int32, top_p_arr: Optional[Pointer[Float32, ImmutAnyOrigin]], top_p_val: Float32, temperature: Optional[Pointer[Float32, ImmutAnyOrigin]], d: Int32)
TopKTopPMaskedProbsKernel with one row spread over a cluster.
Works in the unnormalized domain e_i = exp((logit_i - row_max) / temp):
a token survives the joint constraint iff e > cutoff (recovered by the
same dual-pivot search the sampler uses) and its masked probability is
e / kept_mass. The output row is that masked renormalized distribution
-- the same tensor TopKTopPSamplingFromProbKernel emits under
emit_dist, so a verifier's target-side probabilities and a draft's
proposal distribution are described identically.
One block per row uses only as many SMs as there are rows, which leaves most of the GPU idle at decode batch sizes. Here the CTAs of one cluster share the row. Each CTA reduces its own contiguous slice, and the cluster combines the whole-row reductions that the cutoff needs.
The launch must set cluster_dim to cluster_size.