IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

topk_topp_masked_probs

def topk_topp_masked_probs[dtype: DType, block_size: Int = Int(1024), TopKArrLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TopPArrLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TemperatureLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], ProbsLayoutType: TensorLayout = Layout[TypeList[Int64, Int64](), TypeList[Int64, ComptimeInt[Int(1)]]()], TopKArrEngine: TensorEngine = DefaultEngine, TopPArrEngine: TensorEngine = DefaultEngine, TemperatureEngine: TensorEngine = DefaultEngine](ctx: DeviceContext, logits: TileTensor[dtype, Engine=logits.Engine, address_space=logits.address_space, linear_idx_type=logits.linear_idx_type], probs: TileTensor[.float32, ProbsLayoutType, MutAnyOrigin], top_k_val: Int, top_p_val: Float32 = 1, top_k_arr: Optional[TileTensor[.int64, TopKArrLayoutType, ImmutAnyOrigin, Engine=TopKArrEngine]] = None, top_p_arr: Optional[TileTensor[.float32, TopPArrLayoutType, ImmutAnyOrigin, Engine=TopPArrEngine]] = None, temperature: Optional[TileTensor[.float32, TemperatureLayoutType, ImmutAnyOrigin, Engine=TemperatureEngine]] = None)

Computes per-row top-k/top-p masked softmax.

Dispatches by device: NVIDIA SM90+ takes the cluster-capable launcher, every other target takes the single-block one. The branch is a comptime one, so a target without thread-block clusters never instantiates the cluster kernels.

See topk_fi.topk_topp_masked_probs for the parameters, arguments and output contract; both sides share them.

Was this page helpful?