IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

topk_topp_sampling_from_prob

def topk_topp_sampling_from_prob[dtype: DType, out_idx_type: DType, block_size: Int = Int(1024), from_logits: Bool = False, emit_dist: Bool = False, dist_dtype: DType = .float32, DistLayoutType: TensorLayout = Layout[TypeList[Int64, Int64](), TypeList[Int64, ComptimeInt[Int(1)]]()], TopKArrLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], IndicesLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TopPArrLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], SeedLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TemperatureLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], MinPLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TopKArrEngine: TensorEngine = DefaultEngine, IndicesEngine: TensorEngine = DefaultEngine, TopPArrEngine: TensorEngine = DefaultEngine, SeedEngine: TensorEngine = DefaultEngine, TemperatureEngine: TensorEngine = DefaultEngine, MinPEngine: TensorEngine = DefaultEngine](ctx: DeviceContext, probs: TileTensor[dtype, Engine=probs.Engine, address_space=probs.address_space, linear_idx_type=probs.linear_idx_type], output: TileTensor[out_idx_type, Engine=output.Engine, address_space=output.address_space, linear_idx_type=output.linear_idx_type], top_k_val: Int, top_p_val: Float32 = 1, deterministic: Bool = False, rng_seed: Optional[TileTensor[.uint64, SeedLayoutType, ImmutAnyOrigin, Engine=SeedEngine]] = None, rng_offset: UInt64 = UInt64(0), indices: Optional[TileTensor[out_idx_type, IndicesLayoutType, ImmutAnyOrigin, Engine=IndicesEngine]] = None, top_k_arr: Optional[TileTensor[out_idx_type, TopKArrLayoutType, ImmutAnyOrigin, Engine=TopKArrEngine]] = None, top_p_arr: Optional[TileTensor[.float32, TopPArrLayoutType, ImmutAnyOrigin, Engine=TopPArrEngine]] = None, temperature: Optional[TileTensor[.float32, TemperatureLayoutType, ImmutAnyOrigin, Engine=TemperatureEngine]] = None, min_p: Optional[TileTensor[.float32, MinPLayoutType, ImmutAnyOrigin, Engine=MinPEngine]] = None, out_dist: Optional[TileTensor[dist_dtype, DistLayoutType, MutAnyOrigin]] = None)

Joint top-k + top-p sampling from probability distribution.

Dispatches by device: with emit_dist on NVIDIA SM90+, the cluster-capable launcher builds the emitted distribution across a thread-block cluster; everything else takes the single-block launcher. The branch is a comptime one, so a target without thread-block clusters never instantiates the cluster kernels.

See topk_fi.topk_topp_sampling_from_prob for the parameters, arguments and output contract; both sides share them.

Was this page helpful?