For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
topk_topp_sampling_from_prob
def topk_topp_sampling_from_prob[dtype: DType, out_idx_type: DType, block_size: Int = Int(1024), from_logits: Bool = False, emit_dist: Bool = False, dist_dtype: DType = .float32, DistLayoutType: TensorLayout = Layout[TypeList[Int64, Int64](), TypeList[Int64, ComptimeInt[Int(1)]]()], TopKArrLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], IndicesLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TopPArrLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], SeedLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TemperatureLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], MinPLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TopKArrEngine: TensorEngine = DefaultEngine, IndicesEngine: TensorEngine = DefaultEngine, TopPArrEngine: TensorEngine = DefaultEngine, SeedEngine: TensorEngine = DefaultEngine, TemperatureEngine: TensorEngine = DefaultEngine, MinPEngine: TensorEngine = DefaultEngine](ctx: DeviceContext, probs: TileTensor[dtype, Engine=probs.Engine, address_space=probs.address_space, linear_idx_type=probs.linear_idx_type], output: TileTensor[out_idx_type, Engine=output.Engine, address_space=output.address_space, linear_idx_type=output.linear_idx_type], top_k_val: Int, top_p_val: Float32 = 1, deterministic: Bool = False, rng_seed: Optional[TileTensor[.uint64, SeedLayoutType, ImmutAnyOrigin, Engine=SeedEngine]] = None, rng_offset: UInt64 = UInt64(0), indices: Optional[TileTensor[out_idx_type, IndicesLayoutType, ImmutAnyOrigin, Engine=IndicesEngine]] = None, top_k_arr: Optional[TileTensor[out_idx_type, TopKArrLayoutType, ImmutAnyOrigin, Engine=TopKArrEngine]] = None, top_p_arr: Optional[TileTensor[.float32, TopPArrLayoutType, ImmutAnyOrigin, Engine=TopPArrEngine]] = None, temperature: Optional[TileTensor[.float32, TemperatureLayoutType, ImmutAnyOrigin, Engine=TemperatureEngine]] = None, min_p: Optional[TileTensor[.float32, MinPLayoutType, ImmutAnyOrigin, Engine=MinPEngine]] = None, out_dist: Optional[TileTensor[dist_dtype, DistLayoutType, MutAnyOrigin]] = None)
Joint top-k + top-p sampling from probability distribution.
Dispatches by device: with emit_dist on NVIDIA SM90+, the
cluster-capable launcher builds the emitted distribution across a
thread-block cluster; everything else takes the single-block
launcher. The branch is a comptime one, so a target without
thread-block clusters never instantiates the cluster kernels.
See topk_fi.topk_topp_sampling_from_prob for the parameters,
arguments and output contract; both sides share them.