For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
fused_token_sampling_gpu
def fused_token_sampling_gpu[dtype: DType, out_idx_type: DType, //, KLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TemperatureLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TopPLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], MinPLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], SeedLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], KEngine: TensorEngine = DefaultEngine, TemperatureEngine: TensorEngine = DefaultEngine, TopPEngine: TensorEngine = DefaultEngine, MinPEngine: TensorEngine = DefaultEngine, SeedEngine: TensorEngine = DefaultEngine](ctx: DeviceContext, max_k: Int, min_top_p: Float32, input: TileTensor[dtype, Engine=input.Engine, address_space=input.address_space, linear_idx_type=input.linear_idx_type], out_idxs: TileTensor[out_idx_type, Engine=out_idxs.Engine, address_space=out_idxs.address_space, linear_idx_type=out_idxs.linear_idx_type], block_size: Optional[Int] = None, num_blocks_per_input: Optional[Int] = None, k: Optional[TileTensor[.int64, KLayoutType, ImmutAnyOrigin, Engine=KEngine]] = None, temperature: Optional[TileTensor[.float32, TemperatureLayoutType, ImmutAnyOrigin, Engine=TemperatureEngine]] = None, top_p: Optional[TileTensor[.float32, TopPLayoutType, ImmutAnyOrigin, Engine=TopPEngine]] = None, min_p: Optional[TileTensor[.float32, MinPLayoutType, ImmutAnyOrigin, Engine=MinPEngine]] = None, seed: Optional[TileTensor[.uint64, SeedLayoutType, ImmutAnyOrigin, Engine=SeedEngine]] = None)
Top K algorithm with fused sampling. Returns the sampled indices from the Top-K of the innermost dimension of the input tensor for each row/subvolume.