For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
fused_token_sampling_cpu
def fused_token_sampling_cpu[dtype: DType, out_idx_type: DType, KLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TemperatureLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], TopPLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], SeedLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], KEngine: TensorEngine = DefaultEngine, TemperatureEngine: TensorEngine = DefaultEngine, TopPEngine: TensorEngine = DefaultEngine, SeedEngine: TensorEngine = DefaultEngine](max_k: Int, input: TileTensor[dtype, Engine=input.Engine, address_space=input.address_space, linear_idx_type=input.linear_idx_type], out_idxs: TileTensor[out_idx_type, Engine=out_idxs.Engine, address_space=out_idxs.address_space, linear_idx_type=out_idxs.linear_idx_type], k: Optional[TileTensor[.int64, KLayoutType, ImmutAnyOrigin, Engine=KEngine]] = None, temperature: Optional[TileTensor[.float32, TemperatureLayoutType, ImmutAnyOrigin, Engine=TemperatureEngine]] = None, top_p: Optional[TileTensor[.float32, TopPLayoutType, ImmutAnyOrigin, Engine=TopPEngine]] = None, seed: Optional[TileTensor[.uint64, SeedLayoutType, ImmutAnyOrigin, Engine=SeedEngine]] = None)
Generalized implementation of the Top K algorithm with sampling. Returns the sampled index from the innermost dimension of the input tensor for each row/subvolume.
Parameters:
- dtype (
DType): Data type of the input buffer. - out_idx_type (
DType): Data type of the output indices. - KLayoutType (
TensorLayout): Layout type of the k buffer. - TemperatureLayoutType (
TensorLayout): Layout type of the temperature buffer. - TopPLayoutType (
TensorLayout): Layout type of the top_p buffer. - SeedLayoutType (
TensorLayout): Layout type of the seed buffer. - KEngine (
TensorEngine): Engine policy of the k buffer. - TemperatureEngine (
TensorEngine): Engine policy of the temperature buffer. - TopPEngine (
TensorEngine): Engine policy of the top_p buffer. - SeedEngine (
TensorEngine): Engine policy of the seed buffer.
Args:
- max_k (
Int): Largest number of top elements. - input (
TileTensor[dtype, Engine=input.Engine, address_space=input.address_space, linear_idx_type=input.linear_idx_type]): TileTensor[dtype] (Any shape)- The input tensor. - out_idxs (
TileTensor[out_idx_type, Engine=out_idxs.Engine, address_space=out_idxs.address_space, linear_idx_type=out_idxs.linear_idx_type]): TileTensor[out_idx_type] (shape of [input_shape[:-1]] + [1]) - The output indices. - k (
Optional[TileTensor[.int64, KLayoutType, ImmutAnyOrigin, Engine=KEngine]]): Optional device buffer of top elements to keep for each batch element. - temperature (
Optional[TileTensor[.float32, TemperatureLayoutType, ImmutAnyOrigin, Engine=TemperatureEngine]]): The temperature based scaling. - top_p (
Optional[TileTensor[.float32, TopPLayoutType, ImmutAnyOrigin, Engine=TopPEngine]]): Only use the tokens whose cumulative probability exceeds this threshold. - seed (
Optional[TileTensor[.uint64, SeedLayoutType, ImmutAnyOrigin, Engine=SeedEngine]]): The seed to use for the random number generator.