For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).
Mojo function
compute_log_probabilities_ragged_shape
def compute_log_probabilities_ragged_shape[levels: Int](logits: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=logits.static_spec], tokens: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=tokens.static_spec], sampled_tokens: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=sampled_tokens.static_spec], logit_row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=logit_row_offsets.static_spec], token_row_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=token_row_offsets.static_spec], lp_output_offsets: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=lp_output_offsets.static_spec], lp_output_offsets_host: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=lp_output_offsets_host.static_spec]) -> IndexList[Int(2)]
Computes the output shapes for the ragged log-probabilities op.
Parameters:
- βlevels (
Int): Number of heap levels; the output second dimension is2**levels.
Args:
- βlogits (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=logits.static_spec]): Input logits ragged by batch, shape[total_rows, vocab_size]. - βtokens (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=tokens.static_spec]): Previously generated tokens across all batches, ragged. - βsampled_tokens (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=sampled_tokens.static_spec]): Most recently sampled token for each batch. - βlogit_row_offsets (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=logit_row_offsets.static_spec]): Per-batch start offsets into the first axis oflogits. - βtoken_row_offsets (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=token_row_offsets.static_spec]): Per-batch start offsets intotokens. - βlp_output_offsets (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=lp_output_offsets.static_spec]): Per-batch start offsets into the output row axis. - βlp_output_offsets_host (
ManagedTensorSlice[IOSpec[_, _].Input, static_spec=lp_output_offsets_host.static_spec]): Host-resident copy oflp_output_offsetswhose last element gives the total number of output tokens.
Returns:
IndexList[Int(2)]: The output shape [num_output_tokens, 2**levels].
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!