For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python module
max.pipelines.speculative.depth_schedule
How many drafted tokens to verify, as a function of decode batch size.
For now we draft K tokens but only verify a subset of them depending on the schedule.
Type aliases
DepthScheduleEntry | One inclusive batch-size range and the draft depth to use across it. |
|---|
Functions
build_depth_lookup | Expands a schedule into a dense batch_size -> depth lookup. |
|---|---|
normalize_depth_schedule | Validates a schedule and returns it sorted by batch size. |