For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Benchmark MAX on NVIDIA or AMD GPUs
In this tutorial, you'll deploy a MAX inference endpoint using our
GPU-enabled containers and evaluate endpoint performance with
the max benchmark CLI command. You'll collect key
metrics such as request throughput, latency, and token throughput to benchmark
under production-like workloads and establish a baseline before scaling or
integrating MAX into your deployments.
Deploying AI inference workloads means balancing accuracy, latency, and cost.
This tutorial benchmarks a
Gemma 3 endpoint with the
max benchmark CLI command, which reports the following metrics:
- Request throughput
- Input and output token throughput
- Time-to-first-token (TTFT)
- Time per output token (TPOT)
The benchmark is adapted from vLLM, with additions such as client-side GPU metric collection tailored to MAX. You can see the benchmark script source.
An AI coding agent can run the benchmark for you with the benchmark-model
skill, which drives load against an endpoint you're already serving and
reports the same throughput and latency metrics, plus GPU utilization when it
runs on the same NVIDIA host as the endpoint. See
Modular skills to install it.
System requirements:
Linux
WSL
GPU
Docker
Get access to the modelβ
From here on, you should be running commands on the system with your GPU. If you haven't already, open a shell to that system now.
You'll first need to authorize your Hugging Face account to access the Gemma model:
-
Obtain a Hugging Face access token and set it as an environment variable:
export HF_TOKEN="hf_..." -
Agree to the Gemma 3 license on Hugging Face.
Start the model endpointβ
We provide a pre-configured GPU-enabled Docker container that simplifies deploying an endpoint with MAX. For more information, see MAX container.
Use this command to pull the MAX container and start the model endpoint:
- NVIDIA
- AMD
docker run --rm --gpus=all \
--ipc=host \
-p 8000:8000 \
--env "HF_TOKEN=${HF_TOKEN}" \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-v ~/.cache/max_cache:/opt/venv/share/max/.max_cache \
modular/max-nvidia-full:latest \
--model google/gemma-3-27b-itdocker run \
--device /dev/kfd \
--device /dev/dri \
--group-add video \
--ipc=host \
-p 8000:8000 \
--env "HF_TOKEN=${HF_TOKEN}" \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-v ~/.cache/max_cache:/opt/venv/share/max/.max_cache \
modular/max-amd:latest \
--model google/gemma-3-27b-itIf you want to try a different model, see our supported models.
The endpoint is running when you see the following terminal message (Docker prints JSON logs by default):
π Server ready on http://0.0.0.0:8000 (Press CTRL+C to quit)Start benchmarkingβ
Open a second terminal and install the max[all] or max[benchmark] package
to get the max CLI.
Set up your environmentβ
- pixi
- uv
- If you don't have it, install
pixi:curl -fsSL https://pixi.sh/install.sh | shThen restart your terminal for the changes to take effect.
- Create a project:
pixi init max-benchmark \ -c https://conda.modular.com/max-nightly/ -c conda-forge \ && cd max-benchmark - Install
maxwith all dependencies (nightly):pixi add max-all - Start the virtual environment:
pixi shell
- If you don't have it, install
uv:curl -LsSf https://astral.sh/uv/install.sh | shThen restart your terminal to make
uvaccessible. - Create a project:
uv init max-benchmark && cd max-benchmark - Create and start a virtual environment:
uv venv && source .venv/bin/activate - Install
maxwith all dependencies (nightly):uv add "max[all]" \ --index https://whl.modular.com/nightly/simple/ \ --prerelease allow
Benchmark the modelβ
To benchmark MAX with the sonnet dataset, use this command:
max benchmark \
--model google/gemma-3-27b-it \
--backend modular \
--endpoint /v1/chat/completions \
--dataset-name sonnet \
--num-prompts 500 \
--sonnet-input-len 550 \
--output-lengths 256 \
--sonnet-prefix-len 200When you want to save your own benchmark configurations, you can define one in a
YAML file and pass it to the --config-file option. For example, copy our
gemma-3-27b-sonnet-decode-heavy-prefix200.yaml
from GitHub, then benchmark the same model with this command:
max benchmark --config-file gemma-3-27b-sonnet-decode-heavy-prefix200.yamlFor more information, including other datasets and configuration options, see
the max benchmark documentation.
Use your own datasetβ
The command above uses the sonnet dataset from Hugging Face, but you can
also provide a path to your own dataset.
For example, you can download the ShareGPT dataset with this command:
wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.jsonYou can then use the local dataset with the --dataset-path argument:
max benchmark \
...
--dataset-path ShareGPT_V3_unfiltered_cleaned_split.json \Benchmark with image inputsβ
Vision models such as Gemma 3 and Qwen3-VL accept image inputs, and a single
image can add hundreds of prefill tokens, depending on the model's vision
encoder. max benchmark can generate images and mix them into any text
dataset's requests.
To attach one generated 512x512 image to half of the requests in the sonnet
workload, add the image-mixing flags to the benchmark command:
max benchmark \
--model google/gemma-3-27b-it \
--backend modular \
--endpoint /v1/chat/completions \
--dataset-name sonnet \
--num-prompts 500 \
--image-fraction 0.5 \
--image-count 1 \
--image-long-side 512 \
--image-aspect-ratio 1.0This run selects half of the requests for images, and each selected request carries one 512x512 image.
--image-fraction controls image mixing: at the default 0.0 the workload
stays text-only, and raising it starts mixing images into the selected
requests. --image-count sets how many images each selected request carries,
--image-long-side sets each image's longer pixel dimension, and
--image-aspect-ratio sets width divided by height. 1.0 produces square
images, values above 1.0 produce landscape images, and values below 1.0
produce portrait images.
For multi-turn datasets such as ShareGPT, --image-turn picks which user
turn carries the image: first, last, or every. --image-count then
applies to each selected turn.
--image-count, --image-long-side, and --image-aspect-ratio each accept
a distribution string, so image counts and sizes can vary from request to
request. Cat() samples from a list of values with explicit weights:
max benchmark \
--model google/gemma-3-27b-it \
--backend modular \
--endpoint /v1/chat/completions \
--dataset-name sonnet \
--num-prompts 500 \
--image-fraction 0.5 \
--image-long-side "Cat(1024:0.7, 512:0.2, 2048:0.1)"In this example, 70% of images have a 1024px long side, 20% have 512px,
and 10% have 2048px. Weights are normalized, so Cat(1024:7, 512:2, 2048:1)
behaves the same, and values without weights are equally likely.
The random dataset has its own image flags, --random-image-count and
--random-image-size, which attach the same fixed image count and size to
every request. The random dataset's image flags and the image-mixing flags
are mutually exclusive: max benchmark exits with an error at startup when
both are set.
To preview the workload before sending traffic, add --dry-run: it samples
the generated requests and prints distribution statistics, including the
image count per request and the long-side pixel distribution, without
contacting the server.
For the full flag reference and the distribution grammar, see the
max benchmark documentation. For a deeper walkthrough, see
Mix generated images into any benchmark workload.
Interpret the resultsβ
Your results depend on your hardware, but the structure of the output should look like this:
============ Serving Benchmark Result ============
Successful requests: 50
Failed requests: 0
Benchmark duration (s): 25.27
Total input tokens: 12415
Total generated tokens: 11010
Total nonempty serving response chunks: 11010
Input request rate (req/s): inf
Request throughput (req/s): 1.97837
------------Client Experience Metrics-------------
Max Concurrency: 50
Mean input token throughput (tok/s): 282.37
Std input token throughput (tok/s): 304.38
Median input token throughput (tok/s): 140.81
P90 input token throughput (tok/s): 9.76
P95 input token throughput (tok/s): 7.44
P99 input token throughput (tok/s): 4.94
Mean output token throughput (tok/s): 27.31
Std output token throughput (tok/s): 8.08
Median output token throughput (tok/s): 30.64
P90 output token throughput (tok/s): 12.84
P95 output token throughput (tok/s): 9.11
P99 output token throughput (tok/s): 4.71
---------------Time to First Token----------------
Mean TTFT (ms): 860.54
Std TTFT (ms): 228.57
Median TTFT (ms): 809.41
P90 TTFT (ms): 1214.68
P95 TTFT (ms): 1215.34
P99 TTFT (ms): 1215.82
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 46.72
Std TPOT (ms): 39.77
Median TPOT (ms): 32.63
P90 TPOT (ms): 78.24
P95 TPOT (ms): 111.87
P99 TPOT (ms): 216.31
---------------Inter-token Latency----------------
Mean ITL (ms): 31.16
Std ITL (ms): 91.79
Median ITL (ms): 1.04
P90 ITL (ms): 176.93
P95 ITL (ms): 272.52
P99 ITL (ms): 276.72
-------------Per-Request E2E Latency--------------
Mean Request Latency (ms): 7694.01
Std Request Latency (ms): 6284.40
Median Request Latency (ms): 5667.19
P90 Request Latency (ms): 16636.07
P95 Request Latency (ms): 21380.10
P99 Request Latency (ms): 25251.18For more information about each metric, see the MAX benchmarking key metrics.
Measure latency with finite request ratesβ
Latency metrics like time-to-first-token (TTFT) and time per output token (TPOT) are most meaningful when the endpoint keeps up with the request rate. When the endpoint is overloaded, requests queue, and queue time dominates the results: a benchmark with more prompts produces a deeper queue and higher measured latency.
To control the queue depth, set the average request rate with the
--request-rate flag. Requests then arrive at randomized intervals,
averaging N requests per second.
Comparing to alternativesβ
You can run max benchmark against the Modular or vLLM backends to compare
performance with alternative LLM serving frameworks. Before running the
benchmark, set up and launch the corresponding inference engine so the
benchmark can send requests to it.
Next stepsβ
Now that you have detailed benchmarking results for Gemma 3 on MAX using an NVIDIA or AMD GPU, you can explore more advanced scaling optimizations:
- Deploy MAX on GPU with self-hosted endpoints: Learn how to take a serving endpoint from local testing to production on AWS, GCP, or Azure.
- Using LoRA adapters: Use LoRA adapters with MAX to serve task-specific, fine-tuned variants of LLMs.
- The
benchmark-modelskill: An AI coding agent skill that runs this benchmark workflow for you, then hands off to theprofile-modelskill when the numbers show a bottleneck. See GPU system profiling for that next step.
To read more about our performance methodology, see our blog post, MAX GPU: State of the Art Throughput on a New GenAI platform.
You can also share your experience on the Modular Forum and in our Discord Community.
Read the LLM Inference Handbook to learn more about inference metrics and performance benchmarks.