IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Tutorial

Benchmark MAX on NVIDIA or AMD GPUs

In this tutorial, you'll deploy a MAX inference endpoint using our GPU-enabled containers and evaluate endpoint performance with the max benchmark CLI command. You'll collect key metrics such as request throughput, latency, and token throughput to benchmark under production-like workloads and establish a baseline before scaling or integrating MAX into your deployments.

Deploying AI inference workloads means balancing accuracy, latency, and cost. This tutorial benchmarks a Gemma 3 endpoint with the max benchmark CLI command, which reports the following metrics:

  • Request throughput
  • Input and output token throughput
  • Time-to-first-token (TTFT)
  • Time per output token (TPOT)

The benchmark is adapted from vLLM, with additions such as client-side GPU metric collection tailored to MAX. You can see the benchmark script source.

An AI coding agent can run the benchmark for you with the benchmark-model skill, which drives load against an endpoint you're already serving and reports the same throughput and latency metrics, plus GPU utilization when it runs on the same NVIDIA host as the endpoint. See Modular skills to install it.

System requirements:

Get access to the model​

From here on, you should be running commands on the system with your GPU. If you haven't already, open a shell to that system now.

You'll first need to authorize your Hugging Face account to access the Gemma model:

  1. Obtain a Hugging Face access token and set it as an environment variable:

    export HF_TOKEN="hf_..."
  2. Agree to the Gemma 3 license on Hugging Face.

Start the model endpoint​

We provide a pre-configured GPU-enabled Docker container that simplifies deploying an endpoint with MAX. For more information, see MAX container.

Use this command to pull the MAX container and start the model endpoint:

docker run --rm --gpus=all \
  --ipc=host \
  -p 8000:8000 \
  --env "HF_TOKEN=${HF_TOKEN}" \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -v ~/.cache/max_cache:/opt/venv/share/max/.max_cache \
  modular/max-nvidia-full:latest \
  --model google/gemma-3-27b-it

If you want to try a different model, see our supported models.

The endpoint is running when you see the following terminal message (Docker prints JSON logs by default):

πŸš€ Server ready on http://0.0.0.0:8000 (Press CTRL+C to quit)

Start benchmarking​

Open a second terminal and install the max[all] or max[benchmark] package to get the max CLI.

Set up your environment​

  1. If you don't have it, install pixi:
    curl -fsSL https://pixi.sh/install.sh | sh

    Then restart your terminal for the changes to take effect.

  2. Create a project:
    pixi init max-benchmark \
      -c https://conda.modular.com/max-nightly/ -c conda-forge \
      && cd max-benchmark
  3. Install max with all dependencies (nightlyTo get the stable build, change the version in the website header.):
    pixi add max-all
  4. Start the virtual environment:
    pixi shell

Benchmark the model​

To benchmark MAX with the sonnet dataset, use this command:

max benchmark \
  --model google/gemma-3-27b-it \
  --backend modular \
  --endpoint /v1/chat/completions \
  --dataset-name sonnet \
  --num-prompts 500 \
  --sonnet-input-len 550 \
  --output-lengths 256 \
  --sonnet-prefix-len 200

When you want to save your own benchmark configurations, you can define one in a YAML file and pass it to the --config-file option. For example, copy our gemma-3-27b-sonnet-decode-heavy-prefix200.yaml from GitHub, then benchmark the same model with this command:

max benchmark --config-file gemma-3-27b-sonnet-decode-heavy-prefix200.yaml

For more information, including other datasets and configuration options, see the max benchmark documentation.

Use your own dataset​

The command above uses the sonnet dataset from Hugging Face, but you can also provide a path to your own dataset.

For example, you can download the ShareGPT dataset with this command:

wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json

You can then use the local dataset with the --dataset-path argument:

max benchmark \
  ...
  --dataset-path ShareGPT_V3_unfiltered_cleaned_split.json \

Benchmark with image inputs​

Vision models such as Gemma 3 and Qwen3-VL accept image inputs, and a single image can add hundreds of prefill tokens, depending on the model's vision encoder. max benchmark can generate images and mix them into any text dataset's requests.

To attach one generated 512x512 image to half of the requests in the sonnet workload, add the image-mixing flags to the benchmark command:

max benchmark \
  --model google/gemma-3-27b-it \
  --backend modular \
  --endpoint /v1/chat/completions \
  --dataset-name sonnet \
  --num-prompts 500 \
  --image-fraction 0.5 \
  --image-count 1 \
  --image-long-side 512 \
  --image-aspect-ratio 1.0

This run selects half of the requests for images, and each selected request carries one 512x512 image.

--image-fraction controls image mixing: at the default 0.0 the workload stays text-only, and raising it starts mixing images into the selected requests. --image-count sets how many images each selected request carries, --image-long-side sets each image's longer pixel dimension, and --image-aspect-ratio sets width divided by height. 1.0 produces square images, values above 1.0 produce landscape images, and values below 1.0 produce portrait images.

For multi-turn datasets such as ShareGPT, --image-turn picks which user turn carries the image: first, last, or every. --image-count then applies to each selected turn.

--image-count, --image-long-side, and --image-aspect-ratio each accept a distribution string, so image counts and sizes can vary from request to request. Cat() samples from a list of values with explicit weights:

max benchmark \
  --model google/gemma-3-27b-it \
  --backend modular \
  --endpoint /v1/chat/completions \
  --dataset-name sonnet \
  --num-prompts 500 \
  --image-fraction 0.5 \
  --image-long-side "Cat(1024:0.7, 512:0.2, 2048:0.1)"

In this example, 70% of images have a 1024px long side, 20% have 512px, and 10% have 2048px. Weights are normalized, so Cat(1024:7, 512:2, 2048:1) behaves the same, and values without weights are equally likely.

The random dataset has its own image flags, --random-image-count and --random-image-size, which attach the same fixed image count and size to every request. The random dataset's image flags and the image-mixing flags are mutually exclusive: max benchmark exits with an error at startup when both are set.

To preview the workload before sending traffic, add --dry-run: it samples the generated requests and prints distribution statistics, including the image count per request and the long-side pixel distribution, without contacting the server.

For the full flag reference and the distribution grammar, see the max benchmark documentation. For a deeper walkthrough, see Mix generated images into any benchmark workload.

Interpret the results​

Your results depend on your hardware, but the structure of the output should look like this:

============ Serving Benchmark Result ============
Successful requests:                     50
Failed requests:                         0
Benchmark duration (s):                  25.27
Total input tokens:                      12415
Total generated tokens:                  11010
Total nonempty serving response chunks:  11010
Input request rate (req/s):              inf
Request throughput (req/s):              1.97837
------------Client Experience Metrics-------------
Max Concurrency:                         50
Mean input token throughput (tok/s):     282.37
Std input token throughput (tok/s):      304.38
Median input token throughput (tok/s):   140.81
P90 input token throughput (tok/s):      9.76
P95 input token throughput (tok/s):      7.44
P99 input token throughput (tok/s):      4.94
Mean output token throughput (tok/s):    27.31
Std output token throughput (tok/s):     8.08
Median output token throughput (tok/s):  30.64
P90 output token throughput (tok/s):     12.84
P95 output token throughput (tok/s):     9.11
P99 output token throughput (tok/s):     4.71
---------------Time to First Token----------------
Mean TTFT (ms):                          860.54
Std TTFT (ms):                           228.57
Median TTFT (ms):                        809.41
P90 TTFT (ms):                           1214.68
P95 TTFT (ms):                           1215.34
P99 TTFT (ms):                           1215.82
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          46.72
Std TPOT (ms):                           39.77
Median TPOT (ms):                        32.63
P90 TPOT (ms):                           78.24
P95 TPOT (ms):                           111.87
P99 TPOT (ms):                           216.31
---------------Inter-token Latency----------------
Mean ITL (ms):                           31.16
Std ITL (ms):                            91.79
Median ITL (ms):                         1.04
P90 ITL (ms):                            176.93
P95 ITL (ms):                            272.52
P99 ITL (ms):                            276.72
-------------Per-Request E2E Latency--------------
Mean Request Latency (ms):               7694.01
Std Request Latency (ms):                6284.40
Median Request Latency (ms):             5667.19
P90 Request Latency (ms):                16636.07
P95 Request Latency (ms):                21380.10
P99 Request Latency (ms):                25251.18

For more information about each metric, see the MAX benchmarking key metrics.

Measure latency with finite request rates​

Latency metrics like time-to-first-token (TTFT) and time per output token (TPOT) are most meaningful when the endpoint keeps up with the request rate. When the endpoint is overloaded, requests queue, and queue time dominates the results: a benchmark with more prompts produces a deeper queue and higher measured latency.

To control the queue depth, set the average request rate with the --request-rate flag. Requests then arrive at randomized intervals, averaging N requests per second.

Comparing to alternatives​

You can run max benchmark against the Modular or vLLM backends to compare performance with alternative LLM serving frameworks. Before running the benchmark, set up and launch the corresponding inference engine so the benchmark can send requests to it.

Next steps​

Now that you have detailed benchmarking results for Gemma 3 on MAX using an NVIDIA or AMD GPU, you can explore more advanced scaling optimizations:

To read more about our performance methodology, see our blog post, MAX GPU: State of the Art Throughput on a New GenAI platform.

You can also share your experience on the Modular Forum and in our Discord Community.

Read the LLM Inference Handbook to learn more about inference metrics and performance benchmarks.

Was this page helpful?