For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Environment variables
This page documents all environment variables you can use to configure MAX behavior. These variables control server settings, logging, telemetry, performance, and integrations.
How to set environment variables
You can set environment variables in several ways:
# Export in your shell
export MAX_SERVE_HOST="0.0.0.0"
# Pass to Docker container
docker run --env "MAX_SERVE_HOST=0.0.0.0" modular/max-nvidia-full:latest ...
# Use a .env file in your working directory
echo "MAX_SERVE_HOST=0.0.0.0" >> .envConfiguration precedence
When you configure the same setting in multiple places, the following precedence applies (highest to lowest):
- CLI flags or direct Python initialization: For example,
--port 8080orSettings(MAX_SERVE_PORT=8080). CLI flags pass directly to theSettingsconstructor, so they have the same precedence as direct Python initialization. - Environment variables:
export MAX_SERVE_HOST="0.0.0.0" .envfile values: Values defined in a.envfile in your working directory
Serving
These variables configure the MAX model serving behavior.
For more information on serving a model with MAX, explore the text to text and image and video to text guides.
| Variable | Description | Values | Default |
|---|---|---|---|
MAX_SERVE_HOST | Hostname for the MAX server | String | 0.0.0.0 |
MAX_SERVE_PORT | Port for serving MAX | Integer | 8000 |
MAX_SERVE_METRICS_ENDPOINT_PORT | Port for the Prometheus metrics endpoint | Integer | 8001 |
MAX_SERVE_ALLOWED_IMAGE_ROOTS | Allowed root directories for file:// URI access, as a JSON array (for example, '["/srv/images"]') | JSON array string | Empty |
MAX_SERVE_MAX_LOCAL_IMAGE_BYTES | Maximum size in bytes for local image files | Integer | 20971520 (20 MiB) |
MAX_SERVE_MAX_REQUEST_BYTES | Maximum size in bytes of an accepted HTTP request body. MAX rejects a larger request with HTTP 413, based on the declared Content-Length or on the bytes counted as the body streams in. Raise it for requests that inline large base64 media, or set to 0 to disable the limit. | Integer (≥ 0) | 104857600 (100 MiB) |
MAX_SERVE_MAX_MEDIA_BYTES | Maximum total size in bytes of the media (images and videos) one request may pull in, across every http(s)://, data:, and file: reference it names. MAX bounds both what it fetches and what a single image may decode to, rejecting an oversized decode from the image header before it allocates the pixel buffer. Separate from MAX_SERVE_MAX_REQUEST_BYTES, which bounds only the request body. Set to 0 to disable. | Integer (≥ 0) | 104857600 (100 MiB) |
MAX_SERVE_MEDIA_KIND | Default media kind used in size-limit error messages when a resolver caller doesn't specify one | image, video | image |
MAX_SERVE_MEDIA_URL_ALLOWED_HOSTS | Allowlist of otherwise-blocked internal hosts that MAX can fetch media from while SSRF protection stays on. Each entry is an exact hostname (case-insensitive) or an IP address or CIDR range, such as '["minio.internal", "10.0.0.0/8"]' | JSON array string | Empty |
MAX_SERVE_MEDIA_URL_SSRF_PROTECTION_ENABLED | Protects against server-side request forgery (SSRF) when MAX fetches media from a client-supplied http(s):// URL. MAX checks the host at each redirect and rejects any host that resolves to a private, loopback, or otherwise non-public address. Turn this off only as a last resort, to restore the earlier unchecked behavior. To allow specific internal hosts while the protection stays on, use MAX_SERVE_MEDIA_URL_ALLOWED_HOSTS. | true, false | true |
MAX_SERVE_API_TYPES | Configures which API types MAX serve exposes. Accepts a JSON array of API type strings (for example, '["responses"]'). Use this to enable the Responses API for tasks like image generation. | JSON array string | None |
MAX_SERVE_GRACEFUL_SHUTDOWN_TIMEOUT_S | Seconds to wait for in-flight requests to finish after SIGTERM before canceling them and exiting | Integer | 5 |
MAX_SERVE_MAX_QUEUE_SIZE | Cap on the request queue to the model worker. Once full, the server rejects new requests with HTTP 429 instead of enqueuing them. Pair with MAX_SERVE_MAX_PENDING_REQUESTS for effective backpressure | Integer | None (unbounded) |
MAX_SERVE_MAX_PENDING_REQUESTS | Cap on the scheduler's pending prefill queue depth. When set, the model worker stops pulling new requests from the request queue once it holds this many not-yet-running requests | Integer (≥ 1) | None (unbounded) |
MAX_SERVE_STREAM_MIN_CHUNK_TOKENS | Minimum number of tokens per streamed server-sent events (SSE) chunk. Larger values coalesce streaming output into bigger chunks without affecting time to first token | Integer | 1 |
MAX_SERVE_BATCH_PRIORITY | Batch scheduling strategy that controls how replicas prioritize prefill (context encoding) versus decode (token generation) requests | prefill_first, decode_first, balanced, per_replica | per_replica |
MAX_SERVE_DI_BIND_ADDRESS | Bind address for the disaggregated-inference dispatcher, which carries communication between the decode and prefill workers | String | tcp://127.0.0.1:5555 |
MODULAR_NIXL_TRANSFER_BACKEND | NIXL transfer backend for moving KV cache blocks between prefill and decode workers. Case-insensitive; auto is not accepted | ucx, libfabric, uccl | ucx |
MODULAR_DEVICE_CONTEXT_MEMORY_MANAGER_SIZE_PERCENT | Percentage of available GPU memory the device memory manager may allocate, after the driver takes its share. Required for GPU KV cache transfers in disaggregated serving: it enables the memory manager backing the fast CUDA IPC transport, and startup fails without it. A value of 99 is suggested for serving | Integer (0-100) | 90 |
Logging
These variables control logging behavior and verbosity.
You can read more about logs when using the MAX container.
| Variable | Description | Values | Default |
|---|---|---|---|
MAX_SERVE_LOGS_CONSOLE_LEVEL | Console log verbosity level | CRITICAL, ERROR, WARNING, INFO, DEBUG | INFO |
MODULAR_STRUCTURED_LOGGING | Enable JSON-formatted structured logging for deployed services | 0, 1 | 1 |
MAX_SERVE_LOGS_FILE_PATH | Path to write log files | File path | None |
MAX_SERVE_LOG_PREFIX | Prefix to prepend to all log messages | String | None |
Telemetry and metrics
These variables control telemetry collection and metrics reporting.
For more information, read about telemetry.
| Variable | Description | Values | Default |
|---|---|---|---|
MAX_SERVE_DISABLE_TELEMETRY | Disable remote telemetry collection | 0, 1 | 0 |
MODULAR_USER_ID | User identifier for telemetry (for example, your company name) | String | None |
MAX_SERVE_DEPLOYMENT_ID | Deployment identifier for telemetry (for example, your application name) | String | None |
MAX_SERVE_OTLP_METRICS_ENDPOINT | Push a self-calibrating exponential-histogram shadow (<metric>.exponential) of every histogram metric to an OTLP endpoint, alongside the Prometheus /metrics histograms | Endpoint URL | None |
MAX_SERVE_EPLB_PROFILE | Enable expert-parallel load balancing (EPLB) statistics profiling in max serve | 0, 1 | 0 |
Standard OpenTelemetry variables
MAX reads the standard OpenTelemetry SDK variables, so you can point it at your own collector without MAX-specific configuration.
| Variable | Description | Values | Default |
|---|---|---|---|
OTEL_SERVICE_NAME | Reported as service.name | String | unknown_service |
OTEL_RESOURCE_ATTRIBUTES | Extra resource attributes, see below | key=value,... | None |
OTEL_EXPORTER_OTLP_ENDPOINT | Base metrics endpoint, path appended | URL | Modular's |
OTEL_EXPORTER_OTLP_TRACES_ENDPOINT | Trace endpoint; turns span export on | URL | None |
OTEL_EXPORTER_OTLP_METRICS_ENDPOINT | Metrics endpoint, used as given | URL | Modular's |
OTEL_EXPORTER_OTLP_TRACES_PROTOCOL | Trace protocol, see below | Protocol name | http/protobuf |
OTEL_TRACES_SAMPLER | Trace sampling strategy | Sampler name | parentbased_always_on |
OTEL_TRACES_SAMPLER_ARG | Argument to the sampler | String | None |
OTEL_SDK_DISABLED | Disable OTLP export | true, false | false |
Leave MAX_SERVE_OTLP_METRICS_ENDPOINT unset when you use these variables.
It is a separate opt-in reader, so aiming both at one collector sends every
metric twice, with counters and histograms arriving in two temporalities.
MAX exports spans only when OTEL_EXPORTER_OTLP_TRACES_ENDPOINT is set, and
uses it exactly as given; the base endpoint alone never turns tracing on. For
metrics, a signal-specific endpoint takes precedence over the base endpoint:
setting OTEL_EXPORTER_OTLP_ENDPOINT to http://collector:4318 sends metrics
to http://collector:4318/v1/metrics.
MAX exports metrics over OTLP HTTP only, so their endpoint must be an HTTP
receiver, conventionally port 4318; a gRPC receiver on port 4317 drops them.
Spans go over HTTP too unless OTEL_EXPORTER_OTLP_TRACES_PROTOCOL (or the
generic OTEL_EXPORTER_OTLP_PROTOCOL) is grpc. For gRPC, write the traces
endpoint with an http:// scheme, such as http://collector:4317, to get a
plaintext channel: without one the exporter uses TLS unless
OTEL_EXPORTER_OTLP_TRACES_INSECURE is true.
The published container images set OTEL_SERVICE_NAME to max-serve, so
override it explicitly rather than leaving it unset:
docker run --env "OTEL_SERVICE_NAME=my-service" modular/max-nvidia-full:latest ...OTEL_SDK_DISABLED is enabled only by the value true, ignoring case and
surrounding space, per the OpenTelemetry specification. Any other value, 1
included, leaves telemetry on. OTEL_SDK_DISABLED stops OTLP export only: the
local Prometheus /metrics endpoint keeps serving, and so does an explicitly
set MAX_SERVE_OTLP_METRICS_ENDPOINT.
OTEL_RESOURCE_ATTRIBUTES reaches the resource, but MAX sets deployment.id
and enduser.id itself and an attribute set here loses to that. Use
MAX_SERVE_DEPLOYMENT_ID and MODULAR_USER_ID for those two.
Continuous OTLP log export is off unless MAX_SERVE_LOGS_OTLP_LEVEL is set,
but MAX still sends one record to Modular's collector at startup unless
telemetry is disabled. Either way logs go to Modular:
OTEL_EXPORTER_OTLP_LOGS_ENDPOINT has no effect, and
OTEL_EXPORTER_OTLP_PROTOCOL is not read.
The generic OTEL_EXPORTER_OTLP_HEADERS, OTEL_EXPORTER_OTLP_CERTIFICATE,
OTEL_EXPORTER_OTLP_TIMEOUT and OTEL_EXPORTER_OTLP_COMPRESSION settings are
read by every exporter that gets built, and any signal still pointed at
Modular's collector carries them. If you set headers to authenticate against
your own collector but redirect only traces, your metrics reach Modular with
those headers attached.
Debugging
The MODULAR_DEBUG environment variable enables one or more MAX debugging
options. Set it to a comma-separated list of option names, using name=value
for options that take a value:
export MODULAR_DEBUG=nan-check,assert-level=allThis table includes all options you can add to MODULAR_DEBUG:
| Option | Description | Type | Accepted values | Default |
|---|---|---|---|---|
sensible | Enable a curated default debugging set | Boolean | N/A | false |
nan-check | Insert NaN/Inf checks on a sampled subset of floating-point kernel outputs (see nan-check-stride) | Boolean | N/A | false |
nan-check-stride | When nan-check is enabled, check one of every N floating-point kernel outputs. Set to 1 for full coverage | Integer | Positive integer (≥1) | 20 |
uninitialized-read-check | Detect reads of uninitialized memory | Boolean | N/A | false |
uninitialized-read-mode | report prints each uninitialized-read-check match and continues; output is unbounded. Apple GPUs abort | String | abort, report | abort |
device-sync-mode | Force synchronous GPU execution for debugging | Boolean | N/A | false |
stack-trace-on-error | Show Mojo stack traces on runtime errors | Boolean | N/A | false |
stack-trace-on-crash | Show Mojo stack traces on crashes | Boolean | N/A | false |
source-tracebacks | Include Python source locations in error messages | Boolean | N/A | false |
op-log-level | Log level for op-level execution tracing | String | trace, debug, info, warning, error, critical | off |
assert-level | Assertion level for the Mojo standard library | String | none, warn, safe, all | none |
print-style | Output format for tensor debug printing | String | compact, full, binary, binary_max_checkpoint | compact |
ir-output-dir | Directory where MAX dumps intermediate compiler IR | Path | Filesystem path | unset |
The following variables control memory-allocation debugging:
| Variable | Description | Values | Default |
|---|---|---|---|
MODULAR_DEBUG_DEVICE_ALLOCATOR | Comma-separated allocator debugging modes. poison-all fills every memory-manager allocation with a NaN-pattern byte; the uninitialized-read-check debug option automatically enables uninitialized-poison | Comma-separated modes | None |
MODULAR_DEVICE_CONTEXT_MEMORY_MANAGER_POISON_PATTERN | Byte pattern written by the poison-all allocator mode | Byte (0–255) | Built-in NaN pattern |
MODULAR_DEBUG_ALLOC_POISON_SELECT | Which allocations the uninitialized-poison allocator mode poisons; it zero-fills the rest. See Select allocations to poison | Selection string | all |
MODULAR_DEBUG_ALLOC_POISON_SELECT_FILE | File holding a selection string, checked on every allocation. Each change resets allocation ordinals to 0. Overrides MODULAR_DEBUG_ALLOC_POISON_SELECT | Path | Unset |
The following variable overrides the GPU compiler MAX uses:
| Variable | Description | Values | Default |
|---|---|---|---|
MODULAR_NVPTX_COMPILER_PATH | Path to a custom NVIDIA ptxas binary, useful as an escape hatch when the installed NVIDIA driver is too old for the bundled compiler | Path | Unset |
Select allocations to poison
MODULAR_DEBUG_ALLOC_POISON_SELECT narrows which allocations the
uninitialized-poison mode poisons, so you can bisect for the one an
uninitialized read comes from. Unselected allocations are zero-filled, so both
sides of a comparison stay deterministic. Set it to all (the default),
none, sync-only (which keeps the per-allocation device synchronization but
writes nothing, to tell a race apart from an uninitialized read), or to
dtype=si64,ui32, ord=100-2000, or both joined by &. Ordinals count
allocations from 0, and MAX logs each one with its data type to stderr on a
DEBUG_ALLOC_POISON: line. MODULAR_DEBUG_ALLOC_POISON_SELECT_FILE reads the
selection from a file instead and resets ordinals whenever the file changes;
replace the file with a rename so every change registers.
Profiling
For GPU profiling details, see GPU profiling with Nsight Systems.
The following variable controls runtime profiling and tracing.
| Variable | Description | Values | Default |
|---|---|---|---|
MODULAR_ENABLE_PROFILING | Enable runtime profiling and tracing | off, on, detailed | off |
Performance and caching
The following variables configure caching and memory behavior.
| Variable | Description | Values | Default |
|---|---|---|---|
MODULAR_MAX_CACHE_DIR | Directory to save MAX model cache for reuse | Path | $MODULAR_CACHE_DIR/.max_cache |
MODULAR_CACHE_DIR | Configure cache directory for all MAX filesystems | Path | See note below |
MODULAR_MAX_SHM_WATERMARK | Percentage of /dev/shm to allocate for shared memory. Set to 0.0 to disable shared memory. | Float (0.0–1.0) | 0.9 |
MAX_EAGER_OP_PRECOMPILE | Controls how the eager interpreter compiles its built-in op targets. The default compiles each target lazily on its first dispatch, which avoids a long cold-cache compile when a program touches only a few targets. Set to 1 to precompile the full (device, dtype) matrix at import so steady-state dispatch never recompiles. MAX reads the value when the sweep runs rather than at import. | 0, 1 | 0 |
MAX_EAGER_EXECUTOR | Select the eager execution backend | composite, jit, interpreter, compile | composite |
MAX_EAGER_ALLOW_LAZY_COMPILE | Allow max serve to compile eager interpreter models on demand instead of requiring a warm cache from max warm-interpreter-cache | 0, 1 | 0 |
MODULAR_DEVICE_CONTEXT_MEMORY_MANAGER_VMM | Toggle the VMM defragmenting allocator, which avoids external-fragmentation OOMs. Enabled by default on NVIDIA GPUs; opt in on AMD MI300-series GPUs with 1 | 0, 1 | 1 (NVIDIA), 0 (AMD) |
MODULAR_ENABLE_APPLE_NAIVE_FA_DECODE | Set to 0 to opt out of the split-K decode attention kernel on Apple GPUs (paged KV cache MHA and GQA decode) | 0, 1 | 1 |
MODULAR_MAX_RELEASE_FREE_HOST_MEMORY | Return freed host memory to the operating system after model loading completes | Any non-empty value | Unset (disabled) |
Weight loading
The following variables control how MAX handles model weights when loading from a checkpoint.
| Variable | Description | Values | Default |
|---|---|---|---|
MODULAR_AUTO_CAST_WEIGHTS | Auto-cast loaded weights between float32 and bfloat16 when checkpoint and module dtypes mismatch (shape must match). Other dtype mismatches still raise. | true, false | true |
APPLE_FLUX2_INT8_W8A8 | On Apple M5 GPUs, controls int8 W8A8 quantization for FLUX.2 checkpoints: FLUX.2-klein bf16 checkpoints default to int8 W8A8 (set 0 to opt out); NVFP4 checkpoints opt into an int8 W8A8 requant at load with 1. | 0, 1 | 1 (FLUX.2-klein bf16) |
MODULAR_MAX_RELEASE_HOST_WEIGHTS | Set to 1 to release the host copies of model weights once the model is loaded on GPU, reducing host memory usage. GPU deployments only | 1 | Unset (disabled) |
Hugging Face
Configure your Hugging Face integration with the following environment variable:
| Variable | Description | Values | Default |
|---|---|---|---|
HF_TOKEN | Hugging Face authentication token for accessing gated models | String | None |
Related resources
- MAX container: Deploy MAX with Docker
max serveCLI: Command-line options for serving