IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo struct

GPUInfo

struct GPUInfo

Comprehensive information about a GPU architecture.

This struct contains detailed specifications about GPU capabilities, including compute units, memory, thread organization, and performance characteristics.

Fields​

  • ​name (StaticString): The model name of the GPU.
  • ​api (StaticString): The graphics/compute API the GPU is programmed through. Each vendor has its own API, so this also identifies the vendor: "cuda" for NVIDIA, "hip" for AMD, "metal" for Apple, and "none" for NoGPU. A stdlib plugin contributes its own name for hardware the stdlib has no built-in knowledge of. Comparisons must use these exact spellings.
  • ​arch_name (StaticString): The architecture name of the GPU (e.g., sm_80, gfx942).
  • ​compute (Float32): Compute capability version number for NVIDIA GPUs.
  • ​version (StaticString): Version string of the GPU architecture.
  • ​sm_count (Int): Number of streaming multiprocessors (SMs) on the GPU.
  • ​warp_size (Int): Number of threads in a warp/wavefront.
  • ​threads_per_multiprocessor (Int): Maximum number of threads per streaming multiprocessor.
  • ​shared_memory_per_multiprocessor (Int): Size of shared memory available per multiprocessor in bytes.
  • ​max_registers_per_block (Int): Maximum number of registers that can be allocated to a thread block.
  • ​max_thread_block_size (Int): Maximum number of threads allowed in a thread block.

Implemented traits​

AnyType, Copyable, Deinitable, Equatable, Movable, RegisterPassable, Writable

Methods​

__eq__​

def __eq__(self, other: Self) -> Bool

Checks if two GPUInfo instances represent the same GPU model.

Args:

  • ​other (Self): Another GPUInfo instance to compare against.

Returns:

Bool: True if both instances represent the same GPU model.

target​

def target(self) -> __mlir_type.`!kgen.target`

Gets the MLIR target configuration for this GPU.

Returns:

__mlir_type.`!kgen.target`: MLIR target configuration for the GPU.

from_target​

static def from_target[target: __mlir_type.`!kgen.target`]() -> Self

Creates a GPUInfo instance from an MLIR target.

Parameters:

  • ​target (__mlir_type.`!kgen.target`): MLIR target configuration.

Returns:

Self: GPU info corresponding to the target.

from_name​

static def from_name[name: StringSpan[ImmStaticOrigin]]() -> Self

Creates a GPUInfo instance from a GPU architecture name.

Parameters:

  • ​name (StringSpan[ImmStaticOrigin]): GPU architecture name (e.g., "sm_80", "gfx942").

Returns:

Self: GPU info corresponding to the architecture name.

from_family​

static def from_family(family: AcceleratorArchitectureFamily, name: StringSpan[ImmStaticOrigin], api: StringSpan[ImmStaticOrigin], arch_name: StringSpan[ImmStaticOrigin], compute: Float32, version: StringSpan[ImmStaticOrigin], sm_count: Int) -> Self

Creates a GPUInfo instance using architecture family defaults.

This constructor simplifies GPU definition by inheriting common characteristics from an architecture family while allowing specific values to be overridden.

Args:

  • ​family (AcceleratorArchitectureFamily): Architecture family providing default values.
  • ​name (StringSpan[ImmStaticOrigin]): The model name of the GPU.
  • ​api (StringSpan[ImmStaticOrigin]): The graphics/compute API supported by the GPU.
  • ​arch_name (StringSpan[ImmStaticOrigin]): The architecture name of the GPU.
  • ​compute (Float32): Compute capability version number.
  • ​version (StringSpan[ImmStaticOrigin]): Version string of the GPU architecture.
  • ​sm_count (Int): Number of streaming multiprocessors.

Returns:

Self: A fully configured GPUInfo instance.

write_to​

def write_to(self, mut writer: T)

Writes GPU information to a writer.

Outputs all GPU specifications and capabilities to the provided writer in a human-readable format.

Args:

  • ​writer (T): A Writer instance to output the GPU information.

Was this page helpful?