The config file¶
The config file serves as the input to the mlperf-sysinfo tool. The fields inside it determines which fields get collected, execution path of the tool , and what shape the output file takes(As determined by the particular benchmark rules).
Which profile¶
| Profile | For |
|---|---|
endpoints |
MLPerf Endpoints submissions |
inference |
MLPerf Inference submissions |
training |
MLPerf Training submissions |
profile: defaults to endpoints.
Getting one¶
init writes a commented starter config for the profile you name, containing
only the fields that profile reads:
check names every field still missing or still holding starter text, and
reaches every machine the config mentions, before you spend a capture finding
out.
Every CHANGEME below is a field check will stop on.
What init endpoints writes
# MLPerf system description -- endpoints profile
#
# mlperf-sysinfo check -c sysinfo.yaml # validate + reach everything
# mlperf-sysinfo capture -c sysinfo.yaml # collect and write the file
#
# Anything the tool can detect for itself -- CPU, memory, accelerators, node
# counts, framework version, parallelism and batch size -- is deliberately
# absent below. Only what cannot be detected belongs in this file.
profile: endpoints
output:
dir: results/sysinfo
system:
name: CHANGEME # e.g. H100x8_vLLM
shortened_name: CHANGEME # at most 20 characters, e.g. H100x8
availability: available # available | preview | rdi
accelerator: cuda # cuda | rocm | xpu | tpu | none
cooling: air # air | liquid | passive
nodes:
include_local: false # is this machine part of the system under test?
ssh:
- user@node1
# - user@node2:2222
ssh_key_preconfigured: false
# Optional. Only needed for disaggregated setups, where nodes serve
# different functions. 'match' is compared case-insensitively against the
# detected accelerator model name.
# groups:
# prefill:
# - { match: NVIDIA H100, count: 2 }
# decode:
# - { match: NVIDIA H100, count: 5 }
# Optional. What the SSH nodes are left holding. By default mlcflow installs
# itself at ~/mlcflow on each node and keeps its cache under ~/MLC, both of
# which survive the run. Turn this on where $HOME is shared, small, or simply
# not the volume you want written to: the run gets its own directory under
# /tmp instead, and deletes it on the way out.
# remote:
# isolated: true
serving:
# The endpoint under test, and also probed for the framework name/version.
# Rules 8.2 allow a description instead of a URL, for a hosted endpoint with
# no public address -- e.g. "Managed endpoint, us-east-1, no public URL".
url: http://node1:8000
node: user@node1 # where the server process runs
log: /tmp/serving.log # server stdout/stderr must be redirected here
framework: auto # auto | vllm | sglang | trtllm
submission:
division: standardized # standardized | serviced | rdi
# container_link: https://... # container the submission ran in
notes:
hardware: "" # becomes hw_notes on every node type
software: "" # becomes sw_notes on every node type
# other_hardware: ""
# How the stack was configured for this run. The parallelism degrees and batch
# size are read from serving.log, so they are not listed here -- these three
# cannot be detected from anything on the machine.
run:
# node_config: "prefill: 2x H100; decode: 6x H100" # defaults to a summary of nodes.groups
# config_summary_notes: "" # anything the parallelism fields do not capture
# link_config: https://github.com/myorg/submission/tree/main/configs
What init inference writes
# MLPerf system description -- inference profile
#
# mlperf-sysinfo check -c sysinfo.yaml # validate + reach everything
# mlperf-sysinfo capture -c sysinfo.yaml # collect and write the file
#
# Writes the flat field set the MLPerf Inference submission checker expects.
# Anything detectable is absent below on purpose -- only what cannot be
# detected belongs in this file.
profile: inference
output:
dir: results/sysinfo
system:
name: CHANGEME # e.g. 8xH100_TRT
category: datacenter # datacenter | edge
availability: available # available | preview | rdi
accelerator: cuda # cuda | rocm | xpu | tpu | none
cooling: air
# type_detail: ""
nodes:
include_local: true # single machine: describe the one running this
ssh: [] # add user@host entries for a multi-node system
ssh_key_preconfigured: false
submission:
submitter: CHANGEME
contact: CHANGEME@example.com
division: closed # closed | open
notes:
hardware: ""
software: ""
What init training writes
# MLPerf system description -- training profile
#
# mlperf-sysinfo check -c sysinfo.yaml # validate + reach everything
# mlperf-sysinfo capture -c sysinfo.yaml # collect and write the file
#
# Writes the flat field set mlperf_logging/system_desc_checker validates.
# Anything detectable -- CPU, memory, accelerators, node count, OS, software
# stack -- is deliberately absent below. Only what cannot be detected belongs
# in this file.
#
# The output is named after system.name, because a training submission stores
# it as <submitter>/systems/<system_name>.json.
profile: training
output:
dir: results/sysinfo
system:
name: CHANGEME # e.g. dgx-h100-n8 -- also the output filename
# Training accepts exactly these four, and nothing else. Plain "available"
# is not one of them: training splits it into on-premise and cloud.
availability: Available on-premise
# Available on-premise
# Available cloud
# Preview
# Research, Development, or Internal (RDI)
accelerator: cuda # cuda | rocm | xpu | tpu | none
cooling: air # air | liquid | passive
# How the nodes are wired to each other. No probe can see past the local
# NIC. A single-node system should say so rather than leave it blank.
networking_topology: CHANGEME # e.g. "rail-optimized fat tree, 8x400G per node"
nodes:
include_local: true # single machine: describe the one running this
ssh: [] # add user@host entries for a multi-node system
ssh_key_preconfigured: false
training:
# The framework and version the run used. It lives inside the container
# image, which this tool never opens.
framework: CHANGEME # e.g. "NVIDIA PyTorch Release 25.04"
# Optional short tag for that build. Written only when set.
# framework_name: ngc25.04_pytorch
submission:
submitter: CHANGEME
division: closed # closed | open
notes:
hardware: "" # becomes hw_notes
software: "" # becomes sw_notes
output — where the file lands¶
| Key | Type | Default | Notes |
|---|---|---|---|
dir |
path | . |
Relative paths resolve against the config file, not the cwd |
file |
string | profile's own | Output filename |
Only the deliverable is written to dir. Everything the collection layer
produces goes into dir/.mlperf-sysinfo/. See
Architecture.
remote — what the nodes are left holding¶
Collecting from an SSH node means installing mlcflow on it and leaving an MLC
tree behind. By default both land in the remote user's $HOME — the
virtualenv at ~/mlcflow, the cache under ~/MLC — and stay there, so the
next run reuses them. On your own hardware that is exactly what you want. On
a shared login node, on a home directory with a quota, or anywhere a cached
answer could outlive the hardware it describes, it is not.
| Key | Type | Default | Notes |
|---|---|---|---|
isolated |
bool | false |
Throwaway tree per run, under /tmp, deleted on the way out |
That is the whole setting. mlcflow picks the location: a fresh
/tmp/mlcflow-isolated-<id> on each node, created chmod 700, with the
virtualenv inside it and both MLC roots pointed at it. A trap removes the lot
when the run ends, however it ends.
What it costs, and what it still leaves¶
Nothing is reused between runs, so every capture reinstalls mlcflow on every node. Expect a slower run in exchange for a node that ends as it started, and for a capture that cannot be answered by a stale cache entry.
It is not yet a run that leaves nothing behind. Isolation redirects where
MLC keeps its state; it does not change the directory the remote commands run
in, which is the SSH login directory — normally $HOME. Measured on a real
node, an isolated run still leaves:
| Path (on each node) | What it is |
|---|---|
~/system-info.json |
~180 KB of raw platform detail, written relative to the login directory |
/tmp/mlperf-system-info-single-node/ |
The per-node JSON, at a fixed path under a fixed name |
~/.cache/pip |
pip's own cache, from installing into the throwaway venv |
The last is arguably a feature — it is why the second isolated run is much
faster than the first. The first two are not, and the per-node file is the one
to watch: the name is mlperf-system-info-single-node-<index>.json, so files
from previous runs sit alongside this run's and are indistinguishable by name.
Clear that directory between runs if a node has ever failed mid-capture.
What isolation does remove is the virtualenv and the MLC tree — the large ones, and the ones that go stale.
Where the throwaway tree goes is not configurable here
mlcflow accepts an explicit base directory and an explicit virtualenv
path, and this tool does not pass either. /tmp is writable, private per
run and cleaned up, which covers the case the setting exists for. If your
nodes have a /tmp too small or too locked down for a virtualenv, say so
on the issue tracker — the plumbing is the same and it is a small change.
This applies to the SSH nodes only. The machine running the command writes
where output.dir says, isolated or not.
system.accelerator — which accelerators are probed¶
system.accelerator names the type of accelerator in the system. It selects
the detection step that runs on each node during capture, and the command
check runs to list the accelerators on each node. One value applies to every
node in the config.
| Value | Hardware | What check runs on each node |
|---|---|---|
cuda |
NVIDIA GPUs | nvidia-smi |
rocm |
AMD GPUs | rocm-smi |
xpu |
Intel GPUs | xpu-smi |
tpu |
Google Cloud TPUs | Reads the PCI device list under /sys/bus/pci/devices |
none |
No accelerator | Nothing. No accelerator fields are collected |
What check prints for a node, for example TPU v5p x 4, is for information
only. A reachable node with nothing listed is still collected.
TPU¶
Supported chips are TPU v4, v5e, v5p, v6e and TPU7x. TPU v2 and v3 are not supported.
Detection reads the PCI device list on each node. It does not need sudo,
JAX or libtpu, and it works while a benchmark is running on the TPUs.
accelerators_per_node is the number of TPU chips. A TPU7x chip, which has
two TensorCores, counts as one.
Not every field can be read from the node. The table shows where each accelerator field comes from and which ones you may need to complete yourself:
| Field | Source | When to complete it yourself |
|---|---|---|
accelerator_model_name |
PCI device ID | — |
accelerators_per_node |
Count of TPU chips | — |
accelerator_memory_capacity |
Published HBM size for the chip | — |
accelerator_memory_type |
Published memory type for the chip | TPU v6e: written empty, because the memory type is not yet confirmed |
accelerator_host_interconnect |
PCIe link speed and width | On Cloud TPU VMs the link is not visible to the VM, so the field comes back N/A |
accelerator_interconnect |
ICI for every supported chip |
endpoints profile: comes back N/A |
accelerator_interconnect_topology |
Slice shape from the Cloud TPU or GKE metadata, for example 2x2x1 (v5p-8) |
Outside Cloud TPU and GKE it is empty. inference and training profiles only |
For a multislice job, accelerator_interconnect_topology also gives the
number of slices, for example 2x2x1 (tpu7x-8), 4 slices, when the job
launcher sets MEGASCALE_NUM_SLICES. The network between slices is not
detected. Describe it in submission.notes.hardware.
The libtpu version is recorded in the software fields when the libtpu or
libtpu-nightly package is installed in the Python environment the collection
runs in.
validate lists every field that came back N/A. Edit those in the written
file before you submit.
${VAR} — secrets stay out of the file¶
Any ${VAR} anywhere in the config is replaced from the environment, including
inside lists. An unset variable is a config problem, reported against its path:
extends — share org defaults¶
# ~/.mlperf/org.yaml
submission:
division: standardized
run:
link_config: https://github.com/myorg/submission/tree/main/configs
The child wins. Nested maps merge; lists replace wholesale. A chain that loops back on itself stops at a depth limit rather than following it round forever:
Empty, unset, and absent¶
Three states that look alike in YAML and are not the same thing:
| Written as | Means |
|---|---|
cooling: air |
Set |
cooling: |
Unset. Stays unset; nothing is guessed |
ssh: with every entry commented out |
Absent. Treated as if the key were not there |
The last one is why deleting the final entry under nodes.ssh, or under run:,
is not an error.
Emptying nodes.ssh is an error when nothing else is left to look at
An absent ssh list is only fine while some other machine is still named.
With include_local: false and no serving.node either, there is nothing
to collect from, and the config is rejected before any node is contacted:
Leftover starter text¶
A config still carrying starter text has not been filled in, and check
refuses it. Two patterns are recognised anywhere in the file:
- anything containing
changeme(soCHANGEME@example.comcounts) - anything matching
Insert … hereor<…>
The insert rule is anchored on purpose, so genuine prose survives:
Every string is scanned, not only the fields your profile requires, because
starter text in any field still reaches the submission file. The check report calls
these placeholder.
Unknown and misspelled options¶
check reports an error for any option it does not recognise, and names a
real option where one is close enough:
$ mlperf-sysinfo check -c sysinfo.yaml
error sysinfo.yaml: config is not valid
system.categry: unknown option. Did you mean "category"?
nodes.include-local: unknown option. Did you mean "include_local"?
submission.divison: unknown option. Did you mean "division"?
Matching ignores case, and treats - and _ as the same character.