User Guide
User Guide
User Manual
TABLE OF CONTENTS
[Link]
User Guide v2021.2.1 | ii
Configuring Your Android Device........................................................................ 84
Application...................................................................................................85
2.4. Profiling QNX Targets from the GUI.................................................................. 86
Chapter 3. Export Formats................................................................................... 87
3.1. SQLite Schema Reference..............................................................................87
3.2. JSON and Text Format Description................................................................... 88
Chapter 4. Report Scripts.....................................................................................89
Report Scripts Shipped With Nsight Systems............................................................. 89
apigpusum[:base] -- CUDA API & GPU Summary (CUDA API + kernels + memory ops)........... 89
cudaapisum -- CUDA API Summary...................................................................... 90
cudaapitrace -- CUDA API Trace......................................................................... 90
gpukernsum[:base] -- CUDA GPU Kernel Summary................................................... 90
gpumemsizesum -- GPU Memory Operations Summary (by Size)................................... 91
gpumemtimesum -- GPU Memory Operations Summary (by Time)................................. 91
gpusum[:base] -- GPU Summary (kernels + memory operations)................................... 92
gputrace -- CUDA GPU Trace.............................................................................92
nvtxppsum -- NVTX Push/Pop Range Summary........................................................93
openmpevtsum -- OpenMP Event Summary............................................................ 93
osrtsum -- OS Runtime Summary........................................................................ 93
vulkanmarkerssum -- Vulkan Range Summary......................................................... 94
pixsum -- PIX Range Summary........................................................................... 94
pixsum -- OpenGL KHR_debug Range Summary....................................................... 95
Report Formatters Shipped With Nsight Systems........................................................ 95
Column....................................................................................................... 95
Table.......................................................................................................... 96
CSV............................................................................................................ 96
TSV............................................................................................................ 97
JSON.......................................................................................................... 97
HDoc.......................................................................................................... 97
HTable........................................................................................................ 98
Chapter 5. Migrating from NVIDIA nvprof................................................................. 99
Using the Nsight Systems CLI nvprof Command..........................................................99
CLI nvprof Command Switch Options...................................................................... 99
Next Steps.....................................................................................................102
Chapter 6. Profiling in a Docker on Linux Devices.................................................... 103
Chapter 7. Direct3D Trace.................................................................................. 105
7.1. D3D11 API trace........................................................................................105
7.2. D3D12 API Trace....................................................................................... 105
Chapter 8. WDDM Queues................................................................................... 109
Chapter 9. Vulkan API Trace................................................................................ 111
9.1. Vulkan Overview....................................................................................... 111
9.2. Pipeline Creation Feedback.......................................................................... 112
9.3. Vulkan GPU Trace Notes.............................................................................. 113
[Link]
User Guide v2021.2.1 | iii
Chapter 10. Stutter Analysis................................................................................114
10.1. FPS Overview..........................................................................................114
10.2. Frame Health..........................................................................................117
10.3. GPU Memory Utilization............................................................................. 118
10.4. Vertical Synchronization.............................................................................118
Chapter 11. MPI API Trace.................................................................................. 119
Chapter 12. OpenMP Trace..................................................................................121
Chapter 13. OS Runtime Libraries Trace................................................................. 123
13.1. Locking a Resource...................................................................................124
13.2. Limitations............................................................................................. 124
13.3. OS Runtime Libraries Trace Filters................................................................ 125
13.4. OS Runtime Default Function List................................................................. 126
Chapter 14. NVTX Trace..................................................................................... 129
Chapter 15. CUDA Trace..................................................................................... 132
15.1. CUDA GPU Memory Allocation Graph............................................................. 135
15.2. Unified Memory Transfer Trace.................................................................... 135
Unified Memory CPU Page Faults...................................................................... 137
Unified Memory GPU Page Faults...................................................................... 138
15.3. CUDA Default Function List for CLI............................................................... 140
15.4. cuDNN Function List for X86 CLI...................................................................142
Chapter 16. OpenACC Trace................................................................................ 144
Chapter 17. OpenGL Trace.................................................................................. 146
17.1. OpenGL Trace Using Command Line...............................................................148
Chapter 18. Custom ETW Trace............................................................................150
Chapter 19. GPU Metric Sampling......................................................................... 152
Requirements................................................................................................. 152
Launching GPU Metric Sampling from the GUI..........................................................152
Exporting and Querying Data.............................................................................. 153
Limitations.................................................................................................... 154
Chapter 20. Debug Versions of ELF Files................................................................ 156
Chapter 21. Reading Your Report in GUI.................................................................157
21.1. Generating a New Report........................................................................... 157
21.2. Opening an Existing Report......................................................................... 157
21.3. Sharing a Report File................................................................................ 157
21.4. Report Tab............................................................................................. 157
21.5. Analysis Summary View..............................................................................158
21.6. Timeline View......................................................................................... 158
21.6.1. Timeline...........................................................................................158
Row Height.............................................................................................. 159
21.6.2. Events View...................................................................................... 159
21.6.3. Function Table Modes.......................................................................... 160
21.6.4. Filter Dialog...................................................................................... 163
21.7. Diagnostics Summary View..........................................................................163
[Link]
User Guide v2021.2.1 | iv
21.8. Symbol Resolution Logs View....................................................................... 164
Chapter 22. Broken Backtraces on Tegra................................................................ 165
Chapter 23. Launch Processes in Stopped State....................................................... 167
23.1. LD_PRELOAD........................................................................................... 167
23.2. Launcher............................................................................................... 168
Chapter 24. Import NVTXT..................................................................................170
Commands.....................................................................................................171
Chapter 25. Visual Studio Integration.................................................................... 173
Chapter 26. Troubleshooting............................................................................... 175
GUI Troubleshooting......................................................................................... 175
Android Targets...............................................................................................176
Symbol Resolution........................................................................................... 176
Verbose Logging on Linux Targets.........................................................................178
Verbose Logging on Windows Targets.....................................................................178
QNX Troubleshooting........................................................................................ 179
Chapter 27. Other Resources...............................................................................180
Feature Videos............................................................................................... 180
Blog Posts..................................................................................................... 180
Training Seminars............................................................................................ 180
Conference Presentations.................................................................................. 181
For More Support............................................................................................ 181
[Link]
User Guide v2021.2.1 | v
[Link]
User Guide v2021.2.1 | vi
Chapter 1.
PROFILING FROM THE CLI
or
nsys [command_switch][optional command_switch_options][application] [optional
application_options]
All command line options are case sensitive. For command switch options, when short
options are used, the parameters should follow the switch after a space; e.g. -s cpu.
When long options are used, the switch should be followed by an equal sign and then
the parameter(s); e.g. --sample=cpu.
For this version of Nsight Systems, you must launch a process from the command line
to begin analysis. If an instance of the requested process is already running when the
CLI command is issued, the collection will fail. The launched process will be terminated
when collection is complete unless the user specifies the --kill none option (details
below).
[Link]
User Guide v2021.2.1 | 1
Profiling from the CLI
The Nsight Systems CLI supports concurrent analysis by using sessions. Each Nsight
Systems session is defined by a sequence of CLI commands that define one or more
collections (e.g. when and what data is collected). A session begins with either a start,
launch, or profile command. A session ends with a shutdown command, when a profile
command terminates, or, if requested, when all the process tree(s) launched in the
session exit. Multiple sessions can run concurrently on the same system.
A couple of notes about the use of paths in your command line.
‣ The Nsight Systems command line interface does not handle paths with spaces
properly. Please use paths without spaces
‣ If you run a command (like python X Y Z) from a directory where the command is
not located (like /home/mystuff), and the directory includes a sub-directory with
the same name as the command (like /home/mystuff/python), the command line
parser will interpret that as "/home/mystuff/python X Y Z". This will not work
because python, in this context, would reference the directory, not an executable.
Please either run from the command's home directory or use the full path to the
command.
Command Description
profile A fully formed profiling description
requiring and accepting no further input.
The command switch options used
(see below table) determine when the
collection starts, stops, what collectors are
used (e.g. API trace, IP sampling, etc.),
what processes are monitored, etc.
[Link]
User Guide v2021.2.1 | 2
Profiling from the CLI
Command Description
start Start a collection in interactive mode. The
start command can be executed before or
after a launch command.
stop Stop a collection that was started in
interactive mode. When executed, all
active collections stop, the CLI process
terminates but the application continues
running.
cancel Cancels an existing collection started
in interactive mode. All data already
collected in the current collection is
discarded.
launch In interactive mode, launches an
application in an environment that
supports the requested options. The
launch command can be executed before
or after a start command.
shutdown Disconnects the CLI process from the
launched application and forces the CLI
process to exit. If a collection is pending or
active, it is cancelled
export Generates an export file from an
existing .qdrep file. For more information
about the exported formats see the /
documentation/nsys-exporter directory in
your Nsight Systems installation directory.
stats Post process existing Nsight Systems
result, either in .qdrep or SQLite format,
to generate statistical information. This
option is not available in the Windows CLI
in this release.
status Reports on the status of a CLI-based
collection or the suitability of the profiing
environment.
sessions Gives information about all sessions
running on the system.
nvprof Special option to help with transition
from legacy NVIDIA nvprof tool. Calling
nsys nvprof [options] will provide
the best available translation of nvprof
[options] See Migrating from NVIDIA
nvprof topic for details. No additional
[Link]
User Guide v2021.2.1 | 3
Profiling from the CLI
Command Description
functionality of nsys will be available
when using this option. Note: Not
available on IBM Power targets.
[Link]
User Guide v2021.2.1 | 4
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 5
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 6
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 7
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 8
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 9
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 10
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 11
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 12
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 13
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 14
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 15
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 16
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 17
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 18
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 19
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 20
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 21
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 22
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 23
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 24
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 25
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 26
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 27
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 28
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 29
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 30
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 31
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 32
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 33
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 34
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 35
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 36
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 37
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 38
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 39
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 40
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 41
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 42
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 43
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 44
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 45
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 46
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 47
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 48
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 49
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 50
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 51
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 52
Profiling from the CLI
Reports are generated from an SQLite export of a .qdrep file. If a .qdrep file is specified,
Nsight Systems will look for an accompanying SQLite file and use it. If no SQLite file
exists, one will be exported and created.
Individual reports are generated by calling out to scripts that read data from the SQLite
file and return their report data in CSV format. Nsight Systems ingests this data and
formats it as requested, then displays the data to the console, writes it to a file, or pipes
it to an external process. Adding new reports is as simple as writing a script that can
read the SQLite file and generate the required CSV output. See the shipped scripts as an
example. Both reports and formatters may take arguments to tweak their processing. For
details on shipped scripts and formatters, see Report Scripts topic.
Reports are processed using a three-tuple that consists of 1) the requested report (and
any arguments), 2) the presentation format (and any arguments), and 3) the output
(filename, console, or external process). The first report specified uses the first format
specified, and is presented via the first output specified. The second report uses the
second format for the second output, and so forth. If more reports are specified than
formats or outputs, the format and/or output list is expanded to match the number of
provided reports by repeating the last specified element of the list (or the default, if
nothing was specified).
nsys stats is a very powerful command and can handle complex argument structures,
please see the topic below on Example Stats Command Sequences.
After choosing the stats command switch, the following options are available. Usage:
nsys [global-options] stats [options] [input-file]
[Link]
User Guide v2021.2.1 | 53
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 54
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 55
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 56
Profiling from the CLI
[Link]
User Guide v2021.2.1 | 57
Profiling from the CLI
Subcommand Description
list List all active sessions including ID, name,
and state information
[Link]
User Guide v2021.2.1 | 58
Profiling from the CLI
Effect: Launch the application using the given arguments. Start collecting immediately
and end collection when the application stops. Trace CUDA, OpenGL, NVTX, and
OS runtime libraries APIs. Collect CPU sampling information and thread scheduling
information. With Nsight Systems Embedded Platforms Edition this will only analysis
the single process. With Nsight Systems Workstation Edition this will trace the process
tree. Generate the report#.qdrep file in the default location, incrementing the report
number if needed to avoid overwriting any existing output files.
Limited trace only run
nsys profile --trace=cuda,nvtx -d 20
--sample=none --cpuctxsw=none -o my_test <application>
[application-arguments]
Effect: Launch the application using the given arguments. Start collecting immediately
and end collection after 20 seconds or when the application ends. Trace CUDA and
NVTX APIs. Do not collect CPU sampling information or thread scheduling information.
Profile any child processes. Generate the output file as my_test.qdrep in the current
working directory.
Delayed start run
nsys profile -e TEST_ONLY=0 -y 20
<application> [application-arguments]
Effect: Set environment variable TEST_ONLY=0. Launch the application using the given
arguments. Start collecting after 20 seconds and end collection at application exit. Trace
CUDA, OpenGL, NVTX, and OS runtime libraries APIs. Collect CPU sampling and
thread schedule information. Profile any child processes. Generate the report#.qdrep file
in the default location, incrementing if needed to avoid overwriting any existing output
files.
Collect ftrace events
nsys profile --ftrace=drm/drm_vblank_event
-d 20
[Link]
User Guide v2021.2.1 | 59
Profiling from the CLI
Effect: Launch application. Collect default options and GPU metrics for the first GPU
(a TU10x), using the tu10x-gfxt metric set at the default frequency (10 kHz). Profile any
child processes. Generate the report#.qdrep file in the default location, incrementing if
needed to avoid overwriting any existing output files.
Run GPU metric sampling on all GPUs at a set frequency
nsys profile --gpu-metrics-device=all
--gpu-metrics-frequency=20000 <application>
Effect: Launch application. Collect default options and GPU metrics for all available
GPUs using the first suitable metric set for each and sampling at 20 kHz. Profile any
child processes. Generate the report#.qdrep file in the default location, incrementing if
needed to avoid overwriting any existing output files.
Collect custom ETW trace using configuration file
nsys profile --etw-provider=[Link]
Effect: Configure custom ETW collectors using the contents of [Link]. Collect data for
20 seconds. Generate the report#.qdrep file in the current working directory.
A template JSON configuration file is located at in the Nsight Systems installation
directory as \target-windows-x64\etw_providers_template.json. This path will show up
automatically if you call
nsys profile --help
[Link]
User Guide v2021.2.1 | 60
Profiling from the CLI
‣ EVENT_TRACE_FLAG_JOB
‣ EVENT_TRACE_FLAG_MEMORY_HARD_FAULTS
‣ EVENT_TRACE_FLAG_MEMORY_PAGE_FAULTS
‣ EVENT_TRACE_FLAG_NETWORK_TCPIP
‣ EVENT_TRACE_FLAG_NO_SYSCONFIG
‣ EVENT_TRACE_FLAG_PROCESS
‣ EVENT_TRACE_FLAG_PROCESS_COUNTERS
‣ EVENT_TRACE_FLAG_PROFILE
‣ EVENT_TRACE_FLAG_REGISTRY
‣ EVENT_TRACE_FLAG_SPLIT_IO
‣ EVENT_TRACE_FLAG_SYSTEMCALL
‣ EVENT_TRACE_FLAG_THREAD
‣ EVENT_TRACE_FLAG_VAMAP
‣ EVENT_TRACE_FLAG_VIRTUAL_ALLOC
Typical case: profile a Python script that uses CUDA
nsys profile --trace=cuda,cudnn,cublas,osrt,nvtx
--delay=60 python my_dnn_script.py
Effect: Launch a Python script and start profiling it 60 seconds after the launch, tracing
CUDA, cuDNN, cuBLAS, OS runtime APIs, and NVTX as well as collecting thread
schedule information.
Typical case: profile an app that uses Vulkan
nsys profile --trace=vulkan,osrt,nvtx
--delay=60 ./myapp
Effect: Launch an app and start profiling it 60 seconds after the launch, tracing Vulkan,
OS runtime APIs, and NVTX as well as collecting CPU sampling and thread schedule
information.
Effect: Create interactive CLI process and set it up to begin collecting as soon as an
application is launched. Launch the application, set up to allow tracing of CUDA and
NVTX as well as collection of thread schedule information. Stop only when explicitly
requested. Generate the report#.qdrep in the default location.
If
you
Note: start
a
collection
and
[Link]
User Guide v2021.2.1 | 61
Profiling from the CLI
fail
to
stop
the
collection
(or
if
you
are
allowing
it
to
stop
on
exit,
and
the
application
runs
for
too
long)
your
system’s
storage
space
may
be
filled
with
collected
data
causing
significant
issues
for
the
system.
Nsight
Systems
will
collect
a
different
amount
of
data/
sec
depending
on
options,
but
in
general
Nsight
Systems
does
not
support
runs
[Link]
User Guide v2021.2.1 | 62
Profiling from the CLI
of
more
than
5
minutes
duration.
Effect: Create interactive CLI and launch an application set up for default analysis.
Send application output to the terminal. No data is collected until you manually
start collection at area of interest. Profile until the application ends. Generate the
report#.qdrep in the default location.
If
you
launch
an
application
and
that
application
and
any
descendants
exit
before
Note: start
is
called
Nsight
Systems
will
create
a
fully
formed .qdrep
file
containing
no
data.
Effect: Create interactive CLI process and set it up to begin collecting as soon as
a cudaProfileStart() is detected. Launch application for default analysis, sending
application output to the terminal. Stop collection at next call to cudaProfilerStop,
when the user calls nsys stop, or when the root process terminates. Generate the
report#.qdrep in the default location.
[Link]
User Guide v2021.2.1 | 63
Profiling from the CLI
If
you
call
nsys
launch
before
nsys
start
-
c
cudaProfilerApi
and
the
code
contains
a
large
number
of
short
duration
cudaProfilerStart/
Note:
Stop
pairs,
Nsight
Systems
may
be
unable
to
process
them
correctly,
causing
a
fault.
This
will
be
corrected
in
a
future
version.
The
Nsight
Systems
CLI
does
Note: not
support
multiple
calls
to
the
[Link]
User Guide v2021.2.1 | 64
Profiling from the CLI
cudaProfilerStart/
Stop
API
at
this
time.
Effect: Create interactive CLI process and set it up to begin collecting as soon as an
NVTX range with given message in given domain (capture range) is opened. Launch
application for default analysis, sending application output to the terminal. Stop
collection when all capture ranges are closed, when the user calls nsys stop, or when
the root process terminates. Generate the report#.qdrep in the default location.
The
Nsight
Systems
CLI
only
triggers
the
Note:
profiling
session
for
the
first
capture
range.
This would make the profiling start when the first range with message "profiler" is
opened in domain "service".
‣ Message@*: All ranges with given message in all domains are capture ranges. For
example:
nsys launch -w true -p profiler@* ./app
This would make the profiling start when the first range with message "profiler" is
opened in any domain.
‣ Message: All ranges with given message in default domain are capture ranges. For
example:
nsys launch -w true -p profiler ./app
[Link]
User Guide v2021.2.1 | 65
Profiling from the CLI
This would make the profiling start when the first range with message "profiler" is
opened in the default domain.
‣ By default only messages, provided by NVTX registered strings are considered to
avoid additional overhead. To enable non-registered strings check please launch
your application with NSYS_NVTX_PROFILER_REGISTER_ONLY=0 environment:
nsys launch -w true -p profiler@service -e
NSYS_NVTX_PROFILER_REGISTER_ONLY=0 ./app
Effect: Create interactive CLI and launch an application set up for default analysis.
Send application output to the terminal. No data is collected until the start command
is executed. Collect data from start until stop requested, generate report#.qstrm in
the current working directory. Collect data from second start until the secont stop
request, generate report#.qdrep (incremented by one) in the current working directory.
Shutdown the interactive CLI and send sigkill to the target application's process group.
Calling
nsys
cancel
after
nsys
start
will
Note:
cancel
the
collection
without
generating
a
report.
[Link]
User Guide v2021.2.1 | 66
Profiling from the CLI
or
nsys profile <application>
nsys stats [Link]
Display specific data from a report
nsys stats --report gputrace [Link]
Effect: Export an SQLite file named [Link] from [Link] (assuming it does
not already exist). Print the report generated by the gputrace script to the console in
column format.
Generate multiple reports, in multiple formats, output multiple places
nsys stats --report gputrace --report gpukernsum --report cudaapisum
--format csv,column --output .,- [Link]
Effect: Export an SQLite file named [Link] from [Link] (assuming it does
not already exist). Generate three reports. The first, the gputrace report, will be output
to the file report1_gputrace.csv in CSV format. The other two reports, gpukernsum
and cudaapisum, will be output to the console as columns of data. Although three
reports were given, only two formats and outputs are given. To reconcile this, both the
list of formats and outputs is expanded to match the list of reports by repeating the last
element.
Submit report data to a command
nsys stats --report cudaapisum --format table \ --output @"grep -E
(-|Name|cudaFree)" [Link]
Effect: Open [Link] and run the cudaapisum script on that file. Generate table data
and feed that into the command grep -E (-|Name|cudaFree). The grep command
will filter out everything but the header, formatting, and the cudaFree data, and display
the results to the console.
Note: When the output name starts with @, it is defined as a command. The command
is run, and the output of the report is piped to the command's stdin (standard-input).
The command's stdout and stderr remain attached to the console, so any output will be
displayed directly to the console.
Be aware there are some limitations in how the command string is parsed. No shell
expansions (including *, ?, [], and ~) are supported. The command cannot be piped
to another command, nor redirected to a file using shell syntax. The command and
command arguments are split on whitespace, and no quotes (within the command
syntax) are supported. For commands that require complex command line syntax, it is
suggested that the command be put into a shell script file, and the script designated as
the output command
[Link]
User Guide v2021.2.1 | 67
Profiling from the CLI
generated, you can use the --stats option with the nsys profile or nsys start
command to generate a fixed set of useful summary statistics.
If your run traces CUDA, these include CUDA API, Kernel, and Memory Operation
statistics:
[Link]
User Guide v2021.2.1 | 68
Profiling from the CLI
If your run traces graphics debug markers these include DX11 debug markers, DX12
debug markers, Vulkan debug markers or KHR debug markers:
[Link]
User Guide v2021.2.1 | 69
Profiling from the CLI
Recipes for these statistics as well as documentation on how to create your own metrics
will be available in a future version of the tool.
[Link]
User Guide v2021.2.1 | 70
Profiling from the CLI
The import of really large, multi-gigabyte, .qdstrm files may take up all of the memory
on the host computer and lock up the system. This will be fixed in a later version.
Importing Windows ETL files
For Windows targets, ETL files captured with Xperf or the [Link] command supplied
with GPUView in the Windows Performance Toolkit can be imported to create reports
as if they were captured with Nsight Systems's "WDDM trace" and "Custom ETW trace"
features. Simply choose the .etl file from the Import dialog to convert it to a .qdrep file.
Create .qdrep Using QdstrmImporter
The CLI and QdstrmImporter versions must match to convert a .qdstrm file into a .qdrep
file. This .qdrep file can then be opened in the same version or more recent versions of
the GUI.
To run QdstrmImporter on the host system, find the QdstrmImporter binary in the Host-
x86_64 directory in your installation. QdstrmImporter is available for all host platforms.
See options below.
To run QdstrmImporter on the target system, copy the Linux Host-x86_64 directory to
the target Linux system or install Nsight Systems for Linux host directly on the target.
The Windows or MacOS host QdstrmImporter will not work on a Linux Target. See
options below.
[Link]
User Guide v2021.2.1 | 71
Profiling from the CLI
To profile everything putting the data from each rank into a separate file:
mpirun [mpi options] nsys profile [nsys options]
To profile a single MPI process use a wrapper script. The following script(called
"[Link]") runs nsys on rank 0 only:
#!/bin/bash
if [[ $OMPI_COMM_WORLD_RANK == 0 ]]; then
~/nsys/nsys profile ./myapp "$@" --mydummyargument
else
./myapp "$@"
fi
[Link]
User Guide v2021.2.1 | 72
Profiling from the CLI
Currently
you
will
need
a
dummy
argument
to
the
process,
so
that
Nsight
Systems
can
decide
which
process
to
profile.
This
means
that
your
process
must
Note:
accept
dummy
arguments
to
take
advantage
of
this
workaround.
This
script
as
written
is
for
Open
MPI,
but
should
be
easily
adaptable
to
other
MPI
implementations.
[Link]
User Guide v2021.2.1 | 73
Chapter 2.
PROFILING FROM THE GUI
On Tegra:
[Link]
User Guide v2021.2.1 | 74
Profiling from the GUI
The dialog has simple controls that allow adding, removing, and modifying connections:
Security notice: SSH is only used to establish the initial connection to a target device,
perform checks, and upload necessary files. The actual profiling commands and data
are transferred through a raw, unencrypted socket. Nsight Systems should not be used
in a network setup where attacker-in-the-middle attack is possible, or where untrusted
parties may have network access to the target device.
While connecting to the target device, you will be prompted to input the user's
password. Please note that if you choose to remember the password, it will be stored in
plain text in the configuration file on the host. Stored passwords are bound to the public
key fingerprint of the remote device.
The No authentication option is useful for devices configured for passwordless
login using root username. To enable such a configuration, edit the file /etc/ssh/
sshd_config on the target and specify the following option:
PermitRootLogin yes
Then set empty password using passwd and restart the SSH service with service ssh
restart.
Open ports: The Nsight Systems daemon requires port 22 and port 45555 to be open for
listening. You can confirm that these ports are open with the following command:
sudo firewall-cmd --list-ports --permanent
sudo firewall-cmd --reload
To open a port use the following command, skip --permanent option to open only for
this session:
sudo firewall-cmd --permanent --add-port 45555/tcp
sudo firewall-cmd --reload
[Link]
User Guide v2021.2.1 | 75
Profiling from the GUI
Likewise, if you are running on a cloud system, you must open port 22 and port 45555
for ingress.
Kernel Version Number - To check for the version number of the kernel support of
Nsight Systems on a target device, run the following command on the remote device:
cat /proc/quadd/version
[Link]
User Guide v2021.2.1 | 76
Profiling from the GUI
[Link]
User Guide v2021.2.1 | 77
Profiling from the GUI
Trace all processes – On compatible devices (with kernel module support version 1.107
or higher), this enables trace of all processes and threads in the system. Scheduler events
from all tasks will be recorded.
Collect PMU counters – This allows you to choose which PMU (Performance
Monitoring Unit) counters Nsight Systems will sample. Enable specific counters when
interested in correlating cache misses to functions in your application.
Three different backtrace collections options are available when sampling CPU
instruction pointers. Backtraces can be generated using Intel (c) Last Branch Record
(LBR) registers. LBR backtraces generate minimal overhead but the backtraces have
[Link]
User Guide v2021.2.1 | 78
Profiling from the GUI
limited depth. Backtraces can also be generated using DWARF debug data. DWARF
backtraces incur more overhead than LBR backtraces but have much better depth.
Finally, backtraces can be generated using frame pointers. Frame pointer backtraces
incur medium overhead and have good depth but only resolve frames in the portions
of the application and its libraries (including 3rd party libraries) that were compiled
with frame pointers enabled. Normally, frame pointers are disabled by default during
compilation.
By default, Nsight Systems will use Intel(c) LBRs if available and fall back to using dwarf
unwind if they are not. Choose modes... will allow you to override the default.
The Include child processes switch controls whether API tracing is only for the
launched process, or for all existing and new child processes of the launched process. If
you are running your application through a script, for example a bash script, you need
to set this checkbox.
The Include child processes switch does not control sampling in this version of Nsight
Systems. The full process tree will be sampled regardless of this setting. This will be
fixed in a future version of the product.
Nsight Systems can sample one process tree. Sampling here means interrupting each
processor after a certain number of events and collecting an instruction pointer (IP)/
backtrace sample if the processor is executing the profilee.
When sampling the CPU on a workstation target, Nsight Systems traces thread
context switches and infers thread state as either Running or Blocked. Note that
Blocked in the timeline indicates the thread may be Blocked (Interruptible) or Blocked
(Uninterruptible). Blocked (Uninterruptible) often occurs when a thread has transitioned
into the kernel and cannot be interrupted by a signal. Sampling can be enhanced with
OS runtime libraries tracing; see OS Runtime Libraries Trace for more information.
[Link]
User Guide v2021.2.1 | 79
Profiling from the GUI
Currently Nsight Systems can only sample one process. Sampling here means that the
profilee will be stopped periodically, and backtraces of active threads will be recorded.
Most applications use stripped libraries. In this case, many symbols may stay
unresolved. If unstripped libraries exist, paths to them can be specified using the
Symbol locations... button. Symbol resolution happens on host, and therefore does not
affect performance of profiling on the target.
Additionally, debug versions of ELF files may be picked up from the target system. Refer
to Debug Versions of ELF Files for more information.
[Link]
User Guide v2021.2.1 | 80
Profiling from the GUI
In Attach or launch mode, the process is to first search as if in the Attach only mode,
but if it is not found, the process is launched using the same path and command line
arguments. If NVTX, CUDA, or other trace settings are selected, the process will be
automatically launched with appropriate environment variables.
Note that in some cases, the capabilities of Nsight Systems are not sufficient to correctly
launch the application; for example, if certain environment variables have to be
corrected. In this case, the application has to be started manually and Nsight Systems
should be used in Attach only mode.
The Edit arguments... link will open an editor window, where every command line
argument is edited on a separate line. This is convenient when arguments contain spaces
or quotes.
To properly populate the Search criteria field based on a currently running process on
the target system, use the Select a process button on the right, which has ellipsis as the
caption. The list of processes is automatically refreshed upon opening.
[Link]
User Guide v2021.2.1 | 81
Profiling from the GUI
window. This is useful when tracing games and graphic applications that use fullscreen
display. In these scenarios switching to Nsight Systems' UI would unnecessarily
introduce the window manager's footprint into the trace. To enable the use of Hotkey
check the Hotkey checkbox in the project settings page:
Nsight Systems can sample one process tree. Sampling here means interrupting each
processor periodically. The sampling rate is defined in the project settings and is either
100Hz, 1KHz (default value), 2Khz, 4KHz, or 8KHz.
[Link]
User Guide v2021.2.1 | 82
Profiling from the GUI
On Windows, Nsight Systems can collect thread activity of one process tree. Collecting
thread activity means that each thread context switch event is logged and (optionally) a
backtrace is collected at the point that the thread is scheduled back for execution. Thread
states are displayed on the timeline.
If it was collected, the thread backtrace is displayed when hovering over a region where
the thread execution is blocked.
Symbol Locations
Symbol resolution happens on host, and therefore does not affect performance of
profiling on the target.
Press the Symbol locations... button to open the Configure debug symbols location
dialog.
[Link]
User Guide v2021.2.1 | 83
Profiling from the GUI
When connecting to the target device, Nsight Systems will validate it and install its
daemon into the following location on the device:
/data/local/tmp/[Link]/
[Link]
User Guide v2021.2.1 | 84
Profiling from the GUI
Once the daemon and all required files are installed correctly, a green check mark will
appear and Device is ready text will be displayed:
Application
This section allows you to choose which application to profile. All information will be
collected about the main process of the selected application, except when the Trace all
processes checkbox is enabled.
For non-rooted Android devices, the list of applications only shows information about
debuggable applications. By default, applications that are being developed using the
Android SDK already contain the debuggable option in their manifests.
On rooted Android devices, profiling of all applications is allowed.
For convenience, the application list also shows the process identifiers (PID) of processes
correlated to the listed packages. To refresh this information, use the button in the upper
right corner of the list.
The two checkboxes below the application list are important to ensure that the correct
launch or attach behavior is configured.
Allow sending intent to launch the default activity, when unselected, forces the
profiler to attach to a running process. If no processes are found to correlate to the
specified application name, the profiling session fails to start with an error message.
When selected, Nsight Systems may launch the default intent of the selected application
to make sure it is running and appears on top of the screen on the target device.
In some applications, especially in early stages of development, common bugs related to
handling the lifecycle of activities can be found. In such cases, sending the default intent
may lead to undesired behavior or even crashes of the profilee. Leaving the checkbox
unselected ensures that the profiler does not affect the application.
Restart application if running is a convenient option in two cases:
1. When profiling from the very beginning of the application is desired.
2. When using some of the trace features described below. They require that a
special library is injected into the application in runtime, which happens when
the application is paused by the Android runtime's virtual machine just after
starting. In this case, enabling this option helps ensure that the application is always
restarted and the injection always happens, as opposed to potentially attaching to
the application's process without injection.
Collect NVTX trace. See NVTX Trace for more information.
Collect OpenGL trace. See OpenGL Trace for more information.
[Link]
User Guide v2021.2.1 | 85
Profiling from the GUI
[Link]
User Guide v2021.2.1 | 86
Chapter 3.
EXPORT FORMATS
sqlite> .schema
CREATE TABLE StringIds (id INTEGER PRIMARY KEY, value TEXT NOT NULL);
CREATE TABLE SCHED_EVENTS (id INTEGER PRIMARY KEY AUTOINCREMENT, start INT NOT
NULL, cpu INT NOT NULL, isSchedIn INT NOT NULL, globalTid INT NOT NULL);
CREATE TABLE sqlite_sequence(name,seq);
CREATE TABLE COMPOSITE_EVENTS (id INT NOT NULL PRIMARY KEY, start INT NOT NULL,
cpu INT NOT NULL, threadState INT NOT NULL, globalTid INT NOT NULL, cpuCycles
INT NOT NULL);
...
Note: Currently tables are created lazily therefore not every table described in the
documentation will be present in a particular database. This will change in a future
version of the product
Currently, a table is created for each data type in the exported database. Since usage
patterns for exported data may vary greatly and no default use cases have been
established, no indexes or extra constraints are created. Instead, refer to the Examples
section in your installed documentation directory for a list of common recipes. This may
change in a future version of the product.
[Link]
User Guide v2021.2.1 | 87
Export Formats
Due to current limitations, all fields are declared as NOT NULL, even if the actual value
may be missing. If the value is missing, the cell is set to the default value for that field.
This might change in future versions.
{Event #1}
{Event #2}
...
{Event #N}
{Strings}
{Streams}
{Threads}
For easier grepping of JSON output, the --separate-strings switch may be used to
force manual splitting of strings, streams and thread names data.
Example line split: nsys export --export-json --separate-strings
[Link] -- -
Note, that only last few lines are shown here for clarity and that carriage returns and
indents were added to avoid wrapping documentation.
[Link]
User Guide v2021.2.1 | 88
Chapter 4.
REPORT SCRIPTS
[Link]
User Guide v2021.2.1 | 89
Report Scripts
This report combines data from the cudaapisum, gpukernsum, and gpumemsizesum
reports. It is very similar to profile section of nvprof --dependency-analysis.
[Link]
User Guide v2021.2.1 | 90
Report Scripts
‣ Total Time : The total time used by all executions of this kernel
‣ Instances : The number of calls to this kernal
‣ Average : The average execution time of this kernal
‣ Minimum : The smallest execution time of this kernal
‣ Maximum : The largest execution time of this kernal
‣ Name : The name of the kernal
This report provides a summary of CUDA kernels and their execution times. Note that
the Time(%) column is calculated using a summation of the Total Time column, and
represents that kernel's percent of the execution time of the kernels listed, and not a
percentage of the application wall or CPU execution time.
[Link]
User Guide v2021.2.1 | 91
Report Scripts
[Link]
User Guide v2021.2.1 | 92
Report Scripts
‣ Strm : Stream ID
‣ Name : Trace event name
This report displays a trace of CUDA kernels and memory operations. Items are sorted
by start time.
[Link]
User Guide v2021.2.1 | 93
Report Scripts
‣ Total Time : The total time used by all executions of this function
‣ Num Calls : The number of calls to this function
‣ Average : The average execution time of this function
‣ Minimum : The smallest execution time of this function
‣ Maximum : The largest execution time of this function
‣ Name : The name of the function
This report provides a summary of operating system functions and their execution
times. Note that the Time(%) column is calculated using a summation of the Total Time
column, and represents that function's percent of the execution time of the functions
listed, and not a percentage of the application wall or CPU execution time.
[Link]
User Guide v2021.2.1 | 94
Report Scripts
column, and represents that function's percent of the execution time of the functions
listed, and not a percentage of the application wall or CPU execution time.
Column
Usage:
column[:nohdr][:nolimit][:nofmt][:<width>[:<width>]...]
Arguments
‣ nohdr : Do not display the header
‣ nolimit : Remove 100 character limit from auto-width columns Note: This can result
in extremely wide columns.
‣ nofmt : Do not reformat numbers.
‣ <width>... : Define the explicit width of one or more columns. If the value "." is
given, the column will auto-adjust. If a width of 0 is given, the column will not be
displayed.
The column formatter presents data in vertical text columns. It is primarily designed to
be a human-readable format for displaying data on a console display.
Text data will be left-justified, while numeric data will be right-justified. If the data
overflows the available column width, it will be marked with a "…" character, to indicate
[Link]
User Guide v2021.2.1 | 95
Report Scripts
the data values were clipped. Clipping always occurs on the right-hand side, even for
numeric data.
Numbers will be reformatted to make easier to visually scan and understand.
This includes adding thousands-separators. This process requires that the string
representation of the number is converted into its native representation (integer or
floating point) and then converted back into a string representation to print. This
conversion process attempts to preserve elements of number presentation, such as the
number of decimal places, or the use of scientific notation, but the conversion is not
always perfect (the number should always be the same, but the presentation may not
be). To disable the reformatting process, use the argument nofmt.
If no explicit width is given, the columns auto-adjust their width based off the header
size and the first 100 lines of data. This auto-adjustment is limited to a maximum
width of 100 characters. To allow larger auto-width columns, pass the initial argument
nolimit. If the first 100 lines do not calculate the correct column width, it is suggested
that explicit column widths be provided.
Table
Usage:
table[:nohdr][:nolimit][:nofmt][:<width>[:<width>]...]
Arguments
‣ nohdr : Do not display the header
‣ nolimit : Remove 100 character limit from auto-width columns Note: This can result
in extremely wide columns.
‣ nofmt : Do not reformat numbers.
‣ <width>... : Define the explicit width of one or more columns. If the value "." is
given, the column will auto-adjust. If a width of 0 is given, the column will not be
displayed.
The table formatter presents data in vertical text columns inside text boxes. Other than
the lines between columns, it is identical to the column formatter.
CSV
Usage:
csv[:nohdr]
Arguments
‣ nohdr : Do not display the header
The csv formatter outputs data as comma-separated values. This format is commonly
used for import into other data applications, such as spread-sheets and databases.
There are many different standards for CSV files. Most differences are in how escapes
are handled, meaning data values that contain a comma or space.
[Link]
User Guide v2021.2.1 | 96
Report Scripts
This CSV formatter will escape commas by surrounding the whole value in double-
quotes.
TSV
Usage:
tsv[:nohdr][:esc]
Arguments
‣ nohdr : Do not display the header
‣ esc : escape tab characters, rather than removing them
The tsv formatter outputs data as tab-separated values. This format is sometimes used
for import into other data applications, such as spreadsheets and databases.
Most TSV import/export systems disallow the tab character in data values. The formatter
will normally replace any tab characters with a single space. If the esc argument has
been provided, any tab characters will be replaced with the literal characters "\t".
JSON
Usage:
json
Arguments: no arguments
The json formatter outputs data as an array of JSON objects. Each object represents one
line of data, and uses the column names as field labels. All objects have the same fields.
The formatter attempts to recognize numeric values, as well as JSON keywords, and
converts them. Empty values are passed as an empty string (and not nil, or as a missing
field).
At this time the formatter does not escape quotes, so if a data value includes double-
quotation marks, it will corrupt the JSON file.
HDoc
Usage:
hdoc[:title=<title>][:css=<URL>]
Arguments:
‣ title : string for HTML document title
‣ css : URL of CSS document to include
The hdoc formatter generates a complete, verifiable (mostly), standalone HTML
document. It is designed to be opened in a web browser, or included in a larger
document via an <iframe>.
[Link]
User Guide v2021.2.1 | 97
Report Scripts
HTable
Usage:
htable
Arguments: no arguments
The htable formatter outputs a raw HTML <table> without any of the surrounding
HTML document. It is designed to be included into a larger HTML document. Although
most web browsers will open and display the document, it is better to use the hdoc
format for this type of use.
[Link]
User Guide v2021.2.1 | 98
Chapter 5.
MIGRATING FROM NVIDIA NVPROF
[Link]
User Guide v2021.2.1 | 99
Migrating from NVIDIA nvprof
[Link]
User Guide v2021.2.1 | 100
Migrating from NVIDIA nvprof
[Link]
User Guide v2021.2.1 | 101
Migrating from NVIDIA nvprof
Next Steps
NVIDIA Visual Profiler (NVVP) and NVIDIA nvprof are deprecated. New GPUs and
features will not be supported by those tools. We encourage you to make the move to
Nsight Systems now. For additional information, suggestions, and rationale, see the blog
series in Other Resources.
[Link]
User Guide v2021.2.1 | 102
Chapter 6.
PROFILING IN A DOCKER ON LINUX
DEVICES
Download the default seccomp profile file, [Link], relevant to your Docker version.
If perf_event_open is already listed in the file as guarded by CAP_SYS_ADMIN, then
remove the perf_event_open line. Add the following lines under "syscalls" and save
the resulting file as default_with_perf.json.
{
"name": "perf_event_open",
"action": "SCMP_ACT_ALLOW",
"args": []
},
[Link]
User Guide v2021.2.1 | 103
Profiling in a Docker on Linux Devices
Then you will be able to use the following switch when starting the Docker to apply the
new seccomp profile.
--security-opt seccomp=default_with_perf.json
There is a known issue where Docker collections terminate prematurely with older
versions of the driver and the CUDA Toolkit. If collection is ending unexpectedly, please
update to the latest versions.
After the Docker has been started, use the Nsight Systems CLI to launch a collection
within the Docker. The resulting .qdstrm file can be imported into the Nsight Systems
host like any other CLI result.
[Link]
User Guide v2021.2.1 | 104
Chapter 7.
DIRECT3D TRACE
Nsight Systems has the ability to trace both the Direct3D 11 API and the Direct3D 12 API
on Windows targets.
SLI Trace
Trace SLI queries and peer-to-peer transfers of D3D11 applications. Requires SLI
hardware and an active SLI profile definition in the NVIDIA console.
[Link]
User Guide v2021.2.1 | 105
Direct3D Trace
The Command List Creation row displays time periods when command lists
were being created. This enables developers to improve their application’s
multithreaded command list creation. Command list creation time period is
measured between the call to ID3D12GraphicsCommandList::Reset and the call to
ID3D12GraphicsCommandList::Close.
The GPU row shows an aggregated view of D3D12 API calls and GPU workloads. Note
that not all D3D12 API calls are logged.
A Command Queue row is displayed for each D3D12 command queue created by the
profiled application. The row’s header displays the queue's running index and its type
(Direct, Compute, Copy).
[Link]
User Guide v2021.2.1 | 106
Direct3D Trace
In addition, you can see the PIX command queue CPU-side performance markers, GPU-
side performance markers and the GPU Command List performance markers, each in
their row.
Detecting which CPU thread was blocked by a fence can be difficult in complex apps
that run tens of CPU threads. The timeline view displays the 3 operations involved:
‣ The CPU thread pushing a signal command and fence value into the command
queue. This is displayed on the DX12 Synchronization sub-row of the calling thread.
‣ The GPU executing that command, setting the fence value and signaling the fence.
This is displayed on the GPU Queue Synchronization sub-row.
‣ The CPU thread calling a Win32 wait API to block-wait until the fence is signaled.
This is displayed on the Thread's OS runtime libraries row.
Clicking one of these will highlight it and the corresponding other two calls.
[Link]
User Guide v2021.2.1 | 107
Direct3D Trace
[Link]
User Guide v2021.2.1 | 108
Chapter 8.
WDDM QUEUES
The Windows Display Driver Model (WDDM) architecture uses queues to send work
packets from the CPU to the GPU. Each D3D device in each process is associated
with one or more contexts. Graphics, compute, and copy commands that the profiled
application uses are associated with a context, batched in a command buffer, and pushed
into the relevant queue associated with that context.
Nsight Systems can capture the state of these queues during the trace session.
Enabling the "Collect additional range of ETW events" option will also capture extended
DxgKrnl events such as context status, allocations, sync wait, signal events, etc.
A command buffer in a WDDM queues may have one the following types:
‣ Render
‣ Deferred
‣ System
‣ MMIOFlip
‣ Wait
‣ Signal
‣ Device
‣ Software
It may also be marked as a Present buffer, indicating that the application has finished
rendering and requests to display the source surface.
[Link]
User Guide v2021.2.1 | 109
WDDM Queues
See the Microsoft documentation for the WDDM architecture and the
DXGKETW_QUEUE_PACKET_TYPE enumeration.
To retain the .etl trace files captured, so that they can be viewed in other tools (e.g.
GPUView), change the "Save ETW log files in project folder" option under "Profile
Behavior" in Nsight Systems's global Options dialog. The .etl files will appear in the
same folder as the .qdrep file, accessible by right-clicking the report in the Project
Explorer and choosing "Show in Folder...". Data collected from each ETW provider will
appear in its own .etl file, and an additional .etl file named "Report XX-Merged-*.etl",
containing the events from all captured sources, will be created as well.
[Link]
User Guide v2021.2.1 | 110
Chapter 9.
VULKAN API TRACE
The Command Buffer Creation row displays time periods when command buffers were
being created. This enables developers to improve their application’s multi-threaded
command buffer creation. Command buffer creation time period is measured between
the call to vkBeginCommandBuffer and the call to vkEndCommandBuffer.
The Swap chains row displays the available swap chains and the time periods where
vkQueuePresentKHR was executed on each swap chain.
[Link]
User Guide v2021.2.1 | 111
Vulkan API Trace
A Queue row is displayed for each Vulkan queue created by the profiled application.
The API sub-row displays time periods where vkQueueSubmit was called. The GPU
Workload sub-row displays time periods where workloads were executed by the GPU.
In addition, you can see Vulkan debug util labels on both the CPU and the GPU.
[Link]
User Guide v2021.2.1 | 112
Vulkan API Trace
[Link]
User Guide v2021.2.1 | 113
Chapter 10.
STUTTER ANALYSIS
The frame duration row displays live FPS statistics for the current timeline viewport.
Values shown are:
1. Number of CPU frames shown of the total number captured
2. Average, minimal, and maximal CPU frame time of the currently displayed time
range
3. Average FPS value for the currently displayed frames
4. The 99th percentile value of the frame lengths (such that only 1% of the frames in the
range are longer than this value).
The values will update automatically when scrolling, zooming or filtering the timeline
view.
[Link]
User Guide v2021.2.1 | 114
Stutter Analysis
The stutter row highlights frames that are significantly longer than the other frames in
their immediate vicinity.
The stutter row uses an algorithm that compares the duration of each frame to the
median duration of the surrounding 19 frames. Duration difference under 4 milliseconds
is never considered a stutter, to avoid cluttering the display with frames whose absolute
stutter is small and not noticeable to the user.
For example, if the stutter threshold is set at 20%:
1. Median duration is 10 ms. Frame with 13 ms time will not be reported (relative
difference > 20%, absolute difference < 4 ms)
2. Median duration is 60 ms. Frame with 71 ms time will not be reported (relative
difference < 20%, absolute difference > 4 ms)
3. Median duration is 60 ms. Frame with 80 ms is a stutter (relative difference > 20%,
absolute difference > 4 ms, both conditions met)
OSC detection
The "19 frame window median" algorithm by itself may not work well with some cases
of "oscillation" (consecutive fast and slow frames), resulting in some false positives. The
median duration is not meaningful in cases of oscillation and can be misleading.
To address the issue and identify if oscillating frames, the following method is applied:
1. For every frame, calculate the median duration, 1st and 3rd quartiles of 19-frames
window.
2. Calculate the delta and ratio between 1st and 3rd quartiles.
3. If the 90th percentile of 3rd – 1st quartile delta array > 4 ms AND the 90th percentile
of 3rd/1st quartile array > 1.2 (120%) then mark the results with "OSC" text.
Right-clicking the Frame Duration row caption lets you choose the target frame rate (30,
60, 90 or custom frames per second).
By clicking the Customize FPS Display option, a customization dialog pops up. In the
dialog, you can now define the frame duration threshold to customize the view of the
potentially problematic frames. In addition, you can define the threshold for the stutter
analysis frames.
[Link]
User Guide v2021.2.1 | 115
Stutter Analysis
The CPU Frame Duration row displays the CPU frame duration measured between the
ends of consecutive frame boundary calls:
‣ The OpenGL frame boundaries are eglSwapBuffers/glXSwapBuffers/
SwapBuffers calls.
‣ The D3D11 and D3D12 frame boundaries are IDXGISwapChainX::Present calls.
‣ The Vulkan frame boundaries are vkQueuePresentKHR calls.
The GPU Frame Duration row displays the time measured between
‣ The start time of the first GPU workload execution of this frame.
‣ The start time of the first GPU workload execution of the next frame.
Reflex SDK
NVIDIA Reflex SDK is a series of NVAPI calls that allow applications to integrate the
Ultra Low Latency driver feature more directly into their game to further optimize
synchronization between simulation and rendering stages and lower the latency
between user input and final image rendering. For more details about Reflex SDK, see
Reflex SDK Site.
Nsight Systems will automatically capture NVAPI functions when either Direct3D 11,
Direct3D 12, or Vulkan API trace are enabled.
The Reflex SDK row displays timeline ranges for the following types of latency markers:
‣ RenderSubmit.
‣ Simulation.
‣ Present.
‣ Driver.
‣ OS Render Queue.
[Link]
User Guide v2021.2.1 | 116
Stutter Analysis
‣ GPU Render.
[Link]
User Guide v2021.2.1 | 117
Stutter Analysis
Note that this is not the same as the CUDA kernel memory allocation graph, see CUDA
GPU Memory Graph for that functionality.
[Link]
User Guide v2021.2.1 | 118
Chapter 11.
MPI API TRACE
For Linux x86_64 and Power targets, Nsight Systems is capable of capturing information
about the MPI APIs executed in the profiled process. It has built-in API trace support
only for the OpenMPI and MPICH implementations of MPI and only for a default list of
synchronous APIs.
If you require more control over the list of traced APIs or if you are using a different
MPI implementation, see github nvtx pmpi wrappers. You can use this documentation
to generate a shared object to wrap a list of synchronous MPI APIs with NVTX using
the MPI profiling interface (PMPI). If you set your LD_PRELOAD environment variable
to the path of that object, Nsight Systems will capture and report the MPI API trace
information when NVTX tracing is enabled.
[Link]
User Guide v2021.2.1 | 119
MPI API Trace
[Link]
User Guide v2021.2.1 | 120
Chapter 12.
OPENMP TRACE
Nsight Systems for Linux x86_64 and Power targets is capable of capturing information
about OpenMP events. This functionality is built on the OpenMP Tools Interface
(OMPT), full support is available only for runtime libraries supporting tools interface
defined in OpenMP 5.0 or greater.
As an example, LLVM OpenMP runtime library partially implements tools interface.
If you use PGI compiler <= 20.4 to build your OpenMP applications, add -mp=libomp
switch to use LLVM OpenMP runtime and enable OMPT based tracing. If you use
Clang, make sure the LLVM OpenMP runtime library you link to was compiled with
tools interface enabled.
The
raw
Note: OMPT
events
are
[Link]
User Guide v2021.2.1 | 121
OpenMP Trace
used
to
generate
ranges
indicating
the
runtime
of
OpenMP
operations
and
constructs.
Example screenshot:
[Link]
User Guide v2021.2.1 | 122
Chapter 13.
OS RUNTIME LIBRARIES TRACE
OS runtime libraries can be traced to gather information about low-level userspace APIs.
This traces the system call wrappers and thread synchronization interfaces exposed by
the C runtime and POSIX Threads (pthread) libraries. This does not perform a complete
runtime library API trace, but instead focuses on the functions that can take a long time
to execute, or could potentially cause your thread be unscheduled from the CPU while
waiting for an event to complete.
OS runtime tracing complements and enhances sampling information by:
1. Visualizing when the process is communicating with the hardware, controlling
resources, performing multi-threading synchronization or interacting with the
kernel scheduler.
2. Adding additional thread states by correlating how OS runtime libraries traces affect
the thread scheduling:
‣ Waiting — the thread is not scheduled on a CPU, it is inside of an OS runtime
libraries trace and is believed to be waiting on the firmware to complete a
request.
‣ In OS runtime library function — the thread is scheduled on a CPU and inside
of an OS runtime libraries trace. If the trace represents a system call, the process
is likely running in kernel mode.
3. Collecting backtraces for long OS runtime libraries call. This provides a way to
gather blocked-state backtraces, allowing you to gain more context about why the
thread was blocked so long, yet avoiding unnecessary overhead for short events.
[Link]
User Guide v2021.2.1 | 123
OS Runtime Libraries Trace
You can also use Skip if shorter than. This will skip calls shorter than the given
threshold. Enabling this option will improve performances as well as reduce noise on
thetimeline. We strongly encourage you to skip OS runtime libraries call shorter than 1
μs.
Note that even if a call is determined as potentially blocking, there is a chance that it
may not actually block after a few cycles have elapsed. The call will still be traced in this
scenario.
13.2. Limitations
‣ Nsight Systems only traces syscall wrappers exposed by the C runtime. It is not able
to trace syscall invoked through assembly code.
[Link]
User Guide v2021.2.1 | 124
OS Runtime Libraries Trace
‣ Additional thread states, as well as backtrace collection on long calls, are only
enabled if sampling is turned on.
‣ It is not possible to configure the depth and duration threshold when collecting
backtraces. Currently, only OS runtime libraries calls longer than 80 μs will generate
a backtrace with a maximum of 24 frames. This limitation will be removed in a
future version of the product.
‣ It is required to compile your application and libraries with the -funwind-tables
compiler flag in order for Nsight Systems to unwind the backtraces correctly.
[Link]
User Guide v2021.2.1 | 125
OS Runtime Libraries Trace
POSIX Threads
pthread_barrier_wait
pthread_cancel
pthread_cond_broadcast
pthread_cond_signal
pthread_cond_timedwait
pthread_cond_wait
pthread_create
pthread_join
pthread_kill
pthread_mutex_lock
pthread_mutex_timedlock
pthread_mutex_trylock
pthread_rwlock_rdlock
pthread_rwlock_timedrdlock
pthread_rwlock_timedwrlock
pthread_rwlock_tryrdlock
pthread_rwlock_trywrlock
pthread_rwlock_wrlock
pthread_spin_lock
pthread_spin_trylock
pthread_timedjoin_np
pthread_tryjoin_np
pthread_yield
sem_timedwait
sem_trywait
sem_wait
[Link]
User Guide v2021.2.1 | 127
OS Runtime Libraries Trace
I/O
aio_fsync
aio_fsync64
aio_suspend
aio_suspend64
fclose
fcloseall
fflush
fflush_unlocked
fgetc
fgetc_unlocked
fgets
fgets_unlocked
fgetwc
fgetwc_unlocked
fgetws
fgetws_unlocked
flockfile
fopen
fopen64
fputc
fputc_unlocked
fputs
fputs_unlocked
fputwc
fputwc_unlocked
fputws
fputws_unlocked
fread
fread_unlocked
freopen
freopen64
ftrylockfile
fwrite
fwrite_unlocked
getc
getc_unlocked
getdelim
getline
getw
getwc
getwc_unlocked
lockf
lockf64
mkfifo
mkfifoat
posix_fallocate
posix_fallocate64
putc
putc_unlocked
putwc
putwc_unlocked
Miscellaneous
forkpty
popen
posix_spawn
posix_spawnp
sigwait
sigwaitinfo
sleep
system
usleep
[Link]
User Guide v2021.2.1 | 128
Chapter 14.
NVTX TRACE
The NVIDIA Tools Extension Library (NVTX) is a powerful mechanism that allows
users to manually instrument their application. Nsight Systems can then collect the
information and present it on the timeline.
Nsight Systems supports version 3.0 of the NVTX specification.
The following features are supported:
‣ Domains
nvtxDomainCreate(), nvtxDomainDestroy()
nvtxDomainRegisterString()
‣ Push-pop ranges (nested ranges that start and end in the same thread).
nvtxRangePush(), nvtxRangePushEx()
nvtxRangePop()
nvtxDomainRangePushEx()
nvtxDomainRangePop()
‣ Start-end ranges (ranges that are global to the process and are not restricted to a
single thread)
nvtxRangeStart(), nvtxRangeStartEx()
nvtxRangeEnd()
nvtxDomainRangeStartEx()
nvtxDomainRangeEnd()
‣ Marks
nvtxMark(), nvtxMarkEx()
nvtxDomainMarkEx()
‣ Thread names
nvtxNameOsThread()
‣ Categories
nvtxNameCategory()
nvtxDomainNameCategory()
To learn more about specific features of NVTX, please refer to the NVTX header file:
nvToolsExt.h or the NVTX documentation.
[Link]
User Guide v2021.2.1 | 129
NVTX Trace
In addition, by enabling the "Insert NVTX Marker hotkey" option it is possible to add
NVTX markers to a running non-console applications by pressing the F11 key. These will
appear in the report under the NVTX Domain named "HotKey markers".
Typically calls to NVTX functions can be left in the source code even if the application is
not being built for profiling purposes, since the overhead is very low when the profiler is
not attached.
NVTX is not intended to annotate very small pieces of code that are being called very
frequently. A good rule of thumb to use: if code being annotated usually takes less than
1 microsecond to execute, adding an NVTX range around this code should be done
carefully.
Range
annotations
should
be
matched
carefully.
Note: If
many
ranges
are
opened
but
not
closed,
[Link]
User Guide v2021.2.1 | 130
NVTX Trace
Nsight
Systems
has
no
meaningful
way
to
visualize
it.
A
rule
of
thumb
is
to
not
have
more
than
a
couple
dozen
ranges
open
at
any
point
in
time.
Nsight
Systems
does
not
support
reports
with
many
unclosed
ranges.
[Link]
User Guide v2021.2.1 | 131
Chapter 15.
CUDA TRACE
Near the bottom of the timeline row tree, the GPU node will appear and contain a
CUDA node. Within the CUDA node, each CUDA context used within the process will
be shown along with its corresponding CUDA streams. Steams will contain memory
operations and kernel launches on the GPU. Kernel launches are represented by blue,
while memory transfers are displayed in red.
[Link]
User Guide v2021.2.1 | 132
CUDA Trace
The easiest way to capture CUDA information is to launch the process from Nsight
Systems, and it will setup the environment for you. To do so, simply set up a normal
launch and select the Collect CUDA trace checkbox.
For Nsight Systems Workstation Edition this looks like:
[Link]
User Guide v2021.2.1 | 133
CUDA Trace
cudaDeviceReset(), and then let the application gracefully exit (as opposed to
crashing).
This option allows flushing CUDA trace data even before the device is finalized.
However, it might introduce additional overhead to a random CUDA Driver or
CUDA Runtime API call.
‣ Skip some API calls — avoids tracing insignificant CUDA Runtime
API calls (namely, cudaConfigureCall(), cudaSetupArgument(),
cudaHostGetDevicePointers()). Not tracing these functions allows Nsight
Systems to significantly reduce the profiling overhead, without losing any
interesting data. (See CUDA Trace Filters, below)
‣ Collect GPU Memory Usage - collects information used to generate a graph of
CUDA allocated memory across time. Note that this will increase overhead. See
section on CUDA GPU Memory Allocation Graph below.
‣ Collect Unified Memory CPU page faults - collects information on page faults that
occur when CPU code tries to access a memory page that resides on the device. See
section on Unified Memory CPU Page Faults in the Unified Memory Transfer
Trace documentation below.
‣ Collect Unified Memory GPU page faults - collects information on page faults that
occur when GPU code tries to access a memory page that resides on the CPU. See
section on Unified Memory GPU Page Faults in the Unified Memory Transfer
Trace documentation below.
‣ For Nsight Systems Workstation Edition, Collect cuDNN trace, Collect cuBLAS
trace, Collect OpenACC trace - selects which (if any) extra libraries that depend on
CUDA to trace.
OpenACC versions 2.0, 2.5, and 2.6 are supported when using PGI runtime version
15.7 or greater and not compiling statically. In order to differentiate constructs, a PGI
runtime of 16.1 or later is required. Note that Nsight Systems Workstation Edition
does not support the GCC implementation of OpenACC at this time.
‣ For Nsight Systems Embedded Platforms Edition if desired, the target application
can be manually set up to collect CUDA trace. To capture information about CUDA
execution, the following requirements should be satisfied:
‣ The profiled process should be started with the specified environment variable,
depending on the architecture of the process:
‣ For ARMv7 (32-bit) processes: CUDA_INJECTION32_PATH, which should
point to the injection library:
/opt/nvidia/nsight_systems/[Link]
‣ For ARMv8 (64-bit) processes: CUDA_INJECTION64_PATH, which should
point to the injection library:
/opt/nvidia/nsight_systems/[Link]
‣ If the application is started by Nsight Systems, all required environment
variables will be set automatically.
Please note that if your application crashes before all collected CUDA trace data has
been copied out, some or all data might be lost and not present in the report.
[Link]
User Guide v2021.2.1 | 134
CUDA Trace
[Link]
User Guide v2021.2.1 | 135
CUDA Trace
HtoD transfer indicates the CUDA kernel accessed managed memory that was residing
on the host, so the kernel execution paused and transferred the data to the device. Heavy
traffic here will incur performance penalties in CUDA kernels, so consider using manual
cudaMemcpy operations from pinned host memory instead.
PtoP transfer indicates the CUDA kernel accessed managed memory that was residing
on a different device, so the kernel execution paused and transferred the data to this
device. Heavy traffic here will incur performance penalties, so consider using manual
cudaMemcpyPeer operations to transfer from other devices' memory instead. The row
showing these events is for the destination device -- the source device is shown in the
tooltip for each transfer event.
DtoH transfer indicates the CPU accessed managed memory that was residing on a
CUDA device, so the CPU execution paused and transferred the data to system memory.
Heavy traffic here will incur performance penalties in CPU code, so consider using
manual cudaMemcpy operations from pinned host memory instead.
Some Unified Memory transfers are highlighted with red to indicate potential
performance issues:
[Link]
User Guide v2021.2.1 | 136
CUDA Trace
Collecting
Unified
Memory
CPU
page
faults
can
cause
overhead
of
up
Note:
to
70%
in
testing.
Please
use
this
functionality
only
when
needed.
[Link]
User Guide v2021.2.1 | 137
CUDA Trace
Collecting
Unified
Memory
GPU
page
faults
can
cause
overhead
Note: of
up
to
70%
in
testing.
Please
use
this
functionality
only
[Link]
User Guide v2021.2.1 | 138
CUDA Trace
when
needed.
[Link]
User Guide v2021.2.1 | 139
CUDA Trace
[Link]
User Guide v2021.2.1 | 143
Chapter 16.
OPENACC TRACE
Nsight Systems for Linux x86_64 and Power targets is capable of capturing information
about OpenACC execution in the profiled process.
OpenACC versions 2.0, 2.5, and 2.6 are supported when using PGI runtime version 15.7
or later. In order to differentiate constructs (see tooltip below), a PGI runtime of 16.0 or
later is required. Note that Nsight Systems does not support the GCC implementation of
OpenACC at this time.
Under the CPU rows in the timeline tree, each thread that uses OpenACC will show
OpenACC trace information. You can click on a OpenACC API call to see correlation
with the underlying CUDA API calls (highlighted in teal):
If the OpenACC API results in GPU work, that will also be highlighted:
[Link]
User Guide v2021.2.1 | 144
OpenACC Trace
Hovering over a particular OpenACC construct will bring up a tooltip with details about
that construct:
To capture OpenACC information from the Nsight Systems GUI, select the Collect
OpenACC trace checkbox under Collect CUDA trace configurations. Note that turning
on OpenACC tracing will also turn on CUDA tracing.
Please note that if your application crashes before all collected OpenACC trace data has
been copied out, some or all data might be lost and not present in the report.
[Link]
User Guide v2021.2.1 | 145
Chapter 17.
OPENGL TRACE
OpenGL and OpenGL ES APIs can be traced to assist in the analysis of CPU and GPU
interactions.
A few usage examples are:
1. Visualize how long eglSwapBuffers (or similar) is taking.
2. API trace can easily show correlations between thread state and graphics driver's
behavior, uncovering where the CPU may be waiting on the GPU.
3. Spot bubbles of opportunity on the GPU, where more GPU workload could be
created.
4. Use KHR_debug extension to trace GL events on both the CPU and GPU.
OpenGL trace feature in Nsight Systems consists of two different activities which will be
shown in the CPU rows for those threads
‣ CPU trace: interception of API calls that an application does to APIs (such as
OpenGL, OpenGL ES, EGL, GLX, WGL, etc.).
‣ GPU trace (or workload trace): trace of GPU workload (activity) triggered by use
of OpenGL or OpenGL ES. Since draw calls are executed back-to-back, the GPU
workload trace ranges include many OpenGL draw calls and operations in order to
optimize performance overhead, rather than tracing each individual operation.
To collect GPU trace, the glQueryCounter() function is used to measure how much
time batches of GPU workload take to complete.
[Link]
User Guide v2021.2.1 | 146
OpenGL Trace
Ranges defined by the KHR_debug calls are represented similarly to OpenGL API and
OpenGL GPU workload trace. GPU ranges in this case represent incremental draw cost.
They cannot fully account for GPUs that can execute multiple draw calls in parallel. In
this case, Nsight Systems will not show overlapping GPU ranges.
[Link]
User Guide v2021.2.1 | 147
OpenGL Trace
[Link]
User Guide v2021.2.1 | 149
Chapter 18.
CUSTOM ETW TRACE
Use the custom ETW trace feature to enable and collect any manifest-based ETW log.
The collected events are displayed on the timeline on dedicated rows for each event
type.
Custom ETW is available on Windows target machines.
[Link]
User Guide v2021.2.1 | 150
Custom ETW Trace
To retain the .etl trace files captured, so that they can be viewed in other tools (e.g.
GPUView), change the "Save ETW log files in project folder" option under "Profile
Behavior" in Nsight Systems's global Options dialog. The .etl files will appear in the
same folder as the .qdrep file, accessible by right-clicking the report in the Project
Explorer and choosing "Show in Folder...". Data collected from each ETW provider will
appear in its own .etl file, and an additional .etl file named "Report XX-Merged-*.etl",
containing the events from all captured sources, will be created as well.
[Link]
User Guide v2021.2.1 | 151
Chapter 19.
GPU METRIC SAMPLING
Requirements
Nsight Systems has support for GPU metrics sampling. These metrics provide an
overview of GPU efficiency over time within compute, graphics, and input/output (IO)
activities such as:
‣ IO throughputs: PCIe, NVLink, and GPU memory bandwidth
‣ SM utilization: SMs activity, tensor core activity, instructions issued, warp
occupancy, and unassigned warp slots
It is designed to help users answer the common questions:
‣ Is my GPU idle?
‣ Is my GPU full? Enough kernel grids size and streams? Are my SMs and warp slots
full?
‣ Am I using TensorCores?
‣ Is my instruction rate high?
‣ Am I possibly blocked on IO, or number of warps, etc
Nsight Systems GPU metric sampling is only available for Linux targets on x86-64 and
for Windows targets. It requires NVIDIA Turing architecture or newer with minimum
driver version r460.
Note:Elevated permissions are required. On Linux use sudo to elevate privileges.
On Windows the user must run from an admin command prompt or accept the
UAC escalation dialog. See Permissions Issues and Performance Counters for more
information.
[Link]
User Guide v2021.2.1 | 152
GPU Metric Sampling
Select the GPUs dropdown to pick which GPUs you wish to sample.
Select the Metric set: dropdown to choose which available metric set you would like to
sample.
Note that metric sets for GPUs that are not being sampled will be greyed out.
309277039|80
309301295|99
309325583|99
309349776|99
309373872|60
309397872|19
309421840|100
309446000|100
309470096|100
309494161|99
[Link]
User Guide v2021.2.1 | 153
GPU Metric Sampling
Limitations
‣ If metrics sets with NVLink are used but the links are not active, they may appear as
fully utilized.
‣ Only one tool that subscribes to these counters can be used at a time, therefore,
Nsight Systems GPU metric sampling cannot be used at the same time as the
following tools:
‣ Nsight Graphics
‣ Nsight Compute
‣ DCGM (Data Center GPU Manager)
Use the following command:
‣ dcgmi profile --pause
‣ dcgmi profile --resume
Or API:
‣ dcgmProfPause
‣ dcgmProfResume
‣ Non-NVIDIA products which use:
‣CUPTI sampling used directly in the application. CUPTI trace is okay
(although it will block Nsight Systems CUDA trace)
‣ DCGM library
‣ Nsight Systems limits the amount of memory that can be used to store GPU metrics
sampling data. Analysis with higher sampling rates or on GPUs with more SMs has
a risk of filling these buffers. This will lead to gaps with long samples on timeline.
If you select that area on the timeline you will see that the counters will pause and
[Link]
User Guide v2021.2.1 | 154
GPU Metric Sampling
remain at a steady state for a while. Future releases will reduce the frequency of this
happening and better present these periods.
[Link]
User Guide v2021.2.1 | 155
Chapter 20.
DEBUG VERSIONS OF ELF FILES
Often, after a binary is built, especially if it is built with debug information (-g compiler
flag), it gets stripped before deploying or installing. In this case, ELF sections that
contain useful information, such as non-export function names or unwind information,
can get stripped as well.
One solution is to deploy or install the original unstripped library instead of the stripped
one, but in many cases this would be inconvenient. Nsight Systems can use missing
information from alternative locations.
For target devices with Ubuntu, see Debug Symbol Packages. These packages typically
install debug ELF files with /usr/lib/debug prefix. Nsight Systems can find debug
libraries there, and if it matches the original library (e.g., the built-in BuildID is the
same), it will be picked up and used to provide symbol names and unwind information.
Many packages have debug companions in the same repository and can be directly
installed with APT (apt-get). Look for packages with the -dbg suffix. For other
packages, refer to the Debug Symbol Packages wiki page on how to add the debs
package repository. After setting up the repository and running apt-get update, look for
packages with -dbgsym suffix.
To verify that a debug version of a library has been picked up and downloaded from the
target device, look in the Module Summary section of Analysis Summary:
[Link]
User Guide v2021.2.1 | 156
Chapter 21.
READING YOUR REPORT IN GUI
[Link]
User Guide v2021.2.1 | 157
Reading Your Report in GUI
21.6.1. Timeline
Timeline is a versatile control that contains a tree-like hierarchy on the left, and
corresponding charts on the right.
Contents of the hierarchy depend on the project settings used to collect the report. For
example, if a certain feature has not been enabled, corresponding rows will not be show
on the timeline.
To display trace events in the Events View right-click a timeline row and select the
“Show in Events View” command. The events of the selected row and all of its sub-rows
will be displayed in the Events View.
[Link]
User Guide v2021.2.1 | 158
Reading Your Report in GUI
If a timeline row has been selected for display in the Events View then double-clicking
a timeline item on that row will automatically scroll the content of the Events View to
make the corresponding Events View item visible and select it.
Row Height
Several of the rows in the timeline use height as a way to model the percent utilization
of resources. This gives the user insight into what is going on even when the timeline is
zoomed all the way out.
In this picture you see that for kernel occupation there is a colored bar of variable height.
Nsight Systems calculates the average occupancy for the period of time represented by
particular pixel width of screen. It then uses that average to set the top of the colored
section. So, for instance, if 25% of that timeslice the kernel is active, the bar goes 25% of
the distance to the top of the row.
In order to make the difference clear, if the percentage of the row height is non-zero, but
would be represented by less than one vertical pixel, Nsight Systems displays it as one
pixel high. The gray height represents the maximum usage in that time range.
This row height coding is used in the CPU utilization, thread and process occupancy,
kernel occupancy, and memory transfer activity rows.
[Link]
User Guide v2021.2.1 | 159
Reading Your Report in GUI
‣ NVTX
‣ Vulkan VK_EXT_debug_marker markers, VK_EXT_debug_utils labels
‣ PIX events and markers
‣ OpenGL KHR_debug markers
[Link]
User Guide v2021.2.1 | 160
Reading Your Report in GUI
‣ To navigate the call tree of the application and while generally searching for
algorithms and parts of the code that consume unexpectedly large amount of CPU
time, the Top-Down view should be used.
‣ To quickly assess which parts of the application, or high level parts of an algorithm,
consume significant amount of CPU time, use the Flat view.
The Top-Down and Bottom-Up views have Self and Total columns, while the Flat view
has a Flat column. It is important to understand the meaning of each of the columns:
‣ Top-Down view
‣ Self column denotes the relative amount of time spent executing instructions of
this particular function.
‣ Total column shows how much time has been spent executing this function,
including all other functions called from this one. Total values of sibling rows
sum up to the Total value of the parent row, or 100% for the top-level rows.
‣ Bottom-Up view
‣ Self column for top-level rows, as in the Top-Down view, shows how much time
has been spent directly in this function. Self times of all top-level rows add up to
100%.
‣ Self column for children rows breaks down the value of the parent row based on
the various call chains leading to that function. Self times of sibling rows add up
to the value of the parent row.
‣ Flat view
‣ Flat column shows how much time this function has been anywhere on the
call stack. Values in this column do not add up or have other significant
relationships.
If
low-
impact
functions
have
been
filtered
out,
values
may
not
Note: add
up
correctly
to
100%,
or
to
the
value
of
the
parent
row.
[Link]
User Guide v2021.2.1 | 161
Reading Your Report in GUI
This
filtering
can
be
disabled.
Contents of the symbols table is tightly related to the timeline. Users can apply and
modify filters on the timeline, and they will affect which information is displayed in
the symbols table:
‣ Per-thread filtering — Each thread that has sampling information associated with it
has a checkbox next to it on the timeline. Only threads with selected checkboxes are
represented in the symbols table.
‣ Time filtering — A time filter can be setup on the timeline by pressing the left
mouse button, dragging over a region of interest on the timeline, and then choosing
Filter by selection in the dropdown menu. In this case, only sampling information
collected during the selected time range will be used to build the symbols table.
If
too
little
sampling
data
is
being
used
to
build
the
symbols
table
(for
example,
when
the
sampling
rate
Note: is
configured
to
be
low,
and
a
short
period
of
time
is
used
for
time-
based
filtering),
the
numbers
in
[Link]
User Guide v2021.2.1 | 162
Reading Your Report in GUI
the
symbols
table
might
not
be
representative
or
accurate
in
some
cases.
‣ Collapse unresolved lines is useful if some of the binary code does not have
symbols. In this case, subtrees that consist of only unresolved symbols get collapsed
in the Top-Down view, since they provide very little useful information.
‣ Hide functions with CPU usage below X% is useful for large applications, where
the sampling profiler hits lots of function just a few times. To filter out the "long
tail," which is typically not important for CPU performance bottleneck analysis, this
checkbox should be selected.
[Link]
User Guide v2021.2.1 | 163
Reading Your Report in GUI
Information from this view can be selected and copied using the mouse cursor.
[Link]
User Guide v2021.2.1 | 164
Chapter 22.
BROKEN BACKTRACES ON TEGRA
In Nsight Systems Embedded Platforms Edition, in the symbols table there is a special
entry called Broken backtraces. This entry is used to denote the point in the call chain
where the unwinding algorithms used by Nsight Systems could not determine what is
the next (caller) function.
Broken backtraces happen because there is no information related to the current function
that the unwinding algorithms can use. In the Top-Down view, these functions are
immediate children of the Broken backtraces row.
One can eliminate broken backtraces by modifying the build system to provide at
least one kind of unwind information. The types of unwind information, used by the
algorithms in Nsight Systems, include the following:
For ARMv7 binaries:
‣ DWARF information in ELF sections: .debug_frame, .zdebug_frame, .eh_frame,
.eh_frame_hdr. This information is the most precise. .zdebug_frame is a
compressed version of .debug_frame, so at most one of them is typically present.
.eh_frame_hdr is a companion section for .eh_frame and might be absent.
Compiler flag: -g.
‣ Exception handling information in EHABI format provided in .[Link] and
.[Link] ELF sections. .[Link] might be absent if all information is
compact enough to be encoded into .[Link].
Compiler flag: -funwind-tables.
‣ Frame pointers (built into the .text section).
Compiler flag: -fno-omit-frame-pointer.
For Aarch64 binaries:
‣ DWARF information in ELF sections: .debug_frame, .zdebug_frame, .eh_frame,
.eh_frame_hdr. See additional comments above.
Compiler flag: -g.
‣ Frame pointers (built into the .text section).
Compiler flag: -fno-omit-frame-pointer.
[Link]
User Guide v2021.2.1 | 165
Broken Backtraces on Tegra
The following ELF sections should be considered empty if they have size of 4 bytes:
.debug_frame, .eh_frame, .[Link]. In this case, these sections only contain
termination records and no useful information.
For GCC, use the following compiler invocation to see which compiler flags are enabled
in your toolchain by default (for example, to check if -funwind-tables is enabled by
default):
$ gcc -Q --help=common
For GCC and Clang, add -### to the compiler invocation command to see which
compiler flags are actually being used.
Since EHABI and DWARF information is compiled on per-unit basis (every .cpp or
.c file, as well as every static library, can be built with or without this information),
presence of the ELF sections does not guarantee that every function has necessary
unwind information.
Frame pointers are required by the Aarch64 Procedure Call Standard. Adding frame
pointers slows down execution time, but in most cases the difference is negligible.
[Link]
User Guide v2021.2.1 | 166
Chapter 23.
LAUNCH PROCESSES IN STOPPED STATE
In many cases, it is important to profile an application from the very beginning of its
execution. When launching processes, Nsight Systems takes care of it by making sure
that the profiling session is fully initialized before making the exec() system call on
Linux, and by using the JDWP protocol on Android.
If the process launch capabilities of Nsight Systems are not sufficient, the application
should be launched manually, and the profiler should be configured to attach to the
already launched process. One approach would be to call sleep() somewhere early in
the application code, which would provide time for the user to attach to the process in
Nsight Systems Embedded Platforms Edition, but there are two other more convenient
mechanisms that can be used on Linux, without the need to recompile the application.
(Note that the rest of this section is only applicable to Linux-based target devices, not
Windows or Android.)
Both mechanisms ensure that between the time the process is created (and therefore its
PID is known) and the time any of the application's code is called, the process is stopped
and waits for a signal to be delivered before continuing.
23.1. LD_PRELOAD
The first mechanism uses LD_PRELOAD environment variable. It only works with
dynamically linked binaries, since static binaries do not invoke the runtime linker, and
therefore are not affected by the LD_PRELOAD environment variable.
‣ For ARMv7 binaries, preload
/opt/nvidia/nsight_systems/[Link]
‣ Otherwise if running from host, preload
/opt/nvidia/nsight_systems/[Link]
‣ Otherwise if running from CLI, preload
[installation_directory]/[Link]
The most common way to do that is to specify the environment variable as part of the
process launch command, for example:
$ LD_PRELOAD=/opt/nvidia/nsight_systems/[Link] ./my-aarch64-binary --
arguments
[Link]
User Guide v2021.2.1 | 167
Launch Processes in Stopped State
When loaded, this library will send itself a SIGSTOP signal, which is equivalent to typing
Ctrl+Z in the terminal. The process is now a background job, and you can use standard
commands like jobs, fg and bg to control them. Use jobs -l to see the PID of the
launched process.
When attaching to a stopped process, Nsight Systems will send SIGCONT signal, which is
equivalent to using the bg command.
23.2. Launcher
The second mechanism can be used with any binary. Use
[installation_directory]/launcher to launch your application, for example:
$ /opt/nvidia/nsight_systems/launcher ./my-binary --arguments
The process will be launched, daemonized, and wait for SIGUSR1 signal. After attaching
to the process with Nsight Systems, the user needs to manually resume execution of the
process from command line:
$ pkill -USR1 launcher
Note
that
pkill
will
send
the
signal
to
any
process
with
the
matching
name.
If
that
is
Note: not
desirable,
use
kill
to
send
it
to
a
specific
process.
The
standard
output
and
error
streams
are
redirected
[Link]
User Guide v2021.2.1 | 168
Launch Processes in Stopped State
to
/
tmp/
stdout_<PID>.tx
and
/
tmp/
stderr_<PID>.tx
The launcher mechanism is more complex and less automated than the LD_PRELOAD
option, but gives more control to the user.
[Link]
User Guide v2021.2.1 | 169
Chapter 24.
IMPORT NVTXT
Modes description:
‣ lerp - Insert with linear interpolation
--mode lerp --ns_a arg --ns_b arg [--nvtxt_a arg --nvtxt_b arg]
‣ lin - insert with linear equation
--mode lin --ns_a arg --freq arg [--nvtxt_a arg]
Modes' parameters:
[Link]
User Guide v2021.2.1 | 170
Import NVTXT
Commands
Info
To find out report's start and end time use info command.
Usage:
ImportNvtxt --cmd info -i [--input] arg
Example:
ImportNvtxt info [Link]
Analysis start (ns) 83501026500000
Analysis end (ns) 83506375000000
Create
You can create a report file using existing NVTXT with create command.
Usage:
ImportNvtxt --cmd create -n [--nvtxt] arg -o [--output] arg [-m [--mode]
mode_name mode_args]
with:
‣ ns_a — a nanoseconds value.
‣ ns_b — a nanoseconds value (greater than ns_a).
‣ nvtxt_a — an nvtxt file's time unit value corresponding to ns_a nanoseconds.
‣ nvtxt_b — an nvtxt file's time unit value corresponding to ns_b nanoseconds.
If nvtxt_a and nvtxt_b are not specified, they are repectively set to nvtxt file's
minimum and maximum time value.
Usage for lin mode is:
--mode lin --ns_a arg --freq arg [--nvtxt_a arg]
with:
[Link]
User Guide v2021.2.1 | 171
Import NVTXT
The output will be a new generated report file which can be opened and viewed by
Nsight Systems.
Merge
To merge NVTXT file with an existing report file use merge command.
Usage:
ImportNvtxt --cmd merge -i [--input] arg -n [--nvtxt] arg -o [--output] arg [-m
[--mode] mode_name mode_args]
with:
‣ ns_a — a nanoseconds value.
‣ ns_b — a nanoseconds value (greater than ns_a).
‣ nvtxt_a — an nvtxt file's time unit value corresponding to ns_a nanoseconds.
‣ nvtxt_b — an nvtxt file's time unit value corresponding to ns_b nanoseconds.
If nvtxt_a and nvtxt_b are not specified, they are repectively set to nvtxt file's
minimum and maximum time value.
Usage for lin mode is:
--mode lin --ns_a arg --freq arg [--nvtxt_a arg]
with:
‣ ns_a — a nanoseconds value.
‣ freq — the nvtxt file's timer frequency.
‣ nvtxt_a — an nvtxt file's time unit value corresponding to ns_a nanoseconds.
If nvtxt_a is not specified, it is set to nvtxt file's minimum time value.
Time values in <[Link]> are assumed to be nanoseconds if no mode
specified.
Example
ImportNvtxt --cmd merge -i [Link] -n [Link] -o [Link]
[Link]
User Guide v2021.2.1 | 172
Chapter 25.
VISUAL STUDIO INTEGRATION
NVIDIA Nsight Integration is a Visual Studio extension that allows you to access the
power of Nsight Systems from within Visual Studio.
When Nsight Systems is installed along with NVIDIA Nsight Integration, Nsight
Systems activities will appear under the NVIDIA Nsight menu in the Visual Studio
menu bar. These activities launch Nsight Systems with the current project settings and
executable.
Selecting the "Trace" command will launch Nsight Systems, create a new Nsight Systems
project and apply settings from the current Visual Studio project:
‣ Target application path
‣ Command line parameters
‣ Working folder
If the "Trace" command has already been used with this Visual Studio project then
Nsight Systems will load the respective Nsight Systems project and any previously
captured trace sessions will be available for review using the Nsight Systems project
explorer tree.
[Link]
User Guide v2021.2.1 | 173
Visual Studio Integration
For more information about using Nsight Systems from within Visual Studio, please
visit
‣ NVIDIA Nsight Integration Overview
‣ NVIDIA Nsight Integration User Guide
[Link]
User Guide v2021.2.1 | 174
Chapter 26.
TROUBLESHOOTING
If the profiler behaves unexpectedly during the profiling session, or the profiling session
fails to start, try the following steps:
‣ Close the host application.
‣ Restart the target device.
‣ Start the host application and connect to the target device.
To enable logging on the host, refer to this config file:
host-linux-x64/[Link]
When reporting any bugs please include the build version number as described in
the Help → About dialog. If possible, attach log files and report (.qdrep) files, as they
already contain necessary version information.
Nsight Systems uses a settings file (NVIDIA Nsight [Link]) on the host to
store information about loaded projects, report files, window layout configuration,
etc. Location of the settings file is described in the Help → About dialog. Deleting the
settings file will restore Nsight Systems to a fresh state, but all projects and reports will
disappear from the Project Explorer.
GUI Troubleshooting
If opening the Nsight Systems Linux GUI fails with the following error, you may be
missing some required libraries:
This application failed to start because it could not find or load the Qt
platform plugin "xcb" in "". Available platform plugins are: xcb. Reinstalling
the application may fix this problem.
Launch Nsight Systems using the following command line to determine which libraries
are missing and install them.
$ QT_DEBUG_PLUGINS=1 ./nsys-ui
If the workload does not run when launched via Nsight Systems or the timeline is
empty, check the [Link] and [Link] (click on drop-down menu showing Timeline
View and click on Files) to see the errors encountered by the app.
[Link]
User Guide v2021.2.1 | 175
Troubleshooting
Android Targets
When connecting to an Android-based device, Nsight Systems installs its executable and
library files into the following directory:
/data/local/tmp/[Link]/
On rooted Android devices, the command above should be started from superuser
(e.g., adb shell su -c ...).
5. Start the host application and connect to the target device.
Please note that in some cases, debug logging can significantly slow down the profiler
Symbol Resolution
If stack trace information is missing symbols and you have a symbol file, you can
manually re-resolve using the ResolveSymbols utility. This can be done by right-clicking
the report file in the Project Explorer window and selecting "Resolve Symbols...".
Alternatively, you can find the utility as a separate executable in the
[installation_path]\Host directory. This utility works with ELF format files, with
Windows PDB directories and symbol servers, or with files where each line is in the
format <start><length><name>.
[Link]
User Guide v2021.2.1 | 176
Troubleshooting
[Link]
User Guide v2021.2.1 | 177
Troubleshooting
To enable verbose logging on the target device, when launched from the host, follow
these steps:
1. Close the host application.
2. Restart the target device.
3. Place [Link] from host directory to the /opt/nvidia/nsight_systems
directory on target.
4. From SSH console, launch the following command:
sudo /opt/nvidia/nsight_systems/nsys --daemon --debug
5. Start the host application and connect to the target device.
Logs on the target devices are collected into this file (if enabled):
[Link]
To enable verbose logging on the target device, when launched from the host, follow
these steps:
1. Close the host application.
2. Terminate the nsys process.
3. Place [Link] from host directory next to Nsight Systems Windows agent on
the target device
‣ Local Windows target:
C:\Program Files\NVIDIA Corporation\Nsight Systems 2021.2\target-
windows-x64
[Link]
User Guide v2021.2.1 | 178
Troubleshooting
QNX Troubleshooting
Common issues with QNX targets:
‣ Make sure that tracelogger utility is available and can be run on the target.
‣ Make sure that /tmp directory is accessible and supports sub-directories.
‣ When switching between Nsight Systems versions, processes related to the previous
version, including profiled applications forked by the daemon, must be killed before
the new version is used. If you experience issues after switching between Nsight
Systems versions, try rebooting the target.
[Link]
User Guide v2021.2.1 | 179
Chapter 27.
OTHER RESOURCES
Looking for information to help you use Nsight Systems the most effectively? Here are
some more resources you might want to review:
Feature Videos
Short videos, only a minute or two, to introduce new features.
‣ OpenMP Trace Feature Spotlight
‣ Command Line Sessions Video Spotlight
‣ Direct3D11 Feature Spotlight
‣ Vulkan Trace
‣ Statistics Driven Profiling
Blog Posts
NVIDIA developer blogs, these are longer form, technical pieces written by tool and
domain experts.
‣ 2019 - Migrating to NVIDIA Nsight Tools from NVVP and nvprof
‣ 2019 - Transitioning to Nsight Systems from NVIDIA Visual Profiler / nvprof
‣ 2019 - NVIDIA Nsight Systems Add Vulkan Support
‣ 2019 - TensorFlow Performance Logging Plugin nvtx-plugins-tf Goes Public
‣ 2020 - NVIDIA Nsight Systems in Containers and the Cloud
‣ 2020 - Understanding the Visualization of Overhead and Latency in Nsight Systems
Training Seminars
2018 NCSA Blue Waters Webinar - Introduction to NVIDIA Nsight Systems
[Link]
User Guide v2021.2.1 | 180
Other Resources
Conference Presentations
‣ GTC 2020 - Rebalancing the Load: Profile-Guided Optimization of the NAMD
Molecular Dynamics Program for Modern GPUs using Nsight Systems
‣ GTC 2020 - Scaling the Transformer Model Implementation in PyTorch Across
Multiple Nodes
‣ GTC 2019 - Using Nsight Tools to Optimize the NAMD Molecular Dynamics
Simulation Program
‣ GTC 2019 - Optimizing Facebook AI Workloads for NVIDIA GPUs
‣ GTC 2018 - Optimizing HPC Simulation and Visualization Codes Using NVIDIA
Nsight Systems
‣ GTC 2018 - Israel - Boost DNN Training Performance using NVIDIA Tools
‣ Siggraph 2018 - Taming the Beast; Using NVIDIA Tools to Unlock Hidden GPU
Performance
[Link]
User Guide v2021.2.1 | 181
Other Resources
[Link]
User Guide v2021.2.1 | 182