CUDA Math API Reference Guide
CUDA Math API Reference Guide
Chapter 1. Modules.............................................................................................. 1
1.1. Half Precision Intrinsics..................................................................................2
Half Arithmetic Functions.................................................................................. 2
Half2 Arithmetic Functions................................................................................ 2
Half Comparison Functions.................................................................................2
Half2 Comparison Functions............................................................................... 2
Half Precision Conversion And Data Movement.........................................................2
Half Math Functions......................................................................................... 2
Half2 Math Functions....................................................................................... 2
1.1.1. Half Arithmetic Functions..........................................................................2
__habs...................................................................................................... 2
__hadd...................................................................................................... 3
__hadd_sat................................................................................................. 3
__hdiv....................................................................................................... 3
__hfma...................................................................................................... 4
__hfma_relu................................................................................................4
__hfma_sat................................................................................................. 4
__hmul...................................................................................................... 5
__hmul_sat................................................................................................. 5
__hneg...................................................................................................... 6
__hsub...................................................................................................... 6
__hsub_sat................................................................................................. 6
1.1.2. Half2 Arithmetic Functions........................................................................ 6
__h2div..................................................................................................... 7
__habs2..................................................................................................... 7
__hadd2.....................................................................................................7
__hadd2_sat................................................................................................7
__hfma2.................................................................................................... 8
__hfma2_relu.............................................................................................. 8
__hfma2_sat............................................................................................... 9
__hmul2.....................................................................................................9
__hmul2_sat.............................................................................................. 10
__hneg2................................................................................................... 10
__hsub2....................................................................................................10
__hsub2_sat...............................................................................................10
1.1.3. Half Comparison Functions....................................................................... 11
__heq...................................................................................................... 11
__hequ.....................................................................................................12
__hge...................................................................................................... 12
__hgeu.....................................................................................................13
[Link]
CUDA Math API vRelease Version | ii
__hgt.......................................................................................................13
__hgtu..................................................................................................... 14
__hisinf.................................................................................................... 14
__hisnan................................................................................................... 15
__hle....................................................................................................... 15
__hleu..................................................................................................... 16
__hlt....................................................................................................... 16
__hltu......................................................................................................17
__hmax.................................................................................................... 17
__hmax_nan.............................................................................................. 17
__hmin.....................................................................................................18
__hmin_nan...............................................................................................18
__hne...................................................................................................... 18
__hneu.....................................................................................................19
1.1.4. Half2 Comparison Functions......................................................................19
__hbeq2................................................................................................... 19
__hbequ2..................................................................................................20
__hbge2................................................................................................... 20
__hbgeu2.................................................................................................. 21
__hbgt2....................................................................................................21
__hbgtu2.................................................................................................. 22
__hble2.................................................................................................... 23
__hbleu2.................................................................................................. 23
__hblt2.................................................................................................... 24
__hbltu2................................................................................................... 24
__hbne2................................................................................................... 25
__hbneu2..................................................................................................25
__heq2.....................................................................................................26
__hequ2................................................................................................... 26
__hge2..................................................................................................... 27
__hgeu2................................................................................................... 27
__hgt2..................................................................................................... 28
__hgtu2....................................................................................................28
__hisnan2................................................................................................. 29
__hle2..................................................................................................... 29
__hleu2.................................................................................................... 30
__hlt2...................................................................................................... 30
__hltu2.................................................................................................... 31
__hmax2...................................................................................................31
__hmax2_nan............................................................................................. 31
__hmin2................................................................................................... 32
__hmin2_nan............................................................................................. 32
__hne2.....................................................................................................32
[Link]
CUDA Math API vRelease Version | iii
__hneu2................................................................................................... 33
1.1.5. Half Precision Conversion And Data Movement............................................... 33
__double2half............................................................................................ 33
__float22half2_rn........................................................................................ 34
__float2half...............................................................................................34
__float2half2_rn......................................................................................... 35
__float2half_rd...........................................................................................35
__float2half_rn...........................................................................................36
__float2half_ru...........................................................................................36
__float2half_rz........................................................................................... 37
__floats2half2_rn........................................................................................ 37
__half22float2............................................................................................ 38
__half2float...............................................................................................38
__half2half2.............................................................................................. 38
__half2int_rd............................................................................................. 39
__half2int_rn............................................................................................. 39
__half2int_ru............................................................................................. 40
__half2int_rz............................................................................................. 40
__half2ll_rd............................................................................................... 41
__half2ll_rn............................................................................................... 41
__half2ll_ru............................................................................................... 42
__half2ll_rz............................................................................................... 42
__half2short_rd.......................................................................................... 43
__half2short_rn.......................................................................................... 43
__half2short_ru.......................................................................................... 44
__half2short_rz.......................................................................................... 44
__half2uint_rd............................................................................................45
__half2uint_rn............................................................................................45
__half2uint_ru............................................................................................46
__half2uint_rz............................................................................................ 46
__half2ull_rd............................................................................................. 47
__half2ull_rn............................................................................................. 47
__half2ull_ru............................................................................................. 48
__half2ull_rz..............................................................................................48
__half2ushort_rd.........................................................................................49
__half2ushort_rn......................................................................................... 49
__half2ushort_ru......................................................................................... 50
__half2ushort_rz......................................................................................... 50
__half_as_short.......................................................................................... 51
__half_as_ushort......................................................................................... 51
__halves2half2........................................................................................... 51
__high2float.............................................................................................. 52
__high2half............................................................................................... 52
[Link]
CUDA Math API vRelease Version | iv
__high2half2.............................................................................................. 53
__highs2half2............................................................................................. 53
__int2half_rd............................................................................................. 54
__int2half_rn............................................................................................. 54
__int2half_ru............................................................................................. 55
__int2half_rz............................................................................................. 55
__ldca..................................................................................................... 56
__ldca..................................................................................................... 56
__ldcg......................................................................................................56
__ldcg......................................................................................................56
__ldcs...................................................................................................... 57
__ldcs...................................................................................................... 57
__ldcv......................................................................................................57
__ldcv......................................................................................................57
__ldg....................................................................................................... 58
__ldg....................................................................................................... 58
__ldlu...................................................................................................... 58
__ldlu...................................................................................................... 58
__ll2half_rd............................................................................................... 59
__ll2half_rn............................................................................................... 59
__ll2half_ru............................................................................................... 60
__ll2half_rz............................................................................................... 60
__low2float............................................................................................... 61
__low2half................................................................................................ 61
__low2half2...............................................................................................61
__lowhigh2highlow...................................................................................... 62
__lows2half2..............................................................................................62
__shfl_down_sync........................................................................................63
__shfl_down_sync........................................................................................64
__shfl_sync............................................................................................... 64
__shfl_sync............................................................................................... 65
__shfl_up_sync........................................................................................... 66
__shfl_up_sync........................................................................................... 66
__shfl_xor_sync.......................................................................................... 67
__shfl_xor_sync.......................................................................................... 68
__short2half_rd.......................................................................................... 69
__short2half_rn.......................................................................................... 69
__short2half_ru.......................................................................................... 70
__short2half_rz.......................................................................................... 70
__short_as_half.......................................................................................... 71
__stcg......................................................................................................71
__stcg......................................................................................................71
__stcs...................................................................................................... 72
[Link]
CUDA Math API vRelease Version | v
__stcs...................................................................................................... 72
__stwb..................................................................................................... 72
__stwb..................................................................................................... 72
__stwt..................................................................................................... 73
__stwt..................................................................................................... 73
__uint2half_rd............................................................................................73
__uint2half_rn............................................................................................74
__uint2half_ru............................................................................................74
__uint2half_rz............................................................................................ 75
__ull2half_rd............................................................................................. 75
__ull2half_rn............................................................................................. 76
__ull2half_ru............................................................................................. 76
__ull2half_rz..............................................................................................77
__ushort2half_rd.........................................................................................77
__ushort2half_rn......................................................................................... 78
__ushort2half_ru......................................................................................... 78
__ushort2half_rz......................................................................................... 79
__ushort_as_half......................................................................................... 79
1.1.6. Half Math Functions............................................................................... 79
hceil........................................................................................................80
hcos........................................................................................................ 80
hexp........................................................................................................80
hexp10.....................................................................................................81
hexp2...................................................................................................... 81
hfloor...................................................................................................... 82
hlog........................................................................................................ 82
hlog10..................................................................................................... 82
hlog2....................................................................................................... 83
hrcp........................................................................................................ 83
hrint........................................................................................................84
hrsqrt...................................................................................................... 84
hsin.........................................................................................................84
hsqrt....................................................................................................... 85
htrunc......................................................................................................85
1.1.7. Half2 Math Functions..............................................................................86
h2ceil...................................................................................................... 86
h2cos.......................................................................................................86
h2exp...................................................................................................... 87
h2exp10................................................................................................... 87
h2exp2.....................................................................................................87
h2floor.....................................................................................................88
h2log....................................................................................................... 88
h2log10.................................................................................................... 89
[Link]
CUDA Math API vRelease Version | vi
h2log2..................................................................................................... 89
h2rcp.......................................................................................................90
h2rint...................................................................................................... 90
h2rsqrt.....................................................................................................90
h2sin....................................................................................................... 91
h2sqrt...................................................................................................... 91
h2trunc.................................................................................................... 92
1.2. Bfloat16 Precision Intrinsics........................................................................... 92
Bfloat16 Arithmetic Functions........................................................................... 92
Bfloat162 Arithmetic Functions.......................................................................... 92
Bfloat16 Comparison Functions.......................................................................... 92
Bfloat162 Comparison Functions.........................................................................92
Bfloat16 Precision Conversion And Data Movement.................................................. 92
Bfloat16 Math Functions.................................................................................. 92
Bfloat162 Math Functions................................................................................. 93
1.2.1. Bfloat16 Arithmetic Functions................................................................... 93
__h2div.................................................................................................... 93
__habs..................................................................................................... 93
__hadd.....................................................................................................93
__hadd_sat................................................................................................94
__hdiv..................................................................................................... 94
__hfma.................................................................................................... 94
__hfma_relu.............................................................................................. 95
__hfma_sat............................................................................................... 95
__hmul.....................................................................................................96
__hmul_sat................................................................................................96
__hneg.....................................................................................................96
__hsub..................................................................................................... 97
__hsub_sat................................................................................................ 97
1.2.2. Bfloat162 Arithmetic Functions..................................................................97
__habs2....................................................................................................97
__hadd2................................................................................................... 98
__hadd2_sat.............................................................................................. 98
__hfma2................................................................................................... 99
__hfma2_relu.............................................................................................99
__hfma2_sat.............................................................................................100
__hmul2..................................................................................................100
__hmul2_sat.............................................................................................101
__hneg2.................................................................................................. 101
__hsub2.................................................................................................. 101
__hsub2_sat............................................................................................. 102
1.2.3. Bfloat16 Comparison Functions................................................................ 102
__heq.....................................................................................................102
[Link]
CUDA Math API vRelease Version | vii
__hequ................................................................................................... 103
__hge..................................................................................................... 103
__hgeu................................................................................................... 104
__hgt..................................................................................................... 104
__hgtu....................................................................................................105
__hisinf...................................................................................................105
__hisnan................................................................................................. 106
__hle..................................................................................................... 106
__hleu.................................................................................................... 107
__hlt...................................................................................................... 107
__hltu.................................................................................................... 108
__hmax...................................................................................................108
__hmax_nan............................................................................................. 109
__hmin................................................................................................... 109
__hmin_nan............................................................................................. 109
__hne.....................................................................................................109
__hneu................................................................................................... 110
1.2.4. Bfloat162 Comparison Functions............................................................... 110
__hbeq2..................................................................................................111
__hbequ2................................................................................................ 111
__hbge2.................................................................................................. 112
__hbgeu2................................................................................................ 112
__hbgt2.................................................................................................. 113
__hbgtu2................................................................................................. 114
__hble2.................................................................................................. 114
__hbleu2................................................................................................. 115
__hblt2................................................................................................... 115
__hbltu2................................................................................................. 116
__hbne2..................................................................................................117
__hbneu2................................................................................................ 117
__heq2................................................................................................... 118
__hequ2..................................................................................................118
__hge2................................................................................................... 119
__hgeu2.................................................................................................. 119
__hgt2.................................................................................................... 120
__hgtu2.................................................................................................. 121
__hisnan2................................................................................................ 121
__hle2.................................................................................................... 122
__hleu2.................................................................................................. 122
__hlt2.................................................................................................... 123
__hltu2................................................................................................... 123
__hmax2................................................................................................. 124
__hmax2_nan........................................................................................... 124
[Link]
CUDA Math API vRelease Version | viii
__hmin2..................................................................................................124
__hmin2_nan............................................................................................ 125
__hne2................................................................................................... 125
__hneu2.................................................................................................. 125
1.2.5. Bfloat16 Precision Conversion And Data Movement.........................................126
__bfloat1622float2..................................................................................... 126
__bfloat162bfloat162.................................................................................. 127
__bfloat162float........................................................................................ 127
__bfloat162int_rd...................................................................................... 128
__bfloat162int_rn...................................................................................... 128
__bfloat162int_ru...................................................................................... 129
__bfloat162int_rz...................................................................................... 129
__bfloat162ll_rd........................................................................................ 130
__bfloat162ll_rn........................................................................................ 130
__bfloat162ll_ru........................................................................................ 131
__bfloat162ll_rz........................................................................................ 131
__bfloat162short_rd................................................................................... 132
__bfloat162short_rn................................................................................... 132
__bfloat162short_ru................................................................................... 133
__bfloat162short_rz....................................................................................133
__bfloat162uint_rd.....................................................................................134
__bfloat162uint_rn..................................................................................... 134
__bfloat162uint_ru..................................................................................... 135
__bfloat162uint_rz..................................................................................... 135
__bfloat162ull_rd...................................................................................... 136
__bfloat162ull_rn.......................................................................................136
__bfloat162ull_ru.......................................................................................137
__bfloat162ull_rz.......................................................................................137
__bfloat162ushort_rd.................................................................................. 138
__bfloat162ushort_rn.................................................................................. 138
__bfloat162ushort_ru.................................................................................. 139
__bfloat162ushort_rz.................................................................................. 139
__bfloat16_as_short................................................................................... 140
__bfloat16_as_ushort.................................................................................. 140
__double2bfloat16..................................................................................... 141
__float22bfloat162_rn................................................................................. 141
__float2bfloat16........................................................................................ 142
__float2bfloat162_rn.................................................................................. 142
__float2bfloat16_rd.................................................................................... 143
__float2bfloat16_rn.................................................................................... 143
__float2bfloat16_ru.................................................................................... 144
__float2bfloat16_rz.................................................................................... 144
__floats2bfloat162_rn................................................................................. 145
[Link]
CUDA Math API vRelease Version | ix
__halves2bfloat162.....................................................................................145
__high2bfloat16........................................................................................ 146
__high2bfloat162....................................................................................... 146
__high2float............................................................................................. 147
__highs2bfloat162...................................................................................... 147
__int2bfloat16_rd...................................................................................... 148
__int2bfloat16_rn...................................................................................... 148
__int2bfloat16_ru...................................................................................... 149
__int2bfloat16_rz...................................................................................... 149
__ldca.................................................................................................... 150
__ldca.................................................................................................... 150
__ldcg.................................................................................................... 150
__ldcg.................................................................................................... 150
__ldcs.................................................................................................... 151
__ldcs.................................................................................................... 151
__ldcv.................................................................................................... 151
__ldcv.................................................................................................... 151
__ldg..................................................................................................... 152
__ldg..................................................................................................... 152
__ldlu.....................................................................................................152
__ldlu.....................................................................................................152
__ll2bfloat16_rd........................................................................................ 153
__ll2bfloat16_rn........................................................................................ 153
__ll2bfloat16_ru........................................................................................ 154
__ll2bfloat16_rz........................................................................................ 154
__low2bfloat16......................................................................................... 155
__low2bfloat162........................................................................................ 155
__low2float..............................................................................................156
__lowhigh2highlow..................................................................................... 156
__lows2bfloat162.......................................................................................157
__shfl_down_sync...................................................................................... 157
__shfl_down_sync...................................................................................... 158
__shfl_sync.............................................................................................. 159
__shfl_sync.............................................................................................. 159
__shfl_up_sync..........................................................................................160
__shfl_up_sync..........................................................................................161
__shfl_xor_sync.........................................................................................161
__shfl_xor_sync.........................................................................................162
__short2bfloat16_rd................................................................................... 163
__short2bfloat16_rn................................................................................... 163
__short2bfloat16_ru................................................................................... 164
__short2bfloat16_rz....................................................................................164
__short_as_bfloat16................................................................................... 165
[Link]
CUDA Math API vRelease Version | x
__stcg.................................................................................................... 165
__stcg.................................................................................................... 165
__stcs.....................................................................................................166
__stcs.....................................................................................................166
__stwb................................................................................................... 166
__stwb................................................................................................... 166
__stwt.................................................................................................... 167
__stwt.................................................................................................... 167
__uint2bfloat16_rd.....................................................................................167
__uint2bfloat16_rn..................................................................................... 168
__uint2bfloat16_ru..................................................................................... 168
__uint2bfloat16_rz..................................................................................... 169
__ull2bfloat16_rd...................................................................................... 169
__ull2bfloat16_rn.......................................................................................170
__ull2bfloat16_ru.......................................................................................170
__ull2bfloat16_rz.......................................................................................171
__ushort2bfloat16_rd.................................................................................. 171
__ushort2bfloat16_rn.................................................................................. 172
__ushort2bfloat16_ru.................................................................................. 172
__ushort2bfloat16_rz.................................................................................. 173
__ushort_as_bfloat16.................................................................................. 173
1.2.6. Bfloat16 Math Functions.........................................................................173
hceil...................................................................................................... 174
hcos.......................................................................................................174
hexp...................................................................................................... 174
hexp10................................................................................................... 175
hexp2.....................................................................................................175
hfloor.....................................................................................................176
hlog....................................................................................................... 176
hlog10.................................................................................................... 177
hlog2..................................................................................................... 177
hrcp.......................................................................................................177
hrint...................................................................................................... 178
hrsqrt.....................................................................................................178
hsin....................................................................................................... 179
hsqrt...................................................................................................... 179
htrunc.................................................................................................... 180
1.2.7. Bfloat162 Math Functions....................................................................... 180
h2ceil.....................................................................................................180
h2cos..................................................................................................... 181
h2exp.....................................................................................................181
h2exp10.................................................................................................. 182
h2exp2................................................................................................... 182
[Link]
CUDA Math API vRelease Version | xi
h2floor................................................................................................... 183
h2log..................................................................................................... 183
h2log10...................................................................................................184
h2log2.................................................................................................... 184
h2rcp..................................................................................................... 184
h2rint.....................................................................................................185
h2rsqrt................................................................................................... 185
h2sin......................................................................................................186
h2sqrt.................................................................................................... 186
h2trunc...................................................................................................187
1.3. Mathematical Functions...............................................................................187
1.4. Single Precision Mathematical Functions...........................................................187
acosf.........................................................................................................187
acoshf....................................................................................................... 188
asinf......................................................................................................... 188
asinhf........................................................................................................189
atan2f....................................................................................................... 189
atanf.........................................................................................................189
atanhf....................................................................................................... 190
cbrtf......................................................................................................... 190
ceilf..........................................................................................................191
copysignf....................................................................................................191
cosf.......................................................................................................... 191
coshf.........................................................................................................192
cospif........................................................................................................192
cyl_bessel_i0f.............................................................................................. 192
cyl_bessel_i1f.............................................................................................. 193
erfcf......................................................................................................... 193
erfcinvf..................................................................................................... 193
erfcxf........................................................................................................194
erff.......................................................................................................... 194
erfinvf....................................................................................................... 195
exp10f.......................................................................................................195
exp2f........................................................................................................ 196
expf..........................................................................................................196
expm1f...................................................................................................... 196
fabsf......................................................................................................... 197
fdimf........................................................................................................ 197
fdividef......................................................................................................198
floorf........................................................................................................ 198
fmaf......................................................................................................... 198
fmaxf........................................................................................................ 199
fminf........................................................................................................ 199
[Link]
CUDA Math API vRelease Version | xii
fmodf........................................................................................................200
frexpf........................................................................................................200
hypotf....................................................................................................... 201
ilogbf........................................................................................................ 201
isfinite...................................................................................................... 202
isinf.......................................................................................................... 202
isnan.........................................................................................................202
j0f............................................................................................................203
j1f............................................................................................................203
jnf........................................................................................................... 204
ldexpf....................................................................................................... 204
lgammaf.................................................................................................... 205
llrintf........................................................................................................ 205
llroundf..................................................................................................... 205
log10f....................................................................................................... 206
log1pf....................................................................................................... 206
log2f......................................................................................................... 207
logbf......................................................................................................... 207
logf.......................................................................................................... 207
lrintf.........................................................................................................208
lroundf...................................................................................................... 208
modff........................................................................................................208
nanf..........................................................................................................209
nearbyintf.................................................................................................. 209
nextafterf.................................................................................................. 210
norm3df.....................................................................................................210
norm4df.....................................................................................................210
normcdff.................................................................................................... 211
normcdfinvf................................................................................................ 211
normf........................................................................................................212
powf......................................................................................................... 212
rcbrtf........................................................................................................ 213
remainderf................................................................................................. 213
remquof.....................................................................................................214
rhypotf...................................................................................................... 214
rintf..........................................................................................................215
rnorm3df....................................................................................................215
rnorm4df....................................................................................................215
rnormf.......................................................................................................216
roundf....................................................................................................... 216
rsqrtf........................................................................................................ 217
scalblnf..................................................................................................... 217
scalbnf...................................................................................................... 217
[Link]
CUDA Math API vRelease Version | xiii
signbit....................................................................................................... 218
sincosf.......................................................................................................218
sincospif.................................................................................................... 219
sinf...........................................................................................................219
sinhf......................................................................................................... 220
sinpif........................................................................................................ 220
sqrtf......................................................................................................... 220
tanf.......................................................................................................... 221
tanhf........................................................................................................ 221
tgammaf.................................................................................................... 222
truncf........................................................................................................222
y0f........................................................................................................... 222
y1f........................................................................................................... 223
ynf........................................................................................................... 223
1.5. Double Precision Mathematical Functions......................................................... 224
acos..........................................................................................................224
acosh........................................................................................................ 224
asin.......................................................................................................... 225
asinh.........................................................................................................225
atan..........................................................................................................226
atan2........................................................................................................ 226
atanh........................................................................................................ 226
cbrt.......................................................................................................... 227
ceil...........................................................................................................227
copysign.....................................................................................................228
cos........................................................................................................... 228
cosh..........................................................................................................228
cospi.........................................................................................................229
cyl_bessel_i0............................................................................................... 229
cyl_bessel_i1............................................................................................... 229
erf........................................................................................................... 230
erfc.......................................................................................................... 230
erfcinv...................................................................................................... 231
erfcx.........................................................................................................231
erfinv........................................................................................................ 231
exp...........................................................................................................232
exp10........................................................................................................232
exp2......................................................................................................... 233
expm1....................................................................................................... 233
fabs.......................................................................................................... 233
fdim......................................................................................................... 234
floor......................................................................................................... 234
fma.......................................................................................................... 235
[Link]
CUDA Math API vRelease Version | xiv
fmax......................................................................................................... 235
fmin......................................................................................................... 236
fmod.........................................................................................................236
frexp.........................................................................................................237
hypot........................................................................................................ 237
ilogb......................................................................................................... 238
isfinite...................................................................................................... 238
isinf.......................................................................................................... 238
isnan.........................................................................................................239
j0.............................................................................................................239
j1.............................................................................................................239
jn............................................................................................................ 240
ldexp........................................................................................................ 240
lgamma..................................................................................................... 241
llrint......................................................................................................... 241
llround...................................................................................................... 242
log........................................................................................................... 242
log10........................................................................................................ 242
log1p........................................................................................................ 243
log2.......................................................................................................... 243
logb.......................................................................................................... 244
lrint..........................................................................................................244
lround....................................................................................................... 244
modf.........................................................................................................245
nan...........................................................................................................245
nearbyint................................................................................................... 245
nextafter................................................................................................... 246
norm.........................................................................................................246
norm3d......................................................................................................247
norm4d......................................................................................................247
normcdf..................................................................................................... 247
normcdfinv................................................................................................. 248
pow.......................................................................................................... 248
rcbrt......................................................................................................... 249
remainder.................................................................................................. 249
remquo......................................................................................................250
rhypot....................................................................................................... 250
rint...........................................................................................................251
rnorm........................................................................................................251
rnorm3d.....................................................................................................252
rnorm4d.....................................................................................................252
round........................................................................................................ 253
rsqrt......................................................................................................... 253
[Link]
CUDA Math API vRelease Version | xv
scalbln...................................................................................................... 253
scalbn....................................................................................................... 254
signbit....................................................................................................... 254
sin............................................................................................................254
sincos........................................................................................................255
sincospi..................................................................................................... 255
sinh.......................................................................................................... 256
sinpi......................................................................................................... 256
sqrt.......................................................................................................... 256
tan........................................................................................................... 257
tanh......................................................................................................... 257
tgamma..................................................................................................... 258
trunc.........................................................................................................258
y0............................................................................................................ 258
y1............................................................................................................ 259
yn............................................................................................................ 259
1.6. Single Precision Intrinsics.............................................................................260
__cosf....................................................................................................... 260
__exp10f.................................................................................................... 260
__expf.......................................................................................................261
__fadd_rd...................................................................................................261
__fadd_rn...................................................................................................262
__fadd_ru...................................................................................................262
__fadd_rz................................................................................................... 262
__fdiv_rd................................................................................................... 263
__fdiv_rn....................................................................................................263
__fdiv_ru....................................................................................................263
__fdiv_rz....................................................................................................264
__fdividef...................................................................................................264
__fmaf_rd.................................................................................................. 265
__fmaf_rn.................................................................................................. 265
__fmaf_ru.................................................................................................. 266
__fmaf_rz...................................................................................................266
__fmul_rd...................................................................................................267
__fmul_rn...................................................................................................267
__fmul_ru...................................................................................................267
__fmul_rz................................................................................................... 268
__frcp_rd................................................................................................... 268
__frcp_rn................................................................................................... 268
__frcp_ru................................................................................................... 269
__frcp_rz................................................................................................... 269
__frsqrt_rn................................................................................................. 270
__fsqrt_rd.................................................................................................. 270
[Link]
CUDA Math API vRelease Version | xvi
__fsqrt_rn.................................................................................................. 270
__fsqrt_ru.................................................................................................. 271
__fsqrt_rz...................................................................................................271
__fsub_rd................................................................................................... 271
__fsub_rn................................................................................................... 272
__fsub_ru................................................................................................... 272
__fsub_rz................................................................................................... 273
__log10f.....................................................................................................273
__log2f...................................................................................................... 273
__logf....................................................................................................... 274
__powf...................................................................................................... 274
__saturatef................................................................................................. 275
__sincosf.................................................................................................... 275
__sinf........................................................................................................275
__tanf....................................................................................................... 276
1.7. Double Precision Intrinsics........................................................................... 276
__dadd_rd.................................................................................................. 276
__dadd_rn.................................................................................................. 277
__dadd_ru.................................................................................................. 277
__dadd_rz.................................................................................................. 277
__ddiv_rd................................................................................................... 278
__ddiv_rn................................................................................................... 278
__ddiv_ru................................................................................................... 278
__ddiv_rz................................................................................................... 279
__dmul_rd.................................................................................................. 279
__dmul_rn.................................................................................................. 280
__dmul_ru.................................................................................................. 280
__dmul_rz.................................................................................................. 280
__drcp_rd...................................................................................................281
__drcp_rn...................................................................................................281
__drcp_ru...................................................................................................282
__drcp_rz................................................................................................... 282
__dsqrt_rd.................................................................................................. 282
__dsqrt_rn.................................................................................................. 283
__dsqrt_ru.................................................................................................. 283
__dsqrt_rz.................................................................................................. 284
__dsub_rd.................................................................................................. 284
__dsub_rn...................................................................................................284
__dsub_ru...................................................................................................285
__dsub_rz...................................................................................................285
__fma_rd................................................................................................... 286
__fma_rn................................................................................................... 286
__fma_ru................................................................................................... 287
[Link]
CUDA Math API vRelease Version | xvii
__fma_rz....................................................................................................287
1.8. Integer Intrinsics....................................................................................... 288
__brev.......................................................................................................288
__brevll..................................................................................................... 288
__byte_perm............................................................................................... 288
__clz.........................................................................................................289
__clzll....................................................................................................... 289
__ffs......................................................................................................... 289
__ffsll....................................................................................................... 290
__funnelshift_l.............................................................................................290
__funnelshift_lc........................................................................................... 290
__funnelshift_r............................................................................................ 291
__funnelshift_rc........................................................................................... 291
__hadd...................................................................................................... 291
__mul24.....................................................................................................292
__mul64hi.................................................................................................. 292
__mulhi..................................................................................................... 292
__popc...................................................................................................... 293
__popcll.....................................................................................................293
__rhadd..................................................................................................... 293
__sad........................................................................................................ 293
__uhadd.....................................................................................................294
__umul24................................................................................................... 294
__umul64hi................................................................................................. 294
__umulhi.................................................................................................... 295
__urhadd....................................................................................................295
__usad.......................................................................................................295
1.9. Type Casting Intrinsics................................................................................ 296
__double2float_rd.........................................................................................296
__double2float_rn.........................................................................................296
__double2float_ru.........................................................................................296
__double2float_rz......................................................................................... 297
__double2hiint............................................................................................. 297
__double2int_rd........................................................................................... 297
__double2int_rn........................................................................................... 297
__double2int_ru........................................................................................... 298
__double2int_rz........................................................................................... 298
__double2ll_rd............................................................................................. 298
__double2ll_rn............................................................................................. 299
__double2ll_ru............................................................................................. 299
__double2ll_rz............................................................................................. 299
__double2loint............................................................................................. 299
__double2uint_rd..........................................................................................300
[Link]
CUDA Math API vRelease Version | xviii
__double2uint_rn..........................................................................................300
__double2uint_ru..........................................................................................300
__double2uint_rz.......................................................................................... 301
__double2ull_rd........................................................................................... 301
__double2ull_rn........................................................................................... 301
__double2ull_ru........................................................................................... 302
__double2ull_rz............................................................................................302
__double_as_longlong.................................................................................... 302
__float2int_rd..............................................................................................303
__float2int_rn..............................................................................................303
__float2int_ru..............................................................................................303
__float2int_rz.............................................................................................. 303
__float2ll_rd............................................................................................... 304
__float2ll_rn............................................................................................... 304
__float2ll_ru............................................................................................... 304
__float2ll_rz................................................................................................305
__float2uint_rd............................................................................................ 305
__float2uint_rn............................................................................................ 305
__float2uint_ru............................................................................................ 305
__float2uint_rz............................................................................................ 306
__float2ull_rd.............................................................................................. 306
__float2ull_rn.............................................................................................. 306
__float2ull_ru.............................................................................................. 307
__float2ull_rz.............................................................................................. 307
__float_as_int..............................................................................................307
__float_as_uint............................................................................................ 307
__hiloint2double...........................................................................................308
__int2double_rn........................................................................................... 308
__int2float_rd..............................................................................................308
__int2float_rn..............................................................................................308
__int2float_ru..............................................................................................309
__int2float_rz.............................................................................................. 309
__int_as_float..............................................................................................309
__ll2double_rd............................................................................................. 310
__ll2double_rn............................................................................................. 310
__ll2double_ru............................................................................................. 310
__ll2double_rz............................................................................................. 310
__ll2float_rd............................................................................................... 311
__ll2float_rn............................................................................................... 311
__ll2float_ru............................................................................................... 311
__ll2float_rz................................................................................................312
__longlong_as_double.................................................................................... 312
__uint2double_rn..........................................................................................312
[Link]
CUDA Math API vRelease Version | xix
__uint2float_rd............................................................................................ 312
__uint2float_rn............................................................................................ 313
__uint2float_ru............................................................................................ 313
__uint2float_rz............................................................................................ 313
__uint_as_float............................................................................................ 314
__ull2double_rd........................................................................................... 314
__ull2double_rn........................................................................................... 314
__ull2double_ru........................................................................................... 315
__ull2double_rz............................................................................................315
__ull2float_rd.............................................................................................. 315
__ull2float_rn.............................................................................................. 316
__ull2float_ru.............................................................................................. 316
__ull2float_rz.............................................................................................. 316
1.10. SIMD Intrinsics.........................................................................................317
__vabs2..................................................................................................... 317
__vabs4..................................................................................................... 317
__vabsdiffs2................................................................................................ 317
__vabsdiffs4................................................................................................ 318
__vabsdiffu2................................................................................................318
__vabsdiffu4................................................................................................318
__vabsss2................................................................................................... 319
__vabsss4................................................................................................... 319
__vadd2..................................................................................................... 319
__vadd4..................................................................................................... 320
__vaddss2...................................................................................................320
__vaddss4...................................................................................................320
__vaddus2.................................................................................................. 321
__vaddus4.................................................................................................. 321
__vavgs2.................................................................................................... 321
__vavgs4.................................................................................................... 322
__vavgu2....................................................................................................322
__vavgu4....................................................................................................322
__vcmpeq2................................................................................................. 323
__vcmpeq4................................................................................................. 323
__vcmpges2................................................................................................ 323
__vcmpges4................................................................................................ 324
__vcmpgeu2................................................................................................ 324
__vcmpgeu4................................................................................................ 324
__vcmpgts2.................................................................................................325
__vcmpgts4.................................................................................................325
__vcmpgtu2................................................................................................ 325
__vcmpgtu4................................................................................................ 326
__vcmples2................................................................................................. 326
[Link]
CUDA Math API vRelease Version | xx
__vcmples4................................................................................................. 326
__vcmpleu2................................................................................................ 327
__vcmpleu4................................................................................................ 327
__vcmplts2................................................................................................. 327
__vcmplts4................................................................................................. 328
__vcmpltu2................................................................................................. 328
__vcmpltu4................................................................................................. 328
__vcmpne2................................................................................................. 329
__vcmpne4................................................................................................. 329
__vhaddu2.................................................................................................. 329
__vhaddu4.................................................................................................. 330
__vmaxs2................................................................................................... 330
__vmaxs4................................................................................................... 330
__vmaxu2................................................................................................... 331
__vmaxu4................................................................................................... 331
__vmins2....................................................................................................331
__vmins4....................................................................................................332
__vminu2................................................................................................... 332
__vminu4................................................................................................... 332
__vneg2..................................................................................................... 333
__vneg4..................................................................................................... 333
__vnegss2................................................................................................... 333
__vnegss4................................................................................................... 333
__vsads2.................................................................................................... 334
__vsads4.................................................................................................... 334
__vsadu2.................................................................................................... 334
__vsadu4.................................................................................................... 335
__vseteq2...................................................................................................335
__vseteq4...................................................................................................335
__vsetges2..................................................................................................336
__vsetges4..................................................................................................336
__vsetgeu2................................................................................................. 336
__vsetgeu4................................................................................................. 337
__vsetgts2.................................................................................................. 337
__vsetgts4.................................................................................................. 337
__vsetgtu2..................................................................................................338
__vsetgtu4..................................................................................................338
__vsetles2.................................................................................................. 338
__vsetles4.................................................................................................. 339
__vsetleu2.................................................................................................. 339
__vsetleu4.................................................................................................. 339
__vsetlts2...................................................................................................340
__vsetlts4...................................................................................................340
[Link]
CUDA Math API vRelease Version | xxi
__vsetltu2.................................................................................................. 340
__vsetltu4.................................................................................................. 341
__vsetne2...................................................................................................341
__vsetne4...................................................................................................341
__vsub2..................................................................................................... 342
__vsub4..................................................................................................... 342
__vsubss2................................................................................................... 342
__vsubss4................................................................................................... 343
__vsubus2...................................................................................................343
__vsubus4...................................................................................................343
[Link]
CUDA Math API vRelease Version | xxii
Chapter 1.
MODULES
[Link]
CUDA Math API vRelease Version | 1
Modules
Parameters
a
- half. Is only being read.
Returns
half
‣ The
absolute value of a.
[Link]
CUDA Math API vRelease Version | 2
Modules
Description
Calculates the absolute value of input half number and returns the result.
Description
Performs half addition of inputs a and b, in round-to-nearest-even mode.
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
half
‣ The
sum of a and b, with respect to saturation.
Description
Performs half add of inputs a and b, in round-to-nearest-even mode, and clamps the
result to range [0.0, 1.0]. NaN results are flushed to +0.0.
Description
Divides half input a by input b in round-to-nearest mode.
[Link]
CUDA Math API vRelease Version | 3
Modules
Description
Performs half multiply on inputs a and b, then performs a half add of the result with
c, rounding the result once in round-to-nearest-even mode.
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
c
- half. Is only being read.
Returns
half
‣ The
result of fused multiply-add operation on a, b, and c with relu saturation.
Description
Performs half multiply on inputs a and b, then performs a half add of the result
with c, rounding the result once in round-to-nearest-even mode. Then negative result is
clamped to 0. NaN result is converted to canonical NaN.
Parameters
a
- half. Is only being read.
[Link]
CUDA Math API vRelease Version | 4
Modules
b
- half. Is only being read.
c
- half. Is only being read.
Returns
half
‣ The
result of fused multiply-add operation on a, b, and c, with respect to saturation.
Description
Performs half multiply on inputs a and b, then performs a half add of the result with
c, rounding the result once in round-to-nearest-even mode, and clamps the result to
range [0.0, 1.0]. NaN results are flushed to +0.0.
Description
Performs half multiplication of inputs a and b, in round-to-nearest mode.
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
half
‣ The
result of multiplying a and b, with respect to saturation.
Description
Performs half multiplication of inputs a and b, in round-to-nearest mode, and clamps
the result to range [0.0, 1.0]. NaN results are flushed to +0.0.
[Link]
CUDA Math API vRelease Version | 5
Modules
Description
Negates input half number and returns the result.
Description
Subtracts half input b from input a in round-to-nearest mode.
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
half
‣ The
result of subtraction of b from a, with respect to saturation.
Description
Subtracts half input b from input a in round-to-nearest mode, and clamps the result to
range [0.0, 1.0]. NaN results are flushed to +0.0.
[Link]
CUDA Math API vRelease Version | 6
Modules
Description
Divides half2 input vector a by input vector b in round-to-nearest mode.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ Returns
a with the absolute value of both halves.
Description
Calculates the absolute value of both halves of the input half2 number and returns the
result.
Description
Performs half2 vector add of inputs a and b, in round-to-nearest mode.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
[Link]
CUDA Math API vRelease Version | 7
Modules
Returns
half2
‣ The
sum of a and b, with respect to saturation.
Description
Performs half2 vector add of inputs a and b, in round-to-nearest mode, and clamps the
results to range [0.0, 1.0]. NaN results are flushed to +0.0.
Description
Performs half2 vector multiply on inputs a and b, then performs a half2 vector add
of the result with c, rounding the result once in round-to-nearest-even mode.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
c
- half2. Is only being read.
Returns
half2
‣ The
result of elementwise fused multiply-add operation on vectors a, b, and c with relu
saturation.
[Link]
CUDA Math API vRelease Version | 8
Modules
Description
Performs half2 vector multiply on inputs a and b, then performs a half2 vector add
of the result with c, rounding the result once in round-to-nearest-even mode. Then
negative result is clamped to 0. NaN result is converted to canonical NaN.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
c
- half2. Is only being read.
Returns
half2
‣ The
result of elementwise fused multiply-add operation on vectors a, b, and c, with
respect to saturation.
Description
Performs half2 vector multiply on inputs a and b, then performs a half2 vector
add of the result with c, rounding the result once in round-to-nearest-even mode, and
clamps the results to range [0.0, 1.0]. NaN results are flushed to +0.0.
Description
Performs half2 vector multiplication of inputs a and b, in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 9
Modules
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
result of elementwise multiplication of vectors a and b, with respect to saturation.
Description
Performs half2 vector multiplication of inputs a and b, in round-to-nearest-even mode,
and clamps the results to range [0.0, 1.0]. NaN results are flushed to +0.0.
Description
Negates both halves of the input half2 number a and returns the result.
Description
Subtracts half2 input vector b from input vector a in round-to-nearest-even mode.
Parameters
a
- half2. Is only being read.
[Link]
CUDA Math API vRelease Version | 10
Modules
b
- half2. Is only being read.
Returns
half2
‣ The
subtraction of vector b from a, with respect to saturation.
Description
Subtracts half2 input vector b from input vector a in round-to-nearest-even mode, and
clamps the results to range [0.0, 1.0]. NaN results are flushed to +0.0.
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of if-equal comparison of a and b.
Description
Performs half if-equal comparison of inputs a and b. NaN inputs generate false results.
[Link]
CUDA Math API vRelease Version | 11
Modules
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of unordered if-equal comparison of a and b.
Description
Performs half if-equal comparison of inputs a and b. NaN inputs generate true results.
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of greater-equal comparison of a and b.
Description
Performs half greater-equal comparison of inputs a and b. NaN inputs generate false
results.
[Link]
CUDA Math API vRelease Version | 12
Modules
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of unordered greater-equal comparison of a and b.
Description
Performs half greater-equal comparison of inputs a and b. NaN inputs generate true
results.
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of greater-than comparison of a and b.
Description
Performs half greater-than comparison of inputs a and b. NaN inputs generate false
results.
[Link]
CUDA Math API vRelease Version | 13
Modules
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of unordered greater-than comparison of a and b.
Description
Performs half greater-than comparison of inputs a and b. NaN inputs generate true
results.
Parameters
a
- half. Is only being read.
Returns
int
‣ -1
iff a is equal to negative infinity,
‣ 1
iff a is equal to positive infinity,
‣ 0
otherwise.
Description
Checks if the input half number a is infinite.
[Link]
CUDA Math API vRelease Version | 14
Modules
Parameters
a
- half. Is only being read.
Returns
bool
‣ true
iff argument is NaN.
Description
Determine whether half value a is a NaN.
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of less-equal comparison of a and b.
Description
Performs half less-equal comparison of inputs a and b. NaN inputs generate false
results.
[Link]
CUDA Math API vRelease Version | 15
Modules
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of unordered less-equal comparison of a and b.
Description
Performs half less-equal comparison of inputs a and b. NaN inputs generate true
results.
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of less-than comparison of a and b.
Description
Performs half less-than comparison of inputs a and b. NaN inputs generate false
results.
[Link]
CUDA Math API vRelease Version | 16
Modules
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of unordered less-than comparison of a and b.
Description
Performs half less-than comparison of inputs a and b. NaN inputs generate true
results.
Description
Calculates half max(a, b) defined as (a > b) ? a : b.
Description
Calculates half max(a, b) defined as (a > b) ? a : b.
[Link]
CUDA Math API vRelease Version | 17
Modules
Description
Calculates half min(a, b) defined as (a < b) ? a : b.
Description
Calculates half min(a, b) defined as (a < b) ? a : b.
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of not-equal comparison of a and b.
Description
Performs half not-equal comparison of inputs a and b. NaN inputs generate false
results.
[Link]
CUDA Math API vRelease Version | 18
Modules
Parameters
a
- half. Is only being read.
b
- half. Is only being read.
Returns
bool
‣ The
boolean result of unordered not-equal comparison of a and b.
Description
Performs half not-equal comparison of inputs a and b. NaN inputs generate true
results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
bool
‣ trueif
both half results of if-equal comparison of vectors a and b are true;
‣ falseotherwise.
[Link]
CUDA Math API vRelease Version | 19
Modules
Description
Performs half2 vector if-equal comparison of inputs a and b. The bool result is set to
true only if both half if-equal comparisons evaluate to true, or false otherwise. NaN
inputs generate false results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
bool
‣ trueif
both half results of unordered if-equal comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs half2 vector if-equal comparison of inputs a and b. The bool result is set to
true only if both half if-equal comparisons evaluate to true, or false otherwise. NaN
inputs generate true results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
bool
‣ trueif
[Link]
CUDA Math API vRelease Version | 20
Modules
Description
Performs half2 vector greater-equal comparison of inputs a and b. The bool result
is set to true only if both half greater-equal comparisons evaluate to true, or false
otherwise. NaN inputs generate false results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
bool
‣ trueif
both half results of unordered greater-equal comparison of vectors a and b are
true;
‣ falseotherwise.
Description
Performs half2 vector greater-equal comparison of inputs a and b. The bool result
is set to true only if both half greater-equal comparisons evaluate to true, or false
otherwise. NaN inputs generate true results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
[Link]
CUDA Math API vRelease Version | 21
Modules
Returns
bool
‣ trueif
both half results of greater-than comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs half2 vector greater-than comparison of inputs a and b. The bool result is set
to true only if both half greater-than comparisons evaluate to true, or false otherwise.
NaN inputs generate false results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
bool
‣ trueif
both half results of unordered greater-than comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs half2 vector greater-than comparison of inputs a and b. The bool result is set
to true only if both half greater-than comparisons evaluate to true, or false otherwise.
NaN inputs generate true results.
[Link]
CUDA Math API vRelease Version | 22
Modules
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
bool
‣ trueif
both half results of less-equal comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs half2 vector less-equal comparison of inputs a and b. The bool result is set to
true only if both half less-equal comparisons evaluate to true, or false otherwise. NaN
inputs generate false results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
bool
‣ trueif
both half results of unordered less-equal comparison of vectors a and b are true;
‣ falseotherwise.
[Link]
CUDA Math API vRelease Version | 23
Modules
Description
Performs half2 vector less-equal comparison of inputs a and b. The bool result is set to
true only if both half less-equal comparisons evaluate to true, or false otherwise. NaN
inputs generate true results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
bool
‣ trueif
both half results of less-than comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs half2 vector less-than comparison of inputs a and b. The bool result is set to
true only if both half less-than comparisons evaluate to true, or false otherwise. NaN
inputs generate false results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
bool
‣ trueif
[Link]
CUDA Math API vRelease Version | 24
Modules
both half results of unordered less-than comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs half2 vector less-than comparison of inputs a and b. The bool result is set to
true only if both half less-than comparisons evaluate to true, or false otherwise. NaN
inputs generate true results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
bool
‣ trueif
both half results of not-equal comparison of vectors a and b are true,
‣ false
otherwise.
Description
Performs half2 vector not-equal comparison of inputs a and b. The bool result is set to
true only if both half not-equal comparisons evaluate to true, or false otherwise. NaN
inputs generate false results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
[Link]
CUDA Math API vRelease Version | 25
Modules
Returns
bool
‣ trueif
both half results of unordered not-equal comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs half2 vector not-equal comparison of inputs a and b. The bool result is set to
true only if both half not-equal comparisons evaluate to true, or false otherwise. NaN
inputs generate true results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
vector result of if-equal comparison of vectors a and b.
Description
Performs half2 vector if-equal comparison of inputs a and b. The corresponding half
results are set to 1.0 for true, or 0.0 for false. NaN inputs generate false results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
[Link]
CUDA Math API vRelease Version | 26
Modules
Returns
half2
‣ The
vector result of unordered if-equal comparison of vectors a and b.
Description
Performs half2 vector if-equal comparison of inputs a and b. The corresponding half
results are set to 1.0 for true, or 0.0 for false. NaN inputs generate true results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
vector result of greater-equal comparison of vectors a and b.
Description
Performs half2 vector greater-equal comparison of inputs a and b. The corresponding
half results are set to 1.0 for true, or 0.0 for false. NaN inputs generate false results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
[Link]
CUDA Math API vRelease Version | 27
Modules
‣ The
half2 vector result of unordered greater-equal comparison of vectors a and b.
Description
Performs half2 vector greater-equal comparison of inputs a and b. The corresponding
half results are set to 1.0 for true, or 0.0 for false. NaN inputs generate true results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
vector result of greater-than comparison of vectors a and b.
Description
Performs half2 vector greater-than comparison of inputs a and b. The corresponding
half results are set to 1.0 for true, or 0.0 for false. NaN inputs generate false results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
half2 vector result of unordered greater-than comparison of vectors a and b.
[Link]
CUDA Math API vRelease Version | 28
Modules
Description
Performs half2 vector greater-than comparison of inputs a and b. The corresponding
half results are set to 1.0 for true, or 0.0 for false. NaN inputs generate true results.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
half2 with the corresponding half results set to 1.0 for for NaN, 0.0 otherwise.
Description
Determine whether each half of input half2 number a is a NaN.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
half2 result of less-equal comparison of vectors a and b.
Description
Performs half2 vector less-equal comparison of inputs a and b. The corresponding
half results are set to 1.0 for true, or 0.0 for false. NaN inputs generate false results.
[Link]
CUDA Math API vRelease Version | 29
Modules
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
vector result of unordered less-equal comparison of vectors a and b.
Description
Performs half2 vector less-equal comparison of inputs a and b. The corresponding
half results are set to 1.0 for true, or 0.0 for false. NaN inputs generate true results.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
half2 vector result of less-than comparison of vectors a and b.
Description
Performs half2 vector less-than comparison of inputs a and b. The corresponding half
results are set to 1.0 for true, or 0.0 for false. NaN inputs generate false results.
[Link]
CUDA Math API vRelease Version | 30
Modules
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
vector result of unordered less-than comparison of vectors a and b.
Description
Performs half2 vector less-than comparison of inputs a and b. The corresponding half
results are set to 1.0 for true, or 0.0 for false. NaN inputs generate true results.
Description
Calculates half2 vector max(a, b) Elementwise half operation is defined as (a > b) ? a
: b.
Description
Calculates half2 vector max(a, b) Elementwise half operation is defined as (a > b) ? a
: b.
[Link]
CUDA Math API vRelease Version | 31
Modules
Description
Calculates half2 vector min(a, b) Elementwise half operation is defined as (a < b) ? a :
b.
Description
Calculates half2 vector min(a, b) Elementwise half operation is defined as (a < b) ? a :
b.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
vector result of not-equal comparison of vectors a and b.
Description
Performs half2 vector not-equal comparison of inputs a and b. The corresponding
half results are set to 1.0 for true, or 0.0 for false. NaN inputs generate false results.
[Link]
CUDA Math API vRelease Version | 32
Modules
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
vector result of unordered not-equal comparison of vectors a and b.
Description
Performs half2 vector not-equal comparison of inputs a and b. The corresponding
half results are set to 1.0 for true, or 0.0 for false. NaN inputs generate true results.
Parameters
a
- double. Is only being read.
Returns
half
‣ \p
a converted to half.
Description
Converts double number a to half precision in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 33
Modules
Parameters
a
- float2. Is only being read.
Returns
half2
‣ The
half2 which has corresponding halves equal to the converted float2 components.
Description
Converts both components of float2 to half precision in round-to-nearest mode and
combines the results into one half2 number. Low 16 bits of the return value correspond
to a.x and high 16 bits of the return value correspond to a.y.
Parameters
a
- float. Is only being read.
Returns
half
‣ \p
a converted to half.
Description
Converts float number a to half precision in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 34
Modules
Parameters
a
- float. Is only being read.
Returns
half2
‣ The
half2 value with both halves equal to the converted half precision number.
Description
Converts input a to half precision in round-to-nearest-even mode and populates both
halves of half2 with converted value.
Parameters
a
- float. Is only being read.
Returns
half
‣ \p
a converted to half.
Description
Converts float number a to half precision in round-down mode.
[Link]
CUDA Math API vRelease Version | 35
Modules
Parameters
a
- float. Is only being read.
Returns
half
‣ \p
a converted to half.
Description
Converts float number a to half precision in round-to-nearest-even mode.
Parameters
a
- float. Is only being read.
Returns
half
‣ \p
a converted to half.
Description
Converts float number a to half precision in round-up mode.
[Link]
CUDA Math API vRelease Version | 36
Modules
Parameters
a
- float. Is only being read.
Returns
half
‣ \p
a converted to half.
Description
Converts float number a to half precision in round-towards-zero mode.
Parameters
a
- float. Is only being read.
b
- float. Is only being read.
Returns
half2
‣ The
half2 value with corresponding halves equal to the converted input floats.
Description
Converts both input floats to half precision in round-to-nearest-even mode and
combines the results into one half2 number. Low 16 bits of the return value correspond
to the input a, high 16 bits correspond to the input b.
[Link]
CUDA Math API vRelease Version | 37
Modules
Parameters
a
- half2. Is only being read.
Returns
float2
‣ \p
a converted to float2.
Description
Converts both halves of half2 input a to float2 and returns the result.
Parameters
a
- float. Is only being read.
Returns
float
‣ \p
a converted to float.
Description
Converts half number a to float.
Parameters
a
- half. Is only being read.
[Link]
CUDA Math API vRelease Version | 38
Modules
Returns
half2
‣ The
vector which has both its halves equal to the input a.
Description
Returns half2 number with both halves equal to the input a half number.
Parameters
h
- half. Is only being read.
Returns
int
‣ \p
h converted to a signed integer.
Description
Convert the half-precision floating point value h to a signed integer in round-down
mode.
Parameters
h
- half. Is only being read.
Returns
int
‣ \p
h converted to a signed integer.
[Link]
CUDA Math API vRelease Version | 39
Modules
Description
Convert the half-precision floating point value h to a signed integer in round-to-nearest-
even mode.
Parameters
h
- half. Is only being read.
Returns
int
‣ \p
h converted to a signed integer.
Description
Convert the half-precision floating point value h to a signed integer in round-up mode.
Parameters
h
- half. Is only being read.
Returns
int
‣ \p
h converted to a signed integer.
Description
Convert the half-precision floating point value h to a signed integer in round-towards-
zero mode.
[Link]
CUDA Math API vRelease Version | 40
Modules
Parameters
h
- half. Is only being read.
Returns
long long int
‣ \p
h converted to a signed 64-bit integer.
Description
Convert the half-precision floating point value h to a signed 64-bit integer in round-
down mode.
Parameters
h
- half. Is only being read.
Returns
long long int
‣ \p
h converted to a signed 64-bit integer.
Description
Convert the half-precision floating point value h to a signed 64-bit integer in round-to-
nearest-even mode.
[Link]
CUDA Math API vRelease Version | 41
Modules
Parameters
h
- half. Is only being read.
Returns
long long int
‣ \p
h converted to a signed 64-bit integer.
Description
Convert the half-precision floating point value h to a signed 64-bit integer in round-up
mode.
Parameters
h
- half. Is only being read.
Returns
long long int
‣ \p
h converted to a signed 64-bit integer.
Description
Convert the half-precision floating point value h to a signed 64-bit integer in round-
towards-zero mode.
[Link]
CUDA Math API vRelease Version | 42
Modules
Parameters
h
- half. Is only being read.
Returns
short int
‣ \p
h converted to a signed short integer.
Description
Convert the half-precision floating point value h to a signed short integer in round-
down mode.
Parameters
h
- half. Is only being read.
Returns
short int
‣ \p
h converted to a signed short integer.
Description
Convert the half-precision floating point value h to a signed short integer in round-to-
nearest-even mode.
[Link]
CUDA Math API vRelease Version | 43
Modules
Parameters
h
- half. Is only being read.
Returns
short int
‣ \p
h converted to a signed short integer.
Description
Convert the half-precision floating point value h to a signed short integer in round-up
mode.
Parameters
h
- half. Is only being read.
Returns
short int
‣ \p
h converted to a signed short integer.
Description
Convert the half-precision floating point value h to a signed short integer in round-
towards-zero mode.
[Link]
CUDA Math API vRelease Version | 44
Modules
Parameters
h
- half. Is only being read.
Returns
unsigned int
‣ \p
h converted to an unsigned integer.
Description
Convert the half-precision floating point value h to an unsigned integer in round-down
mode.
Parameters
h
- half. Is only being read.
Returns
unsigned int
‣ \p
h converted to an unsigned integer.
Description
Convert the half-precision floating point value h to an unsigned integer in round-to-
nearest-even mode.
[Link]
CUDA Math API vRelease Version | 45
Modules
Parameters
h
- half. Is only being read.
Returns
unsigned int
‣ \p
h converted to an unsigned integer.
Description
Convert the half-precision floating point value h to an unsigned integer in round-up
mode.
Parameters
h
- half. Is only being read.
Returns
unsigned int
‣ \p
h converted to an unsigned integer.
Description
Convert the half-precision floating point value h to an unsigned integer in round-
towards-zero mode.
[Link]
CUDA Math API vRelease Version | 46
Modules
Parameters
h
- half. Is only being read.
Returns
unsigned long long int
‣ \p
h converted to an unsigned 64-bit integer.
Description
Convert the half-precision floating point value h to an unsigned 64-bit integer in round-
down mode.
Parameters
h
- half. Is only being read.
Returns
unsigned long long int
‣ \p
h converted to an unsigned 64-bit integer.
Description
Convert the half-precision floating point value h to an unsigned 64-bit integer in round-
to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 47
Modules
Parameters
h
- half. Is only being read.
Returns
unsigned long long int
‣ \p
h converted to an unsigned 64-bit integer.
Description
Convert the half-precision floating point value h to an unsigned 64-bit integer in round-
up mode.
Parameters
h
- half. Is only being read.
Returns
unsigned long long int
‣ \p
h converted to an unsigned 64-bit integer.
Description
Convert the half-precision floating point value h to an unsigned 64-bit integer in round-
towards-zero mode.
[Link]
CUDA Math API vRelease Version | 48
Modules
Parameters
h
- half. Is only being read.
Returns
unsigned short int
‣ \p
h converted to an unsigned short integer.
Description
Convert the half-precision floating point value h to an unsigned short integer in round-
down mode.
Parameters
h
- half. Is only being read.
Returns
unsigned short int
‣ \p
h converted to an unsigned short integer.
Description
Convert the half-precision floating point value h to an unsigned short integer in round-
to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 49
Modules
Parameters
h
- half. Is only being read.
Returns
unsigned short int
‣ \p
h converted to an an unsigned short integer.
Description
Convert the half-precision floating point value h to an unsigned short integer in round-
up mode.
Parameters
h
- half. Is only being read.
Returns
unsigned short int
‣ \p
h converted to an unsigned short integer.
Description
Convert the half-precision floating point value h to an unsigned short integer in round-
towards-zero mode.
[Link]
CUDA Math API vRelease Version | 50
Modules
Parameters
h
- half. Is only being read.
Returns
short int
‣ The
reinterpreted value.
Description
Reinterprets the bits in the half-precision floating point number h as a signed short
integer.
Parameters
h
- half. Is only being read.
Returns
unsigned short int
‣ The
reinterpreted value.
Description
Reinterprets the bits in the half-precision floating point h as an unsigned short number.
Parameters
a
- half. Is only being read.
[Link]
CUDA Math API vRelease Version | 51
Modules
b
- half. Is only being read.
Returns
half2
‣ The
half2 with one half equal to a and the other to b.
Description
Combines two input half number a and b into one half2 number. Input a is stored in
low 16 bits of the return value, input b is stored in high 16 bits of the return value.
Parameters
a
- half2. Is only being read.
Returns
float
‣ The
high 16 bits of a converted to float.
Description
Converts high 16 bits of half2 input a to 32 bit floating point number and returns the
result.
Parameters
a
- half2. Is only being read.
Returns
half
‣ The
[Link]
CUDA Math API vRelease Version | 52
Modules
Description
Returns high 16 bits of half2 input a.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
half2 with both halves equal to the high 16 bits of the input.
Description
Extracts high 16 bits from half2 input a and returns a new half2 number which has
both halves equal to the extracted bits.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
high 16 bits of a and of b.
[Link]
CUDA Math API vRelease Version | 53
Modules
Description
Extracts high 16 bits from each of the two half2 inputs and combines into one half2
number. High 16 bits from input a is stored in low 16 bits of the return value, high 16
bits from input b is stored in high 16 bits of the return value.
Parameters
i
- int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed integer value i to a half-precision floating point value in round-
down mode.
Parameters
i
- int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed integer value i to a half-precision floating point value in round-to-
nearest-even mode.
[Link]
CUDA Math API vRelease Version | 54
Modules
Parameters
i
- int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed integer value i to a half-precision floating point value in round-up
mode.
Parameters
i
- int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed integer value i to a half-precision floating point value in round-
towards-zero mode.
[Link]
CUDA Math API vRelease Version | 55
Modules
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
[Link]
CUDA Math API vRelease Version | 56
Modules
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
[Link]
CUDA Math API vRelease Version | 57
Modules
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
[Link]
CUDA Math API vRelease Version | 58
Modules
Parameters
i
- long long int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed 64-bit integer value i to a half-precision floating point value in
round-down mode.
Parameters
i
- long long int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed 64-bit integer value i to a half-precision floating point value in
round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 59
Modules
Parameters
i
- long long int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed 64-bit integer value i to a half-precision floating point value in
round-up mode.
Parameters
i
- long long int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed 64-bit integer value i to a half-precision floating point value in
round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 60
Modules
Parameters
a
- half2. Is only being read.
Returns
float
‣ The
low 16 bits of a converted to float.
Description
Converts low 16 bits of half2 input a to 32 bit floating point number and returns the
result.
Parameters
a
- half2. Is only being read.
Returns
half
‣ Returns
half which contains low 16 bits of the input a.
Description
Returns low 16 bits of half2 input a.
Parameters
a
- half2. Is only being read.
[Link]
CUDA Math API vRelease Version | 61
Modules
Returns
half2
‣ The
half2 with both halves equal to the low 16 bits of the input.
Description
Extracts low 16 bits from half2 input a and returns a new half2 number which has
both halves equal to the extracted bits.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ \p
a with its halves being swapped.
Description
Swaps both halves of the half2 input and returns a new half2 number with swapped
halves.
Parameters
a
- half2. Is only being read.
b
- half2. Is only being read.
Returns
half2
‣ The
[Link]
CUDA Math API vRelease Version | 62
Modules
Description
Extracts low 16 bits from each of the two half2 inputs and combines into one half2
number. Low 16 bits from input a is stored in low 16 bits of the return value, low 16 bits
from input b is stored in high 16 bits of the return value.
Parameters
mask
- unsigned int. Is only being read.
var
- half. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 2-byte word referenced by var from the source thread ID as half. If the
source thread ID is out of range or the source thread has exited, the calling thread's own
var is returned.
Description
Calculates a source thread ID by adding delta to the caller's thread ID. The value
of var held by the resulting thread ID is returned: this has the effect of shifting var
down the warp by delta threads. If width is less than warpSize then each subsection
of the warp behaves as a separate entity with a starting logical thread ID of 0. As for
__shfl_up_sync(), the ID number of the source thread will not wrap around the value of
width and so the upper delta threads will remain unchanged.
[Link]
CUDA Math API vRelease Version | 63
Modules
Parameters
mask
- unsigned int. Is only being read.
var
- half2. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 4-byte word referenced by var from the source thread ID as half2. If the
source thread ID is out of range or the source thread has exited, the calling thread's own
var is returned.
Description
Calculates a source thread ID by adding delta to the caller's thread ID. The value
of var held by the resulting thread ID is returned: this has the effect of shifting var
down the warp by delta threads. If width is less than warpSize then each subsection
of the warp behaves as a separate entity with a starting logical thread ID of 0. As for
__shfl_up_sync(), the ID number of the source thread will not wrap around the value of
width and so the upper delta threads will remain unchanged.
Parameters
mask
- unsigned int. Is only being read.
var
- half. Is only being read.
delta
- int. Is only being read.
[Link]
CUDA Math API vRelease Version | 64
Modules
width
- int. Is only being read.
Returns
Returns the 2-byte word referenced by var from the source thread ID as half. If the
source thread ID is out of range or the source thread has exited, the calling thread's own
var is returned.
Description
Returns the value of var held by the thread whose ID is given by delta. If width is less
than warpSize then each subsection of the warp behaves as a separate entity with a
starting logical thread ID of 0. If delta is outside the range [0:width-1], the value returned
corresponds to the value of var held by the delta modulo width (i.e. ithin the same
subsection). width must have a value which is a power of 2; results are undefined if
width is not a power of 2, or is a number greater than warpSize.
Parameters
mask
- unsigned int. Is only being read.
var
- half2. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 4-byte word referenced by var from the source thread ID as half2. If the
source thread ID is out of range or the source thread has exited, the calling thread's own
var is returned.
Description
Returns the value of var held by the thread whose ID is given by delta. If width is less
than warpSize then each subsection of the warp behaves as a separate entity with a
starting logical thread ID of 0. If delta is outside the range [0:width-1], the value returned
corresponds to the value of var held by the delta modulo width (i.e. ithin the same
[Link]
CUDA Math API vRelease Version | 65
Modules
subsection). width must have a value which is a power of 2; results are undefined if
width is not a power of 2, or is a number greater than warpSize.
Parameters
mask
- unsigned int. Is only being read.
var
- half. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 2-byte word referenced by var from the source thread ID as half. If the
source thread ID is out of range or the source thread has exited, the calling thread's own
var is returned.
Description
Calculates a source thread ID by subtracting delta from the caller's lane ID. The value
of var held by the resulting lane ID is returned: in effect, var is shifted up the warp by
delta threads. If width is less than warpSize then each subsection of the warp behaves
as a separate entity with a starting logical thread ID of 0. The source thread index
will not wrap around the value of width, so effectively the lower delta threads will be
unchanged. width must have a value which is a power of 2; results are undefined if
width is not a power of 2, or is a number greater than warpSize.
Parameters
mask
- unsigned int. Is only being read.
[Link]
CUDA Math API vRelease Version | 66
Modules
var
- half2. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 4-byte word referenced by var from the source thread ID as half2. If the
source thread ID is out of range or the source thread has exited, the calling thread's own
var is returned.
Description
Calculates a source thread ID by subtracting delta from the caller's lane ID. The value
of var held by the resulting lane ID is returned: in effect, var is shifted up the warp by
delta threads. If width is less than warpSize then each subsection of the warp behaves
as a separate entity with a starting logical thread ID of 0. The source thread index
will not wrap around the value of width, so effectively the lower delta threads will be
unchanged. width must have a value which is a power of 2; results are undefined if
width is not a power of 2, or is a number greater than warpSize.
Parameters
mask
- unsigned int. Is only being read.
var
- half. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 2-byte word referenced by var from the source thread ID as half. If the
source thread ID is out of range or the source thread has exited, the calling thread's own
var is returned.
[Link]
CUDA Math API vRelease Version | 67
Modules
Description
Calculates a source thread ID by performing a bitwise XOR of the caller's thread ID with
mask: the value of var held by the resulting thread ID is returned. If width is less than
warpSize then each group of width consecutive threads are able to access elements from
earlier groups of threads, however if they attempt to access elements from later groups
of threads their own value of var will be returned. This mode implements a butterfly
addressing pattern such as is used in tree reduction and broadcast.
Parameters
mask
- unsigned int. Is only being read.
var
- half2. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 4-byte word referenced by var from the source thread ID as half2. If the
source thread ID is out of range or the source thread has exited, the calling thread's own
var is returned.
Description
Calculates a source thread ID by performing a bitwise XOR of the caller's thread ID with
mask: the value of var held by the resulting thread ID is returned. If width is less than
warpSize then each group of width consecutive threads are able to access elements from
earlier groups of threads, however if they attempt to access elements from later groups
of threads their own value of var will be returned. This mode implements a butterfly
addressing pattern such as is used in tree reduction and broadcast.
[Link]
CUDA Math API vRelease Version | 68
Modules
Parameters
i
- short int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed short integer value i to a half-precision floating point value in
round-down mode.
Parameters
i
- short int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed short integer value i to a half-precision floating point value in
round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 69
Modules
Parameters
i
- short int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed short integer value i to a half-precision floating point value in
round-up mode.
Parameters
i
- short int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the signed short integer value i to a half-precision floating point value in
round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 70
Modules
Parameters
i
- short int. Is only being read.
Returns
half
‣ The
reinterpreted value.
Description
Reinterprets the bits in the signed short integer i as a half-precision floating point
number.
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
ptr
- memory location
value
- the value to be stored
[Link]
CUDA Math API vRelease Version | 71
Modules
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
ptr
- memory location
value
- the value to be stored
[Link]
CUDA Math API vRelease Version | 72
Modules
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
i
- unsigned int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned integer value i to a half-precision floating point value in round-
down mode.
[Link]
CUDA Math API vRelease Version | 73
Modules
Parameters
i
- unsigned int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned integer value i to a half-precision floating point value in round-to-
nearest-even mode.
Parameters
i
- unsigned int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned integer value i to a half-precision floating point value in round-up
mode.
[Link]
CUDA Math API vRelease Version | 74
Modules
Parameters
i
- unsigned int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned integer value i to a half-precision floating point value in round-
towards-zero mode.
Parameters
i
- unsigned long long int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned 64-bit integer value i to a half-precision floating point value in
round-down mode.
[Link]
CUDA Math API vRelease Version | 75
Modules
Parameters
i
- unsigned long long int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned 64-bit integer value i to a half-precision floating point value in
round-to-nearest-even mode.
Parameters
i
- unsigned long long int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned 64-bit integer value i to a half-precision floating point value in
round-up mode.
[Link]
CUDA Math API vRelease Version | 76
Modules
Parameters
i
- unsigned long long int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned 64-bit integer value i to a half-precision floating point value in
round-towards-zero mode.
Parameters
i
- unsigned short int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned short integer value i to a half-precision floating point value in
round-down mode.
[Link]
CUDA Math API vRelease Version | 77
Modules
Parameters
i
- unsigned short int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned short integer value i to a half-precision floating point value in
round-to-nearest-even mode.
Parameters
i
- unsigned short int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned short integer value i to a half-precision floating point value in
round-up mode.
[Link]
CUDA Math API vRelease Version | 78
Modules
Parameters
i
- unsigned short int. Is only being read.
Returns
half
‣ \p
i converted to half.
Description
Convert the unsigned short integer value i to a half-precision floating point value in
round-towards-zero mode.
Parameters
i
- unsigned short int. Is only being read.
Returns
half
‣ The
reinterpreted value.
Description
Reinterprets the bits in the unsigned short integer i as a half-precision floating point
number.
[Link]
CUDA Math API vRelease Version | 79
Modules
Parameters
h
- half. Is only being read.
Returns
half
‣ The
smallest integer value not less than h.
Description
Compute the smallest integer value not less than h.
Parameters
a
- half. Is only being read.
Returns
half
‣ The
cosine of a.
Description
Calculates half cosine of input a in round-to-nearest-even mode.
Parameters
a
- half. Is only being read.
[Link]
CUDA Math API vRelease Version | 80
Modules
Returns
half
‣ The
natural exponential function on a.
Description
Calculates half natural exponential function of input a in round-to-nearest-even mode.
Parameters
a
- half. Is only being read.
Returns
half
‣ The
decimal exponential function on a.
Description
Calculates half decimal exponential function of input a in round-to-nearest-even mode.
Parameters
a
- half. Is only being read.
Returns
half
‣ The
binary exponential function on a.
Description
Calculates half binary exponential function of input a in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 81
Modules
Parameters
h
- half. Is only being read.
Returns
half
‣ The
largest integer value which is less than or equal to h.
Description
Calculate the largest integer value which is less than or equal to h.
Parameters
a
- half. Is only being read.
Returns
half
‣ The
natural logarithm of a.
Description
Calculates half natural logarithm of input a in round-to-nearest-even mode.
Parameters
a
- half. Is only being read.
[Link]
CUDA Math API vRelease Version | 82
Modules
Returns
half
‣ The
decimal logarithm of a.
Description
Calculates half decimal logarithm of input a in round-to-nearest-even mode.
Parameters
a
- half. Is only being read.
Returns
half
‣ The
binary logarithm of a.
Description
Calculates half binary logarithm of input a in round-to-nearest-even mode.
Parameters
a
- half. Is only being read.
Returns
half
‣ The
reciprocal of a.
Description
Calculates half reciprocal of input a in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 83
Modules
Parameters
h
- half. Is only being read.
Returns
half
‣ The
nearest integer to h.
Description
Round h to the nearest integer value in half-precision floating point format, with
halfway cases rounded to the nearest even integer value.
Parameters
a
- half. Is only being read.
Returns
half
‣ The
reciprocal square root of a.
Description
Calculates half reciprocal square root of input a in round-to-nearest mode.
Parameters
a
- half. Is only being read.
[Link]
CUDA Math API vRelease Version | 84
Modules
Returns
half
‣ The
sine of a.
Description
Calculates half sine of input a in round-to-nearest-even mode.
Parameters
a
- half. Is only being read.
Returns
half
‣ The
square root of a.
Description
Calculates half square root of input a in round-to-nearest-even mode.
Parameters
h
- half. Is only being read.
Returns
half
‣ The
truncated integer value.
Description
Round h to the nearest integer value that does not exceed h in magnitude.
[Link]
CUDA Math API vRelease Version | 85
Modules
Parameters
h
- half2. Is only being read.
Returns
half2
‣ The
vector of smallest integers not less than h.
Description
For each component of vector h compute the smallest integer value not less than h.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
elementwise cosine on vector a.
Description
Calculates half2 cosine of input vector a in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 86
Modules
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
elementwise exponential function on vector a.
Description
Calculates half2 exponential function of input vector a in round-to-nearest-even mode.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
elementwise decimal exponential function on vector a.
Description
Calculates half2 decimal exponential function of input vector a in round-to-nearest-
even mode.
Parameters
a
- half2. Is only being read.
[Link]
CUDA Math API vRelease Version | 87
Modules
Returns
half2
‣ The
elementwise binary exponential function on vector a.
Description
Calculates half2 binary exponential function of input vector a in round-to-nearest-even
mode.
Parameters
h
- half2. Is only being read.
Returns
half2
‣ The
vector of largest integers which is less than or equal to h.
Description
For each component of vector h calculate the largest integer value which is less than or
equal to h.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
elementwise natural logarithm on vector a.
[Link]
CUDA Math API vRelease Version | 88
Modules
Description
Calculates half2 natural logarithm of input vector a in round-to-nearest-even mode.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
elementwise decimal logarithm on vector a.
Description
Calculates half2 decimal logarithm of input vector a in round-to-nearest-even mode.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
elementwise binary logarithm on vector a.
Description
Calculates half2 binary logarithm of input vector a in round-to-nearest mode.
[Link]
CUDA Math API vRelease Version | 89
Modules
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
elementwise reciprocal on vector a.
Description
Calculates half2 reciprocal of input vector a in round-to-nearest-even mode.
Parameters
h
- half2. Is only being read.
Returns
half2
‣ The
vector of rounded integer values.
Description
Round each component of half2 vector h to the nearest integer value in half-precision
floating point format, with halfway cases rounded to the nearest even integer value.
Parameters
a
- half2. Is only being read.
[Link]
CUDA Math API vRelease Version | 90
Modules
Returns
half2
‣ The
elementwise reciprocal square root on vector a.
Description
Calculates half2 reciprocal square root of input vector a in round-to-nearest-even
mode.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
elementwise sine on vector a.
Description
Calculates half2 sine of input vector a in round-to-nearest-even mode.
Parameters
a
- half2. Is only being read.
Returns
half2
‣ The
elementwise square root on vector a.
[Link]
CUDA Math API vRelease Version | 91
Modules
Description
Calculates half2 square root of input vector a in round-to-nearest mode.
Parameters
h
- half2. Is only being read.
Returns
half2
‣ The
truncated h.
Description
Round each component of vector h to the nearest integer value that does not exceed h in
magnitude.
[Link]
CUDA Math API vRelease Version | 92
Modules
Description
Divides nv_bfloat162 input vector a by input vector b in round-to-nearest mode.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
absolute value of a.
Description
Calculates the absolute value of input nv_bfloat16 number and returns the result.
Description
Performs nv_bfloat16 addition of inputs a and b, in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 93
Modules
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
sum of a and b, with respect to saturation.
Description
Performs nv_bfloat16 add of inputs a and b, in round-to-nearest-even mode, and
clamps the result to range [0.0, 1.0]. NaN results are flushed to +0.0.
Description
Divides nv_bfloat16 input a by input b in round-to-nearest mode.
Description
Performs nv_bfloat16 multiply on inputs a and b, then performs a nv_bfloat16
add of the result with c, rounding the result once in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 94
Modules
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
c
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
result of fused multiply-add operation on a, b, and c with relu saturation.
Description
Performs nv_bfloat16 multiply on inputs a and b, then performs a nv_bfloat16
add of the result with c, rounding the result once in round-to-nearest-even mode. Then
negative result is clamped to 0. NaN result is converted to canonical NaN.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
c
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
[Link]
CUDA Math API vRelease Version | 95
Modules
Description
Performs nv_bfloat16 multiply on inputs a and b, then performs a nv_bfloat16
add of the result with c, rounding the result once in round-to-nearest-even mode, and
clamps the result to range [0.0, 1.0]. NaN results are flushed to +0.0.
Description
Performs nv_bfloat16 multiplication of inputs a and b, in round-to-nearest mode.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
result of multiplying a and b, with respect to saturation.
Description
Performs nv_bfloat16 multiplication of inputs a and b, in round-to-nearest mode, and
clamps the result to range [0.0, 1.0]. NaN results are flushed to +0.0.
Description
Negates input nv_bfloat16 number and returns the result.
[Link]
CUDA Math API vRelease Version | 96
Modules
Description
Subtracts nv_bfloat16 input b from input a in round-to-nearest mode.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
result of subtraction of b from a, with respect to saturation.
Description
Subtracts nv_bfloat16 input b from input a in round-to-nearest mode, and clamps the
result to range [0.0, 1.0]. NaN results are flushed to +0.0.
Parameters
a
- nv_bfloat162. Is only being read.
[Link]
CUDA Math API vRelease Version | 97
Modules
Returns
bfloat2
‣ Returns
a with the absolute value of both halves.
Description
Calculates the absolute value of both halves of the input nv_bfloat162 number and
returns the result.
Description
Performs nv_bfloat162 vector add of inputs a and b, in round-to-nearest mode.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
sum of a and b, with respect to saturation.
Description
Performs nv_bfloat162 vector add of inputs a and b, in round-to-nearest mode, and
clamps the results to range [0.0, 1.0]. NaN results are flushed to +0.0.
[Link]
CUDA Math API vRelease Version | 98
Modules
Description
Performs nv_bfloat162 vector multiply on inputs a and b, then performs a
nv_bfloat162 vector add of the result with c, rounding the result once in round-to-
nearest-even mode.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
c
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
result of elementwise fused multiply-add operation on vectors a, b, and c with relu
saturation.
Description
Performs nv_bfloat162 vector multiply on inputs a and b, then performs a
nv_bfloat162 vector add of the result with c, rounding the result once in round-to-
nearest-even mode. Then negative result is clamped to 0. NaN result is converted to
canonical NaN.
[Link]
CUDA Math API vRelease Version | 99
Modules
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
c
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
result of elementwise fused multiply-add operation on vectors a, b, and c, with
respect to saturation.
Description
Performs nv_bfloat162 vector multiply on inputs a and b, then performs a
nv_bfloat162 vector add of the result with c, rounding the result once in round-to-
nearest-even mode, and clamps the results to range [0.0, 1.0]. NaN results are flushed to
+0.0.
Description
Performs nv_bfloat162 vector multiplication of inputs a and b, in round-to-nearest-
even mode.
[Link]
CUDA Math API vRelease Version | 100
Modules
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
result of elementwise multiplication of vectors a and b, with respect to saturation.
Description
Performs nv_bfloat162 vector multiplication of inputs a and b, in round-to-nearest-
even mode, and clamps the results to range [0.0, 1.0]. NaN results are flushed to +0.0.
Description
Negates both halves of the input nv_bfloat162 number a and returns the result.
Description
Subtracts nv_bfloat162 input vector b from input vector a in round-to-nearest-even
mode.
[Link]
CUDA Math API vRelease Version | 101
Modules
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
subtraction of vector b from a, with respect to saturation.
Description
Subtracts nv_bfloat162 input vector b from input vector a in round-to-nearest-even
mode, and clamps the results to range [0.0, 1.0]. NaN results are flushed to +0.0.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
[Link]
CUDA Math API vRelease Version | 102
Modules
Description
Performs nv_bfloat16 if-equal comparison of inputs a and b. NaN inputs generate
false results.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
boolean result of unordered if-equal comparison of a and b.
Description
Performs nv_bfloat16 if-equal comparison of inputs a and b. NaN inputs generate
true results.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
[Link]
CUDA Math API vRelease Version | 103
Modules
Description
Performs nv_bfloat16 greater-equal comparison of inputs a and b. NaN inputs
generate false results.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
boolean result of unordered greater-equal comparison of a and b.
Description
Performs nv_bfloat16 greater-equal comparison of inputs a and b. NaN inputs
generate true results.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
[Link]
CUDA Math API vRelease Version | 104
Modules
Description
Performs nv_bfloat16 greater-than comparison of inputs a and b. NaN inputs
generate false results.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
boolean result of unordered greater-than comparison of a and b.
Description
Performs nv_bfloat16 greater-than comparison of inputs a and b. NaN inputs
generate true results.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
int
‣ -1
iff a is equal to negative infinity,
‣ 1
[Link]
CUDA Math API vRelease Version | 105
Modules
Description
Checks if the input nv_bfloat16 number a is infinite.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
bool
‣ true
iff argument is NaN.
Description
Determine whether nv_bfloat16 value a is a NaN.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
boolean result of less-equal comparison of a and b.
[Link]
CUDA Math API vRelease Version | 106
Modules
Description
Performs nv_bfloat16 less-equal comparison of inputs a and b. NaN inputs generate
false results.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
boolean result of unordered less-equal comparison of a and b.
Description
Performs nv_bfloat16 less-equal comparison of inputs a and b. NaN inputs generate
true results.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
boolean result of less-than comparison of a and b.
[Link]
CUDA Math API vRelease Version | 107
Modules
Description
Performs nv_bfloat16 less-than comparison of inputs a and b. NaN inputs generate
false results.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
boolean result of unordered less-than comparison of a and b.
Description
Performs nv_bfloat16 less-than comparison of inputs a and b. NaN inputs generate
true results.
Description
Calculates nv_bfloat16 max(a, b) defined as (a > b) ? a : b.
[Link]
CUDA Math API vRelease Version | 108
Modules
Description
Calculates nv_bfloat16 max(a, b) defined as (a > b) ? a : b.
Description
Calculates nv_bfloat16 min(a, b) defined as (a < b) ? a : b.
Description
Calculates nv_bfloat16 min(a, b) defined as (a < b) ? a : b.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
[Link]
CUDA Math API vRelease Version | 109
Modules
Returns
bool
‣ The
boolean result of not-equal comparison of a and b.
Description
Performs nv_bfloat16 not-equal comparison of inputs a and b. NaN inputs generate
false results.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
bool
‣ The
boolean result of unordered not-equal comparison of a and b.
Description
Performs nv_bfloat16 not-equal comparison of inputs a and b. NaN inputs generate
true results.
[Link]
CUDA Math API vRelease Version | 110
Modules
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
both nv_bfloat16 results of if-equal comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs nv_bfloat162 vector if-equal comparison of inputs a and b. The bool result
is set to true only if both nv_bfloat16 if-equal comparisons evaluate to true, or false
otherwise. NaN inputs generate false results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
both nv_bfloat16 results of unordered if-equal comparison of vectors a and b are
true;
[Link]
CUDA Math API vRelease Version | 111
Modules
‣ falseotherwise.
Description
Performs nv_bfloat162 vector if-equal comparison of inputs a and b. The bool result
is set to true only if both nv_bfloat16 if-equal comparisons evaluate to true, or false
otherwise. NaN inputs generate true results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
both nv_bfloat16 results of greater-equal comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs nv_bfloat162 vector greater-equal comparison of inputs a and b. The bool
result is set to true only if both nv_bfloat16 greater-equal comparisons evaluate to
true, or false otherwise. NaN inputs generate false results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
[Link]
CUDA Math API vRelease Version | 112
Modules
Returns
bool
‣ trueif
both nv_bfloat16 results of unordered greater-equal comparison of vectors a and
b are true;
‣ falseotherwise.
Description
Performs nv_bfloat162 vector greater-equal comparison of inputs a and b. The bool
result is set to true only if both nv_bfloat16 greater-equal comparisons evaluate to
true, or false otherwise. NaN inputs generate true results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
both nv_bfloat16 results of greater-than comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs nv_bfloat162 vector greater-than comparison of inputs a and b. The bool
result is set to true only if both nv_bfloat16 greater-than comparisons evaluate to
true, or false otherwise. NaN inputs generate false results.
[Link]
CUDA Math API vRelease Version | 113
Modules
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
both nv_bfloat16 results of unordered greater-than comparison of vectors a and
b are true;
‣ falseotherwise.
Description
Performs nv_bfloat162 vector greater-than comparison of inputs a and b. The bool
result is set to true only if both nv_bfloat16 greater-than comparisons evaluate to
true, or false otherwise. NaN inputs generate true results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
both nv_bfloat16 results of less-equal comparison of vectors a and b are true;
[Link]
CUDA Math API vRelease Version | 114
Modules
‣ falseotherwise.
Description
Performs nv_bfloat162 vector less-equal comparison of inputs a and b. The bool
result is set to true only if both nv_bfloat16 less-equal comparisons evaluate to true,
or false otherwise. NaN inputs generate false results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
both nv_bfloat16 results of unordered less-equal comparison of vectors a and b
are true;
‣ falseotherwise.
Description
Performs nv_bfloat162 vector less-equal comparison of inputs a and b. The bool
result is set to true only if both nv_bfloat16 less-equal comparisons evaluate to true,
or false otherwise. NaN inputs generate true results.
Parameters
a
- nv_bfloat162. Is only being read.
[Link]
CUDA Math API vRelease Version | 115
Modules
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
both nv_bfloat16 results of less-than comparison of vectors a and b are true;
‣ falseotherwise.
Description
Performs nv_bfloat162 vector less-than comparison of inputs a and b. The bool result
is set to true only if both nv_bfloat16 less-than comparisons evaluate to true, or false
otherwise. NaN inputs generate false results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
both nv_bfloat16 results of unordered less-than comparison of vectors a and b
are true;
‣ falseotherwise.
Description
Performs nv_bfloat162 vector less-than comparison of inputs a and b. The bool result
is set to true only if both nv_bfloat16 less-than comparisons evaluate to true, or false
otherwise. NaN inputs generate true results.
[Link]
CUDA Math API vRelease Version | 116
Modules
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
both nv_bfloat16 results of not-equal comparison of vectors a and b are true,
‣ false
otherwise.
Description
Performs nv_bfloat162 vector not-equal comparison of inputs a and b. The bool
result is set to true only if both nv_bfloat16 not-equal comparisons evaluate to true, or
false otherwise. NaN inputs generate false results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
bool
‣ trueif
[Link]
CUDA Math API vRelease Version | 117
Modules
Description
Performs nv_bfloat162 vector not-equal comparison of inputs a and b. The bool
result is set to true only if both nv_bfloat16 not-equal comparisons evaluate to true, or
false otherwise. NaN inputs generate true results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
vector result of if-equal comparison of vectors a and b.
Description
Performs nv_bfloat162 vector if-equal comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate false results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
[Link]
CUDA Math API vRelease Version | 118
Modules
Returns
nv_bfloat162
‣ The
vector result of unordered if-equal comparison of vectors a and b.
Description
Performs nv_bfloat162 vector if-equal comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate true results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
vector result of greater-equal comparison of vectors a and b.
Description
Performs nv_bfloat162 vector greater-equal comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate false results.
Parameters
a
- nv_bfloat162. Is only being read.
[Link]
CUDA Math API vRelease Version | 119
Modules
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 vector result of unordered greater-equal comparison of vectors a
and b.
Description
Performs nv_bfloat162 vector greater-equal comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate true results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
vector result of greater-than comparison of vectors a and b.
Description
Performs nv_bfloat162 vector greater-than comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate false results.
[Link]
CUDA Math API vRelease Version | 120
Modules
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 vector result of unordered greater-than comparison of vectors a
and b.
Description
Performs nv_bfloat162 vector greater-than comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate true results.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 with the corresponding nv_bfloat16 results set to 1.0 for for NaN,
0.0 otherwise.
Description
Determine whether each nv_bfloat16 of input nv_bfloat162 number a is a NaN.
[Link]
CUDA Math API vRelease Version | 121
Modules
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 result of less-equal comparison of vectors a and b.
Description
Performs nv_bfloat162 vector less-equal comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate false results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
vector result of unordered less-equal comparison of vectors a and b.
[Link]
CUDA Math API vRelease Version | 122
Modules
Description
Performs nv_bfloat162 vector less-equal comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate true results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 vector result of less-than comparison of vectors a and b.
Description
Performs nv_bfloat162 vector less-than comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate false results.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
[Link]
CUDA Math API vRelease Version | 123
Modules
Description
Performs nv_bfloat162 vector less-than comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate true results.
Description
Calculates nv_bfloat162 vector max(a, b) Elementwise nv_bfloat16 operation is
defined as (a > b) ? a : b.
Description
Calculates nv_bfloat162 vector max(a, b) Elementwise nv_bfloat16 operation is
defined as (a > b) ? a : b.
Description
Calculates nv_bfloat162 vector min(a, b) Elementwise nv_bfloat16 operation is
defined as (a < b) ? a : b.
[Link]
CUDA Math API vRelease Version | 124
Modules
Description
Calculates nv_bfloat162 vector min(a, b) Elementwise nv_bfloat16 operation is
defined as (a < b) ? a : b.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
vector result of not-equal comparison of vectors a and b.
Description
Performs nv_bfloat162 vector not-equal comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate false results.
Parameters
a
- nv_bfloat162. Is only being read.
[Link]
CUDA Math API vRelease Version | 125
Modules
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
vector result of unordered not-equal comparison of vectors a and b.
Description
Performs nv_bfloat162 vector not-equal comparison of inputs a and b. The
corresponding nv_bfloat16 results are set to 1.0 for true, or 0.0 for false. NaN inputs
generate true results.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
float2
‣ \p
a converted to float2.
Description
Converts both halves of nv_bfloat162 input a to float2 and returns the result.
[Link]
CUDA Math API vRelease Version | 126
Modules
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat162
‣ The
vector which has both its halves equal to the input a.
Description
Returns nv_bfloat162 number with both halves equal to the input a nv_bfloat16
number.
Parameters
a
- float. Is only being read.
Returns
float
‣ \p
a converted to float.
Description
Converts nv_bfloat16 number a to float.
[Link]
CUDA Math API vRelease Version | 127
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
int
‣ \p
h converted to a signed integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed integer in round-
down mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
int
‣ \p
h converted to a signed integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed integer in round-to-
nearest-even mode.
[Link]
CUDA Math API vRelease Version | 128
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
int
‣ \p
h converted to a signed integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed integer in round-up
mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
int
‣ \p
h converted to a signed integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed integer in round-
towards-zero mode.
[Link]
CUDA Math API vRelease Version | 129
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
long long int
‣ \p
h converted to a signed 64-bit integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed 64-bit integer in
round-down mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
long long int
‣ \p
h converted to a signed 64-bit integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed 64-bit integer in
round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 130
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
long long int
‣ \p
h converted to a signed 64-bit integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed 64-bit integer in
round-up mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
long long int
‣ \p
h converted to a signed 64-bit integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed 64-bit integer in
round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 131
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
short int
‣ \p
h converted to a signed short integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed short integer in
round-down mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
short int
‣ \p
h converted to a signed short integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed short integer in
round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 132
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
short int
‣ \p
h converted to a signed short integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed short integer in
round-up mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
short int
‣ \p
h converted to a signed short integer.
Description
Convert the nv_bfloat16-precision floating point value h to a signed short integer in
round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 133
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned int
‣ \p
h converted to an unsigned integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned integer in
round-down mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned int
‣ \p
h converted to an unsigned integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned integer in
round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 134
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned int
‣ \p
h converted to an unsigned integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned integer in
round-up mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned int
‣ \p
h converted to an unsigned integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned integer in
round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 135
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned long long int
‣ \p
h converted to an unsigned 64-bit integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned 64-bit integer in
round-down mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned long long int
‣ \p
h converted to an unsigned 64-bit integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned 64-bit integer in
round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 136
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned long long int
‣ \p
h converted to an unsigned 64-bit integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned 64-bit integer in
round-up mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned long long int
‣ \p
h converted to an unsigned 64-bit integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned 64-bit integer in
round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 137
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned short int
‣ \p
h converted to an unsigned short integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned short integer in
round-down mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned short int
‣ \p
h converted to an unsigned short integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned short integer in
round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 138
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned short int
‣ \p
h converted to an an unsigned short integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned short integer in
round-up mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned short int
‣ \p
h converted to an unsigned short integer.
Description
Convert the nv_bfloat16-precision floating point value h to an unsigned short integer in
round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 139
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
short int
‣ The
reinterpreted value.
Description
Reinterprets the bits in the nv_bfloat16-precision floating point number h as a signed
short integer.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
unsigned short int
‣ The
reinterpreted value.
Description
Reinterprets the bits in the nv_bfloat16-precision floating point h as an unsigned short
number.
[Link]
CUDA Math API vRelease Version | 140
Modules
Parameters
a
- double. Is only being read.
Returns
nv_bfloat16
‣ \p
a converted to nv_bfloat16.
Description
Converts double number a to nv_bfloat16 precision in round-to-nearest-even mode.
Parameters
a
- float2. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 which has corresponding halves equal to the converted float2
components.
Description
Converts both components of float2 to nv_bfloat16 precision in round-to-nearest mode
and combines the results into one nv_bfloat162 number. Low 16 bits of the return
value correspond to a.x and high 16 bits of the return value correspond to a.y.
[Link]
CUDA Math API vRelease Version | 141
Modules
Parameters
a
- float. Is only being read.
Returns
nv_bfloat16
‣ \p
a converted to nv_bfloat16.
Description
Converts float number a to nv_bfloat16 precision in round-to-nearest-even mode.
Parameters
a
- float. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 value with both halves equal to the converted nv_bfloat16 precision
number.
Description
Converts input a to nv_bfloat16 precision in round-to-nearest-even mode and populates
both halves of nv_bfloat162 with converted value.
[Link]
CUDA Math API vRelease Version | 142
Modules
Parameters
a
- float. Is only being read.
Returns
nv_bfloat16
‣ \p
a converted to nv_bfloat16.
Description
Converts float number a to nv_bfloat16 precision in round-down mode.
Parameters
a
- float. Is only being read.
Returns
nv_bfloat16
‣ \p
a converted to nv_bfloat16.
Description
Converts float number a to nv_bfloat16 precision in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 143
Modules
Parameters
a
- float. Is only being read.
Returns
nv_bfloat16
‣ \p
a converted to nv_bfloat16.
Description
Converts float number a to nv_bfloat16 precision in round-up mode.
Parameters
a
- float. Is only being read.
Returns
nv_bfloat16
‣ \p
a converted to nv_bfloat16.
Description
Converts float number a to nv_bfloat16 precision in round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 144
Modules
Parameters
a
- float. Is only being read.
b
- float. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 value with corresponding halves equal to the converted input
floats.
Description
Converts both input floats to nv_bfloat16 precision in round-to-nearest-even mode and
combines the results into one nv_bfloat162 number. Low 16 bits of the return value
correspond to the input a, high 16 bits correspond to the input b.
Parameters
a
- nv_bfloat16. Is only being read.
b
- nv_bfloat16. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 with one nv_bfloat16 equal to a and the other to b.
[Link]
CUDA Math API vRelease Version | 145
Modules
Description
Combines two input nv_bfloat16 number a and b into one nv_bfloat162 number.
Input a is stored in low 16 bits of the return value, input b is stored in high 16 bits of the
return value.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat16
‣ The
high 16 bits of the input.
Description
Returns high 16 bits of nv_bfloat162 input a.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 with both halves equal to the high 16 bits of the input.
Description
Extracts high 16 bits from nv_bfloat162 input a and returns a new nv_bfloat162
number which has both halves equal to the extracted bits.
[Link]
CUDA Math API vRelease Version | 146
Modules
Parameters
a
- nv_bfloat162. Is only being read.
Returns
float
‣ The
high 16 bits of a converted to float.
Description
Converts high 16 bits of nv_bfloat162 input a to 32 bit floating point number and
returns the result.
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
high 16 bits of a and of b.
Description
Extracts high 16 bits from each of the two nv_bfloat162 inputs and combines into one
nv_bfloat162 number. High 16 bits from input a is stored in low 16 bits of the return
value, high 16 bits from input b is stored in high 16 bits of the return value.
[Link]
CUDA Math API vRelease Version | 147
Modules
Parameters
i
- int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed integer value i to a nv_bfloat16-precision floating point value in
round-down mode.
Parameters
i
- int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed integer value i to a nv_bfloat16-precision floating point value in
round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 148
Modules
Parameters
i
- int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed integer value i to a nv_bfloat16-precision floating point value in
round-up mode.
Parameters
i
- int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed integer value i to a nv_bfloat16-precision floating point value in
round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 149
Modules
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
[Link]
CUDA Math API vRelease Version | 150
Modules
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
[Link]
CUDA Math API vRelease Version | 151
Modules
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
Parameters
ptr
- memory location
Returns
The value pointed by `ptr`
[Link]
CUDA Math API vRelease Version | 152
Modules
Parameters
i
- long long int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed 64-bit integer value i to a nv_bfloat16-precision floating point value
in round-down mode.
Parameters
i
- long long int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed 64-bit integer value i to a nv_bfloat16-precision floating point value
in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 153
Modules
Parameters
i
- long long int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed 64-bit integer value i to a nv_bfloat16-precision floating point value
in round-up mode.
Parameters
i
- long long int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed 64-bit integer value i to a nv_bfloat16-precision floating point value
in round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 154
Modules
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat16
‣ Returns
nv_bfloat16 which contains low 16 bits of the input a.
Description
Returns low 16 bits of nv_bfloat162 input a.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
nv_bfloat162 with both halves equal to the low 16 bits of the input.
Description
Extracts low 16 bits from nv_bfloat162 input a and returns a new nv_bfloat162
number which has both halves equal to the extracted bits.
[Link]
CUDA Math API vRelease Version | 155
Modules
Parameters
a
- nv_bfloat162. Is only being read.
Returns
float
‣ The
low 16 bits of a converted to float.
Description
Converts low 16 bits of nv_bfloat162 input a to 32 bit floating point number and
returns the result.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ \p
a with its halves being swapped.
Description
Swaps both halves of the nv_bfloat162 input and returns a new nv_bfloat162
number with swapped halves.
[Link]
CUDA Math API vRelease Version | 156
Modules
Parameters
a
- nv_bfloat162. Is only being read.
b
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
low 16 bits of a and of b.
Description
Extracts low 16 bits from each of the two nv_bfloat162 inputs and combines into one
nv_bfloat162 number. Low 16 bits from input a is stored in low 16 bits of the return
value, low 16 bits from input b is stored in high 16 bits of the return value.
Parameters
mask
- unsigned int. Is only being read.
var
- nv_bfloat16. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
[Link]
CUDA Math API vRelease Version | 157
Modules
Returns
Returns the 2-byte word referenced by var from the source thread ID as nv_bfloat16. If
the source thread ID is out of range or the source thread has exited, the calling thread's
own var is returned.
Description
Calculates a source thread ID by adding delta to the caller's thread ID. The value
of var held by the resulting thread ID is returned: this has the effect of shifting var
down the warp by delta threads. If width is less than warpSize then each subsection
of the warp behaves as a separate entity with a starting logical thread ID of 0. As for
__shfl_up_sync(), the ID number of the source thread will not wrap around the value of
width and so the upper delta threads will remain unchanged.
Parameters
mask
- unsigned int. Is only being read.
var
- nv_bfloat162. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 4-byte word referenced by var from the source thread ID as nv_bfloat162. If
the source thread ID is out of range or the source thread has exited, the calling thread's
own var is returned.
Description
Calculates a source thread ID by adding delta to the caller's thread ID. The value
of var held by the resulting thread ID is returned: this has the effect of shifting var
down the warp by delta threads. If width is less than warpSize then each subsection
of the warp behaves as a separate entity with a starting logical thread ID of 0. As for
__shfl_up_sync(), the ID number of the source thread will not wrap around the value of
width and so the upper delta threads will remain unchanged.
[Link]
CUDA Math API vRelease Version | 158
Modules
Parameters
mask
- unsigned int. Is only being read.
var
- nv_bfloat16. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 2-byte word referenced by var from the source thread ID as nv_bfloat16. If
the source thread ID is out of range or the source thread has exited, the calling thread's
own var is returned.
Description
Returns the value of var held by the thread whose ID is given by delta. If width is less
than warpSize then each subsection of the warp behaves as a separate entity with a
starting logical thread ID of 0. If delta is outside the range [0:width-1], the value returned
corresponds to the value of var held by the delta modulo width (i.e. ithin the same
subsection). width must have a value which is a power of 2; results are undefined if
width is not a power of 2, or is a number greater than warpSize.
Parameters
mask
- unsigned int. Is only being read.
var
- nv_bfloat162. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
[Link]
CUDA Math API vRelease Version | 159
Modules
Returns
Returns the 4-byte word referenced by var from the source thread ID as nv_bfloat162. If
the source thread ID is out of range or the source thread has exited, the calling thread's
own var is returned.
Description
Returns the value of var held by the thread whose ID is given by delta. If width is less
than warpSize then each subsection of the warp behaves as a separate entity with a
starting logical thread ID of 0. If delta is outside the range [0:width-1], the value returned
corresponds to the value of var held by the delta modulo width (i.e. ithin the same
subsection). width must have a value which is a power of 2; results are undefined if
width is not a power of 2, or is a number greater than warpSize.
Parameters
mask
- unsigned int. Is only being read.
var
- nv_bfloat16. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 2-byte word referenced by var from the source thread ID as nv_bfloat16. If
the source thread ID is out of range or the source thread has exited, the calling thread's
own var is returned.
Description
Calculates a source thread ID by subtracting delta from the caller's lane ID. The value
of var held by the resulting lane ID is returned: in effect, var is shifted up the warp by
delta threads. If width is less than warpSize then each subsection of the warp behaves
as a separate entity with a starting logical thread ID of 0. The source thread index
will not wrap around the value of width, so effectively the lower delta threads will be
unchanged. width must have a value which is a power of 2; results are undefined if
width is not a power of 2, or is a number greater than warpSize.
[Link]
CUDA Math API vRelease Version | 160
Modules
Parameters
mask
- unsigned int. Is only being read.
var
- nv_bfloat162. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 4-byte word referenced by var from the source thread ID as nv_bfloat162. If
the source thread ID is out of range or the source thread has exited, the calling thread's
own var is returned.
Description
Calculates a source thread ID by subtracting delta from the caller's lane ID. The value
of var held by the resulting lane ID is returned: in effect, var is shifted up the warp by
delta threads. If width is less than warpSize then each subsection of the warp behaves
as a separate entity with a starting logical thread ID of 0. The source thread index
will not wrap around the value of width, so effectively the lower delta threads will be
unchanged. width must have a value which is a power of 2; results are undefined if
width is not a power of 2, or is a number greater than warpSize.
Parameters
mask
- unsigned int. Is only being read.
var
- nv_bfloat16. Is only being read.
delta
- int. Is only being read.
[Link]
CUDA Math API vRelease Version | 161
Modules
width
- int. Is only being read.
Returns
Returns the 2-byte word referenced by var from the source thread ID as nv_bfloat16. If
the source thread ID is out of range or the source thread has exited, the calling thread's
own var is returned.
Description
Calculates a source thread ID by performing a bitwise XOR of the caller's thread ID with
mask: the value of var held by the resulting thread ID is returned. If width is less than
warpSize then each group of width consecutive threads are able to access elements from
earlier groups of threads, however if they attempt to access elements from later groups
of threads their own value of var will be returned. This mode implements a butterfly
addressing pattern such as is used in tree reduction and broadcast.
Parameters
mask
- unsigned int. Is only being read.
var
- nv_bfloat162. Is only being read.
delta
- int. Is only being read.
width
- int. Is only being read.
Returns
Returns the 4-byte word referenced by var from the source thread ID as nv_bfloat162. If
the source thread ID is out of range or the source thread has exited, the calling thread's
own var is returned.
Description
Calculates a source thread ID by performing a bitwise XOR of the caller's thread ID with
mask: the value of var held by the resulting thread ID is returned. If width is less than
warpSize then each group of width consecutive threads are able to access elements from
earlier groups of threads, however if they attempt to access elements from later groups
[Link]
CUDA Math API vRelease Version | 162
Modules
of threads their own value of var will be returned. This mode implements a butterfly
addressing pattern such as is used in tree reduction and broadcast.
Parameters
i
- short int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed short integer value i to a nv_bfloat16-precision floating point value
in round-down mode.
Parameters
i
- short int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed short integer value i to a nv_bfloat16-precision floating point value
in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 163
Modules
Parameters
i
- short int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed short integer value i to a nv_bfloat16-precision floating point value
in round-up mode.
Parameters
i
- short int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the signed short integer value i to a nv_bfloat16-precision floating point value
in round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 164
Modules
Parameters
i
- short int. Is only being read.
Returns
nv_bfloat16
‣ The
reinterpreted value.
Description
Reinterprets the bits in the signed short integer i as a nv_bfloat16-precision floating
point number.
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
ptr
- memory location
value
- the value to be stored
[Link]
CUDA Math API vRelease Version | 165
Modules
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
ptr
- memory location
value
- the value to be stored
[Link]
CUDA Math API vRelease Version | 166
Modules
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
ptr
- memory location
value
- the value to be stored
Parameters
i
- unsigned int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned integer value i to a nv_bfloat16-precision floating point value in
round-down mode.
[Link]
CUDA Math API vRelease Version | 167
Modules
Parameters
i
- unsigned int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned integer value i to a nv_bfloat16-precision floating point value in
round-to-nearest-even mode.
Parameters
i
- unsigned int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned integer value i to a nv_bfloat16-precision floating point value in
round-up mode.
[Link]
CUDA Math API vRelease Version | 168
Modules
Parameters
i
- unsigned int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned integer value i to a nv_bfloat16-precision floating point value in
round-towards-zero mode.
Parameters
i
- unsigned long long int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned 64-bit integer value i to a nv_bfloat16-precision floating point
value in round-down mode.
[Link]
CUDA Math API vRelease Version | 169
Modules
Parameters
i
- unsigned long long int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned 64-bit integer value i to a nv_bfloat16-precision floating point
value in round-to-nearest-even mode.
Parameters
i
- unsigned long long int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned 64-bit integer value i to a nv_bfloat16-precision floating point
value in round-up mode.
[Link]
CUDA Math API vRelease Version | 170
Modules
Parameters
i
- unsigned long long int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned 64-bit integer value i to a nv_bfloat16-precision floating point
value in round-towards-zero mode.
Parameters
i
- unsigned short int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned short integer value i to a nv_bfloat16-precision floating point
value in round-down mode.
[Link]
CUDA Math API vRelease Version | 171
Modules
Parameters
i
- unsigned short int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned short integer value i to a nv_bfloat16-precision floating point
value in round-to-nearest-even mode.
Parameters
i
- unsigned short int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned short integer value i to a nv_bfloat16-precision floating point
value in round-up mode.
[Link]
CUDA Math API vRelease Version | 172
Modules
Parameters
i
- unsigned short int. Is only being read.
Returns
nv_bfloat16
‣ \p
i converted to nv_bfloat16.
Description
Convert the unsigned short integer value i to a nv_bfloat16-precision floating point
value in round-towards-zero mode.
Parameters
i
- unsigned short int. Is only being read.
Returns
nv_bfloat16
‣ The
reinterpreted value.
Description
Reinterprets the bits in the unsigned short integer i as a nv_bfloat16-precision floating
point number.
[Link]
CUDA Math API vRelease Version | 173
Modules
To use these functions include the header file cuda_bf16.h in your program.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
smallest integer value not less than h.
Description
Compute the smallest integer value not less than h.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
cosine of a.
Description
Calculates nv_bfloat16 cosine of input a in round-to-nearest-even mode.
Parameters
a
- nv_bfloat16. Is only being read.
[Link]
CUDA Math API vRelease Version | 174
Modules
Returns
nv_bfloat16
‣ The
natural exponential function on a.
Description
Calculates nv_bfloat16 natural exponential function of input a in round-to-nearest-
even mode.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
decimal exponential function on a.
Description
Calculates nv_bfloat16 decimal exponential function of input a in round-to-nearest-
even mode.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
binary exponential function on a.
[Link]
CUDA Math API vRelease Version | 175
Modules
Description
Calculates nv_bfloat16 binary exponential function of input a in round-to-nearest-
even mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
largest integer value which is less than or equal to h.
Description
Calculate the largest integer value which is less than or equal to h.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
natural logarithm of a.
Description
Calculates nv_bfloat16 natural logarithm of input a in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 176
Modules
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
decimal logarithm of a.
Description
Calculates nv_bfloat16 decimal logarithm of input a in round-to-nearest-even mode.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
binary logarithm of a.
Description
Calculates nv_bfloat16 binary logarithm of input a in round-to-nearest-even mode.
Parameters
a
- nv_bfloat16. Is only being read.
[Link]
CUDA Math API vRelease Version | 177
Modules
Returns
nv_bfloat16
‣ The
reciprocal of a.
Description
Calculates nv_bfloat16 reciprocal of input a in round-to-nearest-even mode.
Parameters
h
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
nearest integer to h.
Description
Round h to the nearest integer value in nv_bfloat16-precision floating point format, with
bfloat16way cases rounded to the nearest even integer value.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
reciprocal square root of a.
[Link]
CUDA Math API vRelease Version | 178
Modules
Description
Calculates nv_bfloat16 reciprocal square root of input a in round-to-nearest mode.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
sine of a.
Description
Calculates nv_bfloat16 sine of input a in round-to-nearest-even mode.
Parameters
a
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
square root of a.
Description
Calculates nv_bfloat16 square root of input a in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 179
Modules
Parameters
h
- nv_bfloat16. Is only being read.
Returns
nv_bfloat16
‣ The
truncated integer value.
Description
Round h to the nearest integer value that does not exceed h in magnitude.
Parameters
h
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
vector of smallest integers not less than h.
Description
For each component of vector h compute the smallest integer value not less than h.
[Link]
CUDA Math API vRelease Version | 180
Modules
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
elementwise cosine on vector a.
Description
Calculates nv_bfloat162 cosine of input vector a in round-to-nearest-even mode.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
elementwise exponential function on vector a.
Description
Calculates nv_bfloat162 exponential function of input vector a in round-to-nearest-
even mode.
[Link]
CUDA Math API vRelease Version | 181
Modules
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
elementwise decimal exponential function on vector a.
Description
Calculates nv_bfloat162 decimal exponential function of input vector a in round-to-
nearest-even mode.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
elementwise binary exponential function on vector a.
Description
Calculates nv_bfloat162 binary exponential function of input vector a in round-to-
nearest-even mode.
[Link]
CUDA Math API vRelease Version | 182
Modules
Parameters
h
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
vector of largest integers which is less than or equal to h.
Description
For each component of vector h calculate the largest integer value which is less than or
equal to h.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
elementwise natural logarithm on vector a.
Description
Calculates nv_bfloat162 natural logarithm of input vector a in round-to-nearest-even
mode.
[Link]
CUDA Math API vRelease Version | 183
Modules
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
elementwise decimal logarithm on vector a.
Description
Calculates nv_bfloat162 decimal logarithm of input vector a in round-to-nearest-even
mode.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
elementwise binary logarithm on vector a.
Description
Calculates nv_bfloat162 binary logarithm of input vector a in round-to-nearest mode.
Parameters
a
- nv_bfloat162. Is only being read.
[Link]
CUDA Math API vRelease Version | 184
Modules
Returns
nv_bfloat162
‣ The
elementwise reciprocal on vector a.
Description
Calculates nv_bfloat162 reciprocal of input vector a in round-to-nearest-even mode.
Parameters
h
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
vector of rounded integer values.
Description
Round each component of nv_bfloat162 vector h to the nearest integer value in
nv_bfloat16-precision floating point format, with bfloat16way cases rounded to the
nearest even integer value.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
elementwise reciprocal square root on vector a.
[Link]
CUDA Math API vRelease Version | 185
Modules
Description
Calculates nv_bfloat162 reciprocal square root of input vector a in round-to-nearest-
even mode.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
elementwise sine on vector a.
Description
Calculates nv_bfloat162 sine of input vector a in round-to-nearest-even mode.
Parameters
a
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
elementwise square root on vector a.
Description
Calculates nv_bfloat162 square root of input vector a in round-to-nearest mode.
[Link]
CUDA Math API vRelease Version | 186
Modules
Parameters
h
- nv_bfloat162. Is only being read.
Returns
nv_bfloat162
‣ The
truncated h.
Description
Round each component of vector h to the nearest integer value that does not exceed h in
magnitude.
Returns
Result will be in radians, in the interval [0, ] for x inside [-1, +1].
[Link]
CUDA Math API vRelease Version | 187
Modules
Description
Calculate the principal value of the arc cosine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Result will be in the interval [0, ].
‣ acoshf(1) returns 0.
‣ acoshf(x) returns NaN for x in the interval [ , 1).
Description
Calculate the nonnegative arc hyperbolic cosine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Description
Calculate the principal value of the arc sine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 188
Modules
Returns
‣ asinhf(0) returns 1.
Description
Calculate the arc hyperbolic sine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Result will be in radians, in the interval [- ,+ ].
Description
Calculate the principal value of the arc tangent of the ratio of first and second input
arguments y / x. The quadrant of the result is determined by the signs of inputs y and x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Description
Calculate the principal value of the arc tangent of the input argument x.
[Link]
CUDA Math API vRelease Version | 189
Modules
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ atanhf( ) returns .
‣ atanhf( ) returns .
‣ atanhf(x) returns NaN for x outside interval [-1, 1].
Description
Calculate the arc hyperbolic tangent of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
‣ cbrtf( ) returns .
‣ cbrtf( ) returns .
Description
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 190
Modules
Returns
Returns expressed as a floating-point number.
‣ ceilf( ) returns .
‣ ceilf( ) returns .
Description
Compute the smallest integer value not less than x.
Returns
Returns a value with the magnitude of x and the sign of y.
Description
Create a floating-point value with the magnitude x and the sign of y.
Returns
‣ cosf(0) returns 1.
‣ cosf( ) returns NaN.
Description
Calculate the cosine of the input argument x (measured in radians).
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This function is affected by the --use_fast_math compiler flag. See
the CUDA C Programming Guide, Appendix D.2, Table 8 for a complete list of
functions affected.
[Link]
CUDA Math API vRelease Version | 191
Modules
Returns
‣ coshf(0) returns 1.
‣ coshf( ) returns NaN.
Description
Calculate the hyperbolic cosine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ cospif( ) returns 1.
‣ cospif( ) returns NaN.
Description
Calculate the cosine of x (measured in radians), where x is the input argument.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the value of the regular modified cylindrical Bessel function of order 0.
Description
Calculate the value of the regular modified cylindrical Bessel function of order 0 for the
input argument x, .
[Link]
CUDA Math API vRelease Version | 192
Modules
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the value of the regular modified cylindrical Bessel function of order 1.
Description
Calculate the value of the regular modified cylindrical Bessel function of order 1 for the
input argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ erfcf( ) returns 2.
‣ erfcf( ) returns +0.
Description
Calculate the complementary error function of the input argument x, 1 - erf(x).
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ erfcinvf(0) returns .
[Link]
CUDA Math API vRelease Version | 193
Modules
‣ erfcinvf(2) returns .
Description
Calculate the inverse complementary error function of the input argument y, for y in the
interval [0, 2]. The inverse complementary error function find the value x that satisfies
the equation y = erfc(x), for , and .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ erfcxf( ) returns
‣ erfcxf( ) returns +0
‣ erfcxf(x) returns if the correctly calculated value is outside the single floating
point range.
Description
Calculate the scaled complementary error function of the input argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ erff( ) returns .
‣ erff( ) returns .
Description
Calculate the value of the error function for the input argument x, .
[Link]
CUDA Math API vRelease Version | 194
Modules
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ erfinvf(1) returns .
‣ erfinvf(-1) returns .
Description
Calculate the inverse error function of the input argument y, for y in the interval [-1,
1]. The inverse error function finds the value x that satisfies the equation y = erf(x), for
, and .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
Description
Calculate the base 10 exponential of the input argument x.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This function is affected by the --use_fast_math compiler flag. See
the CUDA C Programming Guide, Appendix D.2, Table 8 for a complete list of
functions affected.
[Link]
CUDA Math API vRelease Version | 195
Modules
Returns
Returns .
Description
Calculate the base 2 exponential of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
Description
Calculate the base exponential of the input argument x, .
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This function is affected by the --use_fast_math compiler flag. See
the CUDA C Programming Guide, Appendix D.2, Table 8 for a complete list of
functions affected.
Returns
Returns .
Description
Calculate the base exponential of the input argument x, minus 1.
[Link]
CUDA Math API vRelease Version | 196
Modules
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the absolute value of its argument.
‣ fabs( ) returns .
‣ fabs( ) returns 0.
Description
Calculate the absolute value of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the positive difference between x and y.
Description
Compute the positive difference between x and y. The positive difference is x - y when x
> y and +0 otherwise.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 197
Modules
Returns
Returns x / y.
Description
Compute x divided by y. If --use_fast_math is specified, use __fdividef() for higher
performance, otherwise use normal division.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This function is affected by the --use_fast_math compiler flag. See
the CUDA C Programming Guide, Appendix D.2, Table 8 for a complete list of
functions affected.
Returns
Returns expressed as a floating-point number.
‣ floorf( ) returns .
‣ floorf( ) returns .
Description
Calculate the largest integer value which is less than or equal to x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the rounded value of as a single operation.
[Link]
CUDA Math API vRelease Version | 198
Modules
Description
Compute the value of as a single ternary operation. After computing the value
to infinite precision, the value is rounded once.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the maximum numeric values of the arguments x and y.
Description
Determines the maximum numeric value of the arguments x and y. Treats NaN
arguments as missing data. If one argument is a NaN and the other is legitimate numeric
value, the numeric value is chosen.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the minimum numeric values of the arguments x and y.
[Link]
CUDA Math API vRelease Version | 199
Modules
Description
Determines the minimum numeric value of the arguments x and y. Treats NaN
arguments as missing data. If one argument is a NaN and the other is legitimate numeric
value, the numeric value is chosen.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ Returns the floating point remainder of x / y.
‣ fmodf( , y) returns if y is not zero.
‣ fmodf(x, ) returns x if x is finite.
‣ fmodf(x, y) returns NaN if x is or y is zero.
‣ If either argument is NaN, NaN is returned.
Description
Calculate the floating-point remainder of x / y. The floating-point remainder of the
division operation x / y calculated by this function is exactly the value x - n*y, where
n is x / y with its fractional part truncated. The computed value will have the same sign
as x, and it's magnitude will be less than the magnitude of y.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the fractional component m.
‣ frexp(0, nptr) returns 0 for the fractional component and zero for the integer
component.
‣ frexp( , nptr) returns and stores zero in the location pointed to by nptr.
‣ frexp( , nptr) returns and stores an unspecified value in the location to
which nptr points.
[Link]
CUDA Math API vRelease Version | 200
Modules
Description
Decomposes the floating-point value x into a component m for the normalized fraction
element and another term n for the exponent. The absolute value of m will be greater
than or equal to 0.5 and less than 1.0 or it will be equal to 0; . The integer
exponent n will be stored in the location to which nptr points.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the length of the hypotenuse . If the correct value would overflow,
returns . If the correct value would underflow, returns 0.
Description
Calculates the length of the hypotenuse of a right triangle whose two sides have lengths
x and y without undue overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ If successful, returns the unbiased exponent of the argument.
‣ ilogbf(0) returns INT_MIN.
‣ ilogbf(NaN) returns INT_MIN.
‣ ilogbf(x) returns INT_MAX if x is or the correct value is greater than INT_MAX.
‣ ilogbf(x) return INT_MIN if the correct value is less than INT_MIN.
[Link]
CUDA Math API vRelease Version | 201
Modules
Description
Calculates the unbiased integer exponent of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ With Visual Studio 2013 host compiler: __RETURN_TYPE is 'bool'. Returns true if
and only if a is a finite value.
‣ With other host compilers: __RETURN_TYPE is 'int'. Returns a nonzero value if and
only if a is a finite value.
Description
Determine whether the floating-point value a is a finite value (zero, subnormal, or
normal and not infinity or NaN).
Returns
‣ With Visual Studio 2013 host compiler: __RETURN_TYPE is 'bool'. Returns true if
and only if a is a infinite value.
‣ With other host compilers: __RETURN_TYPE is 'int'. Returns a nonzero value if and
only if a is a infinite value.
Description
Determine whether the floating-point value a is an infinite value (positive or negative).
Returns
‣ With Visual Studio 2013 host compiler: __RETURN_TYPE is 'bool'. Returns true if
and only if a is a NaN value.
[Link]
CUDA Math API vRelease Version | 202
Modules
‣ With other host compilers: __RETURN_TYPE is 'int'. Returns a nonzero value if and
only if a is a NaN value.
Description
Determine whether the floating-point value a is a NaN.
Returns
Returns the value of the Bessel function of the first kind of order 0.
Description
Calculate the value of the Bessel function of the first kind of order 0 for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the value of the Bessel function of the first kind of order 1.
‣ j1f( ) returns .
‣ j1f( ) returns .
‣ j1f(NaN) returns NaN.
Description
Calculate the value of the Bessel function of the first kind of order 1 for the input
argument x, .
[Link]
CUDA Math API vRelease Version | 203
Modules
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the value of the Bessel function of the first kind of order n.
Description
Calculate the value of the Bessel function of the first kind of order n for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ ldexpf(x) returns if the correctly calculated value is outside the single floating
point range.
Description
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 204
Modules
Returns
‣ lgammaf(1) returns +0.
‣ lgammaf(2) returns +0.
‣ lgammaf(x) returns if the correctly calculated value is outside the single floating
point range.
‣ lgammaf(x) returns if x 0 and x is an integer.
‣ lgammaf( ) returns .
‣ lgammaf( ) returns .
Description
Calculate the natural logarithm of the absolute value of the gamma function of the input
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value, with halfway cases rounded to the nearest even
integer value. If the result is outside the range of the return type, the result is undefined.
Returns
Returns rounded integer value.
[Link]
CUDA Math API vRelease Version | 205
Modules
Description
Round x to the nearest integer value, with halfway cases rounded away from zero. If the
result is outside the range of the return type, the result is undefined.
This function may be slower than alternate rounding methods. See llrintf().
Returns
‣ log10f( ) returns .
‣ log10f(1) returns +0.
‣ log10f(x) returns NaN for x < 0.
‣ log10f( ) returns .
Description
Calculate the base 10 logarithm of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ log1pf( ) returns .
‣ log1pf(-1) returns +0.
‣ log1pf(x) returns NaN for x < -1.
‣ log1pf( ) returns .
Description
Calculate the value of of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 206
Modules
Returns
‣ log2f( ) returns .
‣ log2f(1) returns +0.
‣ log2f(x) returns NaN for x < 0.
‣ log2f( ) returns .
Description
Calculate the base 2 logarithm of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ logbf returns
‣ logbf returns
Description
Calculate the floating point representation of the exponent of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ logf( ) returns .
‣ logf(1) returns +0.
‣ logf(x) returns NaN for x < 0.
‣ logf( ) returns .
[Link]
CUDA Math API vRelease Version | 207
Modules
Description
Calculate the natural logarithm of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value, with halfway cases rounded to the nearest even
integer value. If the result is outside the range of the return type, the result is undefined.
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value, with halfway cases rounded away from zero. If the
result is outside the range of the return type, the result is undefined.
This function may be slower than alternate rounding methods. See lrintf().
Returns
‣ modff( , iptr) returns a result with the same sign as x.
‣ modff( , iptr) returns and stores in the object pointed to by iptr.
‣ modff(NaN, iptr) stores a NaN in the object pointed to by iptr and returns a
NaN.
[Link]
CUDA Math API vRelease Version | 208
Modules
Description
Break down the argument x into fractional and integral parts. The integral part is stored
in the argument iptr. Fractional and integral parts are given the same sign as the
argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ nanf(tagp) returns NaN.
Description
Return a representation of a quiet NaN. Argument tagp selects one of the possible
representations.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ nearbyintf( ) returns .
‣ nearbyintf( ) returns .
Description
Round argument x to an integer value in single precision floating-point format.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 209
Modules
Returns
‣ nextafterf( , y) returns .
Description
Calculate the next representable single-precision floating-point value following x in
the direction of y. For example, if y is greater than x, nextafterf() returns the smallest
representable number greater than x
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Description
Calculates the length of three dimensional vector p in euclidean space without undue
overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
[Link]
CUDA Math API vRelease Version | 210
Modules
Description
Calculates the length of four dimensional vector p in euclidean space without undue
overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ normcdff( ) returns 1
‣ normcdff( ) returns +0
Description
Calculate the cumulative distribution function of the standard normal distribution for
input argument y, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ normcdfinvf(0) returns .
‣ normcdfinvf(1) returns .
‣ normcdfinvf(x) returns NaN if x is not in the interval [0,1].
Description
Calculate the inverse of the standard normal cumulative distribution function for input
argument y, . The function is defined for input values in the interval .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 211
Modules
Returns
Description
Calculates the length of a vector p, dimension of which is passed as an agument without
undue overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ powf( , y) returns for y an integer less than 0.
‣ powf( , y) returns for y an odd integer greater than 0.
‣ powf( , y) returns +0 for y > 0 and not and odd integer.
‣ powf(-1, ) returns 1.
‣ powf(+1, y) returns 1 for any y, even a NaN.
‣ powf(x, ) returns 1 for any x, even a NaN.
‣ powf(x, y) returns a NaN for finite x < 0 and finite non-integer y.
‣ powf(x, ) returns for .
‣ powf(x, ) returns +0 for .
‣ powf(x, ) returns +0 for .
‣ powf(x, ) returns for .
‣ powf( , y) returns -0 for y an odd integer less than 0.
‣ powf( , y) returns +0 for y < 0 and not an odd integer.
‣ powf( , y) returns for y an odd integer greater than 0.
‣ powf( , y) returns for y > 0 and not an odd integer.
‣ powf( , y) returns +0 for y < 0.
‣ powf( , y) returns for y > 0.
[Link]
CUDA Math API vRelease Version | 212
Modules
Description
Calculate the value of x to the power of y.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ rcbrt( ) returns .
‣ rcbrt( ) returns .
Description
Calculate reciprocal cube root function of x
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ remainderf(x, 0) returns NaN.
‣ remainderf( , y) returns NaN.
‣ remainderf(x, ) returns x for finite x.
Description
Compute single-precision floating-point remainder r of dividing x by y for nonzero y.
Thus . The value n is the integer value nearest . In the case when ,
the even n value is chosen.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 213
Modules
Returns
Returns the remainder.
Description
Compute a double-precision floating-point remainder in the same way as the
remainderf() function. Argument quo returns part of quotient upon division of x by y.
Value quo has the same sign as and may not be the exact quotient but agrees with the
exact quotient in the low order 3 bits.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns one over the length of the hypotenuse . If the square root would
Description
Calculates one over the length of the hypotenuse of a right triangle whose two sides
have lengths x and y without undue overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 214
Modules
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value in floating-point format, with halfway cases
rounded to the nearest even integer value.
Returns
Returns one over the length of the 3D vector . If the square root
Description
Calculates one over the length of three dimension vector p in euclidean space without
undue overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns one over the length of the 3D vector . If the square root
[Link]
CUDA Math API vRelease Version | 215
Modules
Description
Calculates one over the length of four dimension vector p in euclidean space without
undue overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns one over the length of the vector . If the square root
Description
Calculates one over the length of vector p, dimension of which is passed as an agument,
in euclidean space without undue overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value in floating-point format, with halfway cases
rounded away from zero.
This function may be slower than alternate rounding methods. See rintf().
[Link]
CUDA Math API vRelease Version | 216
Modules
Returns
Returns .
Description
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns x * .
‣ scalblnf( , n) returns .
‣ scalblnf(x, 0) returns x.
‣ scalblnf( , n) returns .
Description
Returns
Returns x * .
‣ scalbnf( , n) returns .
‣ scalbnf(x, 0) returns x.
‣ scalbnf( , n) returns .
[Link]
CUDA Math API vRelease Version | 217
Modules
Description
Returns
Reports the sign bit of all values including infinities, zeros, and NaNs.
‣ With Visual Studio 2013 host compiler: __RETURN_TYPE is 'bool'. Returns true if
and only if a is negative.
‣ With other host compilers: __RETURN_TYPE is 'int'. Returns a nonzero value if and
only if a is negative.
Description
Determine whether the floating-point value a is negative.
Returns
‣ none
Description
Calculate the sine and cosine of the first input argument x (measured in radians). The
results for sine and cosine are written into the second argument, sptr, and, respectively,
third argument, cptr.
See also:
sinf() and cosf().
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This function is affected by the --use_fast_math compiler flag. See
the CUDA C Programming Guide, Appendix D.2, Table 8 for a complete list of
functions affected.
[Link]
CUDA Math API vRelease Version | 218
Modules
Returns
‣ none
Description
Calculate the sine and cosine of the first input argument, x (measured in radians),
. The results for sine and cosine are written into the second argument, sptr, and,
respectively, third argument, cptr.
See also:
sinpif() and cospif().
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ sinf( ) returns .
‣ sinf( ) returns NaN.
Description
Calculate the sine of the input argument x (measured in radians).
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This function is affected by the --use_fast_math compiler flag. See
the CUDA C Programming Guide, Appendix D.2, Table 8 for a complete list of
functions affected.
[Link]
CUDA Math API vRelease Version | 219
Modules
Returns
‣ sinhf( ) returns .
‣ sinhf( ) returns NaN.
Description
Calculate the hyperbolic sine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ sinpif( ) returns .
‣ sinpif( ) returns NaN.
Description
Calculate the sine of x (measured in radians), where x is the input argument.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
‣ sqrtf( ) returns .
‣ sqrtf( ) returns .
‣ sqrtf(x) returns NaN if x is less than 0.
[Link]
CUDA Math API vRelease Version | 220
Modules
Description
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
‣ tanf( ) returns .
‣ tanf( ) returns NaN.
Description
Calculate the tangent of the input argument x (measured in radians).
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This function is affected by the --use_fast_math compiler flag. See
the CUDA C Programming Guide, Appendix D.2, Table 8 for a complete list of
functions affected.
Returns
‣ tanhf( ) returns .
Description
Calculate the hyperbolic tangent of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 221
Modules
Returns
‣ tgammaf( ) returns .
‣ tgammaf(2) returns +1.
‣ tgammaf(x) returns if the correctly calculated value is outside the single floating
point range.
‣ tgammaf(x) returns NaN if x < 0 and x is an integer.
‣ tgammaf( ) returns NaN.
‣ tgammaf( ) returns .
Description
Calculate the gamma function of the input argument x, namely the value of .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns truncated integer value.
Description
Round x to the nearest integer value that does not exceed x in magnitude.
Returns
Returns the value of the Bessel function of the second kind of order 0.
‣ y0f(0) returns .
‣ y0f(x) returns NaN for x < 0.
‣ y0f( ) returns +0.
[Link]
CUDA Math API vRelease Version | 222
Modules
Description
Calculate the value of the Bessel function of the second kind of order 0 for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the value of the Bessel function of the second kind of order 1.
‣ y1f(0) returns .
‣ y1f(x) returns NaN for x < 0.
‣ y1f( ) returns +0.
‣ y1f(NaN) returns NaN.
Description
Calculate the value of the Bessel function of the second kind of order 1 for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the value of the Bessel function of the second kind of order n.
[Link]
CUDA Math API vRelease Version | 223
Modules
Description
Calculate the value of the Bessel function of the second kind of order n for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Result will be in radians, in the interval [0, ] for x inside [-1, +1].
Description
Calculate the principal value of the arc cosine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Result will be in the interval [0, ].
‣ acosh(1) returns 0.
‣ acosh(x) returns NaN for x in the interval [ , 1).
[Link]
CUDA Math API vRelease Version | 224
Modules
Description
Calculate the nonnegative arc hyperbolic cosine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Result will be in radians, in the interval [- /2, + /2] for x inside [-1, +1].
Description
Calculate the principal value of the arc sine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ asinh(0) returns 1.
Description
Calculate the arc hyperbolic sine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 225
Modules
Returns
Result will be in radians, in the interval [- /2, + /2].
Description
Calculate the principal value of the arc tangent of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Result will be in radians, in the interval [- /, + ].
Description
Calculate the principal value of the arc tangent of the ratio of first and second input
arguments y / x. The quadrant of the result is determined by the signs of inputs y and x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ atanh( ) returns .
‣ atanh( ) returns .
‣ atanh(x) returns NaN for x outside interval [-1, 1].
[Link]
CUDA Math API vRelease Version | 226
Modules
Description
Calculate the arc hyperbolic tangent of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns .
‣ cbrt( ) returns .
‣ cbrt( ) returns .
Description
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns expressed as a floating-point number.
‣ ceil( ) returns .
‣ ceil( ) returns .
Description
Compute the smallest integer value not less than x.
[Link]
CUDA Math API vRelease Version | 227
Modules
Returns
Returns a value with the magnitude of x and the sign of y.
Description
Create a floating-point value with the magnitude x and the sign of y.
Returns
‣ cos( ) returns 1.
‣ cos( ) returns NaN.
Description
Calculate the cosine of the input argument x (measured in radians).
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ cosh(0) returns 1.
‣ cosh( ) returns .
Description
Calculate the hyperbolic cosine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 228
Modules
Returns
‣ cospi( ) returns 1.
‣ cospi( ) returns NaN.
Description
Calculate the cosine of x (measured in radians), where x is the input argument.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the value of the regular modified cylindrical Bessel function of order 0.
Description
Calculate the value of the regular modified cylindrical Bessel function of order 0 for the
input argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the value of the regular modified cylindrical Bessel function of order 1.
[Link]
CUDA Math API vRelease Version | 229
Modules
Description
Calculate the value of the regular modified cylindrical Bessel function of order 1 for the
input argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ erf( ) returns .
‣ erf( ) returns .
Description
Calculate the value of the error function for the input argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ erfc( ) returns 2.
‣ erfc( ) returns +0.
Description
Calculate the complementary error function of the input argument x, 1 - erf(x).
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 230
Modules
Returns
‣ erfcinv(0) returns .
‣ erfcinv(2) returns .
Description
Calculate the inverse complementary error function of the input argument y, for y in the
interval [0, 2]. The inverse complementary error function find the value x that satisfies
the equation y = erfc(x), for , and .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ erfcx( ) returns
‣ erfcx( ) returns +0
‣ erfcx(x) returns if the correctly calculated value is outside the double floating
point range.
Description
Calculate the scaled complementary error function of the input argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ erfinv(1) returns .
‣ erfinv(-1) returns .
[Link]
CUDA Math API vRelease Version | 231
Modules
Description
Calculate the inverse error function of the input argument y, for y in the interval [-1,
1]. The inverse error function finds the value x that satisfies the equation y = erf(x), for
, and .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns .
Description
Calculate the base exponential of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns .
Description
Calculate the base 10 exponential of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 232
Modules
Returns
Returns .
Description
Calculate the base 2 exponential of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns .
Description
Calculate the base exponential of the input argument x, minus 1.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the absolute value of the input argument.
‣ fabs( ) returns .
‣ fabs( ) returns 0.
Description
Calculate the absolute value of the input argument x.
[Link]
CUDA Math API vRelease Version | 233
Modules
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the positive difference between x and y.
Description
Compute the positive difference between x and y. The positive difference is x - y when x
> y and +0 otherwise.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns expressed as a floating-point number.
‣ floor( ) returns .
‣ floor( ) returns .
Description
Calculates the largest integer value which is less than or equal to x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 234
Modules
Returns
Returns the rounded value of as a single operation.
Description
Compute the value of as a single ternary operation. After computing the value
to infinite precision, the value is rounded once.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the maximum numeric values of the arguments x and y.
Description
Determines the maximum numeric value of the arguments x and y. Treats NaN
arguments as missing data. If one argument is a NaN and the other is legitimate numeric
value, the numeric value is chosen.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 235
Modules
Returns
Returns the minimum numeric values of the arguments x and y.
Description
Determines the minimum numeric value of the arguments x and y. Treats NaN
arguments as missing data. If one argument is a NaN and the other is legitimate numeric
value, the numeric value is chosen.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ Returns the floating point remainder of x / y.
‣ fmod( , y) returns if y is not zero.
‣ fmod(x, ) returns x if x is finite.
‣ fmod(x, y) returns NaN if x is or y is zero.
‣ If either argument is NaN, NaN is returned.
Description
Calculate the double-precision floating-point remainder of x / y. The floating-point
remainder of the division operation x / y calculated by this function is exactly the value
x - n*y, where n is x / y with its fractional part truncated. The computed value will
have the same sign as x, and it's magnitude will be less than the magnitude of y.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 236
Modules
Returns
Returns the fractional component m.
‣ frexp(0, nptr) returns 0 for the fractional component and zero for the integer
component.
‣ frexp( , nptr) returns and stores zero in the location pointed to by nptr.
‣ frexp( , nptr) returns and stores an unspecified value in the location to
which nptr points.
‣ frexp(NaN, y) returns a NaN and stores an unspecified value in the location to
which nptr points.
Description
Decompose the floating-point value x into a component m for the normalized fraction
element and another term n for the exponent. The absolute value of m will be greater
than or equal to 0.5 and less than 1.0 or it will be equal to 0; . The integer
exponent n will be stored in the location to which nptr points.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the length of the hypotenuse . If the correct value would overflow,
returns . If the correct value would underflow, returns 0.
Description
Calculate the length of the hypotenuse of a right triangle whose two sides have lengths x
and y without undue overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 237
Modules
Returns
‣ If successful, returns the unbiased exponent of the argument.
‣ ilogb(0) returns INT_MIN.
‣ ilogb(NaN) returns INT_MIN.
‣ ilogb(x) returns INT_MAX if x is or the correct value is greater than INT_MAX.
‣ ilogb(x) return INT_MIN if the correct value is less than INT_MIN.
Description
Calculates the unbiased integer exponent of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ With Visual Studio 2013 host compiler: __RETURN_TYPE is 'bool'. Returns true if
and only if a is a finite value.
‣ With other host compilers: __RETURN_TYPE is 'int'. Returns a nonzero value if and
only if a is a finite value.
Description
Determine whether the floating-point value a is a finite value (zero, subnormal, or
normal and not infinity or NaN).
Returns
‣ With Visual Studio 2013 host compiler: Returns true if and only if a is a infinite
value.
‣ With other host compilers: Returns a nonzero value if and only if a is a infinite
value.
[Link]
CUDA Math API vRelease Version | 238
Modules
Description
Determine whether the floating-point value a is an infinite value (positive or negative).
Returns
‣ With Visual Studio 2013 host compiler: __RETURN_TYPE is 'bool'. Returns true if
and only if a is a NaN value.
‣ With other host compilers: __RETURN_TYPE is 'int'. Returns a nonzero value if and
only if a is a NaN value.
Description
Determine whether the floating-point value a is a NaN.
Returns
Returns the value of the Bessel function of the first kind of order 0.
Description
Calculate the value of the Bessel function of the first kind of order 0 for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the value of the Bessel function of the first kind of order 1.
[Link]
CUDA Math API vRelease Version | 239
Modules
‣ j1( ) returns .
‣ j1( ) returns .
‣ j1(NaN) returns NaN.
Description
Calculate the value of the Bessel function of the first kind of order 1 for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the value of the Bessel function of the first kind of order n.
Description
Calculate the value of the Bessel function of the first kind of order n for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ ldexp(x) returns if the correctly calculated value is outside the double floating
point range.
[Link]
CUDA Math API vRelease Version | 240
Modules
Description
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ lgamma(1) returns +0.
‣ lgamma(2) returns +0.
‣ lgamma(x) returns if the correctly calculated value is outside the double
floating point range.
‣ lgamma(x) returns if x 0 and x is an integer.
‣ lgamma( ) returns .
‣ lgamma( ) returns .
Description
Calculate the natural logarithm of the absolute value of the gamma function of the input
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value, with halfway cases rounded to the nearest even
integer value. If the result is outside the range of the return type, the result is undefined.
[Link]
CUDA Math API vRelease Version | 241
Modules
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value, with halfway cases rounded away from zero. If the
result is outside the range of the return type, the result is undefined.
This function may be slower than alternate rounding methods. See llrint().
Returns
‣ log( ) returns .
‣ log(1) returns +0.
‣ log(x) returns NaN for x < 0.
‣ log( ) returns
Description
Calculate the base logarithm of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ log10( ) returns .
‣ log10(1) returns +0.
‣ log10(x) returns NaN for x < 0.
‣ log10( ) returns .
[Link]
CUDA Math API vRelease Version | 242
Modules
Description
Calculate the base 10 logarithm of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ log1p( ) returns .
‣ log1p(-1) returns +0.
‣ log1p(x) returns NaN for x < -1.
‣ log1p( ) returns .
Description
Calculate the value of of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ log2( ) returns .
‣ log2(1) returns +0.
‣ log2(x) returns NaN for x < 0.
‣ log2( ) returns .
Description
Calculate the base 2 logarithm of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 243
Modules
Returns
‣ logb returns
‣ logb returns
Description
Calculate the floating point representation of the exponent of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value, with halfway cases rounded to the nearest even
integer value. If the result is outside the range of the return type, the result is undefined.
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value, with halfway cases rounded away from zero. If the
result is outside the range of the return type, the result is undefined.
This function may be slower than alternate rounding methods. See lrint().
[Link]
CUDA Math API vRelease Version | 244
Modules
Returns
‣ modf( , iptr) returns a result with the same sign as x.
‣ modf( , iptr) returns and stores in the object pointed to by iptr.
‣ modf(NaN, iptr) stores a NaN in the object pointed to by iptr and returns a NaN.
Description
Break down the argument x into fractional and integral parts. The integral part is stored
in the argument iptr. Fractional and integral parts are given the same sign as the
argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ nan(tagp) returns NaN.
Description
Return a representation of a quiet NaN. Argument tagp selects one of the possible
representations.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ nearbyint( ) returns .
‣ nearbyint( ) returns .
[Link]
CUDA Math API vRelease Version | 245
Modules
Description
Round argument x to an integer value in double precision floating-point format.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ nextafter( , y) returns .
Description
Calculate the next representable double-precision floating-point value following x in
the direction of y. For example, if y is greater than x, nextafter() returns the smallest
representable number greater than x
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Description
Calculate the length of a vector p, dimension of which is passed as an argument
without undue overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 246
Modules
Returns
Description
Calculate the length of three dimensional vector p in euclidean space without undue
overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Description
Calculate the length of four dimensional vector p in euclidean space without undue
overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ normcdf( ) returns 1
[Link]
CUDA Math API vRelease Version | 247
Modules
‣ normcdf( ) returns +0
Description
Calculate the cumulative distribution function of the standard normal distribution for
input argument y, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ normcdfinv(0) returns .
‣ normcdfinv(1) returns .
‣ normcdfinv(x) returns NaN if x is not in the interval [0,1].
Description
Calculate the inverse of the standard normal cumulative distribution function for input
argument y, . The function is defined for input values in the interval .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ pow( , y) returns for y an integer less than 0.
‣ pow( , y) returns for y an odd integer greater than 0.
‣ pow( , y) returns +0 for y > 0 and not and odd integer.
‣ pow(-1, ) returns 1.
‣ pow(+1, y) returns 1 for any y, even a NaN.
‣ pow(x, ) returns 1 for any x, even a NaN.
‣ pow(x, y) returns a NaN for finite x < 0 and finite non-integer y.
‣ pow(x, ) returns for .
‣ pow(x, ) returns +0 for .
[Link]
CUDA Math API vRelease Version | 248
Modules
Description
Calculate the value of x to the power of y
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ rcbrt( ) returns .
‣ rcbrt( ) returns .
Description
Calculate reciprocal cube root function of x
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ remainder(x, 0) returns NaN.
‣ remainder( , y) returns NaN.
‣ remainder(x, ) returns x for finite x.
[Link]
CUDA Math API vRelease Version | 249
Modules
Description
Compute double-precision floating-point remainder r of dividing x by y for nonzero y.
Thus . The value n is the integer value nearest . In the case when ,
the even n value is chosen.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the remainder.
Description
Compute a double-precision floating-point remainder in the same way as the
remainder() function. Argument quo returns part of quotient upon division of x by y.
Value quo has the same sign as and may not be the exact quotient but agrees with the
exact quotient in the low order 3 bits.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns one over the length of the hypotenuse . If the square root would
[Link]
CUDA Math API vRelease Version | 250
Modules
Description
Calculate one over the length of the hypotenuse of a right triangle whose two sides have
lengths x and y without undue overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value in floating-point format, with halfway cases
rounded to the nearest even integer value.
Returns
Returns one over the length of the vector . If the square root
Description
Calculates one over the length of vector p, dimension of which is passed as an agument,
in euclidean space without undue overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 251
Modules
Returns
Returns one over the length of the 3D vetor . If the square root would
Description
Calculate one over the length of three dimensional vector p in euclidean space undue
overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns one over the length of the 3D vetor . If the square root
Description
Calculate one over the length of four dimensional vector p in euclidean space undue
overflow or underflow.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 252
Modules
Returns
Returns rounded integer value.
Description
Round x to the nearest integer value in floating-point format, with halfway cases
rounded away from zero.
This function may be slower than alternate rounding methods. See rint().
Returns
Returns .
Description
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns x * .
‣ scalbln( , n) returns .
‣ scalbln(x, 0) returns x.
[Link]
CUDA Math API vRelease Version | 253
Modules
‣ scalbln( , n) returns .
Description
Returns
Returns x * .
‣ scalbn( , n) returns .
‣ scalbn(x, 0) returns x.
‣ scalbn( , n) returns .
Description
Returns
Reports the sign bit of all values including infinities, zeros, and NaNs.
‣ With Visual Studio 2013 host compiler: __RETURN_TYPE is 'bool'. Returns true if
and only if a is negative.
‣ With other host compilers: __RETURN_TYPE is 'int'. Returns a nonzero value if and
only if a is negative.
Description
Determine whether the floating-point value a is negative.
Returns
‣ sin( ) returns .
‣ sin( ) returns NaN.
[Link]
CUDA Math API vRelease Version | 254
Modules
Description
Calculate the sine of the input argument x (measured in radians).
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ none
Description
Calculate the sine and cosine of the first input argument x (measured in radians). The
results for sine and cosine are written into the second argument, sptr, and, respectively,
third argument, cptr.
See also:
sin() and cos().
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ none
Description
Calculate the sine and cosine of the first input argument, x (measured in radians),
. The results for sine and cosine are written into the second argument, sptr, and,
respectively, third argument, cptr.
See also:
[Link]
CUDA Math API vRelease Version | 255
Modules
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ sinh( ) returns .
Description
Calculate the hyperbolic sine of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ sinpi( ) returns .
‣ sinpi( ) returns NaN.
Description
Calculate the sine of x (measured in radians), where x is the input argument.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns .
[Link]
CUDA Math API vRelease Version | 256
Modules
‣ sqrt( ) returns .
‣ sqrt( ) returns .
‣ sqrt(x) returns NaN if x is less than 0.
Description
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ tan( ) returns .
‣ tan( ) returns NaN.
Description
Calculate the tangent of the input argument x (measured in radians).
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
‣ tanh( ) returns .
Description
Calculate the hyperbolic tangent of the input argument x.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 257
Modules
Returns
‣ tgamma( ) returns .
‣ tgamma(2) returns +1.
‣ tgamma(x) returns if the correctly calculated value is outside the double
floating point range.
‣ tgamma(x) returns NaN if x < 0 and x is an integer.
‣ tgamma( ) returns NaN.
‣ tgamma( ) returns .
Description
Calculate the gamma function of the input argument x, namely the value of .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns truncated integer value.
Description
Round x to the nearest integer value that does not exceed x in magnitude.
Returns
Returns the value of the Bessel function of the second kind of order 0.
‣ y0(0) returns .
‣ y0(x) returns NaN for x < 0.
‣ y0( ) returns +0.
[Link]
CUDA Math API vRelease Version | 258
Modules
Description
Calculate the value of the Bessel function of the second kind of order 0 for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the value of the Bessel function of the second kind of order 1.
‣ y1(0) returns .
‣ y1(x) returns NaN for x < 0.
‣ y1( ) returns +0.
‣ y1(NaN) returns NaN.
Description
Calculate the value of the Bessel function of the second kind of order 1 for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the value of the Bessel function of the second kind of order n.
[Link]
CUDA Math API vRelease Version | 259
Modules
Description
Calculate the value of the Bessel function of the second kind of order n for the input
argument x, .
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the approximate cosine of x.
Description
Calculate the fast approximate cosine of the input argument x, measured in radians.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
‣ Input and output in the denormal range is flushed to sign preserving 0.0.
Returns
Returns an approximation to .
[Link]
CUDA Math API vRelease Version | 260
Modules
Description
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
‣ Most input and output values around denormal range are flushed to sign
preserving 0.0.
Returns
Returns an approximation to .
Description
Calculate the fast approximate base exponential of the input argument x, .
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
‣ Most input and output values around denormal range are flushed to sign
preserving 0.0.
Returns
Returns x + y.
Description
Compute the sum of x and y in round-down (to negative infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
[Link]
CUDA Math API vRelease Version | 261
Modules
Returns
Returns x + y.
Description
Compute the sum of x and y in round-to-nearest-even rounding mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x + y.
Description
Compute the sum of x and y in round-up (to positive infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x + y.
Description
Compute the sum of x and y in round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 262
Modules
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x / y.
Description
Divide two floating point values x by y in round-down (to negative infinity) mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns x / y.
Description
Divide two floating point values x by y in round-to-nearest-even mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns x / y.
[Link]
CUDA Math API vRelease Version | 263
Modules
Description
Divide two floating point values x by y in round-up (to positive infinity) mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns x / y.
Description
Divide two floating point values x by y in round-towards-zero mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns x / y.
Description
Calculate the fast approximate division of x by y.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
[Link]
CUDA Math API vRelease Version | 264
Modules
Returns
Returns the rounded value of as a single operation.
Description
Computes the value of as a single ternary operation, rounding the result once
in round-down (to negative infinity) mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the rounded value of as a single operation.
Description
Computes the value of as a single ternary operation, rounding the result once
in round-to-nearest-even mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 265
Modules
Returns
Returns the rounded value of as a single operation.
Description
Computes the value of as a single ternary operation, rounding the result once
in round-up (to positive infinity) mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns the rounded value of as a single operation.
Description
Computes the value of as a single ternary operation, rounding the result once
in round-towards-zero mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 266
Modules
Returns
Returns x * y.
Description
Compute the product of x and y in round-down (to negative infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x * y.
Description
Compute the product of x and y in round-to-nearest-even mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x * y.
Description
Compute the product of x and y in round-up (to positive infinity) mode.
[Link]
CUDA Math API vRelease Version | 267
Modules
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x * y.
Description
Compute the product of x and y in round-towards-zero mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns .
Description
Compute the reciprocal of x in round-down (to negative infinity) mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
[Link]
CUDA Math API vRelease Version | 268
Modules
Description
Compute the reciprocal of x in round-to-nearest-even mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
Description
Compute the reciprocal of x in round-up (to positive infinity) mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
Description
Compute the reciprocal of x in round-towards-zero mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
[Link]
CUDA Math API vRelease Version | 269
Modules
Returns
Returns .
Description
Compute the reciprocal square root of x in round-to-nearest-even mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
Description
Compute the square root of x in round-down (to negative infinity) mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
Description
Compute the square root of x in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 270
Modules
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
Description
Compute the square root of x in round-up (to positive infinity) mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns .
Description
Compute the square root of x in round-towards-zero mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
Returns
Returns x - y.
[Link]
CUDA Math API vRelease Version | 271
Modules
Description
Compute the difference of x and y in round-down (to negative infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x - y.
Description
Compute the difference of x and y in round-to-nearest-even rounding mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x - y.
Description
Compute the difference of x and y in round-up (to positive infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
[Link]
CUDA Math API vRelease Version | 272
Modules
Returns
Returns x - y.
Description
Compute the difference of x and y in round-towards-zero mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 6.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns an approximation to .
Description
Calculate the fast approximate base 10 logarithm of the input argument x.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
‣ Most input and output values around denormal range are flushed to sign
preserving 0.0.
Returns
Returns an approximation to .
Description
Calculate the fast approximate base 2 logarithm of the input argument x.
[Link]
CUDA Math API vRelease Version | 273
Modules
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
‣ Input and output in the denormal range is flushed to sign preserving 0.0.
Returns
Returns an approximation to .
Description
Calculate the fast approximate base logarithm of the input argument x.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
‣ Most input and output values around denormal range are flushed to sign
preserving 0.0.
Returns
Returns an approximation to .
Description
Calculate the fast approximate of x, the first input argument, raised to the power of y,
the second input argument, .
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
‣ Most input and output values around denormal range are flushed to sign
preserving 0.0.
[Link]
CUDA Math API vRelease Version | 274
Modules
Returns
‣ __saturatef(x) returns 0 if x < 0.
‣ __saturatef(x) returns 1 if x > 1.
‣ __saturatef(x) returns x if .
‣ __saturatef(NaN) returns 0.
Description
Clamp the input argument x to be within the interval [+0.0, 1.0].
Returns
‣ none
Description
Calculate the fast approximate of sine and cosine of the first input argument x
(measured in radians). The results for sine and cosine are written into the second
argument, sptr, and, respectively, third argument, cptr.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
‣ Denorm input/output is flushed to sign preserving 0.0.
Returns
Returns the approximate sine of x.
Description
Calculate the fast approximate sine of the input argument x, measured in radians.
[Link]
CUDA Math API vRelease Version | 275
Modules
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
‣ Input and output in the denormal range is flushed to sign preserving 0.0.
Returns
Returns the approximate tangent of x.
Description
Calculate the fast approximate tangent of the input argument x, measured in radians.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.2, Table 9.
‣ The result is computed as the fast divide of __sinf() by __cosf(). Denormal input
and output are flushed to sign-preserving 0.0 at each step of the computation.
Returns
Returns x + y.
Description
Adds two floating point values x and y in round-down (to negative infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 276
Modules
Returns
Returns x + y.
Description
Adds two floating point values x and y in round-to-nearest-even mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x + y.
Description
Adds two floating point values x and y in round-up (to positive infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x + y.
Description
Adds two floating point values x and y in round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 277
Modules
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x / y.
Description
Divides two floating point values x by y in round-down (to negative infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
Returns
Returns x / y.
Description
Divides two floating point values x by y in round-to-nearest-even mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
Returns
Returns x / y.
[Link]
CUDA Math API vRelease Version | 278
Modules
Description
Divides two floating point values x by y in round-up (to positive infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
Returns
Returns x / y.
Description
Divides two floating point values x by y in round-towards-zero mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
Returns
Returns x * y.
Description
Multiplies two floating point values x and y in round-down (to negative infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
[Link]
CUDA Math API vRelease Version | 279
Modules
Returns
Returns x * y.
Description
Multiplies two floating point values x and y in round-to-nearest-even mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x * y.
Description
Multiplies two floating point values x and y in round-up (to positive infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x * y.
Description
Multiplies two floating point values x and y in round-towards-zero mode.
[Link]
CUDA Math API vRelease Version | 280
Modules
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns .
Description
Compute the reciprocal of x in round-down (to negative infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
Returns
Returns .
Description
Compute the reciprocal of x in round-to-nearest-even mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
[Link]
CUDA Math API vRelease Version | 281
Modules
Returns
Returns .
Description
Compute the reciprocal of x in round-up (to positive infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
Returns
Returns .
Description
Compute the reciprocal of x in round-towards-zero mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
Returns
Returns .
Description
Compute the square root of x in round-down (to negative infinity) mode.
[Link]
CUDA Math API vRelease Version | 282
Modules
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
Returns
Returns .
Description
Compute the square root of x in round-to-nearest-even mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
Returns
Returns .
Description
Compute the square root of x in round-up (to positive infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
[Link]
CUDA Math API vRelease Version | 283
Modules
Returns
Returns .
Description
Compute the square root of x in round-towards-zero mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ Requires compute capability >= 2.0.
Returns
Returns x - y.
Description
Subtracts two floating point values x and y in round-down (to negative infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x - y.
Description
Subtracts two floating point values x and y in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 284
Modules
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x - y.
Description
Subtracts two floating point values x and y in round-up (to positive infinity) mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
Returns
Returns x - y.
Description
Subtracts two floating point values x and y in round-towards-zero mode.
‣ For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
‣ This operation will never be merged into a single multiply-add instruction.
[Link]
CUDA Math API vRelease Version | 285
Modules
Returns
Returns the rounded value of as a single operation.
Description
Computes the value of as a single ternary operation, rounding the result once
in round-down (to negative infinity) mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the rounded value of as a single operation.
Description
Computes the value of as a single ternary operation, rounding the result once
in round-to-nearest-even mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 286
Modules
Returns
Returns the rounded value of as a single operation.
Description
Computes the value of as a single ternary operation, rounding the result once
in round-up (to positive infinity) mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
Returns
Returns the rounded value of as a single operation.
Description
Computes the value of as a single ternary operation, rounding the result once
in round-towards-zero mode.
For accuracy information for this function see the CUDA C Programming Guide,
Appendix D.1, Table 7.
[Link]
CUDA Math API vRelease Version | 287
Modules
Returns
Returns the bit-reversed value of x. i.e. bit N of the return value corresponds to bit 31-N
of x.
Description
Reverses the bit order of the 32 bit unsigned integer x.
Returns
Returns the bit-reversed value of x. i.e. bit N of the return value corresponds to bit 63-N
of x.
Description
Reverses the bit order of the 64 bit unsigned integer x.
Returns
The returned value r is computed to be: result[n] := input[selector[n]]
where result[n] is the nth byte of r.
[Link]
CUDA Math API vRelease Version | 288
Modules
Description
byte_perm(x,y,s) returns a 32-bit integer consisting of four bytes from eight input bytes
provided in the two input integers x and y, as specified by a selector, s.
The input bytes are indexed as follows: input[0] = x<7:0> input[1] = x<15:8> input[2]
= x<23:16> input[3] = x<31:24> input[4] = y<7:0> input[5] = y<15:8> input[6] = y<23:16>
input[7] = y<31:24> The selector indices are as follows (the upper 16-bits of the selector
are not used): selector[0] = s<2:0> selector[1] = s<6:4> selector[2] = s<10:8> selector[3] =
s<14:12>
Returns
Returns a value between 0 and 32 inclusive representing the number of zero bits.
Description
Count the number of consecutive leading zero bits, starting at the most significant bit
(bit 31) of x.
Returns
Returns a value between 0 and 64 inclusive representing the number of zero bits.
Description
Count the number of consecutive leading zero bits, starting at the most significant bit
(bit 63) of x.
Returns
Returns a value between 0 and 32 inclusive representing the position of the first bit set.
‣ __ffs(0) returns 0.
[Link]
CUDA Math API vRelease Version | 289
Modules
Description
Find the position of the first (least significant) bit set to 1 in x, where the least significant
bit position is 1.
Returns
Returns a value between 0 and 64 inclusive representing the position of the first bit set.
‣ __ffsll(0) returns 0.
Description
Find the position of the first (least significant) bit set to 1 in x, where the least significant
bit position is 1.
Returns
Returns the most significant 32 bits of the shifted 64-bit value.
Description
Shift the 64-bit value formed by concatenating argument lo and hi left by the amount
specified by the argument shift. Argument lo holds bits 31:0 and argument hi holds
bits 63:32 of the 64-bit source value. The source is shifted left by the wrapped value of
shift (shift & 31). The most significant 32-bits of the result are returned.
Returns
Returns the most significant 32 bits of the shifted 64-bit value.
Description
Shift the 64-bit value formed by concatenating argument lo and hi left by the amount
specified by the argument shift. Argument lo holds bits 31:0 and argument hi holds
[Link]
CUDA Math API vRelease Version | 290
Modules
bits 63:32 of the 64-bit source value. The source is shifted left by the clamped value of
shift (min(shift, 32)). The most significant 32-bits of the result are returned.
Returns
Returns the least significant 32 bits of the shifted 64-bit value.
Description
Shift the 64-bit value formed by concatenating argument lo and hi right by the amount
specified by the argument shift. Argument lo holds bits 31:0 and argument hi holds
bits 63:32 of the 64-bit source value. The source is shifted right by the wrapped value of
shift (shift & 31). The least significant 32-bits of the result are returned.
Returns
Returns the least significant 32 bits of the shifted 64-bit value.
Description
Shift the 64-bit value formed by concatenating argument lo and hi right by the amount
specified by the argument shift. Argument lo holds bits 31:0 and argument hi holds
bits 63:32 of the 64-bit source value. The source is shifted right by the clamped value of
shift (min(shift, 32)). The least significant 32-bits of the result are returned.
Returns
Returns a signed integer value representing the signed average value of the two inputs.
[Link]
CUDA Math API vRelease Version | 291
Modules
Description
Compute average of signed input arguments x and y as ( x + y ) >> 1, avoiding overflow
in the intermediate sum.
Returns
Returns the least significant 32 bits of the product x * y.
Description
Calculate the least significant 32 bits of the product of the least significant 24 bits of x
and y. The high order 8 bits of x and y are ignored.
Returns
Returns the most significant 64 bits of the product x * y.
Description
Calculate the most significant 64 bits of the 128-bit product x * y, where x and y are 64-
bit integers.
Returns
Returns the most significant 32 bits of the product x * y.
Description
Calculate the most significant 32 bits of the 64-bit product x * y, where x and y are 32-bit
integers.
[Link]
CUDA Math API vRelease Version | 292
Modules
Returns
Returns a value between 0 and 32 inclusive representing the number of set bits.
Description
Count the number of bits that are set to 1 in x.
Returns
Returns a value between 0 and 64 inclusive representing the number of set bits.
Description
Count the number of bits that are set to 1 in x.
Returns
Returns a signed integer value representing the signed rounded average value of the two
inputs.
Description
Compute average of signed input arguments x and y as ( x + y + 1 ) >> 1, avoiding
overflow in the intermediate sum.
Returns
Returns .
[Link]
CUDA Math API vRelease Version | 293
Modules
Description
Calculate , the 32-bit sum of the third argument z plus and the absolute value
of the difference between the first argument, x, and second argument, y.
Inputs x and y are signed 32-bit integers, input z is a 32-bit unsigned integer.
Returns
Returns an unsigned integer value representing the unsigned average value of the two
inputs.
Description
Compute average of unsigned input arguments x and y as ( x + y ) >> 1, avoiding
overflow in the intermediate sum.
Returns
Returns the least significant 32 bits of the product x * y.
Description
Calculate the least significant 32 bits of the product of the least significant 24 bits of x
and y. The high order 8 bits of x and y are ignored.
Returns
Returns the most significant 64 bits of the product x * y.
[Link]
CUDA Math API vRelease Version | 294
Modules
Description
Calculate the most significant 64 bits of the 128-bit product x * y, where x and y are 64-
bit unsigned integers.
Returns
Returns the most significant 32 bits of the product x * y.
Description
Calculate the most significant 32 bits of the 64-bit product x * y, where x and y are 32-bit
unsigned integers.
Returns
Returns an unsigned integer value representing the unsigned rounded average value of
the two inputs.
Description
Compute average of unsigned input arguments x and y as ( x + y + 1 ) >> 1, avoiding
overflow in the intermediate sum.
Returns
Returns .
Description
Calculate , the 32-bit sum of the third argument z plus and the absolute value
of the difference between the first argument, x, and second argument, y.
[Link]
CUDA Math API vRelease Version | 295
Modules
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a single-precision floating point
value in round-down (to negative infinity) mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a single-precision floating point
value in round-to-nearest-even mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a single-precision floating point
value in round-up (to positive infinity) mode.
[Link]
CUDA Math API vRelease Version | 296
Modules
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a single-precision floating point
value in round-towards-zero mode.
Returns
Returns reinterpreted value.
Description
Reinterpret the high 32 bits in the double-precision floating point value x as a signed
integer.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a signed integer value in round-
down (to negative infinity) mode.
Returns
Returns converted value.
[Link]
CUDA Math API vRelease Version | 297
Modules
Description
Convert the double-precision floating point value x to a signed integer value in round-
to-nearest-even mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a signed integer value in round-
up (to positive infinity) mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a signed integer value in round-
towards-zero mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a signed 64-bit integer value in
round-down (to negative infinity) mode.
[Link]
CUDA Math API vRelease Version | 298
Modules
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a signed 64-bit integer value in
round-to-nearest-even mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a signed 64-bit integer value in
round-up (to positive infinity) mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to a signed 64-bit integer value in
round-towards-zero mode.
Returns
Returns reinterpreted value.
[Link]
CUDA Math API vRelease Version | 299
Modules
Description
Reinterpret the low 32 bits in the double-precision floating point value x as a signed
integer.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to an unsigned integer value in
round-down (to negative infinity) mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to an unsigned integer value in
round-to-nearest-even mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to an unsigned integer value in
round-up (to positive infinity) mode.
[Link]
CUDA Math API vRelease Version | 300
Modules
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to an unsigned integer value in
round-towards-zero mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to an unsigned 64-bit integer value
in round-down (to negative infinity) mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to an unsigned 64-bit integer value
in round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 301
Modules
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to an unsigned 64-bit integer value
in round-up (to positive infinity) mode.
Returns
Returns converted value.
Description
Convert the double-precision floating point value x to an unsigned 64-bit integer value
in round-towards-zero mode.
Returns
Returns reinterpreted value.
Description
Reinterpret the bits in the double-precision floating point value x as a signed 64-bit
integer.
[Link]
CUDA Math API vRelease Version | 302
Modules
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to a signed integer in round-down (to
negative infinity) mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to a signed integer in round-to-
nearest-even mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to a signed integer in round-up (to
positive infinity) mode.
Returns
Returns converted value.
[Link]
CUDA Math API vRelease Version | 303
Modules
Description
Convert the single-precision floating point value x to a signed integer in round-towards-
zero mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to a signed 64-bit integer in round-
down (to negative infinity) mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to a signed 64-bit integer in round-to-
nearest-even mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to a signed 64-bit integer in round-up
(to positive infinity) mode.
[Link]
CUDA Math API vRelease Version | 304
Modules
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to a signed 64-bit integer in round-
towards-zero mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to an unsigned integer in round-
down (to negative infinity) mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to an unsigned integer in round-to-
nearest-even mode.
Returns
Returns converted value.
[Link]
CUDA Math API vRelease Version | 305
Modules
Description
Convert the single-precision floating point value x to an unsigned integer in round-up
(to positive infinity) mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to an unsigned integer in round-
towards-zero mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to an unsigned 64-bit integer in
round-down (to negative infinity) mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to an unsigned 64-bit integer in
round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 306
Modules
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to an unsigned 64-bit integer in
round-up (to positive infinity) mode.
Returns
Returns converted value.
Description
Convert the single-precision floating point value x to an unsigned 64-bit integer in
round-towards_zero mode.
Returns
Returns reinterpreted value.
Description
Reinterpret the bits in the single-precision floating point value x as a signed integer.
Returns
Returns reinterpreted value.
[Link]
CUDA Math API vRelease Version | 307
Modules
Description
Reinterpret the bits in the single-precision floating point value x as a unsigned integer.
Returns
Returns reinterpreted value.
Description
Reinterpret the integer value of hi as the high 32 bits of a double-precision floating
point value and the integer value of lo as the low 32 bits of the same double-precision
floating point value.
Returns
Returns converted value.
Description
Convert the signed integer value x to a double-precision floating point value.
Returns
Returns converted value.
Description
Convert the signed integer value x to a single-precision floating point value in round-
down (to negative infinity) mode.
Returns
Returns converted value.
[Link]
CUDA Math API vRelease Version | 308
Modules
Description
Convert the signed integer value x to a single-precision floating point value in round-to-
nearest-even mode.
Returns
Returns converted value.
Description
Convert the signed integer value x to a single-precision floating point value in round-up
(to positive infinity) mode.
Returns
Returns converted value.
Description
Convert the signed integer value x to a single-precision floating point value in round-
towards-zero mode.
Returns
Returns reinterpreted value.
Description
Reinterpret the bits in the signed integer value x as a single-precision floating point
value.
[Link]
CUDA Math API vRelease Version | 309
Modules
Returns
Returns converted value.
Description
Convert the signed 64-bit integer value x to a double-precision floating point value in
round-down (to negative infinity) mode.
Returns
Returns converted value.
Description
Convert the signed 64-bit integer value x to a double-precision floating point value in
round-to-nearest-even mode.
Returns
Returns converted value.
Description
Convert the signed 64-bit integer value x to a double-precision floating point value in
round-up (to positive infinity) mode.
Returns
Returns converted value.
[Link]
CUDA Math API vRelease Version | 310
Modules
Description
Convert the signed 64-bit integer value x to a double-precision floating point value in
round-towards-zero mode.
Returns
Returns converted value.
Description
Convert the signed integer value x to a single-precision floating point value in round-
down (to negative infinity) mode.
Returns
Returns converted value.
Description
Convert the signed 64-bit integer value x to a single-precision floating point value in
round-to-nearest-even mode.
Returns
Returns converted value.
Description
Convert the signed integer value x to a single-precision floating point value in round-up
(to positive infinity) mode.
[Link]
CUDA Math API vRelease Version | 311
Modules
Returns
Returns converted value.
Description
Convert the signed integer value x to a single-precision floating point value in round-
towards-zero mode.
Returns
Returns reinterpreted value.
Description
Reinterpret the bits in the 64-bit signed integer value x as a double-precision floating
point value.
Returns
Returns converted value.
Description
Convert the unsigned integer value x to a double-precision floating point value.
Returns
Returns converted value.
[Link]
CUDA Math API vRelease Version | 312
Modules
Description
Convert the unsigned integer value x to a single-precision floating point value in round-
down (to negative infinity) mode.
Returns
Returns converted value.
Description
Convert the unsigned integer value x to a single-precision floating point value in round-
to-nearest-even mode.
Returns
Returns converted value.
Description
Convert the unsigned integer value x to a single-precision floating point value in round-
up (to positive infinity) mode.
Returns
Returns converted value.
Description
Convert the unsigned integer value x to a single-precision floating point value in round-
towards-zero mode.
[Link]
CUDA Math API vRelease Version | 313
Modules
Returns
Returns reinterpreted value.
Description
Reinterpret the bits in the unsigned integer value x as a single-precision floating point
value.
Returns
Returns converted value.
Description
Convert the unsigned 64-bit integer value x to a double-precision floating point value in
round-down (to negative infinity) mode.
Returns
Returns converted value.
Description
Convert the unsigned 64-bit integer value x to a double-precision floating point value in
round-to-nearest-even mode.
[Link]
CUDA Math API vRelease Version | 314
Modules
Returns
Returns converted value.
Description
Convert the unsigned 64-bit integer value x to a double-precision floating point value in
round-up (to positive infinity) mode.
Returns
Returns converted value.
Description
Convert the unsigned 64-bit integer value x to a double-precision floating point value in
round-towards-zero mode.
Returns
Returns converted value.
Description
Convert the unsigned integer value x to a single-precision floating point value in round-
down (to negative infinity) mode.
[Link]
CUDA Math API vRelease Version | 315
Modules
Returns
Returns converted value.
Description
Convert the unsigned integer value x to a single-precision floating point value in round-
to-nearest-even mode.
Returns
Returns converted value.
Description
Convert the unsigned integer value x to a single-precision floating point value in round-
up (to positive infinity) mode.
Returns
Returns converted value.
Description
Convert the unsigned integer value x to a single-precision floating point value in round-
towards-zero mode.
[Link]
CUDA Math API vRelease Version | 316
Modules
Returns
Returns computed value.
Description
Splits 4 bytes of argument into 2 parts, each consisting of 2 bytes, then computes
absolute value for each of parts. Result is stored as unsigned int and returned.
Returns
Returns computed value.
Description
Splits argument by bytes. Computes absolute value of each byte. Result is stored as
unsigned int.
Returns
Returns computed value.
Description
Splits 4 bytes of each into 2 parts, each consisting of 2 bytes. For corresponding parts
function computes absolute difference. Result is stored as unsigned int and returned.
[Link]
CUDA Math API vRelease Version | 317
Modules
Returns
Returns computed value.
Description
Splits 4 bytes of each into 4 parts, each consisting of 1 byte. For corresponding parts
function computes absolute difference. Result is stored as unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function computes absolute difference. Result is stored as unsigned
int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function computes absolute difference. Result is stored as unsigned int and
returned.
[Link]
CUDA Math API vRelease Version | 318
Modules
Returns
Returns computed value.
Description
Splits 4 bytes of argument into 2 parts, each consisting of 2 bytes, then computes
absolute value with signed saturation for each of parts. Result is stored as unsigned int
and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of argument into 4 parts, each consisting of 1 byte, then computes absolute
value with signed saturation for each of parts. Result is stored as unsigned int and
returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes, then performs
unsigned addition on corresponding parts. Result is stored as unsigned int and
returned.
[Link]
CUDA Math API vRelease Version | 319
Modules
Returns
Returns computed value.
Description
Splits 'a' into 4 bytes, then performs unsigned addition on each of these bytes with the
corresponding byte from 'b', ignoring overflow. Result is stored as unsigned int and
returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes, then performs
addition with signed saturation on corresponding parts. Result is stored as unsigned int
and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte, then performs
addition with signed saturation on corresponding parts. Result is stored as unsigned int
and returned.
[Link]
CUDA Math API vRelease Version | 320
Modules
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes, then performs
addition with unsigned saturation on corresponding parts.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte, then performs
addition with unsigned saturation on corresponding parts.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. then computes
signed rounded avarege of corresponding parts. Result is stored as unsigned int and
returned.
[Link]
CUDA Math API vRelease Version | 321
Modules
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. then computes
signed rounded avarege of corresponding parts. Result is stored as unsigned int and
returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. then computes
unsigned rounded avarege of corresponding parts. Result is stored as unsigned int and
returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. then computes
unsigned rounded avarege of corresponding parts. Result is stored as unsigned int and
returned.
[Link]
CUDA Math API vRelease Version | 322
Modules
Returns
Returns 0xffff computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts result is ffff if they are equal, and 0000 otherwise. For example
__vcmpeq2(0x1234aba5, 0x1234aba6) returns 0xffff0000.
Returns
Returns 0xff if a = b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts result is ff if they are equal, and 00 otherwise. For example __vcmpeq4(0x1234aba5,
0x1234aba6) returns 0xffffff00.
Returns
Returns 0xffff if a >= b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts result is ffff if 'a' part >= 'b' part, and 0000 otherwise. For example
__vcmpges2(0x1234aba5, 0x1234aba6) returns 0xffff0000.
[Link]
CUDA Math API vRelease Version | 323
Modules
Returns
Returns 0xff if a >= b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For
corresponding parts result is ff if 'a' part >= 'b' part, and 00 otherwise. For example
__vcmpges4(0x1234aba5, 0x1234aba6) returns 0xffffff00.
Returns
Returns 0xffff if a >= b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts result is ffff if 'a' part >= 'b' part, and 0000 otherwise. For example
__vcmpgeu2(0x1234aba5, 0x1234aba6) returns 0xffff0000.
Returns
Returns 0xff if a = b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For
corresponding parts result is ff if 'a' part >= 'b' part, and 00 otherwise. For example
__vcmpgeu4(0x1234aba5, 0x1234aba6) returns 0xffffff00.
[Link]
CUDA Math API vRelease Version | 324
Modules
Returns
Returns 0xffff if a > b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts result is ffff if 'a' part > 'b' part, and 0000 otherwise. For example
__vcmpgts2(0x1234aba5, 0x1234aba6) returns 0x00000000.
Returns
Returns 0xff if a > b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For
corresponding parts result is ff if 'a' part > 'b' part, and 00 otherwise. For example
__vcmpgts4(0x1234aba5, 0x1234aba6) returns 0x00000000.
Returns
Returns 0xffff if a > b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts result is ffff if 'a' part > 'b' part, and 0000 otherwise. For example
__vcmpgtu2(0x1234aba5, 0x1234aba6) returns 0x00000000.
[Link]
CUDA Math API vRelease Version | 325
Modules
Returns
Returns 0xff if a > b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For
corresponding parts result is ff if 'a' part > 'b' part, and 00 otherwise. For example
__vcmpgtu4(0x1234aba5, 0x1234aba6) returns 0x00000000.
Returns
Returns 0xffff if a <= b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts result is ffff if 'a' part <= 'b' part, and 0000 otherwise. For example
__vcmples2(0x1234aba5, 0x1234aba6) returns 0xffffffff.
Returns
Returns 0xff if a <= b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For
corresponding parts result is ff if 'a' part <= 'b' part, and 00 otherwise. For example
__vcmples4(0x1234aba5, 0x1234aba6) returns 0xffffffff.
[Link]
CUDA Math API vRelease Version | 326
Modules
Returns
Returns 0xffff if a <= b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts result is ffff if 'a' part <= 'b' part, and 0000 otherwise. For example
__vcmpleu2(0x1234aba5, 0x1234aba6) returns 0xffffffff.
Returns
Returns 0xff if a <= b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For
corresponding parts result is ff if 'a' part <= 'b' part, and 00 otherwise. For example
__vcmpleu4(0x1234aba5, 0x1234aba6) returns 0xffffffff.
Returns
Returns 0xffff if a < b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts result is ffff if 'a' part < 'b' part, and 0000 otherwise. For example
__vcmplts2(0x1234aba5, 0x1234aba6) returns 0x0000ffff.
[Link]
CUDA Math API vRelease Version | 327
Modules
Returns
Returns 0xff if a < b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For
corresponding parts result is ff if 'a' part < 'b' part, and 00 otherwise. For example
__vcmplts4(0x1234aba5, 0x1234aba6) returns 0x000000ff.
Returns
Returns 0xffff if a < b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts result is ffff if 'a' part < 'b' part, and 0000 otherwise. For example
__vcmpltu2(0x1234aba5, 0x1234aba6) returns 0x0000ffff.
Returns
Returns 0xff if a < b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For
corresponding parts result is ff if 'a' part < 'b' part, and 00 otherwise. For example
__vcmpltu4(0x1234aba5, 0x1234aba6) returns 0x000000ff.
[Link]
CUDA Math API vRelease Version | 328
Modules
Returns
Returns 0xffff if a != b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts result is ffff if 'a' part != 'b' part, and 0000 otherwise. For example
__vcmplts2(0x1234aba5, 0x1234aba6) returns 0x0000ffff.
Returns
Returns 0xff if a != b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For
corresponding parts result is ff if 'a' part != 'b' part, and 00 otherwise. For example
__vcmplts4(0x1234aba5, 0x1234aba6) returns 0x000000ff.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. then computes
unsigned avarege of corresponding parts. Result is stored as unsigned int and returned.
[Link]
CUDA Math API vRelease Version | 329
Modules
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. then computes
unsigned avarege of corresponding parts. Result is stored as unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function computes signed maximum. Result is stored as unsigned
int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function computes signed maximum. Result is stored as unsigned int and
returned.
[Link]
CUDA Math API vRelease Version | 330
Modules
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function computes unsigned maximum. Result is stored as
unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function computes unsigned maximum. Result is stored as unsigned int and
returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function computes signed minimum. Result is stored as unsigned
int and returned.
[Link]
CUDA Math API vRelease Version | 331
Modules
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function computes signed minimum. Result is stored as unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function computes unsigned minimum. Result is stored as
unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function computes unsigned minimum. Result is stored as unsigned int and
returned.
[Link]
CUDA Math API vRelease Version | 332
Modules
Returns
Returns computed value.
Description
Splits 4 bytes of argument into 2 parts, each consisting of 2 bytes. For each part function
computes negation. Result is stored as unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of argument into 4 parts, each consisting of 1 byte. For each part function
computes negation. Result is stored as unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of argument into 2 parts, each consisting of 2 bytes. For each part function
computes negation. Result is stored as unsigned int and returned.
Returns
Returns computed value.
[Link]
CUDA Math API vRelease Version | 333
Modules
Description
Splits 4 bytes of argument into 4 parts, each consisting of 1 byte. For each part function
computes negation. Result is stored as unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts functions computes absolute difference and sum it up. Result is
stored as unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts functions computes absolute difference and sum it up. Result is stored as unsigned
int and returned.
Returns
Returns computed value.
[Link]
CUDA Math API vRelease Version | 334
Modules
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function computes absolute differences, and returns sum of those
differences.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function computes absolute differences, and returns sum of those differences.
Returns
Returns 1 if a = b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function performs comparison 'a' part == 'b' part. If both equalities
are satisfiad, function returns 1.
Returns
Returns 1 if a = b, else returns 0.
[Link]
CUDA Math API vRelease Version | 335
Modules
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function performs comparison 'a' part == 'b' part. If both equalities are satisfiad,
function returns 1.
Returns
Returns 1 if a >= b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function performs comparison 'a' part >= 'b' part. If both inequalities
are satisfied, function returns 1.
Returns
Returns 1 if a >= b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function performs comparison 'a' part >= 'b' part. If both inequalities are satisfied,
function returns 1.
Returns
Returns 1 if a >= b, else returns 0.
[Link]
CUDA Math API vRelease Version | 336
Modules
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function performs comparison 'a' part >= 'b' part. If both inequalities
are satisfied, function returns 1.
Returns
Returns 1 if a >= b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function performs comparison 'a' part >= 'b' part. If both inequalities are satisfied,
function returns 1.
Returns
Returns 1 if a > b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function performs comparison 'a' part > 'b' part. If both inequalities
are satisfied, function returns 1.
Returns
Returns 1 if a > b, else returns 0.
[Link]
CUDA Math API vRelease Version | 337
Modules
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function performs comparison 'a' part > 'b' part. If both inequalities are satisfied,
function returns 1.
Returns
Returns 1 if a > b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function performs comparison 'a' part > 'b' part. If both inequalities
are satisfied, function returns 1.
Returns
Returns 1 if a > b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function performs comparison 'a' part > 'b' part. If both inequalities are satisfied,
function returns 1.
Returns
Returns 1 if a <= b, else returns 0.
[Link]
CUDA Math API vRelease Version | 338
Modules
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function performs comparison 'a' part <= 'b' part. If both inequalities
are satisfied, function returns 1.
Returns
Returns 1 if a <= b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function performs comparison 'a' part <= 'b' part. If both inequalities are satisfied,
function returns 1.
Returns
Returns 1 if a <= b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function performs comparison 'a' part <= 'b' part. If both inequalities
are satisfied, function returns 1.
Returns
Returns 1 if a <= b, else returns 0.
[Link]
CUDA Math API vRelease Version | 339
Modules
Description
Splits 4 bytes of each argument into 4 part, each consisting of 1 byte. For corresponding
parts function performs comparison 'a' part <= 'b' part. If both inequalities are satisfied,
function returns 1.
Returns
Returns 1 if a < b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function performs comparison 'a' part <= 'b' part. If both inequalities
are satisfied, function returns 1.
Returns
Returns 1 if a < b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function performs comparison 'a' part <= 'b' part. If both inequalities are satisfied,
function returns 1.
Returns
Returns 1 if a < b, else returns 0.
[Link]
CUDA Math API vRelease Version | 340
Modules
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function performs comparison 'a' part <= 'b' part. If both inequalities
are satisfied, function returns 1.
Returns
Returns 1 if a < b, else returns 0.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function performs comparison 'a' part <= 'b' part. If both inequalities are satisfied,
function returns 1.
Returns
Returns 1 if a != b, else returns 0.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts function performs comparison 'a' part != 'b' part. If both conditions
are satisfied, function returns 1.
Returns
Returns 1 if a != b, else returns 0.
[Link]
CUDA Math API vRelease Version | 341
Modules
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts function performs comparison 'a' part != 'b' part. If both conditions are satisfied,
function returns 1.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts functions performs substraction. Result is stored as unsigned int
and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts functions performs substraction. Result is stored as unsigned int and returned.
Returns
Returns computed value.
[Link]
CUDA Math API vRelease Version | 342
Modules
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts functions performs substraction with signed saturation. Result is
stored as unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts functions performs substraction with signed saturation. Result is stored as
unsigned int and returned.
Returns
Returns computed value.
Description
Splits 4 bytes of each argument into 2 parts, each consisting of 2 bytes. For
corresponding parts functions performs substraction with unsigned saturation. Result is
stored as unsigned int and returned.
Returns
Returns computed value.
[Link]
CUDA Math API vRelease Version | 343
Modules
Description
Splits 4 bytes of each argument into 4 parts, each consisting of 1 byte. For corresponding
parts functions performs substraction with unsigned saturation. Result is stored as
unsigned int and returned.
[Link]
CUDA Math API vRelease Version | 344
Notice
ALL NVIDIA DESIGN SPECIFICATIONS, REFERENCE BOARDS, FILES, DRAWINGS,
DIAGNOSTICS, LISTS, AND OTHER DOCUMENTS (TOGETHER AND SEPARATELY,
"MATERIALS") ARE BEING PROVIDED "AS IS." NVIDIA MAKES NO WARRANTIES,
EXPRESSED, IMPLIED, STATUTORY, OR OTHERWISE WITH RESPECT TO THE
MATERIALS, AND EXPRESSLY DISCLAIMS ALL IMPLIED WARRANTIES OF
NONINFRINGEMENT, MERCHANTABILITY, AND FITNESS FOR A PARTICULAR
PURPOSE.
Information furnished is believed to be accurate and reliable. However, NVIDIA
Corporation assumes no responsibility for the consequences of use of such
information or for any infringement of patents or other rights of third parties
that may result from its use. No license is granted by implication of otherwise
under any patent rights of NVIDIA Corporation. Specifications mentioned in this
publication are subject to change without notice. This publication supersedes and
replaces all other information previously supplied. NVIDIA Corporation products
are not authorized as critical components in life support devices or systems
without express written approval of NVIDIA Corporation.
Trademarks
NVIDIA and the NVIDIA logo are trademarks or registered trademarks of NVIDIA
Corporation in the U.S. and other countries. Other company and product names
may be trademarks of the respective companies with which they are associated.
Copyright
© 2007-2019 NVIDIA Corporation. All rights reserved.
[Link]