Index
The page numbers follow Appendix B of the PDF and link to the corresponding web location.
A¶
action density: 320
additively homogeneous map: 163
ADMM: 318
adversarial
affine
affine-difference cost: 59
affine-invariant metric: 39
Alexandrov solution: 24
algorithm
Birkhoff–von Neumann decomposition: 49
algorithmic convergence: 159
alternating
alternating minimization: 301
antisymmetric flux: 375
antisymmetric velocity: 327
Assouad lemma: 187
atom-preserving map: 78
atomic measure: 326
auction algorithm: 7, 92, 94, 167
coordinate ascent: 92
augmented Lagrangian: 335
B¶
Baker-Campbell-Hausdorff formula: 308
Bakry–Émery criterion: 361
balanced
transport between relaxed marginals: 205
balanced interpolation: 332
balanced OT: 143
Banach fixed-point theorem: 212
Banach space: 12
Banach-space embedding: 30
barrier
barycentric
Beckmann
belief propagation: 260
Benamou–Brenier: 139, 141, 277–280, 313, 315, 316, 319, 321, 324, 332
Bernstein–Schnabl operator: 76
Berry-Esseen
Bessel function: 113
bias: 195
biconvex
relaxation: 299
bilevel problem: 266
bilinear cost: 269
Birkhoff
block-coordinate
ascent: 249
block-coordinate descent: 262
block-coordinate method: 231
block-stationary point: 262
Bobkov-Ledoux formula: 30
Bochner spectral measure: 115
Bochner theorem: 198
Boltzmann entropy: 121
Boltzmann equation: 384
bounded confidence: 405
box constraint: 264
Bregman
Brenier
Brenier theorem: 299
Brenier–Strassen projection: 275
bridge
Bures
Burg entropy: 118
C¶
c-convexity: 27
c-cyclical monotonicity: 61
c-splitting set: 255
capacity-constrained OT feasibility: 264
capacity-constrained Sinkhorn: 264
capped transportation polytope: 264
Carathéodory theorem: 48
Catalan number: 7
causality: 404
CDF formula: 30
centered fluctuation: 195
centered measure: 66
central path: 53
centroidal Voronoi: 104
centroidal Voronoi tessellation: 102
chain graph: 259
characteristic kernel: 342
Choi matrix: 308
Cholesky factorization: 146
circle
classical convexity: 369
clustering: 405
co-motion function: 256
co-motion map: 256
column normalization: 125
commutator: 308
compact sublevel set: 207
compactness: 328
complete graph: 259
complete metric space: 12
completely positive cone: 200
completely positive kernel: 200
completely positive matrix: 200
complex
complexity
Birkhoff–von Neumann decomposition: 49
composition constraint: 347
computational complexity: 167
computer vision: 51
concave
minimization: 290
concave cost: 2
concave envelope: 153
concavity
conditional
conditionally positive definite kernel: 113
cone
congruence normalization: 308
conic lifting: 209
conjugate moment measure: 404
consensus: 405
Dobrushin contraction: 407
contact set: 27
continuity equation: 141, 277, 279, 280, 313, 315–318, 331, 336, 339, 373, 387–389, 391, 396, 399, 400, 410
continuous
continuous DTW: 310
continuous-depth neural network: 404
contraction ratio: 171
contractive Gaussian projection: 418
controlled transport: 410
convergence
convergence-determining class: 181
convex
convex envelope
Kantorovich cost: 57
convexity
convolution: 81
Wasserstein contraction: 81
convolutional Sinkhorn: 146
copositive certificate: 202
copositive matrix: 200
cost
cost-sectional curvature: 27
counterexample: 256
coupling measure: 64
covariance
Cramér condition: 183
Cramer-von Mises distance: 31
created mass: 206
cross-covariance: 414
cross-curvature: 90
cross-moment: 295
Csiszár divergence: 350
cumulant
fourth: 183
cycle graph: 259
cyclic
cylindrical Gaussian: 251
D¶
Dacorogna–Moser
damped iteration: 246
damped second-order equation: 382
Danskin theorem: 266
data-processing inequality: 119
DCA algorithm: 293
de la Vallée–Poussin criterion: 182
Delaunay graph: 110
demographic parity: 242
denoising score matching: 387
dense linear system: 123
density
depth evolution: 404
depth variable: 380
derivative
destroyed mass: 206
deterministic interpolant: 387
diagonal coupling: 274
diameter
contraction: 407
difference-of-convex programming: 293
diffusion model: 78, 81, 387, 391, 394
Gaussian noising: 132
diffusive map: 81
dimension dependence: 186
Dirac mass: 11, 14, 17, 29, 34, 41, 63, 68, 78, 243, 261, 336, 340, 347, 388
Dirac measure: 255
directional derivative: 266
Dirichlet energy: 341
discrepancy minimization: 395
discrete
disintegration: 12, 32, 65, 141, 210, 238, 255, 271–273, 389, 396
displacement convexity: 22, 26, 349, 359, 365, 368, 401, 402
distance profile: 289
distance residual: 288
distributional robustness: 73
divergence-free field: 411
Dobrushin coefficient: 406
domain adaptation: 51
Doob transform: 140
doubly nonnegative cone: 200
doubly nonnegative matrix: 200
doubly positive kernel: 200
doubly stochastic normalization: 410
Douglas-Rachford algorithm: 318
drifting
dual
duality
Dudley entropy: 192
dyadic cube: 186
dyadic partition: 185
dynamic
dynamic programming: 309
dynamic time warping: 309
dynamical plan: 239
E¶
earth mover’s distance: 51
edge flux: 330
edge imbalance: 31
Edgeworth expansion: 183
effective cost: 262
effective dimension: 200
effective rank: 199
eikonal equation: 419
electron density: 256
elliptical law: 36
elliptically contoured distribution: 36
embedding
alignment: 230
empirical
energy
energy-dissipation identity: 353
entropic
entropy: 402
$$-scaling: 95
equal marginal: 256
essential continuity: 401
Euclidean group: 230
Eulerian
extreme minimizer: 47
F¶
factored coupling: 261
fairness: 242
farthest-point sampling: 385
fast diffusion equation: 372
feature cost: 302
feature sketch: 197
feature term: 303
Fenchel inequality: 139
Fenchel-Young loss: 268
FFT: 148
fiber: 239
fiberwise optimal plan: 239
fiberwise transport: 238
fill-in edge: 259
filtered back-projection: 253
finite metric length: 353
finite space: 196
finite-sample bias: 396
finite-state transport: 329
first variation: 150, 266, 267, 307, 336, 338, 354, 367–369, 381, 399, 400, 410, 413, 415, 417
fixed-point
Gaussian barycenter: 246
flattening map: 250
flow
Fokker–Planck
Fokker–Planck particle closure: 399
Fortet iteration: 164
Fourier features: 197
Fourier multiplier: 113
Fourier random features: 198
Fourier transform: 252
Fourier-slice theorem: 253
fractional diffusion: 376
fractional heat equation: 376
fractional PDE: 375
Frechet inception distance: 115
Frobenius inner product: 37
frozen surrogate: 400
functional
G¶
Gamma convergence: 132
GAN: 119
gauge fixing: 178
Gaussian
blurring: 132
complex Sinkhorn: 158
convex order: 275
denoiser: 391
kernel: 198
kernel attention: 405
kernel rank: 200
mean: 412
mixture: 3, 97, 103, 125, 126, 129, 153, 206, 244, 263, 265, 282, 392, 394
one-dimensional: 415
optimal map: 29
preserving flow: 417
push-forward: 35
scale mixture: 202
Sinkhorn divergence: 177
sliced Wasserstein: 224
unbalanced: 213
Gaussian Hellinger-Kantorovich: 213
Gaussian measure
Gromov–Wasserstein: 295
Gelbrich inequality: 418
Gelbrich theorem: 418
generalized
generalized quantile: 32
generalized Wasserstein flow: 370
generator manifold: 396
geodesic: 327
Gibbs
Girsanov theorem: 140
global
global invariance: 230
gradient
graph
graphical model: 258
Grassmann manifold: 225
gravity interaction: 384
greedy sweep: 45
Gromov–Wasserstein
Gaussian measures: 295
ground cost: 63, 106, 250, 266, 268, 283, 289
adversarial: 293
ground metric learning: 66
group action: 229
growth term: 331
H¶
Hall theorem: 49
Hamilton-Jacobi-Bellman equation: 334
hard congestion: 334
Hausdorff measure: 20
heavy-tailed noise: 376
Hegselmann–Krause model: 405
Hellinger
Hellinger-Kantorovich: 332
Helmholtz decomposition: 411
hidden convexity: 402
Hilbert
Hilbert space: 352
Hilbertian
Hilbertian tangent norm: 320
Hilbertian Wasserstein space: 30
histogram: 4, 11, 30, 32, 63, 64, 122, 157, 205, 242, 277, 283
Hoeffding inequality: 199
Holder curve: 353
Hölder inequality: 236
holomorphic continuation: 156
holomorphic implicit function theorem: 157
homogeneous action: 322
homogeneous Sobolev norm: 156
homogenization: 209
Horn matrix: 202
hyperbolic geodesic: 26
hyperbolic metric: 34
hypersurface: 20
I¶
ice-cream cone: 37
imitation learning: 120
in-context map: 404
incompressible Euler: 260
increasing rearrangement: 60
induced width: 259
inertia: 384
inertial particle method: 383
infinite-depth limit: 404
infinite-dimensional linear programming: 56
inner product
Gromov–Wasserstein: 295
integer multiplicity: 51
interaction energy: 358
interaction kernel: 384
interaction polynomial: 76
intermediate measure: 261
internal energy: 349
interpolant: 394
intrinsic length metric: 223
inverse OT gap loss: 268
inverse square root: 308
irreducibility: 328
isometric embedding: 304
isotonic regression: 60
iterative closest point: 230
J¶
Jensen inequality: 69, 116, 128, 182, 190, 194, 273, 316, 325
JKO scheme: 143, 336, 339, 340, 346, 370, 373, 377, 378, 382, 416, 418
joint
Jordan decomposition: 13
jump kernel: 326
junction tree: 259
K¶
Kantorovich
Kantorovich–Rubinstein
kernel
kernelized self-interaction: 342
kernelized Stein discrepancy: 397
kinetic action: 141
kinetic equation: 384
kinetic mean-field limit: 383
KL
KL barycenter: 246
Kolmogorov forward equation: 377
Kolmogorov-Smirnov distance: 31
Kullback-Leibler divergence: 34, 39, 117, 127, 128, 130, 143, 356
L¶
L2 attention: 405
L2 self-attention: 405
Lagrange multiplier: 345
Lagrangian
Lagrangian velocity: 383
Laguerre cell: 18, 87, 96–98, 104
discrete: 92
Langevin
Laplace method: 179
Laplace-Beltrami operator: 145
large-temperature collapse: 154
large-temperature limit: 156
latent
latent measure: 261
lattice distribution: 183
Lavenant criterion: 418
law over measures: 250
layer-cake formula: 30
layerwise matching: 381
least squares
Radon reconstruction: 252
least-square
Legendre function: 159
length metric: 223
length space: 320
Lévy process: 376
linear
linear-time Sinkhorn: 197
linearized Wasserstein metric: 363
Liouville equation: 383
Lipschitz
Lloyd
local distance distribution: 283
local matching indicator: 2
local metric slope: 339
local profile: 283
local tangent action: 320
localization: 24
log-concave
log-concave generator: 404
log-concavity: 361
log-domain Sinkhorn: 122
logarithmic error: 199
logarithmic kernel approximation: 199
LogDet divergence: 213
look-ahead point: 382
Lorentz cone: 37
low-rank
lower semicontinuity: 56, 115, 116, 155, 207, 208, 250, 272, 277, 328, 332
lower semicontinuous relaxation: 322
lower-dimensional set: 20
Lyapunov equation: 411
M¶
Maas distance: 329
many-token regime: 79
map estimation: 194
map extrapolation: 192
marginal
market clearing: 164
Markov
averaging: 405
Markov kernel: 32, 81, 168, 283
Wasserstein stability: 81
masking: 404
mass
mass flux: 334
matching
matrix
matrix-vector iteration: 123
matrix-vector product: 199
max-flow min-cut theorem: 264
maximal density constraint: 344
maximum mean discrepancy: 112–115, 119, 120, 184, 190, 233, 342, 348, 349, 363, 384, 396, 412, 413
McCann
mean
Wasserstein decomposition: 66
mean field game: 333
mean shift: 405
mean-field momentum: 383
mean-preserving splitting: 273
mean-shift PDE: 405
measure
measure convolution: 71
measure-function duality: 83
measure-preserving
measure-to-measure map: 77
measure-to-vector map: 75
measure-valued Radon transform: 219
mechanical energy: 383
mesh size: 146
message passing: 259
metric
metric-measure space: 277, 283, 284, 286, 288, 289, 358, 359
min-plus algebra: 89
minibatch: 395
minibatch bias: 395
minibatch noise: 377
minimal velocity: 340
minimax lower bound: 187
minimax optimality: 187
minimizer
sparse: 44
minimizing movement scheme: 336
Minkowski problem: 401
mixed Hessian: 27
MMD-GAN: 395
modulo isometry: 286
modulus of continuity: 86
moment
convergence: 72
moment closure: 419
moment condition: 181
moment energy: 358
momentum variable: 320
Monge
Monge gap: 59
Monge map
Gromov–Wasserstein: 299
Monge–Ampère
monotone
Moore-Penrose pseudo-inverse: 252
MTW condition: 27
multi-marginal
multi-omics: 303
multi-species
multi-valued subdifferential: 20
multinomial distribution: 197
multiplicative scaling: 144
multiscale coupling: 185
Muon polar factor: 373
mutual information: 130
N¶
Nash equilibrium: 333
natural gradient: 396
natural language processing: 67
near-linear Sinkhorn: 199
nearest neighbor: 230
negative Sobolev norm: 364
Nesterov acceleration: 382
neural
neural optimal transport: 90
Newton dynamics: 384
Newtonian dynamics: 384
Newton–Schulz iteration: 374
node feature: 303
noising path: 394
noising schedule: 392
noisy gradient descent: 369
non-convex domain: 24
non-convexity: 352
non-variational scaling: 164
nonconvex
optimization: 242
nonlinear flattening: 250
nonlocal
nonlocal Wasserstein
nonnegative
orthant: 161
nonnegative factorization: 200
nonnegative rank: 261
nonnegative transport matrix: 49
norm
normal approximation: 182
normal cone: 344
normalized dynamics: 372
normalized SGD: 372
north-west corner
Nystrom approximation: 200
O¶
ODE limit: 382
one-dimensional
operator bound: 237
operator norm: 13
operator-valued coupling: 308
opinion dynamics: 405
optimal coupling: 28–30, 44, 48, 50, 55, 57, 58, 62, 65, 67, 82, 83, 88, 101, 243, 248, 257, 266, 268, 271, 286, 288, 316, 347
optimal plan: 20, 27, 42, 44, 45, 62, 84, 88, 125, 149, 271, 286, 347
optimal transport map
optimality
orbit space: 229
order-constrained coupling: 274
order-preserving map: 163
OT map estimation: 192
OT on trees: 31
OU bridge: 394
overshooting bridge: 394
P¶
pair space: 327
pair-space action: 326
pairwise distance: 284
parabolic optimal transport: 178
parameter space: 395
partial matching: 8
particle method: 78
particle polynomial: 337
particle splitting: 326
particle Wasserstein descent: 382
path action: 320
path entropy: 311
path metric: 223
path-space problem: 140
Pearson divergence: 118
penalized minimization oracle: 370
permutation
equivariance: 79
permutation matrix: 46
perspective recession: 277
perturbation response: 396
phase space: 383
phase-space density: 386
phase-space equation: 385
phase-space measure: 383
phi-divergence: 111, 115, 116, 118, 119, 127, 149–151, 351, 363
PL convergence: 354
plan
PMO: 370
Poincaré disk: 26
Poisson equation: 146
polar factor: 374
polar formula: 237
Polyak momentum: 382
Polyak-Lojasiewicz inequality: 354
population law: 333
Portmanteau theorem: 56
positional encoding: 79
positive
potential energy: 358
potential game: 333
power divergence: 350
power-law jump kernel: 376
Prékopa
primal-dual
method: 92
principal value: 376
probability integral transform: 28
probability measure: 11, 12, 14, 17, 19, 30, 33, 53, 56, 63, 70, 75, 84, 112–114, 118, 151, 219, 224, 232, 241, 254, 281, 285, 289, 315, 316, 342, 368
probability path: 388
Procrustes
product
product Wasserstein metric: 346
projected formulation: 234
projected measure: 219
projection direction: 219
projection-free attention: 405
projection-robust Wasserstein: 225
projective
propagation of chaos: 344
proximal
push-forward: 11, 14–18, 28, 29, 35, 36, 41, 314, 338, 387, 389
Q¶
quadratic
quantile: 28
quantization
quantum
quotient
R¶
Rademacher bound: 192
radial coupling: 21
radial generator: 36
radial measure: 20
Radon barycenter: 252
Radon pseudo-inverse: 252
Radon sinogram: 252
Radon structure: 221
Radon Wasserstein barycenter: 252
Radon Wasserstein distance: 220
Radon-Nikodym derivative: 389
ramp filter: 253
random measure: 250
random walk: 330
rank: 44
rank constraint: 234
rank-deficient covariance: 39
rate
raw second moment: 37
reaction term: 378
reaction–transport action: 331
reaction-transport geometry: 331
rectifiable set: 20
reference
reflow: 388
regression loss: 387
regular conditional distribution: 32
regular vector field: 326
regularization
Monge gap: 60
reinforcement learning: 120
relative
relative entropy
regularity: 76
rescaled convolution: 71
residual attention: 404
residual block: 379
residual network: 381
ResNet: 379
resolvent: 146
reverse formulation: 208
reverse KL divergence: 118
reversible jump kernel: 326
reversible Markov chain: 329
Ricci curvature: 359
Richardson–Romberg extrapolation: 337
Riemann sum: 70
Riemannian metric: 320
Riemannian tensor: 320
Riesz potential: 113
rigid motion: 230
rigid registration: 230
rigid update: 231
robust
row normalization: 125
running-intersection property: 259
S¶
Sakoe–Chiba band: 309
sampling: 78
scalar bridge: 394
scaling
Schatten norm: 236
Schoenberg theorem: 202
Schrödinger
score-based generative modeling: 387
score-SDE: 387
second-moment quotient: 37
second-order limit: 197
sectional curvature: 28
self Sinkhorn: 156
self-concordance: 123
self-corrected field: 400
semi-coupling: 209
semi-discrete
semi-relaxed problem: 86
semidefinite program: 305
separable space: 12
separator: 260
Shampoo optimizer: 372
Shannon entropy: 116, 117, 121, 306, 329, 341, 342, 347, 351
Shannon-Boltzmann entropy: 121
shape space: 229
signed
simplex: 330
projection: 162
simplex network: 52
Sinkhorn
acceleration: 171
closed-form coupling: 131
continuous dual iteration: 138
continuous iteration: 130
continuous limit: 179
coupling: 142
divergence: 121, 153–156, 177, 184, 187, 190, 197, 344, 396, 397, 414, 415
full-cycle map: 170
GAN: 395
Jacobian: 172
kernel: 199
linear-time: 197
local convergence: 171
over-relaxation: 173
potential: 178
scaling: 121, 123, 125, 143, 149, 152, 160, 170, 212, 308, 399
sketching: 197
statistical convergence: 181
update: 307
skip connection: 404
slack variable: 8
sliced barycenter: 252
sliced length metric: 223
sliced tangent norm: 222
sliced Wasserstein
slope
small-jump limit: 329
smooth density: 188
smooth OT estimation: 189
smoothed plug-in estimator: 188
smoothing estimator: 188
smoothness: 188
Sobolev norm: 156
Sobolev smoothness: 188
Sobolev space: 189
learned cost: 294
soft
soft-DTW: 310
expected alignment: 312
softmax attention: 204
source term: 331
sparse linear program: 109
sparsity: 44
spectral
spectral Wasserstein
sphere
c-concave potential: 90
spherical average: 219
spherical geodesic: 26
squared local action: 320
stable process: 376
stationary condition: 368
stationary density: 368
stationary point: 290
stationary potential: 179
statistical bias: 195
Stein force: 397
Stein geometry: 325
Stein method: 182
Stein operator: 397
step-size normalization: 374
Stiefel manifold: 225
stochastic
Strang splitting: 307
Strassen theorem: 274
strict
strictly correlated electrons: 256
strong law of large numbers: 181
strong MTW condition: 27
strong upper gradient: 340
structured data: 283
subgaussian measure: 190
subgradient: 267
subgradient inequality: 369
sublinear convergence: 357
subspace
subspace-sliced Wasserstein: 225
successive over-relaxation: 173
sum-product algorithm: 259
supergradient: 267
SVGD: 397
symmetric plan: 256
synthetic Ricci curvature: 359
T¶
tail index: 377
tangent coordinate: 251
tangent norm: 320
teacher distribution: 342
teacher kernel mean: 342
temporal noise: 387
tensor product coupling: 54
Thompson metric: 212
time change: 394
time reparametrization: 394
time series
alignment: 309
tomography: 253
topical map: 163
topology
total unimodularity: 50
total variation: 11, 13, 63, 70, 73, 111, 112, 116–118, 206, 207, 330
vector-valued measure: 107
trace
trace-class operator: 308
transfer learning: 51
transport
transport-entropy inequality: 365
transportation polytope: 41, 42, 45, 50, 52–54, 122, 128, 160, 161, 166, 264, 283
tree
tree-sliced Wasserstein: 32
tree-Wasserstein distance: 31
treewidth: 259
triangle inequality: 6, 19, 23, 39, 64, 65, 107, 167, 186, 190, 211, 220, 237, 284–286, 288, 328
triangular map: 32
trivial coupling: 388
tropical
barycenter: 89
truncated marginal: 218
twist condition: 25, 26, 254, 255
learned bilinear cost: 299
two-sample testing: 115
U¶
V¶
Varadhan formula: 145
variable elimination: 259
variable projection: 174
variance: 195
variational
mean field game: 333
variational dual formula: 119
vector quantile: 232
velocity covariance: 324
Vlasov equation: 384
von Neumann entropy: 306
W¶
Waddington landscape: 340
Waddington-OT: 51
warping path: 309
Wasserstein
barycenter: 143, 233, 241, 242, 244, 248, 250, 251, 257, 258
convergence: 181
coordinate: 232
distance: 30, 63–67, 69–73, 118, 119, 182, 205, 211, 219, 220, 234, 236, 237, 281–283, 285, 289, 315, 319, 323, 354, 402, 418
empirical rate: 185
flow: 354
flow of discrepancy: 396
GAN: 395
Gaussian geometry: 415
gradient: 315, 336, 338–340, 342, 344, 354, 367, 368, 382, 387, 397, 399, 400, 410, 411, 416–418
gradient flow: 22, 104, 150, 315, 336, 339–345, 354, 357, 368, 379, 381, 387, 396, 397, 410, 418
Hölder functional: 75
infinity distance: 74
infinity robustness: 74
Kurdyka–Łojasiewicz inequality: 356
Lipschitz functional: 338
minimax rate: 187
p-action: 323
p-distance: 25
polynomial: 76
space CLT: 251
squared distance: 352
Taylor regularity: 337
tower: 90
training: 367
weak
weight clipping: 120
weighted
wide stencil: 24
word mover’s distance: 67