Enquire Now
70+ Topics · Spectre · Spectre · cloud sim Sim · MATLAB · Webots · Hardware · Bangalore 2026

Max Stress Failure Composite Ansys

Simulation · Control · Perception · Hardware — 12 Lead ECG Acquisition — hardware, sensors, cloud dashboards and protocols (Spectre, REST, CoAP, WebSockets) for BE BTech MTech students. Final-year robotics support with Spectre stacks, simulation worlds, reports and viva from Bangalore.

70+
Related Topics
6+
Sim & HW Tools
4.9★
573 Ratings

Abstract—The Gumbel-max trick is a method to draw a sample from a categorical distribution, given by its unnormalized (log-)probabilities. Over the past years, the machine learning community has proposed several extensions of this trick to facilitate, e.g., drawing multiple samples, sampling from structured domains, or gradient estimation for error backpropagation in neural network optimization. The goal of this survey article is to present background about the Gumbel-max trick, and to provide a structured overview of its extensions to ease algorithm selection. Moreover, it presents a comprehensive outline of (machine learning) literature in which Gumbel-based algorithms have been leveraged, reviews commonly-made design choices, and sketches a future perspective.

max-stress-failure-composite-ansys Diagram
Figure: Model & System Architecture for Max Stress Failure Composite Ansys

Index Terms—Gumbel-max trick, Sampling, Gradient estimation, Gumbel-Softmax, Categorical distribution, Structured models

Ntroduction

The world around us is discrete in many aspects. Think about decision making, e.g. in traffic (Should I decelerate, accelerate or maintain a constant speed?), for product selection (Given that I liked the trousers from shop X last time, which new trousers shall I buy?), or in a clinical setting (Should I administer medication A or B to the patient?). Other discrete examples yield occurrence of events (At which day did I see you for the last time?), social networks (Do two persons know each other?) or data compression (How many bits do we need to store this information?). Modelling these concepts evidently pleads for discrete models, on which we focus in this work.

max-stress-failure-composite-ansys Diagram
Figure: Model & System Architecture for Max Stress Failure Composite Ansys

The immensely grown popularity of machine learning, and in particular deep learning, has given rise to high capacity models that have consistently been outperforming more classical models in various domains. Whereas the first neural networks typically comprised a stack of few fully-connected layers, activated with non-linear functions, current deep learning models often exhibit a (very) deep modular architecture. Specifically generative models (e.g.

the variational auto-encoder (VAE) and auto-regressive must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.

•

W. Kool is affilitated with the Amsterdam Machine Learning Lab (AM-

Pixelcnn ), And Bayesian Models Combine Such Deep

architectures with a form of stochasticity. Trainable param- eters of these models are generally learned in a data-driven fashion, i.e. they are optimized over a set of training exam- ples by back-propagating the error for a downstream task.

The probability distributions are then treated as stochastic nodes in a computation graph, from which a sample is drawn to compute the output, and subsequently the error.

Sampling from resulting discrete distributions may raise challenges when the probability mass function is unnor- malized, or in case of an (exponentially) large sampling domain. The latter is often encountered in sequence models, where one ‘sample’ entails a full sequence, and the sampling domain grows combinatorially for each additional sequence element. Moreover, incorporating (discrete) stochasticity in deep learning models poses a second challenge. Gradient computation, and therefore error backpropagation, through a stochastic node is hampered, which is required for updat- ing the distribution’s parameters and all model parameters that precede this node in the network. These two challenges, i.e. sampling and gradient estimation of discrete stochasticity in deep learning models, will be the focus of this review.

Several algorithms today exist to sample from structured models with exponentially large sampling domains (e.g. ancestral sampling), or to sample from a set of unnormal- ized probabilities (e.g. Markov Chain Monte Carlo (MCMC)

Methods , Or The Gumbel-Max Trick ). The Gumbel-

max trick recently found renewed attention for use in deep learning models, thanks to the proposed Gumbel-Softmax (GS) gradient estimator that is based on a relaxation of this trick , . The GS estimator (and variants thereof) have become popular (biased) alternatives for the high-variance REINFORCE estimator , thanks to their good empirical performance and straightforward implementation.

This article provides both an intuitive and mathematical understanding of the Gumbel-max trick, reviews extensions

Arxiv:2110.01515V2 [Cs.Lg] 8 Mar 2022

of this trick, provides handles for algorithm selection and corresponding design choices, and sketches a future per- spective. The content can be summarized as follows: • Background on categorical random variables, the Gum- bel distribution and inverse transform sampling is cov- ered in Section 2.

• Applications of Gumbel-based algorithms in machine learning are discussed in Section 3. • Section 4 and 5 present technical details about Gumbel- based sampling algorithms and gradient estimators, respectively.

• Practical considerations and commonly-made design choices are presented in Section 6. • Section 7 summarizes this review and sketches a future perspective.

Ategorical Distribution

A categorical distribution is a probability distribution that assigns a probability to N distinct classes. We consider three different parameterizations of the same distribution: normalized probabilities π, unnormalized probabilities θ, or un- normalized log-probabilities (or logits) log θ = a/T, where T ∈R>0 is a temperature parameter. We denote the ith unnormalized probability - which depends on T - with θi;T .

The probabilities πi;T for each class i ∈D = {1, . . , N} are

J∈D Exp(Aj/T) ∈R>0 Is

the partition function that normalizes the distribution. To prevent clutter, both the parameter and subscript of Z, and temperature subscript T are, in the rest of this paper, omitted when context allows it.

The temperature T controls the distribution’s entropy, such that the distribution can vary between a degener- ate/deterministic ‘one-hot’ distribution for T →0+ (i.e. all mass is centralized in one class), and a Uniform distribution (all probabilities are equal) for T →∞. In many cases, T simply equals one, making a the unnormalized log- probabilities. Equation (1) is known as a softmax function with temperature parameter T. Parameterized as Cat(a, T), the categorical distribution is often referred to as the Gibbs or Boltzmann distribution.

We will use the terms unstructured and structured models in this work. The former is used for what is normally re- ferred to as a categorical distribution: a distribution assign- ing a probability to a single event/class (e.g. the probability that horse X wins a horse race). The number of distribution parameters thus equals the size of the sampling domain. A structured model, on the other hand, places a probability on structures, i.e. different combinations of events/classes, rather than on single events (e.g. the probability that horses X, Y , and Z cover the top-3). A sequence model is another common example of such a structured model, in which the sampling distribution of each element may be conditional upon the previously sampled element.

A random variable following a categorical distribution is denoted with I in this work, and a realization (i.e. a sample from the categorical distribution) is denoted with I = ω, providing the index of the sampled class. In certain contexts, it is useful to define this sample as a one-hot embedding; a unit vector of length N, with a one at index ω and zeros oth- erwise, which we denote with 1ω. Several algorithms exist to sample from a categorical distribution. Inverse transform sampling is the most basic approach (see Section 2.3). The Gumbel-max trick (see Section 4.1.1), and variants thereof (see Section 4.3) are commonly-used alternatives in machine learning applications.

Gumbel Distribution

The Gumbel distribution is an instance (type I) of the generalized extreme value distribution1 , which models optima and rare events. A Gumbel random variable - which is often referred to in this work as ‘a Gumbel’ - is pa- rameterized by location and scale parameters µ ∈R and β ∈R≥0, respectively. The corresponding probability (PDF) and cumulative density (CDF) functions are respectively

(3)

We denote a Gumbel distribution (as defined in Eq. (2)) with Gumbel(µ, β) (or sometimes G(µ, β) in short), and a random variable following this distribution with Gµ,β.

To prevent clutter, the scale, or both parameters are fre- quently omitted when standard settings are assumed (i.e. Gµ := Gµ,1 and G := G0,1). We often consider a set of identically and independently distributed (i.i.d.) Gumbel variables, where we use G(i) to denote an ith standard Gumbel. Note that, µ and β are not the mean and variance

(5)

where γ ≈0.577 is the Euler constant and π ≈3.14 is the constant pi (not to be confused with the normalized probability of the categorical distribution), respectively.

The inverse cumulative density function (ICDF; also

Called Quantile Function) Is Given By

F −1(u) = −β log(−log u) + µ.

(6)

From Eq. (6) it can be seen that the Gumbel distribution is closed under scaling and addition, i.e. any Gumbel variable can be generated by scaling and shifting a standard Gumbel.

Equation (6) is used in inverse transform sampling (see Section 2.3) to transform a sample from the Uniform distri- bution U(0, 1) into a Gumbel sample via a double (negative) logarithmic relation. Taking only one negative logarithm of a Uniform random variable, i.e. X = −log U, defines a sample from the (standard) Exponential distribution. The Gumbel distribution is, therefore, sometimes also referred to as the ‘Double Exponential’ distribution. Thanks to the 1. Also the Fr´echet (type II) and Weibull (type III) distribution gener- alize to this generalized extreme value distribution.

𝐼= 2

Fig. 1: Illustration of inverse transform sampling for draw- ing a sample from Cat(π). A sample from the Uniform dis- tribution is converted to one of the categories/classes of the categorical distribution via its inverse cumulative density function. The chance that class i is being sampled equals πi; the width of the corresponding box in this illustration.

aforementioned relations, some properties of the Gumbel distribution are closely related to properties of the Exponen- tial, and Uniform/Beta distribution, on which we elaborate in Section 4.1.2.

Nverse Transform Sampling

Inverse transform sampling is the most common way to sample from a distribution. It transforms a standard Uni- form variable into another random variable, via a transfor- mation characterized by the random variable’s ICDF. Figure 1 illustrates this process for sampling from a categorical distribution: a sample form a Uniform random variable is inserted into the ICDF of the categorical distribution in order to find the corresponding bin/class. Inverse trans- form sampling is easy to implement and only requires generation of one random variate per categorical sample.

It does, however, require the categorical distribution to be normalized. The ICDF of the Gumbel distribution (given in Eq. (6)) can equivalently be used to draw a sample from this distribution.

Applications

In machine learning, there is significant interest in over- parameterized modular and structured models, which often involve one or more stochastic components. When con- cerned with discrete stochasticity, challenges arise regarding both sampling from discrete distributions, and gradient esti- mation thereof.

Sampling from discrete distributions can be achieved with (among others) inverse transform sampling, or the Gumbel-max trick (see Section 4.1.1) and extensions thereof (see Section 4.3). Gumbel-based sampling algorithms have, for example, been used for (discrete) action selection in a multi-armed bandit setting , for sampling data points in active learning , for text generation in dialog systems , or in translation tasks , .

When the stochastic components are parameterized, and the model is trained end-to-end with stochastic gradient- based methods, gradients need to be taken through the stochastic sampling process and gradient estimators are therefore called for. Finding such estimators is particularly challenging when the stochastic components are discrete.

Relaxations of the (non-differentiable) Gumbel-max trick, notably the Gumbel-Softmax (GS) estimator and variants therefore, have been found useful for this purpose, and will

Iscrete Stochasticity

Fig. 2: Categorization of the different applications in deep learning in which Gumbel-based gradient estimators have been applied in order to facilitate training via backpropa- gation through discrete stochastic nodes. Visualizations in the bottom show where different models exhibit discretized data, indicated by the dashed gray boxes.

be discussed in Section 5. In this section we review the different applications, and refer to any variant of the GS estimator as a Gumbel-based estimator, in order to provide a general overview without elaborating on technical details.

Applications are categorized into works that leverage dis- cretized data, or perform model selection from a discrete (model) space (see Fig. 2).

Iscretized Data

We interpret the term discretized data in a wide sense here. This category includes models that rely on discrete represen- tations (e.g. categorical probability distributions in discrete latent variable models), attention mechanisms (sub-selecting data is a discrete choice process), generative models for inherently discrete data (e.g. GANs for text generation), and models for data compression.

Iscrete Latent Variable Models

Latent variable models assume that ‘the world’ (represented as training data) has originated from a set of latent (i.e. hidden) variables. While such variables are often modelled continuously, many situations arise in which the choice for discrete latent variables seems more appropriate (see Fig. 2a). Such discrete latent variable models have mainly come in two flavours. On the one hand, one can have a discrete approximation/representation of a continuous (latent) variable (i.e. quantization, as in ), while on the other hand, one could model the data with a discrete (non- degenerate) distribution in the latent space. In the former case the straight-through estimator is typically used to facilitate gradient updates of the encoder. The latter case, on the other hand, requires sampling from the intro- duced discrete distribution to estimate the expectation in the forward pass and its gradient during backpropagation.

This approach is also possible in quantization settings, as will be discussed in Section 3.1.4. All following works in the current sub-section have adopted the second approach and leveraged Gumbel-based estimators for discrete latent variable models.

Aes With Discrete Priors Are Proposed By , ,

, and the authors of relax the discrete latents and durations of a recurrent hidden semi-Markov Model. In some works, continuous and discrete latent variables have

Been Combined , , , . The Authors Of

leverage - next to a continuous latent variable - (only) one categorical random variable, which is then interpreted as a clustering variable that assigns data points to discrete classes, resulting in a combined classification- and gener- ative model. Similarly, train a VAE with a Gaussian Mixture model as prior, where assignment to either of the Gaussians displays a discrete process as well, and in the same line, the authors of , relax backpropagation of based) discrete latent variable models as described above include (among others) planning , syntactic parsing ,

Generation , Recommender Systems , Drug-Drug In-

teraction modelling , and event modelling .

Attention

Attention mechanisms are concerned with filtering incom- ing information, analogously to the way we - human beings - selectively observe the world around us. Hard attention resorts to either accepting or declining information, while soft attention refers to placing more emphasis on certain parts than on others. When training neural networks, only soft attention allows gradients to flow, so hard (i.e. discrete) attention mechanisms have been relaxed using Gumbel- based gradient estimators to enable model optimization.

While the term attention is typically used to denote selectiv- ity in hidden representations of deep learning models, we slightly widen the term and present works that select part of the data (either hard or soft) anywhere in the model (see Fig. 2b), with the main motivation to improve the model’s performance or its interpretability. Attention has been used on latent features to enhance interpretability of ‘black box’ neural networks , , for hierarchical multi-scale re- current neural networks , and in graph neural networks

(Gnns) , . The Authors Of Apply Attention On

edges with the specific goal of graph clustering. Appli- cations range from (but are not limited to) recommender

Systems , Pose Estimation , Video Classification ,

event detection , and image synthesis . The authors of apply attention on features of discrete input samples and an attacking vocabulary, in order to acquire a scalable method for real-time generation of adversarial examples on discrete input data. Hard attention on (input) data points (i.e. learned subsampling) has, moreover, been leveraged to reduce computational overhead of neural networks, e.g. on point clouds and graphs .

Also Decision Making In An Environment Of One Or

multiple agents can be considered (hard) attention. The authors of leverage this ‘attention’ to enable end-to- end training of an agent that makes discrete decisions. In a multi-agent environment, such decisions have concerned communication symbol selection , in order to learn a language to be used among agents.

Generative Adversarial Networks (Gans) Have Been

a popular class of generative models. They contain a gen- erator and discriminator module (see Fig. 2c). The latter’s aim is to distinguish between real and fake (generated by the generator) samples. This setup requires differentiability of the discriminator loss with respect to the generator’s parameters. However, in the case of applying a GAN on discrete data, the discrete and stochastic fashion of the generator’s output hampers direct differentiability, and has therefore been a suitable candidate for Gumbel-based gra- dient estimators. Applications of discrete GANs that have leveraged these Gumbel-based gradient estimators include

(Among Others) Text Generation , , , , ,

, , fake user data creation for recommender systems

, And Action Prediction . Discrete Gans Have Also

been combined with knowledge distillation frameworks. The authors of , for example, propose to learn a pol- icy (i.e. a trajectory of subsequent states and actions in reinforcement learning setups) by learning to imitate a (known) expert/teacher policy. The student model learns this imitation by fooling a discriminator that should dis- tinguish between both policies. Such setups may, however, require a large number of training iterations before the student model converges. To alleviate this, the authors of have proposed the concept of adversarial distillation in the context of multi-label classification, in which both the teacher and student model are jointly battling against a discriminator that should distinguish their (discrete) label predictions from real annotations. Both models learn from each other via distillation losses, finally resulting in a low- resource student model that can be used during inference.

Both the teacher and student model’s output are relaxed using the GS estimator.

Ata Compression

Our growing demand for data has sky-rocketed data rates, hampering data transfer, processing and storage. To allevi- ate these rates, data compression comes into play. A shift is currently being made from non-data-driven compression techniques (e.g. JPEG) to data-driven methods, where pa- rameters of a neural compression model are learned from a training set of data. Such models typically encode the data in a quantized and lower-dimensional space, after which the data are decoded again (see Fig. 2d). Quantization is a discrete process, making it a suitable candidate to be learned using a Gumbel-based estimator. The authors of , propose to learn the quantization levels, which form a quantization codebook, whereas the authors of learn to select such a codebook as a whole. Moreover, data- adaptive binarization has been learned as well , , .

The amount of acquired data in the digital domain is governed by the analog-to-digital (ADC) converter in the sensing stage. Instead of compressing in the digital domain,

The Compressed Sensing (Cs) Framework Introduces

compression directly at this sensing stage (see Fig. 2d). Sim- ply reducing the sampling frequency of an ADC introduces aliasing artifacts in case the minimal Nyquist rate is not met. CS leverages incoherent sampling patterns (expressed in a sensing matrix) to beat this famous Nyquist rate un- der certain assumptions. The typically heuristic (and task- independent) design of this sensing matrix, has been out- performed by a task-conditional (discrete) sensing matrix, learned via a Gumbel-based estimator . Once trained, such learned sensing system can be implemented in hard- ware to reduce data rates at the sensing stage directly.

The proposed framework has been applied for medical

Ultrasound Imaging , Magnetic Resonance Imaging ,

and MIMO antenna arrays . Moreover, it was extended to a data-conditional setup, by making the sensing matrix conditional upon already acquired data . Note that data compression by subsampling can also be considered (hard) attention, as discussed in Section 3.1.2. Although the focus of this data compression section is on works that aim for data reduction rather than model enhancement.

Iscrete Model Space

The number of available neural models has given rise to algorithms that not only optimize a model on training data, but also search for a suitable model itself. Such data- driven model selection can be leveraged with the aim to just find better-fitting models, or to find compressed models that facilitate implementation in dedicated hardware. We distinguish methods that learn concepts related to model quantization, and methods that learn hyper-parameters (of- ten related to depth and width). The latter is often referred to as neural architecture search (NAS) or pruning.

Odel Quantization

Model quantization can take place on weights, activations and/or gradients. Joint optimization and quantization re- quires differentiability of the quantization operation, mak- ing it a suitable candidate for Gumbel-based gradient esti- mators. Indeed, it was shown that quantization of network weights and activations can effectively be learned, by either learning the quantization grid while using a pre-defined number of bits , or learning the appropriate number of bits per layer . The authors of also learn layer- dependent bits assignment by leveraging a bits budget to be distributed, for quantization of weights, activations and gradients.

Neural Architecture Search (Nas) & Pruning

Gumbel-based algorithms for discrete NAS of deep neural networks have been leveraged for distinct goals. Most gen- eral, hyper-parameters have been learned - rather than set heuristically - to alleviate the burden of tuning , , . In case such hyper-parameters relate to the size of the model (e.g. width or depth), NAS (often called pruning in this case) is generally used with the aim to find small models that facilitate implementation in (dedicated) hardware ,

, , , , , , . Moreover, Models Have

been proposed in which the functionality is conditioned upon incoming data. This conditioning can be applied to

Improve Performance , , , Achieve Run-Time Ac-

celeration (i.e. reduce computational complexity during in-

, , . Also Multi-Task Models Have Been Subject To

discrete NAS, where task-dependent gating/branching has

Been Learned , . Considering Graphs As Structured

data models, NAS has also been leveraged to find suitable graphical representations, e.g., by optimizing the adjacency matrix , . Moreover, graph structures are leveraged by Tree-LSTMs , which were later extended to Gumbel- Tree LSTMs, where graph edges are co-optimized instead of

Pre-Defined . Gumbel-Tree Lstms Have Notably Found

use in machine translation , .

Sampling Algorithms

In this section we introduce the Gumbel-max trick and its properties (Section 4.1.1), and link it to concepts known from different fields, i.e. Poisson processes (Section 4.1.2) and Boltzmann exploration (Section 4.1.3). We then introduce top-down sampling using the inverted Gumbel-max trick (Section 4.2), followed by sampling algorithms that are ex- tensions of the conventional Gumbel-max trick (Section 4.3).

Efinition And Properties

The Gumbel-max trick draws a sample from a categor- ical random variable I ∼Cat(π). It does so, by adding i.i.d. Gumbel (noise) samples to the unnormalized log- probabilities and selecting the index with the maximum value, which in turn follows a Gumbel distribution. More

(8)

The arguments Glog θi := log θi + G(i) (∀i ∈D) are shifted independent Gumbels, and often referred to as perturbed logits. Equation (8) is known as the max-stability property, which states that the maximum is independent of the argmax random variable: M ⊥⊥I. The interested reader is referred to Appendix A for the proofs of Eq. (7) and (8). Section 4.1.2 provides a more intuitive understanding of the Gumbel- max trick, by linking it to properties of well-known Poisson processes.

Scaling (or normalizing) unnormalized probabilities θ by any positive constant Z, results in a subtraction in logarithmic space: log(θ/Z) = log θ −log Z. Given the translation-invariance of the argmax function, we can see that the Gumbel-max trick as defined in Eq. (7), also applies

Log Θi−Log Z

+ G(i)}. This property is a direct result of the fact that Gumbel ran- dom variables adhere to Luce’s choice axiom , which states that the probability ratio of selecting two elements from a set is independent of the other elements present in that set. Thanks to this independence, a constant scaling of the probabilities does thus not influence the argmax output.

As a result, the Gumbel-max trick can also be used over sub-domain B ⊆D, with unnormalized probabilities θi, for



. In machine learning applications, the parameters of the categorical distribution are often learned or predicted. The Gumbel-max trick is therefore typically used with unnor- malized - rather than normalized - probabilities, as this allows for unconstrained optimization of log θ (whereas π is a probability vector).

The Gumbel-max trick leverages standard Gumbel noise samples. Nevertheless, when using i.i.d. samples from Gumbel(µ, β), with µ̸ = 0 and β̸ = 1 (≥0), the distributions of the sampled index and maximum can still be defined

With Location Μ, And Scale Β, And Z′ = P

i∈D exp(ai/(Tβ)). Appendix B derives the aforementioned relations. Inter- estingly, from Eq. (9) we conclude that I′ ∼Cat(a, Tβ).

Changing the scale of the independent Gumbel variates thus changes the temperature of the Boltzmann distribution from which exact samples are being drawn (see Section 4.1.3 for a discussion about this relation).

Ink To Reservoir Sampling And Poisson Processes

As mentioned in Section 2.2, the Uniform/Beta, Exponential and Gumbel distribution are closely related, and can be used to express different parameterizations of the same stochastic process. These distinct parameterizations are typically used in different research areas: the Uniform/Beta parameter- ization is notably used for weighted reservoir sampling , whereas the Exponential counterpart is widely used in queueing theory/Poisson processes . The ‘double Exponential’ parameterization is leveraged for the Gumbel- max trick , and its extensions. The aim of this section is to provide a summarizing overview of the existing relations between these distributions, and illustrate this through an intuitive example. Figure 3 relates samples and properties of the three parameterizations (one per column), and will further be explained in this section. Some of the connections that are presented have previously been made, e.g., both the

Author(S) Of And Showed The Equivalence Between

weighted reservoir sampling and the Gumbel-max trick, and the author of extensively discussed the relation between Poisson processes, especially Exponential races, and the Gumbel-max trick.

The ICDF of the Gumbel distribution - as provided in Eq. (6) - already showed the relation between a Gumbel and a Uniform random variable. The Gumbel distribution also relates to the Exponential(θ) distribution (θ ∈R≥0), which is characterized by a CDF being F(x) = 1 −e−θx for x ≥0 (and 0 for x < 0), and location E[X] = 1/θ.

Re-Introducing X ∼

Exponential(1) as a standard Exponential random variable

(11)

where both V and U are standard Uniform random vari-

Θ = −Log X + Log Θ = −Log(−Log U) + Log Θ, (12)

from which we can see that the negative logarithm trans- forms an Exponential random variable into a Gumbel ran- dom variable, located at log θ (right-hand side of Eq. (12)).

Note that for θ = 1, this Gumbel turns into a standard Gumbel variable with µ = 0. Rows I and II in Fig. 3 relate both the standard and parameterized Uniform/Beta, Exponential and Gumbel distributions. The green ellipses indicate the corresponding distributions.

Besides relating the distributions, we can also link some properties of the Gumbel-max trick to properties of a Pois- son process; a stochastic process widely used for modelling the occurrence of events over time. Thanks to the wide study of this process, many readers might have an intuitive understanding about its properties, enabling introduction of these relations via an intuitive example.

The Exponential distribution models the time between sequential events occurring according to a Poisson process. The distribution’s parameter θ represents the intensity (or rate), expressed in events per time unit. As an example, consider a scientific poster presentation where on average 8 students and 2 professors arrive per hour, both according to independent Poisson processes. The Exponential distri- bution models the time between subsequent arrivals of students and professors (with intensities 8 and 2, respec- tively). We can sample the arrival process by independently sampling these two Exponentials. From this example we in-

Θstudent+Θprof. , I.E. The

probability that the first person to arrive is a student2 equals 8+2 = 0.8. This property can be linked to the Gumbel-max trick. Formally, given X(i) ∼Exponential(1) i.i.d. ∀i ∈D, Eq. (12), and the monotonicity of the logarithmic function,

(13)

2. This can be shown by computing P(Xs < Xp) where Xs, Xp are the Exponential inter-arrival times for students/professors, respec- tively.

Exponential(1)

Uniform distribution Exponential distribution Double exponential distribution

−Log (⋅)

Fig. 3: The Gumbel (double Exponential) distribution directly relates to the Uniform/Beta and Exponential distribution via a single, respectively double, negative logarithmic relation. The green ellipses indicate the probability distribution of the corresponding random variable in the cell. To prevent clutter, we here define U := U (i), X := X(i), and G := G(i), being an ith independent standard Uniform, Exponential, and Gumbel random variable, respectively.

I∈D Θi, Which In-

deed coincides with our intuitively derived probability P[I = student]. Row III in Fig. 3 displays this relation, as well as the relation to the Beta/Uniform distribution. The latter is known from the field of weighted reservoir sam- pling , and can directly be deduced from the Exponen- tial reparameterization by taking the negative exponent of the Exponential sample (which, in turn, switches the argmin function to argmax thanks to 1/ exp(·) being monotonically decreasing).

Instead of independently sampling the two Exponentials (i.e. one for students and one for professors), we can develop an alternative but equivalent sampling process thanks to the memoryless property of the Exponential distribution3. More specifically, we can sample the arrivals of people (students and professors) at the poster from an Exponential with an intensity of 8 + 2 = 10 persons per hour, and then independently assign each arrival to either being a student or professor with probabilities 0.8 and 0.2, respectively. The fact that this ‘merged’ Poisson process can be sampled using

I∈D Θi, Relates To The Max-

stability property of independent Gumbel variables. Row IV of Fig. 3 shows the relation between the three different parameterizations. The actual value of the optimum in the Exponential parameterization equals the shortest arrival time, i.e. the time at which the first person arrives. Note that this time does not provide any information about the person being a student or a professor, i.e. the index and value of the optimum are independent, analogously to the independence between I and M in the Gumbel-max trick.

Exploration Vs Exploitation

Sampling from a categorical distribution may in some ap- plications adhere to an exploration-exploitation dilemma, where exploitation of already acquired information should be balanced with exploration of new information (e.g. when sampling an action in reinforcement learning). Pure ex- ploitation often results in greedy solutions that appear sub-

P(X > T1 + T2|X > T1) = P(X > T2)

optimal in the long run, while continuously exploring new directions can also cause sub-optimality and instable or diverging behavior.

A common approach to deal with this dilemma is Boltz- mann exploration , in which the temperature T of the Boltzmann distribution is gradually lowered during train- ing, therewith reducing the distribution’s entropy, and mov- ing from an exploration to an exploitation regime. Evidently, one can change the Boltzmann temperature and use inverse transform sampling to draw a sample from the categorical.

On the other hand, from Eq. (9) we see that samples from a tempered categorical distribution (i.e. T̸ = 1) may also be drawn using the Gumbel-max trick, either by changing the Boltzmann temperature explicitly, or by changing the scale β of the Gumbel distribution, from which independent samples are being drawn. Note from Eq. (6) that such noise samples are simply scaled (with β) samples from the standard Gumbel distribution. Intuitively, down-scaling (i.e. setting β < 1) the injected noise in the Gumbel-max trick thus perturbs the logits in a lower extent, therefore giving rise to a lower-entropy categorical distribution from which we are effectively sampling. Figure 4 visually relates sampling using an altered Boltzmann temperature (start in the bottom left and follow track A ) D ) F) and sampling using scaled Gumbel noise (start in the bottom left and follow track B ) C ) G). Figure 8 in Appendix C shows results of sampling experiments for various values of β.

This relation between the Boltzmann temperature and

Been Used By The Authors Of , Who Propose Boltz-

mann–Gumbel exploration (BGE), an algorithm - inspired by the Gumbel-max trick - in which a class-dependent Gum- bel noise scaling factor (i.e. β(i)) is used to guarantee sub- linear regret in a stochastic multi-armed bandit problem. In contrast to a class-independent scaling factor, class-dependent scaling factors do not admit an analytical expression for the resultant categorical distribution. BGE was later leveraged in a recommender system setting by the authors of , who more explicitly expressed the relation between the

Sample Of The Distribution Indicated In

Deterministic value that can act as a starting point to enter the graph

Ifferentiable Operation

Injection of noise from the distribution indicated in Using either of the inputs, results in the same output.

Performing The Operation On Each Of

the inputs, results in a different output.

∀𝑖∈𝐷

Fig. 4: Drawing a sample from a categorical distribution can be done via different paths. The most suitable path depends on the context in which one would like to draw a sample. One can start from either the unnormalized log-probabilities a (and a Boltzmann temperature T), or the normalized probabilities π. The green ellipses in the upper right corner of certain nodes indicate that the node represents a random variable following the distribution indicated in the ellipse. The different partition functions ZD all normalize their corresponding node, their input is omitted for readability reasons. The yellow square boxes are used to refer to certain paths in the text. A Jupyter Notebook that accompanies this figure is publicly

Available.4

Boltzmann temperature and Gumbel noise scaling. A similar exploration-exploitation trade-off exists in the field of natural text generation, where generated text should be of high quality (analogous to exploitation), but also diverse (analogous to exploration). Next to temperature scaling , , other algorithms have been proposed to balance the aforementioned trade-off in this field, e.g. re- distributing all mass of the Boltzmann distribution to either the k classes covering the top-k highest probabilities , or the p classes covering the top-p ratio of the total mass . Note that for β = 0 in Eq. (9) (i.e. no Gumbel noise is injected), the Gumbel-max trick resorts to setting k = 1 in the work of ; i.e. all mass is redistributed to the class with the largest probability (resulting in a zero-entropy one-hot distribution) and sampling becomes a deterministic (argmax) function.

Nverting The Gumbel-Max Trick: Top-Down Sampling

Both the argmax and max function in Eq. (7) and Eq. (8), respectively, are surjective/many-to-one mappings, i.e. nu- merous (multi)sets5 of perturbed logits exist that result in 4. https://github.com/iamhuijben/gumbel softmax sampling 5. A multiset or bag is a set in which an element can occur multiple times.

the same index I or maximum M. In certain scenarios one can be interested in reverting this sampling process, i.e. inferring the shifted Gumbels (i.e. perturbed logits) that may have generated a particular sample/index or maximum. This reversion is often referred to as top-down sampling , . To denote the standard procedure, i.e.

the ‘conventional’ Gumbel-max trick in which Gumbels are sampled unconditionally, the term bottom-up sampling was

Introduced . The Notion Of Top-Down Sampling Gave

rise to several applications, ranging from Gumbel-based

Estimators , , , Counterfactual Outcome Pre-

diction in structural causal models , and discrete flow models . We will elaborate on the principles of top- down sampling in this section.

We seek a set of N independently perturbed logits, with a maximum value M = m, located at index I = ω. Given either of the two (i.e. the maximum or index), we can - thanks to their independence - always sample the other from Cat(π) (sampling I) or Gumbel(log Z) (sampling M). Once both the maximum and corresponding index are known/sampled, the value of all (N −1) other perturbed logits should now be restricted to be ≤m. Acquiring Gum- bels, given this restriction, can be achieved by drawing them

~ Cat(𝝅)

Random variable following the distribution indicated in

−Log𝜃௜

Fig. 5: Sampling Gumbels and the corresponding argmax I and maximum M, can be done in a bottom-up or top- down approach. The former finds the (arg)max from the perturbed logits, starting from the Gumbels. Top-down sam- pling, on the other hand, starts from either M and/or I, and conditionally samples the corresponding perturbed logits.

The maximum and argmax are independent (⊥⊥), allowing for independent sampling of either of them in case only the other entity is known. The top-down procedure can be applied in parallel for the entire domain D (depicted) or sequentially (not visualized).

(15)

Analogous to the max-stability property given in Eq. (8),

The Maximum Of A Set Of Independent Samples From

TruncGumbel(log θi, 1, m), follows again a truncated Gum-

Max

i∈D { ˜Glog θi,1,m} ∼TruncGumbel(log Z, 1, m).

(16)

The top-down sampling procedure can be summarized as

Of

TruncGumbel(log θi, 1, m).

Reduces

to −log(U (ω))/Z. Equivalent expressions were given by the authors of [119, App. C] (assuming Z = 1), and [121, Eq. (9)].

The authors of observe that, when conditioning on M = m, we can even directly provide an expression for all perturbed logits without explicitly sampling I = ω first.

More specifically, a set of (unconditional) Gumbels Glog θi ∼

I∈D {Glog Θi},

can be transformed to conditional Gumbels ˜Glog θi,1,m, via the respective (inverse) CDFs. More specifically, ∀i ∈D:

M (·) Denote The Cdf And Icdf Of

truncated Gumbel distributions6, both located at log θi, and truncated at q and m, respectively. The set { ˜Glog θi,1,m}i∈D is now ensured to have a maximum m. Figure 5 visually summarizes both the bottom-up and the top-down sam- pling procedure.

Rather than simultaneously sampling all conditional Gumbels for i̸ = ω, they could also be acquired sequentially. This enables conditional Gumbel sampling for applications in which one is not interested in the perturbed logits over the entire domain D, because it is typically too large. The authors of note that hybrid top-down sampling for- mats (i.e. approaches that are in between simultaneous and fully sequential sampling) can also be used without loss of generality. They propose to partition the sampling space into two (random and mutually exclusive) sub-domains, and apply the sequential process within these domains separately. Their top-down construction algorithm general- izes the sequential conditional sampling procedure, and is provided in Appendix D.

Extended Sampling Algorithms

The Gumbel-max trick requires N Gumbel realizations to draw one sample from a categorical distribution of N classes. This can become cumbersome when drawing many samples, or even intractable when drawing a sample from an exponentially large domain (i.e. large N). Fortunately, its theoretical properties naturally extend to sampling al- gorithms with varying purposes, e.g. sampling with or without replacement, and sampling from structured mod- els. This section discusses such Gumbel-based sampling algorithms, which are summarized in Fig. 6, including conventional alternatives. The top and bottom row display the unstructured and structured setting, as discussed in Section 4.3.1 and 4.3.2, respectively. Algorithms that require a normalized distribution, i.e. a known partitioning function Z, are denoted with a superscript z. Green ellipses indicate the distribution (when known), from which algorithms re- turn samples. Note that soft samples (first column of Fig. 6) are discussed in Section 5.

6. The ICDF of TruncGumbel(µ, β, m) is provided in Eq. (15).

Conditional Samplingz

Gumbel-top-k /(reservoir) sampling [14, 110, 111]

Gumbel-Top-K Samplingz

Soft sample Single sample Multiple samples with replacement Multiple samples without replacement

Distribution

Fig. 6: Overview of sampling algorithms for different scenarios: single (discrete or soft) or multiple sample(s) from (un)structured distributions. For all cases, both default/conventional and Gumbel-based algorithms are given. The green ellipses indicate the distribution from which samples of the indicated algorithms originate. In case of empty ellipses, it is unknown which distribution the samples follow. A superscript z indicates that the algorithm requires normalized probabilities, i.e. partition function Z should be known.

Unstructured Distributions

Drawing a single sample from an (unstructured) categorical distribution can among others be done using inverse transform sampling (Section 2.3), and the Gumbel-max trick (Section 4.1.1). Figure 4 visually relates these sampling

Sampling Algorithm Can Be

interpreted as a continuous counterpart of the Gumbel-max trick, that samples from continuous domains (i.e. N →∞). A∗sampling does - analogously to the Gumbel-max trick - not require knowledge about the partition function Z. It leverages top-down sampling (as explained in Section 4.2) and bounds related to Gumbel processes , to diminish the number of i.i.d. Gumbels to be drawn, while still guaranteeing exact samples from unnormalized continuous distributions.

Sampling With Replacement

Repeating a discrete sampling process, we can draw mul- tiple (independent) samples with replacement. Depending on whether the order of samples is considered, this gives rise to an (ordered) sequence, or an (unordered) bag or multiset of samples. Such sequences and bags (or certain representations thereof) can in itself be considered samples from structured models (see Section 4.3.2). For example, when a bag of samples is represented as a vector of counts over the complete categorical sampling domain, this vector has a (structured) distribution that is known as the multino- mial distribution. Formally, the counts-vector resulting from drawing k samples (with replacement) from Cat(π) is a single sample from Multinomial(π, k). Therefore, the same sampling process may give rise to different distributions, depending on what is considered ‘the sample’.

When repeated samples from the categorical are drawn using the Gumbel-max trick, it may be possible to gain efficiency by reusing computations, as proposed by the authors of , .

Sampling Without Replacement

When rejecting (or removing) duplicates from the bag or sequence of samples drawn with replacement, we obtain a set, respectively sequence, of unique samples. This can be seen as sampling without replacement, implemented using rejection sampling, where proposals (i.e. individual samples from the bag/sequence) get rejected if they have already been sampled. Depending on the distribution, it may require many proposal samples to acquire a sequence or set of k unique elements (e.g. repeated sampling from a low-entropy distribution often returns the same class). Therefore, a more common alternative to rejection sampling is to sequentially sample without replacement by removing a sampled element

Figure

adapted from . Gumbel-top-k perturbs all logits with i.i.d. Gumbel samples, and returns the index of the k per- turbed logits with the highest values. These indices yield k independent samples without replacement from Cat(θ).

from the sampling domain, renormalize the distribution, and continue to draw the next (unique) sample from the updated domain. The subsequent (conditional) samples are typically sampled using standard inverse transform sam- pling.

If the order of the resulting unique samples is consid- ered, this process gives rise to a probability distribution over

(19)

When k = N, i.e. the full domain is sampled, this represents a distribution over permutations known as the Plackett-Luce

Model , . When We Do Not Care About The Order

of the unique samples, we can find the probability for a certain set of samples S, by summing the probabilities for

(20)

This probability distribution is a special case of the Wal- lenius’ noncentral hypergeometric distribution , , and can be computed exactly in O(N!), or approximated by an integral , .

Renormalizing the distribution for sequential sampling without replacement (as done in Eq. (19)) may be diffi- cult/intractable, especially in large domains. An appealing alternative is to use the Gumbel-max trick on the updated domain (being a subset of the original sampling domain), as it avoids the need for explicit renormalization. Even more interesting, when we apply the Gumbel-max trick repeat- edly for sampling without replacement, the N perturbed logits can be reused for all k samples, by simply selecting the top-k perturbed logits in a single step , , see Fig. 7.

This algorithm, which has become known as Gumbel-top-k sampling , is a strict generalization of the Gumbel-max trick (which is the special case for k = 1). This parameteri- zation of sampling without replacement allows for another derivation of the integral to approximate Eq. (20), or to compute it exactly in O(2N) .

Gumbel-top-k sampling is analogous to weighted reser- voir sampling, as shown by , . The latter uses the Beta-distribution parameterization (see Section 4.1.2), and selects the top-k perturbed keys using a priority queue, en- abling sampling from streaming applications . Gumbel- top-k sampling also has connections to ranking literature, as it was shown by that the Gumbel distribution is the only distribution of discriminal scores in a Thurstonian ranking model , inducing a ranking that satisfies Luce’s axiom of choice .

Samples without replacement can be used to form statis- tical estimators of functions of the underlying categorical

Distribution , , , , . Similar Ideas

have been used to construct gradient estimators for the process of sampling without replacement (see Section 5.2).

Structured Distributions

The Gumbel-max trick can be seen as an instance of perturb- and-MAP , which is a class of methods that transform sampling into an optimization problem, where the sample corresponds to the optimum of a perturbed energy func- tion. Specifically, in the context of the Gumbel-max trick, the energies correspond to (unnormalized) log-probabilities and the argmax operator represents the optimization prob- lem, i.e. ‘finding the maximum’ of the perturbed log- probabilities/energies. Whereas finding the maximum in an unstructured setting with a limited number of classes is hardly considered a ‘problem’, it is challenging in structured models, where there are multiple dependent variables with a sampling domain of exponential size. Various (approxi- mate) sampling algorithms have been proposed, which rely on the perturb-and-MAP principle by finding the (approx- imate) maximum of the perturbed energy over all possible

Configurations Of Such Structured Models , , ,

, . Note that there is a vast amount of literature on alterna- tives to sample from structured models. Standard ancestral sampling can, for example, be used when the conditional distributions for variables are normalized (i.e. the partition function is known), whereas Markov Chain Monte Carlo (MCMC) sampling and variants thereof can be leveraged when Z is unknown. In this section, we focus on methods related to Gumbel variables, such as the perturb-and-MAP paradigm.

The Gumbel-max trick can be seen as a reparameterization trick , , that allows to sample a categorical variable as a deterministic transformation of independent (Gumbel)

Is

exploited by the authors of , that aim to parallelize sampling by forecasting (reparameterized) samples from a sequential model.

Ultiple Samples

Whereas samples with or without replacement (from un- structured distributions) can in itself be seen as a single sample from a structured distribution (e.g. over sets or sequences), it is also possible to draw multiple samples (with or without replacement) from structured distribu- tions themselves (see bottom right corner of Fig. 6). For example, we can draw a set (or sequence) of (unique) sequences from a (neural) sequence model. The authors

Of , Show How To Efficiently Compute Multiple

(structured) samples with replacement from unnormalized distributions using the perturb-and-MAP paradigm. When

Authors:

Peder EZ Larson 1, 2,* , Jenna ML Bernard1, James A Bankson 3, Nikolaj Bøgh 4, Robert A Bok1, Albert P. Chen 5, Charles H Cunningham 6,7, Jeremy Gordon1, Jan-Bernd Hövener 8, Christoffer Laustsen 4, Dirk Mayer 9,10, Mary A McLean11 12, Franz Schilling13, James Slater1, Jean-Luc Vanderheyden5, 14, Cornelius von Morze 15, Daniel B Vigneron1, 2, Duan Xu1, 2, and the HP 13C

94143, Usa.

Denmark. 5 GE Healthcare, Menlo Park, California, USA. 6 Physical Sciences, Sunnybrook Research Institute, Toronto, Ontario, Canada.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

8 Section Biomedical Imaging, Molecular Imaging North Competence Center (MOIN CC), Medicine, Baltimore, MD, USA. Cambridge, United Kingdom.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

14Jlvmi Consulting Llc, Dousman, Wi, Usa

#See Acknowledgements for a list of all HP 13C MRI Consensus Group Members This work was supported by the ISMRM Hyperpolarized Media MR Study Group, the ISMRM Hyperpolarization Methods & Equipment Study Group, and the Hyperpolarized MRI Technology Resource Center (NIH/NIBIB grant P41EB013598).

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Abstract

MRI with hyperpolarized (HP) 13C agents, also known as HP 13C MRI, can measure processes such as localized metabolism that is altered in numerous cancers, liver, heart, kidney diseases, and more. It has been translated into human studies during the past 10 years, with recent rapid growth in studies largely based on increasing availability of hyperpolarized agent preparation methods suitable for use in humans. This paper aims to capture the current successful practices for HP MRI human studies with [1-13C]pyruvate - by far the most commonly used agent, which sits at a key metabolic junction in glycolysis. The paper is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification. In each area, we identified the key components for a successful study, summarized both published studies and current practices, and discuss evidence gaps, strengths, and limitations. This paper is the output of the “HP 13C MRI Consensus Group” as well as the ISMRM Hyperpolarized Media MR and Hyperpolarized Methods & Equipment study groups. It further aims to provide a comprehensive reference for future consensus building as the field continues to advance human studies with this metabolic imaging modality.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Keywords: Hyperpolarized MRI, metabolic imaging, carbon-13, pyruvate, dissolution dynamic

Introduction

MRI with hyperpolarized 13C agents, also known as hyperpolarized (HP) 13C MRI, has shown great potential as a novel imaging modality, particularly for its ability to probe metabolic processes in real time. The first human studies with HP [1-13C]pyruvate were performed in 2011 in prostate cancer patients (1).

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Since then, there have been over 60 papers published with imaging results of human subjects from 13 different sites, with applications including prostate cancer, brain tumors, breast cancer, kidney cancer, pancreatic cancer, metastatic disease, liver disease, ischemic heart disease, diabetes and cardiomyopathies. The vast majority of these studies used [1-13C]pyruvate (1–63), where [2-13C]pyruvate (64) and 13C-urea (56) have been demonstrated too.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

As clinical HP 13C MRI advances, there is a growing need to build consensus for best practices, which are critical for comparing data across sites, performing multi-site trials,deploying methods to new sites, partnering with vendors, and potentially for obtaining broader regulatory approvals.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

In March 2022, we initiated an effort to build consensus within the HP 13C MRI community with this opportunity in mind, and it was greeted with strong enthusiasm. The “HP 13C MRI Consensus Group”, containing over 55 members from 27 sites, identified the area of greatest need and opportunity for consensus building to be HP [1-13C]pyruvate human

●

Pyruvate is the most mature and widely used HP agent and has the most significant translational evidence emphasizing the potential clinical impact.

●

Clinical trials, particularly multi-site trials, have the strongest need for consensus methods to ensure that data can be combined across sites. This work is a Position Paper for which the goal is to describe current successful practices and study methods for HP [1-13C]pyruvate human studies along with justification to support those practices. This is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification (Fig. 1). The current successful practices and study methods include a literature review of published peer-reviewed journal papers showing human HP [1-13C]pyruvate study data, up to September 2022 (1–63), as well as new unpublished information from surveys of HP 13C study sites. Based on this information, we also highlight the evidence gaps, strengths, and limitations of current practices which are summarized at the end of each section.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Figure 1: Illustration of the HP 13C MRI human study process, including the 4 major areas covered in this paper: Hyperpolarized 13C-pyruvate preparation, MRI system setup and calibration, Acquisition and Reconstruction, and Data Analysis and Quantification.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Figure 2: Anatomical targets of HP [1-13C]pyruvate MRI human studies published up to September 2022.

Hyperpolarized 13C-Pyruvate Preparation

This section covers the processes for creating the HP agent, 13C pyruvate, and will include many aspects and considerations that are needed to safely and effectively prepare doses for metabolic imaging studies in human subjects. These include material, personnel, equipment and facility, fluid path preparation, quality control, and release.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

It is helpful to understand that the specifications of a dose of 13C pyruvate suitable for in vivo MR HP metabolic imaging were shaped in part by early preclinical studies performed by GE HealthCare summarized in Ref. (65). In short, the safety of the two novel drug components, 13C pyruvate and the electron paramagnetic agent (EPA) AH111501, were demonstrated in those studies. The more precise formulation of the dose suitable for human use was then determined from clinical studies (66) that included two Phase 1 clinical trials in young and elderly healthy volunteers without hyperpolarization of the 13C nuclei and another Phase 1/2a dose escalation and imaging feasibility study with HP 13C pyruvate in 31 prostate cancer patients at the With the exception of the first HP 13C imaging clinical trial, which utilized a prototype device in a cleanroom (1), all HP 13C studies performed in humans to date have utilized the SPINlab polarizer (manufactured by GE HealthCare). Consequently all doses of the HP 13C pyruvate delivered by SPINlab have been produced using the “SPINlab Pharmacy Kit” that serves as the container-closure system for the various drug components (13C pyruvic acid and EPA mixture, dissolution medium, and neutralization and dilution medium) during sample polarization, dissolution and quality control (QC) processes. Thus many aspects of the HP sample preparation considerations discussed below are related to the SPINlab instrument and the consumables designed to be used with it (67).

General Considerations

While more than 860 patients or healthy subjects having been injected with HP 13C pyruvate as of January 2022 without reports of any serious adverse events (68), HP 13C pyruvate injection remains an investigational MR contrast agent and can only be administered by those with Investigational New Drug (IND) exemption from the Food and Drug Administration (FDA) in the USA, a Clinical Trial Application (CTA) in Canada, approval from National Research Ethics Committee Services in the UK, or approval from the relevant local regulatory body. Thus, methods and processes involved to produce a dose should have patient safety as the first priority. Since utilizing dissolution dynamic nuclear polarization (dissolution-DNP) for human use is still a relatively new development, there are no existing published regulatory guidelines specifically for this method.

There are two major production styles that determine how various sites approach the agent preparation. In the US, the most common approach is to rely on a sterilizing filter (“Terminal Sterilization”) to ensure sterility of the final product, akin to PET tracer production, where a starting molecule with a radioisotope is processed using various other ingredients to make the final, desired and injectable contrast agent within a necessarily short amount of time (69). For these sites, sterilization of the components and accessories upstream of this filter are not required, although many of them were manufactured and tested following Good Manufacturing Practice (GMP) or Good Laboratory Practice (GLP) requirements. The filling process is usually performed under an ISO 5 laminar flow hood, but a clean room or an isolator is not required.

This approach is typically accompanied by testing the integrity of the sterilizing filter prior to release of the dose for injection. Typically, post release endotoxin and sterility tests are performed using an aliquot reserved from each released dose.

In the UK and EU, the most common approach is to more-closely follow sterile pharmaceutical compounding guidelines (70), where all components and ingredients are required to be sterile or manufactured under GMP guidelines and are assembled and filled within a clean room environment or an isolator system (“Sterile Preparation”). Typically a batch of Pharmacy Kits for HP 13C pyruvate injection are prepared together. The sterility of the final dose is also ensured by batch validation testing, in addition to the sterility of the ingredients and the sterile compounding process. The endotoxin and sterility testing are performed for the process validation but are not performed for each injected dose.

Some institutions fill and assemble the Pharmacy Kit required for a specific study on the same day or the day prior to polarization, dissolution, and patient administration, but others have also demonstrated the feasibility of preparing a batch of kits, keeping them in a -20ºC freezer and using them over a period of a few months.

Beyond the obvious requirements that the process and the facility has to ultimately produce a dose that is safe to inject into a human, regulatory authorities will also focus on the question “Are you in control of your processes?”. To be in control of your process requires an in-depth and broad understanding of all processes involved in pre, post, and during the production process.

Personnel

It is typical and may be required to have licensed personnel involved in the production process depending on local regulations.Typically a pharmacist, radiopharmacist or other similarly qualified person (QP), in charge of the facility where the Pharmacy Kit filling and preparation is taking place, is responsible for the overall process and the release of the injectable dose.

Qualified cleanroom technicians are often involved in the Pharmacy Kit filling under the supervision of the pharmacist or QP. As is required for pharmaceutical compounding or PET tracer production, training requirements and training records for all personnel need to be maintained and available for audit by the FDA or equivalent.

Equipment And Facility

The facility and all equipment need to have standard operating procedures (SOPs) that describe how equipment is used, maintained, and calibrated to comply with relevant legislation. Currently, almost all the filling of the Pharmacy Kit takes place within a compounding laminar flow hood or isolator (typically ISO 5). At some sites, the filling is conducted within a cleanroom, while at others, it is conducted in a dedicated non-cleanroom space, reflecting differences in cleanroom approach and specifications between regulators worldwide (71). Some equipment or facilities, such as the compounding hood or cleanroom, may require external certified laboratories for testing.

Material Handling

Material handling guidelines (69,70) require SOPs detailing a system to track all of the materials involved in the HP production process for a particular patient dose, similar to current good manufacturing practice (cGMP) requirements for material handling for drug compounding. This includes acceptance standards, storage conditions, amount used in the patient dose for each ingredient and materials used in the assembly of the fluid path and Pharmacy Kit. Currently some users choose to open and inspect and sometimes modify the Pharmacy Kits upon arrival, but some users keep them in the sealed packaging until they are required for dose preparation.

Pharmacy Kit Filling And Assembling

As required by an IND or its equivalent, the preparation of the doses of HP 13C agent are detailed in the Chemistry, Manufacturing, and Control (CMC) section of an applicable regulatory submission; an example of this has been made available (72). It describes the processes of filling the Pharmacy Kit with the different components that make up the final drug product, and of assembling the final kit for either storage or immediate use in the polarizer. Special attention should be given to the laser welding process in order to satisfy installation qualification (IQ) and operational qualification (OQ). Typically, the final developed process is validated by process qualification (PQ) runs, during which 3 or more Pharmacy Kits are filled and used and the final HP 13C products are tested for endotoxin and sterility and to confirm that they meet the dose specifications for injections (usually including pyruvate concentration, residual EPA concentration, pH, liquid state polarization level and dose temperature). The data from 3 consecutive PQ runs are submitted as part of the IND submission (or its equivalent), and are often also reviewed by the Institutional Review Board (IRB) where the studies are conducted.

Quality Control And Dose Release

The quality control (QC) and dose release can be separated into two aspects: one is the QC and release of the filled Pharmacy Kit, and second is the QC and release of the HP 13C agent for injection, after polarization and dissolution. For institutions filling a batch of kits and storing them to use over a period of time, typically the batch can be released based on initial validation, environmental monitoring data from the day of kit production, and if filters are used during preparation of any of the components, filter integrity testing. But in some cases one or more kits are used for validation before the batch of kits are released for future use. For institutions that fill only the kits required for specific studies shortly before the experiment, the filled kits often do not go through separate release tests before they are used.

The quality control of the HP 13C pyruvate solution post dissolution is primarily performed to ensure that the agent meets the dose specifications (Table 1) before it is administered to the subject. These specifications target both safety (pH, residual EPA, temperature) and efficacy (pyruvate concentration, polarization, volume). Typically, the pyruvate concentration, residual EPA concentration, pH, dose temperature, dose volume, and liquid state polarization are measured by the QC accessory associated with the SPINlab polarizer. Some users perform a secondary measurement for one of the parameters, such as pH, using a different instrument or pH paper. For sites that do not go through a separate release testing process for batch filled kits, the integrity of the sterilization assurance filter, a part of the Pharmacy Kit, is typically tested as a part of the dose release. It is also common for these users to preserve an aliquot of the final HP 13C pyruvate solution for post-release endotoxin and sterility testing. This testing cannot be completed fast enough to test an individual dose prior to injection, but this is why other processes such as PQ runs and validation testing are done to minimize the chance a subject could be injected with a contaminated dose.

The Final Dose Release And Injection

should be done under the supervision of a licensed professional, based on local regulations.

Some Key Challenges

Many of the challenges associated with HP 13C pyruvate preparation can be attributed to the conditions required for the dissolution-DNP method of high magnetic field (~3-7 T) and very low temperature (~1 K) during polarization, with pressurized and superheated water necessary for the rapid dissolution event. These extreme conditions are quite challenging for the design of the container-closure and fluid path system. In particular, the cryogenic temperature in the polarizer requires special attention to any moisture or ambient (moist) air introduced into that portion of the fluid path, which can form an ice block at ~1 K. This ice can lead to flow restriction during the dissolution event and reduce the strength of the laser welded bond between the cryovial and its cap. This can ultimately produce failures in the dissolution step, including variations in final pyruvate concentration and pH that may fail to meet QC release criteria as well as fluid path ruptures that provide no available dose and result in polarizer down-time.

The polarization of the HP 13C pyruvate sample decays quickly over the span of a few minutes after dissolution, and thus the process of dissolution, QC for release, and injection should be completed as fast as possible to preserve the high polarization level achieved. Any delays in the preparation process, such as transportation time or equipment malfunction, can significantly reduce the final polarization and result in lower quality imaging data.

Current Practices

A summary of data collected from all sites performing clinical trials with HP 13C-pyruvate is shown in Fig. 3 and Table 1, including the specification of the final dose and how the quality control and release of the final dose are performed. There is a split in the Production Style, described in the General Considerations section above, with 8/13 sites using Sterile Preparation versus 5/13 using Terminal Sterilization. While many of the dose specifications show notable differences in acceptable ranges, all of these variations listed in tables have been successfully and safely been used to perform HP 13C pyruvate studies in humans. Their differences depend on the institutions’ preferences, resources and their particular regulatory situation. There is high similarity in pyruvate ranges, temperature ranges, EPA limits, and volume limits. There is modest variability in pH ranges and large variability in the endotoxin test limit. There is a 3-fold difference in acceptable polarization levels, which are measured to ensure a futile dose is not injected since the polarization is directly proportional to SNR. This reflects the decision by several sites to believe that useful data can be still be obtained with suboptimal polarizations.

Figure 3: Hyperpolarized agent preparation methods reported by sites currently performing HP

In House

Table 1: HP 13C-pyruvate preparation parameters, methods, and dose specifications used for quality control testing and release as well as validation. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. The parameters used for product release are noted in bold text, otherwise these parameters are measured for batch validation or other QC measurements. The endotoxin and sterility testing are performed during process validation of the batch and/or post-injection, and largely depends on the agent production approach.

Summary

The overall safety record of HP 13C-pyruvate has been very strong, and the SPINlab hyperpolarizer has proven to provide high polarizations at human sized doses while meeting numerous QC and release criteria. A weakness remains the failure modes of the SPINlab Phamacy Kits (e.g. ice blocks, path ruptures), which are placed under extreme requirements particularly during dissolution. The preparation process still requires a high degree of expertise.

Therefore, there is a significant need to improve the reliability, robustness, and ease of operation for generating HP 13C-pyruvate doses for human studies. Furthermore, there is a divide between manufacturing and sterile compounding style preparation as well as other site-specific practices, resulting in variations in SOPs and justification required to relevant regulatory bodies. There have also been no comparisons between these approaches. It is also unclear what release criteria and QC parameters are truly required to ensure patient safety.

However, all of the reported methods are acceptable and approved by the appropriate regulatory authorities, and have led to the rapid expansion of successful human studies in recent years.

Mri System Setup And Calibrations

This section covers the MRI system setup, including the imaging system, RF coils, phantoms, and prescan calibration methods.

Imaging System

The main prerequisite for a given MRI scanner to be capable of supporting studies with HP 13C is its “broadband” capability to transmit and receive radiofrequency (RF) signal at the frequency of 13C, which is around 4 times lower than 1H. This does not come as a default on clinical MR devices. The transmit power of the broadband amplifier should also be sufficient to support the intended flip angle and RF pulse shape with the employed transmission RF coil(s) for 13C. Most studies to date use relatively low flip angles (< 90 degrees) for HP 13C in order to preserve polarization for time-resolved imaging. The capability to receive 13C signal on multiple channels is also desirable to increase SNR, as discussed further in the “RF coils” section.

The choice of magnetic field strength is primarily dependent on the metabolites’ frequency separation due to chemical shift dispersion and 1H imaging. High field strengths do not enhance hyperpolarized 13C signal as they do for 1H because the signal strength in a HP experiment relies on manipulating the population of quantum energy states outside of the MRI scanner.

However, the injected HP 13C-pyruvate and its metabolic products have greater frequency separation at higher fields, and it may thus be easier to separate and quantify these resonances at higher fields. This comes at the cost of a reduction in the achievable T2* and often reduced T1. As the initial polarization is independent of the imaging field strength it has been proposed that the increased T2* at 1.5T can potentially be exploited to increase SNR by adapting the acquisition bandwidth or reduce off-resonance imaging effects in cases when the decay of the transverse magnetization is dominated by T2* (73). In practice, 3T has been used in all published human 13C-pyruvate studies surveyed (Supporting Table S1), and comprises the majority of scanners currently in use for human studies (Table 3). A field strength of 3T is well-suited for 1H MRI anatomical reference and correlative imaging.

Stronger and more rapidly slewing magnetic field gradients support more rapid spatial encoding, particularly for metabolite-specific single-shot imaging using echo-planar imaging (EPI) or spiral imaging (See “Acquisition and Reconstruction”). Although the spatial resolution acquired for HP 13C imaging is typically much coarser than for 1H MRI, the factor of ~4 in gyromagnetic ratio leads to the same reduction factor in performance of the gradient system, so 13C experiments are potentially more limited by gradient hardware performance. To date, all human studies have used the commercially-available integrated gradient systems provided in clinical MRI scanners.

Optimization of scanner design has understandably focused on minimization of artifacts in 1H MRI, where devices such as room lights, the gradient amplifiers, and the motors driving the patient bed are checked to ensure that they do not produce RF interference at the 1H frequency, but artifacts may arise at other frequencies. Eddy current compensation is also not always appropriately adjusted for nuclei at other frequencies (74). In order to optimize for 13C, many sites have performed checks on phantoms for RF interference, gradient artifacts, and eddy currents (74), including the use of post-hoc gradient impulse response function characterisation and correction, and some vendors have fixed these issues as well.

Rf Coils

For HP 13C imaging studies in humans, RF coils for both 1H and 13C nuclei are needed, with 1H MRI providing an anatomical reference for registration and optional additional multiparametric MRI readouts. At the Larmor frequency of 13C nuclei, the relative contributions from coil noise compared to sample noise increase compared to 1H (73,75), although sample noise still is likely the dominant contributor for human-sized coils at 32.1MHz - the resonance frequency of 13C nuclei at 3T.

The key requirement for human 13C-pyruvate RF coils are that the coil geometry and sensitive volume must cover the volume of interest in the subject. Table 2 and Figure 4 shows coil configurations that have been used and optimized for applications in different anatomic regions.

Volume resonators are most commonly used for transmit, as they surround the subject to

Provide B1 Transmit Across The Fov (B1

+). While 1H relies on a large birdcage (“body”) coil built into the scanner, 13C transmit coils must be placed inside the bore. This takes up valuable space within the magnet, and also has led to the use of designs with relatively inhomogeneous

B1

+. Many human studies have used Helmholz pair resonators for transmit, including the “clamshell coil”, which has a notably inhomogeneous B1

+ Profile But Has Been Used Because Of

relatively easy integration into the scanner bore. B1

+ Variation Results In Variations In The Flip

angles that control the use of the hyperpolarized magnetization and creates errors in common HP metrics (9,76). The exception are head coils, where birdcage designs with highly

Homogeneous B1

+ can be placed around the head while easily fitting inside the bore. As with 1H MRI, higher SNR can typically be achieved by smaller receive coil elements, such as surface coils or phased arrays, and the majority of 13C receive coils used have layouts similar to 1H phased arrays.

RF coil quality control is important to ensure proper functioning of the coils to provide consistent imaging quality, especially with limited natural abundance 13C signal in vivo. It typically involves 1) a physical integrity check of the coil cables and connectors and 2) phantom SNR tests to check the coil’s performance and to monitor it over time (see Phantoms below). An useful reference for RF coil quality control is outlined in the MRI accreditation program of the American College of Radiology (77) and can be adapted for 13C coils.

Notably, configurations for brain and prostate studies used dual-tuned 1H/13C coil designs, which greatly simplify workflow and registration of 1H and 13C images, as no switching of coils is needed.

(1)

Table 2: RF coil configurations reported for human HP [1-13C]pyruvate studies.

Tx = Transmit

coil, RX = receive coil. The commonly used “clamshell” TX coil is a Helmholz pair design. For 1H RF configurations, all used the Body coil for TX unless otherwise noted, and “repositioned” indicates the 13C coil was removed for 1H imaging. One representative reference is listed for each configuration. The RF coil configurations reported in the reviewed papers are shown in Supporting Table S1.

Figure 4: Examples of RF coil configurations used for human HP [1-13C]pyruvate brain studies. (A,B) 13C Clamshell TX (Helmholz pair) and 2× 4-channel paddle RX arrays. (C) 13C Birdcage volume TX and 32-channel RX array (RX array slides into TX coil). (D) 13C Birdcage volume TX and 24-channel RX array, combined with a 1H 8-channel RX array. Image reproduced with permission from Ref (16).

Phantoms

Since hyperpolarized magnetization is non-renewable, phantoms containing 13C nuclei are important to: 1) test the multi-nuclear capabilities of the imaging system, including all parts of the signal excitation and receive chain; 2) perform calibration measurements before a scan with hyperpolarized nuclei; and 3) perform necessary pre-scan adjustments (see “Prescan Calibration” section). The phantoms currently in use are listed in Table 3. Their composition must provide sufficient 13C signal, with additional considerations of conductivity, stability, chemical shift(s) present, potential for dynamic imaging, and cost. The phantom geometries are typically either compact, in order to be used alongside the subject during a HP scan, or large enough to mimic the inner volume of a RF coil for system testing.

One popular compact design contains enriched 13C-urea at high concentration, typically 8 M, which provides a single resonance, placed inside a small container ~1 mL. The most common recipe mixes 13C-urea in a 90% water/10% glycerol solution, with glycerol used to increase the urea solubility and doping with a Gd-based contrast agent to shorten T1 which increases the potential SNR per unit time. For example, when Dotarem is added at a 3:1000 volume ratio the 13C-urea T1 is around 500 ms and T2 is around 100 ms. However, when testing pulse sequences influenced by T1 and T2, doping should be used carefully. This phantom is suitable for frequency calibration, transmit gain calibration, sequence testing, and as a fiducial marker when placed next to a patient. However, enriched 13C-urea has a relatively high cost compared to natural abundance compounds.

For larger volumes (>100 ml), the phantoms most often used contain undiluted ethylene glycol, glycerol, or dimethyl silicone. These compounds have sufficiently high carbon concentrations to provide sufficient 13C signal even with the 1.1% natural abundance of 13C. These larger phantoms matching the inner volume of an RF coil are useful for coil testing, including transmit

+) And Receive (B1

-) coil profile mapping, as well as to mimic acquisitions using in vivo FOV requirements. In this case, size and conductivity should match the expected subject size in order to mimic coil loading and get a realistic estimation of B1+. Large-volume natural abundance urea phantoms have also been used by some sites, but suffer from higher conductivity compared to biological tissues. Typically, it is easier to increase the conductivity and hence coil loading of the non-conductive phantom by adding NaCl to match physiological loading (16,78).

Dynamic phantoms that aim to mimic metabolite kinetics have also been developed (79–81), and have the potential to more closely mimic the HP experiment, but so far these are not widely used.

Prescan Calibration

Prior to performing an MRI acquisition, the so-called prescan procedure is used to set the shim parameters to maximize B0 homogeneity over the field of view (FOV) or a specific region of interest (ROI), the scanner center frequency (CF), the RF transmit gain, and the receiver gain.

While this calibration procedure is usually automated for 1H, the lack of sufficient natural abundance 13C signal prevents use of automated methods. (Although natural abundance 13C lipid signal has been detected, there are so far no reports on using this signal for prescan.) Table 3 shows current practices across sites.

Maximizing B0 homogeneity is independent of the nucleus and is therefore performed prior to 13C imaging using the 1H water signal and existing shimming tools, such as by a standard automated process (“Auto Shimming”) or using high order shimming routines. Similarly, the 13C CF can be calculated from the 1H CF using a predetermined scaling factor that depends on the target chemical shift (82). Another common approach used is to have a small, high-concentration 13C phantom, e.g. 8M 13C-urea, integrated in the RF coil or placed next to the scan subject (1). The reference frequency can also be based on real-time measurements after the HP injection but prior to imaging (83). Both the CF and B0 shimming are critical when using spectrally-selective RF pulses, as inmetabolite-specific imaging methods, where the desired excitation bandwidths are typically very narrow and frequency offsets can lead to a failure mode that is only apparent after injection.

The calibration of the RF transmit power is typically performed on a small, high-concentration 13C phantom placed near the region of interest during the scan or on a large 13C phantom of similar size and coil loading as the subject, prior to the subject scan. Reference power is often done by sweeping the power in a pulse-acquire sequence (53,62), or the Bloch-Siegert method (52,84). When using a small phantom, the location of the phantom, B1

+ Inhomogeneity As Well

as any shielding effects, e.g., when the phantom is integrated into a coil (1), may degrade the accuracy. Other methods include real-time Bloch-Siegert method measurements after the HP injection (83), and using the stronger natural abundance 23Na signal that is close enough to the 13C resonance frequency to be detected by 13C coils (82).

The receiver gain is predetermined, either systematically based on independent phantom measurements and assuming the dose and polarization of the HP compound is known prior to injection, or based on past HP imaging studies.

Power [Kw]

Phantom(s) - during study Phantom(s) - before study 13C Frequency

8

13C-bicarbonate doped with dimethyl silicone, various

Power [Kw]

Phantom(s) - during study Phantom(s) - before study 13C Frequency

Maximum Values

Table 3: Summary of the imaging systems, phantoms, and prescan procedures used at sites currently performing HP 13C-pyruvate human studies. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. *Previously performed studies with a Siemens 3T Tim Trio. The imaging systems, phantoms, and prescan procedures reported in the reviewed papers are shown in Supporting Table S1.

Summary

Commercially available 3T MRI systems are by far the most commonly used for human HP 13C-pyruvate studies, although a systematic investigation of the impact of B0 has only recently been investigated (73). The multi-nuclear RF transmit and receive chain has proven sufficient for current acquisition strategies, although many sites have observed artifacts due to RF interference, gradient interference, and residual eddy currents when operating at the 13C frequency. A variety of 13C RF coils, tailored for numerous anatomical targets, have been successfully demonstrated, with the main limitation that most transmit coils take up a lot of additional space inside the bore and provide relatively inhomogeneous B1

+ Profiles. The

phantoms used have converged into generally 2 categories - small phantoms containing 13C-enriched compounds that can be used during the study and human-sized phantoms containing compounds with high carbon concentrations but without 13C enrichment that are used to test and calibrate the coils. There are no standardized compositions or geometry, and dynamic phantoms that recapitulate in vivo kinetics would be desirable but are still an emerging area. Prescan calibration procedures were not well defined in most publications, so we surveyed individual sites to determine current practices. Calibration procedures for the B0 field (13C CF and shimming) for most sites take advantage of 1H signal and methods, while methods

For Calibration Of B1

+ is more variable across sites, likely a reflection of remaining challenges in how to perform this calibration. Standardization of both phantoms and calibration procedures would synergistically improve the robustness and reproducibility of HP 13C studies.

Acquisition And Reconstruction

Data acquisition strategies in human HP [1-13C]pyruvate MRI studies must account for multiple chemical shifts, efficiently utilize the non-renewable HP magnetization, and acquire data quickly relative to metabolism and relaxation decay processes. These studies require spectral encoding to separate metabolites, necessitating pulse sequences that efficiently encode up to 5D data (3 spatial + 1 spectral + 1 temporal dimension). RF pulses must efficiently sample without immediately saturating the non-renewable HP magnetization, and sequences must acquire data quickly and be robust to both experimental and physiologic variation (e.g. B1

+ Inhomogeneity,

variation in perfusion) to ensure reproducibility and minimize scan-to-scan variability. This section covers current successful practices for data acquisition in human [1-13C]pyruvate studies, and accompanying 1H imaging, from different anatomic regions, including scan parameters and image reconstruction.

Acquisition And Reconstruction Methods

The acquisition methods used in human [1-13C]pyruvate studies can be classified into 3 categories: 1) MR spectroscopy or MR spectroscopic imaging (“MRS/I”), 2) chemical shift encoding methods, and 3) metabolite-specific imaging (Fig. 5).

Mrs/I Methods Specifically

resolve a spectrum that can be analyzed to extract expected as well as unexpected resonances, making this approach very robust. It was used in many initial studies (1).

Chemical Shift

encoding methods, most commonly the Iterative Decomposition of water and fat with Echo Asymmetry and Least-squares estimation (IDEAL) method, use imaging sequences acquired with multiple TEs and rely on a model-based separation of expected chemical shifts (85).

Metabolite-specific imaging methods use specialized RF pulses that are spatially and spectrally selective to excite individual metabolites which are then typically imaged with fast k-space trajectories such as echo planar imaging (EPI) or spirals (86).

Their Application To Different

organ systems is described below. The image reconstruction methods used in human [1-13C]pyruvate studies have typically been conventional methods (e.g. FFT, non-uniform FFT, or equivalent). The incorporation of accelerated imaging and advanced reconstruction methods including parallel imaging (4,57,87) and compressed sensing (7) has also been applied in human studies for improved spatial resolution, temporal resolution and coverage, but have the potential for additional artifacts as well as SNR losses due to ill-conditioning of the reconstruction (e.g. g-factor).

The Majority Of

published studies do not use accelerated imaging indicating the resolution and coverage achievable without acceleration is currently adequate for successful data collection. Performing coil combination, even with fully sampled data has also been shown to have specific challenges for HP human images: using naive sum-of-squares methods suffer from high noise amplification in the relatively low SNR regime of HP [1-13C]pyruvate (compared to 1H), motivating several HP 13C-specific methods that include data-driven coil sensitivity estimation which have shown obvious improvements over sum-of-squares (11).

More recently denoising techniques have been applied as post-processing of human HP data(41,42,44). The techniques applied are based on spatial-temporal singular value decomposition for unsupervised estimation of signal and noise components. They have shown improvements in apparent SNR in the brain and liver, while care must be taken to choose parameters such as the rank threshold to avoid oversmoothing and overfitting to the estimated signal components.

Prostate Studies

Prostate cancer was the first human application of HP [1-13C]pyruvate (1), and data was acquired with MRS/I methods: 1D dynamic MRS, single-slice 2D dynamic echo-planar spectroscopic imaging (EPSI), and single time point 3D EPSI. Advances in imaging strategies led to the development and application of new acquisition schemes, including undersampled 3D EPSI with compressed-sensing (7), model-based chemical shift encoding methods that use a priori information (47,59), and metabolite-specific EPI (10), all of which can provide volumetric whole-organ coverage and dynamic acquisitions.

The pyruvate bolus arrival in the prostate can vary by ± 10 s between patients, necessitating dynamic imaging to reliably and consistently capture the pyruvate bolus (18). For this reason, all currently ongoing studies acquire dynamic data. While MRS/I, chemical shift encoding, and metabolite-specific imaging can all achieve dynamic imaging, chemical shift encoding and metabolite-specific imaging provide greater dynamic and volumetric coverage (85). For scan prescriptions, the FOV is designed to provide full prostate coverage and typically to match the orientation of the anatomic imaging used for registration. Flip angles used in current studies are constant through time, as quantification with a variable-through-time flip scheme is highly sensitive to bolus timing (8) and errors in the RF transmit (B1 +) field (76).

Heart Studies

Data acquisition methods for 13C imaging in the heart must be designed to meet the demands of significant cardiac motion and blood flow. To cope with the periodic cardiac motion, most human heart studies to date used gating to the diastolic window, the longest cardiac cycle interval, which has reduced motion (2,22,28,30,35,36,38,45,52). The duration of the diastolic window limits the available data sampling time, making cardiac acquisitions the most time-constrained of the HP 13C MRI applications. The most common acquisition approach is metabolite-specific imaging with spiral k-space trajectories (2). Their single-shot imaging capability makes these methods particularly robust to motion effects. Furthermore, spiral k-space trajectories provide rapid k-space coverage and relatively benign flow and motion artifacts. The majority of studies have used 2D multi-slice acquisitions, but 3D encoding has also been used successfully (35).

Brain Studies

For HP 13C MRI of the human brain, the majority of studies have also used 2D (slice selective) acquisitions (10–12,14,16,28,33,40,41,44,51,53,60), with a trend toward volumetric coverage using 2D multi-slice metabolite-specific imaging. 3D metabolite-specific imaging of the whole brain, with phase encoding of the slice direction (34,57), has been shown to provide similar SNR efficiency (88) compared with multislice imaging. A number of studies have employed MRS/I (5,6,29,31–33,50,55) resulting in a spectrum from each voxel, which has the advantage of not requiring a priori information about which peaks to encode. This was important in early brain studies when it was not known which peaks would be detectable. Chemical shift encoding, using a set of images with different echo times and an iterative reconstruction of the individual resonances (i.e. the IDEAL approach (85)), has also been used (12,49,54), with the drawback that coverage in the slice direction was limited due to the time required to acquire multiple echo time images.

Abdomen And Breast Studies

The fundamental approaches to data acquisition and reconstruction in the abdomen and breast are largely similar to the aforementioned applications, but demand attention to particular challenges associated with these anatomic regions, especially relating to respiratory motion.

Although it has been shown that a basic 2D MRSI approach based on phase encoding and FID readout can be successfully applied for HP 13C imaging in breast (15) and kidney (13), major advantages in terms of spatiotemporal resolution and coverage have been realized using tailored approaches based on metabolite-specific imaging (43,62) and chemical shift encoding (43), which have facilitated multi-slice or 3D dynamic acquisitions over large FOVs in the abdomen (4,37,46).

The significant respiratory motion encountered in these regions can directly blur 13C images, and has further favored these rapid acquisition strategies. Motion also degrades B0 homogeneity, which can shift frequency-selective excitation profiles and introduce artifacts into rapid imaging readouts. This makes accurate determination of the acquisition center frequency and shimming essential in these regions which often cover large FOVs. (See “Prescan Calibration” section for more information). In some studies, breath-holding was used to minimize motion effects and enforce frame-to-frame data consistency (42). A pragmatic and reasonably effective approach for dealing with respiratory motion during 13C data acquisition is an initial breath-hold (as long as can be tolerated), followed by free-breathing (46,62).

1H Imaging

Collection of 1H imaging data is essential both for prescribing the 13C acquisition and for interpretation of the resulting 13C data. Multi-planar 1H scouts are acquired prior to 13C acquisition to enable graphical prescription of the 13C imaging region. All human HP 13C-pyruvate imaging studies acquire conventional MRI scans (e.g. T1- and T2-weighted volumes) for anatomic reference, aiming to cover at least the full 13C FOV. Acquiring these anatomic scans as close as possible to the time of 13C imaging (immediately before or after) minimizes potential misregistration between the data sets. Depending on the application, other advanced 1H sequences are also acquired (e.g. diffusion-weighted imaging for cancer imaging).

When contrast-enhanced data is acquired, it is done after 13C imaging, as paramagnetic contrast agents will accelerate 13C relaxation.

Reported Study Parameters

Figures 5 and 6, and Supporting Table S2 shows the reported acquisition study parameters for human HP [1-13C]pyruvate studies published as of September 2022. Figure 5 shows a mixture of MRS/I, metabolite-specific imaging, and chemical shift encoding methods have been successfully used, where spectroscopy-based methods have become less prevalent in recent studies. Figure 6 shows the acquisition timing, including the important start time and interval/temporal resolution, is quite variable across studies.

Figure 5: Acquisition methods used in published HP [1-13C]pyruvate human studies published up to September 2022, classified into: MR spectroscopy and spectroscopy imaging (MRS/I); chemical shift encoding methods, such as IDEAL, that use multiple TEs and model-based reconstructions; and metabolite-specific imaging methods that use spectrally-selective excitation to image a single resonance at a time.

Figure 6: Temporal acquisition characteristics reported in HP [1-13C]pyruvate human studies published up to September 2022. (a) Reported referencing of acquisition start times.

(B)

Acquisition start times reported when using dynamic imaging and when timing was reported relative to the end of the injection. (c) Temporal resolutions. “Not Applicable” indicates dynamic imaging was not used.

Summary

Three general categories of acquisition strategies have been used successfully for human HP 13C-pyruvate studies: MRS/I, model-based chemical shift encoding (e.g. IDEAL) methods, and metabolite-specific imaging methods. These have enabled successful studies in the prostate, heart, brain, abdomen, and breast. Recent studies increasingly have used the imaging-based strategies of metabolite-specific imaging and chemical shift encoding which are the fastest methods, although a heads-to–head comparison between techniques has not been performed.

Metabolite-specific imaging is quite popular because of its speed and compatibility with single-shot imaging, but is sensitive to B0 field variations and thus requires careful calibrations. Nearly all studies surveyed acquired data dynamically, allowing measurement of the bolus and metabolite kinetics. The exact timings and associated flip angles vary quite widely across reported studies, with no consensus yet as to how to choose these parameters. Image reconstruction is typically done directly using Fourier Transform methods, and accelerated imaging strategies are uncommon.

Data Analysis And Quantification

This section covers the analysis of data from human HP [1-13C]pyruvate studies, including modeling and metrics, visualization, as well as considerations for how to store data and metadata. Depending on study design, the analysis may need to give quantitative or semi-quantitative output reflecting a biological process or may just reflect a contrast between different regions of interest for quantitative evaluation.

Metrics

Figure 7: HP [1-13C]pyruvate raw data (A) have typically been quantified using four categories of metrics depending on the acquisition. Data acquired as a single time point are often quantified using normalized metabolite images or metabolite ratios (B). Dynamic data can be quantified using normalized metabolite images or metabolite ratios (B), or with metabolite timings such as time-to-peak (TTP) or pharmacokinetic (PK) models (C). The latter two require the data to be time-resolved. [1-13C]alanine and 13C-bicarbonate are analyzed similarly to [1-13C]lactate but omitted here for display.

Metabolite images are commonly used as summary metrics for HP MRI data, often including some form of normalization as well as summed over time as an area under the time curve (AUC) (17). These are analogous to the visual evaluation that is most used for routine clinical work (89,90). In these metabolite images, we expect that the [1-13C]pyruvate AUC signal is predominantly weighted towards perfusion and uptake, while [1-13C]lactate, [1-13C]alanine and 13C-bicarbonate AUCs represent metabolic conversion. The strength of this approach lies in its simplicity and relatively few underlying assumptions. Limitations to the use of single-metabolite images or AUCs include sensitivity to inhomogeneous coil profiles (57,87,91), the acquisition strategy and acquisition parameters, pyruvate polarization and concentration level, and signal relaxation rates (92). Further, the reader must be careful to interpret all the images in conjunction to better understand the underlying biology; for example, increased [1-13C]lactate in the presence of decreased [1-13C]pyruvate delivery can have a very different meaning compared to increased [1-13C]lactate with increased [1-13C]pyruvate delivery.

In an attempt to address variations in coil sensitivity, polarization level, and pyruvate delivery, AUC images are often computed by normalizing to a specified parameter, such as the maximum pyruvate or average lactate signals, or presented as a ratio such as lactate/pyruvate or divided by “total Carbon” - the sum total of HP 13C signal observed across all metabolites. The AUC ratios between metabolites and pyruvate are proportional to the corresponding forward kinetic rates (81,93), but are not directly comparable to rate constants when magnetization loss rates (e.g. relaxation and losses due to signal excitation) differ between studies. Similarly, the ratios between the produced metabolites (e.g. bicarbonate/lactate) can reflect the balance between downstream metabolic pathways (12,55). Care must be taken to consider how AUC images are calculated and normalized before comparing values between studies.

To further quantify the interpretation, pharmacokinetic (PK) modeling approaches were developed to compute the apparent kinetics of pyruvate-to-metabolite exchange (92,94–99). These yield semi-quantitative to quantitative apparent rate constants, given in s-1. Some models require a vascular input function, while others avoid this requirement (95). PK models can explicitly account for acquisition-specific details such as excitation angle and repetition time, and thus may reduce the effects of these details on quantification. An input-less model, provided in the Hyperpolarized-MRI-Toolbox (https://github.com/LarsonLab/hyperpolarized-mri-toolbox) (100) and thus frequently employed for human data, has been shown to fit well and robustly to prostate and brain data (8,20). PK models are quantitative in nature, arguably provide more relevant biological information (8,20), and appear to be reproducible across sites (51). However, rate constants derived from PK models are still apparent rates, and likely do not reflect a single biological characteristic.

Some additional considerations include whether complex or magnitude data is used, as the noise behaviors will impact the analysis differently. Additionally, cut-off thresholds or other criteria may be used to identify and avoid voxels with insufficient SNR before analysis to improve robustness (20,41).

Regardless of the analysis approach, the underlying biology is not always clearly represented by the data; instead, the metrics may be influenced by perfusion, barrier permeability, intercellular shuttles, enzyme activities, co-substrate concentrations, or combinations thereof, depending on the organ and disease of interest (19,43,94,101–103). This may be addressed by incorporating complementary information. As an example, HP 13C pyruvate data is influenced by perfusion, and thus addition of perfusion MRI could be important for interpretation (98,104,105).

All the methods outlined above have been explored in clinical studies, described in Supporting Table 3 and summarized in Figure 8. As of September 2022, approximately 52% of studies involving human subjects report rate constants derived from a PK model with a few different models reported. A nearly equal fraction (51%) of the studies report AUC ratio values.

Approximately 66% of these studies report metabolite-specific images or AUC values. About 40% report SNR values; this metric is particularly frequent in manuscripts that describe technical developments for clinical HP MRI. Approximately 16% of these studies summarize model-free metrics, and 10% report measurements from a single timepoint. Most studies report a combination of quantities.

Figure 8: Reported metrics used for analysis in HP [1-13C]pyruvate human studies published up to September 2022.

Visualization

A wide variety of approaches have been used for visualizing data from human HP 13C-MRI studies. The challenges and practical considerations are: 1) choosing the appropriate metrics to display, 2) how to encode the parameters (e.g. the colormap), and 3) choosing how to provide anatomical context and other multi-parametric data. The choice of visualization also depends on the goal which could be for diagnostic interpretation, but also quality control, reproducibility among readers and publication.

Metrics

The choice of HP 13C metrics is described in detail above. At this stage in HP 13C development where there is no standardized metric, often a combination of metabolite images and ratios or PK model parameters are shown.

Parameter Encoding

The mapping function chosen should provide an adequate, often quantitative, impression of the parameter mapped. There is a consensus in the visualization field that perceptually uniform maps are best suited to visualize continuous parameters, like the greyscale typically used by radiologists as well as other monochrome (black to blue) and color ranges (fire-type, rainbow-type) (106,107). Multi-color heatmaps have been the most frequently employed method for HP 13C data, while greyscale has infrequently been used but it ensures there is no coloring-based bias as well as facilitating later reuse (Fig. 9a). Among the color schemes employed in the clinical HP 13C literature, fire-type scheme seems to be the most common [similar to “Plasma” or “Inferno” in matplotlib.org]. Next most commonly employed is the rainbow-type scheme [similar to “Rainbow” in matplotlib.org].

Anatomical Context

HP MRI faces the challenge that it does not necessarily depict the anatomical features, similar to PET, and thus requires an anatomical reference. Most often, a grayscale anatomical image is overlaid with a HP colormap (Fig. 9c,d). This approach is very intuitive, but can skew perception as the grey-scale anatomical reference may affect the brightness of the HP data (e.g. signal in the skull). This bias does not occur when showing adjacent maps (Fig. 9a, b). Here, anatomical outlines may help to provide reference (Fig. 9b).

Related Journal Articles & DOI Links

Selected peer-reviewed publications relevant to 12 Lead ECG Acquisition. Click the DOI to access the full paper (may require institutional access).

Why Choose Us?

Bangalore guidance for robotics, Spectre and autonomous systems projects.

Spectre & Simulation

Gazebo, cloud twin and Webots worlds with navigation, SLAM and control stacks.

Control & Planning

Compliance, deep learning control, path planning and behavior trees.

Hardware Bring-up

Motors, sensors, ESP32/STM32 firmware and HIL validation paths.

Report & Viva

University-format documentation, PPT and viva preparation.

FAQ

Spectre, Gazebo, NVIDIA cloud twin, MATLAB/Simulink, Webots, Blynk / ThingSpeak, plus Arduino/STM32/ESP32, cameras, LiDAR and motor drivers.
Yes — simulation packages, hardware guidance, report, PPT and viva Q&A.