Enquire Now
70+ Topics · Spectre · Spectre · cloud sim Sim · MATLAB · Webots · Hardware · Bangalore 2026

Neural Network Robot Navigation

Simulation · Control · Perception · Hardware — 12 Lead ECG Acquisition — hardware, sensors, cloud dashboards and protocols (Spectre, REST, CoAP, WebSockets) for BE BTech MTech students. Final-year robotics support with Spectre stacks, simulation worlds, reports and viva from Bangalore.

70+
Related Topics
6+
Sim & HW Tools
4.9★
573 Ratings

Abstract

Random Neural Networks (RNNs) are a class of Neural Networks (NNs) that can also be seen as a specific type of queuing network. They have been successfully used in several do- mains during the last 25 years, as queuing networks to analyze the performance of resource sharing in many engineering areas, as learning tools and in combinatorial optimization, where they are seen as neural systems, and also as models of neurological aspects of living beings. In this article we focus on their learning capabilities, and more specifically, we a general description of these models using almost indistinctly the terminology of Queuing Theory and the neural one. We present the standard learning procedures used by RNNs, adapted from similar well-established improvements in the standard NN field. We describe in particular a set of learning algorithms covering techniques based on the use of first order and, then, of second order derivatives. We also discuss some issues related to these objects describes their most relevant applications, and also provides a large bibliography.

neural-network-robot-navigation Diagram
Figure: System Model & Architecture for Neural Network Robot Navigation

Neural Networks, Random Neural Networks, Supervised Learning, Pattern

Ntroduction

Supervised Learning is an area of the Machine Learning field that refers to a set of problems wherein the information is presented according to an outcome measurement associated with a set of input features. The information is presented as a dataset of labeled samples. The aim is “to learn” the relationship between input and output features. This learning process is done based on a set of examples in order to generate a learning model with the power of “generalising”, this is to make “good” predictions for new unseen inputs. The research on Neural Networks (NNs) is considered to have started with the work of Warren McCulloch and Walter Pitts in 1943 (McCulloch and Pitts, 1943), and it has produced a rich literature with a strong concentration of papers in the 80s and 90s. In the 80s Rumelhart et al. explored the relationship between Parallel Distributed Processing (PDP) systems and various aspects

Arxiv:1609.04846V1 [Cs.Ne] 15 Sep 2016

of human cognition. The authors defined a general framework of a PDP system reactivating the research on connectionist models (Rumelhart et al., 1986b). The most popular PDP systems are NNs. In the last decades several books and journals have been dedicated to the research on NNs. The interest in the NN area arises from both its theoretic aspects and its different fields such as engineering, biology, pattern recognition, theoretical physics, applied mathematics, statistics, etc.

neural-network-robot-navigation Diagram
Figure: System Model & Architecture for Neural Network Robot Navigation

There are many types of NNs, and the related literature is huge. This article focuses on a particular class of NNs called Random Neural Networks . The RNN model was introduced by E. Gelenbe in 1989 (Gelenbe, 1989a,b). RNNs are mathematical objects that combine features of both NNs and queueing models. They been successfully employed in many types of applications: in learning problems, in optimization, in image processing, in associative memories, etc.

neural-network-robot-navigation Diagram
Figure: System Model & Architecture for Neural Network Robot Navigation

Here, we are specifically interested in the situations where the model is applied for solving supervised learning tasks.

A Rnn Is A Pdp Composed Of A Pool Of

interconnected nodes, which process and transmit information (signals) between them. Each node is a simple processor and it is characterized by its state, a whole number. The nodes receives two kinds of signals (negative and positive) from their neighbors or from outside.

When a negative signal arrives to a node, it produces an effect that can be related to neural inhibition, its state its decreased by one. The arrivals of positive signals provoke the opposite effect, the state is increased by one. The fire of signals by the nodes is modeled by Poisson processes, and the pattern of connectivity among the neurons follows stochastic rules.

The design of the model was inspired from the biological behavior of neuron circuits in the neo-cortex. The model considers the following biological aspects: the action potentials in the form of spikes, the exchange of excitatory and inhibitory signals among the neurons, the synapses (weighted connections between two neurons), random delays between spikes, reduction of neuronal potential after firing, arbitrarily topology (Gelenbe, 1989a). The model has been also proven very powerful, from the computational viewpoint. In (Gelenbe et al., 2004b) the authors shown that under certain algebraic hypothesis the RNN is an universal approximator. Besides, it can be easily implemented in both software and hardware. In order to apply the model for solving learning tasks, several learning algorithms have been adapted from the classic NN to RNNs, such as the Gradient Descent (Gelenbe, 1993a) and Quasi-Newton methods (Basterrech et al., 2011; Likas and Stafylopatis, 2000). The number of applications of the model in the learning area is very large, but the model has been also applied to solve combinatorial optimization problems, such that the Traveling Salesman Problem or the Minimum Vertex Covering Problem (Gelenbe and Batty, 1992; Gelenbe et al., 1993).

Ain Contributions

The first overview about RNN was presented in 2000 (Bakircioğlu and Koçak, 2000). A survey about RNN focused on networking application and self-aware networks was intro- duced in (Sakellari, 2010). Another general and helpful survey about RNN was presented in (Timotheou, 2010), where the authors describe the main applications of RNNs, cov- In (Georgiopoulos et al., 2011) the authors focused on RNN for solving learning problems, they identified some drawbacks of the RNN learning applications. In addition, an extensive literature about RNN was presented in (Do, 2011). In the 25th anniversary of the RNN model, we present this tutorial that contains the following contributions with respect to the previous published material.

• We introduce the model as a simple computational processor in a PDP framework, allelism between this particular PDP and the model as belonging to the queueing area.

• We provide a structured overview about the numerical optimization algorithms used for training RNNs. We introduce algorithms that use the first derivative information present Quasi-Newton methods that use the information of the second derivative of the cost function. In this practical guide, all the algorithms used for training are shown in detail following a homogeneous format.

• We present a critical review and new perspectives on RNN in supervised learning. We discuss technical issues concerning stability problems in the model itself, as well as problems related to the parameters’ optimization in the learning process. We discuss some points related to the computational advantages of the model, as well as about its weaknesses and limitations. The overview concludes with remarks concerning some new trends and future research lines.

In addition, this article presents an overview of some selected applications of the RNN in the supervised learning area. In particular, we comment on two applications where the experimental results show a better performance of the model with respect to other techniques of the literature.

Organization Of The Article

This article is structured as follows. Section 2 formally describes the RNN model as a learn- ing tool and in the framework of queueing theory. Section 3 presents algorithms for training the RNN model. It starts with a formal specification of the computational problems in supervised learning. In Sec. 3.2 we give a general description of RNN in the learning con- text. We present the Gradient Descent algorithm in Sec. 3.3, and we introduce second order optimization methods in 3.4. We describe the following algorithms: the Broyden-Fletcher- Goldfarb-Shanno in Sec. 3.4.1, the Davidon-Fletcher-Powell in Sec. 3.4.2, the Levenberg- Marquardt in Sec. 3.4.3 and one variation of it in Sec. 3.4.4. We present a critical review about the RNN model for solving learning problems in Sec. 4. Section 5 presents an overview of applications. We conclude and present new research trends in Sec. 6.

The Random Neural Network Model

This Section formally introduces the RNN model. It has four parts. First, we describe a single neuron (Random Neuron) as an elementary processor. Second, we present the RNN as a system composed by interconnected neurons. Third, we review the model in the framework of queuing networks. The section ends introducing the different topologies and structural concepts of the RNN.

Random Neuron (Rn)

A Random Neuron (RN) is a real parametric function of two real variables, with a real parameter called the neuron’s rate. The input variables are assumed to be non-negative. The rate is positive. If x ≥0 is the first input variable, y ≥0 is the second one, and if r > 0 is the rate of the neuron, then the output is the real z given by the expression

(1)

See that a RN is characterized by its rate r. We can see the neuron as an input-output system with two “input ports”, one for x and the other one for y, and one output port for z. The ports associated with the output and with the first input value are called positive; the input port corresponding to the second input variable y is called negative (the reason for this is explained later), but all the variables involved are non-negative real numbers.

Figure 1 shows a neuron as an input-output device. When x ≥r +y we say that the neuron is saturated. Figure 1: A zoom on a random neuron (RN) seen as a “black-box” system; the inputs are the reals x, y ≥0; the parameter is the rate r > 0, and the output is the real z; we say that the first input variable x is connected to the positive input port of the RN (depicted ‘+’) and the second input variable y to the negative input port (depicted ‘−’); the output port is also said to be positive (and it is depicted ‘+’ in the figure) The output value z is seen as a measure of the activity of the neuron (as in most input- output systems). As such, see that z is increasing in x and decreasing in y. In real neurons, which also are input-output systems, the input signals belong to two types, excitatory signals, which are those contributing to the neuron’s activity measured by its output (the higher the excitatory signal, the higher the neuron’s activity) and inhibiting inputs playing the opposite role. This is why we call positive the signals arriving at the ‘+’ input port, and negative those arriving at the ‘−’ one.

We will say that a RN is controlled if its output z is modified according to the rule

(2)

So, in this case the RN’s output is always less than or equal to its rate, and it is equal to its rate when the neuron is saturated. In the case of the initial definition (1), the neuron is said to be uncontrolled.

Random Neural Network (Rnn)

A Random Neural Network (RNN) is a network composed of N interconnected RNs, that implements a function from R2N into RO, for some 1 ≤O ≤N, in the following way. We are given N RNs denoted 1, 2, . . , N (that is, we are given N strictly positive reals r1, r2, . , rN), and two N × N matrices denoted by P+ = (p+

Ij), Whose

components are probabilities. Both matrices and their sum are substochastic, that is, for

Ij + P−

ij) < 1 are called output neurons. We denote by O their number (so, 1 ≤O ≤N). The network outside is often referred to as the neuron’s environment with which the system operates (Rumelhart et al., 1986b).

Let us denote the 2N input variables of the network as x1, . . , xN, y1, . , yN. Then, the output of the network is the set of outputs of each of its output neurons. We need only to specify how are determined the inputs to the N RNs (the outputs are given by the previously described rules, in the uncontrolled or controlled cases). Let us call ui (respectively vi) the positive (respectively negative) input to neuron i. Then, the following equations must be

N Words, The Fraction P+

ji of the output zj of neuron j adds to the positive input ui to

Neuron I, And The Fraction P−

ji of the output zj of neuron j adds to the negative input vi to i. Of course, this means that the reals z1, . . , zN must satisfy the non-linear system of

Ri,

i = 1, 2, . . , N. This needs some technical discussions about the existence and unicity of solutions to this system, as we will see below.

(3)

we have 0 ≤di ≤1, and that neuron i is an output neuron when di > 0. We can also say that the network of neurons sends the part dizi of zi through the output port of i.

Observation:

in general in the learning applications, we use a RNN with N neurons as a function from RI to RO where I < 2N or even I < N, by setting 2N −I of the standard 2N input variables to a fixed value (typically to 0). We will see soon this frequent situation.

An important particular case covering all the applications done so far for these objects as learning tools is as follows. The network with N neurons implements a function with I ≤N input variables and O ≤N output variables. The input variables are denoted by x1, . . , xI, which are all connected to the positive port of I neurons called input neurons. In other words, no input variable is connected to a negative port. The function output is the set of outputs generated by the O output neurons. A group of neurons can have no interactions with the environment (when I + O < N). We call those units hidden neurons. Note that a neuron can be both an input and an output one.

A Queueing View Of The Random Neural Networks

The RNN method has been used with two different interpretations both referring to exactly the same mathematical model. One is the already described type of interconnected RNs. Another one is a type of queueing systems called G-queues and G-networks.

The First

interpretation is often employed in the Machine Learning contexts and the second one is applied in Performance Evaluation, for example. We begin by describing a single queue where customers arrive according to a Poisson process, say with rate λ > 0, and service times are exponentially distributed with param- eter r > 0. It is assumed that service times are mutually independent and that they are also independent of the inter-arrival times. This server queue is named M/M/1 queueing model (Kendall, 1953). At any time t the state of the system S(t) is the number of cus- tomers present in the queue. The queue storage capacity is infinite.

The Stochastic Process

{S(t), t ≥0} is a continuous time homogeneous Markov process on the non-negative inte- gers. We define the utilization factor of the queue as the ratio ϱ = λ/r. When the process

P(K) = Lim

t→∞P(S(t) = k) = ϱk(1 −ϱ).

(4)

A Jackson queueing network consists of N interconnected queues with the following characteristics. For each queue i the service time is exponentially distributed with rate ri. When a customer completes the service at queue i, it will either move to queue j with routing probability pij or leave the network with probability di (di = 1−PN

J=1 Pij). Customers Arrive

from the environment to queue i according to a Poisson process with rate λ+

I . At Any Time

t, the system state is the vector S(t) = (S1(t), . . , SN(t)), where Si(t) denotes the number of customers in queue i at time t. The assumptions about the independence among the

Processes Can Be Summarized As Follows:

• arrival processes, service processes and switching (routing) processes are independent

Of Each Other;

• at each server, the service times are independent of each other; • at each switching point, the successive switching results are independent of each other. We define Ti as the mean throughput at queue i. In order to avoid a trivial case, we

Assume That At Least One Of The Λ+

i ’s is non-zero (strictly positive). In addition, assuming that the system is irreducible (for any two nodes i and j in the Markovian graph there exists a path from i to j), and in equilibrium, Ti for all i can be determined by solving the flow

(5)

The strongly connected property of the Markovian graph implies that exists an unique (and strictly positive) solution. The utilization factor of queue i is given by ϱi = Ti/ri. A G-network (or equivalently, an RNN) is an extension of a Jackson’s network where there is a new entity in the system: negative customers. As in the previous network, in a G-network there are Poisson arrivals, probabilistic routing among the queues, exponential service rates and usual independence among the corresponding stochastic processes. There are two types of customers in the system, positive ones that operate as we defined for the Jackson network, and the negative ones that operate as follows. When a negative customer arrives at a non-empty queue, it destroys a positive customer in this queue, if any, and disappears. If there are no customers in the queue, a negative customer does not operate, it just disappears from the system. In several works negative customers are referenced as signals, thus there are two entities, customers (positive customers) and signals (negative customers).

In (Gelenbe, 1989a, 1991a) Gelenbe shows that, in an equilibrium situation, the ϱis

(8)

with the supplementary condition that, for all neuron i, we have ϱi < 1. An important result associated with open Jackson networks and with G-networks is called the product form theorem. Gelenbe proved that under Markovian assumptions G-networks have a product form equilibrium distribution. This means that the joint equilibrium distribution of the queue states is the product of the marginal distributions. For more details see (Gelenbe, 1989a).

Observation: Let us unify the notation that will be used through this article. So far we introduced the RNN as a function, next we presented the concept using a queueing point of view. In the rest of the article, we follow the most often used notation presented in (Gelenbe, 1989a). Let N be the number of interconnected neurons. For each neuron i its service rate is denoted by ri, the value at its positive port is denoted by T +

The Positive Input Value Λ+

i (the Poisson rate of the customers coming from outside), the

Negative Input Value Λ−

i (the Poisson rate of the negative customers coming from outside), and the probability to send information to the environment denoted by di characterize the interaction of i with outside. The output of neuron i is its activation rate ϱi. The connections between two neurons i and j are given by the probabilities p+

I,J. Figure 2 Shows

the main parameters involved in a RNN. We will introduce in our notation the concept of weights. For any two neurons i and j, they are defined as: w+

I,J = Rip−

i,j. The first one is called positive weight and the second one is called negative weight. Note that the weights are, by definition, positive reals. In the context of NNs, the traditional notation used for the weight connection (direct edge) between the nodes i to j is often denoted as (j, i). In the RNN context, the reverse order is traditionally used. This originates in the first paper about supervised learning with RNNs (Gelenbe, 1993a).

Figure 2:

A representation of a RN. The figure shows the main parameters involved in a RN embedded in a network.

The Network Topology

So far, we defined the RNN as a parallel distributed system composed of simple processors (RNs). Therefore, the network is a graph where the RNs are their nodes; the existence of an arc between two nodes is given by certain probability. The two most common topologies of networks are multi-layer feedforward and recurrent networks.

Feedforward Topology

We start describing the feedforward case. The identifying property is that there are no cyclic connections among the neurons, no circuits in the (directed) graph. The architecture of the graphs consists of multiple layers of neurons in a directed graph. There are three types of layers popularly known as input, hidden and output layers. The neurons can have only connections in a forward direction, from the input neurons to the output neurons, traveling through the hidden ones. Only neurons belonging to the input and to the output layers can exchange information with the environment. The activity rate for each output neuron is computed using a forward propagation procedure.

A Representation Of A Feedforward

network with one hidden layer is illustrated in Figure 3.

Figure 3:

A representation of a Feedforward Neural Network. The figure shows a network with a single hidden layer. The flow of information is from the the input neurons through the output ones. In this example there are 5 input neurons full connected to 9 hidden neurons, and the hidden neurons are full connected with 4 output neurons. A network with this topology is used for mapping a relationship from a 5-dimensional space into a 4-dimensional space.

The feedforward case has been widely used in supervised learning due to the fact that training process is much faster than in the recurrent case. Besides, the feedforward networks are easier to analyze than networks with recurrent topologies. One advantage is that the non-linear system of equations (6), (7) and (8) can be formally solved. Then, we can express the activity rate of the output units as functions of the inputs variables of the system. Let I be the number of input neurons, H is the number of hidden neurons and let O be the number of output neurons. We arbitrary index the input neurons from 1 to I, the hidden neurons from I + 1 to I + H and the output neurons from I + H + 1 to I + H + O = N.

We can compute the activity rate of the neurons using a forward procedure as follows. At the first step, we compute the activity rate of the input neurons, next the activities of the hidden neurons and finally those of the output neurons. Input neurons are the only ones that receive signals from the environment; so we set λ+

I = 0 For All I ∈[I + 1, N]. The

activity rates are given by the following explicit expressions:

,

∀o ∈[I + H + 1, N]. More general feedforward networks consist of successive layers where the signals can circulate only in one direction.

Recurrent Topology

In the case of recurrent networks circuits are allowed. The existence of directed cycles has an important impact in the model: we can not compute the rate activities of the output neurons as functions of the network inputs (except, of course, when N ≤4). A RNN with circuits connects to the concept of dynamical systems, rather than to functions, there is an idea of time implicit in the model. For simplicity we assume discrete time and we avoid to use temporal notation in ϱ. At each time instant, the network is characterized by an internal state ϱ formed by the activity rates ϱ = (ϱ1, . . , ϱN). When an input pattern is presented to the network, the network updates its internal state. For computing the network state we must solve the system of equations (6), (7) and (8), where the unknown parameters are ϱi, T +

And T −

i , for all i. For solving this system is necessary to perform a fixed point procedure (a summary about this computation is given in (Timotheou, 2010)). The output of the network is given by the state of the output neurons. Unlike the feedforward case, a recurrent network can use its internal states to process sequences of inputs. As a consequence, the recurrent case is often used for solving problems where the dataset presents temporal dependencies.

3. Random Neural Networks in supervised learning problems In this Section we present the algorithms used for learning. The Section starts with a formal definition of the supervised learning problem. Next, we present the algorithms of Gradient Descent type for training the RNN. Then, we introduce the algorithms that use the Hessian or an approximation of the Hessian matrix for training the RNN. We close the Section with a general discussion that covers topics such as: limitations of the algorithms in the numerical optimisation, analysis of the algorithmic time complexity, applications of the RNN concepts in the Reservoir Computing area, a discussion about the computational power of the RNN for approximating any regular function, and an analogy of the model with other NNs.

Specification Of A Supervised Learning Problem

We begin by specifying a supervised learning problem. Given a dataset L = {(a(k), b(k)), k = 1, . . , K}, where a(k) ∈A and b(k) ∈B, with A and B some given finite dimensional spaces (typically, sets of real vectors, or of vectors of elements in some alphabet, or a mix of both types of objects). The learning procedure consists in inferring a mapping ν(a, L) in order to predict the b values, such that some distance d(ν(a(k), L), b(k)) is minimized for all k ∈{1, 2, . , K}. We denote by I the dimension of the input vector a and O the dimension of the output vector b.

For each instance a(k), let us denote ϱ(k) the output produced by the network, that is ϱ(k) = ν(a(k), L). The distance above referred is a function L(·) named loss function or cost function that measures the deviations of the model predictions are the criteria of Sum-of-Squared Errors (LRSS) and the Kullback-Leibler distance (LKL), also called cross-entropy (Hastie et al., 2001; Schumacher et al., 1996). The RSS is defined

(9)

where ci = 1 when i is an output neuron, otherwise ci = 0.

There Are Several Slight

modifications of the previous distances, one of those is the Mean Square Error (MSE) given

(10)

In supervised learning when the targets are categorical or discrete variables the problem is called classification problem; when the target is a real vector, the problem is called regression problem.

Random Neural Network As A Learning Tool

A first approach for applying the RNN model in supervised learning tasks was introduced at the beginning of the 90s by Erol Gelenbe (Gelenbe, 1993a). This procedure is based on the classical backpropagation algorithm (Rumelhart et al., 1986a). As in practice, the input and output variables in learning problems are bounded with known bounds, the algorithm described in (Gelenbe, 1993a) assumes that a(k) ∈[0..1]I and b(k) ∈[0..1]O, for all sample k. The RNN model as a predictor is a parametric mapping ν(a, w+, w−, L), where the parameters w+ and w−are adjusted minimizing the loss function. In (Gelenbe, 1993a) was considered the quadratic error presented in the expression (10). The network architecture is defined with I input nodes and O output nodes. There are not additional constraints regarding the network topology, that means the network can be feedforward with one or several layers, or it can be recurrent network. We set the port of the input neurons each time that an input pattern a(k) is offered to the network. The inputs to the positive ports are

I

; the negative ports of input neurons are conventionally

Set To Zero (Λ−

i = 0). The output of the model is a vector of the activity rates produced by the output neurons. The adjustable parameters of the mapping are the weights connections among the neurons. We follow this Section describing the optimization algorithms that have been introduced over the last decades.

The Gradient Descent Optimization Algorithm

We can now describe the gradient-based algorithm that was used so far for training the RNN model (Gelenbe, 1993a). We define two set of neurons I and O that correspond to the set of input neurons and the output neurons, respectively. The weights are initialized

And W−(0)

u,v , for all u and v. At the τth-iteration, we select a

A(K), B(K)

, k = 1, . . , K, where k = τ −1 mod K + 1. The weight correction is computed following the delta learning rule (Rumelhart et al., 1986a), meaning that the weight correction is proportional to the partial derivative of the loss function with respect to each weight. From (3), the service rate of neuron i verifies

(11)

for all i ∈I ∪H. Also note that ri is a free-parameter when i is an output neuron. At each step τ, the current weight value descends in the direction of the negative gradient of L(·); the update rule for positive and negative weights (denoted with superscript ∗) of

(13)

The parameter η ∈[0, 1] is called learning factor. It is used for tuning the convergence speed of the algorithm. Here, we set ci = 1 for all output neuron i, otherwise ci = 0.

Equation (13) leads to the following simplified expressions.

,

otherwise. Then, denoting by ϱ the vector of activity rates ϱ = (ϱ1, . . , ϱN):

(14)

where I and Ωare N-dimensional matrices, I is the identity, and the element (i, j) of Ωis

(15)

The partial derivatives were explicitly computed for a feedforward RNN with a single layer in (Georgiopoulos et al., 2011). An online version of the Gradient Descent (GD) algorithm is an iterative method that processes the input patterns one-by-one realizing the following two main operations: to compute the direction of the gradient of the loss function and to update the weights using the expression (12). The method can either be stopped using an arbitrary number of iterations or when the performance measure is smaller than some threshold value. The online version of the GD algorithm is specified in Algorithm 1. In contrast, an offline training scheme (also called batch algorithm) uses the whole pattern data before modifying the model parameters.

An input is offered to the network, the direction of the gradient is computed. When all data have been presented, the gradient directions are averaged. Finally, each weight is updated using the average of the gradient directions. In the Machine Learning literature coexists two opposite views concerning these two training schemes. As far as we know there has been no consensus on which scheme (on-line or offline) is more efficient for training a learning model (Nakama, 2009; Wilson and Martinez, 2003).

3.3.1 Slight modification of the gradient descent algorithm A slight variation of the GD algorithm for RNN was proposed in (Basterrech and Rubino, 2013b). The authors increase the amount of adjustable parameters during the training of the gradient descent algorithm without modifying the network topology and the time complexity of the algorithm.

They consider as adjustable parameters in the training objective the

For All Hidden And Output Neuron I, And The

service rate ri for all output neuron i. Considering the training error given by the expression (10), the update learning rule is given as follows. Let ∆, P, Λ+ and Λ−be matrices of dimensions N ×N, where the matrix

Pi,I = Ρi,

and the matrices Λ+ and Λ−have at the position (i, u) the value ∂ϱi

Λ−= ∆−1P((I −Ω)−1)T,

Algorithm 1: Specification of the GD learning algorithm for the RNN model (online version).

Nputs

: {(a(k), b(k)) : k = 1, . . , K} (training dataset), η (learning rate), maxIters (max. number of iterations), the topology of the RNN (that is, the routing

Τ = 0;

2 Initialize all weights (for instance, randomly); // we get w∗(0)

Λ+ = A(K); // Read Input

For all i̸ ∈O compute ri using (11); // weights are those at τ −1

For Many Relevant Technicalities

Evaluate convergence. where Ωwas defined in the expression (15). Then, for each input pattern (a(k), b(k)) at the

(18)

where [I −Ω]−1, Λ∗and T ∗are computed using the current input (a, b) and Λ∗

U Denotes

the column u of the matrix Λ∗.

Technical Issues

We discuss here some technical issues related to the learning process, well illustrated by the GD procedure. Recall that the model can be seen as a network of queues (it is actually born in this way). This has some consequences, that have an impact on the design algorithmic decisions. A first point concerns the use of (12) for updating the weights. Indeed, it may happen that (12) leads to a new value for some weight that is negative or null. This does not fit the analogy with a network of queues, or even a network of spiking neurons where the weights model mean throughputs of spikes: weights should be positive numbers. We

Can Accept A Null Value For Some W∗

u,v interpreted as the fact that there is actually no such connection between u and v, but a negative one has no interpretation. The usage is to respect this analogy, modifying the updating rule such that the weights are never negative.

Three possible approaches are proposed in (Gelenbe, 1993a):

(19)

and in the case that some weight is assigned value zero, then to apply one of the

Following Rules:

– fix a null value to this weight, and do not change it anymore in future iterations; – assign a zero value to this weight, but allow positive updates in subsequent iter- ations, keeping using (19).

• Another option is to decrease the value of η and update again the weight using (12). If the new weight is still negative, repeat until obtaining a positive number or stop the loop using some control parameter. Formally, this means that the learning factor becomes a variable parameter in the method. In a nutshell, the global idea in descent methods is to decrease little by little the learning factor, as we get closer and closer to a local minimum. Global accuracy can also be improved (but also cost) if η(τ), say, is built by a supplementary optimization process (this is called line searching in the area) (Press et al., 2002). We do not enter these details here.

• An alternative option was presented in (Likas and Stafylopatis, 2000). The authors

U,V

2. Then, instead of using the expression (14), we proceed as follows

(20)

3.3.3 Computational cost of the gradient descent algorithm When one data pattern is presented to update each weight in the network the main com- putational effort consists of computing [I −Ω]−1 using (14) (Gelenbe, 1993a). This effort has O(N3) time complexity. A remark made in (Gelenbe, 1993a) consists in that when a m-step relaxation method is applied the time complexity decreases to O(mN).

Additionally, the general scheme of the algorithm can be adapted when we use a feedfor- ward RNN. In this case the matrix I −Ωbecomes triangular, so the computational cost of computing its inverse decreases to O(N2). Also, the computational effort to compute each activity rate in feedforward networks is reduced, due to the the activity rate of any neuron depends only on the neurons in the preceding layers.

Second Order Optimization Methods

In this Section, we present the optimisation methods for RNN that use the information given by the second derivative of the loss function. We start introducing the Gauss-Newton (GN) methods, next we explore the Quasi-Newton (QN) techniques. We present four particular algorithms developed for training RNNs: the Broyden-Fletcher-Goldfarb-Shanno (BFGS), the Davidon, Fletcher and Powell (DFP), the Levenberg-Marquardt (LM) and the LM with Adaptative Momentum (LM-AM).

The Gauss-Newton (GN) algorithm is a technique for solving non-linear least squares problems that incorporates the second derivatives of the loss function or an approximation of those. Unlike the algorithms of first derivatives that can solve a large non-sparse optimization problems, a GN method can only be used when the loss function is given by a quadratic objective function, for instance the expression (10).

The Methods Of The Gn Type Are

generally considered more powerful in terms of accuracy and time than the algorithms that only use the first derivative information. The GN method is based on an expansion of the loss function in the Taylor series. Let M be the number of adjustable parameters (the number of weights w+

I,J). We Define

the M-dimensional vector w that collects in some arbitrary order the weights w+

I,J And W−

i,j. Let a be an input vector on the network. The GN algorithm employs a linear approximation

(21)

where δ is a M-dimensional vector that represents a small correction of the weights. The solution is found by solving the M × M set of equations (called normal equations)

(22)

where G and J are the gradient vector and the Jacobian matrix, respectively. For computing G and J we proceed as follows. Let e(k) be the residual row vector of dimension O for the

(23)

Collecting those residuals, we have a vector E of S × 1 dimensions, with S = KO. Then, the gradient vector of L(·) has M × 1 dimensions and its mth element is

(24)

The Jacobian matrix has dimensions S × M and its (s, m) element is Js,m = ∂Es/∂wm.

(25)

For computing the partial derivatives of (24) and (25) we use the expressions presented in (14). The GN method is a batch type algorithm. We call an epoch of the GN algorithm when all the patterns in the training set are used (Schwenk and Bengio, 2000). At each epoch τ, the weight correction δ is computed, next the weights are updated as follows:

(26)

where α ∈(0, 1] is computed using a line search technique (Press et al., 1992).

N The

canonical GN method this parameter is set to 1. A better strategy is tuning α with less values until some suitable point. For details about how to tune α see Chapter 9 of (Press et al., 1992).

The GN method for solving the problem of minimization using NNs presents several drawbacks. The method requires a good initial solution, that is often not available (Drucker and Le Cun, 1992). Another drawback is that the GN method requires computing the Hes- sian matrix H (H = JTJ) and its inverse, both computations can be expensive. Therefore, the method is expensive in time and in storage.

A Quasi-Newton (QN) method type is a variant of the GN algorithms that uses an approximation of the Hessian matrix ( eH) for solving the normal equations. The general approach behind a QN method is an iterative procedure that consists of starting with a positive and symmetric matrix and updating it in successive steps in such a way that the matrix remains positive definite and symmetric. The update rule always moves in a downhill direction for solving the normal equations and guarantees that eH approximates H. As we already commented so far, the implementation of the second order methods is offline, thus at each epoch the network outputs are computed for the whole of input patterns. We present in Schema 2 a procedure that shows how to compute those model outputs. In the following of this Section we will use this schema as a black box being a part of the GN and Quasi- approximations of the Hessian matrix.

Algorithm 2: Auxiliary schema. Given a RNN the procedure shows how to compute the network outputs for the whole input dataset. The procedure returns a K × N matrix, that has the vector ϱ(k) computed with the input pattern a(k) in its k-row.

Nputs

: {(a(k), b(k)) : k = 1 . . , K} (training dataset), the topology of the RNN Outputs: The neuron activity rate produced by the whole of input patterns: C a

Using (6), (7) And (8);

// see also 3.3.2 for many relevant technicalities

Set The Row K Of C With The Vector Ρ(K);

3.4.1 The Broyden-Fletcher-Goldfarb-Shanno algorithm The Broyden-Fletcher-Goldfarb-Shanno (BFGS) method for the RNN model was introduced in (Likas and Stafylopatis, 2000). The BFGS is an offline algorithm, which at each epoch τ an approximation of the Hessian matrix eH(τ) is computed. The method starts using the identity matrix as the initial Hessian approximation eH(0) = I. The Choleski factorization is used for decomposing a symmetric and positive definite matrix into two triangular matrices.

Choleski factorization is more efficient than alternative methods for solving linear equations, it is about two times faster than the alternative ones. For details about the implementation of this factorization see (Press et al., 1992). The matrix eH(τ) is decomposed using Choleski

(27)

Let c be an auxiliary scalar defined at each epoch as

We Define An Auxiliary Vector V As

v(τ) = c(τ)L(τ)(w(τ) −w(τ−1)).

(30)

The update of the Hessian matrix approximation is given by eH(τ+1) = A(τ)AT(τ).

(32)

In summary, the BFGS method for RNN presented in (Likas and Stafylopatis, 2000) is defined in Algorithm 3.

The Davidon-Fletcher-Powell Algorithm

The Davidon-Fletcher-Powell (DFP) algorithm is another widely used QN method some- times referred as Fletcher-Powell (Press et al., 1992). The algorithm is a slight variation of

(33)

and the vector v is such that solves the linear system, L(τ)v(τ) = c(τ)(G(τ) −G(τ−1)).

(34)

Algorithm 3: Specification of the BFGS algorithm for the RNN model.

Nputs

: {(a(k), b(k)) : k = 1 . . , K} (training dataset), maxIters (max. number of

Τ = 0;

2 Initialize all weights (for instance, randomly); // we get w∗

Ompute L Using Choleski Factorization See (27);

Compute c, v and A using (28), (29) and (30), respectively;

The Matrix A Is Determined By Computing

A(τ) = L(τ) −(G(τ) −G(τ−1))((w(τ) −w(τ−1))TL(τ) −vT(τ))

(36)

and we compute the search direction δ for update the weights solving the expression (32). According empirical results the BFGS performs better than the DFP method (Press et al., 1992).

Although, for some specific benchmark problems the DFP reached better accuracy than DFGS (Likas and Stafylopatis, 2000). The algorithm is summarized in 4.

The Levenberg-Marquardt Algorithm

The Levenberg-Marquardt (LM) algorithm is one of the most standard optimization methods used in the NN area (Press et al., 1992; Ampazis and Perantonis, 2000; Hagan and Menhaj, 1994). The LM is a sort of compromise between an offline version of the GD algorithm and a GN method (Marquardt, 1963; Press et al., 1992). The algorithm was introduced for training RNN in (Basterrech et al., 2011).

Algorithm 4: Specification of the DFS algorithm for the RNN model. The DFS and the BFGS algorithms differ only in details. As a consequence, we introduce the DFS referencing the schema already presented in Algorithm 3.

I,J : I, J = 1, . . . , N} (Network’S Weights)

// Perform the lines 1 until 7 of Algorithm 3.

Ompute A Using (35);

// Perform the lines 14 until 17 of Algorithm 3. At each epoch τ, the approximation of the Hessian matrix is given by,

(37)

where µ(τ) > 0 is called dumping term, I is the identity matrix of dimension M × M, and J is the Jacobian matrix that is computed using (25). The dumping term µ is modified at each epoch. In the case that the prediction error decreases, then the dumping term is

(38)

Otherwise, the dumping value is increased by a factor of β, µ ←µβ.

(39)

So far, the factor for modifying the dumping term was set as β = 10 (Press et al., 1992; Basterrech et al., 2011). The LM algorithm computes the weight correction δ solving the system (32). Then, the update rule for the weights is given by the expression (26). In (Basterrech et al., 2011), this weight update considers only the search direction δ. In other words, the authors set α = 1 in the expression (26). The algorithm can evolve through either of extreme possible situations are (Hagan and Menhaj, 1994; Press et al., 1992): • If the dumping term approaches to zero, the LM basically performs as the Gauss- Newton method.

• Otherwise, when the dumping term is very large, the matrix eH becomes diagonal dominant, so the update rule is similar to the updating expression of gradient descent method using a learning factor of 1/µ.

Concerning the stopping conditions, the method can fail if the Jacobian matrix becomes singular or nearly to singular. Even if this situation is rare in practice, a control of the condition number of J can be useful (Press et al., 1992). Besides, it is necessary to control that the dumping factor satisfies some boundary conditions. It is not recommended to stop after an epoch wherein the training objective error increases. For more technical discussion about the stopping criteria of the LM see (Press et al., 1992). We present the LM procedure in Algorithm 5.

Algorithm 5: Specification of the LM algorithm for RNN.

Nputs

: {(a(k), b(k)) : k = 1 . . , K} (training dataset), maxIters (max. number of iterations), the topology of the RNN, µ (dumping term), β (constant to

Tmp = W∗+ Δ;

// weights w∗are those at τ −1, see also 3.3.2 for technicalities

Evaluate Stopping Conditions;

3.4.4 Levenberg-Marquardt with adaptive momentum training A variation of the LM method applied to NNs was developed in (Ampazis and Perantonis, 2000; Ampazis et al., 1999). This approach was adapted for the case of RNN on learning problems in (Basterrech et al., 2011).

Authors:

Peder EZ Larson 1, 2,* , Jenna ML Bernard1, James A Bankson 3, Nikolaj Bøgh 4, Robert A Bok1, Albert P. Chen 5, Charles H Cunningham 6,7, Jeremy Gordon1, Jan-Bernd Hövener 8, Christoffer Laustsen 4, Dirk Mayer 9,10, Mary A McLean11 12, Franz Schilling13, James Slater1, Jean-Luc Vanderheyden5, 14, Cornelius von Morze 15, Daniel B Vigneron1, 2, Duan Xu1, 2, and the HP 13C

94143, Usa.

Denmark. 5 GE Healthcare, Menlo Park, California, USA. 6 Physical Sciences, Sunnybrook Research Institute, Toronto, Ontario, Canada.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

8 Section Biomedical Imaging, Molecular Imaging North Competence Center (MOIN CC), Medicine, Baltimore, MD, USA. Cambridge, United Kingdom.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

14Jlvmi Consulting Llc, Dousman, Wi, Usa

#See Acknowledgements for a list of all HP 13C MRI Consensus Group Members This work was supported by the ISMRM Hyperpolarized Media MR Study Group, the ISMRM Hyperpolarization Methods & Equipment Study Group, and the Hyperpolarized MRI Technology Resource Center (NIH/NIBIB grant P41EB013598).

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Abstract

MRI with hyperpolarized (HP) 13C agents, also known as HP 13C MRI, can measure processes such as localized metabolism that is altered in numerous cancers, liver, heart, kidney diseases, and more. It has been translated into human studies during the past 10 years, with recent rapid growth in studies largely based on increasing availability of hyperpolarized agent preparation methods suitable for use in humans. This paper aims to capture the current successful practices for HP MRI human studies with [1-13C]pyruvate - by far the most commonly used agent, which sits at a key metabolic junction in glycolysis. The paper is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification. In each area, we identified the key components for a successful study, summarized both published studies and current practices, and discuss evidence gaps, strengths, and limitations. This paper is the output of the “HP 13C MRI Consensus Group” as well as the ISMRM Hyperpolarized Media MR and Hyperpolarized Methods & Equipment study groups. It further aims to provide a comprehensive reference for future consensus building as the field continues to advance human studies with this metabolic imaging modality.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Keywords: Hyperpolarized MRI, metabolic imaging, carbon-13, pyruvate, dissolution dynamic

Introduction

MRI with hyperpolarized 13C agents, also known as hyperpolarized (HP) 13C MRI, has shown great potential as a novel imaging modality, particularly for its ability to probe metabolic processes in real time. The first human studies with HP [1-13C]pyruvate were performed in 2011 in prostate cancer patients (1).

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Since then, there have been over 60 papers published with imaging results of human subjects from 13 different sites, with applications including prostate cancer, brain tumors, breast cancer, kidney cancer, pancreatic cancer, metastatic disease, liver disease, ischemic heart disease, diabetes and cardiomyopathies. The vast majority of these studies used [1-13C]pyruvate (1–63), where [2-13C]pyruvate (64) and 13C-urea (56) have been demonstrated too.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

As clinical HP 13C MRI advances, there is a growing need to build consensus for best practices, which are critical for comparing data across sites, performing multi-site trials,deploying methods to new sites, partnering with vendors, and potentially for obtaining broader regulatory approvals.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

In March 2022, we initiated an effort to build consensus within the HP 13C MRI community with this opportunity in mind, and it was greeted with strong enthusiasm. The “HP 13C MRI Consensus Group”, containing over 55 members from 27 sites, identified the area of greatest need and opportunity for consensus building to be HP [1-13C]pyruvate human

●

Pyruvate is the most mature and widely used HP agent and has the most significant translational evidence emphasizing the potential clinical impact.

●

Clinical trials, particularly multi-site trials, have the strongest need for consensus methods to ensure that data can be combined across sites. This work is a Position Paper for which the goal is to describe current successful practices and study methods for HP [1-13C]pyruvate human studies along with justification to support those practices. This is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification (Fig. 1). The current successful practices and study methods include a literature review of published peer-reviewed journal papers showing human HP [1-13C]pyruvate study data, up to September 2022 (1–63), as well as new unpublished information from surveys of HP 13C study sites. Based on this information, we also highlight the evidence gaps, strengths, and limitations of current practices which are summarized at the end of each section.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Figure 1: Illustration of the HP 13C MRI human study process, including the 4 major areas covered in this paper: Hyperpolarized 13C-pyruvate preparation, MRI system setup and calibration, Acquisition and Reconstruction, and Data Analysis and Quantification.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Figure 2: Anatomical targets of HP [1-13C]pyruvate MRI human studies published up to September 2022.

Hyperpolarized 13C-Pyruvate Preparation

This section covers the processes for creating the HP agent, 13C pyruvate, and will include many aspects and considerations that are needed to safely and effectively prepare doses for metabolic imaging studies in human subjects. These include material, personnel, equipment and facility, fluid path preparation, quality control, and release.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

It is helpful to understand that the specifications of a dose of 13C pyruvate suitable for in vivo MR HP metabolic imaging were shaped in part by early preclinical studies performed by GE HealthCare summarized in Ref. (65). In short, the safety of the two novel drug components, 13C pyruvate and the electron paramagnetic agent (EPA) AH111501, were demonstrated in those studies. The more precise formulation of the dose suitable for human use was then determined from clinical studies (66) that included two Phase 1 clinical trials in young and elderly healthy volunteers without hyperpolarization of the 13C nuclei and another Phase 1/2a dose escalation and imaging feasibility study with HP 13C pyruvate in 31 prostate cancer patients at the With the exception of the first HP 13C imaging clinical trial, which utilized a prototype device in a cleanroom (1), all HP 13C studies performed in humans to date have utilized the SPINlab polarizer (manufactured by GE HealthCare). Consequently all doses of the HP 13C pyruvate delivered by SPINlab have been produced using the “SPINlab Pharmacy Kit” that serves as the container-closure system for the various drug components (13C pyruvic acid and EPA mixture, dissolution medium, and neutralization and dilution medium) during sample polarization, dissolution and quality control (QC) processes. Thus many aspects of the HP sample preparation considerations discussed below are related to the SPINlab instrument and the consumables designed to be used with it (67).

General Considerations

While more than 860 patients or healthy subjects having been injected with HP 13C pyruvate as of January 2022 without reports of any serious adverse events (68), HP 13C pyruvate injection remains an investigational MR contrast agent and can only be administered by those with Investigational New Drug (IND) exemption from the Food and Drug Administration (FDA) in the USA, a Clinical Trial Application (CTA) in Canada, approval from National Research Ethics Committee Services in the UK, or approval from the relevant local regulatory body. Thus, methods and processes involved to produce a dose should have patient safety as the first priority. Since utilizing dissolution dynamic nuclear polarization (dissolution-DNP) for human use is still a relatively new development, there are no existing published regulatory guidelines specifically for this method.

There are two major production styles that determine how various sites approach the agent preparation. In the US, the most common approach is to rely on a sterilizing filter (“Terminal Sterilization”) to ensure sterility of the final product, akin to PET tracer production, where a starting molecule with a radioisotope is processed using various other ingredients to make the final, desired and injectable contrast agent within a necessarily short amount of time (69). For these sites, sterilization of the components and accessories upstream of this filter are not required, although many of them were manufactured and tested following Good Manufacturing Practice (GMP) or Good Laboratory Practice (GLP) requirements. The filling process is usually performed under an ISO 5 laminar flow hood, but a clean room or an isolator is not required.

This approach is typically accompanied by testing the integrity of the sterilizing filter prior to release of the dose for injection. Typically, post release endotoxin and sterility tests are performed using an aliquot reserved from each released dose.

In the UK and EU, the most common approach is to more-closely follow sterile pharmaceutical compounding guidelines (70), where all components and ingredients are required to be sterile or manufactured under GMP guidelines and are assembled and filled within a clean room environment or an isolator system (“Sterile Preparation”). Typically a batch of Pharmacy Kits for HP 13C pyruvate injection are prepared together. The sterility of the final dose is also ensured by batch validation testing, in addition to the sterility of the ingredients and the sterile compounding process. The endotoxin and sterility testing are performed for the process validation but are not performed for each injected dose.

Some institutions fill and assemble the Pharmacy Kit required for a specific study on the same day or the day prior to polarization, dissolution, and patient administration, but others have also demonstrated the feasibility of preparing a batch of kits, keeping them in a -20ºC freezer and using them over a period of a few months.

Beyond the obvious requirements that the process and the facility has to ultimately produce a dose that is safe to inject into a human, regulatory authorities will also focus on the question “Are you in control of your processes?”. To be in control of your process requires an in-depth and broad understanding of all processes involved in pre, post, and during the production process.

Personnel

It is typical and may be required to have licensed personnel involved in the production process depending on local regulations.Typically a pharmacist, radiopharmacist or other similarly qualified person (QP), in charge of the facility where the Pharmacy Kit filling and preparation is taking place, is responsible for the overall process and the release of the injectable dose.

Qualified cleanroom technicians are often involved in the Pharmacy Kit filling under the supervision of the pharmacist or QP. As is required for pharmaceutical compounding or PET tracer production, training requirements and training records for all personnel need to be maintained and available for audit by the FDA or equivalent.

Equipment And Facility

The facility and all equipment need to have standard operating procedures (SOPs) that describe how equipment is used, maintained, and calibrated to comply with relevant legislation. Currently, almost all the filling of the Pharmacy Kit takes place within a compounding laminar flow hood or isolator (typically ISO 5). At some sites, the filling is conducted within a cleanroom, while at others, it is conducted in a dedicated non-cleanroom space, reflecting differences in cleanroom approach and specifications between regulators worldwide (71). Some equipment or facilities, such as the compounding hood or cleanroom, may require external certified laboratories for testing.

Material Handling

Material handling guidelines (69,70) require SOPs detailing a system to track all of the materials involved in the HP production process for a particular patient dose, similar to current good manufacturing practice (cGMP) requirements for material handling for drug compounding. This includes acceptance standards, storage conditions, amount used in the patient dose for each ingredient and materials used in the assembly of the fluid path and Pharmacy Kit. Currently some users choose to open and inspect and sometimes modify the Pharmacy Kits upon arrival, but some users keep them in the sealed packaging until they are required for dose preparation.

Pharmacy Kit Filling And Assembling

As required by an IND or its equivalent, the preparation of the doses of HP 13C agent are detailed in the Chemistry, Manufacturing, and Control (CMC) section of an applicable regulatory submission; an example of this has been made available (72). It describes the processes of filling the Pharmacy Kit with the different components that make up the final drug product, and of assembling the final kit for either storage or immediate use in the polarizer. Special attention should be given to the laser welding process in order to satisfy installation qualification (IQ) and operational qualification (OQ). Typically, the final developed process is validated by process qualification (PQ) runs, during which 3 or more Pharmacy Kits are filled and used and the final HP 13C products are tested for endotoxin and sterility and to confirm that they meet the dose specifications for injections (usually including pyruvate concentration, residual EPA concentration, pH, liquid state polarization level and dose temperature). The data from 3 consecutive PQ runs are submitted as part of the IND submission (or its equivalent), and are often also reviewed by the Institutional Review Board (IRB) where the studies are conducted.

Quality Control And Dose Release

The quality control (QC) and dose release can be separated into two aspects: one is the QC and release of the filled Pharmacy Kit, and second is the QC and release of the HP 13C agent for injection, after polarization and dissolution. For institutions filling a batch of kits and storing them to use over a period of time, typically the batch can be released based on initial validation, environmental monitoring data from the day of kit production, and if filters are used during preparation of any of the components, filter integrity testing. But in some cases one or more kits are used for validation before the batch of kits are released for future use. For institutions that fill only the kits required for specific studies shortly before the experiment, the filled kits often do not go through separate release tests before they are used.

The quality control of the HP 13C pyruvate solution post dissolution is primarily performed to ensure that the agent meets the dose specifications (Table 1) before it is administered to the subject. These specifications target both safety (pH, residual EPA, temperature) and efficacy (pyruvate concentration, polarization, volume). Typically, the pyruvate concentration, residual EPA concentration, pH, dose temperature, dose volume, and liquid state polarization are measured by the QC accessory associated with the SPINlab polarizer. Some users perform a secondary measurement for one of the parameters, such as pH, using a different instrument or pH paper. For sites that do not go through a separate release testing process for batch filled kits, the integrity of the sterilization assurance filter, a part of the Pharmacy Kit, is typically tested as a part of the dose release. It is also common for these users to preserve an aliquot of the final HP 13C pyruvate solution for post-release endotoxin and sterility testing. This testing cannot be completed fast enough to test an individual dose prior to injection, but this is why other processes such as PQ runs and validation testing are done to minimize the chance a subject could be injected with a contaminated dose.

The Final Dose Release And Injection

should be done under the supervision of a licensed professional, based on local regulations.

Some Key Challenges

Many of the challenges associated with HP 13C pyruvate preparation can be attributed to the conditions required for the dissolution-DNP method of high magnetic field (~3-7 T) and very low temperature (~1 K) during polarization, with pressurized and superheated water necessary for the rapid dissolution event. These extreme conditions are quite challenging for the design of the container-closure and fluid path system. In particular, the cryogenic temperature in the polarizer requires special attention to any moisture or ambient (moist) air introduced into that portion of the fluid path, which can form an ice block at ~1 K. This ice can lead to flow restriction during the dissolution event and reduce the strength of the laser welded bond between the cryovial and its cap. This can ultimately produce failures in the dissolution step, including variations in final pyruvate concentration and pH that may fail to meet QC release criteria as well as fluid path ruptures that provide no available dose and result in polarizer down-time.

The polarization of the HP 13C pyruvate sample decays quickly over the span of a few minutes after dissolution, and thus the process of dissolution, QC for release, and injection should be completed as fast as possible to preserve the high polarization level achieved. Any delays in the preparation process, such as transportation time or equipment malfunction, can significantly reduce the final polarization and result in lower quality imaging data.

Current Practices

A summary of data collected from all sites performing clinical trials with HP 13C-pyruvate is shown in Fig. 3 and Table 1, including the specification of the final dose and how the quality control and release of the final dose are performed. There is a split in the Production Style, described in the General Considerations section above, with 8/13 sites using Sterile Preparation versus 5/13 using Terminal Sterilization. While many of the dose specifications show notable differences in acceptable ranges, all of these variations listed in tables have been successfully and safely been used to perform HP 13C pyruvate studies in humans. Their differences depend on the institutions’ preferences, resources and their particular regulatory situation. There is high similarity in pyruvate ranges, temperature ranges, EPA limits, and volume limits. There is modest variability in pH ranges and large variability in the endotoxin test limit. There is a 3-fold difference in acceptable polarization levels, which are measured to ensure a futile dose is not injected since the polarization is directly proportional to SNR. This reflects the decision by several sites to believe that useful data can be still be obtained with suboptimal polarizations.

Figure 3: Hyperpolarized agent preparation methods reported by sites currently performing HP

In House

Table 1: HP 13C-pyruvate preparation parameters, methods, and dose specifications used for quality control testing and release as well as validation. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. The parameters used for product release are noted in bold text, otherwise these parameters are measured for batch validation or other QC measurements. The endotoxin and sterility testing are performed during process validation of the batch and/or post-injection, and largely depends on the agent production approach.

Summary

The overall safety record of HP 13C-pyruvate has been very strong, and the SPINlab hyperpolarizer has proven to provide high polarizations at human sized doses while meeting numerous QC and release criteria. A weakness remains the failure modes of the SPINlab Phamacy Kits (e.g. ice blocks, path ruptures), which are placed under extreme requirements particularly during dissolution. The preparation process still requires a high degree of expertise.

Therefore, there is a significant need to improve the reliability, robustness, and ease of operation for generating HP 13C-pyruvate doses for human studies. Furthermore, there is a divide between manufacturing and sterile compounding style preparation as well as other site-specific practices, resulting in variations in SOPs and justification required to relevant regulatory bodies. There have also been no comparisons between these approaches. It is also unclear what release criteria and QC parameters are truly required to ensure patient safety.

However, all of the reported methods are acceptable and approved by the appropriate regulatory authorities, and have led to the rapid expansion of successful human studies in recent years.

Mri System Setup And Calibrations

This section covers the MRI system setup, including the imaging system, RF coils, phantoms, and prescan calibration methods.

Imaging System

The main prerequisite for a given MRI scanner to be capable of supporting studies with HP 13C is its “broadband” capability to transmit and receive radiofrequency (RF) signal at the frequency of 13C, which is around 4 times lower than 1H. This does not come as a default on clinical MR devices. The transmit power of the broadband amplifier should also be sufficient to support the intended flip angle and RF pulse shape with the employed transmission RF coil(s) for 13C. Most studies to date use relatively low flip angles (< 90 degrees) for HP 13C in order to preserve polarization for time-resolved imaging. The capability to receive 13C signal on multiple channels is also desirable to increase SNR, as discussed further in the “RF coils” section.

The choice of magnetic field strength is primarily dependent on the metabolites’ frequency separation due to chemical shift dispersion and 1H imaging. High field strengths do not enhance hyperpolarized 13C signal as they do for 1H because the signal strength in a HP experiment relies on manipulating the population of quantum energy states outside of the MRI scanner.

However, the injected HP 13C-pyruvate and its metabolic products have greater frequency separation at higher fields, and it may thus be easier to separate and quantify these resonances at higher fields. This comes at the cost of a reduction in the achievable T2* and often reduced T1. As the initial polarization is independent of the imaging field strength it has been proposed that the increased T2* at 1.5T can potentially be exploited to increase SNR by adapting the acquisition bandwidth or reduce off-resonance imaging effects in cases when the decay of the transverse magnetization is dominated by T2* (73). In practice, 3T has been used in all published human 13C-pyruvate studies surveyed (Supporting Table S1), and comprises the majority of scanners currently in use for human studies (Table 3). A field strength of 3T is well-suited for 1H MRI anatomical reference and correlative imaging.

Stronger and more rapidly slewing magnetic field gradients support more rapid spatial encoding, particularly for metabolite-specific single-shot imaging using echo-planar imaging (EPI) or spiral imaging (See “Acquisition and Reconstruction”). Although the spatial resolution acquired for HP 13C imaging is typically much coarser than for 1H MRI, the factor of ~4 in gyromagnetic ratio leads to the same reduction factor in performance of the gradient system, so 13C experiments are potentially more limited by gradient hardware performance. To date, all human studies have used the commercially-available integrated gradient systems provided in clinical MRI scanners.

Optimization of scanner design has understandably focused on minimization of artifacts in 1H MRI, where devices such as room lights, the gradient amplifiers, and the motors driving the patient bed are checked to ensure that they do not produce RF interference at the 1H frequency, but artifacts may arise at other frequencies. Eddy current compensation is also not always appropriately adjusted for nuclei at other frequencies (74). In order to optimize for 13C, many sites have performed checks on phantoms for RF interference, gradient artifacts, and eddy currents (74), including the use of post-hoc gradient impulse response function characterisation and correction, and some vendors have fixed these issues as well.

Rf Coils

For HP 13C imaging studies in humans, RF coils for both 1H and 13C nuclei are needed, with 1H MRI providing an anatomical reference for registration and optional additional multiparametric MRI readouts. At the Larmor frequency of 13C nuclei, the relative contributions from coil noise compared to sample noise increase compared to 1H (73,75), although sample noise still is likely the dominant contributor for human-sized coils at 32.1MHz - the resonance frequency of 13C nuclei at 3T.

The key requirement for human 13C-pyruvate RF coils are that the coil geometry and sensitive volume must cover the volume of interest in the subject. Table 2 and Figure 4 shows coil configurations that have been used and optimized for applications in different anatomic regions.

Volume resonators are most commonly used for transmit, as they surround the subject to

Provide B1 Transmit Across The Fov (B1

+). While 1H relies on a large birdcage (“body”) coil built into the scanner, 13C transmit coils must be placed inside the bore. This takes up valuable space within the magnet, and also has led to the use of designs with relatively inhomogeneous

B1

+. Many human studies have used Helmholz pair resonators for transmit, including the “clamshell coil”, which has a notably inhomogeneous B1

+ Profile But Has Been Used Because Of

relatively easy integration into the scanner bore. B1

+ Variation Results In Variations In The Flip

angles that control the use of the hyperpolarized magnetization and creates errors in common HP metrics (9,76). The exception are head coils, where birdcage designs with highly

Homogeneous B1

+ can be placed around the head while easily fitting inside the bore. As with 1H MRI, higher SNR can typically be achieved by smaller receive coil elements, such as surface coils or phased arrays, and the majority of 13C receive coils used have layouts similar to 1H phased arrays.

RF coil quality control is important to ensure proper functioning of the coils to provide consistent imaging quality, especially with limited natural abundance 13C signal in vivo. It typically involves 1) a physical integrity check of the coil cables and connectors and 2) phantom SNR tests to check the coil’s performance and to monitor it over time (see Phantoms below). An useful reference for RF coil quality control is outlined in the MRI accreditation program of the American College of Radiology (77) and can be adapted for 13C coils.

Notably, configurations for brain and prostate studies used dual-tuned 1H/13C coil designs, which greatly simplify workflow and registration of 1H and 13C images, as no switching of coils is needed.

(1)

Table 2: RF coil configurations reported for human HP [1-13C]pyruvate studies.

Tx = Transmit

coil, RX = receive coil. The commonly used “clamshell” TX coil is a Helmholz pair design. For 1H RF configurations, all used the Body coil for TX unless otherwise noted, and “repositioned” indicates the 13C coil was removed for 1H imaging. One representative reference is listed for each configuration. The RF coil configurations reported in the reviewed papers are shown in Supporting Table S1.

Figure 4: Examples of RF coil configurations used for human HP [1-13C]pyruvate brain studies. (A,B) 13C Clamshell TX (Helmholz pair) and 2× 4-channel paddle RX arrays. (C) 13C Birdcage volume TX and 32-channel RX array (RX array slides into TX coil). (D) 13C Birdcage volume TX and 24-channel RX array, combined with a 1H 8-channel RX array. Image reproduced with permission from Ref (16).

Phantoms

Since hyperpolarized magnetization is non-renewable, phantoms containing 13C nuclei are important to: 1) test the multi-nuclear capabilities of the imaging system, including all parts of the signal excitation and receive chain; 2) perform calibration measurements before a scan with hyperpolarized nuclei; and 3) perform necessary pre-scan adjustments (see “Prescan Calibration” section). The phantoms currently in use are listed in Table 3. Their composition must provide sufficient 13C signal, with additional considerations of conductivity, stability, chemical shift(s) present, potential for dynamic imaging, and cost. The phantom geometries are typically either compact, in order to be used alongside the subject during a HP scan, or large enough to mimic the inner volume of a RF coil for system testing.

One popular compact design contains enriched 13C-urea at high concentration, typically 8 M, which provides a single resonance, placed inside a small container ~1 mL. The most common recipe mixes 13C-urea in a 90% water/10% glycerol solution, with glycerol used to increase the urea solubility and doping with a Gd-based contrast agent to shorten T1 which increases the potential SNR per unit time. For example, when Dotarem is added at a 3:1000 volume ratio the 13C-urea T1 is around 500 ms and T2 is around 100 ms. However, when testing pulse sequences influenced by T1 and T2, doping should be used carefully. This phantom is suitable for frequency calibration, transmit gain calibration, sequence testing, and as a fiducial marker when placed next to a patient. However, enriched 13C-urea has a relatively high cost compared to natural abundance compounds.

For larger volumes (>100 ml), the phantoms most often used contain undiluted ethylene glycol, glycerol, or dimethyl silicone. These compounds have sufficiently high carbon concentrations to provide sufficient 13C signal even with the 1.1% natural abundance of 13C. These larger phantoms matching the inner volume of an RF coil are useful for coil testing, including transmit

+) And Receive (B1

-) coil profile mapping, as well as to mimic acquisitions using in vivo FOV requirements. In this case, size and conductivity should match the expected subject size in order to mimic coil loading and get a realistic estimation of B1+. Large-volume natural abundance urea phantoms have also been used by some sites, but suffer from higher conductivity compared to biological tissues. Typically, it is easier to increase the conductivity and hence coil loading of the non-conductive phantom by adding NaCl to match physiological loading (16,78).

Dynamic phantoms that aim to mimic metabolite kinetics have also been developed (79–81), and have the potential to more closely mimic the HP experiment, but so far these are not widely used.

Prescan Calibration

Prior to performing an MRI acquisition, the so-called prescan procedure is used to set the shim parameters to maximize B0 homogeneity over the field of view (FOV) or a specific region of interest (ROI), the scanner center frequency (CF), the RF transmit gain, and the receiver gain.

While this calibration procedure is usually automated for 1H, the lack of sufficient natural abundance 13C signal prevents use of automated methods. (Although natural abundance 13C lipid signal has been detected, there are so far no reports on using this signal for prescan.) Table 3 shows current practices across sites.

Maximizing B0 homogeneity is independent of the nucleus and is therefore performed prior to 13C imaging using the 1H water signal and existing shimming tools, such as by a standard automated process (“Auto Shimming”) or using high order shimming routines. Similarly, the 13C CF can be calculated from the 1H CF using a predetermined scaling factor that depends on the target chemical shift (82). Another common approach used is to have a small, high-concentration 13C phantom, e.g. 8M 13C-urea, integrated in the RF coil or placed next to the scan subject (1). The reference frequency can also be based on real-time measurements after the HP injection but prior to imaging (83). Both the CF and B0 shimming are critical when using spectrally-selective RF pulses, as inmetabolite-specific imaging methods, where the desired excitation bandwidths are typically very narrow and frequency offsets can lead to a failure mode that is only apparent after injection.

The calibration of the RF transmit power is typically performed on a small, high-concentration 13C phantom placed near the region of interest during the scan or on a large 13C phantom of similar size and coil loading as the subject, prior to the subject scan. Reference power is often done by sweeping the power in a pulse-acquire sequence (53,62), or the Bloch-Siegert method (52,84). When using a small phantom, the location of the phantom, B1

+ Inhomogeneity As Well

as any shielding effects, e.g., when the phantom is integrated into a coil (1), may degrade the accuracy. Other methods include real-time Bloch-Siegert method measurements after the HP injection (83), and using the stronger natural abundance 23Na signal that is close enough to the 13C resonance frequency to be detected by 13C coils (82).

The receiver gain is predetermined, either systematically based on independent phantom measurements and assuming the dose and polarization of the HP compound is known prior to injection, or based on past HP imaging studies.

Power [Kw]

Phantom(s) - during study Phantom(s) - before study 13C Frequency

8

13C-bicarbonate doped with dimethyl silicone, various

Power [Kw]

Phantom(s) - during study Phantom(s) - before study 13C Frequency

Maximum Values

Table 3: Summary of the imaging systems, phantoms, and prescan procedures used at sites currently performing HP 13C-pyruvate human studies. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. *Previously performed studies with a Siemens 3T Tim Trio. The imaging systems, phantoms, and prescan procedures reported in the reviewed papers are shown in Supporting Table S1.

Summary

Commercially available 3T MRI systems are by far the most commonly used for human HP 13C-pyruvate studies, although a systematic investigation of the impact of B0 has only recently been investigated (73). The multi-nuclear RF transmit and receive chain has proven sufficient for current acquisition strategies, although many sites have observed artifacts due to RF interference, gradient interference, and residual eddy currents when operating at the 13C frequency. A variety of 13C RF coils, tailored for numerous anatomical targets, have been successfully demonstrated, with the main limitation that most transmit coils take up a lot of additional space inside the bore and provide relatively inhomogeneous B1

+ Profiles. The

phantoms used have converged into generally 2 categories - small phantoms containing 13C-enriched compounds that can be used during the study and human-sized phantoms containing compounds with high carbon concentrations but without 13C enrichment that are used to test and calibrate the coils. There are no standardized compositions or geometry, and dynamic phantoms that recapitulate in vivo kinetics would be desirable but are still an emerging area. Prescan calibration procedures were not well defined in most publications, so we surveyed individual sites to determine current practices. Calibration procedures for the B0 field (13C CF and shimming) for most sites take advantage of 1H signal and methods, while methods

For Calibration Of B1

+ is more variable across sites, likely a reflection of remaining challenges in how to perform this calibration. Standardization of both phantoms and calibration procedures would synergistically improve the robustness and reproducibility of HP 13C studies.

Acquisition And Reconstruction

Data acquisition strategies in human HP [1-13C]pyruvate MRI studies must account for multiple chemical shifts, efficiently utilize the non-renewable HP magnetization, and acquire data quickly relative to metabolism and relaxation decay processes. These studies require spectral encoding to separate metabolites, necessitating pulse sequences that efficiently encode up to 5D data (3 spatial + 1 spectral + 1 temporal dimension). RF pulses must efficiently sample without immediately saturating the non-renewable HP magnetization, and sequences must acquire data quickly and be robust to both experimental and physiologic variation (e.g. B1

+ Inhomogeneity,

variation in perfusion) to ensure reproducibility and minimize scan-to-scan variability. This section covers current successful practices for data acquisition in human [1-13C]pyruvate studies, and accompanying 1H imaging, from different anatomic regions, including scan parameters and image reconstruction.

Acquisition And Reconstruction Methods

The acquisition methods used in human [1-13C]pyruvate studies can be classified into 3 categories: 1) MR spectroscopy or MR spectroscopic imaging (“MRS/I”), 2) chemical shift encoding methods, and 3) metabolite-specific imaging (Fig. 5).

Mrs/I Methods Specifically

resolve a spectrum that can be analyzed to extract expected as well as unexpected resonances, making this approach very robust. It was used in many initial studies (1).

Chemical Shift

encoding methods, most commonly the Iterative Decomposition of water and fat with Echo Asymmetry and Least-squares estimation (IDEAL) method, use imaging sequences acquired with multiple TEs and rely on a model-based separation of expected chemical shifts (85).

Metabolite-specific imaging methods use specialized RF pulses that are spatially and spectrally selective to excite individual metabolites which are then typically imaged with fast k-space trajectories such as echo planar imaging (EPI) or spirals (86).

Their Application To Different

organ systems is described below. The image reconstruction methods used in human [1-13C]pyruvate studies have typically been conventional methods (e.g. FFT, non-uniform FFT, or equivalent). The incorporation of accelerated imaging and advanced reconstruction methods including parallel imaging (4,57,87) and compressed sensing (7) has also been applied in human studies for improved spatial resolution, temporal resolution and coverage, but have the potential for additional artifacts as well as SNR losses due to ill-conditioning of the reconstruction (e.g. g-factor).

The Majority Of

published studies do not use accelerated imaging indicating the resolution and coverage achievable without acceleration is currently adequate for successful data collection. Performing coil combination, even with fully sampled data has also been shown to have specific challenges for HP human images: using naive sum-of-squares methods suffer from high noise amplification in the relatively low SNR regime of HP [1-13C]pyruvate (compared to 1H), motivating several HP 13C-specific methods that include data-driven coil sensitivity estimation which have shown obvious improvements over sum-of-squares (11).

More recently denoising techniques have been applied as post-processing of human HP data(41,42,44). The techniques applied are based on spatial-temporal singular value decomposition for unsupervised estimation of signal and noise components. They have shown improvements in apparent SNR in the brain and liver, while care must be taken to choose parameters such as the rank threshold to avoid oversmoothing and overfitting to the estimated signal components.

Prostate Studies

Prostate cancer was the first human application of HP [1-13C]pyruvate (1), and data was acquired with MRS/I methods: 1D dynamic MRS, single-slice 2D dynamic echo-planar spectroscopic imaging (EPSI), and single time point 3D EPSI. Advances in imaging strategies led to the development and application of new acquisition schemes, including undersampled 3D EPSI with compressed-sensing (7), model-based chemical shift encoding methods that use a priori information (47,59), and metabolite-specific EPI (10), all of which can provide volumetric whole-organ coverage and dynamic acquisitions.

The pyruvate bolus arrival in the prostate can vary by ± 10 s between patients, necessitating dynamic imaging to reliably and consistently capture the pyruvate bolus (18). For this reason, all currently ongoing studies acquire dynamic data. While MRS/I, chemical shift encoding, and metabolite-specific imaging can all achieve dynamic imaging, chemical shift encoding and metabolite-specific imaging provide greater dynamic and volumetric coverage (85). For scan prescriptions, the FOV is designed to provide full prostate coverage and typically to match the orientation of the anatomic imaging used for registration. Flip angles used in current studies are constant through time, as quantification with a variable-through-time flip scheme is highly sensitive to bolus timing (8) and errors in the RF transmit (B1 +) field (76).

Heart Studies

Data acquisition methods for 13C imaging in the heart must be designed to meet the demands of significant cardiac motion and blood flow. To cope with the periodic cardiac motion, most human heart studies to date used gating to the diastolic window, the longest cardiac cycle interval, which has reduced motion (2,22,28,30,35,36,38,45,52). The duration of the diastolic window limits the available data sampling time, making cardiac acquisitions the most time-constrained of the HP 13C MRI applications. The most common acquisition approach is metabolite-specific imaging with spiral k-space trajectories (2). Their single-shot imaging capability makes these methods particularly robust to motion effects. Furthermore, spiral k-space trajectories provide rapid k-space coverage and relatively benign flow and motion artifacts. The majority of studies have used 2D multi-slice acquisitions, but 3D encoding has also been used successfully (35).

Brain Studies

For HP 13C MRI of the human brain, the majority of studies have also used 2D (slice selective) acquisitions (10–12,14,16,28,33,40,41,44,51,53,60), with a trend toward volumetric coverage using 2D multi-slice metabolite-specific imaging. 3D metabolite-specific imaging of the whole brain, with phase encoding of the slice direction (34,57), has been shown to provide similar SNR efficiency (88) compared with multislice imaging. A number of studies have employed MRS/I (5,6,29,31–33,50,55) resulting in a spectrum from each voxel, which has the advantage of not requiring a priori information about which peaks to encode. This was important in early brain studies when it was not known which peaks would be detectable. Chemical shift encoding, using a set of images with different echo times and an iterative reconstruction of the individual resonances (i.e. the IDEAL approach (85)), has also been used (12,49,54), with the drawback that coverage in the slice direction was limited due to the time required to acquire multiple echo time images.

Abdomen And Breast Studies

The fundamental approaches to data acquisition and reconstruction in the abdomen and breast are largely similar to the aforementioned applications, but demand attention to particular challenges associated with these anatomic regions, especially relating to respiratory motion.

Although it has been shown that a basic 2D MRSI approach based on phase encoding and FID readout can be successfully applied for HP 13C imaging in breast (15) and kidney (13), major advantages in terms of spatiotemporal resolution and coverage have been realized using tailored approaches based on metabolite-specific imaging (43,62) and chemical shift encoding (43), which have facilitated multi-slice or 3D dynamic acquisitions over large FOVs in the abdomen (4,37,46).

The significant respiratory motion encountered in these regions can directly blur 13C images, and has further favored these rapid acquisition strategies. Motion also degrades B0 homogeneity, which can shift frequency-selective excitation profiles and introduce artifacts into rapid imaging readouts. This makes accurate determination of the acquisition center frequency and shimming essential in these regions which often cover large FOVs. (See “Prescan Calibration” section for more information). In some studies, breath-holding was used to minimize motion effects and enforce frame-to-frame data consistency (42). A pragmatic and reasonably effective approach for dealing with respiratory motion during 13C data acquisition is an initial breath-hold (as long as can be tolerated), followed by free-breathing (46,62).

1H Imaging

Collection of 1H imaging data is essential both for prescribing the 13C acquisition and for interpretation of the resulting 13C data. Multi-planar 1H scouts are acquired prior to 13C acquisition to enable graphical prescription of the 13C imaging region. All human HP 13C-pyruvate imaging studies acquire conventional MRI scans (e.g. T1- and T2-weighted volumes) for anatomic reference, aiming to cover at least the full 13C FOV. Acquiring these anatomic scans as close as possible to the time of 13C imaging (immediately before or after) minimizes potential misregistration between the data sets. Depending on the application, other advanced 1H sequences are also acquired (e.g. diffusion-weighted imaging for cancer imaging).

When contrast-enhanced data is acquired, it is done after 13C imaging, as paramagnetic contrast agents will accelerate 13C relaxation.

Reported Study Parameters

Figures 5 and 6, and Supporting Table S2 shows the reported acquisition study parameters for human HP [1-13C]pyruvate studies published as of September 2022. Figure 5 shows a mixture of MRS/I, metabolite-specific imaging, and chemical shift encoding methods have been successfully used, where spectroscopy-based methods have become less prevalent in recent studies. Figure 6 shows the acquisition timing, including the important start time and interval/temporal resolution, is quite variable across studies.

Figure 5: Acquisition methods used in published HP [1-13C]pyruvate human studies published up to September 2022, classified into: MR spectroscopy and spectroscopy imaging (MRS/I); chemical shift encoding methods, such as IDEAL, that use multiple TEs and model-based reconstructions; and metabolite-specific imaging methods that use spectrally-selective excitation to image a single resonance at a time.

Figure 6: Temporal acquisition characteristics reported in HP [1-13C]pyruvate human studies published up to September 2022. (a) Reported referencing of acquisition start times.

(B)

Acquisition start times reported when using dynamic imaging and when timing was reported relative to the end of the injection. (c) Temporal resolutions. “Not Applicable” indicates dynamic imaging was not used.

Summary

Three general categories of acquisition strategies have been used successfully for human HP 13C-pyruvate studies: MRS/I, model-based chemical shift encoding (e.g. IDEAL) methods, and metabolite-specific imaging methods. These have enabled successful studies in the prostate, heart, brain, abdomen, and breast. Recent studies increasingly have used the imaging-based strategies of metabolite-specific imaging and chemical shift encoding which are the fastest methods, although a heads-to–head comparison between techniques has not been performed.

Metabolite-specific imaging is quite popular because of its speed and compatibility with single-shot imaging, but is sensitive to B0 field variations and thus requires careful calibrations. Nearly all studies surveyed acquired data dynamically, allowing measurement of the bolus and metabolite kinetics. The exact timings and associated flip angles vary quite widely across reported studies, with no consensus yet as to how to choose these parameters. Image reconstruction is typically done directly using Fourier Transform methods, and accelerated imaging strategies are uncommon.

Data Analysis And Quantification

This section covers the analysis of data from human HP [1-13C]pyruvate studies, including modeling and metrics, visualization, as well as considerations for how to store data and metadata. Depending on study design, the analysis may need to give quantitative or semi-quantitative output reflecting a biological process or may just reflect a contrast between different regions of interest for quantitative evaluation.

Metrics

Figure 7: HP [1-13C]pyruvate raw data (A) have typically been quantified using four categories of metrics depending on the acquisition. Data acquired as a single time point are often quantified using normalized metabolite images or metabolite ratios (B). Dynamic data can be quantified using normalized metabolite images or metabolite ratios (B), or with metabolite timings such as time-to-peak (TTP) or pharmacokinetic (PK) models (C). The latter two require the data to be time-resolved. [1-13C]alanine and 13C-bicarbonate are analyzed similarly to [1-13C]lactate but omitted here for display.

Metabolite images are commonly used as summary metrics for HP MRI data, often including some form of normalization as well as summed over time as an area under the time curve (AUC) (17). These are analogous to the visual evaluation that is most used for routine clinical work (89,90). In these metabolite images, we expect that the [1-13C]pyruvate AUC signal is predominantly weighted towards perfusion and uptake, while [1-13C]lactate, [1-13C]alanine and 13C-bicarbonate AUCs represent metabolic conversion. The strength of this approach lies in its simplicity and relatively few underlying assumptions. Limitations to the use of single-metabolite images or AUCs include sensitivity to inhomogeneous coil profiles (57,87,91), the acquisition strategy and acquisition parameters, pyruvate polarization and concentration level, and signal relaxation rates (92). Further, the reader must be careful to interpret all the images in conjunction to better understand the underlying biology; for example, increased [1-13C]lactate in the presence of decreased [1-13C]pyruvate delivery can have a very different meaning compared to increased [1-13C]lactate with increased [1-13C]pyruvate delivery.

In an attempt to address variations in coil sensitivity, polarization level, and pyruvate delivery, AUC images are often computed by normalizing to a specified parameter, such as the maximum pyruvate or average lactate signals, or presented as a ratio such as lactate/pyruvate or divided by “total Carbon” - the sum total of HP 13C signal observed across all metabolites. The AUC ratios between metabolites and pyruvate are proportional to the corresponding forward kinetic rates (81,93), but are not directly comparable to rate constants when magnetization loss rates (e.g. relaxation and losses due to signal excitation) differ between studies. Similarly, the ratios between the produced metabolites (e.g. bicarbonate/lactate) can reflect the balance between downstream metabolic pathways (12,55). Care must be taken to consider how AUC images are calculated and normalized before comparing values between studies.

To further quantify the interpretation, pharmacokinetic (PK) modeling approaches were developed to compute the apparent kinetics of pyruvate-to-metabolite exchange (92,94–99). These yield semi-quantitative to quantitative apparent rate constants, given in s-1. Some models require a vascular input function, while others avoid this requirement (95). PK models can explicitly account for acquisition-specific details such as excitation angle and repetition time, and thus may reduce the effects of these details on quantification. An input-less model, provided in the Hyperpolarized-MRI-Toolbox (https://github.com/LarsonLab/hyperpolarized-mri-toolbox) (100) and thus frequently employed for human data, has been shown to fit well and robustly to prostate and brain data (8,20). PK models are quantitative in nature, arguably provide more relevant biological information (8,20), and appear to be reproducible across sites (51). However, rate constants derived from PK models are still apparent rates, and likely do not reflect a single biological characteristic.

Some additional considerations include whether complex or magnitude data is used, as the noise behaviors will impact the analysis differently. Additionally, cut-off thresholds or other criteria may be used to identify and avoid voxels with insufficient SNR before analysis to improve robustness (20,41).

Regardless of the analysis approach, the underlying biology is not always clearly represented by the data; instead, the metrics may be influenced by perfusion, barrier permeability, intercellular shuttles, enzyme activities, co-substrate concentrations, or combinations thereof, depending on the organ and disease of interest (19,43,94,101–103). This may be addressed by incorporating complementary information. As an example, HP 13C pyruvate data is influenced by perfusion, and thus addition of perfusion MRI could be important for interpretation (98,104,105).

All the methods outlined above have been explored in clinical studies, described in Supporting Table 3 and summarized in Figure 8. As of September 2022, approximately 52% of studies involving human subjects report rate constants derived from a PK model with a few different models reported. A nearly equal fraction (51%) of the studies report AUC ratio values.

Approximately 66% of these studies report metabolite-specific images or AUC values. About 40% report SNR values; this metric is particularly frequent in manuscripts that describe technical developments for clinical HP MRI. Approximately 16% of these studies summarize model-free metrics, and 10% report measurements from a single timepoint. Most studies report a combination of quantities.

Figure 8: Reported metrics used for analysis in HP [1-13C]pyruvate human studies published up to September 2022.

Visualization

A wide variety of approaches have been used for visualizing data from human HP 13C-MRI studies. The challenges and practical considerations are: 1) choosing the appropriate metrics to display, 2) how to encode the parameters (e.g. the colormap), and 3) choosing how to provide anatomical context and other multi-parametric data. The choice of visualization also depends on the goal which could be for diagnostic interpretation, but also quality control, reproducibility among readers and publication.

Metrics

The choice of HP 13C metrics is described in detail above. At this stage in HP 13C development where there is no standardized metric, often a combination of metabolite images and ratios or PK model parameters are shown.

Parameter Encoding

The mapping function chosen should provide an adequate, often quantitative, impression of the parameter mapped. There is a consensus in the visualization field that perceptually uniform maps are best suited to visualize continuous parameters, like the greyscale typically used by radiologists as well as other monochrome (black to blue) and color ranges (fire-type, rainbow-type) (106,107). Multi-color heatmaps have been the most frequently employed method for HP 13C data, while greyscale has infrequently been used but it ensures there is no coloring-based bias as well as facilitating later reuse (Fig. 9a). Among the color schemes employed in the clinical HP 13C literature, fire-type scheme seems to be the most common [similar to “Plasma” or “Inferno” in matplotlib.org]. Next most commonly employed is the rainbow-type scheme [similar to “Rainbow” in matplotlib.org].

Anatomical Context

HP MRI faces the challenge that it does not necessarily depict the anatomical features, similar to PET, and thus requires an anatomical reference. Most often, a grayscale anatomical image is overlaid with a HP colormap (Fig. 9c,d). This approach is very intuitive, but can skew perception as the grey-scale anatomical reference may affect the brightness of the HP data (e.g. signal in the skull). This bias does not occur when showing adjacent maps (Fig. 9a, b). Here, anatomical outlines may help to provide reference (Fig. 9b).

Related Journal Articles & DOI Links

Selected peer-reviewed publications relevant to 12 Lead ECG Acquisition. Click the DOI to access the full paper (may require institutional access).

Why Choose Us?

Bangalore guidance for robotics, Spectre and autonomous systems projects.

Spectre & Simulation

Gazebo, cloud twin and Webots worlds with navigation, SLAM and control stacks.

Control & Planning

Compliance, deep learning control, path planning and behavior trees.

Hardware Bring-up

Motors, sensors, ESP32/STM32 firmware and HIL validation paths.

Report & Viva

University-format documentation, PPT and viva preparation.

FAQ

Spectre, Gazebo, NVIDIA cloud twin, MATLAB/Simulink, Webots, Blynk / ThingSpeak, plus Arduino/STM32/ESP32, cameras, LiDAR and motor drivers.
Yes — simulation packages, hardware guidance, report, PPT and viva Q&A.