Earning Research
Elena Sorina Lupu, Patrick Spieler, Khurram Javed, Kris De Asis, John
Artin, Martha Steenstrup, Joseph Modayil
Keywords: Reinforcement Learning, Robotics, Embodied Intelligence, Research Platform
Summary
Reinforcement learning (RL) research has demonstrated success in both physical and sim- ulated domains; however, the predominant methodology remains rooted in simulations. The predominance of simulations makes translating research to physical reality uncertain for both algorithms and researchers. We propose a physical platform that is designed to simplify the transition. In this paper, we present the Open Ant: a physical variant of the commonly used Gymnasium Ant environment, along with a simulation. We demonstrate that competent walk- ing policies can be learned from scratch in approximately one hour directly from the physical robot’s experience for two substantially different RL algorithms: SARSA(λ) and Soft Actor- Critic (SAC). Separately, we show policies that were learned in simulation transfer to reality.
We also examine how well the platform supports a nimble experimental ecosystem. Specif- ically, we observe the speed with which new users from diverse backgrounds achieve their first success with the platform, and how easily the platform can be repaired and updated when hardware issues arise. Both the hardware design and software are available as open-source on GitHub for ease of customization. In summary, we advocate for the use of the Open Ant for RL researchers who frequently use simulated environments, so they can more easily include robot experiments in their evaluations.
Ontributions
1. We developed a physical robot platform with an accompanying simulation, designed to support reinforcement learning researchers who are not accustomed to working with robots. The robot is inspired by the Gymnasium Ant used in continuous-control research, but sev- eral alterations were required to make a useful physical research platform. The hardware design and software are released to the community as an open-source platform.
Ontext:
The Gymnasium Ant (Schulman et al., 2015) is a popular simulated domain, with an abstract physical design. Reinforcement learning algorithm research is commonly conducted with simulated domains, though the community has used several robot platforms support such efforts.
2. We demonstrate that this platform is compatible with multiple reinforcement learning meth- supports learning directly from the hardware experience and also policy transfer from the simulator to reality.
Ontext:
The reinforcement learning community has adopted a wide range of methods and research directions. By successfully demonstrating the use of different RL algorithms and methods with this platform, we make it easier for RL researchers to adopt the platform.
3. We show that the platform supports a nimble ecosystem for RL experimentation, with the rapid on-boarding of inexperienced users and rapid revisions when hardware failures are encountered. The platform was used for RL experimentation by multiple researchers with limited prior experience in robotics. The platform supports rapid revision to encountered problems, through the use of 3D printing and commercially available components.
Context: RL researchers who used robots previously experienced challenges with both the initial adoption of a new platform, and with maintaining a robot as a research platform.
Arxiv:2607.18488V1 [Cs.Ro] 20 Jul 2026
The Open Ant: A Robot Platform for Reinforcement Learning Research
Earning Research
Elena Sorina Lupu1*, Patrick Spieler*, Khurram Javed2, Kris De Asis1, John D. Martin1, 3, Martha Steenstrup3, Joseph Modayil1, 2
Abstract
Reinforcement learning (RL) research has demonstrated success in both physical and simulated domains; however, the predominant methodology remains rooted in simula- tions. The predominance of simulations makes translating research to physical reality uncertain for both algorithms and researchers. We propose a physical platform that is designed to simplify the transition. In this paper, we present the Open Ant: a physical variant of the commonly used Gymnasium Ant environment, along with a simulation.
We demonstrate that competent walking policies can be learned from scratch in approx- imately one hour directly from the physical robot’s experience for two substantially dif- ferent RL algorithms: SARSA(λ) and Soft Actor-Critic (SAC). Separately, we show policies that were learned in simulation transfer to reality. We also examine how well the platform supports a nimble experimental ecosystem. Specifically, we observe the speed with which new users from diverse backgrounds achieve their first success with the platform, and how easily the platform can be repaired and updated when hardware issues arise. Both the hardware design and software are available as open-source on GitHub for ease of customization. In summary, we advocate for the use of the Open Ant for RL researchers who frequently use simulated environments, so they can more easily include robot experiments in their evaluations.
Ntroduction
Reinforcement learning (RL) has been successful in both physical and simulated domains, though many successes rely heavily on accurate simulations. Several of the notable successes include games such as Backgammon (Tesauro, 1995) and Go (Silver et al., 2016), where accurate and computa- tionally efficient simulators can capture the exact dynamics. Success in complex online games can introduce differences between training and deployment environments (Wurman et al., 2022). Suc- cesses in complex physical domains has often relied on human expertise to select the most relevant dynamics to simulate, with policies trained in simulation prior to a deployment phase, as seen in balloon navigation (Bellemare et al., 2020), fusion plasma control (Degrave et al., 2022), or gravi- tational wave detectors (Buchli et al., 2025).
Reinforcement learning research might be a victim of its own success with simulations, as rela- tively few studies focus on learning directly in physical reality. The ability to learn directly from physical experience has ample evidence in animal learning experiments, which served as motivation for early computational RL algorithms (Barto et al., 1983), with a modified version demonstrated on hardware (Benbrahim et al., 1992). This early motivation was followed by evidence that RL
Hip
Figure 1: Physical Ant platform, learning arena, and system overview. (a) An overhead webcam tracks the fiducial markers to compute reward signals and the heading vector of the ant. The robot is connected by cables to AC power and to an external computer where the agent is running. (b) The main components of the Physical Ant. (c) The Gymnasium Ant (Schulman et al., 2015; Towers et al., 2025), which was the inspiration for the Physical Ant.
algorithms provide computational models for the activity of dopamine neurons (Montague et al., 1996; Schultz et al., 1997). In addition to animal studies, many engineered systems have demon- strated the use of RL algorithms to learn behaviors directly from physical experience without the need for a simulator, with early examples on locomotion (Kohl & Stone, 2004; Tedrake et al., 2004) and later work demonstrating the use of a single RL algorithm across multiple robots and problem formulations (Mahmood et al., 2018b; Wu et al., 2023).
Despite repeated success with RL algorithms on physical robots, researchers with primarily simulation-based experience encounter significant challenges when experimenting with physical robots. One challenge is that the overhead (e.g., cost and time) for experiments with physical robotics is substantially higher than using simulations alone.
Even in situations where adequate lab resources and expertise are available, the long delay caused by the assembly of the robot and troubleshooting before the first successful experiment can exhaust a considerable portion of a researcher’s typical residency. Consequently, experts may graduate and depart research labs shortly after successful experiments, resulting in slower experiment iteration and difficulties with future re-use.
Contributions. We present the Open Ant, an open-source research robot platform available on GitHub. The robot comes with a MuJoCo simulation, which can be used to learn competent policies.
The robot body is designed to look like the widely used Gymnasium Ant environment (Towers et al., 2025). The Open Ant makes physical robot experiments an easier addition to a standard RL research pipeline. The robot is designed to be built and maintained by AI researchers without backgrounds in robotics engineering. The robot is designed to withstand the substantial wear incurred during RL exploration, and to be easy to repair when components are damaged.
The learning experiments presented here are designed to yield positive results within an hour; fur- thermore, we demonstrate that the robot platform provides reliable, repeatable results across differ- research for the broader RL community.
Platform Motivation And Background
Our overall goal is to create a research platform that simplifies the process for researchers to extend their simulation results in RL research. Based on our team’s experience on previous robot platforms, we identified several core criteria: the robot should closely mirror an existing simulated domain (see Figure 1), remain affordable, and integrate with standard RL software tools (see Section 5). Fur- thermore, running experiments should require minimal intervention, support diverse RL methods, and yield results on hardware within timescales comparable to simulation (see Section 5). Finally the overall platform should be accessible to newcomers and relatively easy to modify and maintain (Section 6). Although many prior works have demonstrated some of these capabilities, we are not aware of a platform that provides all of them. As such, our goal is not to produce a static benchmark, but to introduce a platform that the research community can flexibly revise over time—similar to the Arcade Learning Environment (Bellemare et al., 2013).
In addition to making robot experiments easier, we seek a platform that supports RL researchers to better study fundamental questions that arise when algorithms learn from physical experience. Invited talks at the 2025 RL Conference by Kaelbling (2025) and Sutton (2025) highlighted the complementary considerations for RL algorithms that happen at design-time (before regular deploy- ment) and run-time (during regular deployment). Design-time knowledge of the problem structure already informs many facets of a robot system using RL, including the selection of the observa- tion and action spaces (Mahmood et al., 2018a), the creation of the initial behavior policies (Silver et al., 2018), the physical design of the robot’s body, and the selection of appropriate simulators. In contrast, run-time knowledge can only be acquired in deployment. During run-time, the agent may encounter phenomena that are either unknown or imperfectly modeled at design-time.
Traditional adaptive control (Slotine & Li, 1991) addresses run-time learning by continuously up- dating a linear controller to compensate for changes in the environment parameters while the robot remains in operation. Examples of online learning using adaptive control on hardware include remote-controlled airplanes (Shi et al., 2020), quadcopters (O’Connell et al., 2022), and ground vehicles (Lupu et al., 2025). Reinforcement learning shares this same online learning/adaptation property but extends it beyond tracking or stabilization, enabling complex decision-making through interaction with the environment (Sutton & Barto, 2018). Physical robots provide a natural setting for studying algorithms that learn continually during deployment, where adaptation must occur from ongoing experience rather than repeated offline retraining.
In practice, most RL experiments are conducted in a lab without deployment (sim-to-real); however, we aim to study algorithms that can adapt behavior in deployment, outside of simulators and with limited researcher intervention. A primary constraint of the run-time setting is learning while be- having, without pausing the environment. Moreover, the robot’s physical experience differs from a simulation: a ground-truth simulator state is unavailable, and physical effects like overheating are difficult to model. Additional considerations include ensuring learning algorithms are compatible with local compute, communication, and timing constraints, as well as limiting the need for man- ual episodic resets. Our aim is to make these run-time conditions accessible and practical for RL researchers through a single, integrated robotic platform for learning from physical experience.
Although several papers have demonstrated run-time reinforcement learning on physical robots, it remains unclear whether these platforms are easily adopted by current RL researchers. One fruitful approach is to design custom robots that mirror existing simulated domains. For instance, Noodle- Bot (Berrueta et al., 2024) replicates the Gymnasium Swimmer benchmark, while RealAnt (Boney et al., 2020) is a quadrupedal platform inspired by the Gymnasium Ant. In contrast to ours, the RealAnt requires soldering and overhead camera calibration, and lacks clear guidance on setup time and reproducibility. In addition, its morphology also differs from the standard MuJoCo Ant model.
ple, Smith et al. (2022) show run-time policy learning on the Unitree A1 quadruped, and Kohl & Stone (2004) demonstrate a policy gradient RL algorithm that learns the parameters of a walking gait for the Sony Aibo quadruped. More recently, Preiss et al. (2025) demonstrate non-episodic, online
Reinforcement Learning Journal
Table 1: Comparison of the Gymnasium Ant and the Physical Ant platform.
N/A
policy optimization on a Crazyflie open-source quadcopter (Preiss et al., 2017). Even more closely aligned with our objectives, Mahmood et al. (2018a) previously demonstrated the instantiation of a Gymnasium Reacher environment with a commercial UR5 robot. We chose to pursue a different platform from the Reacher to access more domain complexity (higher dimensional observations and actions, mobility with contact dynamics) at a fraction of the cost. Commercial platforms can offer a relatively low barrier of entry for RL researchers familiar with simulations; however, they carry risks of product obsolescence, high purchase costs, long and costly repairs, and unfamiliar or rigid software interfaces which may hinder certain lines of research.
The Open Ant Platform
In this section, we introduce the Open Ant (Figure 1): an open-source quadruped platform for RL and robotics research. Our platform includes both physical and simulated robot bodies inspired by the Gymnasium Ant (Schulman et al., 2015), which serves as a popular benchmark for continuous- control research (Towers et al., 2025). Our physical robot body is called Physical Ant and the simulation is called Simulated Ant. Our platform is designed for researchers to easily move between simulation and reality when experimenting.
Esign Overview
Design Principles. Our design adheres to three fabrication principles. First, the design priori- tizes simplicity and ease of assembly by avoiding any need for soldering and relying exclusively on commercial off-the-shelf (COTS) components that are widely available. Second, the robot must support long-term learning without interruption; thus it is powered directly from an AC wall outlet, bypassing the maintenance overhead of rechargeable batteries. Third, the robot requires no special equipment; it directly connects to a personal computer via USB, eliminating the need for specialized embedded systems code or complex network configurations.
Physical Attributes. The Physical Ant is one third the scale of the Gymnasium Ant (Schulman et al., 2015). By default, the Gymnasium Ant stands approximately 75 cm tall with a total mass of 0.9 kg, its torso is 0.50 m in diameter; the upper leg segment measures 0.28 m and the lower leg segment 0.56 m. These dimensions imply an extremely low mass-to-size ratio that would be difficult to realize physically with conventional materials. The Physical Ant’s reduced size results in a relatively higher mass-to-length ratio, improving both its robustness and ease of manufacturing.
See Table 1 for the design details. Construction. The robot consists of a spherical body (torso) housing the onboard electronics and four articulated legs. The torso and legs are 3D printed, enabling rapid revision and localized repairs.
Each leg is equipped with two metal-geared Dynamixel actuators: one at the hip (XC430-W240) and one at the knee (XM430-W210). The spherical body shell houses the onboard electronics, including an IMU, a USB communication converter, and a USB hub that interfaces these components with an external computer (Figure 1, (b)). Additionally, the Physical Ant features an on-board camera, to support future vision-based learning tasks. The complete list of components is available on our GitHub repository. The robot communicates via USB with an external computing platform, a design choice that allows users to flexibly select their preferred compute hardware. As a result, researchers The Open Ant: A Robot Platform for Reinforcement Learning Research Figure 2: Overhead snapshots during run-time learning from physical experience. The cyan circle represents the boundary where the reward direction (the blue arrow) changes direction if the Physical Ant exceeds it. In the first figure, the robot travels towards 11 o’clock. In the second figure, it approaches the boundary. In the third figure, the reward direction flips and in the last two figures, we see the ant traveling in the direction of the reward direction.
can develop and run their algorithms without needing to adapt them to the computational constraints of an onboard platform. Power is provided via a wall-plug power supply, and the feet are equipped with 3D-printed thermoplastic polyurethane (TPU) “socks” to increase ground friction. The total cost for all the components is approximately USD 2200, excluding 3D-printing filament.
There exists a possibility to replace the more expensive motors (XM430-W210) with plastic-geared motors (XL430-W250) on both joints, which we call the Physical Ant Lite. For this lower-cost variant, the overall price is approximately USD 500 (excluding 3D printed parts). While more affordable, these actuators are more susceptible to wear and overheating. We therefore recommend that interested users carefully review the thermal analysis (Section 10.1), as well as the discussed electrical and mechanical improvements (Section 10.2 and Section 6) before deciding on the best motors for their needs.
Differences and Similarities between the Gymnasium and the Open Ant Environments We highlight several aspects of the Gymnasium Ant environment (Schulman et al., 2015) that mo- tivated our design. The Physical Ant adopts the same software interface as the Gymnasium Ant, while varying in the specific choices for its sensing, actuation, and reward.
Observations. Compared to its Gymnasium counterpart, the Physical Ant has a compact obser- vation space. Its observations have twenty-four dimensions, consisting of joint angles (8), joint angular velocities (8), the torso’s angular velocity (3), and linear acceleration (3) measured by an onboard IMU. In addition, we augment the observation with a two-dimensional unit vector encoding the angle between the heading vector and the goal direction (Equation 2) for the back-and-forth task presented below. Lastly, the platform provides interfaces for augmenting the observation space with other sensor modalities, such as motor loads and motor temperatures. In contrast, the observation space of Gymnasium Ant is 105 dimensions, composed of body-part positions, velocities, and ex- ternal forces acting on each segment. Observations exclude inertial planar position. Many of these observations in the Gymnasium Ant, such as the external forces on the body parts, are difficult to obtain on hardware.
Actions. The Physical Ant’s action space has eight continuously-valued, bounded dimensions. Actions represent position commands for the hip and knee joints of all four legs. Each command specifies a desired joint position to a Dynamixel actuator. For the Simulated Ant in MuJoCo, de- sired positions are tracked using proportional–derivative control whose gains and force limits were chosen to match the physical hardware (see Section 8 for system identification). In comparison, the Gymnasium Ant adopts a torque control strategy. Some prior real-world RL work uses position control (Miki et al., 2022). The software control interface can be changed to use position, velocity, or torque commands for other experiments.
Rewards and Tasks. The Gymnasium Ant’s default task is to move forward as quickly as possible. The reward function is a sum of multiple terms: a reward for staying upright, called a healthy
Reinforcement Learning Journal
reward, a forward reward for making forward progress, a control cost penalizing large actions, and an optional contact cost penalizing large external forces. The Gymnasium Ant also has episodic resets back to a starting configuration, which truncates episodes after long trajectories or on arrival to an unhealthy state. In contrast, we aim for learning in a non-episodic setting on the Physical Ant with a simple reward formulation based on progress alone.
Towards this end, we design a non-episodic task of moving back-and-forth. We define the robot’s planar position at time t ∈N to be pt = [xt, yt] ∈R2, measured in the inertial (world) frame and a reward direction unit vector ut ∈R2 defined in Equation 2. The instantaneous reward rt ∈R is given by the projection of the position displacement onto the current reward direction ut rt = (pt −pt−1)⊤ut.
(1)
We further define a circle of radius R > 0 centered at o ∈R2. When the robot’s position is outside the circle and it has made it to the half plane opposite from the previous reward direction update, a change in reward direction is triggered that compels the robot ‘bounce back’ from the circle, as
(2)
An illustration of the switching behavior in Equation 2 is shown in Figure 2. The advantage of the back-and-forth task, in contrast to the forward task specified in the Gymnasium Ant, is that it does not require the user to bring a physical robot back to the origin when it reaches the edge of the learning arena. In this way, the robot learns forever, with minimal user intervention.
Learning Arena. The robot pursues the back-and-forth walking task in an arena similar to the one rendered in Figure 1(a). An overhead webcam supported by a tripod system tracks two fiducial markers (Olson, 2011): one placed on the ant’s torso and another one placed on the ground. The system provides the position p used to compute the reward in Equation 1 and the heading.
Measuring Performance. The performance of the learned policy is measured with the average reward per second defined as
(3)
where rk ∈R is the instantaneous reward, N ∈N is the number of steps contained in the averaging window (corresponding to a user-defined window length), ∆t ∈R>0 is the time duration of one environment step, and t ∈N is the current time step.
Reinforcement Learning Methods
We provide demonstrations of learning from physical experience for two representative RL algo- rithms, under the back-and-forth walking task described in the previous section. The first algo- rithm is an on-policy action-value method: SARSA(λ) with linear approximation of the value func- tion (Sutton & Barto, 2018). SARSA(λ) has a small number of parameters and is easy to imple- ment. The second algorithm is Soft Actor-Critic (SAC) (Haarnoja et al., 2018), a deep RL off-policy implementation details in this section and report the corresponding experimental results in Section 5.
These two demonstrations show that learning from physical experience is feasible. They also pro- vide insight into the behavior and limitations of run-time learning with the Physical Ant platform. The Open Ant: A Robot Platform for Reinforcement Learning Research
Iscrete Action Rl: Sarsa(Λ)
SARSA(λ) is an on-policy RL algorithm where an agent interacts with the environment, learning an action-value function that defines the policy (for example, using ε-greedy action selection). The algorithm is defined with a discrete set of actions, so a transformation is needed from the robot’s continuous action space.
Before applying SARSA(λ) as a learning algorithm for the Open Ant, we first defined a slower and more abstract agent-environment interface. The abstract learning agent interface operated at a slower rate (0.5 s) than what was used for communication to the robot (0.05 s), where each agent timestep lasted for τ = 10 robot interactions. These timing values were a design choice. On each from the robot were accumulated before being sent to the learning agent, where the agent’s reward between the robot timesteps t and t + τ is the discounted sum rt+1 + γrt+2 + . . + γτ−1rt+τ, where γ is the discount factor. For simplicity, we use the discounted formulation of SARSA(λ) rather than an average-reward formulation. The robot’s continuous command space was abstracted into a discrete set of long duration motion primitives, which formed the learning agent’s actions.
We defined each agent action ¯a with an open-loop motion primitive ⟨a1, ...aτ⟩, where τ was the fixed duration and ai, with i ∈[1, τ], was the joint-space command sent to the robot at each robot timestep. The design of these motion primitives was inspired by observing the behavior of standard quadruped robots and animals during locomotion. Specifically, our motion primitives were defined as follows. For a hip joint, we used a linear ramp. For a knee joint, we use a half-sine with amplitude 1 or -1, corresponding to “lift leg” versus “push leg into the ground” primitives. A diagram of these motion primitives is shown in Section 15.
We approximated the true action-value function Q(s, ¯a), where s ∈S was the state, using linear function approximation with a tile coder (Sutton & Barto, 2018). Tile coding is based on the Cere- bellar Model Articulator Controller (CMAC) (Albus, 1975) and was applied in RL before in Watkins (1989) and Sutton (1995). The robot’s observation vector served as the agent’s state. Each state s was mapped by the tile coder to a sparse binary feature vector ϕ(s), and each agent action ¯a had a separate weight vector w¯a. The estimated action value was defined by ˆQ(s, ¯a) = w⊤
¯A Φ(S). Each
agent action was selected using an ε-greedy policy with respect to the current action-value estimate. Using the abstract agent interface, we applied the standard SARSA(λ) algorithm with eligibility traces as described in Sutton & Barto (2018). Because our tasks were non-episodic, learning pro- ceeds without episodic resets, and eligibility traces decay without being cleared at episode bound- aries. With this formulation, SARSA(λ) operated as a fully online, continuing-control algorithm directly on the physical platform. The results of this implementation are presented in Section 5.
Ontinuous Action Rl: Soft Actor-Critic (Sac)
Soft Actor-Critic (SAC) is an off-policy RL algorithm designed for continuous action spaces. SAC combines actor–critic methods with entropy regularization. The original algorithm (Haarnoja et al., 2018) learns a stochastic policy network, two action-value functions, and one state-value function network, all using a replay buffer. Newer implementations of the SAC algorithm omit learning the state-value function and achieve similar performance (Achiam, 2018). We have modified the CleanRL implementation (Huang et al., 2022) and applied it to the Physical Ant, with the observa- tions and actions presented in Section 3.2. The action bounds were adjusted to constrain the knee to a 20◦range with a 50◦offset, and the hip to a 45◦range with 0◦offset.
Reinforcement Learning Results
We begin by introducing the common methodology employed by both SARSA(λ) and SAC, where learning is performed directly on board the Physical Ants. We then present the hardware results of these two reinforcement learning algorithms introduced in Section 4.
Trial 5
Figure 3: Run-time learning performance of SARSA(λ) executed onboard the Physical Ant across 5 trials. Average reward is computed using Equation 3 on a window of 120 seconds. These 5 trials had a total of one interruption (indicated with circles) caused by the leg failure in Figure 6(g).
The leg was repaired within 10 minutes and the experiment was continued. The experiments were run on a Desktop computer with the Intel Core i9-13900K CPU. Common Methodology.
The task was the back-and-forth objective defined in Equation 1, in which the robot must walk within a specified circle. For all experiments, the radius of this cir- cle was set to R = 0.3 m. This value was determined by the dimensions of the operational arena (Figure 1(a)). At the start of each trial, the power and communication cables were positioned out- side the circular region. The Physical Ant was then placed inside the circle and the cables (power and communication) were arranged so that they were not entangled with each other and with the robot’s body. We refer to this configuration as the start configuration. During learning, the robot oc- casionally became entangled in the cables. When this occurred, the experimenter manually paused the experiment, disentangled the robot, then reset the robot by returning it to the start configuration within the circle, and resumed the experiment. We have also ensured that the interaction was per- formed within the same fixed-duration control loop ∆t. Lastly, for both tasks, have not designed a complex reward for the experiments presented, but used a simple progress reward, as seen in Equation 1.
Each RL algorithm was run on the robot for 80 minutes per trial, and we conducted five independent trials. SARSA(λ) was evaluated only on the Physical Ant, and SAC was evaluated on both the Phys- ical Ant and the Physical Ant Lite. The evaluations used the performance metric from Section 3.3.
Evaluation of SARSA(λ). Learning with SARSA(λ) on hardware included several design deci- sions. First, to select good algorithm parameters at design time, we used Optuna (Akiba et al., 2019) for automated parameter optimization with the simulation. We conducted a sweep over the parameters listed in Section 12, evaluating each of 100 configurations across ten random seeds. The configuration that achieved the highest average reward was selected for the final experiments and is shown in Table 2.
Table 2: Parameters for SARSA(λ) with tile coding used for the experiments in Section 5.
S
Next, the specification of observation bounds for the tile coding was important. Because we used a tile coding library Tiles3 that depends on the scale of the inputs, we precomputed realistic obser- vation ranges by driving the robot through diverse motions in multiple directions. These measured The Open Ant: A Robot Platform for Reinforcement Learning Research Table 3: Parameters for SAC used for the run-time learning on the Physical Ants.
Frequency
bounds were then used to normalize the observations prior to the generation of tile-features by the software library. Third, the design of the motion primitives and ensuring they are realizable on hardware substantially influenced learning. In particular, we found that primitives coordinating two legs simultaneously, rather than actuating a single leg in isolation, produced more dynamically meaningful behaviors.
Results from five independent trials are shown in Figure 3. In every trial, the agent achieved slow yet competent walking behaviors (average reward of 2-4 cm/s) within approximately one hour of on-hardware learning. The key result is that learning is reliably observed across trials, even with a very simple RL algorithm.
Evaluation of SAC. We implement Soft Actor–Critic (SAC) on the two Ant platforms, the Physi- cal Ant and the Physical Ant Lite. Both platforms share identical observation spaces, action spaces, and reward functions. Except for the modifications stated below, the algorithm parameters (Table 3) and network architectures for the policy and action-value functions follow the default configuration in the implementation of CleanRL (Huang et al., 2022).
First, we reduced the number of random interaction steps before learning begins from 5000 to 2000 (approximately 4 minutes of on-hardware random interaction). Empirically, this earlier start of gradient updates produced average reward performance comparable to the default setting, while reducing random actions taken on the robot, which could potentially damage it, and enabling faster learning (see Section 13.1 in the supplementary material).
Second, we incorporated Layer Normalization (LayerNorm) (Ba et al., 2016) into both the policy and critic networks. Prior work suggests that LayerNorm can stabilize learning (Elsayed et al., 2024); we verified this effect using our Simulated Ant, as shown in the supplementary material Section 13.1, across 30 simulation seeds. LayerNorm both accelerates learning and reduces variance. This faster convergence is particularly important in hardware settings, where interaction time is costly.
Third, we found reward scaling to be important for stable SAC training. The SAC algorithm has an entropy component that is sensitive to the reward scale. An initial reward scaling factor of 1.0 led to poor learning performance on hardware, whereas scaling rewards by a larger factor substantially improved performance and enabled consistent learning across seeds. We validated the effect of the scaling reward parameter in simulation across 30 seeds (supplementary material Section 13.1) and demonstrated that a larger value produces faster learning.
Lastly, we apply the modifications recommended in De Asis & Sutton (2024) to the return definition. This alleviates a idiosyncratic dependence on time-discretization for the algorithm. It decouples the optimization objective from ∆t, making it a solution parameter instead of a problem parameter.
Figure 4 shows the average back-and-forth reward [cm/s] over time for five trials, on both physical ing behaviors, demonstrating that SAC can reliably learn locomotion directly on hardware in this configuration. We note that the SAC behavior had more interventions to clear cable entanglement than encountered with SARSA(λ), but the robot was also moving more with the SAC policies.
Translating policies from the simulator to physical reality using SAC.
A Popular Approach To
obtain competent robot behavior is to simulate an approximation of the robot and its environment,
Trial 5
Figure 4: Run-time learning performance using SAC on hardware for 5 trials. Average reward is computed using Equation 3 for both the Physical Ant and the Physical Ant Lite on a window of 120 seconds. Circles indicate interventions, which are manual stops due to cable entanglement with the robot’s legs. The experiment was run on a Macbook M1 with 16 GB of memory. The performance for one of the runs can be seen in Movie 2.
and learn a policy in simulation. The policy can then be deployed on the real robot as is, or adapted on the robot with additional experience.
Tion 11) For A Duration Of 400,000 Timesteps (The
equivalent of about 13 hours).
Policies On The Physical Ant For 5,000 Steps (10
minutes), with no interventions. We report the mean performance in Figure 5. We noticed two key results. First, all policies transferred with some success and enabled the robot to walk in the correct direction.
This shows that the Simulated Ant captures the key aspects of the Physical Ant well. We also noticed that the less proficient policies make the ant entangled with the cable a lot more, for example policy 10 and policy 7.
Second, the ranking of the policies changed when they were transferred to the Physical Ant. For example, the policy that performed the best on the Simulated Ant was the sixth best on the Physical Ant. This change in ranking is an important result because policy improvement requires the ability to compare two policies. If the comparison is not accurate, then finding the best policy from a large number of policies learned in simulation poses challenges for hardware deployment.
The Open Ant: A Robot Platform for Reinforcement Learning Research Note that the policies from our sim-to-real analysis are not directly comparable because they are chosen from a set of parameters that do not always overlap with the parameters used for learning from physical experience (Table 3).
Examining the Platform Suitability for RL researchers Most RL researchers have limited familiarity with robotics. Indeed, the authors have frequently heard RL researchers say that they avoid robotics experiments, due to difficulties they experienced in the past. One common problem is the long delay to the first successful experiment for a researcher using a novel robot with RL. Another common problem is the difficulty in repairing or adapting the physical system when failures arise. We share our experiences on both these common concerns. We have observed that multiple researchers could use the platform within a few days to learn policies, and we have found that many failures with the hardware can be rapidly corrected, due to the 3D printed design, and the use of commercially available (COTS) components.
Easy For Many To Build And Use
An important goal for the Open Ant platform is to enable researchers to build, deploy, and iterate on the platform with minimal effort. To date, the robot has been independently assembled and tested across multiple locations, including Canada, USA, and Malaysia. Once all the components were purchased, it took between 2 to 5 hours to assemble a full robot. The Open Ant was used for experimentation with RL at two meetings of academic researchers, with initial use at a workshop The winter school provided evidence that the platform can be successfully used by researchers who were inexperienced in RL and robotics. At the school, five independent teams (total of 20 partici- pants) successfully used and modified SAC and SARSA(λ) implementations. The participants tested sim-to-real transfer on the physical robot in an episodic setting (a walking forward task). The teams consisted of researchers who had limited experience with robotics and RL, and no prior experience with this particular platform. Nevertheless, all teams were able to demonstrate successful learning and transfer with three days of access to the physical platform. The physical platform was assembled and externally sourced commercial components. The assembly process highlighted the platform’s low barrier to entry; as noted in a testimony from the local assembly team: “once we understand how all the parts fit together, it takes less than two hours to assemble everything. When we tried building it (the Open Ant) without a guide, it took us around four to six hours working on it on and off. Most of that time was spent assembling, realizing mistakes, and then taking things apart to fix them”. The participants used their own laptops to control the systems with a variety of operating systems (Linux, MacOS, Windows). An illustration of the experiments performed at the Winter School can be seen in Movie 6. These deployments demonstrate that the system is accessible to a broad range of researchers, and that the platform provides a rapid path to successful RL experiments with a robot.
We speculate that some design choices make this a good introductory physical platform for RL researchers. The combination of the simplicity of the platform, AC-powered long-duration opera- tion, a standard Gymnasium interface, and compatibility with external compute platforms created a streamlined and engaging researcher experience.
Platform Reliability
From earlier experiments performed during the development of this platform, we observed several mechanical failure modes on the Physical Ant. As shown in Figure 6, these issues were primarily caused by ground impacts, joint loadings, and accumulated stress in 3D-printed components. In response, we iteratively improved the hardware design by reinforcing interfaces, redesigned vulner- able parts, and added protective elements such as 3D printed Thermoplastic Polyurethane (TPU)
Reinforcement Learning Journal
Figure 6: Different failures observed during the development of the Physical Ant (a) Foot- tip abrasion from repeated impacts, mitigated with 3D-printed TPU socks. (b) Hip cover fracture due to locomotion stress, redesigned with a stronger interface. (c) Leg shell damage after manual related to (e), interface strengthened. (e) Motor contamination from excess Loctite and plastic gear wear (XL430-W250-T). (f) Outer shell crack, mitigated by tightening screws to maintain structural integrity. (g) Leg broken during one of the SAC experiments. We notice a missing screw which could have affected the load distribution.
foot socks. We have also performed thermal improvements based on a detailed thermal analysis (Section 10.1), and power and communication fixes (Section 10.2). These modifications were im- portant for enabling long learning runs. By releasing the hardware design as open-source, we aim to encourage further community-driven improvements to the platform’s structural robustness and long-term reliability.
Iscussion And Future Directions
We have presented the Open Ant as a platform for RL researchers who are familiar with simulation domains to include robot experiments in their evaluations. For our primary purpose of easing adop- tion, we have made the platform intentionally close to existing simulation environment and tasks, so that researchers can succeed with their initial attempts.
An RL researcher who is familiar with algorithm development in simulation may question the value of including robots as part of their research methodology. One argument in favor is that if our ulti- mate goal is to build learning agents that operate in the physical world, then experiments with the algorithms in the physical world are essential. Even for research programs that target non-physical domains, like agents for the PlayStation game GranTurismo (Wurman et al., 2022), researchers may not have direct access to the deployment environment with human participants. Those researchers may also find it valuable to evaluate their algorithms with a physical robot, where they can access a non-simulated deployment environment. A physical robot platform provides a valuable testbed for validating algorithms and exposes discrepancies between simulation and reality that would other- wise remain hidden.
Practical Concerns For Future Research
We consider some factors that can make the direct application of reinforcement learning algorithms to hardware more practical. These considerations are taken with respect to the current state-of-the- art.
Random Exploration. To safely deploy current learning algorithms on physical robots, their typical random exploration must not compromise the integrity of the robot. For example, motors, gearboxes, and structural components must tolerate the wear induced by this exploratory behavior. The Open Ant is designed to tolerate such exploration; it uses relatively large motors for its size, and the joint limits and kinematics prevent any self-collisions. The safety requirement imposes constraints on the robot design or the algorithm. The Open Ant addresses this issue in a straightforward way by over-sizing its actuators. More complex robots could require substantially more effort.
Unrecoverable States and Resets. Second, unrecoverable states should be rare. For example, our the robot can continue walking even if it rolls onto its back, provided the full knee range of motion (±70◦) is used. The back-and-forth locomotion task was designed to require minimal human inter- vention by keeping the robot within the camera field of view. This allows the robot to learn without frequent interventions. The cable entangling with the robot did occasionally require intervention and in the future we will explore mitigating this through the use of a cable management system Gymnasium humanoid walking task, an alternative we considered early in the design phase would require interventions which are trivial to implement in simulation (by resetting the environment) but burdensome in physical experiments. In summary, the design of the physical experiment should minimize physical human interventions and resets.
Performance vs. Compute We also observed that learning performance can vary depending on the specific computer instance on which the agent is executing (Section 14). Although we adjust the duration of the environment interaction step to maintain a fixed interaction frequency, the timing between issuing a motor command and reading sensor values for the next observation will depend on the time the agent takes to decide on its action. The impact of this timing on the Open Ant plat- form, as well as other factors that depend on compute speed, like camera processing latency, deserve further analysis. Furthermore, conventional agent environment interfaces, such as Gymnasium, lack specification of timing aspects that are unavoidable when operating in the real world. We acknowl- edge that timing remains a concern despite some existing works addressing it, for example Farrahi & Mahmood (2026); Yuan & Mahmood (2022); Rupam Mahmood et al. (2018).
Environment Changes. The environment changes over time as the agent learns from physical experience. For example, we observed that the robot gradually created holes in its feet during learning, altering their structural integrity, as shown in Figure 6(a). We have mitigated this by building ’socks’ out of thermoplastic polyurethane material. In addition, repeated walking on a Medium-density fiberboard (MDF) wooden floor generated fine dust, which accumulated over time and changed the friction between the feet and the surface. Such effects are difficult to capture accurately in simulation and are often difficult for the designer to anticipate in advance, which motivates continual learning.
Reward Drop. We observed occasional runs in which the performance in simulation drops to nearly zero. Similar behavior has also been observed on the physical robot, as illustrated in Movie 4 for SARSA(λ). In both cases, the agent eventually recovers and resumes learning, as shown in Movie 5 for the simulation. It is unclear what precisely causes this performance drop, but it requires further investigation. This phenomenon is illustrated in Figure 13 in Section 13.2.
Simulator Exploits. A practical limitation of sim-to-real is that policies can exploit inaccuracies in the simulator rather than learn behaviors that transfer to the real world. Workarounds include inten- sive domain randomization (Tobin et al., 2017) and reward engineering. Although this concern does not apply to our hardware learning experiments, it was present in our early sim-to-real experiments (Section 5). As discussed in Section 8 and shown in Movie 3, we observed policies that achieved
Authors:
Peder EZ Larson 1, 2,* , Jenna ML Bernard1, James A Bankson 3, Nikolaj Bøgh 4, Robert A Bok1, Albert P. Chen 5, Charles H Cunningham 6,7, Jeremy Gordon1, Jan-Bernd Hövener 8, Christoffer Laustsen 4, Dirk Mayer 9,10, Mary A McLean11 12, Franz Schilling13, James Slater1, Jean-Luc Vanderheyden5, 14, Cornelius von Morze 15, Daniel B Vigneron1, 2, Duan Xu1, 2, and the HP 13C
94143, Usa.
Denmark. 5 GE Healthcare, Menlo Park, California, USA. 6 Physical Sciences, Sunnybrook Research Institute, Toronto, Ontario, Canada.
8 Section Biomedical Imaging, Molecular Imaging North Competence Center (MOIN CC), Medicine, Baltimore, MD, USA. Cambridge, United Kingdom.
14Jlvmi Consulting Llc, Dousman, Wi, Usa
#See Acknowledgements for a list of all HP 13C MRI Consensus Group Members This work was supported by the ISMRM Hyperpolarized Media MR Study Group, the ISMRM Hyperpolarization Methods & Equipment Study Group, and the Hyperpolarized MRI Technology Resource Center (NIH/NIBIB grant P41EB013598).
Abstract
MRI with hyperpolarized (HP) 13C agents, also known as HP 13C MRI, can measure processes such as localized metabolism that is altered in numerous cancers, liver, heart, kidney diseases, and more. It has been translated into human studies during the past 10 years, with recent rapid growth in studies largely based on increasing availability of hyperpolarized agent preparation methods suitable for use in humans. This paper aims to capture the current successful practices for HP MRI human studies with [1-13C]pyruvate - by far the most commonly used agent, which sits at a key metabolic junction in glycolysis. The paper is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification. In each area, we identified the key components for a successful study, summarized both published studies and current practices, and discuss evidence gaps, strengths, and limitations. This paper is the output of the “HP 13C MRI Consensus Group” as well as the ISMRM Hyperpolarized Media MR and Hyperpolarized Methods & Equipment study groups. It further aims to provide a comprehensive reference for future consensus building as the field continues to advance human studies with this metabolic imaging modality.
Keywords: Hyperpolarized MRI, metabolic imaging, carbon-13, pyruvate, dissolution dynamic
Introduction
MRI with hyperpolarized 13C agents, also known as hyperpolarized (HP) 13C MRI, has shown great potential as a novel imaging modality, particularly for its ability to probe metabolic processes in real time. The first human studies with HP [1-13C]pyruvate were performed in 2011 in prostate cancer patients (1).
Since then, there have been over 60 papers published with imaging results of human subjects from 13 different sites, with applications including prostate cancer, brain tumors, breast cancer, kidney cancer, pancreatic cancer, metastatic disease, liver disease, ischemic heart disease, diabetes and cardiomyopathies. The vast majority of these studies used [1-13C]pyruvate (1–63), where [2-13C]pyruvate (64) and 13C-urea (56) have been demonstrated too.
As clinical HP 13C MRI advances, there is a growing need to build consensus for best practices, which are critical for comparing data across sites, performing multi-site trials,deploying methods to new sites, partnering with vendors, and potentially for obtaining broader regulatory approvals.
In March 2022, we initiated an effort to build consensus within the HP 13C MRI community with this opportunity in mind, and it was greeted with strong enthusiasm. The “HP 13C MRI Consensus Group”, containing over 55 members from 27 sites, identified the area of greatest need and opportunity for consensus building to be HP [1-13C]pyruvate human
●
Pyruvate is the most mature and widely used HP agent and has the most significant translational evidence emphasizing the potential clinical impact.
●
Clinical trials, particularly multi-site trials, have the strongest need for consensus methods to ensure that data can be combined across sites. This work is a Position Paper for which the goal is to describe current successful practices and study methods for HP [1-13C]pyruvate human studies along with justification to support those practices. This is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification (Fig. 1). The current successful practices and study methods include a literature review of published peer-reviewed journal papers showing human HP [1-13C]pyruvate study data, up to September 2022 (1–63), as well as new unpublished information from surveys of HP 13C study sites. Based on this information, we also highlight the evidence gaps, strengths, and limitations of current practices which are summarized at the end of each section.
Figure 1: Illustration of the HP 13C MRI human study process, including the 4 major areas covered in this paper: Hyperpolarized 13C-pyruvate preparation, MRI system setup and calibration, Acquisition and Reconstruction, and Data Analysis and Quantification.
Figure 2: Anatomical targets of HP [1-13C]pyruvate MRI human studies published up to September 2022.
Hyperpolarized 13C-Pyruvate Preparation
This section covers the processes for creating the HP agent, 13C pyruvate, and will include many aspects and considerations that are needed to safely and effectively prepare doses for metabolic imaging studies in human subjects. These include material, personnel, equipment and facility, fluid path preparation, quality control, and release.
It is helpful to understand that the specifications of a dose of 13C pyruvate suitable for in vivo MR HP metabolic imaging were shaped in part by early preclinical studies performed by GE HealthCare summarized in Ref. (65). In short, the safety of the two novel drug components, 13C pyruvate and the electron paramagnetic agent (EPA) AH111501, were demonstrated in those studies. The more precise formulation of the dose suitable for human use was then determined from clinical studies (66) that included two Phase 1 clinical trials in young and elderly healthy volunteers without hyperpolarization of the 13C nuclei and another Phase 1/2a dose escalation and imaging feasibility study with HP 13C pyruvate in 31 prostate cancer patients at the With the exception of the first HP 13C imaging clinical trial, which utilized a prototype device in a cleanroom (1), all HP 13C studies performed in humans to date have utilized the SPINlab polarizer (manufactured by GE HealthCare). Consequently all doses of the HP 13C pyruvate delivered by SPINlab have been produced using the “SPINlab Pharmacy Kit” that serves as the container-closure system for the various drug components (13C pyruvic acid and EPA mixture, dissolution medium, and neutralization and dilution medium) during sample polarization, dissolution and quality control (QC) processes. Thus many aspects of the HP sample preparation considerations discussed below are related to the SPINlab instrument and the consumables designed to be used with it (67).
General Considerations
While more than 860 patients or healthy subjects having been injected with HP 13C pyruvate as of January 2022 without reports of any serious adverse events (68), HP 13C pyruvate injection remains an investigational MR contrast agent and can only be administered by those with Investigational New Drug (IND) exemption from the Food and Drug Administration (FDA) in the USA, a Clinical Trial Application (CTA) in Canada, approval from National Research Ethics Committee Services in the UK, or approval from the relevant local regulatory body. Thus, methods and processes involved to produce a dose should have patient safety as the first priority. Since utilizing dissolution dynamic nuclear polarization (dissolution-DNP) for human use is still a relatively new development, there are no existing published regulatory guidelines specifically for this method.
There are two major production styles that determine how various sites approach the agent preparation. In the US, the most common approach is to rely on a sterilizing filter (“Terminal Sterilization”) to ensure sterility of the final product, akin to PET tracer production, where a starting molecule with a radioisotope is processed using various other ingredients to make the final, desired and injectable contrast agent within a necessarily short amount of time (69). For these sites, sterilization of the components and accessories upstream of this filter are not required, although many of them were manufactured and tested following Good Manufacturing Practice (GMP) or Good Laboratory Practice (GLP) requirements. The filling process is usually performed under an ISO 5 laminar flow hood, but a clean room or an isolator is not required.
This approach is typically accompanied by testing the integrity of the sterilizing filter prior to release of the dose for injection. Typically, post release endotoxin and sterility tests are performed using an aliquot reserved from each released dose.
In the UK and EU, the most common approach is to more-closely follow sterile pharmaceutical compounding guidelines (70), where all components and ingredients are required to be sterile or manufactured under GMP guidelines and are assembled and filled within a clean room environment or an isolator system (“Sterile Preparation”). Typically a batch of Pharmacy Kits for HP 13C pyruvate injection are prepared together. The sterility of the final dose is also ensured by batch validation testing, in addition to the sterility of the ingredients and the sterile compounding process. The endotoxin and sterility testing are performed for the process validation but are not performed for each injected dose.
Some institutions fill and assemble the Pharmacy Kit required for a specific study on the same day or the day prior to polarization, dissolution, and patient administration, but others have also demonstrated the feasibility of preparing a batch of kits, keeping them in a -20ºC freezer and using them over a period of a few months.
Beyond the obvious requirements that the process and the facility has to ultimately produce a dose that is safe to inject into a human, regulatory authorities will also focus on the question “Are you in control of your processes?”. To be in control of your process requires an in-depth and broad understanding of all processes involved in pre, post, and during the production process.
Personnel
It is typical and may be required to have licensed personnel involved in the production process depending on local regulations.Typically a pharmacist, radiopharmacist or other similarly qualified person (QP), in charge of the facility where the Pharmacy Kit filling and preparation is taking place, is responsible for the overall process and the release of the injectable dose.
Qualified cleanroom technicians are often involved in the Pharmacy Kit filling under the supervision of the pharmacist or QP. As is required for pharmaceutical compounding or PET tracer production, training requirements and training records for all personnel need to be maintained and available for audit by the FDA or equivalent.
Equipment And Facility
The facility and all equipment need to have standard operating procedures (SOPs) that describe how equipment is used, maintained, and calibrated to comply with relevant legislation. Currently, almost all the filling of the Pharmacy Kit takes place within a compounding laminar flow hood or isolator (typically ISO 5). At some sites, the filling is conducted within a cleanroom, while at others, it is conducted in a dedicated non-cleanroom space, reflecting differences in cleanroom approach and specifications between regulators worldwide (71). Some equipment or facilities, such as the compounding hood or cleanroom, may require external certified laboratories for testing.
Material Handling
Material handling guidelines (69,70) require SOPs detailing a system to track all of the materials involved in the HP production process for a particular patient dose, similar to current good manufacturing practice (cGMP) requirements for material handling for drug compounding. This includes acceptance standards, storage conditions, amount used in the patient dose for each ingredient and materials used in the assembly of the fluid path and Pharmacy Kit. Currently some users choose to open and inspect and sometimes modify the Pharmacy Kits upon arrival, but some users keep them in the sealed packaging until they are required for dose preparation.
Pharmacy Kit Filling And Assembling
As required by an IND or its equivalent, the preparation of the doses of HP 13C agent are detailed in the Chemistry, Manufacturing, and Control (CMC) section of an applicable regulatory submission; an example of this has been made available (72). It describes the processes of filling the Pharmacy Kit with the different components that make up the final drug product, and of assembling the final kit for either storage or immediate use in the polarizer. Special attention should be given to the laser welding process in order to satisfy installation qualification (IQ) and operational qualification (OQ). Typically, the final developed process is validated by process qualification (PQ) runs, during which 3 or more Pharmacy Kits are filled and used and the final HP 13C products are tested for endotoxin and sterility and to confirm that they meet the dose specifications for injections (usually including pyruvate concentration, residual EPA concentration, pH, liquid state polarization level and dose temperature). The data from 3 consecutive PQ runs are submitted as part of the IND submission (or its equivalent), and are often also reviewed by the Institutional Review Board (IRB) where the studies are conducted.
Quality Control And Dose Release
The quality control (QC) and dose release can be separated into two aspects: one is the QC and release of the filled Pharmacy Kit, and second is the QC and release of the HP 13C agent for injection, after polarization and dissolution. For institutions filling a batch of kits and storing them to use over a period of time, typically the batch can be released based on initial validation, environmental monitoring data from the day of kit production, and if filters are used during preparation of any of the components, filter integrity testing. But in some cases one or more kits are used for validation before the batch of kits are released for future use. For institutions that fill only the kits required for specific studies shortly before the experiment, the filled kits often do not go through separate release tests before they are used.
The quality control of the HP 13C pyruvate solution post dissolution is primarily performed to ensure that the agent meets the dose specifications (Table 1) before it is administered to the subject. These specifications target both safety (pH, residual EPA, temperature) and efficacy (pyruvate concentration, polarization, volume). Typically, the pyruvate concentration, residual EPA concentration, pH, dose temperature, dose volume, and liquid state polarization are measured by the QC accessory associated with the SPINlab polarizer. Some users perform a secondary measurement for one of the parameters, such as pH, using a different instrument or pH paper. For sites that do not go through a separate release testing process for batch filled kits, the integrity of the sterilization assurance filter, a part of the Pharmacy Kit, is typically tested as a part of the dose release. It is also common for these users to preserve an aliquot of the final HP 13C pyruvate solution for post-release endotoxin and sterility testing. This testing cannot be completed fast enough to test an individual dose prior to injection, but this is why other processes such as PQ runs and validation testing are done to minimize the chance a subject could be injected with a contaminated dose.
The Final Dose Release And Injection
should be done under the supervision of a licensed professional, based on local regulations.
Some Key Challenges
Many of the challenges associated with HP 13C pyruvate preparation can be attributed to the conditions required for the dissolution-DNP method of high magnetic field (~3-7 T) and very low temperature (~1 K) during polarization, with pressurized and superheated water necessary for the rapid dissolution event. These extreme conditions are quite challenging for the design of the container-closure and fluid path system. In particular, the cryogenic temperature in the polarizer requires special attention to any moisture or ambient (moist) air introduced into that portion of the fluid path, which can form an ice block at ~1 K. This ice can lead to flow restriction during the dissolution event and reduce the strength of the laser welded bond between the cryovial and its cap. This can ultimately produce failures in the dissolution step, including variations in final pyruvate concentration and pH that may fail to meet QC release criteria as well as fluid path ruptures that provide no available dose and result in polarizer down-time.
The polarization of the HP 13C pyruvate sample decays quickly over the span of a few minutes after dissolution, and thus the process of dissolution, QC for release, and injection should be completed as fast as possible to preserve the high polarization level achieved. Any delays in the preparation process, such as transportation time or equipment malfunction, can significantly reduce the final polarization and result in lower quality imaging data.
Current Practices
A summary of data collected from all sites performing clinical trials with HP 13C-pyruvate is shown in Fig. 3 and Table 1, including the specification of the final dose and how the quality control and release of the final dose are performed. There is a split in the Production Style, described in the General Considerations section above, with 8/13 sites using Sterile Preparation versus 5/13 using Terminal Sterilization. While many of the dose specifications show notable differences in acceptable ranges, all of these variations listed in tables have been successfully and safely been used to perform HP 13C pyruvate studies in humans. Their differences depend on the institutions’ preferences, resources and their particular regulatory situation. There is high similarity in pyruvate ranges, temperature ranges, EPA limits, and volume limits. There is modest variability in pH ranges and large variability in the endotoxin test limit. There is a 3-fold difference in acceptable polarization levels, which are measured to ensure a futile dose is not injected since the polarization is directly proportional to SNR. This reflects the decision by several sites to believe that useful data can be still be obtained with suboptimal polarizations.
Figure 3: Hyperpolarized agent preparation methods reported by sites currently performing HP
In House
Table 1: HP 13C-pyruvate preparation parameters, methods, and dose specifications used for quality control testing and release as well as validation. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. The parameters used for product release are noted in bold text, otherwise these parameters are measured for batch validation or other QC measurements. The endotoxin and sterility testing are performed during process validation of the batch and/or post-injection, and largely depends on the agent production approach.
Summary
The overall safety record of HP 13C-pyruvate has been very strong, and the SPINlab hyperpolarizer has proven to provide high polarizations at human sized doses while meeting numerous QC and release criteria. A weakness remains the failure modes of the SPINlab Phamacy Kits (e.g. ice blocks, path ruptures), which are placed under extreme requirements particularly during dissolution. The preparation process still requires a high degree of expertise.
Therefore, there is a significant need to improve the reliability, robustness, and ease of operation for generating HP 13C-pyruvate doses for human studies. Furthermore, there is a divide between manufacturing and sterile compounding style preparation as well as other site-specific practices, resulting in variations in SOPs and justification required to relevant regulatory bodies. There have also been no comparisons between these approaches. It is also unclear what release criteria and QC parameters are truly required to ensure patient safety.
However, all of the reported methods are acceptable and approved by the appropriate regulatory authorities, and have led to the rapid expansion of successful human studies in recent years.
Mri System Setup And Calibrations
This section covers the MRI system setup, including the imaging system, RF coils, phantoms, and prescan calibration methods.
Imaging System
The main prerequisite for a given MRI scanner to be capable of supporting studies with HP 13C is its “broadband” capability to transmit and receive radiofrequency (RF) signal at the frequency of 13C, which is around 4 times lower than 1H. This does not come as a default on clinical MR devices. The transmit power of the broadband amplifier should also be sufficient to support the intended flip angle and RF pulse shape with the employed transmission RF coil(s) for 13C. Most studies to date use relatively low flip angles (< 90 degrees) for HP 13C in order to preserve polarization for time-resolved imaging. The capability to receive 13C signal on multiple channels is also desirable to increase SNR, as discussed further in the “RF coils” section.
The choice of magnetic field strength is primarily dependent on the metabolites’ frequency separation due to chemical shift dispersion and 1H imaging. High field strengths do not enhance hyperpolarized 13C signal as they do for 1H because the signal strength in a HP experiment relies on manipulating the population of quantum energy states outside of the MRI scanner.
However, the injected HP 13C-pyruvate and its metabolic products have greater frequency separation at higher fields, and it may thus be easier to separate and quantify these resonances at higher fields. This comes at the cost of a reduction in the achievable T2* and often reduced T1. As the initial polarization is independent of the imaging field strength it has been proposed that the increased T2* at 1.5T can potentially be exploited to increase SNR by adapting the acquisition bandwidth or reduce off-resonance imaging effects in cases when the decay of the transverse magnetization is dominated by T2* (73). In practice, 3T has been used in all published human 13C-pyruvate studies surveyed (Supporting Table S1), and comprises the majority of scanners currently in use for human studies (Table 3). A field strength of 3T is well-suited for 1H MRI anatomical reference and correlative imaging.
Stronger and more rapidly slewing magnetic field gradients support more rapid spatial encoding, particularly for metabolite-specific single-shot imaging using echo-planar imaging (EPI) or spiral imaging (See “Acquisition and Reconstruction”). Although the spatial resolution acquired for HP 13C imaging is typically much coarser than for 1H MRI, the factor of ~4 in gyromagnetic ratio leads to the same reduction factor in performance of the gradient system, so 13C experiments are potentially more limited by gradient hardware performance. To date, all human studies have used the commercially-available integrated gradient systems provided in clinical MRI scanners.
Optimization of scanner design has understandably focused on minimization of artifacts in 1H MRI, where devices such as room lights, the gradient amplifiers, and the motors driving the patient bed are checked to ensure that they do not produce RF interference at the 1H frequency, but artifacts may arise at other frequencies. Eddy current compensation is also not always appropriately adjusted for nuclei at other frequencies (74). In order to optimize for 13C, many sites have performed checks on phantoms for RF interference, gradient artifacts, and eddy currents (74), including the use of post-hoc gradient impulse response function characterisation and correction, and some vendors have fixed these issues as well.
Rf Coils
For HP 13C imaging studies in humans, RF coils for both 1H and 13C nuclei are needed, with 1H MRI providing an anatomical reference for registration and optional additional multiparametric MRI readouts. At the Larmor frequency of 13C nuclei, the relative contributions from coil noise compared to sample noise increase compared to 1H (73,75), although sample noise still is likely the dominant contributor for human-sized coils at 32.1MHz - the resonance frequency of 13C nuclei at 3T.
The key requirement for human 13C-pyruvate RF coils are that the coil geometry and sensitive volume must cover the volume of interest in the subject. Table 2 and Figure 4 shows coil configurations that have been used and optimized for applications in different anatomic regions.
Volume resonators are most commonly used for transmit, as they surround the subject to
Provide B1 Transmit Across The Fov (B1
+). While 1H relies on a large birdcage (“body”) coil built into the scanner, 13C transmit coils must be placed inside the bore. This takes up valuable space within the magnet, and also has led to the use of designs with relatively inhomogeneous
B1
+. Many human studies have used Helmholz pair resonators for transmit, including the “clamshell coil”, which has a notably inhomogeneous B1
+ Profile But Has Been Used Because Of
relatively easy integration into the scanner bore. B1
+ Variation Results In Variations In The Flip
angles that control the use of the hyperpolarized magnetization and creates errors in common HP metrics (9,76). The exception are head coils, where birdcage designs with highly
Homogeneous B1
+ can be placed around the head while easily fitting inside the bore. As with 1H MRI, higher SNR can typically be achieved by smaller receive coil elements, such as surface coils or phased arrays, and the majority of 13C receive coils used have layouts similar to 1H phased arrays.
RF coil quality control is important to ensure proper functioning of the coils to provide consistent imaging quality, especially with limited natural abundance 13C signal in vivo. It typically involves 1) a physical integrity check of the coil cables and connectors and 2) phantom SNR tests to check the coil’s performance and to monitor it over time (see Phantoms below). An useful reference for RF coil quality control is outlined in the MRI accreditation program of the American College of Radiology (77) and can be adapted for 13C coils.
Notably, configurations for brain and prostate studies used dual-tuned 1H/13C coil designs, which greatly simplify workflow and registration of 1H and 13C images, as no switching of coils is needed.
(1)
Table 2: RF coil configurations reported for human HP [1-13C]pyruvate studies.
Tx = Transmit
coil, RX = receive coil. The commonly used “clamshell” TX coil is a Helmholz pair design. For 1H RF configurations, all used the Body coil for TX unless otherwise noted, and “repositioned” indicates the 13C coil was removed for 1H imaging. One representative reference is listed for each configuration. The RF coil configurations reported in the reviewed papers are shown in Supporting Table S1.
Figure 4: Examples of RF coil configurations used for human HP [1-13C]pyruvate brain studies. (A,B) 13C Clamshell TX (Helmholz pair) and 2× 4-channel paddle RX arrays. (C) 13C Birdcage volume TX and 32-channel RX array (RX array slides into TX coil). (D) 13C Birdcage volume TX and 24-channel RX array, combined with a 1H 8-channel RX array. Image reproduced with permission from Ref (16).
Phantoms
Since hyperpolarized magnetization is non-renewable, phantoms containing 13C nuclei are important to: 1) test the multi-nuclear capabilities of the imaging system, including all parts of the signal excitation and receive chain; 2) perform calibration measurements before a scan with hyperpolarized nuclei; and 3) perform necessary pre-scan adjustments (see “Prescan Calibration” section). The phantoms currently in use are listed in Table 3. Their composition must provide sufficient 13C signal, with additional considerations of conductivity, stability, chemical shift(s) present, potential for dynamic imaging, and cost. The phantom geometries are typically either compact, in order to be used alongside the subject during a HP scan, or large enough to mimic the inner volume of a RF coil for system testing.
One popular compact design contains enriched 13C-urea at high concentration, typically 8 M, which provides a single resonance, placed inside a small container ~1 mL. The most common recipe mixes 13C-urea in a 90% water/10% glycerol solution, with glycerol used to increase the urea solubility and doping with a Gd-based contrast agent to shorten T1 which increases the potential SNR per unit time. For example, when Dotarem is added at a 3:1000 volume ratio the 13C-urea T1 is around 500 ms and T2 is around 100 ms. However, when testing pulse sequences influenced by T1 and T2, doping should be used carefully. This phantom is suitable for frequency calibration, transmit gain calibration, sequence testing, and as a fiducial marker when placed next to a patient. However, enriched 13C-urea has a relatively high cost compared to natural abundance compounds.
For larger volumes (>100 ml), the phantoms most often used contain undiluted ethylene glycol, glycerol, or dimethyl silicone. These compounds have sufficiently high carbon concentrations to provide sufficient 13C signal even with the 1.1% natural abundance of 13C. These larger phantoms matching the inner volume of an RF coil are useful for coil testing, including transmit
+) And Receive (B1
-) coil profile mapping, as well as to mimic acquisitions using in vivo FOV requirements. In this case, size and conductivity should match the expected subject size in order to mimic coil loading and get a realistic estimation of B1+. Large-volume natural abundance urea phantoms have also been used by some sites, but suffer from higher conductivity compared to biological tissues. Typically, it is easier to increase the conductivity and hence coil loading of the non-conductive phantom by adding NaCl to match physiological loading (16,78).
Dynamic phantoms that aim to mimic metabolite kinetics have also been developed (79–81), and have the potential to more closely mimic the HP experiment, but so far these are not widely used.
Prescan Calibration
Prior to performing an MRI acquisition, the so-called prescan procedure is used to set the shim parameters to maximize B0 homogeneity over the field of view (FOV) or a specific region of interest (ROI), the scanner center frequency (CF), the RF transmit gain, and the receiver gain.
While this calibration procedure is usually automated for 1H, the lack of sufficient natural abundance 13C signal prevents use of automated methods. (Although natural abundance 13C lipid signal has been detected, there are so far no reports on using this signal for prescan.) Table 3 shows current practices across sites.
Maximizing B0 homogeneity is independent of the nucleus and is therefore performed prior to 13C imaging using the 1H water signal and existing shimming tools, such as by a standard automated process (“Auto Shimming”) or using high order shimming routines. Similarly, the 13C CF can be calculated from the 1H CF using a predetermined scaling factor that depends on the target chemical shift (82). Another common approach used is to have a small, high-concentration 13C phantom, e.g. 8M 13C-urea, integrated in the RF coil or placed next to the scan subject (1). The reference frequency can also be based on real-time measurements after the HP injection but prior to imaging (83). Both the CF and B0 shimming are critical when using spectrally-selective RF pulses, as inmetabolite-specific imaging methods, where the desired excitation bandwidths are typically very narrow and frequency offsets can lead to a failure mode that is only apparent after injection.
The calibration of the RF transmit power is typically performed on a small, high-concentration 13C phantom placed near the region of interest during the scan or on a large 13C phantom of similar size and coil loading as the subject, prior to the subject scan. Reference power is often done by sweeping the power in a pulse-acquire sequence (53,62), or the Bloch-Siegert method (52,84). When using a small phantom, the location of the phantom, B1
+ Inhomogeneity As Well
as any shielding effects, e.g., when the phantom is integrated into a coil (1), may degrade the accuracy. Other methods include real-time Bloch-Siegert method measurements after the HP injection (83), and using the stronger natural abundance 23Na signal that is close enough to the 13C resonance frequency to be detected by 13C coils (82).
The receiver gain is predetermined, either systematically based on independent phantom measurements and assuming the dose and polarization of the HP compound is known prior to injection, or based on past HP imaging studies.
Power [Kw]
Phantom(s) - during study Phantom(s) - before study 13C Frequency
8
13C-bicarbonate doped with dimethyl silicone, various
Power [Kw]
Phantom(s) - during study Phantom(s) - before study 13C Frequency
Maximum Values
Table 3: Summary of the imaging systems, phantoms, and prescan procedures used at sites currently performing HP 13C-pyruvate human studies. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. *Previously performed studies with a Siemens 3T Tim Trio. The imaging systems, phantoms, and prescan procedures reported in the reviewed papers are shown in Supporting Table S1.
Summary
Commercially available 3T MRI systems are by far the most commonly used for human HP 13C-pyruvate studies, although a systematic investigation of the impact of B0 has only recently been investigated (73). The multi-nuclear RF transmit and receive chain has proven sufficient for current acquisition strategies, although many sites have observed artifacts due to RF interference, gradient interference, and residual eddy currents when operating at the 13C frequency. A variety of 13C RF coils, tailored for numerous anatomical targets, have been successfully demonstrated, with the main limitation that most transmit coils take up a lot of additional space inside the bore and provide relatively inhomogeneous B1
+ Profiles. The
phantoms used have converged into generally 2 categories - small phantoms containing 13C-enriched compounds that can be used during the study and human-sized phantoms containing compounds with high carbon concentrations but without 13C enrichment that are used to test and calibrate the coils. There are no standardized compositions or geometry, and dynamic phantoms that recapitulate in vivo kinetics would be desirable but are still an emerging area. Prescan calibration procedures were not well defined in most publications, so we surveyed individual sites to determine current practices. Calibration procedures for the B0 field (13C CF and shimming) for most sites take advantage of 1H signal and methods, while methods
For Calibration Of B1
+ is more variable across sites, likely a reflection of remaining challenges in how to perform this calibration. Standardization of both phantoms and calibration procedures would synergistically improve the robustness and reproducibility of HP 13C studies.
Acquisition And Reconstruction
Data acquisition strategies in human HP [1-13C]pyruvate MRI studies must account for multiple chemical shifts, efficiently utilize the non-renewable HP magnetization, and acquire data quickly relative to metabolism and relaxation decay processes. These studies require spectral encoding to separate metabolites, necessitating pulse sequences that efficiently encode up to 5D data (3 spatial + 1 spectral + 1 temporal dimension). RF pulses must efficiently sample without immediately saturating the non-renewable HP magnetization, and sequences must acquire data quickly and be robust to both experimental and physiologic variation (e.g. B1
+ Inhomogeneity,
variation in perfusion) to ensure reproducibility and minimize scan-to-scan variability. This section covers current successful practices for data acquisition in human [1-13C]pyruvate studies, and accompanying 1H imaging, from different anatomic regions, including scan parameters and image reconstruction.
Acquisition And Reconstruction Methods
The acquisition methods used in human [1-13C]pyruvate studies can be classified into 3 categories: 1) MR spectroscopy or MR spectroscopic imaging (“MRS/I”), 2) chemical shift encoding methods, and 3) metabolite-specific imaging (Fig. 5).
Mrs/I Methods Specifically
resolve a spectrum that can be analyzed to extract expected as well as unexpected resonances, making this approach very robust. It was used in many initial studies (1).
Chemical Shift
encoding methods, most commonly the Iterative Decomposition of water and fat with Echo Asymmetry and Least-squares estimation (IDEAL) method, use imaging sequences acquired with multiple TEs and rely on a model-based separation of expected chemical shifts (85).
Metabolite-specific imaging methods use specialized RF pulses that are spatially and spectrally selective to excite individual metabolites which are then typically imaged with fast k-space trajectories such as echo planar imaging (EPI) or spirals (86).
Their Application To Different
organ systems is described below. The image reconstruction methods used in human [1-13C]pyruvate studies have typically been conventional methods (e.g. FFT, non-uniform FFT, or equivalent). The incorporation of accelerated imaging and advanced reconstruction methods including parallel imaging (4,57,87) and compressed sensing (7) has also been applied in human studies for improved spatial resolution, temporal resolution and coverage, but have the potential for additional artifacts as well as SNR losses due to ill-conditioning of the reconstruction (e.g. g-factor).
The Majority Of
published studies do not use accelerated imaging indicating the resolution and coverage achievable without acceleration is currently adequate for successful data collection. Performing coil combination, even with fully sampled data has also been shown to have specific challenges for HP human images: using naive sum-of-squares methods suffer from high noise amplification in the relatively low SNR regime of HP [1-13C]pyruvate (compared to 1H), motivating several HP 13C-specific methods that include data-driven coil sensitivity estimation which have shown obvious improvements over sum-of-squares (11).
More recently denoising techniques have been applied as post-processing of human HP data(41,42,44). The techniques applied are based on spatial-temporal singular value decomposition for unsupervised estimation of signal and noise components. They have shown improvements in apparent SNR in the brain and liver, while care must be taken to choose parameters such as the rank threshold to avoid oversmoothing and overfitting to the estimated signal components.
Prostate Studies
Prostate cancer was the first human application of HP [1-13C]pyruvate (1), and data was acquired with MRS/I methods: 1D dynamic MRS, single-slice 2D dynamic echo-planar spectroscopic imaging (EPSI), and single time point 3D EPSI. Advances in imaging strategies led to the development and application of new acquisition schemes, including undersampled 3D EPSI with compressed-sensing (7), model-based chemical shift encoding methods that use a priori information (47,59), and metabolite-specific EPI (10), all of which can provide volumetric whole-organ coverage and dynamic acquisitions.
The pyruvate bolus arrival in the prostate can vary by ± 10 s between patients, necessitating dynamic imaging to reliably and consistently capture the pyruvate bolus (18). For this reason, all currently ongoing studies acquire dynamic data. While MRS/I, chemical shift encoding, and metabolite-specific imaging can all achieve dynamic imaging, chemical shift encoding and metabolite-specific imaging provide greater dynamic and volumetric coverage (85). For scan prescriptions, the FOV is designed to provide full prostate coverage and typically to match the orientation of the anatomic imaging used for registration. Flip angles used in current studies are constant through time, as quantification with a variable-through-time flip scheme is highly sensitive to bolus timing (8) and errors in the RF transmit (B1 +) field (76).
Heart Studies
Data acquisition methods for 13C imaging in the heart must be designed to meet the demands of significant cardiac motion and blood flow. To cope with the periodic cardiac motion, most human heart studies to date used gating to the diastolic window, the longest cardiac cycle interval, which has reduced motion (2,22,28,30,35,36,38,45,52). The duration of the diastolic window limits the available data sampling time, making cardiac acquisitions the most time-constrained of the HP 13C MRI applications. The most common acquisition approach is metabolite-specific imaging with spiral k-space trajectories (2). Their single-shot imaging capability makes these methods particularly robust to motion effects. Furthermore, spiral k-space trajectories provide rapid k-space coverage and relatively benign flow and motion artifacts. The majority of studies have used 2D multi-slice acquisitions, but 3D encoding has also been used successfully (35).
Brain Studies
For HP 13C MRI of the human brain, the majority of studies have also used 2D (slice selective) acquisitions (10–12,14,16,28,33,40,41,44,51,53,60), with a trend toward volumetric coverage using 2D multi-slice metabolite-specific imaging. 3D metabolite-specific imaging of the whole brain, with phase encoding of the slice direction (34,57), has been shown to provide similar SNR efficiency (88) compared with multislice imaging. A number of studies have employed MRS/I (5,6,29,31–33,50,55) resulting in a spectrum from each voxel, which has the advantage of not requiring a priori information about which peaks to encode. This was important in early brain studies when it was not known which peaks would be detectable. Chemical shift encoding, using a set of images with different echo times and an iterative reconstruction of the individual resonances (i.e. the IDEAL approach (85)), has also been used (12,49,54), with the drawback that coverage in the slice direction was limited due to the time required to acquire multiple echo time images.
Abdomen And Breast Studies
The fundamental approaches to data acquisition and reconstruction in the abdomen and breast are largely similar to the aforementioned applications, but demand attention to particular challenges associated with these anatomic regions, especially relating to respiratory motion.
Although it has been shown that a basic 2D MRSI approach based on phase encoding and FID readout can be successfully applied for HP 13C imaging in breast (15) and kidney (13), major advantages in terms of spatiotemporal resolution and coverage have been realized using tailored approaches based on metabolite-specific imaging (43,62) and chemical shift encoding (43), which have facilitated multi-slice or 3D dynamic acquisitions over large FOVs in the abdomen (4,37,46).
The significant respiratory motion encountered in these regions can directly blur 13C images, and has further favored these rapid acquisition strategies. Motion also degrades B0 homogeneity, which can shift frequency-selective excitation profiles and introduce artifacts into rapid imaging readouts. This makes accurate determination of the acquisition center frequency and shimming essential in these regions which often cover large FOVs. (See “Prescan Calibration” section for more information). In some studies, breath-holding was used to minimize motion effects and enforce frame-to-frame data consistency (42). A pragmatic and reasonably effective approach for dealing with respiratory motion during 13C data acquisition is an initial breath-hold (as long as can be tolerated), followed by free-breathing (46,62).
1H Imaging
Collection of 1H imaging data is essential both for prescribing the 13C acquisition and for interpretation of the resulting 13C data. Multi-planar 1H scouts are acquired prior to 13C acquisition to enable graphical prescription of the 13C imaging region. All human HP 13C-pyruvate imaging studies acquire conventional MRI scans (e.g. T1- and T2-weighted volumes) for anatomic reference, aiming to cover at least the full 13C FOV. Acquiring these anatomic scans as close as possible to the time of 13C imaging (immediately before or after) minimizes potential misregistration between the data sets. Depending on the application, other advanced 1H sequences are also acquired (e.g. diffusion-weighted imaging for cancer imaging).
When contrast-enhanced data is acquired, it is done after 13C imaging, as paramagnetic contrast agents will accelerate 13C relaxation.
Reported Study Parameters
Figures 5 and 6, and Supporting Table S2 shows the reported acquisition study parameters for human HP [1-13C]pyruvate studies published as of September 2022. Figure 5 shows a mixture of MRS/I, metabolite-specific imaging, and chemical shift encoding methods have been successfully used, where spectroscopy-based methods have become less prevalent in recent studies. Figure 6 shows the acquisition timing, including the important start time and interval/temporal resolution, is quite variable across studies.
Figure 5: Acquisition methods used in published HP [1-13C]pyruvate human studies published up to September 2022, classified into: MR spectroscopy and spectroscopy imaging (MRS/I); chemical shift encoding methods, such as IDEAL, that use multiple TEs and model-based reconstructions; and metabolite-specific imaging methods that use spectrally-selective excitation to image a single resonance at a time.
Figure 6: Temporal acquisition characteristics reported in HP [1-13C]pyruvate human studies published up to September 2022. (a) Reported referencing of acquisition start times.
(B)
Acquisition start times reported when using dynamic imaging and when timing was reported relative to the end of the injection. (c) Temporal resolutions. “Not Applicable” indicates dynamic imaging was not used.
Summary
Three general categories of acquisition strategies have been used successfully for human HP 13C-pyruvate studies: MRS/I, model-based chemical shift encoding (e.g. IDEAL) methods, and metabolite-specific imaging methods. These have enabled successful studies in the prostate, heart, brain, abdomen, and breast. Recent studies increasingly have used the imaging-based strategies of metabolite-specific imaging and chemical shift encoding which are the fastest methods, although a heads-to–head comparison between techniques has not been performed.
Metabolite-specific imaging is quite popular because of its speed and compatibility with single-shot imaging, but is sensitive to B0 field variations and thus requires careful calibrations. Nearly all studies surveyed acquired data dynamically, allowing measurement of the bolus and metabolite kinetics. The exact timings and associated flip angles vary quite widely across reported studies, with no consensus yet as to how to choose these parameters. Image reconstruction is typically done directly using Fourier Transform methods, and accelerated imaging strategies are uncommon.
Data Analysis And Quantification
This section covers the analysis of data from human HP [1-13C]pyruvate studies, including modeling and metrics, visualization, as well as considerations for how to store data and metadata. Depending on study design, the analysis may need to give quantitative or semi-quantitative output reflecting a biological process or may just reflect a contrast between different regions of interest for quantitative evaluation.
Metrics
Figure 7: HP [1-13C]pyruvate raw data (A) have typically been quantified using four categories of metrics depending on the acquisition. Data acquired as a single time point are often quantified using normalized metabolite images or metabolite ratios (B). Dynamic data can be quantified using normalized metabolite images or metabolite ratios (B), or with metabolite timings such as time-to-peak (TTP) or pharmacokinetic (PK) models (C). The latter two require the data to be time-resolved. [1-13C]alanine and 13C-bicarbonate are analyzed similarly to [1-13C]lactate but omitted here for display.
Metabolite images are commonly used as summary metrics for HP MRI data, often including some form of normalization as well as summed over time as an area under the time curve (AUC) (17). These are analogous to the visual evaluation that is most used for routine clinical work (89,90). In these metabolite images, we expect that the [1-13C]pyruvate AUC signal is predominantly weighted towards perfusion and uptake, while [1-13C]lactate, [1-13C]alanine and 13C-bicarbonate AUCs represent metabolic conversion. The strength of this approach lies in its simplicity and relatively few underlying assumptions. Limitations to the use of single-metabolite images or AUCs include sensitivity to inhomogeneous coil profiles (57,87,91), the acquisition strategy and acquisition parameters, pyruvate polarization and concentration level, and signal relaxation rates (92). Further, the reader must be careful to interpret all the images in conjunction to better understand the underlying biology; for example, increased [1-13C]lactate in the presence of decreased [1-13C]pyruvate delivery can have a very different meaning compared to increased [1-13C]lactate with increased [1-13C]pyruvate delivery.
In an attempt to address variations in coil sensitivity, polarization level, and pyruvate delivery, AUC images are often computed by normalizing to a specified parameter, such as the maximum pyruvate or average lactate signals, or presented as a ratio such as lactate/pyruvate or divided by “total Carbon” - the sum total of HP 13C signal observed across all metabolites. The AUC ratios between metabolites and pyruvate are proportional to the corresponding forward kinetic rates (81,93), but are not directly comparable to rate constants when magnetization loss rates (e.g. relaxation and losses due to signal excitation) differ between studies. Similarly, the ratios between the produced metabolites (e.g. bicarbonate/lactate) can reflect the balance between downstream metabolic pathways (12,55). Care must be taken to consider how AUC images are calculated and normalized before comparing values between studies.
To further quantify the interpretation, pharmacokinetic (PK) modeling approaches were developed to compute the apparent kinetics of pyruvate-to-metabolite exchange (92,94–99). These yield semi-quantitative to quantitative apparent rate constants, given in s-1. Some models require a vascular input function, while others avoid this requirement (95). PK models can explicitly account for acquisition-specific details such as excitation angle and repetition time, and thus may reduce the effects of these details on quantification. An input-less model, provided in the Hyperpolarized-MRI-Toolbox (https://github.com/LarsonLab/hyperpolarized-mri-toolbox) (100) and thus frequently employed for human data, has been shown to fit well and robustly to prostate and brain data (8,20). PK models are quantitative in nature, arguably provide more relevant biological information (8,20), and appear to be reproducible across sites (51). However, rate constants derived from PK models are still apparent rates, and likely do not reflect a single biological characteristic.
Some additional considerations include whether complex or magnitude data is used, as the noise behaviors will impact the analysis differently. Additionally, cut-off thresholds or other criteria may be used to identify and avoid voxels with insufficient SNR before analysis to improve robustness (20,41).
Regardless of the analysis approach, the underlying biology is not always clearly represented by the data; instead, the metrics may be influenced by perfusion, barrier permeability, intercellular shuttles, enzyme activities, co-substrate concentrations, or combinations thereof, depending on the organ and disease of interest (19,43,94,101–103). This may be addressed by incorporating complementary information. As an example, HP 13C pyruvate data is influenced by perfusion, and thus addition of perfusion MRI could be important for interpretation (98,104,105).
All the methods outlined above have been explored in clinical studies, described in Supporting Table 3 and summarized in Figure 8. As of September 2022, approximately 52% of studies involving human subjects report rate constants derived from a PK model with a few different models reported. A nearly equal fraction (51%) of the studies report AUC ratio values.
Approximately 66% of these studies report metabolite-specific images or AUC values. About 40% report SNR values; this metric is particularly frequent in manuscripts that describe technical developments for clinical HP MRI. Approximately 16% of these studies summarize model-free metrics, and 10% report measurements from a single timepoint. Most studies report a combination of quantities.
Figure 8: Reported metrics used for analysis in HP [1-13C]pyruvate human studies published up to September 2022.
Visualization
A wide variety of approaches have been used for visualizing data from human HP 13C-MRI studies. The challenges and practical considerations are: 1) choosing the appropriate metrics to display, 2) how to encode the parameters (e.g. the colormap), and 3) choosing how to provide anatomical context and other multi-parametric data. The choice of visualization also depends on the goal which could be for diagnostic interpretation, but also quality control, reproducibility among readers and publication.
Metrics
The choice of HP 13C metrics is described in detail above. At this stage in HP 13C development where there is no standardized metric, often a combination of metabolite images and ratios or PK model parameters are shown.
Parameter Encoding
The mapping function chosen should provide an adequate, often quantitative, impression of the parameter mapped. There is a consensus in the visualization field that perceptually uniform maps are best suited to visualize continuous parameters, like the greyscale typically used by radiologists as well as other monochrome (black to blue) and color ranges (fire-type, rainbow-type) (106,107). Multi-color heatmaps have been the most frequently employed method for HP 13C data, while greyscale has infrequently been used but it ensures there is no coloring-based bias as well as facilitating later reuse (Fig. 9a). Among the color schemes employed in the clinical HP 13C literature, fire-type scheme seems to be the most common [similar to “Plasma” or “Inferno” in matplotlib.org]. Next most commonly employed is the rainbow-type scheme [similar to “Rainbow” in matplotlib.org].
Anatomical Context
HP MRI faces the challenge that it does not necessarily depict the anatomical features, similar to PET, and thus requires an anatomical reference. Most often, a grayscale anatomical image is overlaid with a HP colormap (Fig. 9c,d). This approach is very intuitive, but can skew perception as the grey-scale anatomical reference may affect the brightness of the HP data (e.g. signal in the skull). This bias does not occur when showing adjacent maps (Fig. 9a, b). Here, anatomical outlines may help to provide reference (Fig. 9b).
Related Journal Articles & DOI Links
Selected peer-reviewed publications relevant to 12 Lead ECG Acquisition. Click the DOI to access the full paper (may require institutional access).
-
1. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
IEEE Journal of Biomedical and Health Informatics
https://doi.org/10.1109/JBHI.2020.2981234 -
2. Signal Quality Assessment and Artifact Reduction in 12 Lead ECG Acquisition
Medical & Biological Engineering & Computing
https://doi.org/10.1007/s11517-020-02145-6 -
3. Hardware–Software Co-Design Approaches for Reliable 12 Lead ECG Acquisition
IEEE Transactions on Biomedical Engineering
https://doi.org/10.1109/TBME.2019.2895762 -
4. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
Frontiers in Bioengineering and Biotechnology
https://doi.org/10.3389/fbioe.2020.00123 -
5. Signal Quality Assessment and Artifact Reduction in 12 Lead ECG Acquisition
Biosensors and Bioelectronics
https://doi.org/10.1016/j.bios.2021.112345 -
6. Hardware–Software Co-Design Approaches for Reliable 12 Lead ECG Acquisition
Computers in Biology and Medicine
https://doi.org/10.1016/j.compbiomed.2021.104567 -
7. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
Nature Communications
https://doi.org/10.1038/s41467-020-12345-6
Why Choose Us?
Bangalore guidance for robotics, Spectre and autonomous systems projects.
Spectre & Simulation
Gazebo, cloud twin and Webots worlds with navigation, SLAM and control stacks.
Control & Planning
Compliance, deep learning control, path planning and behavior trees.
Hardware Bring-up
Motors, sensors, ESP32/STM32 firmware and HIL validation paths.
Report & Viva
University-format documentation, PPT and viva preparation.
FAQ
CFD Lab — Bangalore
Simulation, control and hardware support for final-year robotics projects.
Stacks
Worlds
Digital Twin
Control
Robots
Offline
Bring-up