Enquire Now
70+ Topics · Spectre · Spectre · cloud sim Sim · MATLAB · Webots · Hardware · Bangalore 2026

Visual Slam Robot

Simulation · Control · Perception · Hardware — 12 Lead ECG Acquisition — hardware, sensors, cloud dashboards and protocols (Spectre, REST, CoAP, WebSockets) for BE BTech MTech students. Final-year robotics support with Spectre stacks, simulation worlds, reports and viva from Bangalore.

70+
Related Topics
6+
Sim & HW Tools
4.9★
573 Ratings

OKVIS2-X: Open Keyframe-based Visual-Inertial SLAM

Onfigurable With Dense Depth Or Lidar, And Gnss

Simon Boche∗, Jaehyung Jung∗, Sebasti´an Barbas Laina∗, and Stefan Leutenegger Abstract—To empower mobile robots with usable maps as well as highest state estimation accuracy and robustness, we present OKVIS2-X: a state-of-the-art multi-sensor Simultaneous Localization and Mapping (SLAM) system building dense volu- metric occupancy maps, while scalable to large environments and operating in realtime. Our unified SLAM framework seamlessly integrates different sensor modalities: visual, inertial, measured or learned depth, LiDAR and Global Navigation Satellite System (GNSS) measurements. Unlike most state-of-the-art SLAM sys- tems, we advocate using dense volumetric map representations when leveraging depth or range-sensing capabilities. We employ an efficient submapping strategy that allows our system to scale to large environments, showcased in sequences of up to 9 kilometers. OKVIS2-X enhances its accuracy and robustness by tightly-coupling the estimator and submaps through map alignment factors. Our system provides globally consistent maps, directly usable for autonomous navigation. To further improve the accuracy of OKVIS2-X, we also incorporate the option of performing online calibration of camera extrinsics. Our system achieves the highest trajectory accuracy in EuRoC against state- of-the-art alternatives, outperforms all competitors in the Hilti22 VI-only benchmark, while also proving competitive in the LiDAR version, and showcases state of the art accuracy in the diverse and large-scale sequences from the VBR dataset. Code available at: https://github.com/ethz-mrl/OKVIS2-X.

visual-slam-robot Diagram
Figure: System Model & Architecture for Visual Slam Robot

Index Terms—SLAM; Mapping; Localization; Sensor Fusion

Ntroduction

While autonomous robots and mobile devices are becoming more ubiquitous, their ability to perceive and spatially relate to their environments remains crucial. To cope with potentially unkown environments, certain such applications demand that the system can simultaneously map the environment while also localizing in it.

visual-slam-robot Diagram
Figure: System Model & Architecture for Visual Slam Robot

Such Simultaneous Localization and Mapping (SLAM) sys- tems typically require heterogeneous sensors for enhanced accuracy and robustness, with visual-inertial configurations being one of the preferred combinations of sensors due to its appealing compromise between cost and accuracy. The setup combines spatial information from the camera with proprioceptive temporal information from the inertial mea- surement units (IMU), whereby rendering the gravity direction observable when tightly coupled and reducing dependency on This work was supported by the EU Horizon Europe program under grant agreement 101070405 (DigiForest) and 101120732 (AUTOASSESS).

visual-slam-robot Diagram
Figure: System Model & Architecture for Visual Slam Robot

Simon Boche, Jaehyung Jung and Sebasti´an Barbas Laina are with the Mobile Robotics Lab, School of Computation, Information and Technol- Stefan Leutenegger is with the Mobile Robotics Lab, ETH Zurich, 8092 ∗Equal contribution.

visual-slam-robot Diagram
Figure: System Model & Architecture for Visual Slam Robot

Fig. 1. 3D reconstruction from a run of OKVIS2-X on the Spagna sequence of the VBR dataset . Reconstruction with a LiDAR sensor (top) or with a depth network (bottom) to showcase the versatility of the presented system to different sensor modalities. The estimated trajectory is visualized in black.

visual-slam-robot Diagram
Figure: System Model & Architecture for Visual Slam Robot

Furthermore, different colors per submap are used. brittle visual data associations. For applications that demand a higher mapping and localization accuracy, LiDARs have emerged as a powerful option due to their unmatched depth sensing resolution and accuracy, at the expense of cost. To improve the robustness of SLAM systems, there has been a recent increase of interest in systems that can combine these three sensor modalities, whereby increasing the challenge of obtaining an accurate calibration of the extrinsics between all the different sensors, a fundamental requirement for accurate positioning, GNSS is the obvious choice for outdoor appli- cations, if reliably received – which cannot be guaranteed in all environments.

visual-slam-robot Diagram
Figure: System Model & Architecture for Visual Slam Robot

Most of the visual-inertial SLAM (VI SLAM) systems proposed in the literature – only use a small subset of the camera information to build a sparse map, therefore lack- ing relevant geometric information required for downstream tasks. To address this limitation, the community has proposed

▪

Relocalization & loop-closure init.

Epth + Uncertainty

Fig. 2. System architecture of our proposed multi-sensor state-estimator OKVIS2-X. Components with a grey background correspond to original elements from OKVIS2 and components with a white background are the extensions for the multi-sensor setup. systems that can leverage visual depth information –

Or Lidar Information – To Build Denser Maps That

can then be reused for downstream tasks – thereby naturally increasing memory requirements. The underlying map repre- sentation plays a crucial role not only regarding memory, but perhaps more importantly to facilitate downstream tasks. We argue that volumetric occupancy maps present an unmatched characteristic in enabling safe path planning and navigation by explicitly representing free space while the popular point clouds (or other commonly employed representations such as meshes) do not.

visual-slam-robot Diagram
Figure: System Model & Architecture for Visual Slam Robot

To address the aforementioned challenges, we propose a multi-sensor SLAM system called OKVIS2-X, a substantial extension to OKVIS2 . The key idea is to leverage submap- based dense volumetric occupancy maps which are tightly- coupled to the state estimator via submap alignment fac- tors. Examples of the 3D reconstruction obtained from our framework are presented in Fig. 1. Furthermore, we present a unified framework to fuse visual, inertial, LiDAR or depth from neural networks, and the global position from GNSS.

visual-slam-robot Diagram
Figure: System Model & Architecture for Visual Slam Robot

Our previous works have augmented OKVIS2 for different sensor modalities: visual-inertial-GNSS , visual-inertial- LiDAR , and visual-inertial-depth . But to the best of the authors’ knowledge, here we present for the first time a unified and configurable multi-sensor SLAM system based on a volumetric map representation, supporting different and novel combinations of sensing modalities (e.g. visual-inertial- LiDAR-GNSS and visual-inertial-depth-GNSS) through a fac- tor graph framework. In a further novel contribution, we seamlessly support online-calibration of camera extrinsics.

visual-slam-robot Diagram
Figure: System Model & Architecture for Visual Slam Robot

Fig. 2 describes the system architecture of OKVIS2-X that is configurable with multiple sensor modalities. The resulting system demonstrates unmatched accuracy and robustness in localization and mapping, as well as scalability in building dense volumetric maps from meter to kilometer-level – which we demonstrate in an extensive series of entirely new evalua- tions. The main contributions of our proposed method are: • We present a state-of-the-art multi-sensor SLAM system able to fuse visual, inertial, GNSS and dense mapping

Factors From Either Lidar Measurements Or A Depth

network in realtime. • Our method tightly couples the state estimator with volumetric occupancy mapping, rendering it suitable for large-scale SLAM thanks to our submap representation.

visual-slam-robot Diagram
Figure: System Model & Architecture for Visual Slam Robot

• Our method supports online calibration of the camera- IMU extrinsics in the visual factors including relative • We perform a thorough evaluation in benchmark datasets including large-scale (up to 9 km) scenarios, where we show state-of-the-art results in VI and VI-LiDAR SLAM.

• We will release our framework as open source to foster research in the field of multi-sensor SLAM.

Related Work

We review the most relevant works under the topics of multi-sensor SLAM, dense SLAM, and submap-based SLAM.

A. Visual-Inertial Slam

VI SLAM systems are generally divided into two main coupled estimators typically compute (relative) poses from images alone first before combining with IMU measurements, while the tightly-coupled alternatives consider all measure- ments simultaneously, typically through IMU kinematics in- tegration with visual observations, e.g. in the form of re- projection errors. The latter approach is known to be more accurate. A second distinction of estimator comes from either recursively estimating a selection of recent states employing filtering, e.g. versus optimizing typically a window of states in a nonlinear least-squares fashion, e.g. , . On the other hand, depending on the type of residual, direct methods minimize the photometric error, while the keypoint-based methods such as , , , minimize the reprojection error to estimate robot poses and sparse landmarks.

Nonlinear

optimization-based VI SLAM by fusing IMU preintegration and reprojection errors. OpenVINS is based on the Multi- State Constraint Kalman Filter (MSCKF ), and extends this with long-term SLAM landmarks, extrinsic and intrinsic online calibration. ORB-SLAM3 is a tightly-coupled VI- SLAM system that supports multi-session SLAM with monoc- ular, stereo and RGB-D configurations. It records the state-of- the-art accuracy in public SLAM datasets. OKVIS2 is a multi-camera SLAM system with loop closure, where pose- graph factors are formulated from marginalized landmarks.

MAVIS is also a multi-camera SLAM system with IMU preintegration on matrix Lie Groups. SVIn2 , is a visual-inertial framework that incorporates loop-closures to the factor-graph optimization, while including other sensor modalities for its deployment in underwater environments.

In the context of deep learning-based VI SLAM, Deep- VIO is a self-supervised visual-inertial odometry (VIO) network that directly regresses trajectory from the input sensor measurements. DVI-SLAM learns a confidence map of feature-metric and reprojection factors with a dense bundle adjustment (DBA) layer. DBA-Fusion also utilizes the DBA to tightly fuse visual-inertial measurements. The authors demonstrated dense point cloud mapping result in a large- scale environment. However, the DBA layer, which is a main component in end-to-end SLAM systems, often requires high GPU memory usages (≈24 GB).

In contrast to previous works, OKVIS2-X provides a gen- eral framework to work with multiple cameras and multi- modalities including depth networks, LiDAR, and GNSS re- ceiver. Furthermore, our method goes beyond the classical sparse landmark map representation since it can build a dense volumetric occupancy representation that scales to large scenes.

B. Lidar(-Visual-Inertial) Slam

Due to their inherently accurate depth sensing capabilities, LiDAR sensors have proven to be beneficial for robot local- ization allowing to reduce the drift significantly, which SLAM systems inevitably suffer from. One of the first systems to successfully demonstrate the applicability of LiDAR sensors to SLAM was LOAM . Its core idea was the extraction of salient geometric features, such as edges and planes, which can be reliably tracked across frames. LOAM has inspired an entire line of research adopting the concept of salient feature extraction , . As in VI SLAM, one can further distinguish feature-based and direct methods. Direct LiDAR

Slam Methods, Such As Kiss-Icp Or Ct-Icp ,

typically rely on variants of the Iterative Closest Point (ICP) algorithm on raw LiDAR point clouds, omitting the necessity to extract salient features. , establish sliding-window bundle adjustment (BA) for LiDAR scans for feature-based and direct formulations.

Fusing complementary IMU measurements enables the de- mulates a tightly-coupled pose graph optimization problem fusing edge and plane error residuals with IMU error residuals.

Also filter-based approaches, such as FAST-LIO or FAST- LIO2 have successfully demonstrated high-accuracy lo- calization fusing LiDAR and IMU. In contrast to FAST-LIO, which still uses plane and edge features, FAST-LIO2 is a direct method achieving a significant speed-up by directly operating on the raw points.

Recently, LiDAR-Visual-Inertial SLAM systems have be- come increasingly popular. A large amount of these ap- proaches still make use of LOAM’s idea of geometric feature extraction. LVI-SAM combines a Visual-Inertial (VI) and a LiDAR-Inertial (LI) subsystem to complement each other in challenging scenarios. The LI system builds a factor graph based on IMU preintegration errors and edge and plane residuals. VILENS also builds an optimisation problem consisting of visual, inertial, robot leg odometry, as well as LiDAR-based line and plane residuals. Another line of research uses MSCKF -based approaches for tightly- coupled fusion of visual, inertial and LiDAR measurements, e.g. LIC-Fusion and its successor LIC-Fusion 2.0 .

R2Live Fuses Visual Features, Imu And Lidar Mea-

surements in an Iterated Kalman Filter. Unlike previously

Mentioned Lvi Slam Systems, The Succeeding R3Live

applies direct methods, increasing its robustness in texture-less or structure-less environments. FAST-LIVO and FAST- LIVO2 also formulate an Iterated Extended Kalman Filter (IEKF) framework. The IEKF update step consists of two sequential state updates: a point-to-plane based LiDAR update and a sparse direct visual update operating on image patches instead of dense pixels. Both updates are computed in a frame- to-map fashion using a local voxel map.

In contrast to previous approaches, the proposed OKVIS2-X can fuse LiDAR measurements omitting the necessity to ex- tract salient geometric features or any other correspondences, usually a computationally expensive step. Instead, based on our previous work , residuals can be directly derived from dense occupancy maps which are incrementally constructed and are directly usable for downstream tasks. These residuals are added to a factor graph optimization problem.

Ulti-Sensor Slam With Gnss Fusion

Early work on fusing global position measurements to reduce the drift in VI SLAM is dominated by filter-based methods, mostly building upon the seminal work in where the authors propose an Extended Kalman Filter (EKF) to tackle realtime visual-inertial navigation. and use an EKF and an Unscented Kalman Filter (UKF), respectively, to fuse different sensor modalities, such as LiDAR and GNSS measurements. Additionally, also estimates the IMU- GNSS spatial extrinsics as well as the sensor time offset

Online In An Msckf Framework. Proposes A Robust

Iterated Error State Kalman Filter (IESKF) to fuse GNSS, IMU and LiDAR measurements in a tightly-coupled approach. A filter-based system that not only fuses Lidar-Visual-Inertial information, but also wheel odometry and GNSS is MINS , a system that extends and demonstrates how multi- sensor information improves the robustness of state-estimation.

Filter-based methods are in general computationally efficient. Nevertheless, as investigated in , optimization-based ap- proaches potentially deliver superior results, thanks to their ability of re-linearizing old state estimates.

Ns-Fusion And Gomsf Are Examples For

optimization-based methods fusing global position measure- ments in a loosely-coupled back-end pose-graph optimization. However, a tightly coupled fusion in a unified optimization problem is desirable to fully exploit the available information given by different sensor modalities. offers the integration of GNSS factors into LiDAR-Inertial Odometry as a factor graph optimization problem. proposes a tightly-coupled approach fusing GNSS measurements in the optimization win- dow of a VIO as global factors leveraging IMU preintegration.

– have also proposed tightly-coupled optimization- based frameworks which consider not only global position measurements but also pseudo range and Doppler shift errors from raw GNSS measurements. Unlike the previous, pure

Odometry Approaches, , Successfully Demonstrated

the fusion of global position measurements into VI SLAM including visual loop closures. While assumes global position measurements given in the visual-inertial reference frame, in practice, a 4-Degree-of- Freedom (DoF) transformation between a global and the local reference fame has to be estimated. Initialization of this global frame alignment has been addressed using SVD based on cor- respondences of local and global position measurements or in a tightly-coupled fashion using least-squares minimiza-

Tion , , . Furthermore Introduces An Heuristic

criterion on the observability of the 4-DoF transformation based on the distance traveled. As in and following our previous work , we will use a SVD-based initialization of the GNSS-VIO extrinsics. However, it will be estimated in a tightly-coupled way. In contrast to the aforementioned approaches, to simplify the problem, the global reference frame will be fixed in our approach as soon as it becomes observable. Instead of applying a heuristic threshold as in , we introduce a fully uncertainty-aware criterion to determine the covariance of the estimated GNSS extrinsics. Furthermore, our framework supports not only visual loop-closures but also provides a loop-closure like optimization approach to effectively tackle the issue of drift during longer periods of GNSS outage as presented in . In new experiments, here, we furthermore present how OKVIS2-X supports Lidar- Visual-Inertial SLAM with intermittently available GNSS.

Ense Slam & Mapping

Most of the previously mentioned SLAM systems achieve high accuracy in localization relying only on sparse map representations, such as sparse 3D landmarks or feature maps.

However, understanding dense geometry is indispensable for deploying fully autonomous robots in downstream tasks. KinectFusion pioneered the area of dense SLAM with Truncated Signed Distance Fields (TSDF) fusion and frame-to- model registration. ElasticFusion models an environment with dense surfels, which are updated non-rigidly after a loop.

Nn-Slam And Deepfusion Employed Dense Depth

from a convolutional neural network and fused depths from motion stereo to further improve the depth quality. Kimera and its extension Kimera2 presented the dynamic scene graph built by visual-inertial observations. In their map rep- resentation, semantically annotated meshes are deformed by loop-closure constraints. TANDEM integrated a deep Multi-View Stereo (MVS) network in visual odometry, ad- ditionally rendering depth from TSDF for image alignment.

Simplemapping Adopted Simplerecon And Fused

depths from the MVS network into a TSDF given poses and sparse landmarks from ORB-SLAM3 but without any feedback from mapping. SigmaFusion predicted dense depth uncertainty based on the information matrix from Droid- SLAM and showed less noisy reconstruction through uncertainty integration. However, their method is memory intensive, and mapping and state estimation are decoupled. A new line of work based on geometric foundation models has recently emerged in the SLAM community, with MASt3R-

Slam , A Monocular Slam System, Being One Of The

most notable examples. These systems leverage point maps for both mapping and localization. In the context of dense neural SLAM, where map geometry and appearance benefits from the neural implicit representa- tion, NICE-SLAM represents the map with hierarchical feature grids that are decoded into a volumetric occupancy map. UncLe-SLAM incorporates pixel-wise depth and color uncertainty in tracking and mapping loss functions in NICE-SLAM to properly account for observation uncertainty.

Point-SLAM refrains from a grid-based representation to save memory by assigning neural points only around surfaces that are later used to obtain the map occupancy and color.

Oopy-Slam Extends Point-Slam By Incorporating

loop-closure constraints in the neural point representation by adopting a submap strategy where neural points are anchored at each submap to build a globally consistent map. Afore- mentioned dense neural SLAM methods have shown promis- ing appearance rendering in (multi-)room-sized environments.

However, it is not clear how those lines of works are scalable to km-level large-scale scenarios, which is the main challenge this paper addresses, given the high dimensionality of neural features – or how these representations efficiently support mobile robot navigation.

In contrast, our proposed system supports tight-coupling between the state estimator and the volumetric free-space mapping for generic sensor setups, as presented in our previous

Works For Depth Cameras And Lidar Sensors . The

proposed approach is fully probabilistic, enabling the consid- eration of uncertainties also in downstream tasks.

E. Submapping Approaches

To minimize drift in long-term scenarios, the latest research commonly applies the concept of submapping. This originates from early SLAM research, such as the Atlas framework .

In this context, additional factors can be derived to align submaps and to reduce drift. uses local point cloud submaps and adds ICP odometry measurements into the fac- tor graph. Wildcat , a sliding-window optimization-based LiDAR-Inertial Odometry system, achieves peak state-of-the- art robustness and accuracy by building local surfel submaps and aligning the submaps. These methods use point or surface- based 3D representations. While these can evidently be lever- aged for high-accuracy localization and 3D reconstruction, the map representation is not suitable for navigation due to the difficulty to distinguish between free and unobserved space.

Submapping on volumetric maps has been addressed in

Various Works. Builds Occupancy Submaps Based On

OctoMap and aligns them by standard ICP registration. Voxgraph uses VIO to provide poses for integration into TSDF maps. Upon completion, Euclidean Signed Distance Fields (ESDF) are generated and submaps are aligned in pose graph optimization. New submaps are created at a fixed frequency resulting in a comparably large memory usage.

instead use occupancy maps as their 3D representation. Using Supereight2 , an adaptive-resolution mapping approach, new submaps are spawned based on the distance traveled.

Submaps are re-arranged based on updates from the visual- inertial estimator. The follow-up work improved the submap creation by evaluating the point cloud overlaps of new scans and alignment of submaps is based on ICP.

In this work, we will also adopt the concept of submapping. In contrast to , , , the global alignment of submaps is not decoupled from the estimator but provides direct feed- back. We formulate correspondence-free residuals as in without the expensive need to extract ESDFs as we directly use the available occupancy information.

A. Notation And Definitions

The classic VI SLAM problem formulation includes several different coordinate frames. A moving body is tracked with

By F

−→Ci for i = 1 . . N cameras. Fusing LiDAR requires

−→L And Fusing Gnss

measurements requires to introduce a global reference frame −→G which is a gravity-aligned East-North-UP (ENU) local Cartesian frame located in the global position of the first received GNSS measurement. The submap frame is denoted

−→S When The New Submap Is

declared. The rigid body transformation T AB ∈SE(3) transforms homogeneous points between two frames: ArP = T AB BrP ,

−→A. The

rotational part of T AB is expressed by CAB ∈SO(3) and ArB denotes the translation component. We also denote the rotation CAB with its unit quaternion form qAB.

(1)

where W rS, qW S and W v denote the position, orientation and velocity of the IMU sensor frame in the fixed world frame. bg and ba stand for gyroscope and accelerometer biases, respectively. To fuse the visual-inertial system and global measurements, we estimate the extrinsic transformation

[Grt

W , qGW T ] between the world reference frame of the

−→G. In Case

of online calibration of the camera-IMU extrinsics, the state

T ] For

i = 1 . . N with N being the number of cameras.

Olumetric Occupancy Mapping

Obtaining an accurate map representation is essential for state-estimation and downstream tasks. Most of the state-of-

The-Art Vi Slam Systems , , , Build A Sparse

map based on image features, which is used for state estima- tion but cannot be leveraged for most downstream tasks due to the lack of dense geometric details. To address this, OKVIS2- X leverages dense volumetric occupancy submaps based on the multi-resolution mapping framework Supereight2 , which are also considered in state estimation and can be used in downstream tasks such as autonomous navigation , exploration or place recognition .

In OKVIS2-X, we refrain from the concept of monolithic mapping and we leverage a submapping strategy for environ- ment representation, obtaining two major benefits. First, data is allocated into smaller maps, resulting in a shallower data- structure and consequently reducing the computational cost of integrating the depth information. Second, each submap is an- chored to individual keyframe states from the state estimator, enabling submap poses to be adjusted during updates.

−→M Is Anchored

to a keyframe state xk, and its body pose T W Sk expresses the

−→W . Upon An Update Of Xk,

the new body pose T W Sk updates the rigid transformation be-

−→W , Ensuring The Local Consistency Between

submaps and improving the overall mapping accuracy. Each submap maintains occupancy log-odds which is updated with either LiDAR point clouds or depth images D. The log-odd of

(2)

where Pocc is the occupancy probability. As in Supereight2, we perform additive Bayesian updates of the log-odds occupancies as we integrate new depth information. The mean of log-odd

(3)

where wk is the number of observations and saturated in wmax. This update alleviates the geometric inconsistencies from non- accurate depth perception, more critical when our system leverages learned depth, while also increasing the robustness to the dynamic entities that are perceived from the scene.

The assumption of our submap strategy is that drift within a submap is negligible, therefore to generate a new submap two criteria have to be met. First, a minimum number of depth measurements has been integrated into the current submap.

Second, the overlap between the current depth or LiDAR measurement and the previously completed submap is lower than a threshold or that a maximum of K keyframes have been generated since the creation of the current submap. The keyframe criterion is a proxy for the drift within a submap and the overlap criteria ensures that submap-based alignment constraints can reliably be added to the state estimator.

By leveraging occupancy probabilities as a map represen- tation, the environment can be classified as free, occupied or unobserved, a distinction that becomes essential for safe autonomous navigation. Similar to , the path planner only traverses areas considered free in a submap, since it considers unknown areas as non-traversable, and these paths are elastically deformed, according to the submap updates.

Supereight2 uses a piecewise linear inverse sensor model for the log-odds occupancy as a function of the mea- sured depth. Hereby, a log-odds value of zero represents the measured surface position. Moreover, the depth uncertainty σ can also be modeled as a function of the measured depth zr.

In our work, we assume a linearly growing uncertainty for the LiDAR sensor and quadratically for RGB-D cameras. However, these heuristic uncertainty models are not suf- ficient for learned depth. For instance, textureless regions or object boundaries yield high depth uncertainty due to the ambiguous correlation volume or viewpoint changes. In response, we will introduce a depth fusion method to predict the pixel-wise uncertainty in the following section.

A. Uncertainty-Aware Depth Fusion

It is pivotal to obtain a reliable dense depth as well as the associated uncertainty for our downstream tasks. We follow the key ideas presented in , our previous work, where depths from static and motion stereo are fused, which are complementary to each other — static stereo provides a reliable small baseline even when stationary, while motion stereo potentially brings a large baseline, depending on the camera motion. Our method does not depend on a specific

Network, But We Found That Unimatch And The Mvs

network work well in real-world scenes. However, the previous networks only predict disparity or depth without any

[M−1]

Fig. 3. Predicted inverse depth (top) and its corresponding standard deviation (bottom) of (a) stereo network with the 11 cm baseline, (b) MVS network with the 50 cm maximum baseline among 8 views, and (c) depth fusion in

The Euroc Dataset. (Adopted From .)

uncertainties. Therefore, we augment the base architecture with an uncertainty decoder and adopt the Laplacian loss function for (aleatoric) uncertainty learning. Specifically, our

(4)

where θ is the network weights, and u stands for the disparity. The disparity loss Lu, modeled in the Laplacian distribution

(5)

where T is a training set including pairs of stereo images and the ground-truth disparity. We additionally add the gradient loss L∇u for sharper uncertainty output, which is analogously defined as Lu. The gradient uncertainty along the horizontal and vertical directions is derived from the disparity uncertainty

P

σ2u(x, y + 1) + σ2u(x, y −1).

(6)

We propagate the uncertainty from the disparity to the depth

(7)

where fc is the rectified focal length, b is the stereo baseline. Likewise, we modify the loss function of the MVS net-

(8)

where the network learns log-depth uncertainty. We transform

(9)

Given the pixel-wise estimates from the networks ( ˆdst, σst), ( ˆdmvs, σmvs) and with the assumption that two estimates are independent, we can optimally fuse two depth estimates ,

(10)

Our stereo and MVS networks with additional uncertainty decoders are only fine-tuned in a synthetic dataset . Fig. 3 shows an example where the fan stand has more detailed depth in the stereo network, while farther objects are sharper in the MVS network. Fused depth naturally maintains the optimal depth based on uncertainty-aware fusion.

A. System Overview

OKVIS2-X proposes a modular multi-sensor SLAM sys- tem, specifically designed for large-scale scenarios and robot navigation. It builds on top of the sparse VI SLAM system OKVIS2 and adds capabilities for online camera-IMU extrinsics calibration, fusion of global position measurements and dense map alignment factors derived from LiDAR or depth sensing modalities. Fig. 2 visualizes the overall system architecture and how it extends the original OKVIS2.

The underlying VI SLAM system is split into the visual frontend and a realtime estimator that process images and IMU messages synchronously whenever a new (multi-) frame arrives. To deal with loop-closures, a full factor graph loop optimization is executed asynchronously. As shown in Fig. 2, the frontend deals with state initialization, keypoint matching, stereo triangulation (of successive frames and from stereo im- ages of the same multi-frame), if enabled, running a segmenta- tion CNN (to filter out observations in the sky), as well as place recognition and, if the latter was successful, relocalization and loop-closure initialization. The realtime estimator will then optimize the respective factor graph, and is also responsible for the creation of posegraph edges by marginalizing old observations, as well as for fixation of old states. Upon loop- closure, it turns posegraph edges back into observations, and then triggers the optimization of the full graph around a loop, which runs asynchronously, and which will be synchronized with the realtime factor graph upon completion.

OKVIS2-X extends this system by three additional modules: Depth Network, Multi-Sensor Processor, and the Submapping Interface. Depth Network takes stereo images for the stereo network, as well as robot poses and sparse landmarks from the realtime estimator for the MVS network. The sparse landmarks are back-projected to image planes to provide a depth prior for the network. Depending on the use case, the Multi-Sensor Processor deals with incoming GNSS measure- ments and geometric measurements coming either from a LiDAR, a depth camera or the presented depth fusion network.

For GNSS measurements, it will formulate global position residuals that are added to the realtime estimator. It further determines the time that has passed since the last received GNSS measurement. In case of long signal outages, it will trigger an asynchronous, loop-closure like optimization of the full graph to compensate for the potentially accumulated drift since the last received measurement. In case of dense submap alignment, it retrieves poses for depth images or applies IMU-based motion undistortion of incoming LiDAR point clouds. For that it keeps an internal representation of the trajectory which is always kept synchronized with the estimator. For every live frame, frame-to-map factors with respect to a previous submap are added to the factor graph.

The Submapping Interface manages the collection of all submaps. The pose of each submap is anchored to a keyframe pose from the state estimator, therefore submap poses are always updated with updates from the state estimator using the same internal trajectory representation as the Multi-Sensor Processor. Each time a visual keyframe is generated, it checks for the overlap ratio of incoming measurements with respect to the last submap to decide whether a new submap is created.

After that decision, depth images or LiDAR measurements are integrated into the current active submap. Upon completion of a submap, dense map-to-map factors are added in the optimization problem to keep consistency of submaps in over- lapping regions. In the following sections, we will describe each module in more detail.

The Visual Frontend Extracts Brisk Keypoints And

descriptors in every image of a multi-frame, and matches them to the 3D landmarks already in the map; hereby both, descriptor distance and reprojected image distance, are consid- ered. New 3D landmarks are then initialized both from stereo triangulation within all images of a keyframe, as well as from triangulation between live frame and any of the keyframes. The decision of whether a new frame is considered a keyframe is taken depending on the fraction of matched landmarks in the live frame, as well as the fraction of current matches visible in any of the existing keyframes. If the overlap falls below a threshold, we set the live (multi-)frame as a new keyframe.

A finetuned version of the semantic segmentation net- work Fast-SCNN can be run on keyframe images. For maximum portability and flexibility, inference is carried out asynchronously on the CPU, which is tractable, due to the efficient network as well as the fact that only keyframes are processed. In our implementation, matches that are in the sky are removed. This scheme significantly improves accuracy in presence of slow-moving scene content, particularly clouds, where observations are not already automatically discarded via the Cauchy robustification.

C. Place Recognition, Relocalization, and loop-closure

Okvis2-X Maintains A Dbow2 Database. Dbow

queries of the current frames return a list of matches as loop- closure candidates. To be considered a valid loop-closure, an additional geometric verification step using 3D-2D RANSAC has to be passed. Then, the posegraph edges connecting the loop-closure state will be “revived” and turned back into landmarks and observations (see Fig. 4 (c)); and observations with the current matching frame will also be created. This may also trigger merging of landmarks, if already existing new landmarks are matched to old landmarks. To optimize the loop inconsistency, the error is equally distributed around the loop using rotation averaging followed by position inconsistency distribution. Then, a background optimization process of the full graph is triggered. After the optimization has finished, a synchronization process imports the optimized states and landmarks into the realtime estimator, and re-aligns new states and landmarks created in the meantime.

Fig. 4. Initially a full batch VI factor graph (a) is created and optimized. Later (b), frames with least overlap with the live frame and current keyframe are turned into posegraph poses by construction of relative pose errors under marginalization of common observations; also, old poses and speed/bias variables are fixed to keep the problem realtime capable. When a loop- closure occurs (c), respective observations and landmarks are re-activated.

The proposed system furthermore supports online calibration of the IMU-

Isual-Inertial Slam

The VI estimator will be minimizing visual, inertial, and relative pose errors (posegraph edges), which are briefly in- troduced in what follows. Fig. 4(a)-(c) visually overviews the construction of the underlying factor graph as time progresses.

Of The J-Th Landmark

in the frame of the i-th camera at a timestamp k and its

(11)

Hereby, h (·) denotes the camera projection. 2) IMU Errors: For the formulation of the IMU residuals, we adopt the IMU preintegration approach in . Between

(12)

where ˜xn is the predicted state at an arbitrary time n as a function of the current state xk and IMU measurements ˜zk,n

S

. The ⊟performs regular subtraction except for the quaternion (see ).

3) Relative Posegraph Errors: Relative pose errors er,c

P

between time steps r (the reference) and c are given by

(13)

with SrrSc and qSrSc being nominal relative position and

Orientations Expressed In The Imu Frame F

−→Sr. In short, these posegraph error terms are computed from joint observations into frames r and c by marginalizing out the landmarks. We refer the reader to for a detailed explanation of how a Maximum Spanning Tree (MST) is employed to select posegraph edges to be created. To compute

S , We

transform the landmarks into the reference frame F

−→Sr And

formulate the Gauss-Newton-System of the form Hδχ = b just from joint observations and respective landmarks, i.e. from standard reprojection errors and respective well-known

(14)

Hereby, only one pose occurs, i.e. the relative pose T SrSc referred to by δp. The variable ordering follows with all the landmarks (Srlj) referred to as δlj. Now, landmarks are

(17)

which is now linearized around the current pose p, therefore

(18)

The supposedly equivalent Gauss-Newton system of a respec-

P

denotes the Jacobian. For this to be equivalent to (18) at the linearization point (i.e. at p = p where Er,c

(21)

Since H∗, b∗, and p are constants, the Jacobians w.r.t. the poses at steps r and c are fairly straightforward to determine

E. Online Camera-Imu Extrinsics Calibration

We have extended to fully support online calibration of camera-IMU extrinsic poses T SCi, i = {1, . . , N}. While this extension is straightforward and well-known for reprojection errors, the relative posegraph factors from Eqn. (13) now also depend on T SCi. Consequently, before marginalizing co- visible landmarks, the two-view Gauss-Newton system will be augmented to include extrinsics poses – which remain after marginalizing out landmarks. Eqns (15) ff. remain the same, Fig. 5. Factor Graph including dense submap alignment. Left: The realtime estimator connects set of current keyframe and non-keyframe states by IMU errors and visual reprojection errors. For every state in the optimization window, frame-to-map factors are formulated between every live state and the keyframe state associated to the last completed submap. Right: Measurements between frames can be aggregated and map-to-map factors can be added to the factor graph between submap keyframe states if the geometric overlap

Surpasses A Threshold. (Adopted From .)

but with a higher-dimensional H and b. Eqn. (13) is then

(22)

where SrCi, qSCi denote the linearization point of the ith ex- trinsics, i.e. their values upon construction of the (augmented) relative pose factor.

The related factor graph structure can be observed in Fig. 4(d). Furthermore, to regularize the estimation problem, we add a pose prior to all camera extrinsics using the same formulation as with the pose prior of the estimated robot state1.

F. Submap Alignment

As stated in Sec. V-A and visualized in Fig. 5, there are two types of dense submap alignment residuals. Frame-to-map factors can be formulated between every live frame and the keyframe associated to the last completed submap. Map-to- map residuals are computed between two keyframes that two submaps are anchored to. However, regarding the actual error formulation, they do not differ. Given a completed submap associated to a keyframe pose T W Sa and a point cloud Pb associated to another state T W Sb , we formulate the map-based

−1T W Sb . The

respective Jacobians can be found in Appendix D. In the case of depth images, points Sbp are obtained by backprojecting pixel depth values. The idea of the map alignment residuals is that every measured point should be on a surface (L = 0) in the 3D map; and the distance d of the point from the nearest surface can be extrapolated from the occupancy value L (·) and 1In case we run a full-batch final BA including extrinsics, this prior is removed beforehand.

Fig. 6. To evaluate the global position residual, the propagated state ˆxj at time step j of the measurement is propagated from the state vector xi at time step i of the camera frame by leveraging IMU preintegration. This intermediate state is only used for evaluation of the global factor. (Adopted from .) the occupancy gradient ∇L (·) assuming a linear behavior near the surface as in the sensor model for mapping. The distance

(24)

Lmin is a configuration parameter and denotes the saturation minimum log-odds occupancy value. With the sensor-specific measurement uncertainty σd, we can formulate the weighted

(25)

While we found that for the typically highly precise LiDAR measurements, the simple assumption of an isotropic sensor uncertainty is sufficient, this does not hold for depth from the network. In that case, the depth uncertainty as in Eqn (10), which often includes noisy observations, should be assigned pixel-wise following the approach descriped in Sec. IV-A to properly weight its contribution to the optimization loss.

G. Gnss Fusion

In order to fuse global position residuals, such as from GNSS, we augment the state vector to also contain the 4-DoF transformation T GW between the VI reference frame F

−→G Which Is A Gravity-Aligned

East-North-Up (ENU) local Cartesian frame at the position of the first measurement. 1) Global Position Errors: The global position residual for a measurement zj at a time step j can be formulated as:

(26)

Hereby, SrA considers the position of the GNSS antenna in the IMU sensor frame which is assumed to be known beforehand. To account for asynchronous arrival of GNSS measurements and camera frames, WˆrSj and ˆCW Sj are predictions of the IMU poses at time step j which can be obtained from a previous state at time step i through IMU preintegration. This is visualized in Fig. 6. Compared to , the formulation in (26) extends the global residual by also considering a 4 DoF transformation T GW between the global and a local reference frame. The error Jacobians are derived in Appendix C.

2) Global Reference Frame initialization: A reliable ini- tialization of the global reference frame with respect to the VI reference frame is crucial. It is desirable to define an uncertainty-aware criterion for the observability of the 4-DoF extrinsic transformation between the two reference frames based on the received measurements. In a first step, an initial solution can be obtained using the SVD-based alignment method presented in from correspondences of global measurements and poses in the world reference frame.

Subsequently, the reliability of this initial solution can be evaluated. For Ng received measurements, global position er-

Rors Ej

g (j = 1, . . , Ng) and the corresponding error Jacobians

Ej

g are computed based on Eqn (26). The covariance matrix P for the 4 DoF transformation can be estimated as the inverse

Σj

g for GNSS residuals, which considers GNSS measurement covariances as well as IMU preintegration covariances (see Appendix B for more details). We examine the variance pθθ corresponding to the yaw angle θ in P, and define a

Θ. When The Yaw Uncertainty Falls

below this threshold, T GW can be assumed to be known and fixed. 3) GNSS Alignment: This mechanism of initializing and fixing the GNSS extrinsics enables a global alignment strat- egy to compensate for drift accumulated throughout potential GNSS signal dropouts. OKVIS2-X limits the computational complexity of the optimization problem by fixing states that are far in the past. Whenever GNSS residuals are added to the graph optimization problem, we check the fixation status of the last state that has an associated global position factor.

Dropouts in receiving GNSS signals over a long period of time can be identified by whether the last state with a GNSS factor in the graph is already frozen. In that case, the trajectory might have accumulated a significant amount of drift since then. Eliminating this drift can be addressed in a loop-closure like global alignment approach. After detecting a dropout, a re-initialization of the GNSS extrinsics is done as described in the previous paragraph. From this re-initialization process, a new estimate for the extrinsic transformation, T GWnew , is obtained. The discrepancy between the originally estimated T GW and the newly initialized T GWnew can be calculated as:

(28)

which also gives an estimate for the trajectory drift. Using this delta transformation and rotation averaging, the position and orientation error can be distributed across all the states during a GNSS dropout. After this alignment, an optimization of the full graph is triggered.

H. Factor Graph Optimization

All of the aforementioned factors are combined in the

(29)

Here, the set K contains the most recent frames as well as keyframes with observations of visible landmarks in J (i, k). P contains all posegraph frames and f denotes the most current frame. C (r) ⊂P is the set of all posegraph frames connected to a frame r. Furthermore, the set Lk denotes the set of all map alignment residuals associated to a frame k. M is the set of all past submaps, and Ab the set of all submap frames connected to a submap frame T W Sb via map-to-map residuals. C denotes the last completed submap frame. G is the set of added GNSS residuals. The Cauchy robustifier ρc (·) and Tukey robustifier ρt (·) are used for reprojection errors and map alignment errors.

Evaluation

The objective of this evaluation is to quantify the trajec- tory accuracy and mapping accuracy and completeness of OKVIS2-X with respect to state-of-the-art methods in small to large-scale scenarios. The large-scale environment (km-level traveling distance) possesses extreme challenges due to huge drift over time and memory usage in dense mapping to cover the large area. To showcase OKVIS2-X under such environ- ments, we selected three public datasets: EuRoC , Hilti- Oxford , and VBR datasets where the traveled distance spans from tens of meters with a flying drone (EuRoC), to hun- dreds of meters with a hand-held sensor rig (Hilti-Oxford) to 9 km for a driving car (VBR). For the proposed GNSS fusion,

Table Ii

ROOT MEAN SQUARE OF ABSOLUTE TRAJECTORY ERROR [METER] IN THE EUROC DATASET

Ours-Vid-Ba

All Ours report median in 10 runs. 1 Results taken from respective papers, and OpenVINS from . 2 We consider this as a failed sequence, thus exclude this in the average.

we further added an evaluation in a challenging sequence from the GVINS dataset including indoor-outdoor transitions and cluttered environments. All evaluations were performed in a serial manner, ensuring no frame drops. We account for the time offset between camera and IMU, if known (e.g. in the Hilti-Oxford dataset), as well as time offsets for both, LiDAR and GNSS, with respect to the IMU.

A. Evaluation Metrics

We evaluate several variations of our method to clearly show the effectiveness of multi-modality depending on the sensor configuration as well as the estimator causality. On the one hand, in Table I, Visual includes the reprojection (11) and relative posegraph residuals (13), Inertial indicates preintegration residuals (12), Depth-network is based on the depth network fusion for the submap alignment residuals (23), LiDAR uses LiDAR point clouds for the submap alignment residuals (23), and GNSS includes GNSS residuals (26). On the other hand, Causal only takes measurements up to the current time, Non-causal is the final loop-closed trajectory, and Full-BA refines the Non-causal trajectory in the entire states by turning relative pose graph edges to reprojection edges. For instance, Ours-vi-c means a visual-inertial configuration with the causal evaluation.

To quantify the accuracy of the estimated poses, the es- timated trajectory is aligned in SE(3) to the ground-truth before evaluation. We report trajectory accuracy as root mean square error (RMSE) of the absolute trajectory error (ATE) in EuRoC and VBR datasets. In the Hilti-Oxford dataset, we report the predefined score, where the position error [1, 10] cm is scaled to [0, 100] points. To evaluate mapping accuracy, submap meshes are reconstructed using marching cubes and combined using submap poses. Estimated and ground- truth vertices are downsampled with a voxel size of 1cm3, then we perform a point-to-plane ICP alignment between both sets of vertices. We report accuracy as an average distance from all estimated vertices to the ground-truth within 0.2 m. The completeness is defined as a fraction of ground-truth vertices that are within 0.2 m to the estimated vertices.

B. Euroc Dataset

This dataset provides stereo images and IMU measurements recorded by a drone, the ground-truth trajectory and mm-level accurate point clouds of the Vicon room . We use the stereo pair for our stereo network and left images for our MVS network. Table II summarizes ATE where all competitors also use a stereo or stereo-inertial configuration. We addi- tionally implemented V-SLAM, Ours-v, to show versatility of OKVIS2-X where the IMU preintegration term is replaced by a constant velocity model. OKVIS2-X shows competitive trajectory accuracy when compared to ORB-SLAM3. It is worth noting that the V203 sequence is challenging for V- SLAM due to the motion blur, image dropouts, and signif- icantly varying exposure. This is the reason why OKVIS2- X and ORB-SLAM3 present a large trajectory error on that specific sequence.

However, we resolve this large error in a stereo-inertial configuration. We compare the VIO implementation of Ours- vi without loop-closure for fair comparison to odometry

Approaches, Openvins And Kimera2 . Our Method

decreases the localization error by 41% when compared to competitors. In the SLAM implementation with loop-closure, Ours-vi outperforms state-of-the-art methods in the non-causal evaluation. It is worth noting that incorporating loop-closures

Authors:

Peder EZ Larson 1, 2,* , Jenna ML Bernard1, James A Bankson 3, Nikolaj Bøgh 4, Robert A Bok1, Albert P. Chen 5, Charles H Cunningham 6,7, Jeremy Gordon1, Jan-Bernd Hövener 8, Christoffer Laustsen 4, Dirk Mayer 9,10, Mary A McLean11 12, Franz Schilling13, James Slater1, Jean-Luc Vanderheyden5, 14, Cornelius von Morze 15, Daniel B Vigneron1, 2, Duan Xu1, 2, and the HP 13C

94143, Usa.

Denmark. 5 GE Healthcare, Menlo Park, California, USA. 6 Physical Sciences, Sunnybrook Research Institute, Toronto, Ontario, Canada.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

8 Section Biomedical Imaging, Molecular Imaging North Competence Center (MOIN CC), Medicine, Baltimore, MD, USA. Cambridge, United Kingdom.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

14Jlvmi Consulting Llc, Dousman, Wi, Usa

#See Acknowledgements for a list of all HP 13C MRI Consensus Group Members This work was supported by the ISMRM Hyperpolarized Media MR Study Group, the ISMRM Hyperpolarization Methods & Equipment Study Group, and the Hyperpolarized MRI Technology Resource Center (NIH/NIBIB grant P41EB013598).

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Abstract

MRI with hyperpolarized (HP) 13C agents, also known as HP 13C MRI, can measure processes such as localized metabolism that is altered in numerous cancers, liver, heart, kidney diseases, and more. It has been translated into human studies during the past 10 years, with recent rapid growth in studies largely based on increasing availability of hyperpolarized agent preparation methods suitable for use in humans. This paper aims to capture the current successful practices for HP MRI human studies with [1-13C]pyruvate - by far the most commonly used agent, which sits at a key metabolic junction in glycolysis. The paper is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification. In each area, we identified the key components for a successful study, summarized both published studies and current practices, and discuss evidence gaps, strengths, and limitations. This paper is the output of the “HP 13C MRI Consensus Group” as well as the ISMRM Hyperpolarized Media MR and Hyperpolarized Methods & Equipment study groups. It further aims to provide a comprehensive reference for future consensus building as the field continues to advance human studies with this metabolic imaging modality.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Keywords: Hyperpolarized MRI, metabolic imaging, carbon-13, pyruvate, dissolution dynamic

Introduction

MRI with hyperpolarized 13C agents, also known as hyperpolarized (HP) 13C MRI, has shown great potential as a novel imaging modality, particularly for its ability to probe metabolic processes in real time. The first human studies with HP [1-13C]pyruvate were performed in 2011 in prostate cancer patients (1).

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Since then, there have been over 60 papers published with imaging results of human subjects from 13 different sites, with applications including prostate cancer, brain tumors, breast cancer, kidney cancer, pancreatic cancer, metastatic disease, liver disease, ischemic heart disease, diabetes and cardiomyopathies. The vast majority of these studies used [1-13C]pyruvate (1–63), where [2-13C]pyruvate (64) and 13C-urea (56) have been demonstrated too.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

As clinical HP 13C MRI advances, there is a growing need to build consensus for best practices, which are critical for comparing data across sites, performing multi-site trials,deploying methods to new sites, partnering with vendors, and potentially for obtaining broader regulatory approvals.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

In March 2022, we initiated an effort to build consensus within the HP 13C MRI community with this opportunity in mind, and it was greeted with strong enthusiasm. The “HP 13C MRI Consensus Group”, containing over 55 members from 27 sites, identified the area of greatest need and opportunity for consensus building to be HP [1-13C]pyruvate human

●

Pyruvate is the most mature and widely used HP agent and has the most significant translational evidence emphasizing the potential clinical impact.

●

Clinical trials, particularly multi-site trials, have the strongest need for consensus methods to ensure that data can be combined across sites. This work is a Position Paper for which the goal is to describe current successful practices and study methods for HP [1-13C]pyruvate human studies along with justification to support those practices. This is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification (Fig. 1). The current successful practices and study methods include a literature review of published peer-reviewed journal papers showing human HP [1-13C]pyruvate study data, up to September 2022 (1–63), as well as new unpublished information from surveys of HP 13C study sites. Based on this information, we also highlight the evidence gaps, strengths, and limitations of current practices which are summarized at the end of each section.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Figure 1: Illustration of the HP 13C MRI human study process, including the 4 major areas covered in this paper: Hyperpolarized 13C-pyruvate preparation, MRI system setup and calibration, Acquisition and Reconstruction, and Data Analysis and Quantification.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

Figure 2: Anatomical targets of HP [1-13C]pyruvate MRI human studies published up to September 2022.

Hyperpolarized 13C-Pyruvate Preparation

This section covers the processes for creating the HP agent, 13C pyruvate, and will include many aspects and considerations that are needed to safely and effectively prepare doses for metabolic imaging studies in human subjects. These include material, personnel, equipment and facility, fluid path preparation, quality control, and release.

ansys-mri-compatible-device Diagram
Figure: System Model & Simulation Flow for Ansys Mri Compatible Device

It is helpful to understand that the specifications of a dose of 13C pyruvate suitable for in vivo MR HP metabolic imaging were shaped in part by early preclinical studies performed by GE HealthCare summarized in Ref. (65). In short, the safety of the two novel drug components, 13C pyruvate and the electron paramagnetic agent (EPA) AH111501, were demonstrated in those studies. The more precise formulation of the dose suitable for human use was then determined from clinical studies (66) that included two Phase 1 clinical trials in young and elderly healthy volunteers without hyperpolarization of the 13C nuclei and another Phase 1/2a dose escalation and imaging feasibility study with HP 13C pyruvate in 31 prostate cancer patients at the With the exception of the first HP 13C imaging clinical trial, which utilized a prototype device in a cleanroom (1), all HP 13C studies performed in humans to date have utilized the SPINlab polarizer (manufactured by GE HealthCare). Consequently all doses of the HP 13C pyruvate delivered by SPINlab have been produced using the “SPINlab Pharmacy Kit” that serves as the container-closure system for the various drug components (13C pyruvic acid and EPA mixture, dissolution medium, and neutralization and dilution medium) during sample polarization, dissolution and quality control (QC) processes. Thus many aspects of the HP sample preparation considerations discussed below are related to the SPINlab instrument and the consumables designed to be used with it (67).

General Considerations

While more than 860 patients or healthy subjects having been injected with HP 13C pyruvate as of January 2022 without reports of any serious adverse events (68), HP 13C pyruvate injection remains an investigational MR contrast agent and can only be administered by those with Investigational New Drug (IND) exemption from the Food and Drug Administration (FDA) in the USA, a Clinical Trial Application (CTA) in Canada, approval from National Research Ethics Committee Services in the UK, or approval from the relevant local regulatory body. Thus, methods and processes involved to produce a dose should have patient safety as the first priority. Since utilizing dissolution dynamic nuclear polarization (dissolution-DNP) for human use is still a relatively new development, there are no existing published regulatory guidelines specifically for this method.

There are two major production styles that determine how various sites approach the agent preparation. In the US, the most common approach is to rely on a sterilizing filter (“Terminal Sterilization”) to ensure sterility of the final product, akin to PET tracer production, where a starting molecule with a radioisotope is processed using various other ingredients to make the final, desired and injectable contrast agent within a necessarily short amount of time (69). For these sites, sterilization of the components and accessories upstream of this filter are not required, although many of them were manufactured and tested following Good Manufacturing Practice (GMP) or Good Laboratory Practice (GLP) requirements. The filling process is usually performed under an ISO 5 laminar flow hood, but a clean room or an isolator is not required.

This approach is typically accompanied by testing the integrity of the sterilizing filter prior to release of the dose for injection. Typically, post release endotoxin and sterility tests are performed using an aliquot reserved from each released dose.

In the UK and EU, the most common approach is to more-closely follow sterile pharmaceutical compounding guidelines (70), where all components and ingredients are required to be sterile or manufactured under GMP guidelines and are assembled and filled within a clean room environment or an isolator system (“Sterile Preparation”). Typically a batch of Pharmacy Kits for HP 13C pyruvate injection are prepared together. The sterility of the final dose is also ensured by batch validation testing, in addition to the sterility of the ingredients and the sterile compounding process. The endotoxin and sterility testing are performed for the process validation but are not performed for each injected dose.

Some institutions fill and assemble the Pharmacy Kit required for a specific study on the same day or the day prior to polarization, dissolution, and patient administration, but others have also demonstrated the feasibility of preparing a batch of kits, keeping them in a -20ºC freezer and using them over a period of a few months.

Beyond the obvious requirements that the process and the facility has to ultimately produce a dose that is safe to inject into a human, regulatory authorities will also focus on the question “Are you in control of your processes?”. To be in control of your process requires an in-depth and broad understanding of all processes involved in pre, post, and during the production process.

Personnel

It is typical and may be required to have licensed personnel involved in the production process depending on local regulations.Typically a pharmacist, radiopharmacist or other similarly qualified person (QP), in charge of the facility where the Pharmacy Kit filling and preparation is taking place, is responsible for the overall process and the release of the injectable dose.

Qualified cleanroom technicians are often involved in the Pharmacy Kit filling under the supervision of the pharmacist or QP. As is required for pharmaceutical compounding or PET tracer production, training requirements and training records for all personnel need to be maintained and available for audit by the FDA or equivalent.

Equipment And Facility

The facility and all equipment need to have standard operating procedures (SOPs) that describe how equipment is used, maintained, and calibrated to comply with relevant legislation. Currently, almost all the filling of the Pharmacy Kit takes place within a compounding laminar flow hood or isolator (typically ISO 5). At some sites, the filling is conducted within a cleanroom, while at others, it is conducted in a dedicated non-cleanroom space, reflecting differences in cleanroom approach and specifications between regulators worldwide (71). Some equipment or facilities, such as the compounding hood or cleanroom, may require external certified laboratories for testing.

Material Handling

Material handling guidelines (69,70) require SOPs detailing a system to track all of the materials involved in the HP production process for a particular patient dose, similar to current good manufacturing practice (cGMP) requirements for material handling for drug compounding. This includes acceptance standards, storage conditions, amount used in the patient dose for each ingredient and materials used in the assembly of the fluid path and Pharmacy Kit. Currently some users choose to open and inspect and sometimes modify the Pharmacy Kits upon arrival, but some users keep them in the sealed packaging until they are required for dose preparation.

Pharmacy Kit Filling And Assembling

As required by an IND or its equivalent, the preparation of the doses of HP 13C agent are detailed in the Chemistry, Manufacturing, and Control (CMC) section of an applicable regulatory submission; an example of this has been made available (72). It describes the processes of filling the Pharmacy Kit with the different components that make up the final drug product, and of assembling the final kit for either storage or immediate use in the polarizer. Special attention should be given to the laser welding process in order to satisfy installation qualification (IQ) and operational qualification (OQ). Typically, the final developed process is validated by process qualification (PQ) runs, during which 3 or more Pharmacy Kits are filled and used and the final HP 13C products are tested for endotoxin and sterility and to confirm that they meet the dose specifications for injections (usually including pyruvate concentration, residual EPA concentration, pH, liquid state polarization level and dose temperature). The data from 3 consecutive PQ runs are submitted as part of the IND submission (or its equivalent), and are often also reviewed by the Institutional Review Board (IRB) where the studies are conducted.

Quality Control And Dose Release

The quality control (QC) and dose release can be separated into two aspects: one is the QC and release of the filled Pharmacy Kit, and second is the QC and release of the HP 13C agent for injection, after polarization and dissolution. For institutions filling a batch of kits and storing them to use over a period of time, typically the batch can be released based on initial validation, environmental monitoring data from the day of kit production, and if filters are used during preparation of any of the components, filter integrity testing. But in some cases one or more kits are used for validation before the batch of kits are released for future use. For institutions that fill only the kits required for specific studies shortly before the experiment, the filled kits often do not go through separate release tests before they are used.

The quality control of the HP 13C pyruvate solution post dissolution is primarily performed to ensure that the agent meets the dose specifications (Table 1) before it is administered to the subject. These specifications target both safety (pH, residual EPA, temperature) and efficacy (pyruvate concentration, polarization, volume). Typically, the pyruvate concentration, residual EPA concentration, pH, dose temperature, dose volume, and liquid state polarization are measured by the QC accessory associated with the SPINlab polarizer. Some users perform a secondary measurement for one of the parameters, such as pH, using a different instrument or pH paper. For sites that do not go through a separate release testing process for batch filled kits, the integrity of the sterilization assurance filter, a part of the Pharmacy Kit, is typically tested as a part of the dose release. It is also common for these users to preserve an aliquot of the final HP 13C pyruvate solution for post-release endotoxin and sterility testing. This testing cannot be completed fast enough to test an individual dose prior to injection, but this is why other processes such as PQ runs and validation testing are done to minimize the chance a subject could be injected with a contaminated dose.

The Final Dose Release And Injection

should be done under the supervision of a licensed professional, based on local regulations.

Some Key Challenges

Many of the challenges associated with HP 13C pyruvate preparation can be attributed to the conditions required for the dissolution-DNP method of high magnetic field (~3-7 T) and very low temperature (~1 K) during polarization, with pressurized and superheated water necessary for the rapid dissolution event. These extreme conditions are quite challenging for the design of the container-closure and fluid path system. In particular, the cryogenic temperature in the polarizer requires special attention to any moisture or ambient (moist) air introduced into that portion of the fluid path, which can form an ice block at ~1 K. This ice can lead to flow restriction during the dissolution event and reduce the strength of the laser welded bond between the cryovial and its cap. This can ultimately produce failures in the dissolution step, including variations in final pyruvate concentration and pH that may fail to meet QC release criteria as well as fluid path ruptures that provide no available dose and result in polarizer down-time.

The polarization of the HP 13C pyruvate sample decays quickly over the span of a few minutes after dissolution, and thus the process of dissolution, QC for release, and injection should be completed as fast as possible to preserve the high polarization level achieved. Any delays in the preparation process, such as transportation time or equipment malfunction, can significantly reduce the final polarization and result in lower quality imaging data.

Current Practices

A summary of data collected from all sites performing clinical trials with HP 13C-pyruvate is shown in Fig. 3 and Table 1, including the specification of the final dose and how the quality control and release of the final dose are performed. There is a split in the Production Style, described in the General Considerations section above, with 8/13 sites using Sterile Preparation versus 5/13 using Terminal Sterilization. While many of the dose specifications show notable differences in acceptable ranges, all of these variations listed in tables have been successfully and safely been used to perform HP 13C pyruvate studies in humans. Their differences depend on the institutions’ preferences, resources and their particular regulatory situation. There is high similarity in pyruvate ranges, temperature ranges, EPA limits, and volume limits. There is modest variability in pH ranges and large variability in the endotoxin test limit. There is a 3-fold difference in acceptable polarization levels, which are measured to ensure a futile dose is not injected since the polarization is directly proportional to SNR. This reflects the decision by several sites to believe that useful data can be still be obtained with suboptimal polarizations.

Figure 3: Hyperpolarized agent preparation methods reported by sites currently performing HP

In House

Table 1: HP 13C-pyruvate preparation parameters, methods, and dose specifications used for quality control testing and release as well as validation. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. The parameters used for product release are noted in bold text, otherwise these parameters are measured for batch validation or other QC measurements. The endotoxin and sterility testing are performed during process validation of the batch and/or post-injection, and largely depends on the agent production approach.

Summary

The overall safety record of HP 13C-pyruvate has been very strong, and the SPINlab hyperpolarizer has proven to provide high polarizations at human sized doses while meeting numerous QC and release criteria. A weakness remains the failure modes of the SPINlab Phamacy Kits (e.g. ice blocks, path ruptures), which are placed under extreme requirements particularly during dissolution. The preparation process still requires a high degree of expertise.

Therefore, there is a significant need to improve the reliability, robustness, and ease of operation for generating HP 13C-pyruvate doses for human studies. Furthermore, there is a divide between manufacturing and sterile compounding style preparation as well as other site-specific practices, resulting in variations in SOPs and justification required to relevant regulatory bodies. There have also been no comparisons between these approaches. It is also unclear what release criteria and QC parameters are truly required to ensure patient safety.

However, all of the reported methods are acceptable and approved by the appropriate regulatory authorities, and have led to the rapid expansion of successful human studies in recent years.

Mri System Setup And Calibrations

This section covers the MRI system setup, including the imaging system, RF coils, phantoms, and prescan calibration methods.

Imaging System

The main prerequisite for a given MRI scanner to be capable of supporting studies with HP 13C is its “broadband” capability to transmit and receive radiofrequency (RF) signal at the frequency of 13C, which is around 4 times lower than 1H. This does not come as a default on clinical MR devices. The transmit power of the broadband amplifier should also be sufficient to support the intended flip angle and RF pulse shape with the employed transmission RF coil(s) for 13C. Most studies to date use relatively low flip angles (< 90 degrees) for HP 13C in order to preserve polarization for time-resolved imaging. The capability to receive 13C signal on multiple channels is also desirable to increase SNR, as discussed further in the “RF coils” section.

The choice of magnetic field strength is primarily dependent on the metabolites’ frequency separation due to chemical shift dispersion and 1H imaging. High field strengths do not enhance hyperpolarized 13C signal as they do for 1H because the signal strength in a HP experiment relies on manipulating the population of quantum energy states outside of the MRI scanner.

However, the injected HP 13C-pyruvate and its metabolic products have greater frequency separation at higher fields, and it may thus be easier to separate and quantify these resonances at higher fields. This comes at the cost of a reduction in the achievable T2* and often reduced T1. As the initial polarization is independent of the imaging field strength it has been proposed that the increased T2* at 1.5T can potentially be exploited to increase SNR by adapting the acquisition bandwidth or reduce off-resonance imaging effects in cases when the decay of the transverse magnetization is dominated by T2* (73). In practice, 3T has been used in all published human 13C-pyruvate studies surveyed (Supporting Table S1), and comprises the majority of scanners currently in use for human studies (Table 3). A field strength of 3T is well-suited for 1H MRI anatomical reference and correlative imaging.

Stronger and more rapidly slewing magnetic field gradients support more rapid spatial encoding, particularly for metabolite-specific single-shot imaging using echo-planar imaging (EPI) or spiral imaging (See “Acquisition and Reconstruction”). Although the spatial resolution acquired for HP 13C imaging is typically much coarser than for 1H MRI, the factor of ~4 in gyromagnetic ratio leads to the same reduction factor in performance of the gradient system, so 13C experiments are potentially more limited by gradient hardware performance. To date, all human studies have used the commercially-available integrated gradient systems provided in clinical MRI scanners.

Optimization of scanner design has understandably focused on minimization of artifacts in 1H MRI, where devices such as room lights, the gradient amplifiers, and the motors driving the patient bed are checked to ensure that they do not produce RF interference at the 1H frequency, but artifacts may arise at other frequencies. Eddy current compensation is also not always appropriately adjusted for nuclei at other frequencies (74). In order to optimize for 13C, many sites have performed checks on phantoms for RF interference, gradient artifacts, and eddy currents (74), including the use of post-hoc gradient impulse response function characterisation and correction, and some vendors have fixed these issues as well.

Rf Coils

For HP 13C imaging studies in humans, RF coils for both 1H and 13C nuclei are needed, with 1H MRI providing an anatomical reference for registration and optional additional multiparametric MRI readouts. At the Larmor frequency of 13C nuclei, the relative contributions from coil noise compared to sample noise increase compared to 1H (73,75), although sample noise still is likely the dominant contributor for human-sized coils at 32.1MHz - the resonance frequency of 13C nuclei at 3T.

The key requirement for human 13C-pyruvate RF coils are that the coil geometry and sensitive volume must cover the volume of interest in the subject. Table 2 and Figure 4 shows coil configurations that have been used and optimized for applications in different anatomic regions.

Volume resonators are most commonly used for transmit, as they surround the subject to

Provide B1 Transmit Across The Fov (B1

+). While 1H relies on a large birdcage (“body”) coil built into the scanner, 13C transmit coils must be placed inside the bore. This takes up valuable space within the magnet, and also has led to the use of designs with relatively inhomogeneous

B1

+. Many human studies have used Helmholz pair resonators for transmit, including the “clamshell coil”, which has a notably inhomogeneous B1

+ Profile But Has Been Used Because Of

relatively easy integration into the scanner bore. B1

+ Variation Results In Variations In The Flip

angles that control the use of the hyperpolarized magnetization and creates errors in common HP metrics (9,76). The exception are head coils, where birdcage designs with highly

Homogeneous B1

+ can be placed around the head while easily fitting inside the bore. As with 1H MRI, higher SNR can typically be achieved by smaller receive coil elements, such as surface coils or phased arrays, and the majority of 13C receive coils used have layouts similar to 1H phased arrays.

RF coil quality control is important to ensure proper functioning of the coils to provide consistent imaging quality, especially with limited natural abundance 13C signal in vivo. It typically involves 1) a physical integrity check of the coil cables and connectors and 2) phantom SNR tests to check the coil’s performance and to monitor it over time (see Phantoms below). An useful reference for RF coil quality control is outlined in the MRI accreditation program of the American College of Radiology (77) and can be adapted for 13C coils.

Notably, configurations for brain and prostate studies used dual-tuned 1H/13C coil designs, which greatly simplify workflow and registration of 1H and 13C images, as no switching of coils is needed.

(1)

Table 2: RF coil configurations reported for human HP [1-13C]pyruvate studies.

Tx = Transmit

coil, RX = receive coil. The commonly used “clamshell” TX coil is a Helmholz pair design. For 1H RF configurations, all used the Body coil for TX unless otherwise noted, and “repositioned” indicates the 13C coil was removed for 1H imaging. One representative reference is listed for each configuration. The RF coil configurations reported in the reviewed papers are shown in Supporting Table S1.

Figure 4: Examples of RF coil configurations used for human HP [1-13C]pyruvate brain studies. (A,B) 13C Clamshell TX (Helmholz pair) and 2× 4-channel paddle RX arrays. (C) 13C Birdcage volume TX and 32-channel RX array (RX array slides into TX coil). (D) 13C Birdcage volume TX and 24-channel RX array, combined with a 1H 8-channel RX array. Image reproduced with permission from Ref (16).

Phantoms

Since hyperpolarized magnetization is non-renewable, phantoms containing 13C nuclei are important to: 1) test the multi-nuclear capabilities of the imaging system, including all parts of the signal excitation and receive chain; 2) perform calibration measurements before a scan with hyperpolarized nuclei; and 3) perform necessary pre-scan adjustments (see “Prescan Calibration” section). The phantoms currently in use are listed in Table 3. Their composition must provide sufficient 13C signal, with additional considerations of conductivity, stability, chemical shift(s) present, potential for dynamic imaging, and cost. The phantom geometries are typically either compact, in order to be used alongside the subject during a HP scan, or large enough to mimic the inner volume of a RF coil for system testing.

One popular compact design contains enriched 13C-urea at high concentration, typically 8 M, which provides a single resonance, placed inside a small container ~1 mL. The most common recipe mixes 13C-urea in a 90% water/10% glycerol solution, with glycerol used to increase the urea solubility and doping with a Gd-based contrast agent to shorten T1 which increases the potential SNR per unit time. For example, when Dotarem is added at a 3:1000 volume ratio the 13C-urea T1 is around 500 ms and T2 is around 100 ms. However, when testing pulse sequences influenced by T1 and T2, doping should be used carefully. This phantom is suitable for frequency calibration, transmit gain calibration, sequence testing, and as a fiducial marker when placed next to a patient. However, enriched 13C-urea has a relatively high cost compared to natural abundance compounds.

For larger volumes (>100 ml), the phantoms most often used contain undiluted ethylene glycol, glycerol, or dimethyl silicone. These compounds have sufficiently high carbon concentrations to provide sufficient 13C signal even with the 1.1% natural abundance of 13C. These larger phantoms matching the inner volume of an RF coil are useful for coil testing, including transmit

+) And Receive (B1

-) coil profile mapping, as well as to mimic acquisitions using in vivo FOV requirements. In this case, size and conductivity should match the expected subject size in order to mimic coil loading and get a realistic estimation of B1+. Large-volume natural abundance urea phantoms have also been used by some sites, but suffer from higher conductivity compared to biological tissues. Typically, it is easier to increase the conductivity and hence coil loading of the non-conductive phantom by adding NaCl to match physiological loading (16,78).

Dynamic phantoms that aim to mimic metabolite kinetics have also been developed (79–81), and have the potential to more closely mimic the HP experiment, but so far these are not widely used.

Prescan Calibration

Prior to performing an MRI acquisition, the so-called prescan procedure is used to set the shim parameters to maximize B0 homogeneity over the field of view (FOV) or a specific region of interest (ROI), the scanner center frequency (CF), the RF transmit gain, and the receiver gain.

While this calibration procedure is usually automated for 1H, the lack of sufficient natural abundance 13C signal prevents use of automated methods. (Although natural abundance 13C lipid signal has been detected, there are so far no reports on using this signal for prescan.) Table 3 shows current practices across sites.

Maximizing B0 homogeneity is independent of the nucleus and is therefore performed prior to 13C imaging using the 1H water signal and existing shimming tools, such as by a standard automated process (“Auto Shimming”) or using high order shimming routines. Similarly, the 13C CF can be calculated from the 1H CF using a predetermined scaling factor that depends on the target chemical shift (82). Another common approach used is to have a small, high-concentration 13C phantom, e.g. 8M 13C-urea, integrated in the RF coil or placed next to the scan subject (1). The reference frequency can also be based on real-time measurements after the HP injection but prior to imaging (83). Both the CF and B0 shimming are critical when using spectrally-selective RF pulses, as inmetabolite-specific imaging methods, where the desired excitation bandwidths are typically very narrow and frequency offsets can lead to a failure mode that is only apparent after injection.

The calibration of the RF transmit power is typically performed on a small, high-concentration 13C phantom placed near the region of interest during the scan or on a large 13C phantom of similar size and coil loading as the subject, prior to the subject scan. Reference power is often done by sweeping the power in a pulse-acquire sequence (53,62), or the Bloch-Siegert method (52,84). When using a small phantom, the location of the phantom, B1

+ Inhomogeneity As Well

as any shielding effects, e.g., when the phantom is integrated into a coil (1), may degrade the accuracy. Other methods include real-time Bloch-Siegert method measurements after the HP injection (83), and using the stronger natural abundance 23Na signal that is close enough to the 13C resonance frequency to be detected by 13C coils (82).

The receiver gain is predetermined, either systematically based on independent phantom measurements and assuming the dose and polarization of the HP compound is known prior to injection, or based on past HP imaging studies.

Power [Kw]

Phantom(s) - during study Phantom(s) - before study 13C Frequency

8

13C-bicarbonate doped with dimethyl silicone, various

Power [Kw]

Phantom(s) - during study Phantom(s) - before study 13C Frequency

Maximum Values

Table 3: Summary of the imaging systems, phantoms, and prescan procedures used at sites currently performing HP 13C-pyruvate human studies. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. *Previously performed studies with a Siemens 3T Tim Trio. The imaging systems, phantoms, and prescan procedures reported in the reviewed papers are shown in Supporting Table S1.

Summary

Commercially available 3T MRI systems are by far the most commonly used for human HP 13C-pyruvate studies, although a systematic investigation of the impact of B0 has only recently been investigated (73). The multi-nuclear RF transmit and receive chain has proven sufficient for current acquisition strategies, although many sites have observed artifacts due to RF interference, gradient interference, and residual eddy currents when operating at the 13C frequency. A variety of 13C RF coils, tailored for numerous anatomical targets, have been successfully demonstrated, with the main limitation that most transmit coils take up a lot of additional space inside the bore and provide relatively inhomogeneous B1

+ Profiles. The

phantoms used have converged into generally 2 categories - small phantoms containing 13C-enriched compounds that can be used during the study and human-sized phantoms containing compounds with high carbon concentrations but without 13C enrichment that are used to test and calibrate the coils. There are no standardized compositions or geometry, and dynamic phantoms that recapitulate in vivo kinetics would be desirable but are still an emerging area. Prescan calibration procedures were not well defined in most publications, so we surveyed individual sites to determine current practices. Calibration procedures for the B0 field (13C CF and shimming) for most sites take advantage of 1H signal and methods, while methods

For Calibration Of B1

+ is more variable across sites, likely a reflection of remaining challenges in how to perform this calibration. Standardization of both phantoms and calibration procedures would synergistically improve the robustness and reproducibility of HP 13C studies.

Acquisition And Reconstruction

Data acquisition strategies in human HP [1-13C]pyruvate MRI studies must account for multiple chemical shifts, efficiently utilize the non-renewable HP magnetization, and acquire data quickly relative to metabolism and relaxation decay processes. These studies require spectral encoding to separate metabolites, necessitating pulse sequences that efficiently encode up to 5D data (3 spatial + 1 spectral + 1 temporal dimension). RF pulses must efficiently sample without immediately saturating the non-renewable HP magnetization, and sequences must acquire data quickly and be robust to both experimental and physiologic variation (e.g. B1

+ Inhomogeneity,

variation in perfusion) to ensure reproducibility and minimize scan-to-scan variability. This section covers current successful practices for data acquisition in human [1-13C]pyruvate studies, and accompanying 1H imaging, from different anatomic regions, including scan parameters and image reconstruction.

Acquisition And Reconstruction Methods

The acquisition methods used in human [1-13C]pyruvate studies can be classified into 3 categories: 1) MR spectroscopy or MR spectroscopic imaging (“MRS/I”), 2) chemical shift encoding methods, and 3) metabolite-specific imaging (Fig. 5).

Mrs/I Methods Specifically

resolve a spectrum that can be analyzed to extract expected as well as unexpected resonances, making this approach very robust. It was used in many initial studies (1).

Chemical Shift

encoding methods, most commonly the Iterative Decomposition of water and fat with Echo Asymmetry and Least-squares estimation (IDEAL) method, use imaging sequences acquired with multiple TEs and rely on a model-based separation of expected chemical shifts (85).

Metabolite-specific imaging methods use specialized RF pulses that are spatially and spectrally selective to excite individual metabolites which are then typically imaged with fast k-space trajectories such as echo planar imaging (EPI) or spirals (86).

Their Application To Different

organ systems is described below. The image reconstruction methods used in human [1-13C]pyruvate studies have typically been conventional methods (e.g. FFT, non-uniform FFT, or equivalent). The incorporation of accelerated imaging and advanced reconstruction methods including parallel imaging (4,57,87) and compressed sensing (7) has also been applied in human studies for improved spatial resolution, temporal resolution and coverage, but have the potential for additional artifacts as well as SNR losses due to ill-conditioning of the reconstruction (e.g. g-factor).

The Majority Of

published studies do not use accelerated imaging indicating the resolution and coverage achievable without acceleration is currently adequate for successful data collection. Performing coil combination, even with fully sampled data has also been shown to have specific challenges for HP human images: using naive sum-of-squares methods suffer from high noise amplification in the relatively low SNR regime of HP [1-13C]pyruvate (compared to 1H), motivating several HP 13C-specific methods that include data-driven coil sensitivity estimation which have shown obvious improvements over sum-of-squares (11).

More recently denoising techniques have been applied as post-processing of human HP data(41,42,44). The techniques applied are based on spatial-temporal singular value decomposition for unsupervised estimation of signal and noise components. They have shown improvements in apparent SNR in the brain and liver, while care must be taken to choose parameters such as the rank threshold to avoid oversmoothing and overfitting to the estimated signal components.

Prostate Studies

Prostate cancer was the first human application of HP [1-13C]pyruvate (1), and data was acquired with MRS/I methods: 1D dynamic MRS, single-slice 2D dynamic echo-planar spectroscopic imaging (EPSI), and single time point 3D EPSI. Advances in imaging strategies led to the development and application of new acquisition schemes, including undersampled 3D EPSI with compressed-sensing (7), model-based chemical shift encoding methods that use a priori information (47,59), and metabolite-specific EPI (10), all of which can provide volumetric whole-organ coverage and dynamic acquisitions.

The pyruvate bolus arrival in the prostate can vary by ± 10 s between patients, necessitating dynamic imaging to reliably and consistently capture the pyruvate bolus (18). For this reason, all currently ongoing studies acquire dynamic data. While MRS/I, chemical shift encoding, and metabolite-specific imaging can all achieve dynamic imaging, chemical shift encoding and metabolite-specific imaging provide greater dynamic and volumetric coverage (85). For scan prescriptions, the FOV is designed to provide full prostate coverage and typically to match the orientation of the anatomic imaging used for registration. Flip angles used in current studies are constant through time, as quantification with a variable-through-time flip scheme is highly sensitive to bolus timing (8) and errors in the RF transmit (B1 +) field (76).

Heart Studies

Data acquisition methods for 13C imaging in the heart must be designed to meet the demands of significant cardiac motion and blood flow. To cope with the periodic cardiac motion, most human heart studies to date used gating to the diastolic window, the longest cardiac cycle interval, which has reduced motion (2,22,28,30,35,36,38,45,52). The duration of the diastolic window limits the available data sampling time, making cardiac acquisitions the most time-constrained of the HP 13C MRI applications. The most common acquisition approach is metabolite-specific imaging with spiral k-space trajectories (2). Their single-shot imaging capability makes these methods particularly robust to motion effects. Furthermore, spiral k-space trajectories provide rapid k-space coverage and relatively benign flow and motion artifacts. The majority of studies have used 2D multi-slice acquisitions, but 3D encoding has also been used successfully (35).

Brain Studies

For HP 13C MRI of the human brain, the majority of studies have also used 2D (slice selective) acquisitions (10–12,14,16,28,33,40,41,44,51,53,60), with a trend toward volumetric coverage using 2D multi-slice metabolite-specific imaging. 3D metabolite-specific imaging of the whole brain, with phase encoding of the slice direction (34,57), has been shown to provide similar SNR efficiency (88) compared with multislice imaging. A number of studies have employed MRS/I (5,6,29,31–33,50,55) resulting in a spectrum from each voxel, which has the advantage of not requiring a priori information about which peaks to encode. This was important in early brain studies when it was not known which peaks would be detectable. Chemical shift encoding, using a set of images with different echo times and an iterative reconstruction of the individual resonances (i.e. the IDEAL approach (85)), has also been used (12,49,54), with the drawback that coverage in the slice direction was limited due to the time required to acquire multiple echo time images.

Abdomen And Breast Studies

The fundamental approaches to data acquisition and reconstruction in the abdomen and breast are largely similar to the aforementioned applications, but demand attention to particular challenges associated with these anatomic regions, especially relating to respiratory motion.

Although it has been shown that a basic 2D MRSI approach based on phase encoding and FID readout can be successfully applied for HP 13C imaging in breast (15) and kidney (13), major advantages in terms of spatiotemporal resolution and coverage have been realized using tailored approaches based on metabolite-specific imaging (43,62) and chemical shift encoding (43), which have facilitated multi-slice or 3D dynamic acquisitions over large FOVs in the abdomen (4,37,46).

The significant respiratory motion encountered in these regions can directly blur 13C images, and has further favored these rapid acquisition strategies. Motion also degrades B0 homogeneity, which can shift frequency-selective excitation profiles and introduce artifacts into rapid imaging readouts. This makes accurate determination of the acquisition center frequency and shimming essential in these regions which often cover large FOVs. (See “Prescan Calibration” section for more information). In some studies, breath-holding was used to minimize motion effects and enforce frame-to-frame data consistency (42). A pragmatic and reasonably effective approach for dealing with respiratory motion during 13C data acquisition is an initial breath-hold (as long as can be tolerated), followed by free-breathing (46,62).

1H Imaging

Collection of 1H imaging data is essential both for prescribing the 13C acquisition and for interpretation of the resulting 13C data. Multi-planar 1H scouts are acquired prior to 13C acquisition to enable graphical prescription of the 13C imaging region. All human HP 13C-pyruvate imaging studies acquire conventional MRI scans (e.g. T1- and T2-weighted volumes) for anatomic reference, aiming to cover at least the full 13C FOV. Acquiring these anatomic scans as close as possible to the time of 13C imaging (immediately before or after) minimizes potential misregistration between the data sets. Depending on the application, other advanced 1H sequences are also acquired (e.g. diffusion-weighted imaging for cancer imaging).

When contrast-enhanced data is acquired, it is done after 13C imaging, as paramagnetic contrast agents will accelerate 13C relaxation.

Reported Study Parameters

Figures 5 and 6, and Supporting Table S2 shows the reported acquisition study parameters for human HP [1-13C]pyruvate studies published as of September 2022. Figure 5 shows a mixture of MRS/I, metabolite-specific imaging, and chemical shift encoding methods have been successfully used, where spectroscopy-based methods have become less prevalent in recent studies. Figure 6 shows the acquisition timing, including the important start time and interval/temporal resolution, is quite variable across studies.

Figure 5: Acquisition methods used in published HP [1-13C]pyruvate human studies published up to September 2022, classified into: MR spectroscopy and spectroscopy imaging (MRS/I); chemical shift encoding methods, such as IDEAL, that use multiple TEs and model-based reconstructions; and metabolite-specific imaging methods that use spectrally-selective excitation to image a single resonance at a time.

Figure 6: Temporal acquisition characteristics reported in HP [1-13C]pyruvate human studies published up to September 2022. (a) Reported referencing of acquisition start times.

(B)

Acquisition start times reported when using dynamic imaging and when timing was reported relative to the end of the injection. (c) Temporal resolutions. “Not Applicable” indicates dynamic imaging was not used.

Summary

Three general categories of acquisition strategies have been used successfully for human HP 13C-pyruvate studies: MRS/I, model-based chemical shift encoding (e.g. IDEAL) methods, and metabolite-specific imaging methods. These have enabled successful studies in the prostate, heart, brain, abdomen, and breast. Recent studies increasingly have used the imaging-based strategies of metabolite-specific imaging and chemical shift encoding which are the fastest methods, although a heads-to–head comparison between techniques has not been performed.

Metabolite-specific imaging is quite popular because of its speed and compatibility with single-shot imaging, but is sensitive to B0 field variations and thus requires careful calibrations. Nearly all studies surveyed acquired data dynamically, allowing measurement of the bolus and metabolite kinetics. The exact timings and associated flip angles vary quite widely across reported studies, with no consensus yet as to how to choose these parameters. Image reconstruction is typically done directly using Fourier Transform methods, and accelerated imaging strategies are uncommon.

Data Analysis And Quantification

This section covers the analysis of data from human HP [1-13C]pyruvate studies, including modeling and metrics, visualization, as well as considerations for how to store data and metadata. Depending on study design, the analysis may need to give quantitative or semi-quantitative output reflecting a biological process or may just reflect a contrast between different regions of interest for quantitative evaluation.

Metrics

Figure 7: HP [1-13C]pyruvate raw data (A) have typically been quantified using four categories of metrics depending on the acquisition. Data acquired as a single time point are often quantified using normalized metabolite images or metabolite ratios (B). Dynamic data can be quantified using normalized metabolite images or metabolite ratios (B), or with metabolite timings such as time-to-peak (TTP) or pharmacokinetic (PK) models (C). The latter two require the data to be time-resolved. [1-13C]alanine and 13C-bicarbonate are analyzed similarly to [1-13C]lactate but omitted here for display.

Metabolite images are commonly used as summary metrics for HP MRI data, often including some form of normalization as well as summed over time as an area under the time curve (AUC) (17). These are analogous to the visual evaluation that is most used for routine clinical work (89,90). In these metabolite images, we expect that the [1-13C]pyruvate AUC signal is predominantly weighted towards perfusion and uptake, while [1-13C]lactate, [1-13C]alanine and 13C-bicarbonate AUCs represent metabolic conversion. The strength of this approach lies in its simplicity and relatively few underlying assumptions. Limitations to the use of single-metabolite images or AUCs include sensitivity to inhomogeneous coil profiles (57,87,91), the acquisition strategy and acquisition parameters, pyruvate polarization and concentration level, and signal relaxation rates (92). Further, the reader must be careful to interpret all the images in conjunction to better understand the underlying biology; for example, increased [1-13C]lactate in the presence of decreased [1-13C]pyruvate delivery can have a very different meaning compared to increased [1-13C]lactate with increased [1-13C]pyruvate delivery.

In an attempt to address variations in coil sensitivity, polarization level, and pyruvate delivery, AUC images are often computed by normalizing to a specified parameter, such as the maximum pyruvate or average lactate signals, or presented as a ratio such as lactate/pyruvate or divided by “total Carbon” - the sum total of HP 13C signal observed across all metabolites. The AUC ratios between metabolites and pyruvate are proportional to the corresponding forward kinetic rates (81,93), but are not directly comparable to rate constants when magnetization loss rates (e.g. relaxation and losses due to signal excitation) differ between studies. Similarly, the ratios between the produced metabolites (e.g. bicarbonate/lactate) can reflect the balance between downstream metabolic pathways (12,55). Care must be taken to consider how AUC images are calculated and normalized before comparing values between studies.

To further quantify the interpretation, pharmacokinetic (PK) modeling approaches were developed to compute the apparent kinetics of pyruvate-to-metabolite exchange (92,94–99). These yield semi-quantitative to quantitative apparent rate constants, given in s-1. Some models require a vascular input function, while others avoid this requirement (95). PK models can explicitly account for acquisition-specific details such as excitation angle and repetition time, and thus may reduce the effects of these details on quantification. An input-less model, provided in the Hyperpolarized-MRI-Toolbox (https://github.com/LarsonLab/hyperpolarized-mri-toolbox) (100) and thus frequently employed for human data, has been shown to fit well and robustly to prostate and brain data (8,20). PK models are quantitative in nature, arguably provide more relevant biological information (8,20), and appear to be reproducible across sites (51). However, rate constants derived from PK models are still apparent rates, and likely do not reflect a single biological characteristic.

Some additional considerations include whether complex or magnitude data is used, as the noise behaviors will impact the analysis differently. Additionally, cut-off thresholds or other criteria may be used to identify and avoid voxels with insufficient SNR before analysis to improve robustness (20,41).

Regardless of the analysis approach, the underlying biology is not always clearly represented by the data; instead, the metrics may be influenced by perfusion, barrier permeability, intercellular shuttles, enzyme activities, co-substrate concentrations, or combinations thereof, depending on the organ and disease of interest (19,43,94,101–103). This may be addressed by incorporating complementary information. As an example, HP 13C pyruvate data is influenced by perfusion, and thus addition of perfusion MRI could be important for interpretation (98,104,105).

All the methods outlined above have been explored in clinical studies, described in Supporting Table 3 and summarized in Figure 8. As of September 2022, approximately 52% of studies involving human subjects report rate constants derived from a PK model with a few different models reported. A nearly equal fraction (51%) of the studies report AUC ratio values.

Approximately 66% of these studies report metabolite-specific images or AUC values. About 40% report SNR values; this metric is particularly frequent in manuscripts that describe technical developments for clinical HP MRI. Approximately 16% of these studies summarize model-free metrics, and 10% report measurements from a single timepoint. Most studies report a combination of quantities.

Figure 8: Reported metrics used for analysis in HP [1-13C]pyruvate human studies published up to September 2022.

Visualization

A wide variety of approaches have been used for visualizing data from human HP 13C-MRI studies. The challenges and practical considerations are: 1) choosing the appropriate metrics to display, 2) how to encode the parameters (e.g. the colormap), and 3) choosing how to provide anatomical context and other multi-parametric data. The choice of visualization also depends on the goal which could be for diagnostic interpretation, but also quality control, reproducibility among readers and publication.

Metrics

The choice of HP 13C metrics is described in detail above. At this stage in HP 13C development where there is no standardized metric, often a combination of metabolite images and ratios or PK model parameters are shown.

Parameter Encoding

The mapping function chosen should provide an adequate, often quantitative, impression of the parameter mapped. There is a consensus in the visualization field that perceptually uniform maps are best suited to visualize continuous parameters, like the greyscale typically used by radiologists as well as other monochrome (black to blue) and color ranges (fire-type, rainbow-type) (106,107). Multi-color heatmaps have been the most frequently employed method for HP 13C data, while greyscale has infrequently been used but it ensures there is no coloring-based bias as well as facilitating later reuse (Fig. 9a). Among the color schemes employed in the clinical HP 13C literature, fire-type scheme seems to be the most common [similar to “Plasma” or “Inferno” in matplotlib.org]. Next most commonly employed is the rainbow-type scheme [similar to “Rainbow” in matplotlib.org].

Anatomical Context

HP MRI faces the challenge that it does not necessarily depict the anatomical features, similar to PET, and thus requires an anatomical reference. Most often, a grayscale anatomical image is overlaid with a HP colormap (Fig. 9c,d). This approach is very intuitive, but can skew perception as the grey-scale anatomical reference may affect the brightness of the HP data (e.g. signal in the skull). This bias does not occur when showing adjacent maps (Fig. 9a, b). Here, anatomical outlines may help to provide reference (Fig. 9b).

Related Journal Articles & DOI Links

Selected peer-reviewed publications relevant to 12 Lead ECG Acquisition. Click the DOI to access the full paper (may require institutional access).

Why Choose Us?

Bangalore guidance for robotics, Spectre and autonomous systems projects.

Spectre & Simulation

Gazebo, cloud twin and Webots worlds with navigation, SLAM and control stacks.

Control & Planning

Compliance, deep learning control, path planning and behavior trees.

Hardware Bring-up

Motors, sensors, ESP32/STM32 firmware and HIL validation paths.

Report & Viva

University-format documentation, PPT and viva preparation.

FAQ

Spectre, Gazebo, NVIDIA cloud twin, MATLAB/Simulink, Webots, Blynk / ThingSpeak, plus Arduino/STM32/ESP32, cameras, LiDAR and motor drivers.
Yes — simulation packages, hardware guidance, report, PPT and viva Q&A.