Abstract
The growth of e-commerce has intensified the need for last-mile delivery systems that can jointly manage customer choice and operational efficiency. We study the Dynamic Offering and Pricing of Mixed Delivery Options (DOPMDO) problem, in which a logistics service provider dynamically selects and prices attended home-delivery time slots and out-of-home pickup op- tions for sequentially arriving customers. Each decision affects immediate revenue, customer acceptance, and the route-dependent fulfillment cost realized at the end of the booking horizon.
We formulate DOPMDO as a finite-horizon Markov decision process and propose State-Value Anchored Pricing via Approximate Dynamic Programming (SVAP–ADP). The method learns a continuation-value approximation on an aggregate state representation of the mixed-delivery system and uses accepted-versus-rejected value differences to estimate option-level opportu- nity costs. These opportunity costs are embedded in an anchored trust-region pricing problem around a calibrated fine-static benchmark. Computational experiments on the real-world Seat- tle instance show that SVAP–ADP increases mean episode profit by 7.0% relative to current practice (95% CI: 6.7–7.2%), primarily by reducing terminal fulfillment cost while maintaining a stable home–locker–opt-out mix. Assortment-control experiments show that dynamic pricing is most effective when the menu exposes operationally valuable locker alternatives, with richer can- didate pools delivering substantially larger gains than restricted nearest-locker menus. These results indicate that anticipatory pricing and assortment control are complementary: pricing steers customers toward lower-cost options, but the menu determines whether high-value con- solidation opportunities are available in the first place. Ablation and sensitivity analyses further show that stable algorithm performance requires forward-looking opportunity-cost estimation, a sufficiently rich aggregate state representation, and anchored price control.
Ntroduction
Logistics service providers (LSPs) face a persistent profitability challenge in last-mile delivery. Al- though parcel volumes continue to grow, thin profit margins are often eroded by the high operating cost of attended home delivery (AHD), which requires vehicles to visit individual residences within promised service windows. Long handling times, urban congestion, parking constraints, and failed delivery attempts further increase the cost burden of AHD operations (Chen et al., 2017; Dalla Chiara and Goodchild, 2020; Ranjbari et al., 2023). With global parcel volumes increasing from 64 billion in 2016 to more than 161 billion in 2022 and expected to reach 225 billion by 2028 (Pitney Bowes, 2023), these inefficiencies might translate into mounting financial and environmental pressure on logistics operations.
Out-of-home (OOH) delivery has gradually become an important operational alternative to miti- gate these inefficiencies. Instead of serving each parcel at the customer’s residence, LSPs consolidate
Arxiv:2609.13539V1 [Math.Oc] 11 Sep 2026
deliveries at shared pickup locations such as parcel lockers, retail outlets, or grocery stores. The adoption of parcel lockers has expanded rapidly across Europe in recent years (Pinchasik et al., 2025); in Sweden, delivery points and parcel lockers account for 75% of parcel deliveries (Vakulenko, 2023). By aggregating multiple parcels at common pickup points, OOH delivery can reduce driving distances (Enthoven et al., 2020), shorten service times, and lower the risk of failed delivery (Akker- man et al., 2025; Savelsbergh and Van Woensel, 2016), while remaining attractive to customers who value flexible pickup times.
Despite these advantages, operating AHD and OOH delivery jointly is far from straightforward. Customers differ in their preferences for delivery modes, time windows, prices, and pickup distances. Moreover, customer requests arrive sequentially during a booking horizon, while the final routing cost is realized only after all accepted orders have been confirmed. Each accepted order changes the spatial and temporal structure of the fulfillment problem and therefore changes the marginal value of future orders. A home-delivery acceptance may increase route fragmentation, while a locker acceptance may improve consolidation but require a discount to attract the customer. Conversely, overly aggressive surcharges or discounts may either induce excessive opt-outs or erode revenue. The LSP must therefore make real-time offering and pricing decisions that balance immediate revenue, customer choice, fleet capacity, and long-horizon route consolidation.
Dynamic offering and pricing have emerged as promising control levers for steering demand toward delivery options that are less costly to fulfil (Akkerman et al., 2025; Galiullina et al., 2024; A. Strauss et al., 2021; Yang and A. K. Strauss, 2017). By selecting a tailored menu of AHD time slots and OOH pickup options, and by attaching option-specific surcharges or discounts, an LSP can influence the distribution of accepted demand across delivery channels. However, designing effective policies for mixed AHD–OOH systems raises three methodological challenges. First, the system state contains all accepted orders, their selected delivery modes, time-window commitments, and tentative route structure, making exact dynamic programming computationally intractable. Second, the economic value of an accepted order depends not only on its immediate revenue but also on its effect on future capacity and route consolidation. Third, learned or myopic price adjustments can be unstable: if they depart too far from a well-calibrated static policy, they may sacrifice revenue, induce excessive opt-outs, or create fragmented route plans that are costly to serve. Existing studies have made important progress on dynamic pricing for home delivery, time-slot management, and OOH delivery, but they typically simplify at least one of the key dimensions: the joint assortment of AHD and OOH options, the sequential nature of customer arrivals, or the route-dependent fulfillment cost induced by mixed delivery choices.
In this paper, we study the Dynamic Offering and Pricing of Mixed Delivery Options (DOPMDO) problem. Customers arrive sequentially, and the LSP must decide which AHD time slots and OOH pickup options to offer and how to price them before the final fulfillment routes are known. Each decision therefore affects immediate revenue, customer choice, and the terminal routing cost of the accepted orders. We formulate this problem as a finite-horizon Markov Decision Process (MDP) and propose State-Value Anchored Pricing via Approximate Dynamic Programming (SVAP–ADP). The method estimates aggregate post-decision state value, translates accepted-versus-rejected value dif- ferences into option-level opportunity costs, and applies these costs within an anchored trust-region pricing problem. In this way, SVAP–ADP uses dynamic prices to exploit consolidation opportunities while keeping customer-facing prices close to a reliable static benchmark.
The main contributions are as follows. 1. We introduce the DOPMDO problem as a sequential decision model for mixed attended home delivery and out-of-home pickup. The formulation integrates assortment selection, option-level pricing, stochastic customer choice, and route-dependent terminal fulfillment cost. It captures the key operational trade-off that an accepted order generates immediate revenue but also reshapes future consolidation and routing opportunities.
2. We develop SVAP–ADP, an anchored value-based pricing method for large-scale mixed delivery versus-rejected value differences into option-level opportunity costs, and combines these costs with immediate routing information inside a trust-region pricing problem. The resulting policy is forward-looking while remaining stable and interpretable.
3. We evaluate SVAP–ADP on a real Seattle last-mile delivery network with both home-delivery and parcel-locker options. The experiments compare against a current-practice Static-uniform policy, a calibrated Static-fine benchmark, a myopic routing-aware policy, a compact value- function variant, and an unanchored FreePrice ablation.
Svap–Adp Improves Profit Over
both static baselines primarily by reducing terminal fulfillment cost, rather than by increasing revenue or suppressing demand. The results show that immediate insertion costs are too local to capture future consolidation value, unrestricted learned prices can destabilize the demand mix, and a richer aggregate state representation is needed for reliable opportunity-cost pricing. The sensitivity analyses further show that pricing and assortment control are complementary: prices steer demand, while the candidate pool determines whether operationally useful locker options are available.
The remainder of this paper is organized as follows. Section 2 reviews related work on dynamic offering, pricing, AHD and OOH delivery. Section 3 introduces the DOPMDO problem and its MDP formulation. Section 4 presents the SVAP–ADP method. Section 5 reports the numerical study on the Seattle network. Section 6 concludes the paper and discusses directions for future research.
Iterature Review
In this section, we first review demand management for AHD, followed by studies on OOH delivery and mixed-channel delivery design. We then discuss sequential decision methods for integrated demand management and vehicle routing problems, and position our paper at the intersection of
Emand Management In Attended Home Delivery
Demand management for AHD studies how an LSP can influence customer choices before the fi- nal delivery routes are constructed. Early work showed that delivery-slot availability, incentives, and acceptance decisions can be used to steer demand toward operationally attractive regions and time windows. Campbell and Savelsbergh (2006) study dynamic acceptance and incentive schemes for home delivery services, highlighting the operational value of influencing customer time-window choices. Agatz et al. (2011) consider time-slot management at the area level and use continuous approximations to anticipate delivery costs. Studies by Ehmke and Campbell (2014) and Visser and Savelsbergh (2019) further demonstrate that controlling slot availability can improve routing feasibility and reduce fulfillment cost. Recent reviews by Waßmuth et al. (2023) and Fleckenstein et al. (2023) provide broader classifications of AHD demand management and integrated demand A second branch of this literature studies pricing rather than pure availability control. Asdemir et al. (2009) formulate dynamic time-slot pricing with a multinomial logit choice model, but rely on simplified capacity representations. Yang et al. (2016) combine a choice-based pricing model with dynamically estimated delivery costs, using real e-grocery data to show that customers can be steered through delivery charges.
Yang and A. K. Strauss (2017) extend this line by using approximate dynamic programming (ADP) to account for future revenue and routing-cost effects. R. Klein et al. (2019) study differentiated static time-slot pricing under routing considerations, while Koch and R. Klein (2020) develop a route-based ADP approach for dynamic pricing in AHD.
Vinsensius et al. (2020) similarly study dynamic incentives for slot management in e-commerce AHD. More recently, Abdollahi et al. (2023) incorporate forecast orders into dynamic routing to support time-slot demand management. Abdolhamidi and Lurkin (2025) further incorporate heterogeneous customer preferences into the design of slot assortments and price discounts. Their numerical results show that using a mixed logit model enables the LSP to tailor its service offerings more effectively, thereby increasing profitability and improving customer retention.
Several papers relax the classical assumption that customers choose exactly one narrow time window. Yildiz and Savelsbergh (2020) study discounts for delivery-time flexibility across multiple periods. A. Strauss et al. (2021) introduce flexible time slots, where customers accept uncertainty about the final delivery window in exchange for a lower delivery charge, and develop a dynamic pricing policy informed by approximate opportunity costs. These studies demonstrate the value of using pricing to obtain operational flexibility. However, they focus mainly on temporal flexibility in home delivery. They do not jointly decide prices and assortments over home-delivery time slots and spatially distinct out-of-home pickup locations.
Out-of-home delivery and mixed delivery-option design Out-of-home (OOH) delivery has received increasing attention as a way to consolidate parcels, reduce failed deliveries, and improve last-mile efficiency. Empirical and optimization studies show that parcel lockers and pickup points can reduce delivery effort and create operational economies of density. Ranjbari et al. (2023) provide field evidence on the effect of parcel lockers on delivery times.
Enthoven et al. (2020) study a two-echelon vehicle routing problem with covering options, in which parcel lockers serve as shared delivery locations. Mancini and Gansterer (2021) introduce vehicle routing with private and shared delivery locations, combining home delivery with alternative delivery points. Lin et al. (2022) and Lyu and Teo (2022) study parcel-locker location and locker alliance network design, while Mancini et al. (2023) address locker-location planning under uncertainty in demand and capacity. Janinhoff et al. (2024) provide a recent review of OOH delivery optimization, covering facility location, routing, location-routing, and emerging operational challenges.
Most OOH studies, however, are static: customer locations, potential delivery options, or aggre- gate demand are known before decisions are made. Some works incorporate customer preferences or incentives, but typically without modeling a fully sequential booking process. Dumez et al. (2021) investigate a setting in which customers can actively specify delivery options, each associated with time windows and preference levels. They formulate the problem as a mixed-integer linear program and develop a large neighborhood search method to solve it. Their results show that the proposed formulation can generate substantial cost savings while maintaining a high quality of service. Zhang et al. (2023) study joint location and pricing optimization for self-service delivery under customer choice, with a focus on strategic facility and price design. Galiullina et al. (2024) examine demand steering with home and pickup-point delivery options, integrating incentive and routing decisions under uncertain customer acceptance.
Zhou et al. (2026) investigate a mixed last-mile delivery problem with heterogeneous customer behavior and formulate it as a two-stage stochastic program, showing that joint assortment and pricing decisions can substantially improve delivery performance.
Together, these studies demonstrate the operational value of steering customers toward pickup points and jointly managing home and OOH delivery options. However, they all assume that the relevant customer demand is known before decisions are made, and therefore do not address the real-time problem of jointly selecting and pricing both home time slots and nearby OOH locations as customers arrive sequentially.
The most closely related OOH study is Akkerman et al. (2025), who introduce dynamic selection and pricing of OOH delivery options. They formulate a sequential decision problem in which each arriving customer can be offered a subset of OOH locations with discounts or charges, and they propose a machine-learning policy using a spatial-temporal state representation.
Their Work Is
important because it moves OOH delivery from static optimization to sequential offering and pricing. Nevertheless, their setting focuses on choosing between home delivery and OOH locations without modeling attended home-delivery time slots.
N Contrast, Our Problem Considers A Mixed Menu
consisting of multiple AHD time slots and multiple nearby lockers, so the provider must jointly manage both the temporal dimension of AHD and the spatial consolidation dimension of OOH delivery.
Sequential Decision For Last-Mile Delivery
Because customers arrive dynamically and each accepted order reshapes both the routing problem and future acceptance opportunities, many last-mile pricing problems are naturally modeled as finite-horizon Markov decision processes. Exact dynamic programming is generally intractable at realistic scales, so the literature relies on decomposition, opportunity-cost approximation, ADP, simulation-based policies, and, more recently, machine learning. Fleckenstein et al. (2023) review integrated demand management and vehicle routing problems and emphasize the common structure in which providers control prices, availability, or acceptance decisions during a booking horizon while fulfillment is performed by a vehicle fleet. (2025) further analyze the concept of opportunity cost in integrated demand management and vehicle routing, showing why opportunity- cost approximation is more difficult than in classical revenue management: the value of a request depends jointly on routing interactions, capacity consumption, and displaced future revenue.
ADP-based pricing has been particularly influential in AHD and same-day delivery. Yang and A. K. Strauss (2017) use ADP to estimate the future value of capacity and routing resources for time- slot pricing. Koch and R. Klein (2020) construct route-based value approximations for dynamic AHD pricing. Ulmer (2020) studies dynamic pricing and routing for same-day delivery, where booking and service horizons overlap, and proposes an anticipatory pricing and routing policy. V. Klein and Steinhardt (2023) combine dynamic demand management with online tour planning for same-day delivery, optimizing delivery spans and prices for incoming requests. Banerjee et al. (2025) consider pricing and demand management for integrated same-day and next-day delivery systems, further illustrating the growing interest in coordinating pricing with fleet operations. These studies confirm the value of anticipatory pricing, but most of them control temporal delivery promises rather than mixed AHD–OOH delivery options. Machine-learning methods have also begun to appear in last-mile demand management. Akkerman et al. (2025) use a convolutional neural network with a spatial- temporal state encoding for dynamic OOH selection and pricing.
These Methods Are Attractive
because they can capture complex spatial and temporal interactions. At the same time, they may be unstable in revenue-management settings if learned adjustments deviate too far from a reliable benchmark policy, especially when value estimates are noisy.
To position our study relative to the most closely related literature, Table 1 classifies representa- tive papers by decision setting, demand-management lever, methodology, choice model, and delivery channel. The table highlights that existing dynamic offering and pricing studies typically focus ei- ther on AHD time slots or on OOH delivery locations, while studies that combine home and OOH delivery are mostly static or do not jointly control both assortment and prices in an online setting.
Our paper addresses this gap by introducing the DOPMDO problem as a sequential decision model for mixed last-mile delivery. Unlike prior studies that focus primarily on AHD time-slot control, OOH selection and pricing, or static mixed-channel design, we jointly consider dynamic pricing and assortment control for both attended home-delivery time slots and out-of-home pickup locations as customers arrive, while accounting for stochastic choice behavior and route-dependent terminal fulfillment cost.
Problem Description And Model
We study the DOPMDO problem faced by a last-mile delivery provider during a booking horizon before route execution. Customers arrive sequentially. At each arrival, the provider observes the Table 1: Comparison of related research on last-mile demand management.
Ahd+Ooh
Note. D = dynamic; S = static; O = option, assortment, location, or assignment decision; P = pricing or incentive decision; O+P = joint option and pricing decision. AHD = attended home delivery; OOH = out-of-home delivery; ADP = approximate dynamic programming; CNN = convolutional neural network; LP = linear programming; MIP = mixed-integer programming; MNL = multinomial logit.
customer’s location and feasible delivery alternatives, then chooses which options to display and what price adjustment to attach to each displayed option.
The Option Set Combines Ahd In A
finite set of delivery time windows and OOH delivery to parcel lockers. The customer’s response is stochastic. Accepted orders are fulfilled only after the booking horizon ends, so each online decision affects both immediate revenue and the route-dependent fulfillment cost realized at the terminal stage.
Problem Narrative
The provider operates a fleet V = {1, . . , V } of homogeneous vehicles with capacity Q from a single depot. The service region is represented by a travel-time network G = (N 0, E), where N 0 contains the depot, customer home locations revealed during the booking horizon, and a fixed set L of parcel lockers. Travel time between locations i and j is denoted by τ(i, j). The set of AHD time windows is denoted by S. A home-delivery option assigns the current order to the customer’s home location in one of these time windows. A locker option assigns the parcel to one of the candidate lockers considered for the current customer.
The booking horizon contains a finite sequence of customer-arrival opportunities.
We Index
decision epochs by k and let D denote the maximum horizon length used in the dynamic program and in the value approximation. The realized arrival process may terminate before D; this is represented by a terminal realization of the exogenous information. At epoch k, the current customer is denoted
K
. The observed customer information consists of the home location hk, the service or load requirement bk, and a finite candidate locker set Lk ⊆L. The candidate locker set may be generated from nearby lockers, active lockers, or another customer-relevance rule specified by the online policy; the MDP only requires it to be finite and known before the decision.
The provider controls two customer-facing levers. First, it chooses the displayed assortment, namely the subset of feasible AHD and OOH options shown to the customer. Second, it chooses an option-specific price adjustment. Price adjustments are measured relative to the base order revenue r: positive values are surcharges and negative values are discounts. The customer chooses from the displayed menu according to a stochastic discrete-choice model with an outside option.
The economic trade-off is dynamic. AHD orders generate revenue but require residential visits within selected time windows, which may fragment routes across both space and time. OOH orders can consolidate multiple parcels at shared pickup locations, but discounts used to stimulate OOH demand reduce immediate revenue. Opening many weakly used lockers may also increase route stops, while reusing an already active locker may add little or no additional travel. Because the final routes are constructed only after all accepted orders have been confirmed, the provider must price and offer current options based on both current revenue and the future operational flexibility left for subsequent customers.
Arkov Decision Process
We formulate DOPMDO as a finite-horizon MDP. The decision epoch begins after the current customer has arrived and revealed the information needed for offering and pricing. Let
(1)
denote the booking state inherited from previous decisions, where k is the epoch index, Ωk is the set of accepted orders, and Pk is the tentative fulfillment plan used for feasibility checks and successor-
K
= (hk, bk, Lk) is the current customer information. Thus, Sk is the state on which the online decision is made, while Bk is the post-response booking state carried forward from the previous epoch.
The feasible option set is generated from Sk. AHD options are (hk, w) for w ∈S, and OOH options are lockers ℓ∈Lk. An option is feasible if the insertion operator can construct a feasible option-contingent plan. For home delivery, the plan must respect the selected time window and vehicle-capacity restrictions; for locker delivery, it must respect vehicle capacity and, when modeled, residual locker capacity. If the route plan already visits locker ℓ, an additional parcel assigned to ℓ can be consolidated at that stop.
A decision at epoch k is denoted by xk = (Uk, pk, {P o
K }O∈Uk). Here, Uk Is The Displayed Option
set, pk(o) is the price adjustment for option o, and P o
K Is The Option-Contingent Plan Obtained If
option o is accepted. The feasible action set is denoted by A(Sk). The empty menu is included in A(Sk) and represents a provider-controlled rejection action: if Uk = ∅, then qk(0 | Sk, xk) = 1 and
The System Follows The Opt-Out Successor B0
k+1. The customer’s choice follows a multinomial logit model. The systematic utility of option o is
(3)
where µw is the baseline utility of AHD slot w, µℓis the OOH baseline utility, βp < 0 is price sensitivity, βd > 0 is locker-distance sensitivity, and d(hk, ℓ) is the distance from the customer’s home to locker ℓ. Normalizing the outside-option utility to zero, the choice probabilities are
(4)
Although the choice probabilities depend on xk only through the displayed menu Uk and prices pk, conditioning on (Sk, xk) keeps the notation aligned with the Bellman recursion and the option- contingent successor states. If Uk = ∅, then qk(0 | Sk, xk) = 1.
Let uk ∈Uk ∪{0} denote the realized choice, where 0 is the outside option. If uk = 0, no order is appended. If uk = o ∈Uk, the current order is added to the accepted-order set and the tentative
Plan Becomes P O
k . The corresponding booking-state successors are
(5)
The next decision state is formed by combining the realized successor booking state with the next customer request, unless the booking horizon has ended. At the terminal stage, the provider executes routes to serve all accepted orders. The realized fulfillment cost is recomputed from the terminal booking state and decomposes into three cost
Τ
is fixed route and vehicle cost, and γ is a fulfillment-cost multiplier. The provider seeks a policy π that maximizes expected profit,
(7)
where pk(0) = 0 by convention and τ denotes the terminal epoch. Thus, revenue r+pk(uk) is earned only when the customer accepts an offered option, an opt-out incurs penalty πout, and fulfillment cost is incurred once at the terminal booking state.
Ecision-State And Post-Decision Value Functions
Let Vk(Sk) denote the optimal expected profit from epoch k onward after the current customer has been observed. The value is defined on the full decision state Sk = (Bk, Cnew
) Because The Current
customer determines the feasible options, choice probabilities, and price adjustments. The exact
(8)
where qk(u | Sk, xk) is the MNL choice probability, Rk(0; xk) = −πout for opt-out, and Rk(o; xk) = r + pk(o) for o ∈Uk.
The Post-Decision Value V X
k+1(B) is the conditional expected continuation value after the current customer’s outcome has been incorporated, given the resulting booking state B, but before the next
(9)
where eCk+1 is either the next customer request or the terminal signal ∅. We use the terminal
(10)
so the route-dependent terminal cost enters the recursion through the terminal realization.
(11)
This sequence separates the transient customer request from the persistent booking state.
The
current customer enters the decision state only to define the feasible options, choice probabilities, and option-contingent route updates. After the response, the system carries forward only the updated
Booking State Buk
k+1, and the next decision state is formed when the next request is realized. The MDP formulation makes clear why exact dynamic programming is not computationally viable: the booking state contains both the accepted-order set and the tentative fulfillment plan, so the state space grows rapidly over the booking horizon. This motivates the solution approach developed next. Rather than evaluating the full Bellman recursion, SVAP–ADP approximates the continuation value of post-response booking states using aggregate summaries of Bk = (k, Ωk, Pk).
Each feasible option is then assessed by comparing the successor state created by accepting that option with the opt-out successor, producing an option-level opportunity-cost signal for the online pricing problem.
Solution approach: State-Value Anchored Pricing via ADP The Bellman equation (8) shows why online offering and pricing require anticipation. If the con- tinuation value of each successor booking state were known, the provider could evaluate the future consequence of accepting each feasible option and solve a one-step menu-pricing problem at every arrival. This is the logic used in dynamic time-slot pricing and same-day pricing-and-routing mod- els: the dynamic program is too large to solve exactly, but value differences quantify the future cost of accepting a current order. In DOPMDO, these value differences must account for both AHD time-window commitments and OOH locker consolidation.
We propose State-Value Anchored Pricing via Approximate Dynamic Programming (SVAP– ADP). The method has three design elements. First, it learns an aggregate approximation of the post-decision continuation value from simulated booking trajectories. Second, it estimates option- level opportunity costs by comparing the value of the opt-out successor state with the value of the accepted successor state. Third, it converts these opportunity costs into customer-facing prices within a trust region around a calibrated fine-static price anchor. The value approximation pro- vides anticipation; the anchor and trust region stabilize the price optimization and prevent extreme reactions to approximation error.
Opportunity costs and the one-step pricing problem The post-decision value in (9) allows the Bellman recursion to be written in the same opportunity- cost form used in dynamic pricing for delivery time slots. An accepted option earns immediate revenue, whereas an opt-out incurs the immediate penalty πout ≥0. Substituting (9) into (8) gives
(12)
For an offered option o ∈Uk, define the exact opportunity cost of acceptance as the loss in conditional post-decision continuation value relative to the opt-out successor:
Because V X
k+1(·) is a future-only post-decision value, ∆k(Sk, o) measures the expected future value displaced by accepting the current customer under option o rather than carrying forward the opt-out successor state. This displacement includes the expected change in terminal fulfillment cost, future revenue opportunities, future opt-out penalties, and operational flexibility. It does not include the current immediate revenue or the current opt-out penalty; these current-epoch terms enter the pricing objective explicitly below. Note that the superscript x in V x
K+1 Denotes The Post-Decision
value, not dependence on the current action. The rearrangement follows by adding and subtracting the opt-out continuation value inside the
(15)
The first term is the conditional continuation value of the opt-out successor booking state and is common to all feasible actions. The maximized term contains the incremental expected contribution of the offered options and the expected opt-out penalty. Each accepted option contributes imme- diate revenue plus the price adjustment, net of the option’s opportunity cost. The outside option contributes no accepted-order revenue and incurs penalty πout. The opt-out term must remain inside the maximization because qk(0 | Sk, xk) depends on the displayed menu and prices. Thus, if the exact opportunity costs were known, the online decision would solve the one-step pricing problem
(16)
The maximization is over the insertion-feasible action set A(Sk) defined in Section 3.2: infeasible home time-window insertions, vehicle-capacity violations, and modeled locker-capacity violations are excluded before pricing, and the empty menu remains feasible as a provider-controlled rejection action. When πout = 0, (16) reduces to the no-penalty opportunity-cost pricing problem. Each accepted option earns immediate revenue r + pk(o) and consumes opportunity cost ∆k(Sk, o). A low opportunity cost indicates that the option fits well with the current AHD route structure or locker consolidation pattern.
A high opportunity cost indicates that acceptance is expected to reduce future revenue opportunities, consume scarce operational flexibility, or increase the terminal fulfillment cost. The opt-out penalty discourages menus and prices that reduce operating cost only by pushing too many customers to the outside option. At the final booking decision, the terminal
+1),
so the last decision directly trades off immediate revenue, the expected opt-out penalty, and the change in terminal fulfillment cost. At earlier epochs, the same terminal cost and downstream demand effects affect pricing through the continuation values.
The exact opportunity cost in (13) is unavailable because V x is intractable. SVAP–ADP therefore learns an approximation bV x on aggregate summaries of the post-decision booking state and computes
(17)
The learned object is the post-decision value function; the pricing input is the accepted-versus-opt- out value difference. Current-epoch revenue and the current opt-out penalty are not absorbed into b∆k(Sk, o); they enter the one-step pricing objective explicitly. The remaining subsections describe the aggregate state representation, the offline value-learning procedure, and the anchored online pricing problem used to implement (16) with approximate opportunity costs.
Ore Aggregate Variables And Value Representation
The value approximation targets the post-decision value V x, not the full decision-state value Vk(Sk). The aggregate representation is therefore derived from the booking state Bk = (k, Ωk, Pk), after the previous customer outcome has been incorporated and before the next customer is observed. The
K
is excluded from the learned state representation because it affects the current feasible options, prices, and choice probabilities, but it does not persist after the customer either accepts an option or opts out.
The exact booking state is too detailed for direct value approximation because it contains all ac- cepted commitments and the tentative fulfillment plan. We therefore use a compact set of operational summaries and let the implementation-level basis be generated as deterministic transformations of these summaries. Let nk be the number of accepted orders in Ωk, and let nH
K Be The Number Of
accepted AHD orders. The number of accepted OOH orders is derived as nL
(K, Nk, Nh
k ) records booking progress, total committed load, and the realized AHD–OOH modal split. To describe where the accepted commitments are concentrated, we use two active-resource maps. Partition the service area into coarse home-delivery cells and combine each cell with each AHD time window. Let a = (g, w) denote a home cell–slot pair, and let mH
Ak Be The Number Of Accepted Ahd
orders assigned to pair a before epoch k. The active AHD map is
ℓk Be The Number Of
accepted OOH parcels assigned to locker ℓ. The active locker map is
(19)
understood together with the active locker loads {mL
K }. Finally, Let Ck Denote A Raw Route-
burden summary computed from the tentative plan Pk, such as the accumulated travel distance used by the insertion routine. This quantity is a state summary used for learning; it is not the terminal fulfillment cost Cful in the exact objective. Note that Pk is constructed first, and Ck is then evaluated as an aggregated feature or a metric of the current route plan. Table 2 summarizes the aggregate variables and operational dimensions used to construct the SVAP–ADP value approximation.
Table 2: Core aggregate variables and derived state dimensions used by SVAP–ADP.
Oncentration Or Fragmentation Of Home
commitments across cell–slot route patterns.
Reuse Of Active Lockers, Opening Of New
lockers, and singleton-locker fragmentation.
Route-Burden Information Not Captured By
counts and active-resource maps alone. The core aggregate information used by SVAP–ADP is
K . The Value
approximation is defined on a basis generated from X(Bk). Let ϕ(Bk) =
be the aggregate basis, where each non-intercept component is a deterministic transformation or normalization of the core variables in (20). The full basis implemented is reported in Appendix B. Let zj(Bk) denote the standardized value of basis component ϕj(Bk). The approximate post-
(21)
The basis functions may be nonlinear transformations of the aggregate state, but the approximation is linear in its coefficients. If a home option o = (hk, w) is accepted, then nk and nH
K Increase By
one, the load of the corresponding home cell–slot pair increases or a new pair is activated, and Ck is updated through the insertion operator. If a locker option o = ℓis accepted, then nk increases by
One, Nh
k is unchanged, the load of locker ℓincreases or a new locker is activated, and Ck is updated. If the customer opts out, only the epoch advances. Consequently, accepted-versus-opt-out value differences are computed by evaluating how a feasible option changes the aggregate booking state carried into the future.
Offline Learning Of The Post-Decision Value
The value model is trained offline using simulated booking trajectories. Let πR denote the rollout policy used to generate continuation-value labels. In the main implementation, πR is the calibrated fine-static policy because it is stable, competitive, and samples states in the region of the state space that the anchored dynamic policy is expected to visit. The seed sets used for static calibration, value-function training, validation, and final evaluation are disjoint.
For each training seed, we generate a complete arrival stream and the corresponding customer- choice random numbers. We simulate the booking horizon under πR and sample post-decision booking states along the trajectory. For a sampled booking state Bi at epoch ki, we clone the state and complete the remaining horizon under πR. Let eBi
Τi Be The Terminal Booking State Reached By
this rollout completion, where τi is the realized terminal epoch. Let Rev(B) denote the cumulative accepted-order revenue earned up to booking state B, and let N out
Denote The Number Of Opt-Out
outcomes observed during the cloned rollout completion after Bi. The continuation-profit label is
(22)
Equivalently, Yi is one Monte Carlo realization of the post-decision continuation value from Bi under the rollout policy πR. The label is future-only.
The Subtraction Of Rev(Bi) Removes All
accepted-order revenue earned before the sampled state, so past revenue is not counted again. The
I
counts only opt-outs that occur during the cloned continuation after Bi, so opt-out penalties incurred before the sampled state are not counted again. In contrast, the terminal cost is not differenced against a cost at Bi, because no fulfillment cost has yet been incurred at the post-decision state; the already accepted orders in Bi still have to be served at the terminal stage.
Thus, the learned value approximation represents future accepted-order revenue minus future opt- out penalties and terminal fulfillment cost, conditional on the sampled post-decision booking state. The labeled training set is D = {(ϕ(Bi), Yi) : i = 1, . . , N}. The coefficients in (21) are estimated from this labeled data using a linear regression estimator on the standardized basis. Appendix C summarizes the offline training process.
Candidate-menu construction and online SVAP–ADP policy The exact action space in (16) allows the provider to choose any insertion-feasible displayed menu. Enumerating all subsets of feasible AHD slots and locker options is computationally expensive and may also generate menus that are operationally irrelevant. We therefore implement SVAP–ADP on a restricted, state-dependent candidate-menu family.
Et Oh
k be the feasible AHD options for the current customer. The displayed AHD sub-menu
K
contains the KH feasible home slots with the smallest raw insertion costs. For OOH delivery, SVAP–ADP constructs a state-dependent locker candidate pool. The pool is built from three layers: the Knear nearest accessible lockers, up to Kact feasible active lockers already used in the current tentative plan, and up to Kinact feasible inactive lockers with high forecast residual catchment demand. Duplicates are removed, and the total pool is capped at Kpool. SCAP-ADP further ranks this pool lexicographically by four keys: The first key places active lockers before inactive ones, the second favors low adjusted opportunity cost, the third breaks remaining ties toward proximity to the customer, and the fourth favors dense locker catchments. The nearest accessible locker is always
Retained. The Displayed Locker Sub-Menu U L
k contains the top Koffer lockers after this ranking. The
K ∪U L
k . Section 5.4 then varies the locker-pool design to isolate the additional value of OOH assortment control. At each arrival, SVAP–ADP solves the approximate one-step problem over the feasible options and candidate menus generated by the construction procedure described above. It therefore prices only options that have already passed the insertion-feasibility filter and associated construction rules.
The candidate-menu step restricts the assortment to this manageable feasible set. For each feasible option o, the policy compares the opt-out successor B0
K+1, Both
constructed on cloned booking states. The estimated opportunity cost is
Algorithm 1 Online Svap–Adp Decision At Epoch K
1: Construct feasible options Ok(Sk), candidate menus Mk(Sk), and the opt-out successor B0 k+1.
Onstruct The Accepted Successor Bo
k+1 on a cloned tentative plan.
:
Compute b∆k(Sk, o) by (23) and ecη,k(o) by (24).
K
(o) for all o ∈U and solve (26) to obtain Jk(U) and p∗(U).
K
), observe uk, and update the actual booking state to Buk k+1. Because bV x is a future-only post-decision value approximation, immediate revenue is not included in (23); it enters only the pricing objective below. To stabilize the learned value difference, SVAP–ADP
(24)
where η ∈[0, 1] controls the ADP correction and Π[c,c] denotes projection onto the clipping interval. Prices are restricted to a trust-region grid around the fine-static anchor:
(25)
For each candidate menu U ∈Mk(Sk), SVAP–ADP solves
(26)
and then deploys the menu-price pair with the largest score. The finite grids make (26) small enough to solve by enumeration. Only the final update changes the actual system state; all earlier successor states are coun- terfactual evaluations.
Thus, SVAP–ADP differs from a myopic insertion-cost policy by replac-
Ing Cins
k (o) with the clipped post-decision value signal ecη,k(o), while the trust-region anchor keeps customer-facing prices close to the calibrated static benchmark. Algorithm 1 summarizes this online implementation.
It is worth mentioning that the online SVAP–ADP decision is a regularized approximation of the exact one-step opportunity-cost pricing problem. Three restrictions separate the implemented action from the full-information one-step optimizer. First, the candidate menu family Mk(Sk) restricts the assortment space. Second, the trust-region grid GTR k (o) restricts price deviations from the fine-static anchor. Third, the exact opportunity cost ∆k(Sk, o) is replaced by the blended and clipped estimate ˜cη,k(o) obtained from the learned post-decision value model.
Seattle Case Study
In this section, we evaluate SVAP–ADP on a real-world case based on the greater Seattle delivery network. Section 5.1 introduces the benchmark policies. Section 5.2 describes the instance design. Sections 5.3–5.6 report the experimental results.
Benchmark Policies
We compare SVAP–ADP with five benchmarks designed to isolate the contribution of each modeling component: current-practice uniform pricing, static price calibration, myopic routing awareness, compact state aggregation, and anchored trust-region pricing. The first benchmark, Static-uniform, represents a common real-world pricing practice in which the provider posts one fixed home-delivery price and one fixed locker-delivery price throughout the booking horizon. The second benchmark, Static-fine, is the calibrated static anchor πF used by the anchored dynamic policies. All policies are evaluated on the same customer streams and choice-uniform streams, i.e., common uniform random- number streams used to sample realized MNL choices. Unless otherwise stated, they use the same feasible-option construction, locker candidate pool, and customer-choice model. Thus, performance differences are attributable to the pricing and opportunity-cost logic rather than to differences in the available assortment. Table 3 summarizes the design.
Table 3: Benchmark design. Each policy is defined by its opportunity-cost estimate, value-function form, and pricing feasible set.
Full Method
1. Static-uniform pricing (πU). This policy represents a simple current-practice pricing rule used in many operational settings: the provider posts one fixed home-delivery price and one fixed locker-delivery price throughout the booking horizon. Formally, pU(h, w) = pU
Slot W ∈S And Pu(ℓ) = Pu
ℓfor every locker option ℓ. 2. Static-fine pricing (πF ). Slot-specific home surcharges and zone-segmented locker prices are calibrated offline on dedicated seeds, and the resulting price map pF (o) is applied at every ar- rival independently of the accumulated state. This is the fine-static anchor of SVAP–ADP; it absorbs the systematic price heterogeneity available to a policy that cannot observe the evolv- ing consolidation and routing structure. Comparing Static-fine with Static-uniform quantifies the value of offline price segmentation and calibration, while comparing SVAP–ADP with Static-fine quantifies the additional value of dynamic state-dependent pricing.
3. Myopic pricing. The opportunity-cost estimate collapses to the immediate raw insertion cost. Prices are optimized using the same anchored trust-region pricing structure as SVAP–ADP, but the value-function component is removed. Pricing therefore responds only to the current routing state, with no anticipation of future consolidation value or capacity scarcity. This policy isolates the contribution of forward-looking opportunity-cost estimation.
4. SVAP–ADP–6F. The anticipation channel is active, but the value function is restricted to a compact six-feature aggregate representation defined in Eq. (20). Compared with SVAP–ADP, this benchmark changes only the value basis: it replaces the final 17-feature aggregate basis with the compact six-feature basis. All other components are kept identical to SVAP–ADP.
This policy isolates the value of the richer state representation used by SVAP–ADP. 5. SVAP–ADP–FreePrice. The final state-aggregation value function and the blended opportu- nity-cost estimate of SVAP–ADP are retained, but the trust-region anchor is removed: prices are optimized over the full mode-specific grids Gκ(o) rather than the local grids GTR
K (O). The
learned correction is then free to post any feasible surcharge or discount, with no restriction to a neighborhood of the static anchor pF (o). With the opportunity-cost estimate held identical to SVAP–ADP, the single change is the pricing feasible set; this policy isolates the stabilizing contribution of the anchored trust region and exposes the customer-facing prices to value- approximation error.
6. SVAP–ADP. The full method described in Section 4, combining the dimension-enriched aggre- gate state-value approximation, the blended opportunity-cost estimate, and anchored trust- region pricing.
All policy comparisons use a paired Monte Carlo design. For each evaluation replication, we generate one customer-arrival stream and one stream of customer-choice random numbers, and then replay the same streams for every policy. Thus, each policy faces the same realized demand and latent choice shocks; only the offered menus, prices, and opportunity-cost logic differ. This common-random-number design reduces simulation noise and makes policy lifts interpretable as paired differences rather than as differences between unrelated simulations. The final performance evaluation uses 400 held-out Monte Carlo episodes. For policies with learned value functions, re- ported averages are taken over the same 400 episodes and 10 independently trained value models.
We report mean episode outcomes and 95% paired or nested confidence intervals for profit lifts. The simulator is implemented in Python. Online insertion costs are produced by a cheapest- insertion heuristic; the offline routing backend adopts a Clarke–Wright savings method.
N The
computational experiments, bθ is estimated by ridge regression to stabilize estimation under correlated with a 1.9 GHz Intel Core i7-1370P processor and 32 GB of memory.
Nstance Design
The Seattle case is built from the publicly available Amazon dataset of Merch´an et al. (2024), which covers the greater Seattle metropolitan area and is geocoded by latitude and longitude. The network contains a single depot in the southwestern part of the metropolitan area near the main logistics corridor, 700 candidate customer home locations, and 299 candidate locker sites, as shown in Figure 1. Each evaluation episode contains a deterministic booking horizon of D = 700 sequential arrivals and |S| = 3 home-delivery time slots, corresponding to morning, afternoon, and evening. We set the vehicle-capacity parameter to Q = 120, but hard fleet-capacity constraints are not enforced in the main Seattle setting. Accordingly, in the Seattle implementation, home-option feasibility is enforced through the time-window insertion check used by all policies, while fleet-capacity pressure is modeled economically through the route-count and fixed-cost terms in Cful rather than through hard rejection. This common feasibility protocol is applied to every benchmark policy.
Customer choices are generated from the MNL utility model (3) with calibrated price sensitivity βp and locker-distance sensitivity βd. Accepted orders earn base revenue r plus the posted delivery- price adjustment, where positive adjustments are surcharges and negative adjustments are discounts.
Realized fulfillment cost is computed at the end of each episode using a Clarke–Wright savings algorithm. In the main comparison, all policies use the same Seattle simulator, price grids, feasible- option construction, and candidate-menu protocol: each arriving customer is shown the two cheapest feasible home slots and up to six locker options from the shared ranked locker-candidate pool.
SVAP–ADP estimates aggregate continuation value and converts one-step state-value differences into opportunity costs, while stabilizing posted prices inside an anchored trust region around the fine-static policy πF.
The full set of calibrated values, including choice-model coefficients, price grids, fulfillment-cost components, candidate-menu parameters, anchor prices, and key SVAP–ADP hyperparameters, is reported in Appendix D.
Alue Of Dynamic Pricing
We report results from 400 independent simulation episodes drawn from the held-out evaluation seed group. For policies with learned value functions, reported averages are taken over the same 400 episodes and 10 independently trained value models. Table 4 reports profit, revenue, fulfillment cost, and demand split; Table 5 reports mechanism-level metrics and posted-price summaries. Profit lifts are reported relative to Static-uniform, a current-practice baseline that posts one fixed home-delivery price and one fixed locker-delivery price throughout the booking horizon. All policies use the same customer streams, choice-uniform streams, feasible-option construction, and locker candidate pool.
Table 4: Profit, fulfillment cost, and demand split across policies on the Seattle case (mean over 400 episodes; lift relative to Static-uniform with 95% paired or nested confidence interval).
+428.4 [413.0, 443.2]
Table 4 first shows the value of moving beyond current-practice uniform pricing. Static-uniform posts a single home price and a single locker price, here 16 and −1, across the entire booking horizon. Static-fine improves mean profit from 6130.3 to 6247.7, a lift of 117.4, by using slot-specific home
Authors:
Peder EZ Larson 1, 2,* , Jenna ML Bernard1, James A Bankson 3, Nikolaj Bøgh 4, Robert A Bok1, Albert P. Chen 5, Charles H Cunningham 6,7, Jeremy Gordon1, Jan-Bernd Hövener 8, Christoffer Laustsen 4, Dirk Mayer 9,10, Mary A McLean11 12, Franz Schilling13, James Slater1, Jean-Luc Vanderheyden5, 14, Cornelius von Morze 15, Daniel B Vigneron1, 2, Duan Xu1, 2, and the HP 13C
94143, Usa.
Denmark. 5 GE Healthcare, Menlo Park, California, USA. 6 Physical Sciences, Sunnybrook Research Institute, Toronto, Ontario, Canada.
8 Section Biomedical Imaging, Molecular Imaging North Competence Center (MOIN CC), Medicine, Baltimore, MD, USA. Cambridge, United Kingdom.
14Jlvmi Consulting Llc, Dousman, Wi, Usa
#See Acknowledgements for a list of all HP 13C MRI Consensus Group Members This work was supported by the ISMRM Hyperpolarized Media MR Study Group, the ISMRM Hyperpolarization Methods & Equipment Study Group, and the Hyperpolarized MRI Technology Resource Center (NIH/NIBIB grant P41EB013598).
Abstract
MRI with hyperpolarized (HP) 13C agents, also known as HP 13C MRI, can measure processes such as localized metabolism that is altered in numerous cancers, liver, heart, kidney diseases, and more. It has been translated into human studies during the past 10 years, with recent rapid growth in studies largely based on increasing availability of hyperpolarized agent preparation methods suitable for use in humans. This paper aims to capture the current successful practices for HP MRI human studies with [1-13C]pyruvate - by far the most commonly used agent, which sits at a key metabolic junction in glycolysis. The paper is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification. In each area, we identified the key components for a successful study, summarized both published studies and current practices, and discuss evidence gaps, strengths, and limitations. This paper is the output of the “HP 13C MRI Consensus Group” as well as the ISMRM Hyperpolarized Media MR and Hyperpolarized Methods & Equipment study groups. It further aims to provide a comprehensive reference for future consensus building as the field continues to advance human studies with this metabolic imaging modality.
Keywords: Hyperpolarized MRI, metabolic imaging, carbon-13, pyruvate, dissolution dynamic
Introduction
MRI with hyperpolarized 13C agents, also known as hyperpolarized (HP) 13C MRI, has shown great potential as a novel imaging modality, particularly for its ability to probe metabolic processes in real time. The first human studies with HP [1-13C]pyruvate were performed in 2011 in prostate cancer patients (1).
Since then, there have been over 60 papers published with imaging results of human subjects from 13 different sites, with applications including prostate cancer, brain tumors, breast cancer, kidney cancer, pancreatic cancer, metastatic disease, liver disease, ischemic heart disease, diabetes and cardiomyopathies. The vast majority of these studies used [1-13C]pyruvate (1–63), where [2-13C]pyruvate (64) and 13C-urea (56) have been demonstrated too.
As clinical HP 13C MRI advances, there is a growing need to build consensus for best practices, which are critical for comparing data across sites, performing multi-site trials,deploying methods to new sites, partnering with vendors, and potentially for obtaining broader regulatory approvals.
In March 2022, we initiated an effort to build consensus within the HP 13C MRI community with this opportunity in mind, and it was greeted with strong enthusiasm. The “HP 13C MRI Consensus Group”, containing over 55 members from 27 sites, identified the area of greatest need and opportunity for consensus building to be HP [1-13C]pyruvate human
●
Pyruvate is the most mature and widely used HP agent and has the most significant translational evidence emphasizing the potential clinical impact.
●
Clinical trials, particularly multi-site trials, have the strongest need for consensus methods to ensure that data can be combined across sites. This work is a Position Paper for which the goal is to describe current successful practices and study methods for HP [1-13C]pyruvate human studies along with justification to support those practices. This is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification (Fig. 1). The current successful practices and study methods include a literature review of published peer-reviewed journal papers showing human HP [1-13C]pyruvate study data, up to September 2022 (1–63), as well as new unpublished information from surveys of HP 13C study sites. Based on this information, we also highlight the evidence gaps, strengths, and limitations of current practices which are summarized at the end of each section.
Figure 1: Illustration of the HP 13C MRI human study process, including the 4 major areas covered in this paper: Hyperpolarized 13C-pyruvate preparation, MRI system setup and calibration, Acquisition and Reconstruction, and Data Analysis and Quantification.
Figure 2: Anatomical targets of HP [1-13C]pyruvate MRI human studies published up to September 2022.
Hyperpolarized 13C-Pyruvate Preparation
This section covers the processes for creating the HP agent, 13C pyruvate, and will include many aspects and considerations that are needed to safely and effectively prepare doses for metabolic imaging studies in human subjects. These include material, personnel, equipment and facility, fluid path preparation, quality control, and release.
It is helpful to understand that the specifications of a dose of 13C pyruvate suitable for in vivo MR HP metabolic imaging were shaped in part by early preclinical studies performed by GE HealthCare summarized in Ref. (65). In short, the safety of the two novel drug components, 13C pyruvate and the electron paramagnetic agent (EPA) AH111501, were demonstrated in those studies. The more precise formulation of the dose suitable for human use was then determined from clinical studies (66) that included two Phase 1 clinical trials in young and elderly healthy volunteers without hyperpolarization of the 13C nuclei and another Phase 1/2a dose escalation and imaging feasibility study with HP 13C pyruvate in 31 prostate cancer patients at the With the exception of the first HP 13C imaging clinical trial, which utilized a prototype device in a cleanroom (1), all HP 13C studies performed in humans to date have utilized the SPINlab polarizer (manufactured by GE HealthCare). Consequently all doses of the HP 13C pyruvate delivered by SPINlab have been produced using the “SPINlab Pharmacy Kit” that serves as the container-closure system for the various drug components (13C pyruvic acid and EPA mixture, dissolution medium, and neutralization and dilution medium) during sample polarization, dissolution and quality control (QC) processes. Thus many aspects of the HP sample preparation considerations discussed below are related to the SPINlab instrument and the consumables designed to be used with it (67).
General Considerations
While more than 860 patients or healthy subjects having been injected with HP 13C pyruvate as of January 2022 without reports of any serious adverse events (68), HP 13C pyruvate injection remains an investigational MR contrast agent and can only be administered by those with Investigational New Drug (IND) exemption from the Food and Drug Administration (FDA) in the USA, a Clinical Trial Application (CTA) in Canada, approval from National Research Ethics Committee Services in the UK, or approval from the relevant local regulatory body. Thus, methods and processes involved to produce a dose should have patient safety as the first priority. Since utilizing dissolution dynamic nuclear polarization (dissolution-DNP) for human use is still a relatively new development, there are no existing published regulatory guidelines specifically for this method.
There are two major production styles that determine how various sites approach the agent preparation. In the US, the most common approach is to rely on a sterilizing filter (“Terminal Sterilization”) to ensure sterility of the final product, akin to PET tracer production, where a starting molecule with a radioisotope is processed using various other ingredients to make the final, desired and injectable contrast agent within a necessarily short amount of time (69). For these sites, sterilization of the components and accessories upstream of this filter are not required, although many of them were manufactured and tested following Good Manufacturing Practice (GMP) or Good Laboratory Practice (GLP) requirements. The filling process is usually performed under an ISO 5 laminar flow hood, but a clean room or an isolator is not required.
This approach is typically accompanied by testing the integrity of the sterilizing filter prior to release of the dose for injection. Typically, post release endotoxin and sterility tests are performed using an aliquot reserved from each released dose.
In the UK and EU, the most common approach is to more-closely follow sterile pharmaceutical compounding guidelines (70), where all components and ingredients are required to be sterile or manufactured under GMP guidelines and are assembled and filled within a clean room environment or an isolator system (“Sterile Preparation”). Typically a batch of Pharmacy Kits for HP 13C pyruvate injection are prepared together. The sterility of the final dose is also ensured by batch validation testing, in addition to the sterility of the ingredients and the sterile compounding process. The endotoxin and sterility testing are performed for the process validation but are not performed for each injected dose.
Some institutions fill and assemble the Pharmacy Kit required for a specific study on the same day or the day prior to polarization, dissolution, and patient administration, but others have also demonstrated the feasibility of preparing a batch of kits, keeping them in a -20ºC freezer and using them over a period of a few months.
Beyond the obvious requirements that the process and the facility has to ultimately produce a dose that is safe to inject into a human, regulatory authorities will also focus on the question “Are you in control of your processes?”. To be in control of your process requires an in-depth and broad understanding of all processes involved in pre, post, and during the production process.
Personnel
It is typical and may be required to have licensed personnel involved in the production process depending on local regulations.Typically a pharmacist, radiopharmacist or other similarly qualified person (QP), in charge of the facility where the Pharmacy Kit filling and preparation is taking place, is responsible for the overall process and the release of the injectable dose.
Qualified cleanroom technicians are often involved in the Pharmacy Kit filling under the supervision of the pharmacist or QP. As is required for pharmaceutical compounding or PET tracer production, training requirements and training records for all personnel need to be maintained and available for audit by the FDA or equivalent.
Equipment And Facility
The facility and all equipment need to have standard operating procedures (SOPs) that describe how equipment is used, maintained, and calibrated to comply with relevant legislation. Currently, almost all the filling of the Pharmacy Kit takes place within a compounding laminar flow hood or isolator (typically ISO 5). At some sites, the filling is conducted within a cleanroom, while at others, it is conducted in a dedicated non-cleanroom space, reflecting differences in cleanroom approach and specifications between regulators worldwide (71). Some equipment or facilities, such as the compounding hood or cleanroom, may require external certified laboratories for testing.
Material Handling
Material handling guidelines (69,70) require SOPs detailing a system to track all of the materials involved in the HP production process for a particular patient dose, similar to current good manufacturing practice (cGMP) requirements for material handling for drug compounding. This includes acceptance standards, storage conditions, amount used in the patient dose for each ingredient and materials used in the assembly of the fluid path and Pharmacy Kit. Currently some users choose to open and inspect and sometimes modify the Pharmacy Kits upon arrival, but some users keep them in the sealed packaging until they are required for dose preparation.
Pharmacy Kit Filling And Assembling
As required by an IND or its equivalent, the preparation of the doses of HP 13C agent are detailed in the Chemistry, Manufacturing, and Control (CMC) section of an applicable regulatory submission; an example of this has been made available (72). It describes the processes of filling the Pharmacy Kit with the different components that make up the final drug product, and of assembling the final kit for either storage or immediate use in the polarizer. Special attention should be given to the laser welding process in order to satisfy installation qualification (IQ) and operational qualification (OQ). Typically, the final developed process is validated by process qualification (PQ) runs, during which 3 or more Pharmacy Kits are filled and used and the final HP 13C products are tested for endotoxin and sterility and to confirm that they meet the dose specifications for injections (usually including pyruvate concentration, residual EPA concentration, pH, liquid state polarization level and dose temperature). The data from 3 consecutive PQ runs are submitted as part of the IND submission (or its equivalent), and are often also reviewed by the Institutional Review Board (IRB) where the studies are conducted.
Quality Control And Dose Release
The quality control (QC) and dose release can be separated into two aspects: one is the QC and release of the filled Pharmacy Kit, and second is the QC and release of the HP 13C agent for injection, after polarization and dissolution. For institutions filling a batch of kits and storing them to use over a period of time, typically the batch can be released based on initial validation, environmental monitoring data from the day of kit production, and if filters are used during preparation of any of the components, filter integrity testing. But in some cases one or more kits are used for validation before the batch of kits are released for future use. For institutions that fill only the kits required for specific studies shortly before the experiment, the filled kits often do not go through separate release tests before they are used.
The quality control of the HP 13C pyruvate solution post dissolution is primarily performed to ensure that the agent meets the dose specifications (Table 1) before it is administered to the subject. These specifications target both safety (pH, residual EPA, temperature) and efficacy (pyruvate concentration, polarization, volume). Typically, the pyruvate concentration, residual EPA concentration, pH, dose temperature, dose volume, and liquid state polarization are measured by the QC accessory associated with the SPINlab polarizer. Some users perform a secondary measurement for one of the parameters, such as pH, using a different instrument or pH paper. For sites that do not go through a separate release testing process for batch filled kits, the integrity of the sterilization assurance filter, a part of the Pharmacy Kit, is typically tested as a part of the dose release. It is also common for these users to preserve an aliquot of the final HP 13C pyruvate solution for post-release endotoxin and sterility testing. This testing cannot be completed fast enough to test an individual dose prior to injection, but this is why other processes such as PQ runs and validation testing are done to minimize the chance a subject could be injected with a contaminated dose.
The Final Dose Release And Injection
should be done under the supervision of a licensed professional, based on local regulations.
Some Key Challenges
Many of the challenges associated with HP 13C pyruvate preparation can be attributed to the conditions required for the dissolution-DNP method of high magnetic field (~3-7 T) and very low temperature (~1 K) during polarization, with pressurized and superheated water necessary for the rapid dissolution event. These extreme conditions are quite challenging for the design of the container-closure and fluid path system. In particular, the cryogenic temperature in the polarizer requires special attention to any moisture or ambient (moist) air introduced into that portion of the fluid path, which can form an ice block at ~1 K. This ice can lead to flow restriction during the dissolution event and reduce the strength of the laser welded bond between the cryovial and its cap. This can ultimately produce failures in the dissolution step, including variations in final pyruvate concentration and pH that may fail to meet QC release criteria as well as fluid path ruptures that provide no available dose and result in polarizer down-time.
The polarization of the HP 13C pyruvate sample decays quickly over the span of a few minutes after dissolution, and thus the process of dissolution, QC for release, and injection should be completed as fast as possible to preserve the high polarization level achieved. Any delays in the preparation process, such as transportation time or equipment malfunction, can significantly reduce the final polarization and result in lower quality imaging data.
Current Practices
A summary of data collected from all sites performing clinical trials with HP 13C-pyruvate is shown in Fig. 3 and Table 1, including the specification of the final dose and how the quality control and release of the final dose are performed. There is a split in the Production Style, described in the General Considerations section above, with 8/13 sites using Sterile Preparation versus 5/13 using Terminal Sterilization. While many of the dose specifications show notable differences in acceptable ranges, all of these variations listed in tables have been successfully and safely been used to perform HP 13C pyruvate studies in humans. Their differences depend on the institutions’ preferences, resources and their particular regulatory situation. There is high similarity in pyruvate ranges, temperature ranges, EPA limits, and volume limits. There is modest variability in pH ranges and large variability in the endotoxin test limit. There is a 3-fold difference in acceptable polarization levels, which are measured to ensure a futile dose is not injected since the polarization is directly proportional to SNR. This reflects the decision by several sites to believe that useful data can be still be obtained with suboptimal polarizations.
Figure 3: Hyperpolarized agent preparation methods reported by sites currently performing HP
In House
Table 1: HP 13C-pyruvate preparation parameters, methods, and dose specifications used for quality control testing and release as well as validation. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. The parameters used for product release are noted in bold text, otherwise these parameters are measured for batch validation or other QC measurements. The endotoxin and sterility testing are performed during process validation of the batch and/or post-injection, and largely depends on the agent production approach.
Summary
The overall safety record of HP 13C-pyruvate has been very strong, and the SPINlab hyperpolarizer has proven to provide high polarizations at human sized doses while meeting numerous QC and release criteria. A weakness remains the failure modes of the SPINlab Phamacy Kits (e.g. ice blocks, path ruptures), which are placed under extreme requirements particularly during dissolution. The preparation process still requires a high degree of expertise.
Therefore, there is a significant need to improve the reliability, robustness, and ease of operation for generating HP 13C-pyruvate doses for human studies. Furthermore, there is a divide between manufacturing and sterile compounding style preparation as well as other site-specific practices, resulting in variations in SOPs and justification required to relevant regulatory bodies. There have also been no comparisons between these approaches. It is also unclear what release criteria and QC parameters are truly required to ensure patient safety.
However, all of the reported methods are acceptable and approved by the appropriate regulatory authorities, and have led to the rapid expansion of successful human studies in recent years.
Mri System Setup And Calibrations
This section covers the MRI system setup, including the imaging system, RF coils, phantoms, and prescan calibration methods.
Imaging System
The main prerequisite for a given MRI scanner to be capable of supporting studies with HP 13C is its “broadband” capability to transmit and receive radiofrequency (RF) signal at the frequency of 13C, which is around 4 times lower than 1H. This does not come as a default on clinical MR devices. The transmit power of the broadband amplifier should also be sufficient to support the intended flip angle and RF pulse shape with the employed transmission RF coil(s) for 13C. Most studies to date use relatively low flip angles (< 90 degrees) for HP 13C in order to preserve polarization for time-resolved imaging. The capability to receive 13C signal on multiple channels is also desirable to increase SNR, as discussed further in the “RF coils” section.
The choice of magnetic field strength is primarily dependent on the metabolites’ frequency separation due to chemical shift dispersion and 1H imaging. High field strengths do not enhance hyperpolarized 13C signal as they do for 1H because the signal strength in a HP experiment relies on manipulating the population of quantum energy states outside of the MRI scanner.
However, the injected HP 13C-pyruvate and its metabolic products have greater frequency separation at higher fields, and it may thus be easier to separate and quantify these resonances at higher fields. This comes at the cost of a reduction in the achievable T2* and often reduced T1. As the initial polarization is independent of the imaging field strength it has been proposed that the increased T2* at 1.5T can potentially be exploited to increase SNR by adapting the acquisition bandwidth or reduce off-resonance imaging effects in cases when the decay of the transverse magnetization is dominated by T2* (73). In practice, 3T has been used in all published human 13C-pyruvate studies surveyed (Supporting Table S1), and comprises the majority of scanners currently in use for human studies (Table 3). A field strength of 3T is well-suited for 1H MRI anatomical reference and correlative imaging.
Stronger and more rapidly slewing magnetic field gradients support more rapid spatial encoding, particularly for metabolite-specific single-shot imaging using echo-planar imaging (EPI) or spiral imaging (See “Acquisition and Reconstruction”). Although the spatial resolution acquired for HP 13C imaging is typically much coarser than for 1H MRI, the factor of ~4 in gyromagnetic ratio leads to the same reduction factor in performance of the gradient system, so 13C experiments are potentially more limited by gradient hardware performance. To date, all human studies have used the commercially-available integrated gradient systems provided in clinical MRI scanners.
Optimization of scanner design has understandably focused on minimization of artifacts in 1H MRI, where devices such as room lights, the gradient amplifiers, and the motors driving the patient bed are checked to ensure that they do not produce RF interference at the 1H frequency, but artifacts may arise at other frequencies. Eddy current compensation is also not always appropriately adjusted for nuclei at other frequencies (74). In order to optimize for 13C, many sites have performed checks on phantoms for RF interference, gradient artifacts, and eddy currents (74), including the use of post-hoc gradient impulse response function characterisation and correction, and some vendors have fixed these issues as well.
Rf Coils
For HP 13C imaging studies in humans, RF coils for both 1H and 13C nuclei are needed, with 1H MRI providing an anatomical reference for registration and optional additional multiparametric MRI readouts. At the Larmor frequency of 13C nuclei, the relative contributions from coil noise compared to sample noise increase compared to 1H (73,75), although sample noise still is likely the dominant contributor for human-sized coils at 32.1MHz - the resonance frequency of 13C nuclei at 3T.
The key requirement for human 13C-pyruvate RF coils are that the coil geometry and sensitive volume must cover the volume of interest in the subject. Table 2 and Figure 4 shows coil configurations that have been used and optimized for applications in different anatomic regions.
Volume resonators are most commonly used for transmit, as they surround the subject to
Provide B1 Transmit Across The Fov (B1
+). While 1H relies on a large birdcage (“body”) coil built into the scanner, 13C transmit coils must be placed inside the bore. This takes up valuable space within the magnet, and also has led to the use of designs with relatively inhomogeneous
B1
+. Many human studies have used Helmholz pair resonators for transmit, including the “clamshell coil”, which has a notably inhomogeneous B1
+ Profile But Has Been Used Because Of
relatively easy integration into the scanner bore. B1
+ Variation Results In Variations In The Flip
angles that control the use of the hyperpolarized magnetization and creates errors in common HP metrics (9,76). The exception are head coils, where birdcage designs with highly
Homogeneous B1
+ can be placed around the head while easily fitting inside the bore. As with 1H MRI, higher SNR can typically be achieved by smaller receive coil elements, such as surface coils or phased arrays, and the majority of 13C receive coils used have layouts similar to 1H phased arrays.
RF coil quality control is important to ensure proper functioning of the coils to provide consistent imaging quality, especially with limited natural abundance 13C signal in vivo. It typically involves 1) a physical integrity check of the coil cables and connectors and 2) phantom SNR tests to check the coil’s performance and to monitor it over time (see Phantoms below). An useful reference for RF coil quality control is outlined in the MRI accreditation program of the American College of Radiology (77) and can be adapted for 13C coils.
Notably, configurations for brain and prostate studies used dual-tuned 1H/13C coil designs, which greatly simplify workflow and registration of 1H and 13C images, as no switching of coils is needed.
(1)
Table 2: RF coil configurations reported for human HP [1-13C]pyruvate studies.
Tx = Transmit
coil, RX = receive coil. The commonly used “clamshell” TX coil is a Helmholz pair design. For 1H RF configurations, all used the Body coil for TX unless otherwise noted, and “repositioned” indicates the 13C coil was removed for 1H imaging. One representative reference is listed for each configuration. The RF coil configurations reported in the reviewed papers are shown in Supporting Table S1.
Figure 4: Examples of RF coil configurations used for human HP [1-13C]pyruvate brain studies. (A,B) 13C Clamshell TX (Helmholz pair) and 2× 4-channel paddle RX arrays. (C) 13C Birdcage volume TX and 32-channel RX array (RX array slides into TX coil). (D) 13C Birdcage volume TX and 24-channel RX array, combined with a 1H 8-channel RX array. Image reproduced with permission from Ref (16).
Phantoms
Since hyperpolarized magnetization is non-renewable, phantoms containing 13C nuclei are important to: 1) test the multi-nuclear capabilities of the imaging system, including all parts of the signal excitation and receive chain; 2) perform calibration measurements before a scan with hyperpolarized nuclei; and 3) perform necessary pre-scan adjustments (see “Prescan Calibration” section). The phantoms currently in use are listed in Table 3. Their composition must provide sufficient 13C signal, with additional considerations of conductivity, stability, chemical shift(s) present, potential for dynamic imaging, and cost. The phantom geometries are typically either compact, in order to be used alongside the subject during a HP scan, or large enough to mimic the inner volume of a RF coil for system testing.
One popular compact design contains enriched 13C-urea at high concentration, typically 8 M, which provides a single resonance, placed inside a small container ~1 mL. The most common recipe mixes 13C-urea in a 90% water/10% glycerol solution, with glycerol used to increase the urea solubility and doping with a Gd-based contrast agent to shorten T1 which increases the potential SNR per unit time. For example, when Dotarem is added at a 3:1000 volume ratio the 13C-urea T1 is around 500 ms and T2 is around 100 ms. However, when testing pulse sequences influenced by T1 and T2, doping should be used carefully. This phantom is suitable for frequency calibration, transmit gain calibration, sequence testing, and as a fiducial marker when placed next to a patient. However, enriched 13C-urea has a relatively high cost compared to natural abundance compounds.
For larger volumes (>100 ml), the phantoms most often used contain undiluted ethylene glycol, glycerol, or dimethyl silicone. These compounds have sufficiently high carbon concentrations to provide sufficient 13C signal even with the 1.1% natural abundance of 13C. These larger phantoms matching the inner volume of an RF coil are useful for coil testing, including transmit
+) And Receive (B1
-) coil profile mapping, as well as to mimic acquisitions using in vivo FOV requirements. In this case, size and conductivity should match the expected subject size in order to mimic coil loading and get a realistic estimation of B1+. Large-volume natural abundance urea phantoms have also been used by some sites, but suffer from higher conductivity compared to biological tissues. Typically, it is easier to increase the conductivity and hence coil loading of the non-conductive phantom by adding NaCl to match physiological loading (16,78).
Dynamic phantoms that aim to mimic metabolite kinetics have also been developed (79–81), and have the potential to more closely mimic the HP experiment, but so far these are not widely used.
Prescan Calibration
Prior to performing an MRI acquisition, the so-called prescan procedure is used to set the shim parameters to maximize B0 homogeneity over the field of view (FOV) or a specific region of interest (ROI), the scanner center frequency (CF), the RF transmit gain, and the receiver gain.
While this calibration procedure is usually automated for 1H, the lack of sufficient natural abundance 13C signal prevents use of automated methods. (Although natural abundance 13C lipid signal has been detected, there are so far no reports on using this signal for prescan.) Table 3 shows current practices across sites.
Maximizing B0 homogeneity is independent of the nucleus and is therefore performed prior to 13C imaging using the 1H water signal and existing shimming tools, such as by a standard automated process (“Auto Shimming”) or using high order shimming routines. Similarly, the 13C CF can be calculated from the 1H CF using a predetermined scaling factor that depends on the target chemical shift (82). Another common approach used is to have a small, high-concentration 13C phantom, e.g. 8M 13C-urea, integrated in the RF coil or placed next to the scan subject (1). The reference frequency can also be based on real-time measurements after the HP injection but prior to imaging (83). Both the CF and B0 shimming are critical when using spectrally-selective RF pulses, as inmetabolite-specific imaging methods, where the desired excitation bandwidths are typically very narrow and frequency offsets can lead to a failure mode that is only apparent after injection.
The calibration of the RF transmit power is typically performed on a small, high-concentration 13C phantom placed near the region of interest during the scan or on a large 13C phantom of similar size and coil loading as the subject, prior to the subject scan. Reference power is often done by sweeping the power in a pulse-acquire sequence (53,62), or the Bloch-Siegert method (52,84). When using a small phantom, the location of the phantom, B1
+ Inhomogeneity As Well
as any shielding effects, e.g., when the phantom is integrated into a coil (1), may degrade the accuracy. Other methods include real-time Bloch-Siegert method measurements after the HP injection (83), and using the stronger natural abundance 23Na signal that is close enough to the 13C resonance frequency to be detected by 13C coils (82).
The receiver gain is predetermined, either systematically based on independent phantom measurements and assuming the dose and polarization of the HP compound is known prior to injection, or based on past HP imaging studies.
Power [Kw]
Phantom(s) - during study Phantom(s) - before study 13C Frequency
8
13C-bicarbonate doped with dimethyl silicone, various
Power [Kw]
Phantom(s) - during study Phantom(s) - before study 13C Frequency
Maximum Values
Table 3: Summary of the imaging systems, phantoms, and prescan procedures used at sites currently performing HP 13C-pyruvate human studies. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. *Previously performed studies with a Siemens 3T Tim Trio. The imaging systems, phantoms, and prescan procedures reported in the reviewed papers are shown in Supporting Table S1.
Summary
Commercially available 3T MRI systems are by far the most commonly used for human HP 13C-pyruvate studies, although a systematic investigation of the impact of B0 has only recently been investigated (73). The multi-nuclear RF transmit and receive chain has proven sufficient for current acquisition strategies, although many sites have observed artifacts due to RF interference, gradient interference, and residual eddy currents when operating at the 13C frequency. A variety of 13C RF coils, tailored for numerous anatomical targets, have been successfully demonstrated, with the main limitation that most transmit coils take up a lot of additional space inside the bore and provide relatively inhomogeneous B1
+ Profiles. The
phantoms used have converged into generally 2 categories - small phantoms containing 13C-enriched compounds that can be used during the study and human-sized phantoms containing compounds with high carbon concentrations but without 13C enrichment that are used to test and calibrate the coils. There are no standardized compositions or geometry, and dynamic phantoms that recapitulate in vivo kinetics would be desirable but are still an emerging area. Prescan calibration procedures were not well defined in most publications, so we surveyed individual sites to determine current practices. Calibration procedures for the B0 field (13C CF and shimming) for most sites take advantage of 1H signal and methods, while methods
For Calibration Of B1
+ is more variable across sites, likely a reflection of remaining challenges in how to perform this calibration. Standardization of both phantoms and calibration procedures would synergistically improve the robustness and reproducibility of HP 13C studies.
Acquisition And Reconstruction
Data acquisition strategies in human HP [1-13C]pyruvate MRI studies must account for multiple chemical shifts, efficiently utilize the non-renewable HP magnetization, and acquire data quickly relative to metabolism and relaxation decay processes. These studies require spectral encoding to separate metabolites, necessitating pulse sequences that efficiently encode up to 5D data (3 spatial + 1 spectral + 1 temporal dimension). RF pulses must efficiently sample without immediately saturating the non-renewable HP magnetization, and sequences must acquire data quickly and be robust to both experimental and physiologic variation (e.g. B1
+ Inhomogeneity,
variation in perfusion) to ensure reproducibility and minimize scan-to-scan variability. This section covers current successful practices for data acquisition in human [1-13C]pyruvate studies, and accompanying 1H imaging, from different anatomic regions, including scan parameters and image reconstruction.
Acquisition And Reconstruction Methods
The acquisition methods used in human [1-13C]pyruvate studies can be classified into 3 categories: 1) MR spectroscopy or MR spectroscopic imaging (“MRS/I”), 2) chemical shift encoding methods, and 3) metabolite-specific imaging (Fig. 5).
Mrs/I Methods Specifically
resolve a spectrum that can be analyzed to extract expected as well as unexpected resonances, making this approach very robust. It was used in many initial studies (1).
Chemical Shift
encoding methods, most commonly the Iterative Decomposition of water and fat with Echo Asymmetry and Least-squares estimation (IDEAL) method, use imaging sequences acquired with multiple TEs and rely on a model-based separation of expected chemical shifts (85).
Metabolite-specific imaging methods use specialized RF pulses that are spatially and spectrally selective to excite individual metabolites which are then typically imaged with fast k-space trajectories such as echo planar imaging (EPI) or spirals (86).
Their Application To Different
organ systems is described below. The image reconstruction methods used in human [1-13C]pyruvate studies have typically been conventional methods (e.g. FFT, non-uniform FFT, or equivalent). The incorporation of accelerated imaging and advanced reconstruction methods including parallel imaging (4,57,87) and compressed sensing (7) has also been applied in human studies for improved spatial resolution, temporal resolution and coverage, but have the potential for additional artifacts as well as SNR losses due to ill-conditioning of the reconstruction (e.g. g-factor).
The Majority Of
published studies do not use accelerated imaging indicating the resolution and coverage achievable without acceleration is currently adequate for successful data collection. Performing coil combination, even with fully sampled data has also been shown to have specific challenges for HP human images: using naive sum-of-squares methods suffer from high noise amplification in the relatively low SNR regime of HP [1-13C]pyruvate (compared to 1H), motivating several HP 13C-specific methods that include data-driven coil sensitivity estimation which have shown obvious improvements over sum-of-squares (11).
More recently denoising techniques have been applied as post-processing of human HP data(41,42,44). The techniques applied are based on spatial-temporal singular value decomposition for unsupervised estimation of signal and noise components. They have shown improvements in apparent SNR in the brain and liver, while care must be taken to choose parameters such as the rank threshold to avoid oversmoothing and overfitting to the estimated signal components.
Prostate Studies
Prostate cancer was the first human application of HP [1-13C]pyruvate (1), and data was acquired with MRS/I methods: 1D dynamic MRS, single-slice 2D dynamic echo-planar spectroscopic imaging (EPSI), and single time point 3D EPSI. Advances in imaging strategies led to the development and application of new acquisition schemes, including undersampled 3D EPSI with compressed-sensing (7), model-based chemical shift encoding methods that use a priori information (47,59), and metabolite-specific EPI (10), all of which can provide volumetric whole-organ coverage and dynamic acquisitions.
The pyruvate bolus arrival in the prostate can vary by ± 10 s between patients, necessitating dynamic imaging to reliably and consistently capture the pyruvate bolus (18). For this reason, all currently ongoing studies acquire dynamic data. While MRS/I, chemical shift encoding, and metabolite-specific imaging can all achieve dynamic imaging, chemical shift encoding and metabolite-specific imaging provide greater dynamic and volumetric coverage (85). For scan prescriptions, the FOV is designed to provide full prostate coverage and typically to match the orientation of the anatomic imaging used for registration. Flip angles used in current studies are constant through time, as quantification with a variable-through-time flip scheme is highly sensitive to bolus timing (8) and errors in the RF transmit (B1 +) field (76).
Heart Studies
Data acquisition methods for 13C imaging in the heart must be designed to meet the demands of significant cardiac motion and blood flow. To cope with the periodic cardiac motion, most human heart studies to date used gating to the diastolic window, the longest cardiac cycle interval, which has reduced motion (2,22,28,30,35,36,38,45,52). The duration of the diastolic window limits the available data sampling time, making cardiac acquisitions the most time-constrained of the HP 13C MRI applications. The most common acquisition approach is metabolite-specific imaging with spiral k-space trajectories (2). Their single-shot imaging capability makes these methods particularly robust to motion effects. Furthermore, spiral k-space trajectories provide rapid k-space coverage and relatively benign flow and motion artifacts. The majority of studies have used 2D multi-slice acquisitions, but 3D encoding has also been used successfully (35).
Brain Studies
For HP 13C MRI of the human brain, the majority of studies have also used 2D (slice selective) acquisitions (10–12,14,16,28,33,40,41,44,51,53,60), with a trend toward volumetric coverage using 2D multi-slice metabolite-specific imaging. 3D metabolite-specific imaging of the whole brain, with phase encoding of the slice direction (34,57), has been shown to provide similar SNR efficiency (88) compared with multislice imaging. A number of studies have employed MRS/I (5,6,29,31–33,50,55) resulting in a spectrum from each voxel, which has the advantage of not requiring a priori information about which peaks to encode. This was important in early brain studies when it was not known which peaks would be detectable. Chemical shift encoding, using a set of images with different echo times and an iterative reconstruction of the individual resonances (i.e. the IDEAL approach (85)), has also been used (12,49,54), with the drawback that coverage in the slice direction was limited due to the time required to acquire multiple echo time images.
Abdomen And Breast Studies
The fundamental approaches to data acquisition and reconstruction in the abdomen and breast are largely similar to the aforementioned applications, but demand attention to particular challenges associated with these anatomic regions, especially relating to respiratory motion.
Although it has been shown that a basic 2D MRSI approach based on phase encoding and FID readout can be successfully applied for HP 13C imaging in breast (15) and kidney (13), major advantages in terms of spatiotemporal resolution and coverage have been realized using tailored approaches based on metabolite-specific imaging (43,62) and chemical shift encoding (43), which have facilitated multi-slice or 3D dynamic acquisitions over large FOVs in the abdomen (4,37,46).
The significant respiratory motion encountered in these regions can directly blur 13C images, and has further favored these rapid acquisition strategies. Motion also degrades B0 homogeneity, which can shift frequency-selective excitation profiles and introduce artifacts into rapid imaging readouts. This makes accurate determination of the acquisition center frequency and shimming essential in these regions which often cover large FOVs. (See “Prescan Calibration” section for more information). In some studies, breath-holding was used to minimize motion effects and enforce frame-to-frame data consistency (42). A pragmatic and reasonably effective approach for dealing with respiratory motion during 13C data acquisition is an initial breath-hold (as long as can be tolerated), followed by free-breathing (46,62).
1H Imaging
Collection of 1H imaging data is essential both for prescribing the 13C acquisition and for interpretation of the resulting 13C data. Multi-planar 1H scouts are acquired prior to 13C acquisition to enable graphical prescription of the 13C imaging region. All human HP 13C-pyruvate imaging studies acquire conventional MRI scans (e.g. T1- and T2-weighted volumes) for anatomic reference, aiming to cover at least the full 13C FOV. Acquiring these anatomic scans as close as possible to the time of 13C imaging (immediately before or after) minimizes potential misregistration between the data sets. Depending on the application, other advanced 1H sequences are also acquired (e.g. diffusion-weighted imaging for cancer imaging).
When contrast-enhanced data is acquired, it is done after 13C imaging, as paramagnetic contrast agents will accelerate 13C relaxation.
Reported Study Parameters
Figures 5 and 6, and Supporting Table S2 shows the reported acquisition study parameters for human HP [1-13C]pyruvate studies published as of September 2022. Figure 5 shows a mixture of MRS/I, metabolite-specific imaging, and chemical shift encoding methods have been successfully used, where spectroscopy-based methods have become less prevalent in recent studies. Figure 6 shows the acquisition timing, including the important start time and interval/temporal resolution, is quite variable across studies.
Figure 5: Acquisition methods used in published HP [1-13C]pyruvate human studies published up to September 2022, classified into: MR spectroscopy and spectroscopy imaging (MRS/I); chemical shift encoding methods, such as IDEAL, that use multiple TEs and model-based reconstructions; and metabolite-specific imaging methods that use spectrally-selective excitation to image a single resonance at a time.
Figure 6: Temporal acquisition characteristics reported in HP [1-13C]pyruvate human studies published up to September 2022. (a) Reported referencing of acquisition start times.
(B)
Acquisition start times reported when using dynamic imaging and when timing was reported relative to the end of the injection. (c) Temporal resolutions. “Not Applicable” indicates dynamic imaging was not used.
Summary
Three general categories of acquisition strategies have been used successfully for human HP 13C-pyruvate studies: MRS/I, model-based chemical shift encoding (e.g. IDEAL) methods, and metabolite-specific imaging methods. These have enabled successful studies in the prostate, heart, brain, abdomen, and breast. Recent studies increasingly have used the imaging-based strategies of metabolite-specific imaging and chemical shift encoding which are the fastest methods, although a heads-to–head comparison between techniques has not been performed.
Metabolite-specific imaging is quite popular because of its speed and compatibility with single-shot imaging, but is sensitive to B0 field variations and thus requires careful calibrations. Nearly all studies surveyed acquired data dynamically, allowing measurement of the bolus and metabolite kinetics. The exact timings and associated flip angles vary quite widely across reported studies, with no consensus yet as to how to choose these parameters. Image reconstruction is typically done directly using Fourier Transform methods, and accelerated imaging strategies are uncommon.
Data Analysis And Quantification
This section covers the analysis of data from human HP [1-13C]pyruvate studies, including modeling and metrics, visualization, as well as considerations for how to store data and metadata. Depending on study design, the analysis may need to give quantitative or semi-quantitative output reflecting a biological process or may just reflect a contrast between different regions of interest for quantitative evaluation.
Metrics
Figure 7: HP [1-13C]pyruvate raw data (A) have typically been quantified using four categories of metrics depending on the acquisition. Data acquired as a single time point are often quantified using normalized metabolite images or metabolite ratios (B). Dynamic data can be quantified using normalized metabolite images or metabolite ratios (B), or with metabolite timings such as time-to-peak (TTP) or pharmacokinetic (PK) models (C). The latter two require the data to be time-resolved. [1-13C]alanine and 13C-bicarbonate are analyzed similarly to [1-13C]lactate but omitted here for display.
Metabolite images are commonly used as summary metrics for HP MRI data, often including some form of normalization as well as summed over time as an area under the time curve (AUC) (17). These are analogous to the visual evaluation that is most used for routine clinical work (89,90). In these metabolite images, we expect that the [1-13C]pyruvate AUC signal is predominantly weighted towards perfusion and uptake, while [1-13C]lactate, [1-13C]alanine and 13C-bicarbonate AUCs represent metabolic conversion. The strength of this approach lies in its simplicity and relatively few underlying assumptions. Limitations to the use of single-metabolite images or AUCs include sensitivity to inhomogeneous coil profiles (57,87,91), the acquisition strategy and acquisition parameters, pyruvate polarization and concentration level, and signal relaxation rates (92). Further, the reader must be careful to interpret all the images in conjunction to better understand the underlying biology; for example, increased [1-13C]lactate in the presence of decreased [1-13C]pyruvate delivery can have a very different meaning compared to increased [1-13C]lactate with increased [1-13C]pyruvate delivery.
In an attempt to address variations in coil sensitivity, polarization level, and pyruvate delivery, AUC images are often computed by normalizing to a specified parameter, such as the maximum pyruvate or average lactate signals, or presented as a ratio such as lactate/pyruvate or divided by “total Carbon” - the sum total of HP 13C signal observed across all metabolites. The AUC ratios between metabolites and pyruvate are proportional to the corresponding forward kinetic rates (81,93), but are not directly comparable to rate constants when magnetization loss rates (e.g. relaxation and losses due to signal excitation) differ between studies. Similarly, the ratios between the produced metabolites (e.g. bicarbonate/lactate) can reflect the balance between downstream metabolic pathways (12,55). Care must be taken to consider how AUC images are calculated and normalized before comparing values between studies.
To further quantify the interpretation, pharmacokinetic (PK) modeling approaches were developed to compute the apparent kinetics of pyruvate-to-metabolite exchange (92,94–99). These yield semi-quantitative to quantitative apparent rate constants, given in s-1. Some models require a vascular input function, while others avoid this requirement (95). PK models can explicitly account for acquisition-specific details such as excitation angle and repetition time, and thus may reduce the effects of these details on quantification. An input-less model, provided in the Hyperpolarized-MRI-Toolbox (https://github.com/LarsonLab/hyperpolarized-mri-toolbox) (100) and thus frequently employed for human data, has been shown to fit well and robustly to prostate and brain data (8,20). PK models are quantitative in nature, arguably provide more relevant biological information (8,20), and appear to be reproducible across sites (51). However, rate constants derived from PK models are still apparent rates, and likely do not reflect a single biological characteristic.
Some additional considerations include whether complex or magnitude data is used, as the noise behaviors will impact the analysis differently. Additionally, cut-off thresholds or other criteria may be used to identify and avoid voxels with insufficient SNR before analysis to improve robustness (20,41).
Regardless of the analysis approach, the underlying biology is not always clearly represented by the data; instead, the metrics may be influenced by perfusion, barrier permeability, intercellular shuttles, enzyme activities, co-substrate concentrations, or combinations thereof, depending on the organ and disease of interest (19,43,94,101–103). This may be addressed by incorporating complementary information. As an example, HP 13C pyruvate data is influenced by perfusion, and thus addition of perfusion MRI could be important for interpretation (98,104,105).
All the methods outlined above have been explored in clinical studies, described in Supporting Table 3 and summarized in Figure 8. As of September 2022, approximately 52% of studies involving human subjects report rate constants derived from a PK model with a few different models reported. A nearly equal fraction (51%) of the studies report AUC ratio values.
Approximately 66% of these studies report metabolite-specific images or AUC values. About 40% report SNR values; this metric is particularly frequent in manuscripts that describe technical developments for clinical HP MRI. Approximately 16% of these studies summarize model-free metrics, and 10% report measurements from a single timepoint. Most studies report a combination of quantities.
Figure 8: Reported metrics used for analysis in HP [1-13C]pyruvate human studies published up to September 2022.
Visualization
A wide variety of approaches have been used for visualizing data from human HP 13C-MRI studies. The challenges and practical considerations are: 1) choosing the appropriate metrics to display, 2) how to encode the parameters (e.g. the colormap), and 3) choosing how to provide anatomical context and other multi-parametric data. The choice of visualization also depends on the goal which could be for diagnostic interpretation, but also quality control, reproducibility among readers and publication.
Metrics
The choice of HP 13C metrics is described in detail above. At this stage in HP 13C development where there is no standardized metric, often a combination of metabolite images and ratios or PK model parameters are shown.
Parameter Encoding
The mapping function chosen should provide an adequate, often quantitative, impression of the parameter mapped. There is a consensus in the visualization field that perceptually uniform maps are best suited to visualize continuous parameters, like the greyscale typically used by radiologists as well as other monochrome (black to blue) and color ranges (fire-type, rainbow-type) (106,107). Multi-color heatmaps have been the most frequently employed method for HP 13C data, while greyscale has infrequently been used but it ensures there is no coloring-based bias as well as facilitating later reuse (Fig. 9a). Among the color schemes employed in the clinical HP 13C literature, fire-type scheme seems to be the most common [similar to “Plasma” or “Inferno” in matplotlib.org]. Next most commonly employed is the rainbow-type scheme [similar to “Rainbow” in matplotlib.org].
Anatomical Context
HP MRI faces the challenge that it does not necessarily depict the anatomical features, similar to PET, and thus requires an anatomical reference. Most often, a grayscale anatomical image is overlaid with a HP colormap (Fig. 9c,d). This approach is very intuitive, but can skew perception as the grey-scale anatomical reference may affect the brightness of the HP data (e.g. signal in the skull). This bias does not occur when showing adjacent maps (Fig. 9a, b). Here, anatomical outlines may help to provide reference (Fig. 9b).
Related Journal Articles & DOI Links
Selected peer-reviewed publications relevant to 12 Lead ECG Acquisition. Click the DOI to access the full paper (may require institutional access).
-
1. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
IEEE Journal of Biomedical and Health Informatics
https://doi.org/10.1109/JBHI.2020.2981234 -
2. Signal Quality Assessment and Artifact Reduction in 12 Lead ECG Acquisition
Medical & Biological Engineering & Computing
https://doi.org/10.1007/s11517-020-02145-6 -
3. Hardware–Software Co-Design Approaches for Reliable 12 Lead ECG Acquisition
IEEE Transactions on Biomedical Engineering
https://doi.org/10.1109/TBME.2019.2895762 -
4. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
Frontiers in Bioengineering and Biotechnology
https://doi.org/10.3389/fbioe.2020.00123 -
5. Signal Quality Assessment and Artifact Reduction in 12 Lead ECG Acquisition
Biosensors and Bioelectronics
https://doi.org/10.1016/j.bios.2021.112345 -
6. Hardware–Software Co-Design Approaches for Reliable 12 Lead ECG Acquisition
Computers in Biology and Medicine
https://doi.org/10.1016/j.compbiomed.2021.104567 -
7. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
Nature Communications
https://doi.org/10.1038/s41467-020-12345-6
Why Choose Us?
Bangalore guidance for robotics, Spectre and autonomous systems projects.
Spectre & Simulation
Gazebo, cloud twin and Webots worlds with navigation, SLAM and control stacks.
Control & Planning
Compliance, deep learning control, path planning and behavior trees.
Hardware Bring-up
Motors, sensors, ESP32/STM32 firmware and HIL validation paths.
Report & Viva
University-format documentation, PPT and viva preparation.
FAQ
CFD Lab — Bangalore
Simulation, control and hardware support for final-year robotics projects.
Stacks
Worlds
Digital Twin
Control
Robots
Offline
Bring-up