Abstract
Despite huge gains in performance in natural language understand- ing via large language models in recent years, voice assistants still often fail to meet user expectations. In this study, we conducted a mixed-methods analysis of how voice assistant failures affect users’ trust in their voice assistants. To illustrate how users have experienced these failures, we contribute a crowdsourced dataset of 199 voice assistant failures, categorized across 12 failure sources.
Relying on interview and survey data, we find that certain failures, such as those due to overcapturing users’ input, derail user trust more than others. We additionally examine how failures impact users’ willingness to rely on voice assistants for future tasks. Users often stop using their voice assistants for specific tasks that result in failures for a short period of time before resuming similar usage.
We demonstrate the importance of low stakes tasks, such as playing music, towards building trust after failures.
S Concepts
• Human-centered computing →Empirical studies in HCI. voice assistants, trust, survey, interview, dataset
Acm Reference Format:
Amanda Baughan, Allison Mercurio, Ariel Liu, Xuezhi Wang, Jilin Chen, and Xiao Ma. 2023. A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures. In Proceedings of the 2023 CHI Confer- ence on Human Factors in Computing Systems (CHI ’23), April 23–28, 2023,
Ntroduction
Voice assistants have received a lot of attention from both indus- try and academia, especially given the recent advances in natural ∗This work was conducted as part of an internship with Google Research.
Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored.
For all other uses, contact the owner/author(s).
Hi ’23, April 23–28, 2023, Hamburg, Germany
© 2023 Copyright held by the owner/author(s). language processing (NLP). Within the past five years, advance- ments in NLP have achieved huge gains in accuracy when tested against standard datasets [8, 18, 34, 56, 58, 60, 63], with state-of- the-art accuracy in natural language processing models as high as 99% for certain tasks [8, 34]. This has led many practitioners and researchers alike to imagine a near future where voice assistants can be used in increasingly complex ways, including supporting healthcare tasks [41, 55], giving mental health advice [52, 66], and high stakes decision-making .
However, despite the increasing accuracy of NLP models and the breadth of their applications, evidences suggest that users remain reluctant and distrusting of using voice assistants [12, 29]. In the U.S., voice assistants are common in homes, with an estimated 72% of Americans having used a voice assistant . However, people primarily use these for basic tasks such as playing music, setting timers, and making shopping lists [12, 29, 35]. This is because when voice assistants fail, such as by incorrectly answering a question, it derails user trust [27, 35]. User trust is pivotal to user adoption of various technologies , and in this case, low user trust results in reluctance to try voice assistants’ novel capabilities.
As voice assistants increasingly rely on large language mod- els [24, 47], we believe the gap between the high accuracy of these models and users’ reluctance to use voice assistants for complex tasks may be explained by differences in how users and NLP practi- tioners evaluate the success of a model. Standard NLP models are often evaluated on large datasets of coherent text-based questions and answers [48, 49] or paired written dialogue . Meanwhile, in practice users’ speech may include disfluencies, such as restarts and filler words, questions not covered in training, or background noise which misconstrues speech. In the case of question answering, NLP models are evaluated based on how many questions are accurately answered on a subset of the training dataset [48, 49]. As one may expect, people can interact with voice assistants in a multitude of ways that fall outside of the scope of training data, which can lead to friction. In the eyes of users, these inaccurate responses, or voice assistant failures, can lead to frustration. For example, only five percent of users report never becoming frustrated when using voice search .
We believe that the gap between how NLP models are evaluated and how users encounter and perceive failures hinders the prac- tical applications of the advancements that voice assistants have made. Therefore, we ask, which types of voice assistant failures
Hi ’23, April 23–28, 2023, Hamburg, Germany
Baughan et al. do users currently experience, and how do these failures affect user trust? A human-centered understanding of the types of NLP failures that occur and their impact on users trust would allow technologists to prioritize and address critical failures and enable long-term adoption of voice assistants for a wider variety of use cases.
Further, while research has started to categorize types of break- downs in communication between users and NLP agents [28, 46], little work has looked into how users perceive these failures and subsequently trust and use their voice assistants. We draw from and extend past research to make the following contributions: • C1: Iterating on the existing taxonomy of NLP failures, we crowdsource a dataset of 199 failures users have experienced across 12 different sources of failure.
• C2: A qualitative and quantitative evaluation on how these different failures affect user trust, specifically along dimen- sions of ability, benevolence, and integrity.
• C3: A qualitative and quantitative analysis on how trust im- pacts intended future use. To accomplish this, we developed a mixed-methods, human- centered investigation into voice assistant failures. We first executed interviews with 12 voice assistant users to understand what types of failures they have experienced and how this affected their trust and subsequent use of their assistant. We concurrently crowdsourced a dataset of failures from voice assistant users on Amazon Mechanical Turk. Finally, we executed a survey to quantify how different types of failures impact users’ trust in their voice assistants and their willingness to use them for various tasks in the future.
We found that different types of voice assistant failures have a differential impact on trust. Our interviews and survey revealed that participants are more forgiving of failures due to spurious triggers or ambiguity of their own request. In the case of spurious triggers, the voice assistant activates due to mishearing the activation phrase when it was not said. Users forgave this more easily, as it did not hinder them from accomplishing a goal. Failures due to ambiguity occurred when there were multiple reasonable interpretations of a request, and the response was misaligned with what the user intended while still accurately answering the question. Users tended to blame themselves for these failures. However, failures due to overcapture more severely reduced users’ trust, as when the voice assistant continued listening without any additional input, users considered their use a waste of time.
We additionally find that on many occasions, users would dis- continue using their voice assistant for a specific task for a short period of time following a failure, and then resume again once trust had been rebuilt. Trust was often rebuilt by using the voice assistant for tasks they considered simple, such as playing music, or alternatively, using the voice assistant for the same general task but in a different use case. In addition to these findings, we re- lease a dataset of 199 voice assistant failures, capturing user input, voice assistant response, and the context for the failure, so that researchers may use these failures for future research on how users respond to voice assistant failures. As voice assistants continue to perform increasingly complex and high stakes tasks across various technologists understand, prioritize, and address natural language failures to increase and maintain user trust in voice assistants.
Related Work
Prior research across many fields has examined the interaction between users and voice assistants, including human-computer in- teraction, human-centered AI, human-robotics interaction, science and technology studies (STS), computer-mediated communication (CMC), and social psychology. In addition, some work in natural lan- guage processing (NLP), especially NLP robustness, has approached technology failures in voice assistants and developed certain techni- cal solutions to address them. Here, we provide an interdisciplinary review of research relevant to voice assistant failures during user interaction across these fields. The literature review is organized as follows: 1) literature on user expectations and trust in voice assistants; 2) human-computer interaction (HCI) approaches to un- derstanding voice assistant failures and strategies for mitigation; 3) natural language processing (NLP) approaches to voice assistant failures, including disfluency and robustness.
Assistants
Researchers have long tried to understand how people interact with automated agents, especially comparing and contrasting these experiences with human-to-human communication. When talking with other humans, conversations can broadly be understood as functional (also known as transactional or task-based) or social (interactional), and many conversations include a mix of both .
Functional conversations serve towards the pursuit of a goal, and those who participate often have understood roles towards the pursuit of that goal. In contrast, social conversations have a goal of building, strengthening, or maintaining a positive relationship with one of the participants. These social conversations can help build trust, rapport, and common ground .
People generally expect to have functional conversations with voice assistants . The lack of social conversations may reduce users’ ability to build trust in their voice assistants. Indeed, past re- search has shown that users trust embodied conversational agents more when they engage in small talk , although this varies by user personality type and level of embodiment of the agent . As it stands, people report not using voice assistants for a broad range of work has illustrated the importance of trust for continued voice assistant use [31, 35], as trust is pivotal to user adoption of voice assistants [33, 45] and willingness to broaden the scope of voice as- sistant tasks . It is especially important to support trust-building between users and voice assistants as researchers continue to imag- ine and develop new capabilities for them, including complex tasks such as supporting healthcare tasks [41, 55], giving mental health advice [52, 66], and other high stakes decision-making .
This then begs the question of how trust is built between users and voice assistants. Trust in machines is an increasingly important topic, as use of automated systems is widespread . Concretely, trust can be conceptualized as a combination of confidence in a system as well as willingness to act on its provided recommenda- tions [37, 54]. Prior researchers have examined trust in machines A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures
Hi ’23, April 23–28, 2023, Hamburg, Germany
in terms of people’s confidence in a machine’s ability to perform as expected, benevolence (well-meaning), and integrity to adhere to ethical standards Broadly, past research has evaluated how various factors such as accuracy and errors affect people’s trust in algorithms [19, 20, 67]. In the case of voice assistants, Nasirian et al. and Lee et al. studied how quality affects trust in and adoption of voice assistants, and found that information and system quality did not impact users’ trust in a voice assistant, but interaction quality did. Interaction quality was captured based on a study by Ekinci and Dawes , in which Likert scale responses were captured regarding competence, attitude, service manner, and responsiveness of the voice assistant. In addition, customizing a voice assistant’s personality to the user can lead to higher trust , while gender does not impact users’ trust in a voice assistant .
Overall, prior research demonstrates the importance of the inter- action quality and social conversations for building trust between users and voice assistants, which in turn affects users’ willingness to continue using them and broaden the scope of their tasks.
Hci Approaches To Voice Assistant Failures
However, there are occasionally unforeseen breaches of trust, as not all interactions go as smoothly as one expects. Prior work has explored the diversity of issues affecting engagement and ongo- ing use of voice assistants and has shown that when users have expectations for voice assistants that surpass its capabilities, voice assistant failures and user frustration ensues [31, 35].
This begs the question, how has prior work defined failures in voice assistants? Some work uses specific scenarios in their stud- ies. For example, Lahoual and Frejus conducted evaluation in domestic and driving situations. They identified failures due to poor voice recognition, limited understanding of a command, and connectivity. Cuadra et al. used failures in specific tasks, such as attempting to give directions to an incorrect location, send a text to the wrong person, play the wrong type of music, or adding a re- minder with an incorrect detail . Mahmood et al. simulated online shopping, in which an AI assistant with a voice component the ambiguous item “bow” could mean a hair bow, archery bow, or bow for gift wrapping. Salem et al. had participants control a robot’s movement, and in the faulty condition, the robot would move erratically, incorrectly responding to the users’ input. Can- dello et al. defined failure as occasions in which someone asked a question that could not be understood or was out of scope of the voice assistants’ knowledge, in which case it would divert the conversation to ask an unrelated question.
assistant failures, drawing from theoretical frameworks of commu- nication between humans [10, 28, 46]. We reference Herbert Clark’s grounding model for human communication, which relies on four different levels to achieve mutual understanding: channel, signal, intention, and conversation . This was expanded by Paek and Horvitz , which applied these four levels to human-machine interactions and failure points. Channel level errors include when an AI fails to attend to a users’ attempt to initiate communication; signal level errors include an error in capturing user input (e.g.
due to transcription); intention level errors include mistakes in making sense of the semantic meaning of the transcribed input; and conversation level errors occur when a user has requested an unknown action to the AI (e.g. asking a weather app to schedule something). Hong et al. built on this model, specifically restrict- ing the context to NLP failures, rather than AI as a whole. Based on interviews with NLP practitioners, they renamed the categories as attention (channel), perception (signal), understanding (intention), and response (conversation). focused on failures that are either very common, or rare but very costly, to cover the most important and frequent failures users encounter when inter- existing taxonomy of NLP failures , narrowing the use case to only voice assistant failures, and evaluating how different failures impact on user trust and future intended use.
There is currently little systematic evaluation of the impact of voice assistant failures on user trust. Salem et al. found that if a robot had faulty performance, this did not influence participants’ decisions to comply with its requests, but it did significantly af- fect their perceptions of the robot’s reliability and trustworthiness.
Mahmood et al. found that voice assistants that accepted blame and apologized for mistakes were thought to be more intelligent, likeable, and effective in recovering from failures than assistants that shifted the blame.
Sometimes after a failure, users will try to reformulate, simplify, or hyper-enunciate their commands as a way to continue using the device [31, 35, 43, 61]. If users are repeatedly unable to repair failures with voice assistant, this weakens their trust and causes them to reduce their scope of commands to simple tasks with low risk of failure [31, 35]. Lahoual and Frejus found that in some situations, voice assistant failures can erode trust to the extent that users abandon voice assistants all together. However, not all failures require self-repair. A study by Cuadra et al. found that when voice assistants make mistakes, voice assistant self-repair greatly improves people’s assessment of an intelligent voice assistant, but it can have the opposite impact if no correction is needed. Thus, understanding which types of failures undermine trust the most may also inform us when failure mitigation strategies should be activated.
Nlp Approaches To Voice Assistant Failures
The NLP community has also examined voice assistant failures from a slightly different angle, focusing on the robustness of different NLP components underlying voice assistants, such as models for tasks in natural language inference , question answering [26, 42], and speech recognition . NLP robustness can be defined as understanding how model performance changes when testing on a new dataset, which has a different distribution from the dataset the model is trained on . In practice, users’ real world interactions with voice assistants could differ from data used in development, which mimics the data distribution shift in NLP robustness research.
Such data distribution shifts are shown to lead to model fail- ures. In the case of question answering, state-of-art models per- form nearly at human-level for reading comprehension on standard benchmarks collected from Wikipedia . However, Miller et al.
found that model performance drops when the question an- swering model is evaluated on different topic domains, such as
Hi ’23, April 23–28, 2023, Hamburg, Germany
Baughan et al. Figure 1: To analyze the impact of voice assistant failures on user trust, we used a mixed-methods approach, including inter- views and a survey. As part of the materials for our survey, we crowdsourced 199 failures from 107 voice assistant users, and include this dataset as part of our contributions.
New York Times articles, Reddit posts, and Amazon product re- views. Noisy input can also harm model performances. Lee et al. showed speech recognition errors have catastrophic impact on machine comprehension. Gupta et al. created a question answering dataset Disflu-QA where humans introduce contextual disfluencies, which also lead to model performance drops.
Although these works do not directly focus on voice assistant failures, topic domain changes, speech recognition errors and dis- fluencies are all very common during user interactions with voice assistants. Such similarities motivate us to draw parallels between the NLP robustness literature and HCI perspectives of system fail- ures. By understanding how different types of failures affect trust in voice assistants overall, we can then try to pinpoint the underlying NLP components that are the root cause of the most critical failures that erode trust . Technical solutions can then be leveraged to improve the robustness of the most critical parts of the system in order to increase user trust and long-term engagement most efficiently.
Ethod Overview
Now that we have established the importance of understanding of how voice assistant failures impact user trust, we proceed to conduct a mixed-method study. First, to prepare for the quantitative evaluation, we reviewed existing datasets in HCI and NLP to find failures that we could use as materials for our survey. Ultimately, the existing datasets were not sufficient for our needs. Therefore, we crowdsourced a dataset of failures from voice assistant users, which we also open source as part of the contributions of this study.
Concurrently, we conducted interviews with 12 voice assistant users to understand which types of failures they have experienced, and how this affected their trust in and subsequent use of the assistant.
These interviews were designed to provide a broad understanding of the thoughts, feelings, and behaviors that users have with regard to voice assistant failures and inform the quantitative survey design.
Finally, we executed a survey to quantify how different types of failures impact user perceptions of trust in their voice assistants and their willingness to use them for various tasks in the future. To report these findings, we first describe our process of collecting the crowdsourced dataset of failures, and how we selected a subset to use in our survey. Next, we present the interviews and survey, first describing our data collection and analysis, and then presenting the results concurrently.
Assistant Failures
The first goal in our investigation was to determine which types of failures users experience when using voice assistants. We first evaluated existing datasets for fit and breadth of failures. We deter- mined they were not sufficient for our purposes, so we proceeded to crowdsource a dataset of failures, adapting a taxonomy from Hong et al. to guide our collection. Finally, we cleaned and open-sourced this dataset as a contribution of our work.
A Review Of Existing Hci And Nlp Datasets
We first explored benchmark datasets in NLP, which contain a large number of either questions and answers , or conversational dialogue [25, 57, 68]. We found that existing NLP datasets do not cover the wide breadth of possible conversational failure cases due to their emphasis on correct data for training. Additionally, their focus on specific task performance, such as answering questions or dialogue generation, is more narrow than the variety of use cases for voice assistants. As training data relies on accurate task completion, A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures
Issed Trigger
Users say something to trigger the voice assistant, but it fails to respond.
Spurious Trigger
Users do not say something to trigger the voice assistant, but it activates anyways.
Elayed Trigger
Similar to system latency, the users say something to trigger the voice assistant, but it replies too late to be useful.
Noisy Channel
User input is incorrectly captured due to background noise.
Overcapture
The voice assistant captures more input than intended by either beginning to capture input too early or ending too late, and acting on external data not relevant to the users’ request.
Truncation
System does not fully capture users’ speech, by either beginning to capture input too late or ending too early.
Transcription
System generates a transcription error, often in the form of similar sounding words.
Ambiguity
There may be several interpretations of the users’ in- tent, and the system responds in a way that is plausibly accurate but not correct for the users’ intent.
Isunderstanding
The system maps the users’ input to an incorrect action, perhaps with some correct inference on the users’ intent, but not fully accurate.
No Understanding
The system fails to map the user’s input to any known action or response.
Action
If the system listens to the full request, but then turns off before giving any type of answer or taking action.
Correct Action
The system gives information that is incorrect. Table 1: Qualitative codebook and description of the various failures that were collected. We checked each failure for failure type sequentially, starting by checking if it could be an attention failure and progressing through the types until we found one that fit. From failure type, we then assessed which failure source applied.
these datasets did not contain failures. While testing these models produces a small percentage of errors (roughly 10%), the types of failures could only fall in the response and understanding categories, as attention and perception failures are excluded from the context of training these types of models. This limited their usefulness for our purpose of understanding voice assistant failures that occur in use and their impact on user trust.
In addition to these benchmark datasets, we investigated datasets that incorporated spoken word speech patterns, such as the Spoken SQuAD dataset and Disflu-QA dataset , as well as human- agent interaction datasets, such as the ACE dataset , the Niki and Julie corpus , and a video dataset of voice assistant failures .
In these cases, we found that the datasets were still restricted to only failures at the understanding and response level [26, 32] or the context for the failures was very specific and did not necessarily capture the breadth of possible failures users experience [1, 2].
Cuadra et al. ’s video dataset was the closest available fit for our needs, but we still found the use case of in-lab question-answering too narrow for our purposes. Therefore, we decided to crowdsource a dataset of voice assistant failures from users, and use these failures when conducting our quantitative survey on user trust.
Ataset Collection
Procedure. Crowd workers were asked to submit three fail- ures they had experienced with a voice assistant. They were asked about three specific types of failures out of a taxonomy of 12, which were randomly chosen and displayed in equal measure across all workers. The taxonomy of failures that we used to ask about spe- cific types of failures was adapted from previous work by Hong et al. , and identifies failures due to attention, perception, under- standing, and response, as shown in Table 1. Each question began by asking users if they could recall a time when their voice assistant had failed, based on the definitions in our taxonomy. For example, to capture missed trigger failures we asked “Has there ever been a time when you intended to activate a voice assistant, but it did not respond?” If so, we asked these workers to include 1. what they had said to the voice assistant, 2. how the voice assistant responded, 3.
the context for the failure, including what happened in the envi- ronment, and 4. the frequency at which the failure occurred from 1 (rarely when I use it) to 5 (every time I use it). These were all presented as text entry boxes except for the frequency question, which was multiple choice. Crowd workers were additionally asked to optionally share an additional failure that they had not had the
Hi ’23, April 23–28, 2023, Hamburg, Germany
Baughan et al. chance to share already. This was included to capture failures that did not fit any of the three the categories they were presented with, and we then categorized these failures according to our taxonomy.
assistant in the submitted examples, replacing activation words with “Voice Assistant” for consistency. We then edited grammatical and spelling errors for clarity. We also removed failures if they were not on-task, unclear, or exact repeats of other submitted failures.
Finally, we noticed that some of the categories the users submitted the failures under were incorrect, so we re-categorized the failures according to the codebook we developed as outlined in Table 1. Two raters iteratively coded 101 submitted failures, with a final coding session achieving an interrater agreement of 70%. One researcher then went back and coded the entire dataset in its entirety. In total, our finalized dataset contains 199 failures across 12 categories, submitted by 107 unique crowd workers.
Crowd Worker Characteristics. We used Amazon Mechanical Turk to recruit the crowd workers. In total, 107 crowd workers contributed to our dataset. We required workers to have the follow- ing qualifications: a HIT Approval Rate over 98%, over 1000 HITs approved, AMT Masters, from the United States, over the age of 18, and voice assistant users on at least a weekly basis. The plurality of users were in the age range of 35-44 (𝑛= 46), followed by 25-34 (𝑛= 32), and 45-54 (𝑛= 16), with the rest falling in 55-64 (𝑛= 8), 18-24 (𝑛= 1), and 1 preferring not to answer. Fifty-eight crowd workers were men, 44 were women, 1 preferred not to answer, and 1 identified as both a man and a woman. They used commercial voice assistants such as Amazon Alexa (𝑛= 59), Google Assistant (𝑛= 62), and Apple’s Siri (𝑛= 40), with many using some combination of the three (𝑛= 47). 91 crowd workers were native English speakers, and 13 were not. The plurality identified as White (𝑛= 58), and 39 identified as Asian. Three crowd workers did not provide any demographic information. The task took 15-20 minutes to complete on average, and they received $5.00 USD compensation.
Final Dataset. In total, our finalized dataset contained 199 failures from 107 users across 12 different types of failures according to the taxonomy based on Hong et al. , as updated in Table 1.
The failures we received most often were due to misunderstanding (𝑛= 38), missed trigger (𝑛= 25), and noisy channel (𝑛= 22). Users least often submitted failures for truncation (𝑛= 7), overcapture (𝑛= 7), and delayed triggers (𝑛= 8). Most crowd workers submitted failures saying that they happened “rarely when I use it” (𝑛= 87) or “sometimes when I use it” (𝑛= 84). Example failures across the 12 categories can be found in Table 2.
On average, the highest frequency of failures occurred for no understanding (𝑚= 2.15, sometimes when I use it, 𝑠𝑑= 0.67) and action execution: incorrect (𝑚= 2.00, sometimes when I use it, 𝑠𝑑= 0.88). The rest of the failure sources had an average reported frequency between 1.0 (rarely when I use it) and 2.0 (sometimes when I use it). The lowest frequency failures were due to delayed triggers (𝑚= 1.25, 𝑠𝑑= 0.46) and ambiguity (𝑚= 1.39, 𝑠𝑑= 0.78).
We then used 60 of the failures from our dataset in our survey to quantify the impact of different failures on user trust. This is outlined in more detail in the following section. This dataset has been open sourced1 for researchers to use to answer future research questions related to voice assistant failures in the future.
Nterview And Survey Methods
Once we had gathered and categorized our dataset of voice assis- tant failures, we were ready to answer our research question: how do voice assistant failures impact user trust? To do so, we first conducted exploratory interviews with 12 people to gather their thoughts, feelings, and behaviors after experiencing voice assistant failures. We used these findings and the failures collected in the dataset to then design and execute a survey. This quantified how various voice assistant failures impact users’ trust, as measured by their perceptions of the voice assistant’s ability, benevolence, integrity, and their willingness to use it for future tasks. Here, we describe the methods for both the interviews and survey, and we follow this by jointly presenting the results from both studies.
Nterview Methods
Interview Procedure. Interviews began with questions about why the participants chose to start using voice assistants and what types of questions they frequently would ask of them. We asked for common times and places they would use their voice assistants to understand their general experience with voice assistants.
Once these were established, we asked participants to tell us about a time they were using their voice assistant and it made a mistake, in as much detail as they could recall. We asked what they had been trying to do and why, if others were present, and if anything else was happening in their environment. We probed for users’ feelings once the failure occurred, and their perceptions about the voice assistant’s ability to understand them and give them accurate information. We asked participants what they did in the moment to respond to the failure. Finally, we asked questions about their use of the voice assistant in the aftermath, including how much they trusted it and if they changed any of their behaviors to mitigate future failures. All interviews were conducted remotely.
Interview Participants. During recruitment, we asked partic- ipants to submit their demographic information, how frequently they used voice assistants and on what types of devices. We addition- ally required participants to write a short (1-3 sentence) summary of a time they encountered a failure while using their voice assis- tant. We selected participants based on demographic distribution and the level of detail they included regarding the failure.
All of our 12 participants lived in the United States. They used voice assistants at least 1-3 times a week (𝑛= 2), with the majority reporting using a voice assistant every day (𝑛= 8), and the rest (𝑛= 2) using it 4-6 times a week. The majority of participants used a voice assistant on their mobile device (𝑛= 11), and five of these participants also used a voice assistant smart home device.
One participant only used a voice assistant smart home device. Participants reported using common commercial voice assistants such as Amazon Alexa (𝑛= 2), Google Assistant (𝑛= 7), and Apple’s Siri (𝑛= 8). Participants’ ages ranged from 18 to 50, with the plurality (𝑛= 5) in the age range of 18-23. 3 of our participants were 41-50, 2 were 31-40, and 2 were 24-30. Six of our participants 1https://www.kaggle.com/datasets/googleai/voice-assistant-failures A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures
Tell Her To Set A Timer For Ten Minutes, I Was
alone and no one present at the moment.
And I Was Calling My Coworker Sherry. The Voice
assistant mistakenly got turned on.
Elayed Trigger
It happened while I was driving a car.
Voice And Try Several Times To Be Heard By My
phone even though it was inches from my face.
Weather?
[It didn’t realize that my request had ended ] and
Overcapture
I was telling it to turn off the lights. I was the only one there. Voice Assistant, turn off the lights.
Asked The Voice Assistant To Calculate A Math
question, but it cut me off.
Asked For The Weather Conditions In The City I
live in. No others were present except for me.
Asked It To Play The Theme From Halloween. I
was sitting with my mother.
Oice Assistant, Play The Theme
song to the movie Halloween. [It plays a scary sounds soundtrack instead of the
Was Trying To Run A Routine To Wake Up My
kids. Voice Assistant, wake up the twins. "Sorry, I don’t know that." [However, I’ve set up a
Hi Come Out In Theatres?
[Pauses for a really long time, then turns its lights
Was At Home, In My Living Room, Alone. I Was
trying to find out how long Taco Bell was open.
Oice Assistant, When Does The
Taco Bell on Glenwood close.
Taco Bell, I Realized It Closed At 11:30Pm.]
Table 2: Table of voice assistant failures users submitted, including the context for the failure, what the user said, and what the voice assistant said. identified as women, five participants identified as men, and one participant identified as non-binary. Three participants identified as Asian, three identified as White, three identified as Black or African American, two identified as Hispanic, Latino, or Spanish origin, and one identified as both White and Black or African American. All of our participants spoke English as a native language. Participants were compensated with a $50 gift card and each interview lasted roughly 30 minutes.
Interview Analysis. Interviews were transcribed in their en- tirety by an automated transcription service and analyzed via a deductive and inductive process . We used deductive analysis to assess which types of failures these participants experienced. To ground our deductive analysis, we used the same codebook as we did for the dataset, as demonstrated in Table 1. We first identified in- stances in which participants were discussing distinct failures, and then applied our codebook to these instances. We used cues such as what was happening in their environment, and when appropriate, users’ own perceptions of why the failure occurred. We began by first identifying if failures belonged in which of the four failure types: attention, perception, understanding, or response. First, to determine if there was an attention failure, we investigated if there was evidence that the voice assistant accurately responded to an ac- tivation phrase, as indicated by visual or auditory cues, or otherwise by the participant’s narrative. Second, we evaluated if there was an error in perception, based on the participants’ assumption of if the voice assistant accurately parsed the input from the participant, our own assessment from their narrative, or other audio/visual cues.
Next, assuming that the input was correctly parsed, we sought to understand if the voice assistant accurately understood the seman- tic meaning of the input (understanding failures), using the same process. Finally, assuming all else had been correctly understood, we assigned response failures, indicating that the voice assistant either did not take action or took the incorrect action in response to an accurately understood command. Once a failure type was determined, we then further specified the failure sources as noted in Table 1. We resolved disagreements both asynchronously and in meetings, through discussion and comparison, over the course of several weeks.
While conducting this analysis, we also inductively identified themes related to these failures’ impact on future tasks and recovery strategies. To conduct this analysis, two researchers reviewed the twelve transcripts in their entirety, and one additional researcher reviewed five of these transcripts to further broaden and diversify themes. These researchers met over the course of several weeks to compare notes and themes, ultimately creating four different themes through inductive analysis. Of these themes, we report two
Hi ’23, April 23–28, 2023, Hamburg, Germany
Baughan et al. due to their novelty, specifically as related to future task orientation and recovery strategies.
Survey Methods
To quantify our findings from interviews, we developed a survey to explore users’ trust in voice assistants following each of the twelve different types of failures from our taxonomy, as well as their willingness to use voice assistants for a variety of tasks in the aftermath.
Procedure. The survey contained a screener, the core task, and a demographic section. We required participants be over 18 years old, use their voice assistant in English, and use a voice as- sistant with some regularity to participate. If participants passed the screener, they were required to review and agree to a digital consent form to continue.
The core task stated, “The following questions will ask you what you think about the abilities of a voice assistant, given that the voice assistant has made a mistake. Imagine these mistakes have been made by a voice assistant you have used before. Please consider each scenario as independent of any that come before or follow it. This survey will take approximately 20 minutes.” Participants were then presented with 12 different failure scenarios, and they were asked to rate their trust in two separate questions.
The first question measured trust in voice assistants as a con- fidence score across three dimensions: ability, benevolence, and integrity. These were selected because prior work on trust has determined these elements explain a large portion of trustworthi- ness [36, 40]. In the context of voice assistants, ability refers to how capable the voice assistant is of accurately responding to users’ input. Benevolence refers to how well-meaning the product is. And finally, integrity represents that it will adhere to ethical standards.
We asked participants to rate their confidence in voice assistants’ ability, benevolence, and integrity, as a percentage on a scale of 0-100, with steps of 10, to replicate how prior work has conceptu- alized trust . This was captured in response to the following
Statements:
• (Ability) This voice assistant is generally capable of accurately responding to commands. • (Benevolence) This voice assistant is designed to satisfy the commands its users give.
• (Integrity) This voice assistant will not cause harm to its users. The second question evaluated users’ trust in the voice assistant to complete tasks that required high, medium, and low trust. To select these tasks, we ran a small survey on Mechanical Turk with 88 voice assistant users. We presented 12 different questions, which first gave an example voice assistant failure (one for each failure source), and then asked “How much would you trust this voice assis- tant to do the following tasks:” give a weather forecast, play music, edit a shopping list, text a coworker, and send money. Users could choose that they would trust it completely, trust it somewhat, or not trust it at all.
There was not a significant difference in how much people trusted the voice assistant to play music compared to forecast the weather (𝑍= 2.06, 𝑝= 0.078). There was also not a significant difference in how much people trusted the voice assistant to edit a shopping cart or text a coworker (𝑍= 1.39, 𝑝= 0.21) as determined by pairwise comparisons, using 𝑍-tests, corrected with Holm’s se- quential Bonferroni procedure on an ANOVA of an ordinal mixed model. We found that there were significant differences between playing music, texting a coworker, and transferring money, with users having the most trust in the voice assistant playing music after a failure, less trust in texting a coworker, and still less in transferring money. Therefore, we selected playing music, texting a coworker, and transferring money to represent low, medium, and high levels of trust required. Therefore, after asking about ability, benevolence, and integrity, we asked participants how much they trusted their voice assistants to execute the following tasks: play music, text a coworker, and transfer money. These questions were displayed on a linear scale of 1 (“I do not trust it at all”) to 5 (“I completely trust it”), with steps of 1.
We completed the survey with an open-ended, optional question for participants to share anything else they would like to add. The survey concluded with demographic questions regarding gender, race, ethnicity, whether they were native English speakers, what type of voice assistants they used, and their general trust tendency as control variables. General trust tendency was measured based on responses to the following: “Generally speaking, would you say that most people can be trusted, or that you need to be very care- ful in dealing with people?” The options ranged from 1 (need to be very careful in dealing with people) to 5 (most people can be trusted). The questionnaire used for the survey has been submitted as supplementary materials.
Materials from our Dataset. To present each of the twelve failure sources in our survey, we drew from the dataset we had created. We selected five failures from each of the twelve categories.
We required that these failures had been coded by two of the team members who were in agreement (see dataset examples in Table 2). We used random selection to determine which of the five possible failures was presented to each user for each failure source. These are denoted in the dataset “Survey” column.
Participants. We recruited participants from Amazon Me- chanical Turk. We first ran a small pilot (𝑛= 27) in which we determined that participants completed the survey in roughly 20 minutes on average, and we set the compensation rate at $9 USD.
After removing participants who did not pass the attention check or straight-lined, meaning they responded to every question with the same answer, we had a total of 268 participants. These participants were required to have the following qualifications: AMT Masters, with over 1000 HITs already approved, over 18 years old, live in the United States, an approval rate greater than 97%, and they must not have participated in any of our prior studies.
The plurality of our participants were in the age range of 35-44 (𝑛= 106), followed by 25-34 (𝑛= 68), 45-54 (𝑛= 52), 55-64 (𝑛= 33), with 2-4 participants in each of the age brackets of 18-24, 65-74, and 75+. 134 of our participants identified as men, 132 identified as women, and 2 identified as non-binary genders. The majority of our participants were White (𝑛= 210), 21 participants were Black, and 15 were Asian. The rest of our participants identified as mixed race or preferred not to answer.
Authors:
Peder EZ Larson 1, 2,* , Jenna ML Bernard1, James A Bankson 3, Nikolaj Bøgh 4, Robert A Bok1, Albert P. Chen 5, Charles H Cunningham 6,7, Jeremy Gordon1, Jan-Bernd Hövener 8, Christoffer Laustsen 4, Dirk Mayer 9,10, Mary A McLean11 12, Franz Schilling13, James Slater1, Jean-Luc Vanderheyden5, 14, Cornelius von Morze 15, Daniel B Vigneron1, 2, Duan Xu1, 2, and the HP 13C
94143, Usa.
Denmark. 5 GE Healthcare, Menlo Park, California, USA. 6 Physical Sciences, Sunnybrook Research Institute, Toronto, Ontario, Canada.
8 Section Biomedical Imaging, Molecular Imaging North Competence Center (MOIN CC), Medicine, Baltimore, MD, USA. Cambridge, United Kingdom.
14Jlvmi Consulting Llc, Dousman, Wi, Usa
#See Acknowledgements for a list of all HP 13C MRI Consensus Group Members This work was supported by the ISMRM Hyperpolarized Media MR Study Group, the ISMRM Hyperpolarization Methods & Equipment Study Group, and the Hyperpolarized MRI Technology Resource Center (NIH/NIBIB grant P41EB013598).
Abstract
MRI with hyperpolarized (HP) 13C agents, also known as HP 13C MRI, can measure processes such as localized metabolism that is altered in numerous cancers, liver, heart, kidney diseases, and more. It has been translated into human studies during the past 10 years, with recent rapid growth in studies largely based on increasing availability of hyperpolarized agent preparation methods suitable for use in humans. This paper aims to capture the current successful practices for HP MRI human studies with [1-13C]pyruvate - by far the most commonly used agent, which sits at a key metabolic junction in glycolysis. The paper is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification. In each area, we identified the key components for a successful study, summarized both published studies and current practices, and discuss evidence gaps, strengths, and limitations. This paper is the output of the “HP 13C MRI Consensus Group” as well as the ISMRM Hyperpolarized Media MR and Hyperpolarized Methods & Equipment study groups. It further aims to provide a comprehensive reference for future consensus building as the field continues to advance human studies with this metabolic imaging modality.
Keywords: Hyperpolarized MRI, metabolic imaging, carbon-13, pyruvate, dissolution dynamic
Introduction
MRI with hyperpolarized 13C agents, also known as hyperpolarized (HP) 13C MRI, has shown great potential as a novel imaging modality, particularly for its ability to probe metabolic processes in real time. The first human studies with HP [1-13C]pyruvate were performed in 2011 in prostate cancer patients (1).
Since then, there have been over 60 papers published with imaging results of human subjects from 13 different sites, with applications including prostate cancer, brain tumors, breast cancer, kidney cancer, pancreatic cancer, metastatic disease, liver disease, ischemic heart disease, diabetes and cardiomyopathies. The vast majority of these studies used [1-13C]pyruvate (1–63), where [2-13C]pyruvate (64) and 13C-urea (56) have been demonstrated too.
As clinical HP 13C MRI advances, there is a growing need to build consensus for best practices, which are critical for comparing data across sites, performing multi-site trials,deploying methods to new sites, partnering with vendors, and potentially for obtaining broader regulatory approvals.
In March 2022, we initiated an effort to build consensus within the HP 13C MRI community with this opportunity in mind, and it was greeted with strong enthusiasm. The “HP 13C MRI Consensus Group”, containing over 55 members from 27 sites, identified the area of greatest need and opportunity for consensus building to be HP [1-13C]pyruvate human
●
Pyruvate is the most mature and widely used HP agent and has the most significant translational evidence emphasizing the potential clinical impact.
●
Clinical trials, particularly multi-site trials, have the strongest need for consensus methods to ensure that data can be combined across sites. This work is a Position Paper for which the goal is to describe current successful practices and study methods for HP [1-13C]pyruvate human studies along with justification to support those practices. This is divided into four major topic areas: (1) HP 13C-pyruvate preparation, (2) MRI system setup and calibrations, (3) data acquisition and image reconstruction, and (4) data analysis and quantification (Fig. 1). The current successful practices and study methods include a literature review of published peer-reviewed journal papers showing human HP [1-13C]pyruvate study data, up to September 2022 (1–63), as well as new unpublished information from surveys of HP 13C study sites. Based on this information, we also highlight the evidence gaps, strengths, and limitations of current practices which are summarized at the end of each section.
Figure 1: Illustration of the HP 13C MRI human study process, including the 4 major areas covered in this paper: Hyperpolarized 13C-pyruvate preparation, MRI system setup and calibration, Acquisition and Reconstruction, and Data Analysis and Quantification.
Figure 2: Anatomical targets of HP [1-13C]pyruvate MRI human studies published up to September 2022.
Hyperpolarized 13C-Pyruvate Preparation
This section covers the processes for creating the HP agent, 13C pyruvate, and will include many aspects and considerations that are needed to safely and effectively prepare doses for metabolic imaging studies in human subjects. These include material, personnel, equipment and facility, fluid path preparation, quality control, and release.
It is helpful to understand that the specifications of a dose of 13C pyruvate suitable for in vivo MR HP metabolic imaging were shaped in part by early preclinical studies performed by GE HealthCare summarized in Ref. (65). In short, the safety of the two novel drug components, 13C pyruvate and the electron paramagnetic agent (EPA) AH111501, were demonstrated in those studies. The more precise formulation of the dose suitable for human use was then determined from clinical studies (66) that included two Phase 1 clinical trials in young and elderly healthy volunteers without hyperpolarization of the 13C nuclei and another Phase 1/2a dose escalation and imaging feasibility study with HP 13C pyruvate in 31 prostate cancer patients at the With the exception of the first HP 13C imaging clinical trial, which utilized a prototype device in a cleanroom (1), all HP 13C studies performed in humans to date have utilized the SPINlab polarizer (manufactured by GE HealthCare). Consequently all doses of the HP 13C pyruvate delivered by SPINlab have been produced using the “SPINlab Pharmacy Kit” that serves as the container-closure system for the various drug components (13C pyruvic acid and EPA mixture, dissolution medium, and neutralization and dilution medium) during sample polarization, dissolution and quality control (QC) processes. Thus many aspects of the HP sample preparation considerations discussed below are related to the SPINlab instrument and the consumables designed to be used with it (67).
General Considerations
While more than 860 patients or healthy subjects having been injected with HP 13C pyruvate as of January 2022 without reports of any serious adverse events (68), HP 13C pyruvate injection remains an investigational MR contrast agent and can only be administered by those with Investigational New Drug (IND) exemption from the Food and Drug Administration (FDA) in the USA, a Clinical Trial Application (CTA) in Canada, approval from National Research Ethics Committee Services in the UK, or approval from the relevant local regulatory body. Thus, methods and processes involved to produce a dose should have patient safety as the first priority. Since utilizing dissolution dynamic nuclear polarization (dissolution-DNP) for human use is still a relatively new development, there are no existing published regulatory guidelines specifically for this method.
There are two major production styles that determine how various sites approach the agent preparation. In the US, the most common approach is to rely on a sterilizing filter (“Terminal Sterilization”) to ensure sterility of the final product, akin to PET tracer production, where a starting molecule with a radioisotope is processed using various other ingredients to make the final, desired and injectable contrast agent within a necessarily short amount of time (69). For these sites, sterilization of the components and accessories upstream of this filter are not required, although many of them were manufactured and tested following Good Manufacturing Practice (GMP) or Good Laboratory Practice (GLP) requirements. The filling process is usually performed under an ISO 5 laminar flow hood, but a clean room or an isolator is not required.
This approach is typically accompanied by testing the integrity of the sterilizing filter prior to release of the dose for injection. Typically, post release endotoxin and sterility tests are performed using an aliquot reserved from each released dose.
In the UK and EU, the most common approach is to more-closely follow sterile pharmaceutical compounding guidelines (70), where all components and ingredients are required to be sterile or manufactured under GMP guidelines and are assembled and filled within a clean room environment or an isolator system (“Sterile Preparation”). Typically a batch of Pharmacy Kits for HP 13C pyruvate injection are prepared together. The sterility of the final dose is also ensured by batch validation testing, in addition to the sterility of the ingredients and the sterile compounding process. The endotoxin and sterility testing are performed for the process validation but are not performed for each injected dose.
Some institutions fill and assemble the Pharmacy Kit required for a specific study on the same day or the day prior to polarization, dissolution, and patient administration, but others have also demonstrated the feasibility of preparing a batch of kits, keeping them in a -20ºC freezer and using them over a period of a few months.
Beyond the obvious requirements that the process and the facility has to ultimately produce a dose that is safe to inject into a human, regulatory authorities will also focus on the question “Are you in control of your processes?”. To be in control of your process requires an in-depth and broad understanding of all processes involved in pre, post, and during the production process.
Personnel
It is typical and may be required to have licensed personnel involved in the production process depending on local regulations.Typically a pharmacist, radiopharmacist or other similarly qualified person (QP), in charge of the facility where the Pharmacy Kit filling and preparation is taking place, is responsible for the overall process and the release of the injectable dose.
Qualified cleanroom technicians are often involved in the Pharmacy Kit filling under the supervision of the pharmacist or QP. As is required for pharmaceutical compounding or PET tracer production, training requirements and training records for all personnel need to be maintained and available for audit by the FDA or equivalent.
Equipment And Facility
The facility and all equipment need to have standard operating procedures (SOPs) that describe how equipment is used, maintained, and calibrated to comply with relevant legislation. Currently, almost all the filling of the Pharmacy Kit takes place within a compounding laminar flow hood or isolator (typically ISO 5). At some sites, the filling is conducted within a cleanroom, while at others, it is conducted in a dedicated non-cleanroom space, reflecting differences in cleanroom approach and specifications between regulators worldwide (71). Some equipment or facilities, such as the compounding hood or cleanroom, may require external certified laboratories for testing.
Material Handling
Material handling guidelines (69,70) require SOPs detailing a system to track all of the materials involved in the HP production process for a particular patient dose, similar to current good manufacturing practice (cGMP) requirements for material handling for drug compounding. This includes acceptance standards, storage conditions, amount used in the patient dose for each ingredient and materials used in the assembly of the fluid path and Pharmacy Kit. Currently some users choose to open and inspect and sometimes modify the Pharmacy Kits upon arrival, but some users keep them in the sealed packaging until they are required for dose preparation.
Pharmacy Kit Filling And Assembling
As required by an IND or its equivalent, the preparation of the doses of HP 13C agent are detailed in the Chemistry, Manufacturing, and Control (CMC) section of an applicable regulatory submission; an example of this has been made available (72). It describes the processes of filling the Pharmacy Kit with the different components that make up the final drug product, and of assembling the final kit for either storage or immediate use in the polarizer. Special attention should be given to the laser welding process in order to satisfy installation qualification (IQ) and operational qualification (OQ). Typically, the final developed process is validated by process qualification (PQ) runs, during which 3 or more Pharmacy Kits are filled and used and the final HP 13C products are tested for endotoxin and sterility and to confirm that they meet the dose specifications for injections (usually including pyruvate concentration, residual EPA concentration, pH, liquid state polarization level and dose temperature). The data from 3 consecutive PQ runs are submitted as part of the IND submission (or its equivalent), and are often also reviewed by the Institutional Review Board (IRB) where the studies are conducted.
Quality Control And Dose Release
The quality control (QC) and dose release can be separated into two aspects: one is the QC and release of the filled Pharmacy Kit, and second is the QC and release of the HP 13C agent for injection, after polarization and dissolution. For institutions filling a batch of kits and storing them to use over a period of time, typically the batch can be released based on initial validation, environmental monitoring data from the day of kit production, and if filters are used during preparation of any of the components, filter integrity testing. But in some cases one or more kits are used for validation before the batch of kits are released for future use. For institutions that fill only the kits required for specific studies shortly before the experiment, the filled kits often do not go through separate release tests before they are used.
The quality control of the HP 13C pyruvate solution post dissolution is primarily performed to ensure that the agent meets the dose specifications (Table 1) before it is administered to the subject. These specifications target both safety (pH, residual EPA, temperature) and efficacy (pyruvate concentration, polarization, volume). Typically, the pyruvate concentration, residual EPA concentration, pH, dose temperature, dose volume, and liquid state polarization are measured by the QC accessory associated with the SPINlab polarizer. Some users perform a secondary measurement for one of the parameters, such as pH, using a different instrument or pH paper. For sites that do not go through a separate release testing process for batch filled kits, the integrity of the sterilization assurance filter, a part of the Pharmacy Kit, is typically tested as a part of the dose release. It is also common for these users to preserve an aliquot of the final HP 13C pyruvate solution for post-release endotoxin and sterility testing. This testing cannot be completed fast enough to test an individual dose prior to injection, but this is why other processes such as PQ runs and validation testing are done to minimize the chance a subject could be injected with a contaminated dose.
The Final Dose Release And Injection
should be done under the supervision of a licensed professional, based on local regulations.
Some Key Challenges
Many of the challenges associated with HP 13C pyruvate preparation can be attributed to the conditions required for the dissolution-DNP method of high magnetic field (~3-7 T) and very low temperature (~1 K) during polarization, with pressurized and superheated water necessary for the rapid dissolution event. These extreme conditions are quite challenging for the design of the container-closure and fluid path system. In particular, the cryogenic temperature in the polarizer requires special attention to any moisture or ambient (moist) air introduced into that portion of the fluid path, which can form an ice block at ~1 K. This ice can lead to flow restriction during the dissolution event and reduce the strength of the laser welded bond between the cryovial and its cap. This can ultimately produce failures in the dissolution step, including variations in final pyruvate concentration and pH that may fail to meet QC release criteria as well as fluid path ruptures that provide no available dose and result in polarizer down-time.
The polarization of the HP 13C pyruvate sample decays quickly over the span of a few minutes after dissolution, and thus the process of dissolution, QC for release, and injection should be completed as fast as possible to preserve the high polarization level achieved. Any delays in the preparation process, such as transportation time or equipment malfunction, can significantly reduce the final polarization and result in lower quality imaging data.
Current Practices
A summary of data collected from all sites performing clinical trials with HP 13C-pyruvate is shown in Fig. 3 and Table 1, including the specification of the final dose and how the quality control and release of the final dose are performed. There is a split in the Production Style, described in the General Considerations section above, with 8/13 sites using Sterile Preparation versus 5/13 using Terminal Sterilization. While many of the dose specifications show notable differences in acceptable ranges, all of these variations listed in tables have been successfully and safely been used to perform HP 13C pyruvate studies in humans. Their differences depend on the institutions’ preferences, resources and their particular regulatory situation. There is high similarity in pyruvate ranges, temperature ranges, EPA limits, and volume limits. There is modest variability in pH ranges and large variability in the endotoxin test limit. There is a 3-fold difference in acceptable polarization levels, which are measured to ensure a futile dose is not injected since the polarization is directly proportional to SNR. This reflects the decision by several sites to believe that useful data can be still be obtained with suboptimal polarizations.
Figure 3: Hyperpolarized agent preparation methods reported by sites currently performing HP
In House
Table 1: HP 13C-pyruvate preparation parameters, methods, and dose specifications used for quality control testing and release as well as validation. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. The parameters used for product release are noted in bold text, otherwise these parameters are measured for batch validation or other QC measurements. The endotoxin and sterility testing are performed during process validation of the batch and/or post-injection, and largely depends on the agent production approach.
Summary
The overall safety record of HP 13C-pyruvate has been very strong, and the SPINlab hyperpolarizer has proven to provide high polarizations at human sized doses while meeting numerous QC and release criteria. A weakness remains the failure modes of the SPINlab Phamacy Kits (e.g. ice blocks, path ruptures), which are placed under extreme requirements particularly during dissolution. The preparation process still requires a high degree of expertise.
Therefore, there is a significant need to improve the reliability, robustness, and ease of operation for generating HP 13C-pyruvate doses for human studies. Furthermore, there is a divide between manufacturing and sterile compounding style preparation as well as other site-specific practices, resulting in variations in SOPs and justification required to relevant regulatory bodies. There have also been no comparisons between these approaches. It is also unclear what release criteria and QC parameters are truly required to ensure patient safety.
However, all of the reported methods are acceptable and approved by the appropriate regulatory authorities, and have led to the rapid expansion of successful human studies in recent years.
Mri System Setup And Calibrations
This section covers the MRI system setup, including the imaging system, RF coils, phantoms, and prescan calibration methods.
Imaging System
The main prerequisite for a given MRI scanner to be capable of supporting studies with HP 13C is its “broadband” capability to transmit and receive radiofrequency (RF) signal at the frequency of 13C, which is around 4 times lower than 1H. This does not come as a default on clinical MR devices. The transmit power of the broadband amplifier should also be sufficient to support the intended flip angle and RF pulse shape with the employed transmission RF coil(s) for 13C. Most studies to date use relatively low flip angles (< 90 degrees) for HP 13C in order to preserve polarization for time-resolved imaging. The capability to receive 13C signal on multiple channels is also desirable to increase SNR, as discussed further in the “RF coils” section.
The choice of magnetic field strength is primarily dependent on the metabolites’ frequency separation due to chemical shift dispersion and 1H imaging. High field strengths do not enhance hyperpolarized 13C signal as they do for 1H because the signal strength in a HP experiment relies on manipulating the population of quantum energy states outside of the MRI scanner.
However, the injected HP 13C-pyruvate and its metabolic products have greater frequency separation at higher fields, and it may thus be easier to separate and quantify these resonances at higher fields. This comes at the cost of a reduction in the achievable T2* and often reduced T1. As the initial polarization is independent of the imaging field strength it has been proposed that the increased T2* at 1.5T can potentially be exploited to increase SNR by adapting the acquisition bandwidth or reduce off-resonance imaging effects in cases when the decay of the transverse magnetization is dominated by T2* (73). In practice, 3T has been used in all published human 13C-pyruvate studies surveyed (Supporting Table S1), and comprises the majority of scanners currently in use for human studies (Table 3). A field strength of 3T is well-suited for 1H MRI anatomical reference and correlative imaging.
Stronger and more rapidly slewing magnetic field gradients support more rapid spatial encoding, particularly for metabolite-specific single-shot imaging using echo-planar imaging (EPI) or spiral imaging (See “Acquisition and Reconstruction”). Although the spatial resolution acquired for HP 13C imaging is typically much coarser than for 1H MRI, the factor of ~4 in gyromagnetic ratio leads to the same reduction factor in performance of the gradient system, so 13C experiments are potentially more limited by gradient hardware performance. To date, all human studies have used the commercially-available integrated gradient systems provided in clinical MRI scanners.
Optimization of scanner design has understandably focused on minimization of artifacts in 1H MRI, where devices such as room lights, the gradient amplifiers, and the motors driving the patient bed are checked to ensure that they do not produce RF interference at the 1H frequency, but artifacts may arise at other frequencies. Eddy current compensation is also not always appropriately adjusted for nuclei at other frequencies (74). In order to optimize for 13C, many sites have performed checks on phantoms for RF interference, gradient artifacts, and eddy currents (74), including the use of post-hoc gradient impulse response function characterisation and correction, and some vendors have fixed these issues as well.
Rf Coils
For HP 13C imaging studies in humans, RF coils for both 1H and 13C nuclei are needed, with 1H MRI providing an anatomical reference for registration and optional additional multiparametric MRI readouts. At the Larmor frequency of 13C nuclei, the relative contributions from coil noise compared to sample noise increase compared to 1H (73,75), although sample noise still is likely the dominant contributor for human-sized coils at 32.1MHz - the resonance frequency of 13C nuclei at 3T.
The key requirement for human 13C-pyruvate RF coils are that the coil geometry and sensitive volume must cover the volume of interest in the subject. Table 2 and Figure 4 shows coil configurations that have been used and optimized for applications in different anatomic regions.
Volume resonators are most commonly used for transmit, as they surround the subject to
Provide B1 Transmit Across The Fov (B1
+). While 1H relies on a large birdcage (“body”) coil built into the scanner, 13C transmit coils must be placed inside the bore. This takes up valuable space within the magnet, and also has led to the use of designs with relatively inhomogeneous
B1
+. Many human studies have used Helmholz pair resonators for transmit, including the “clamshell coil”, which has a notably inhomogeneous B1
+ Profile But Has Been Used Because Of
relatively easy integration into the scanner bore. B1
+ Variation Results In Variations In The Flip
angles that control the use of the hyperpolarized magnetization and creates errors in common HP metrics (9,76). The exception are head coils, where birdcage designs with highly
Homogeneous B1
+ can be placed around the head while easily fitting inside the bore. As with 1H MRI, higher SNR can typically be achieved by smaller receive coil elements, such as surface coils or phased arrays, and the majority of 13C receive coils used have layouts similar to 1H phased arrays.
RF coil quality control is important to ensure proper functioning of the coils to provide consistent imaging quality, especially with limited natural abundance 13C signal in vivo. It typically involves 1) a physical integrity check of the coil cables and connectors and 2) phantom SNR tests to check the coil’s performance and to monitor it over time (see Phantoms below). An useful reference for RF coil quality control is outlined in the MRI accreditation program of the American College of Radiology (77) and can be adapted for 13C coils.
Notably, configurations for brain and prostate studies used dual-tuned 1H/13C coil designs, which greatly simplify workflow and registration of 1H and 13C images, as no switching of coils is needed.
(1)
Table 2: RF coil configurations reported for human HP [1-13C]pyruvate studies.
Tx = Transmit
coil, RX = receive coil. The commonly used “clamshell” TX coil is a Helmholz pair design. For 1H RF configurations, all used the Body coil for TX unless otherwise noted, and “repositioned” indicates the 13C coil was removed for 1H imaging. One representative reference is listed for each configuration. The RF coil configurations reported in the reviewed papers are shown in Supporting Table S1.
Figure 4: Examples of RF coil configurations used for human HP [1-13C]pyruvate brain studies. (A,B) 13C Clamshell TX (Helmholz pair) and 2× 4-channel paddle RX arrays. (C) 13C Birdcage volume TX and 32-channel RX array (RX array slides into TX coil). (D) 13C Birdcage volume TX and 24-channel RX array, combined with a 1H 8-channel RX array. Image reproduced with permission from Ref (16).
Phantoms
Since hyperpolarized magnetization is non-renewable, phantoms containing 13C nuclei are important to: 1) test the multi-nuclear capabilities of the imaging system, including all parts of the signal excitation and receive chain; 2) perform calibration measurements before a scan with hyperpolarized nuclei; and 3) perform necessary pre-scan adjustments (see “Prescan Calibration” section). The phantoms currently in use are listed in Table 3. Their composition must provide sufficient 13C signal, with additional considerations of conductivity, stability, chemical shift(s) present, potential for dynamic imaging, and cost. The phantom geometries are typically either compact, in order to be used alongside the subject during a HP scan, or large enough to mimic the inner volume of a RF coil for system testing.
One popular compact design contains enriched 13C-urea at high concentration, typically 8 M, which provides a single resonance, placed inside a small container ~1 mL. The most common recipe mixes 13C-urea in a 90% water/10% glycerol solution, with glycerol used to increase the urea solubility and doping with a Gd-based contrast agent to shorten T1 which increases the potential SNR per unit time. For example, when Dotarem is added at a 3:1000 volume ratio the 13C-urea T1 is around 500 ms and T2 is around 100 ms. However, when testing pulse sequences influenced by T1 and T2, doping should be used carefully. This phantom is suitable for frequency calibration, transmit gain calibration, sequence testing, and as a fiducial marker when placed next to a patient. However, enriched 13C-urea has a relatively high cost compared to natural abundance compounds.
For larger volumes (>100 ml), the phantoms most often used contain undiluted ethylene glycol, glycerol, or dimethyl silicone. These compounds have sufficiently high carbon concentrations to provide sufficient 13C signal even with the 1.1% natural abundance of 13C. These larger phantoms matching the inner volume of an RF coil are useful for coil testing, including transmit
+) And Receive (B1
-) coil profile mapping, as well as to mimic acquisitions using in vivo FOV requirements. In this case, size and conductivity should match the expected subject size in order to mimic coil loading and get a realistic estimation of B1+. Large-volume natural abundance urea phantoms have also been used by some sites, but suffer from higher conductivity compared to biological tissues. Typically, it is easier to increase the conductivity and hence coil loading of the non-conductive phantom by adding NaCl to match physiological loading (16,78).
Dynamic phantoms that aim to mimic metabolite kinetics have also been developed (79–81), and have the potential to more closely mimic the HP experiment, but so far these are not widely used.
Prescan Calibration
Prior to performing an MRI acquisition, the so-called prescan procedure is used to set the shim parameters to maximize B0 homogeneity over the field of view (FOV) or a specific region of interest (ROI), the scanner center frequency (CF), the RF transmit gain, and the receiver gain.
While this calibration procedure is usually automated for 1H, the lack of sufficient natural abundance 13C signal prevents use of automated methods. (Although natural abundance 13C lipid signal has been detected, there are so far no reports on using this signal for prescan.) Table 3 shows current practices across sites.
Maximizing B0 homogeneity is independent of the nucleus and is therefore performed prior to 13C imaging using the 1H water signal and existing shimming tools, such as by a standard automated process (“Auto Shimming”) or using high order shimming routines. Similarly, the 13C CF can be calculated from the 1H CF using a predetermined scaling factor that depends on the target chemical shift (82). Another common approach used is to have a small, high-concentration 13C phantom, e.g. 8M 13C-urea, integrated in the RF coil or placed next to the scan subject (1). The reference frequency can also be based on real-time measurements after the HP injection but prior to imaging (83). Both the CF and B0 shimming are critical when using spectrally-selective RF pulses, as inmetabolite-specific imaging methods, where the desired excitation bandwidths are typically very narrow and frequency offsets can lead to a failure mode that is only apparent after injection.
The calibration of the RF transmit power is typically performed on a small, high-concentration 13C phantom placed near the region of interest during the scan or on a large 13C phantom of similar size and coil loading as the subject, prior to the subject scan. Reference power is often done by sweeping the power in a pulse-acquire sequence (53,62), or the Bloch-Siegert method (52,84). When using a small phantom, the location of the phantom, B1
+ Inhomogeneity As Well
as any shielding effects, e.g., when the phantom is integrated into a coil (1), may degrade the accuracy. Other methods include real-time Bloch-Siegert method measurements after the HP injection (83), and using the stronger natural abundance 23Na signal that is close enough to the 13C resonance frequency to be detected by 13C coils (82).
The receiver gain is predetermined, either systematically based on independent phantom measurements and assuming the dose and polarization of the HP compound is known prior to injection, or based on past HP imaging studies.
Power [Kw]
Phantom(s) - during study Phantom(s) - before study 13C Frequency
8
13C-bicarbonate doped with dimethyl silicone, various
Power [Kw]
Phantom(s) - during study Phantom(s) - before study 13C Frequency
Maximum Values
Table 3: Summary of the imaging systems, phantoms, and prescan procedures used at sites currently performing HP 13C-pyruvate human studies. These were obtained from a survey of all sites performing clinical trials with HP [1-13C]pyruvate. *Previously performed studies with a Siemens 3T Tim Trio. The imaging systems, phantoms, and prescan procedures reported in the reviewed papers are shown in Supporting Table S1.
Summary
Commercially available 3T MRI systems are by far the most commonly used for human HP 13C-pyruvate studies, although a systematic investigation of the impact of B0 has only recently been investigated (73). The multi-nuclear RF transmit and receive chain has proven sufficient for current acquisition strategies, although many sites have observed artifacts due to RF interference, gradient interference, and residual eddy currents when operating at the 13C frequency. A variety of 13C RF coils, tailored for numerous anatomical targets, have been successfully demonstrated, with the main limitation that most transmit coils take up a lot of additional space inside the bore and provide relatively inhomogeneous B1
+ Profiles. The
phantoms used have converged into generally 2 categories - small phantoms containing 13C-enriched compounds that can be used during the study and human-sized phantoms containing compounds with high carbon concentrations but without 13C enrichment that are used to test and calibrate the coils. There are no standardized compositions or geometry, and dynamic phantoms that recapitulate in vivo kinetics would be desirable but are still an emerging area. Prescan calibration procedures were not well defined in most publications, so we surveyed individual sites to determine current practices. Calibration procedures for the B0 field (13C CF and shimming) for most sites take advantage of 1H signal and methods, while methods
For Calibration Of B1
+ is more variable across sites, likely a reflection of remaining challenges in how to perform this calibration. Standardization of both phantoms and calibration procedures would synergistically improve the robustness and reproducibility of HP 13C studies.
Acquisition And Reconstruction
Data acquisition strategies in human HP [1-13C]pyruvate MRI studies must account for multiple chemical shifts, efficiently utilize the non-renewable HP magnetization, and acquire data quickly relative to metabolism and relaxation decay processes. These studies require spectral encoding to separate metabolites, necessitating pulse sequences that efficiently encode up to 5D data (3 spatial + 1 spectral + 1 temporal dimension). RF pulses must efficiently sample without immediately saturating the non-renewable HP magnetization, and sequences must acquire data quickly and be robust to both experimental and physiologic variation (e.g. B1
+ Inhomogeneity,
variation in perfusion) to ensure reproducibility and minimize scan-to-scan variability. This section covers current successful practices for data acquisition in human [1-13C]pyruvate studies, and accompanying 1H imaging, from different anatomic regions, including scan parameters and image reconstruction.
Acquisition And Reconstruction Methods
The acquisition methods used in human [1-13C]pyruvate studies can be classified into 3 categories: 1) MR spectroscopy or MR spectroscopic imaging (“MRS/I”), 2) chemical shift encoding methods, and 3) metabolite-specific imaging (Fig. 5).
Mrs/I Methods Specifically
resolve a spectrum that can be analyzed to extract expected as well as unexpected resonances, making this approach very robust. It was used in many initial studies (1).
Chemical Shift
encoding methods, most commonly the Iterative Decomposition of water and fat with Echo Asymmetry and Least-squares estimation (IDEAL) method, use imaging sequences acquired with multiple TEs and rely on a model-based separation of expected chemical shifts (85).
Metabolite-specific imaging methods use specialized RF pulses that are spatially and spectrally selective to excite individual metabolites which are then typically imaged with fast k-space trajectories such as echo planar imaging (EPI) or spirals (86).
Their Application To Different
organ systems is described below. The image reconstruction methods used in human [1-13C]pyruvate studies have typically been conventional methods (e.g. FFT, non-uniform FFT, or equivalent). The incorporation of accelerated imaging and advanced reconstruction methods including parallel imaging (4,57,87) and compressed sensing (7) has also been applied in human studies for improved spatial resolution, temporal resolution and coverage, but have the potential for additional artifacts as well as SNR losses due to ill-conditioning of the reconstruction (e.g. g-factor).
The Majority Of
published studies do not use accelerated imaging indicating the resolution and coverage achievable without acceleration is currently adequate for successful data collection. Performing coil combination, even with fully sampled data has also been shown to have specific challenges for HP human images: using naive sum-of-squares methods suffer from high noise amplification in the relatively low SNR regime of HP [1-13C]pyruvate (compared to 1H), motivating several HP 13C-specific methods that include data-driven coil sensitivity estimation which have shown obvious improvements over sum-of-squares (11).
More recently denoising techniques have been applied as post-processing of human HP data(41,42,44). The techniques applied are based on spatial-temporal singular value decomposition for unsupervised estimation of signal and noise components. They have shown improvements in apparent SNR in the brain and liver, while care must be taken to choose parameters such as the rank threshold to avoid oversmoothing and overfitting to the estimated signal components.
Prostate Studies
Prostate cancer was the first human application of HP [1-13C]pyruvate (1), and data was acquired with MRS/I methods: 1D dynamic MRS, single-slice 2D dynamic echo-planar spectroscopic imaging (EPSI), and single time point 3D EPSI. Advances in imaging strategies led to the development and application of new acquisition schemes, including undersampled 3D EPSI with compressed-sensing (7), model-based chemical shift encoding methods that use a priori information (47,59), and metabolite-specific EPI (10), all of which can provide volumetric whole-organ coverage and dynamic acquisitions.
The pyruvate bolus arrival in the prostate can vary by ± 10 s between patients, necessitating dynamic imaging to reliably and consistently capture the pyruvate bolus (18). For this reason, all currently ongoing studies acquire dynamic data. While MRS/I, chemical shift encoding, and metabolite-specific imaging can all achieve dynamic imaging, chemical shift encoding and metabolite-specific imaging provide greater dynamic and volumetric coverage (85). For scan prescriptions, the FOV is designed to provide full prostate coverage and typically to match the orientation of the anatomic imaging used for registration. Flip angles used in current studies are constant through time, as quantification with a variable-through-time flip scheme is highly sensitive to bolus timing (8) and errors in the RF transmit (B1 +) field (76).
Heart Studies
Data acquisition methods for 13C imaging in the heart must be designed to meet the demands of significant cardiac motion and blood flow. To cope with the periodic cardiac motion, most human heart studies to date used gating to the diastolic window, the longest cardiac cycle interval, which has reduced motion (2,22,28,30,35,36,38,45,52). The duration of the diastolic window limits the available data sampling time, making cardiac acquisitions the most time-constrained of the HP 13C MRI applications. The most common acquisition approach is metabolite-specific imaging with spiral k-space trajectories (2). Their single-shot imaging capability makes these methods particularly robust to motion effects. Furthermore, spiral k-space trajectories provide rapid k-space coverage and relatively benign flow and motion artifacts. The majority of studies have used 2D multi-slice acquisitions, but 3D encoding has also been used successfully (35).
Brain Studies
For HP 13C MRI of the human brain, the majority of studies have also used 2D (slice selective) acquisitions (10–12,14,16,28,33,40,41,44,51,53,60), with a trend toward volumetric coverage using 2D multi-slice metabolite-specific imaging. 3D metabolite-specific imaging of the whole brain, with phase encoding of the slice direction (34,57), has been shown to provide similar SNR efficiency (88) compared with multislice imaging. A number of studies have employed MRS/I (5,6,29,31–33,50,55) resulting in a spectrum from each voxel, which has the advantage of not requiring a priori information about which peaks to encode. This was important in early brain studies when it was not known which peaks would be detectable. Chemical shift encoding, using a set of images with different echo times and an iterative reconstruction of the individual resonances (i.e. the IDEAL approach (85)), has also been used (12,49,54), with the drawback that coverage in the slice direction was limited due to the time required to acquire multiple echo time images.
Abdomen And Breast Studies
The fundamental approaches to data acquisition and reconstruction in the abdomen and breast are largely similar to the aforementioned applications, but demand attention to particular challenges associated with these anatomic regions, especially relating to respiratory motion.
Although it has been shown that a basic 2D MRSI approach based on phase encoding and FID readout can be successfully applied for HP 13C imaging in breast (15) and kidney (13), major advantages in terms of spatiotemporal resolution and coverage have been realized using tailored approaches based on metabolite-specific imaging (43,62) and chemical shift encoding (43), which have facilitated multi-slice or 3D dynamic acquisitions over large FOVs in the abdomen (4,37,46).
The significant respiratory motion encountered in these regions can directly blur 13C images, and has further favored these rapid acquisition strategies. Motion also degrades B0 homogeneity, which can shift frequency-selective excitation profiles and introduce artifacts into rapid imaging readouts. This makes accurate determination of the acquisition center frequency and shimming essential in these regions which often cover large FOVs. (See “Prescan Calibration” section for more information). In some studies, breath-holding was used to minimize motion effects and enforce frame-to-frame data consistency (42). A pragmatic and reasonably effective approach for dealing with respiratory motion during 13C data acquisition is an initial breath-hold (as long as can be tolerated), followed by free-breathing (46,62).
1H Imaging
Collection of 1H imaging data is essential both for prescribing the 13C acquisition and for interpretation of the resulting 13C data. Multi-planar 1H scouts are acquired prior to 13C acquisition to enable graphical prescription of the 13C imaging region. All human HP 13C-pyruvate imaging studies acquire conventional MRI scans (e.g. T1- and T2-weighted volumes) for anatomic reference, aiming to cover at least the full 13C FOV. Acquiring these anatomic scans as close as possible to the time of 13C imaging (immediately before or after) minimizes potential misregistration between the data sets. Depending on the application, other advanced 1H sequences are also acquired (e.g. diffusion-weighted imaging for cancer imaging).
When contrast-enhanced data is acquired, it is done after 13C imaging, as paramagnetic contrast agents will accelerate 13C relaxation.
Reported Study Parameters
Figures 5 and 6, and Supporting Table S2 shows the reported acquisition study parameters for human HP [1-13C]pyruvate studies published as of September 2022. Figure 5 shows a mixture of MRS/I, metabolite-specific imaging, and chemical shift encoding methods have been successfully used, where spectroscopy-based methods have become less prevalent in recent studies. Figure 6 shows the acquisition timing, including the important start time and interval/temporal resolution, is quite variable across studies.
Figure 5: Acquisition methods used in published HP [1-13C]pyruvate human studies published up to September 2022, classified into: MR spectroscopy and spectroscopy imaging (MRS/I); chemical shift encoding methods, such as IDEAL, that use multiple TEs and model-based reconstructions; and metabolite-specific imaging methods that use spectrally-selective excitation to image a single resonance at a time.
Figure 6: Temporal acquisition characteristics reported in HP [1-13C]pyruvate human studies published up to September 2022. (a) Reported referencing of acquisition start times.
(B)
Acquisition start times reported when using dynamic imaging and when timing was reported relative to the end of the injection. (c) Temporal resolutions. “Not Applicable” indicates dynamic imaging was not used.
Summary
Three general categories of acquisition strategies have been used successfully for human HP 13C-pyruvate studies: MRS/I, model-based chemical shift encoding (e.g. IDEAL) methods, and metabolite-specific imaging methods. These have enabled successful studies in the prostate, heart, brain, abdomen, and breast. Recent studies increasingly have used the imaging-based strategies of metabolite-specific imaging and chemical shift encoding which are the fastest methods, although a heads-to–head comparison between techniques has not been performed.
Metabolite-specific imaging is quite popular because of its speed and compatibility with single-shot imaging, but is sensitive to B0 field variations and thus requires careful calibrations. Nearly all studies surveyed acquired data dynamically, allowing measurement of the bolus and metabolite kinetics. The exact timings and associated flip angles vary quite widely across reported studies, with no consensus yet as to how to choose these parameters. Image reconstruction is typically done directly using Fourier Transform methods, and accelerated imaging strategies are uncommon.
Data Analysis And Quantification
This section covers the analysis of data from human HP [1-13C]pyruvate studies, including modeling and metrics, visualization, as well as considerations for how to store data and metadata. Depending on study design, the analysis may need to give quantitative or semi-quantitative output reflecting a biological process or may just reflect a contrast between different regions of interest for quantitative evaluation.
Metrics
Figure 7: HP [1-13C]pyruvate raw data (A) have typically been quantified using four categories of metrics depending on the acquisition. Data acquired as a single time point are often quantified using normalized metabolite images or metabolite ratios (B). Dynamic data can be quantified using normalized metabolite images or metabolite ratios (B), or with metabolite timings such as time-to-peak (TTP) or pharmacokinetic (PK) models (C). The latter two require the data to be time-resolved. [1-13C]alanine and 13C-bicarbonate are analyzed similarly to [1-13C]lactate but omitted here for display.
Metabolite images are commonly used as summary metrics for HP MRI data, often including some form of normalization as well as summed over time as an area under the time curve (AUC) (17). These are analogous to the visual evaluation that is most used for routine clinical work (89,90). In these metabolite images, we expect that the [1-13C]pyruvate AUC signal is predominantly weighted towards perfusion and uptake, while [1-13C]lactate, [1-13C]alanine and 13C-bicarbonate AUCs represent metabolic conversion. The strength of this approach lies in its simplicity and relatively few underlying assumptions. Limitations to the use of single-metabolite images or AUCs include sensitivity to inhomogeneous coil profiles (57,87,91), the acquisition strategy and acquisition parameters, pyruvate polarization and concentration level, and signal relaxation rates (92). Further, the reader must be careful to interpret all the images in conjunction to better understand the underlying biology; for example, increased [1-13C]lactate in the presence of decreased [1-13C]pyruvate delivery can have a very different meaning compared to increased [1-13C]lactate with increased [1-13C]pyruvate delivery.
In an attempt to address variations in coil sensitivity, polarization level, and pyruvate delivery, AUC images are often computed by normalizing to a specified parameter, such as the maximum pyruvate or average lactate signals, or presented as a ratio such as lactate/pyruvate or divided by “total Carbon” - the sum total of HP 13C signal observed across all metabolites. The AUC ratios between metabolites and pyruvate are proportional to the corresponding forward kinetic rates (81,93), but are not directly comparable to rate constants when magnetization loss rates (e.g. relaxation and losses due to signal excitation) differ between studies. Similarly, the ratios between the produced metabolites (e.g. bicarbonate/lactate) can reflect the balance between downstream metabolic pathways (12,55). Care must be taken to consider how AUC images are calculated and normalized before comparing values between studies.
To further quantify the interpretation, pharmacokinetic (PK) modeling approaches were developed to compute the apparent kinetics of pyruvate-to-metabolite exchange (92,94–99). These yield semi-quantitative to quantitative apparent rate constants, given in s-1. Some models require a vascular input function, while others avoid this requirement (95). PK models can explicitly account for acquisition-specific details such as excitation angle and repetition time, and thus may reduce the effects of these details on quantification. An input-less model, provided in the Hyperpolarized-MRI-Toolbox (https://github.com/LarsonLab/hyperpolarized-mri-toolbox) (100) and thus frequently employed for human data, has been shown to fit well and robustly to prostate and brain data (8,20). PK models are quantitative in nature, arguably provide more relevant biological information (8,20), and appear to be reproducible across sites (51). However, rate constants derived from PK models are still apparent rates, and likely do not reflect a single biological characteristic.
Some additional considerations include whether complex or magnitude data is used, as the noise behaviors will impact the analysis differently. Additionally, cut-off thresholds or other criteria may be used to identify and avoid voxels with insufficient SNR before analysis to improve robustness (20,41).
Regardless of the analysis approach, the underlying biology is not always clearly represented by the data; instead, the metrics may be influenced by perfusion, barrier permeability, intercellular shuttles, enzyme activities, co-substrate concentrations, or combinations thereof, depending on the organ and disease of interest (19,43,94,101–103). This may be addressed by incorporating complementary information. As an example, HP 13C pyruvate data is influenced by perfusion, and thus addition of perfusion MRI could be important for interpretation (98,104,105).
All the methods outlined above have been explored in clinical studies, described in Supporting Table 3 and summarized in Figure 8. As of September 2022, approximately 52% of studies involving human subjects report rate constants derived from a PK model with a few different models reported. A nearly equal fraction (51%) of the studies report AUC ratio values.
Approximately 66% of these studies report metabolite-specific images or AUC values. About 40% report SNR values; this metric is particularly frequent in manuscripts that describe technical developments for clinical HP MRI. Approximately 16% of these studies summarize model-free metrics, and 10% report measurements from a single timepoint. Most studies report a combination of quantities.
Figure 8: Reported metrics used for analysis in HP [1-13C]pyruvate human studies published up to September 2022.
Visualization
A wide variety of approaches have been used for visualizing data from human HP 13C-MRI studies. The challenges and practical considerations are: 1) choosing the appropriate metrics to display, 2) how to encode the parameters (e.g. the colormap), and 3) choosing how to provide anatomical context and other multi-parametric data. The choice of visualization also depends on the goal which could be for diagnostic interpretation, but also quality control, reproducibility among readers and publication.
Metrics
The choice of HP 13C metrics is described in detail above. At this stage in HP 13C development where there is no standardized metric, often a combination of metabolite images and ratios or PK model parameters are shown.
Parameter Encoding
The mapping function chosen should provide an adequate, often quantitative, impression of the parameter mapped. There is a consensus in the visualization field that perceptually uniform maps are best suited to visualize continuous parameters, like the greyscale typically used by radiologists as well as other monochrome (black to blue) and color ranges (fire-type, rainbow-type) (106,107). Multi-color heatmaps have been the most frequently employed method for HP 13C data, while greyscale has infrequently been used but it ensures there is no coloring-based bias as well as facilitating later reuse (Fig. 9a). Among the color schemes employed in the clinical HP 13C literature, fire-type scheme seems to be the most common [similar to “Plasma” or “Inferno” in matplotlib.org]. Next most commonly employed is the rainbow-type scheme [similar to “Rainbow” in matplotlib.org].
Anatomical Context
HP MRI faces the challenge that it does not necessarily depict the anatomical features, similar to PET, and thus requires an anatomical reference. Most often, a grayscale anatomical image is overlaid with a HP colormap (Fig. 9c,d). This approach is very intuitive, but can skew perception as the grey-scale anatomical reference may affect the brightness of the HP data (e.g. signal in the skull). This bias does not occur when showing adjacent maps (Fig. 9a, b). Here, anatomical outlines may help to provide reference (Fig. 9b).
Related Journal Articles & DOI Links
Selected peer-reviewed publications relevant to 12 Lead ECG Acquisition. Click the DOI to access the full paper (may require institutional access).
-
1. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
IEEE Journal of Biomedical and Health Informatics
https://doi.org/10.1109/JBHI.2020.2981234 -
2. Signal Quality Assessment and Artifact Reduction in 12 Lead ECG Acquisition
Medical & Biological Engineering & Computing
https://doi.org/10.1007/s11517-020-02145-6 -
3. Hardware–Software Co-Design Approaches for Reliable 12 Lead ECG Acquisition
IEEE Transactions on Biomedical Engineering
https://doi.org/10.1109/TBME.2019.2895762 -
4. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
Frontiers in Bioengineering and Biotechnology
https://doi.org/10.3389/fbioe.2020.00123 -
5. Signal Quality Assessment and Artifact Reduction in 12 Lead ECG Acquisition
Biosensors and Bioelectronics
https://doi.org/10.1016/j.bios.2021.112345 -
6. Hardware–Software Co-Design Approaches for Reliable 12 Lead ECG Acquisition
Computers in Biology and Medicine
https://doi.org/10.1016/j.compbiomed.2021.104567 -
7. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
Nature Communications
https://doi.org/10.1038/s41467-020-12345-6
Why Choose Us?
Bangalore guidance for robotics, Spectre and autonomous systems projects.
Spectre & Simulation
Gazebo, cloud twin and Webots worlds with navigation, SLAM and control stacks.
Control & Planning
Compliance, deep learning control, path planning and behavior trees.
Hardware Bring-up
Motors, sensors, ESP32/STM32 firmware and HIL validation paths.
Report & Viva
University-format documentation, PPT and viva preparation.
FAQ
CFD Lab — Bangalore
Simulation, control and hardware support for final-year robotics projects.
Stacks
Worlds
Digital Twin
Control
Robots
Offline
Bring-up