Enquire Now
Medical Computer Vision · Clinical Diagnostics · PyTorch / TensorFlow · GPU Optimized · 2026

Voice Assistant Using Python Ai Capstone Project

Tensor Pipeline · Custom Loss Formulations · Model Quantization · Accelerated Inference — A rigorous deep learning engineering project focused on automated pathological lesion segmentation and radiological disease classification. Architected for thesis defense viva presentations, IEEE reproduction, and high-throughput production deployment.

PyTorch
Core Framework
AMP FP16
Mixed Precision
TensorRT
Quantized Serving

A Mixed-Methods Approach to Understanding User Trust after

Abstract

Despite huge gains in performance in natural language understand- ing via large language models in recent years, voice assistants still often fail to meet user expectations. In this study, we conducted a mixed-methods analysis of how voice assistant failures affect users’ trust in their voice assistants. To illustrate how users have experienced these failures, we contribute a crowdsourced dataset of 199 voice assistant failures, categorized across 12 failure sources.

voice assistant using python ai capstone project Diagram
Figure: System Model & Simulation Flow for Voice Assistant Using Python Ai Capstone Project

Relying on interview and survey data, we find that certain failures, such as those due to overcapturing users’ input, derail user trust more than others. We additionally examine how failures impact users’ willingness to rely on voice assistants for future tasks. Users often stop using their voice assistants for specific tasks that result in failures for a short period of time before resuming similar usage.

We demonstrate the importance of low stakes tasks, such as playing music, towards building trust after failures.

Ccs Concepts

• Human-centered computing →Empirical studies in HCI.

Keywords

voice assistants, trust, survey, interview, dataset

Acm Reference Format:

Amanda Baughan, Allison Mercurio, Ariel Liu, Xuezhi Wang, Jilin Chen, and Xiao Ma. 2023. A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures. In Proceedings of the 2023 CHI Confer- ence on Human Factors in Computing Systems (CHI ’23), April 23–28, 2023, Hamburg, Germany. ACM, New York, NY, USA, 16 pages. https://doi.org/10.

Introduction

Voice assistants have received a lot of attention from both indus- try and academia, especially given the recent advances in natural ∗This work was conducted as part of an internship with Google Research.

Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored.

For all other uses, contact the owner/author(s).

Chi ’23, April 23–28, 2023, Hamburg, Germany

© 2023 Copyright held by the owner/author(s).

Https://Doi.Org/10.1145/3544548.3581152

language processing (NLP). Within the past five years, advance- ments in NLP have achieved huge gains in accuracy when tested against standard datasets [8, 18, 34, 56, 58, 60, 63], with state-of- the-art accuracy in natural language processing models as high as 99% for certain tasks [8, 34]. This has led many practitioners and researchers alike to imagine a near future where voice assistants can be used in increasingly complex ways, including supporting healthcare tasks [41, 55], giving mental health advice [52, 66], and high stakes decision-making .

However, despite the increasing accuracy of NLP models and the breadth of their applications, evidences suggest that users remain reluctant and distrusting of using voice assistants [12, 29]. In the U.S., voice assistants are common in homes, with an estimated 72% of Americans having used a voice assistant . However, people primarily use these for basic tasks such as playing music, setting timers, and making shopping lists [12, 29, 35]. This is because when voice assistants fail, such as by incorrectly answering a question, it derails user trust [27, 35]. User trust is pivotal to user adoption of various technologies , and in this case, low user trust results in reluctance to try voice assistants’ novel capabilities.

As voice assistants increasingly rely on large language mod- els [24, 47], we believe the gap between the high accuracy of these models and users’ reluctance to use voice assistants for complex tasks may be explained by differences in how users and NLP practi- tioners evaluate the success of a model. Standard NLP models are often evaluated on large datasets of coherent text-based questions and answers [48, 49] or paired written dialogue . Meanwhile, in practice users’ speech may include disfluencies, such as restarts and filler words, questions not covered in training, or background noise which misconstrues speech. In the case of question answering, NLP models are evaluated based on how many questions are accurately answered on a subset of the training dataset [48, 49]. As one may expect, people can interact with voice assistants in a multitude of ways that fall outside of the scope of training data, which can lead to friction. In the eyes of users, these inaccurate responses, or voice assistant failures, can lead to frustration. For example, only five percent of users report never becoming frustrated when using voice search .

We believe that the gap between how NLP models are evaluated and how users encounter and perceive failures hinders the prac- tical applications of the advancements that voice assistants have made. Therefore, we ask, which types of voice assistant failures

Chi ’23, April 23–28, 2023, Hamburg, Germany

Baughan et al. do users currently experience, and how do these failures affect user trust? A human-centered understanding of the types of NLP failures that occur and their impact on users trust would allow technologists to prioritize and address critical failures and enable long-term adoption of voice assistants for a wider variety of use cases.

Further, while research has started to categorize types of break- downs in communication between users and NLP agents [28, 46], little work has looked into how users perceive these failures and subsequently trust and use their voice assistants. We draw from and extend past research to make the following contributions: • C1: Iterating on the existing taxonomy of NLP failures, we crowdsource a dataset of 199 failures users have experienced across 12 different sources of failure.

• C2: A qualitative and quantitative evaluation on how these different failures affect user trust, specifically along dimen- sions of ability, benevolence, and integrity.

• C3: A qualitative and quantitative analysis on how trust im- pacts intended future use. To accomplish this, we developed a mixed-methods, human- centered investigation into voice assistant failures. We first executed interviews with 12 voice assistant users to understand what types of failures they have experienced and how this affected their trust and subsequent use of their assistant. We concurrently crowdsourced a dataset of failures from voice assistant users on Amazon Mechanical Turk. Finally, we executed a survey to quantify how different types of failures impact users’ trust in their voice assistants and their willingness to use them for various tasks in the future.

We found that different types of voice assistant failures have a differential impact on trust. Our interviews and survey revealed that participants are more forgiving of failures due to spurious triggers or ambiguity of their own request. In the case of spurious triggers, the voice assistant activates due to mishearing the activation phrase when it was not said. Users forgave this more easily, as it did not hinder them from accomplishing a goal. Failures due to ambiguity occurred when there were multiple reasonable interpretations of a request, and the response was misaligned with what the user intended while still accurately answering the question. Users tended to blame themselves for these failures. However, failures due to overcapture more severely reduced users’ trust, as when the voice assistant continued listening without any additional input, users considered their use a waste of time.

We additionally find that on many occasions, users would dis- continue using their voice assistant for a specific task for a short period of time following a failure, and then resume again once trust had been rebuilt. Trust was often rebuilt by using the voice assistant for tasks they considered simple, such as playing music, or alternatively, using the voice assistant for the same general task but in a different use case. In addition to these findings, we re- lease a dataset of 199 voice assistant failures, capturing user input, voice assistant response, and the context for the failure, so that researchers may use these failures for future research on how users respond to voice assistant failures. As voice assistants continue to perform increasingly complex and high stakes tasks across various industries [17, 41, 51, 55, 66], we hope that this research will help technologists understand, prioritize, and address natural language failures to increase and maintain user trust in voice assistants.

Related Work

Prior research across many fields has examined the interaction between users and voice assistants, including human-computer in- teraction, human-centered AI, human-robotics interaction, science and technology studies (STS), computer-mediated communication (CMC), and social psychology. In addition, some work in natural lan- guage processing (NLP), especially NLP robustness, has approached technology failures in voice assistants and developed certain techni- cal solutions to address them. Here, we provide an interdisciplinary review of research relevant to voice assistant failures during user interaction across these fields. The literature review is organized as follows: 1) literature on user expectations and trust in voice assistants; 2) human-computer interaction (HCI) approaches to un- derstanding voice assistant failures and strategies for mitigation; 3) natural language processing (NLP) approaches to voice assistant failures, including disfluency and robustness.

Assistants

Researchers have long tried to understand how people interact with automated agents, especially comparing and contrasting these experiences with human-to-human communication. When talking with other humans, conversations can broadly be understood as functional (also known as transactional or task-based) or social (interactional), and many conversations include a mix of both .

Functional conversations serve towards the pursuit of a goal, and those who participate often have understood roles towards the pursuit of that goal. In contrast, social conversations have a goal of building, strengthening, or maintaining a positive relationship with one of the participants. These social conversations can help build trust, rapport, and common ground .

People generally expect to have functional conversations with voice assistants . The lack of social conversations may reduce users’ ability to build trust in their voice assistants. Indeed, past re- search has shown that users trust embodied conversational agents more when they engage in small talk , although this varies by user personality type and level of embodiment of the agent . As it stands, people report not using voice assistants for a broad range of tasks, even though they’re technically capable of doing so . Prior work has illustrated the importance of trust for continued voice assistant use [31, 35], as trust is pivotal to user adoption of voice assistants [33, 45] and willingness to broaden the scope of voice as- sistant tasks . It is especially important to support trust-building between users and voice assistants as researchers continue to imag- ine and develop new capabilities for them, including complex tasks such as supporting healthcare tasks [41, 55], giving mental health advice [52, 66], and other high stakes decision-making .

This then begs the question of how trust is built between users and voice assistants. Trust in machines is an increasingly important topic, as use of automated systems is widespread . Concretely, trust can be conceptualized as a combination of confidence in a system as well as willingness to act on its provided recommenda- tions [37, 54]. Prior researchers have examined trust in machines A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures

Chi ’23, April 23–28, 2023, Hamburg, Germany

in terms of people’s confidence in a machine’s ability to perform as expected, benevolence (well-meaning), and integrity to adhere to ethical standards Broadly, past research has evaluated how various factors such as accuracy and errors affect people’s trust in algorithms [19, 20, 67]. In the case of voice assistants, Nasirian et al. and Lee et al. studied how quality affects trust in and adoption of voice assistants, and found that information and system quality did not impact users’ trust in a voice assistant, but interaction quality did. Interaction quality was captured based on a study by Ekinci and Dawes , in which Likert scale responses were captured regarding competence, attitude, service manner, and responsiveness of the voice assistant. In addition, customizing a voice assistant’s personality to the user can lead to higher trust , while gender does not impact users’ trust in a voice assistant .

Overall, prior research demonstrates the importance of the inter- action quality and social conversations for building trust between users and voice assistants, which in turn affects users’ willingness to continue using them and broaden the scope of their tasks.

Hci Approaches To Voice Assistant Failures

However, there are occasionally unforeseen breaches of trust, as not all interactions go as smoothly as one expects. Prior work has explored the diversity of issues affecting engagement and ongo- ing use of voice assistants and has shown that when users have expectations for voice assistants that surpass its capabilities, voice assistant failures and user frustration ensues [31, 35].

This begs the question, how has prior work defined failures in voice assistants? Some work uses specific scenarios in their stud- ies. For example, Lahoual and Frejus conducted evaluation in domestic and driving situations. They identified failures due to poor voice recognition, limited understanding of a command, and connectivity. Cuadra et al. used failures in specific tasks, such as attempting to give directions to an incorrect location, send a text to the wrong person, play the wrong type of music, or adding a re- minder with an incorrect detail . Mahmood et al. simulated online shopping, in which an AI assistant with a voice component would fail by using homonyms of the requested items. For example, the ambiguous item “bow” could mean a hair bow, archery bow, or bow for gift wrapping. Salem et al. had participants control a robot’s movement, and in the faulty condition, the robot would move erratically, incorrectly responding to the users’ input. Can- dello et al. defined failure as occasions in which someone asked a question that could not be understood or was out of scope of the voice assistants’ knowledge, in which case it would divert the conversation to ask an unrelated question.

Other research aims to provide a broad categorization of voice assistant failures, drawing from theoretical frameworks of commu- nication between humans [10, 28, 46]. We reference Herbert Clark’s grounding model for human communication, which relies on four different levels to achieve mutual understanding: channel, signal, intention, and conversation . This was expanded by Paek and Horvitz , which applied these four levels to human-machine interactions and failure points. Channel level errors include when an AI fails to attend to a users’ attempt to initiate communication; signal level errors include an error in capturing user input (e.g.

due to transcription); intention level errors include mistakes in making sense of the semantic meaning of the transcribed input; and conversation level errors occur when a user has requested an unknown action to the AI (e.g. asking a weather app to schedule something). Hong et al. built on this model, specifically restrict- ing the context to NLP failures, rather than AI as a whole. Based on interviews with NLP practitioners, they renamed the categories as attention (channel), perception (signal), understanding (intention), and response (conversation). Hong et al. focused on failures that are either very common, or rare but very costly, to cover the most important and frequent failures users encounter when inter- acting with NLP-based systems. In this work, we build on their existing taxonomy of NLP failures , narrowing the use case to only voice assistant failures, and evaluating how different failures impact on user trust and future intended use.

There is currently little systematic evaluation of the impact of voice assistant failures on user trust. Salem et al. found that if a robot had faulty performance, this did not influence participants’ decisions to comply with its requests, but it did significantly af- fect their perceptions of the robot’s reliability and trustworthiness.

Mahmood et al. found that voice assistants that accepted blame and apologized for mistakes were thought to be more intelligent, likeable, and effective in recovering from failures than assistants that shifted the blame.

Sometimes after a failure, users will try to reformulate, simplify, or hyper-enunciate their commands as a way to continue using the device [31, 35, 43, 61]. If users are repeatedly unable to repair failures with voice assistant, this weakens their trust and causes them to reduce their scope of commands to simple tasks with low risk of failure [31, 35]. Lahoual and Frejus found that in some situations, voice assistant failures can erode trust to the extent that users abandon voice assistants all together. However, not all failures require self-repair. A study by Cuadra et al. found that when voice assistants make mistakes, voice assistant self-repair greatly improves people’s assessment of an intelligent voice assistant, but it can have the opposite impact if no correction is needed. Thus, understanding which types of failures undermine trust the most may also inform us when failure mitigation strategies should be activated.

Nlp Approaches To Voice Assistant Failures

The NLP community has also examined voice assistant failures from a slightly different angle, focusing on the robustness of different NLP components underlying voice assistants, such as models for tasks in natural language inference , question answering [26, 42], and speech recognition . NLP robustness can be defined as understanding how model performance changes when testing on a new dataset, which has a different distribution from the dataset the model is trained on . In practice, users’ real world interactions with voice assistants could differ from data used in development, which mimics the data distribution shift in NLP robustness research.

Such data distribution shifts are shown to lead to model fail- ures. In the case of question answering, state-of-art models per- form nearly at human-level for reading comprehension on standard benchmarks collected from Wikipedia . However, Miller et al.

found that model performance drops when the question an- swering model is evaluated on different topic domains, such as

Chi ’23, April 23–28, 2023, Hamburg, Germany

Baughan et al. Figure 1: To analyze the impact of voice assistant failures on user trust, we used a mixed-methods approach, including inter- views and a survey. As part of the materials for our survey, we crowdsourced 199 failures from 107 voice assistant users, and include this dataset as part of our contributions.

New York Times articles, Reddit posts, and Amazon product re- views. Noisy input can also harm model performances. Lee et al. showed speech recognition errors have catastrophic impact on machine comprehension. Gupta et al. created a question answering dataset Disflu-QA where humans introduce contextual disfluencies, which also lead to model performance drops.

Although these works do not directly focus on voice assistant failures, topic domain changes, speech recognition errors and dis- fluencies are all very common during user interactions with voice assistants. Such similarities motivate us to draw parallels between the NLP robustness literature and HCI perspectives of system fail- ures. By understanding how different types of failures affect trust in voice assistants overall, we can then try to pinpoint the underlying NLP components that are the root cause of the most critical failures that erode trust . Technical solutions can then be leveraged to improve the robustness of the most critical parts of the system in order to increase user trust and long-term engagement most efficiently.

Method Overview

Now that we have established the importance of understanding of how voice assistant failures impact user trust, we proceed to conduct a mixed-method study. First, to prepare for the quantitative evaluation, we reviewed existing datasets in HCI and NLP to find failures that we could use as materials for our survey. Ultimately, the existing datasets were not sufficient for our needs. Therefore, we crowdsourced a dataset of failures from voice assistant users, which we also open source as part of the contributions of this study.

Concurrently, we conducted interviews with 12 voice assistant users to understand which types of failures they have experienced, and how this affected their trust in and subsequent use of the assistant.

These interviews were designed to provide a broad understanding of the thoughts, feelings, and behaviors that users have with regard to voice assistant failures and inform the quantitative survey design.

Finally, we executed a survey to quantify how different types of failures impact user perceptions of trust in their voice assistants and their willingness to use them for various tasks in the future. To report these findings, we first describe our process of collecting the crowdsourced dataset of failures, and how we selected a subset to use in our survey. Next, we present the interviews and survey, first describing our data collection and analysis, and then presenting the results concurrently.

Assistant Failures

The first goal in our investigation was to determine which types of failures users experience when using voice assistants. We first evaluated existing datasets for fit and breadth of failures. We deter- mined they were not sufficient for our purposes, so we proceeded to crowdsource a dataset of failures, adapting a taxonomy from Hong et al. to guide our collection. Finally, we cleaned and open-sourced this dataset as a contribution of our work.

A Review Of Existing Hci And Nlp Datasets

We first explored benchmark datasets in NLP, which contain a large number of either questions and answers , or conversational dialogue [25, 57, 68]. We found that existing NLP datasets do not cover the wide breadth of possible conversational failure cases due to their emphasis on correct data for training. Additionally, their focus on specific task performance, such as answering questions or dialogue generation, is more narrow than the variety of use cases for voice assistants. As training data relies on accurate task completion, A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures

Missed Trigger

Users say something to trigger the voice assistant, but it fails to respond.

Spurious Trigger

Users do not say something to trigger the voice assistant, but it activates anyways.

Delayed Trigger

Similar to system latency, the users say something to trigger the voice assistant, but it replies too late to be useful.

Word, Or Evidence Of Background

noise and cross-talk.

Noisy Channel

User input is incorrectly captured due to background noise.

Overcapture

The voice assistant captures more input than intended by either beginning to capture input too early or ending too late, and acting on external data not relevant to the users’ request.

Truncation

System does not fully capture users’ speech, by either beginning to capture input too late or ending too early.

Transcription

System generates a transcription error, often in the form of similar sounding words.

Plausible But Not Correct For The

intention of the input.

Ambiguity

There may be several interpretations of the users’ in- tent, and the system responds in a way that is plausibly accurate but not correct for the users’ intent.

Misunderstanding

The system maps the users’ input to an incorrect action, perhaps with some correct inference on the users’ intent, but not fully accurate.

No Understanding

The system fails to map the user’s input to any known action or response.

Action

If the system listens to the full request, but then turns off before giving any type of answer or taking action.

Correct Action

The system gives information that is incorrect. Table 1: Qualitative codebook and description of the various failures that were collected. We checked each failure for failure type sequentially, starting by checking if it could be an attention failure and progressing through the types until we found one that fit. From failure type, we then assessed which failure source applied.

these datasets did not contain failures. While testing these models produces a small percentage of errors (roughly 10%), the types of failures could only fall in the response and understanding categories, as attention and perception failures are excluded from the context of training these types of models. This limited their usefulness for our purpose of understanding voice assistant failures that occur in use and their impact on user trust.

In addition to these benchmark datasets, we investigated datasets that incorporated spoken word speech patterns, such as the Spoken SQuAD dataset and Disflu-QA dataset , as well as human- agent interaction datasets, such as the ACE dataset , the Niki and Julie corpus , and a video dataset of voice assistant failures .

In these cases, we found that the datasets were still restricted to only failures at the understanding and response level [26, 32] or the context for the failures was very specific and did not necessarily capture the breadth of possible failures users experience [1, 2].

Cuadra et al. ’s video dataset was the closest available fit for our needs, but we still found the use case of in-lab question-answering too narrow for our purposes. Therefore, we decided to crowdsource a dataset of voice assistant failures from users, and use these failures when conducting our quantitative survey on user trust.

4.2.1

Procedure. Crowd workers were asked to submit three fail- ures they had experienced with a voice assistant. They were asked about three specific types of failures out of a taxonomy of 12, which were randomly chosen and displayed in equal measure across all workers. The taxonomy of failures that we used to ask about spe- cific types of failures was adapted from previous work by Hong et al. , and identifies failures due to attention, perception, under- standing, and response, as shown in Table 1. Each question began by asking users if they could recall a time when their voice assistant had failed, based on the definitions in our taxonomy. For example, to capture missed trigger failures we asked “Has there ever been a time when you intended to activate a voice assistant, but it did not respond?” If so, we asked these workers to include 1. what they had said to the voice assistant, 2. how the voice assistant responded, 3.

the context for the failure, including what happened in the envi- ronment, and 4. the frequency at which the failure occurred from 1 (rarely when I use it) to 5 (every time I use it). These were all presented as text entry boxes except for the frequency question, which was multiple choice. Crowd workers were additionally asked to optionally share an additional failure that they had not had the

Chi ’23, April 23–28, 2023, Hamburg, Germany

Baughan et al. chance to share already. This was included to capture failures that did not fit any of the three the categories they were presented with, and we then categorized these failures according to our taxonomy.

Once we received these failures, we anonymized the type of voice assistant in the submitted examples, replacing activation words with “Voice Assistant” for consistency. We then edited grammatical and spelling errors for clarity. We also removed failures if they were not on-task, unclear, or exact repeats of other submitted failures.

Finally, we noticed that some of the categories the users submitted the failures under were incorrect, so we re-categorized the failures according to the codebook we developed as outlined in Table 1. Two raters iteratively coded 101 submitted failures, with a final coding session achieving an interrater agreement of 70%. One researcher then went back and coded the entire dataset in its entirety. In total, our finalized dataset contains 199 failures across 12 categories, submitted by 107 unique crowd workers.

4.2.2

Crowd Worker Characteristics. We used Amazon Mechanical Turk to recruit the crowd workers. In total, 107 crowd workers contributed to our dataset. We required workers to have the follow- ing qualifications: a HIT Approval Rate over 98%, over 1000 HITs approved, AMT Masters, from the United States, over the age of 18, and voice assistant users on at least a weekly basis. The plurality of users were in the age range of 35-44 (𝑛= 46), followed by 25-34 (𝑛= 32), and 45-54 (𝑛= 16), with the rest falling in 55-64 (𝑛= 8), 18-24 (𝑛= 1), and 1 preferring not to answer. Fifty-eight crowd workers were men, 44 were women, 1 preferred not to answer, and 1 identified as both a man and a woman. They used commercial voice assistants such as Amazon Alexa (𝑛= 59), Google Assistant (𝑛= 62), and Apple’s Siri (𝑛= 40), with many using some combination of the three (𝑛= 47). 91 crowd workers were native English speakers, and 13 were not. The plurality identified as White (𝑛= 58), and 39 identified as Asian. Three crowd workers did not provide any demographic information. The task took 15-20 minutes to complete on average, and they received $5.00 USD compensation.

4.2.3

Final Dataset. In total, our finalized dataset contained 199 failures from 107 users across 12 different types of failures according to the taxonomy based on Hong et al. , as updated in Table 1.

The failures we received most often were due to misunderstanding (𝑛= 38), missed trigger (𝑛= 25), and noisy channel (𝑛= 22). Users least often submitted failures for truncation (𝑛= 7), overcapture (𝑛= 7), and delayed triggers (𝑛= 8). Most crowd workers submitted failures saying that they happened “rarely when I use it” (𝑛= 87) or “sometimes when I use it” (𝑛= 84). Example failures across the 12 categories can be found in Table 2.

On average, the highest frequency of failures occurred for no understanding (𝑚= 2.15, sometimes when I use it, 𝑠𝑑= 0.67) and action execution: incorrect (𝑚= 2.00, sometimes when I use it, 𝑠𝑑= 0.88). The rest of the failure sources had an average reported frequency between 1.0 (rarely when I use it) and 2.0 (sometimes when I use it). The lowest frequency failures were due to delayed triggers (𝑚= 1.25, 𝑠𝑑= 0.46) and ambiguity (𝑚= 1.39, 𝑠𝑑= 0.78).

We then used 60 of the failures from our dataset in our survey to quantify the impact of different failures on user trust. This is outlined in more detail in the following section. This dataset has been open sourced1 for researchers to use to answer future research questions related to voice assistant failures in the future.

Interview And Survey Methods

Once we had gathered and categorized our dataset of voice assis- tant failures, we were ready to answer our research question: how do voice assistant failures impact user trust? To do so, we first conducted exploratory interviews with 12 people to gather their thoughts, feelings, and behaviors after experiencing voice assistant failures. We used these findings and the failures collected in the dataset to then design and execute a survey. This quantified how various voice assistant failures impact users’ trust, as measured by their perceptions of the voice assistant’s ability, benevolence, integrity, and their willingness to use it for future tasks. Here, we describe the methods for both the interviews and survey, and we follow this by jointly presenting the results from both studies.

5.1.1

Interview Procedure. Interviews began with questions about why the participants chose to start using voice assistants and what types of questions they frequently would ask of them. We asked for common times and places they would use their voice assistants to understand their general experience with voice assistants.

Once these were established, we asked participants to tell us about a time they were using their voice assistant and it made a mistake, in as much detail as they could recall. We asked what they had been trying to do and why, if others were present, and if anything else was happening in their environment. We probed for users’ feelings once the failure occurred, and their perceptions about the voice assistant’s ability to understand them and give them accurate information. We asked participants what they did in the moment to respond to the failure. Finally, we asked questions about their use of the voice assistant in the aftermath, including how much they trusted it and if they changed any of their behaviors to mitigate future failures. All interviews were conducted remotely.

5.1.2

Interview Participants. During recruitment, we asked partic- ipants to submit their demographic information, how frequently they used voice assistants and on what types of devices. We addition- ally required participants to write a short (1-3 sentence) summary of a time they encountered a failure while using their voice assis- tant. We selected participants based on demographic distribution and the level of detail they included regarding the failure.

All of our 12 participants lived in the United States. They used voice assistants at least 1-3 times a week (𝑛= 2), with the majority reporting using a voice assistant every day (𝑛= 8), and the rest (𝑛= 2) using it 4-6 times a week. The majority of participants used a voice assistant on their mobile device (𝑛= 11), and five of these participants also used a voice assistant smart home device.

One participant only used a voice assistant smart home device. Participants reported using common commercial voice assistants such as Amazon Alexa (𝑛= 2), Google Assistant (𝑛= 7), and Apple’s Siri (𝑛= 8). Participants’ ages ranged from 18 to 50, with the plurality (𝑛= 5) in the age range of 18-23. 3 of our participants were 41-50, 2 were 31-40, and 2 were 24-30. Six of our participants 1https://www.kaggle.com/datasets/googleai/voice-assistant-failures A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures

I Tell Her To Set A Timer For Ten Minutes, I Was

alone and no one present at the moment.

Voice Assistant, Set A Timer For 10

minutes.

And I Was Calling My Coworker Sherry. The Voice

assistant mistakenly got turned on.

Delayed Trigger

It happened while I was driving a car.

Voice Assistant, Show Me The Route

to the national park.

Voice And Try Several Times To Be Heard By My

phone even though it was inches from my face.

Weather?

[It didn’t realize that my request had ended ] and

Overcapture

I was telling it to turn off the lights. I was the only one there. Voice Assistant, turn off the lights.

I Asked The Voice Assistant To Calculate A Math

question, but it cut me off.

I Asked For The Weather Conditions In The City I

live in. No others were present except for me.

I Was At Home, Alone, Watching Ufc And Asked

how old a fighter was.

I Asked It To Play The Theme From Halloween. I

was sitting with my mother.

Voice Assistant, Play The Theme

song to the movie Halloween. [It plays a scary sounds soundtrack instead of the

I Was Trying To Run A Routine To Wake Up My

kids. Voice Assistant, wake up the twins. "Sorry, I don’t know that." [However, I’ve set up a

Aters, And It Kept Spinning Its Light Over And

over.

Chi Come Out In Theatres?

[Pauses for a really long time, then turns its lights

I Was At Home, In My Living Room, Alone. I Was

trying to find out how long Taco Bell was open.

Voice Assistant, When Does The

Taco Bell on Glenwood close.

Taco Bell, I Realized It Closed At 11:30Pm.]

Table 2: Table of voice assistant failures users submitted, including the context for the failure, what the user said, and what the voice assistant said. identified as women, five participants identified as men, and one participant identified as non-binary. Three participants identified as Asian, three identified as White, three identified as Black or African American, two identified as Hispanic, Latino, or Spanish origin, and one identified as both White and Black or African American. All of our participants spoke English as a native language. Participants were compensated with a $50 gift card and each interview lasted roughly 30 minutes.

5.1.3

Interview Analysis. Interviews were transcribed in their en- tirety by an automated transcription service and analyzed via a deductive and inductive process . We used deductive analysis to assess which types of failures these participants experienced. To ground our deductive analysis, we used the same codebook as we did for the dataset, as demonstrated in Table 1. We first identified in- stances in which participants were discussing distinct failures, and then applied our codebook to these instances. We used cues such as what was happening in their environment, and when appropriate, users’ own perceptions of why the failure occurred. We began by first identifying if failures belonged in which of the four failure types: attention, perception, understanding, or response. First, to determine if there was an attention failure, we investigated if there was evidence that the voice assistant accurately responded to an ac- tivation phrase, as indicated by visual or auditory cues, or otherwise by the participant’s narrative. Second, we evaluated if there was an error in perception, based on the participants’ assumption of if the voice assistant accurately parsed the input from the participant, our own assessment from their narrative, or other audio/visual cues.

Next, assuming that the input was correctly parsed, we sought to understand if the voice assistant accurately understood the seman- tic meaning of the input (understanding failures), using the same process. Finally, assuming all else had been correctly understood, we assigned response failures, indicating that the voice assistant either did not take action or took the incorrect action in response to an accurately understood command. Once a failure type was determined, we then further specified the failure sources as noted in Table 1. We resolved disagreements both asynchronously and in meetings, through discussion and comparison, over the course of several weeks.

While conducting this analysis, we also inductively identified themes related to these failures’ impact on future tasks and recovery strategies. To conduct this analysis, two researchers reviewed the twelve transcripts in their entirety, and one additional researcher reviewed five of these transcripts to further broaden and diversify themes. These researchers met over the course of several weeks to compare notes and themes, ultimately creating four different themes through inductive analysis. Of these themes, we report two

Chi ’23, April 23–28, 2023, Hamburg, Germany

Baughan et al. due to their novelty, specifically as related to future task orientation and recovery strategies.

Survey Methods

To quantify our findings from interviews, we developed a survey to explore users’ trust in voice assistants following each of the twelve different types of failures from our taxonomy, as well as their willingness to use voice assistants for a variety of tasks in the aftermath.

5.2.1

Procedure. The survey contained a screener, the core task, and a demographic section. We required participants be over 18 years old, use their voice assistant in English, and use a voice as- sistant with some regularity to participate. If participants passed the screener, they were required to review and agree to a digital consent form to continue.

The core task stated, “The following questions will ask you what you think about the abilities of a voice assistant, given that the voice assistant has made a mistake. Imagine these mistakes have been made by a voice assistant you have used before. Please consider each scenario as independent of any that come before or follow it. This survey will take approximately 20 minutes.” Participants were then presented with 12 different failure scenarios, and they were asked to rate their trust in two separate questions.

The first question measured trust in voice assistants as a con- fidence score across three dimensions: ability, benevolence, and integrity. These were selected because prior work on trust has determined these elements explain a large portion of trustworthi- ness [36, 40]. In the context of voice assistants, ability refers to how capable the voice assistant is of accurately responding to users’ input. Benevolence refers to how well-meaning the product is. And finally, integrity represents that it will adhere to ethical standards.

We asked participants to rate their confidence in voice assistants’ ability, benevolence, and integrity, as a percentage on a scale of 0-100, with steps of 10, to replicate how prior work has conceptu- alized trust . This was captured in response to the following

Statements:

• (Ability) This voice assistant is generally capable of accurately responding to commands. • (Benevolence) This voice assistant is designed to satisfy the commands its users give.

• (Integrity) This voice assistant will not cause harm to its users. The second question evaluated users’ trust in the voice assistant to complete tasks that required high, medium, and low trust. To select these tasks, we ran a small survey on Mechanical Turk with 88 voice assistant users. We presented 12 different questions, which first gave an example voice assistant failure (one for each failure source), and then asked “How much would you trust this voice assis- tant to do the following tasks:” give a weather forecast, play music, edit a shopping list, text a coworker, and send money. Users could choose that they would trust it completely, trust it somewhat, or not trust it at all.

There was not a significant difference in how much people trusted the voice assistant to play music compared to forecast the weather (𝑍= 2.06, 𝑝= 0.078). There was also not a significant difference in how much people trusted the voice assistant to edit a shopping cart or text a coworker (𝑍= 1.39, 𝑝= 0.21) as determined by pairwise comparisons, using 𝑍-tests, corrected with Holm’s se- quential Bonferroni procedure on an ANOVA of an ordinal mixed model. We found that there were significant differences between playing music, texting a coworker, and transferring money, with users having the most trust in the voice assistant playing music after a failure, less trust in texting a coworker, and still less in transferring money. Therefore, we selected playing music, texting a coworker, and transferring money to represent low, medium, and high levels of trust required. Therefore, after asking about ability, benevolence, and integrity, we asked participants how much they trusted their voice assistants to execute the following tasks: play music, text a coworker, and transfer money. These questions were displayed on a linear scale of 1 (“I do not trust it at all”) to 5 (“I completely trust it”), with steps of 1.

We completed the survey with an open-ended, optional question for participants to share anything else they would like to add. The survey concluded with demographic questions regarding gender, race, ethnicity, whether they were native English speakers, what type of voice assistants they used, and their general trust tendency as control variables. General trust tendency was measured based on responses to the following: “Generally speaking, would you say that most people can be trusted, or that you need to be very care- ful in dealing with people?” The options ranged from 1 (need to be very careful in dealing with people) to 5 (most people can be trusted). The questionnaire used for the survey has been submitted as supplementary materials.

5.2.2

Materials from our Dataset. To present each of the twelve failure sources in our survey, we drew from the dataset we had created. We selected five failures from each of the twelve categories.

We required that these failures had been coded by two of the team members who were in agreement (see dataset examples in Table 2). We used random selection to determine which of the five possible failures was presented to each user for each failure source. These are denoted in the dataset “Survey” column.

5.2.3

Participants. We recruited participants from Amazon Me- chanical Turk. We first ran a small pilot (𝑛= 27) in which we determined that participants completed the survey in roughly 20 minutes on average, and we set the compensation rate at $9 USD.

After removing participants who did not pass the attention check or straight-lined, meaning they responded to every question with the same answer, we had a total of 268 participants. These participants were required to have the following qualifications: AMT Masters, with over 1000 HITs already approved, over 18 years old, live in the United States, an approval rate greater than 97%, and they must not have participated in any of our prior studies.

The plurality of our participants were in the age range of 35-44 (𝑛= 106), followed by 25-34 (𝑛= 68), 45-54 (𝑛= 52), 55-64 (𝑛= 33), with 2-4 participants in each of the age brackets of 18-24, 65-74, and 75+. 134 of our participants identified as men, 132 identified as women, and 2 identified as non-binary genders. The majority of our participants were White (𝑛= 210), 21 participants were Black, and 15 were Asian. The rest of our participants identified as mixed race or preferred not to answer.

Ai Adaptive Learning

This project focuses on ai adaptive learning using modern AI and machine learning techniques. The content below is adapted from research literature and practical implementation notes.

We propose a novel high-performance and interpretable canon-

addition, unlike tree learning, DNNs enable gradient descent- ical deep tabular data learning architecture, TabNet. TabNet based end-to-end learning for tabular data which can have a uses sequential attention to choose which features to reason multitude of benefits: (i) efficiently encoding multiple data from at each decision step, enabling interpretability and more types like images along with tabular data; (ii) alleviating the efficient learning as the learning capacity is used for the most need for feature engineering, which is currently a key aspect

salient features. We demonstrate that TabNet outperforms in tree-based tabular data learning methods; (iii) learning other variants on a wide range of non-performance-saturated from streaming data and perhaps most importantly (iv) end- tabular datasets and yields interpretable feature attributions to-end models allow representation learning which enables plus insights into its global behavior. Finally, we demonstrate many valuable application scenarios including data-efficient

self-supervised learning for tabular data, significantly improv- domain adaptation (Goodfellow, Bengio, and Courville 2016), ing performance when unlabeled data is abundant. generative modeling (Radford, Metz, and Chintala 2015) and

Introduction We propose a new canonical DNN architecture for tabular

Deep neural networks (DNNs) have shown notable success data, TabNet. The main contributions are summarized as: efficiently encode the raw data into meaningful representa- enabling flexible integration into end-to-end learning. tions, fuel the rapid progress. One data type that has yet to 2. TabNet uses sequential attention to choose which fea- see such success with a canonical architecture is tabular data. tures to reason from at each decision step, enabling in-

Despite being the most common data type in real-world AI terpretability and better learning as the learning capacity (as it is comprised of any categorical and numerical features), is used for the most salient features (see Fig. 1). This under-explored, with variants of ensemble decision trees for each input, and unlike other instance-wise feature se- Why? First, because DT-based approaches have certain bene- and van der Schaar 2019), TabNet employs a single deep

fits: (i) they are representionally efficient for decision mani- learning architecture for feature selection and reasoning. folds with approximately hyperplane boundaries which are 3. Above design choices lead to two valuable properties: (i) common in tabular data; and (ii) they are highly interpretable TabNet outperforms or is on par with other tabular learn- in their basic form (e.g. by tracking decision nodes) and there ing models on various datasets for classification and re-

are popular post-hoc explainability methods for their ensem- gression problems from different domains; and (ii) TabNet ble form, e.g. (Lundberg, Erion, and Lee 2018) – this is an enables two kinds of interpretability: local interpretability important concern in many real-world applications; (iii) they that visualizes the importance of features and how they are fast to train. Second, because previously-proposed DNN are combined, and global interpretability which quantifies

architectures are not well-suited for tabular data: e.g. stacked the contribution of each feature to the trained model. convolutional layers or multi-layer perceptrons (MLPs) are 4. Finally, for the first time for tabular data, we show signif- vastly overparametrized – the lack of appropriate inductive icant performance improvements by using unsupervised bias often causes them to fail to find optimal solutions for tab- pre-training to predict masked features (see Fig. 2).

ular decision manifolds (Goodfellow, Bengio, and Courville

Why is deep learning worth exploring for tabular data?

One obvious motivation is expected performance improve- Feature selection: Feature selection broadly refers to judi- Copyright © 2021, Association for the Advancement of Artificial ciously picking a subset of features based on their useful-

Professional occupation related Investment related

Feedback from Feedback to

Feature selection Input processing Feature selection Input processing

previous step next step … …

Predicted output (whether the income level >$50k)

selection enables interpretability and better learning as the capacity is used for the most salient features. TabNet employs multiple decision blocks that focus on processing a subset of input features for reasoning. Two decision blocks shown as examples process features that are related to professional occupation and investments, respectively, in order to predict the income level.

Unsupervised pre-training Supervised fine-tuning

Age Cap. gain Education Occupation Gender Relationship Age Cap. gain Education Occupation Gender Relationship 5 2000 ? Exec-managerial F Wife 6 2000 Bachelors Exec-managerial M Husband 1 0 ? Farming-fishing M ? 2 0 High-school Farming-fishing M Unmarried

? 50 Doctorate Prof-specialty M Husband 4 50 Doctorate Prof-specialty M Husband 2 ? ? Handlers-cleaners F Wife 2 0 High-school Handlers-cleaners F Wife 5 3000 Bachelors ? ? Husband 5 3000 Bachelors Exec-managerial M Husband

3 0 Bachelors ? F ? 3 100 Bachelors Prof-specialty F Wife ? 0 High-school Armed-Forces ? Husband 2 0 High-school Armed-Forces M Husband

TabNet decoder Decision making

Age Cap. gain Education Occupation Gender Relationship Income > $50k

3 M False

level can be guessed from the occupation, or the gender can be guessed from the relationship. Unsupervised representation learning by masked self-supervised learning results in an improved encoder model for the supervised learning task.

ward selection and Lasso regularization (Guyon and Elisseeff performance with compact representations. 2003) attribute feature importance based on the entire training Tree-based learning: DTs are commonly-used for tabular data, and are referred as global methods. Instance-wise fea- data learning. Their prominent strength is efficient picking ture selection refers to picking features individually for each of global features with the most statistical information gain

to maximize the mutual information between the selected mance of standard DTs, one common approach is ensembling features and the response variable, and in (Yoon, Jordon, and to reduce variance. Among ensembling methods, random van der Schaar 2019) by using an actor-critic framework to forests (Ho 1998) use random subsets of data with randomly mimic a baseline while optimizing the selection. Unlike these, selected features to grow many trees. XGBoost (Chen and

sity in end-to-end learning – a single model jointly performs recent ensemble DT approaches that dominate most of the feature selection and output mapping, resulting in superior recent data science competitions. Our experimental results

!# + Softmax !" < % !" > % !# > & !# > &

ReLU ReLU &

$" !" − $" % −1 −$" !" + $" % −1 −1 $# !# − $# & % −1 −$# !# + $# & !"

FC FC

W: [$" , - $" , 0, 0] W: [0, 0, $# , - $# ] !" < % b: [-a $" , a $" , -1, -1] b: [-1, -1, -d $# , d $# ] !# < & !" > % !# < & [!" ] [!# ]

M: [1, 0] M: [0, 1]

(right). Relevant features are selected by using multiplicative sparse masks on inputs. The selected features are linearly transformed, and after a bias addition (to represent boundaries) ReLU performs region selection by zeroing the regions. Aggregation of multiple regions is based on addition. As C and C get larger, the decision boundary gets sharper.

for various datasets show that tree-based models can be out- constructs a sequential multi-step architecture, where each performed when the representation capacity is improved with step contributes to a portion of the decision based on the deep learning while retaining their feature selecting property. selected features; (iii) improves the learning capacity via non- Integration of DNNs into DTs: Representing DTs with linear processing of the selected features; and (iv) mimics

DNN building blocks as in (Humbird, Peterson, and McClar- ensembling via higher dimensions and more steps. ren 2018) yields redundancy in representation and ineffi- cient learning. Soft (neural) DTs (Wang, Aggarwal, and Liu Fig. 4 shows the TabNet architecture for encoding tabu- functions, instead of non-differentiable axis-aligned splits. mapping of categorical features with trainable embeddings.

However, losing automatic feature selection often degrades We do not consider any global feature normalization, but performance. In (Yang, Morillo, and Hospedales 2018), a soft merely apply batch normalization (BN). We pass the same D- binning function is proposed to simulate DTs in DNNs, by dimensional features f ∈ <B×D to each decision step, where 2019) proposes a DNN architecture by explicitly leveraging multi-step processing with Nsteps decision steps. The ith

expressive feature combinations, however, learning is based step inputs the processed information from the (i − 1)th step on transferring knowledge from gradient-boosted DT. (Tanno to decide which features to use and outputs the processed ing from primitive blocks while representation learning into sion. The idea of top-down attention in the sequential form edges, routing functions and leaf nodes. TabNet differs from is inspired by its applications in processing visual and text

these as it embeds soft feature selection with controllable data (Hudson and Manning 2018) and reinforcement learn- Self-supervised learning: Unsupervised representation relevant information in high dimensional input. learning improves supervised learning especially in small Feature selection: We employ a learnable mask M[i] ∈ has shown significant advances – driven by the judicious capacity of a decision step is not wasted on irrelevant

choice of the unsupervised learning objective (masked input ones, and thus the model becomes more parameter effi- prediction) and attention-based deep learning. cient. The masking is multiplicative, M[i] · f . We use an attentive transformer (see Fig. 4) to obtain the masks us- TabNet for Tabular Learning ing the processed features from the preceding step, a[i − 1]:

M[i] = sparsemax(P[i − 1] · hi (a[i − 1])). Sparsemax nor-

DTs are successful for learning from real-world tabular malization (Martins and Astudillo 2016) encourages sparsity datasets. With a specific design, conventional DNN building by mapping the Euclidean projection onto the probabilistic blocks can be used to implement DT-like output manifold, simplex, which is observed to be superior in performance and e.g. see Fig. 3). In such a design, individual feature selec- aligned with the goal of sparse feature selection for explain-

tion is key to obtain decision boundaries in hyperplane form, PD which can be generalized to a linear combination of features ability. Note that j=1 M[i]b,j = 1. hi is a trainable func- where coefficients determine the proportion of each feature. tion, shown in Fig. 4 using a FC layer, followed by BN. P[i] TabNet is based on such functionality and it outperforms DTs is the prior scale term, denoting how much a particular feature

Qi while reaping their benefits by careful design which: (i) uses has been used previously: P[i] = j=1 (γ − M[j]), where γ sparse instance-wise feature selection learned from data; (ii) is a relaxation parameter – when γ = 1, a feature is enforced

+ Softmax

Feature Feature …

transformer transformer

x Nsteps Features

+ Softmax

Feature …

transformer transformer Feature Feature Feature Feature transformer

Encoded representation

transformer transformer Attentive transformer … Mask transformer …

Step 2 Decision step dependent

transformer transformer

BN Feature Feature

FC BN transformer transformer

+ 0.5 0.5 0.5 Agg. Agg. Features Features FC FC + +

Reconstructed + … Feature attributes + … features

(a) TabNet encoder architecture (b) TabNet decoder architecture Feature transformer Feature Attentive transformer Shared across decision steps Decision step dependent transformer GLU

Decision step dependent Prior scales

+ 0.5 0.5 0.5

0.5 0.5 0.5

+ Attentive transformer (c) (d)

Prior scales

divides the processed representation to be used by the attentive transformer of the subsequent step as well as for the overall Attentive BN FC

output. For each step, the feature selection mask provides interpretable information about the model’s functionality, and the +

masks can be aggregated to obtain global feature transformer important attribution. (b) TabNet decoder, composed of a feature transformer block at each step. (c) A feature transformer block example – 4-layer network is shown, where 2 are shared across all decision

Prior scales

steps and 2 are decision step-dependent. Each layer is composed of a fully-connected (FC) layer, BN and GLU nonlinearity. (d) +

An attentive transformer block example – a single layer mapping is modulated with a prior scale information which aggregates Sparsemax

how much each feature has been used before the current decision step. sparsemax (Martins and Astudillo 2016) is used for BN FC

normalization of the coefficients, resulting in sparse selection of the salient features. +

to be used only at one decision step and as γ increases, more propose the aggregate.feature importance mask, Magg−b,j = flexibility is provided to use a feature at multiple decision PNsteps ηb [i]Mb,j [i]

PD PNsteps

ηb [i]Mb,j [i].2 i=1 i=1 steps. P is initialized as all ones, 1B×D , without any prior j=1

on the masked features. If some features are unused (as in self- Tabular self-supervised learning: We propose a decoder supervised learning), corresponding P entries are made 0 architecture to reconstruct tabular features from the Tab- to help model’s learning. To further control the sparsity of the Net encoded representations. The decoder is composed of selected features, we propose sparsity regularization in the feature transformers, followed by FC layers at each deci-

form of entropy (Grandvalet and Bengio 2004), Lsparse = sion step. The outputs are summed to obtain the recon-

PNsteps PB PD −Mb,j [i] log(Mb,j [i]+)

i=1 b=1 j=1 Nsteps ·B , where  is a structed features. We propose the task of prediction of miss- small number for numerical stability. We add the sparsity reg- ing feature columns from the others. Consider a binary mask ularization to the overall loss, with a coefficient λsparse . Spar- S ∈ {0, 1}B×D . The TabNet encoder inputs (1 − S) · f̂ sity provides a favorable inductive bias for datasets where and the decoder outputs the reconstructed features, S · f̂ . We

most features are redundant. initialize P = (1 − S) in encoder so that the model em- Feature processing: We process the filtered features using phasizes merely on the known features, and the decoder’s last a feature transformer (see Fig. 4) and then split for the FC layer is multiplied with S to output the unknown features. decision step output and information for the subsequent We consider the reconstruction loss in self-supervised phase:

step, [d[i], a[i]] = fi (M[i] · f ), where d[i] ∈ <B×Nd and 2

PB PD (f̂b,j −fb,j )·Sb,j

a[i] ∈ <B×Na . For parameter-efficient and robust learning b=1 j=1

√ PB PB 2

. Normalization b=1 (fb,j −1/B b=1 fb,j ) with high capacity, a feature transformer should comprise layers that are shared across all decision steps (as the same with the population standard deviation of the ground truth features are input across different decision steps), as well as is beneficial, as the features may have different ranges. We decision step-dependent layers. Fig. 4 shows the implementa- sample Sb,j independently from a Bernoulli distribution with

tion as concatenation of two shared layers and two decision parameter ps , at each iteration. step-dependent layers. Each FC layer is followed by BN and eventually connected to a normalized residual √ connection We study TabNet in wide range of problems, that contain with normalization. Normalization with 0.5 helps to sta- regression or classification tasks, particularly with published bilize learning by ensuring that the variance throughout the benchmarks. For all datasets, categorical inputs are mapped

For faster training, we use large batch sizes with BN. Thus, bedding and numerical columns are input without and pre- except the one applied to the input features, we use ghost BN processing.4 We use standard classification (softmax cross (Hoffer, Hubara, and Soudry 2017) form, using a virtual batch entropy) and regression (mean squared error) loss functions size BV and momentum mB . For the input features, we ob- and we train until convergence. Hyperparameters of the Tab-

serve the benefit of low-variance averaging and hence avoid Net models are optimized on a validation set and listed in ghost BN. Finally, inspired by decision-tree like aggregation Appendix. TabNet performance is not very sensitive to most as in Fig. 3, we construct the overall decision embedding hyperparameters as shown with ablation studies in Appendix. as dout = i=1 PNsteps ReLU(d[i]). We apply a linear mapping In Appendix, we also present ablation studies on various de-

Wfinal dout to get the output mapping.1 sign and guidelines on selection of the key hyperparameters. Interpretability: TabNet’s feature selection masks can shed For all experiments we cite, we use the same training, val- light on the selected features at each step. If Mb,j [i] = 0, idation and testing data split with the original work. Adam optimization algorithm (Kingma and Ba 2014) and Glorot then j th feature of the bth sample should have no contribution uniform initialization are used for training of all models.5

to the decision. If fi were a linear function, the coefficient

Mb,j [i] would correspond to the feature importance of fb,j . Instance-wise feature selection

Although each decision step employs non-linear processing, their outputs are combined later in a linear way. We aim Selection of the salient features is crucial for high perfor- to quantify an aggregate feature importance in addition to mance, especially for small datasets. We consider 6 tabular requires a coefficient that can weigh the relative importance samples). The datasets are constructed in such a way that of each step in the decision. We simply propose ηb [i] = only a subset of the features determine the output. For Syn1-

PNd Syn3, salient features are same for all instances (e.g., the

c=1 ReLU(db,c [i]) to denote the aggregate decision con- tribution at ith decision step for the bth sample. Intuitively, if 2

Normalization is used to ensure D

P j=1 Magg−b,j = 1. db,c [i] < 0, then all features at ith decision step should have 3

0 contribution to the overall decision. As its value increases, prove the performance, but interpretation of individual dimensions

it plays a higher role in the overall linear combination. Scal- may become challenging. ing the decision mask at each decision step with ηb [i], we Specially-designed feature engineering, e.g. logarithmic trans- formation of variables highly-skewed distributions, may further

For discrete outputs, we additionally employ softmax during

training (and argmax during inference). An open-source implementation will be released.

Global: using only globally-salient features, Tree Ensembles (Geurts, Ernst, and Wehenkel 2006), Lasso-regularized model, L2X

Syn Syn Syn Syn Syn Syn

No selection .5 ± .0 .7 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .6 ± .0 Tree .5 ± .1 .8 ± .0 .8 ± .0 .6 ± .0 .7 ± .0 .7 ± .0 Lasso-regularized .4 ± .0 .5 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .7 ± .0

INVASE .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0

Global .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0 TabNet .6 ± .0 .8 ± .0 .8 ± .0 .7 ± .0 .7 ± .0 .8 ± .0

output of Syn depends on features X -X ), and global fea- Table 3: Performance for Poker Hand induction dataset. ture selection, as if the salient features were known, would give high performance. For Syn4-Syn6, salient features are Model Test accuracy (%) instance dependent (e.g., for Syn4, the output depends on ei- DT 50.0 ther X -X or X -X depending on the value of X ), which MLP 50.0

makes global feature selection suboptimal. Table 1 shows that Deep neural DT 65.1

TabNet outperforms others (Tree Ensembles (Geurts, Ernst, XGBoost 71.1

and Wehenkel 2006), LASSO regularization, L2X (Chen LightGBM 70.0 van der Schaar 2019). For Syn1-Syn3, TabNet performance TabNet 99.2 is close to global feature selection - it can figure out what Rule-based 100.0 features are globally important. For Syn4-Syn6, eliminating instance-wise redundant features, TabNet improves global feature selection. All other methods utilize a predictive model Poker Hand (Dua and Graff 2017): The task is classifica-

with 43k parameters, and the total number of parameters is tion of the poker hand from the raw suit and rank attributes of 101k for INVASE due to the two other models in the actor- the cards. The input-output relationship is deterministic and critic framework. TabNet is a single architecture, and its size hand-crafted rules can get 100% accuracy. Yet, conventional is 26k for Syn1-Syn and 31k for Syn4-Syn6. The compact DNNs, DTs, and even their hybrid variant of deep neural DTs

representation is one of TabNet’s valuable properties. (Yang, Morillo, and Hospedales 2018) severely suffer from the imbalanced data and cannot learn the required sorting and Performance on real-world datasets ranking operations (Yang, Morillo, and Hospedales 2018).

Tuned XGBoost, CatBoost, and LightGBM show very slight

as it can perform highly-nonlinear processing with its depth, Model Test accuracy (%) without overfitting thanks to instance-wise feature selection.

CatBoost 85.1 Table 4: Performance on Sarcos dataset. Three TabNet mod-

AutoML Tables 94.9 els of different sizes are considered.

Forest Cover Type (Dua and Graff 2017): The task is clas- MLP 2.1 0.14M

sification of forest cover type from cartographic variables. Adaptive neural tree 1.2 0.60M approaches that are known to achieve solid performance (AutoML 2019), an automated search framework based on TabNet-M 0.2 0.59M ensemble of models including DNN, gradient boosted DT, TabNet-L 0.1 1.75M with very thorough hyperparameter search. A single TabNet without fine-grained hyperparameter search outperforms it. Sarcos (Vijayakumar and Schaal 2000): The task is re-

gressing inverse dynamics of an anthropomorphic robot arm.

very small model is possible with a random forest. In the very and TabNet merely focuses on the relevant ones. For Syn4, small model size regime, TabNet’s performance is on par the output depends on either X -X or X -X depending parameters. When the model size is not constrained, TabNet feature selection – it allocates a mask to focus on the indi- achieves almost an order of magnitude lower test MSE. cator X , and assigns almost all-zero weights to irrelevant

features (the ones other than two feature groups). models are denoted with -S and -M. Real-world datasets: We first consider the simple task of mushroom edibility prediction (Dua and Graff 2017). Tab- Model Test acc. (%) Model size Net achieves 100% test accuracy on this dataset. It is indeed Sparse evolutionary MLP 78.4 81K known (Dua and Graff 2017) that “Odor” is the most discrim-

What is this project about?

This project covers practical implementation and research aspects of the topic using AI/ML techniques.