Enquire Now
Medical Computer Vision · Clinical Diagnostics · PyTorch / TensorFlow · GPU Optimized · 2026

Ai Powered Resume Parser And Ranking Capstone

Tensor Pipeline · Custom Loss Formulations · Model Quantization · Accelerated Inference — A rigorous deep learning engineering project focused on automated pathological lesion segmentation and radiological disease classification. Architected for thesis defense viva presentations, IEEE reproduction, and high-throughput production deployment.

PyTorch
Core Framework
AMP FP16
Mixed Precision
TensorRT
Quantized Serving

Faith in AI can narrow the futures individuals consider

Aoi Naito1,2 And Hirokazu Shirado1,∗

2School of Environment and Society, Institute of Science Tokyo, Tokyo, 108-0023, Japan. Artificial intelligence (AI) predictions are increasingly used to inform human decisions1,2. Here, using a behavioral implementation of the classic Newcomb’s paradox3 in 1,305 par- ticipants, we show that AI predictions can also shape the reasoning people use to make a decision. In this paradigm, perceived predictive authority can alter how people reason about their future actions, leading them to forgo a guaranteed reward. Over 40% of participants treated AI as such a predictive authority about their own behavior, signifi- cantly increasing the odds of forgoing the guaranteed reward by a factor of 3.39 (95% CI: 2.45–4.70) and reducing earnings by 10.7–42.9%. The effect appeared across AI presen- tations and decision contexts and remained detectable even when predictions repeatedly failed. When people perceive AI as capable of predicting their personal behavior, the mere presence of AI predictions may shape their decision-making, narrowing the futures they consider4.

ai powered resume parser and ranking capstone Diagram
Figure: System Model & Simulation Flow for Ai Powered Resume Parser And Ranking Capstone

Artificial intelligence is increasingly used to predict human behavior, and people are directly interacting with such systems1,2. From recommendation algorithms to large language models, AI routinely predicts what people will choose, say, and do5,6. In most applications, AI predic- tions are used to inform decisions, helping individuals and organizations perform tasks more effectively7–10. However, AI predictions about people’s own actions may become part of the decision environment itself, influencing the very behaviors they are intended to support3,11,12.

ai powered resume parser and ranking capstone Diagram
Figure: System Model & Simulation Flow for Ai Powered Resume Parser And Ranking Capstone

Such predictions may alter how people reason about their own future actions and, in some cases, even narrow the set of futures they consider and lead them to forgo guaranteed rewards. To test this, we conduct behavioral experiments using a decision paradigm based on the clas- sic Newcomb’s paradox, a two-choice task in which different reasoning processes can prescribe different choices.

Originally proposed as a philosophical thought experiment, Newcomb’s

Arxiv:2603.28944V2 [Cs.Hc] 3 Jul 2026

paradox has become a canonical problem in decision theory concerning prediction and rational choice3,13,14. We adapt this paradigm by replacing the abstract predictor with an AI system and deliberately withholding information about its predictive reliability, allowing us to examine how people’s own interpretations of the AI shape their reasoning and choices.

Participants face two boxes, where Box A always contains a guaranteed US$1 and Box B contains either US$0 or US$3 (Fig. 1A). Participants then choose either both boxes (two- boxing) or only Box B (one-boxing), and are paid accordingly. Before expressing their choice, participants are told that an AI system has predicted which option they will choose: if the AI predicts one-boxing, Box B contains US$3; if it predicts two-boxing, Box B contains US$0. At the time of choice, participants are, however, not told the prediction or what Box B contains.

Instead, they are informed that Box B’s content has already been determined by the AI’s prediction and cannot be changed, regardless of what they choose. From a strategic dominance perspective, two-boxing yields a higher payoff regardless of prediction (1 + 𝑋> 𝑋, where 𝑋∈{0, 3}). However, if individuals believe that AI can predict their future actions, they may instead treat the contents of Box B (𝑋) as contingent on the action they anticipate taking. Under this logic, one-boxing can be preferred (1 + 0 < 3), even though it requires forgoing Box A’s guaranteed reward. Unlike cooperation or social dilemma games15–17, this paradigm contains no strategic or social incentives that would otherwise justify choosing a lower-payoff option. Thus, observing one-boxing in this setup indicates that AI prediction can itself alter how individuals reason about their available actions, as if they believe the AI already “knows” what they will do.

Using this design, we conducted four preregistered online studies with 1,305 unique partic- ipants (see Methods). Study 1 (𝑁= 200) tested whether AI prediction increases one-boxing relative to a random control. Study 2 (𝑁= 601) examined the robustness of this effect and its underlying mechanism across different interaction contexts and interfaces. Study 3 (𝑁= 303) tested whether the effect generalizes beyond the economic task using vignette-based scenarios.

Finally, Study 4 (𝑁= 201) investigated how repeated interaction with AI prediction reshapes behavior over time.

2

AI prediction increases forgoing guaranteed rewards Across Studies 1 and 2, participants frequently forwent the guaranteed US$1 by choosing one-boxing when the decision was framed as being predicted by an AI system.

Study 1 tested this effect in a two-condition online experiment (𝑁= 200). Participants were told that the content of Box B would be determined either by an “AI system” predicting their choice or by a “random picker wheel” with the same payoff structure, but without using any (ostensible) prediction of their choice. In the AI condition, 41 of 100 participants (41.0%) chose one-boxing, compared with 26 of 100 (26.0%) in the random condition (𝑝= 0.025; Fig.

1B).

Study 2 (𝑁= 601) confirmed the robustness of this effect across different interaction contexts and system interfaces. In addition to the identity of the predictor (AI versus random mechanism), we manipulated whether participants interacted with the system before making their decision. In the interactive AI condition, participants exchanged brief messages with an AI system powered by OpenAI’s GPT-4.1, whereas the interactive random condition presented a random drawing process of equivalent duration.

Non-Interactive Conditions Removed All

system-specific interaction and described the outcome as determined by either an “AI system”

Or A “Random Generator.”

Under random framing, one-boxing remained uncommon (15.3% in both interactive and non-interactive conditions; Fig. 1C). In contrast, one-boxing was substantially more frequent when decisions were framed as predicted by AI (45.0% in the non-interactive AI condition and 42.0% in the interactive AI condition).

The Overall Increase In One-Boxing Under Ai

prediction was statistically significant (𝑝< 0.001), whereas interaction with the system did not significantly moderate the effect (𝑝= 1.0 for both random and AI conditions; Extended Data Table 1).

A fixed-effect meta-analysis across Studies 1 and 2 confirmed this effect: AI prediction increased the odds of forgoing the guaranteed reward by a factor of 3.39 (95% CI: 2.45–4.70; 𝑝< 0.001). This shift had measurable economic consequences. The observed shift toward one- boxing reduced realized earnings by 10.7–42.9% relative to the two-boxing baseline, depending

3

on the AI prediction regime. To examine whether this effect generalizes beyond the stylized economic task, Study 3 presented participants (𝑁= 303) with three vignette scenarios adapted from Newcomb’s paradox: a job interview decision, a mobile data coupon choice, and a task application on a freelancing platform. Each scenario was presented under three conditions: no prediction, human-expert prediction, and AI prediction.

Across scenarios, participants chose the one-box option in 26.7% of the cases under AI prediction and 36.6% under human-expert prediction, compared to 10.6% without prediction (Fig. 1D). Pairwise contrasts confirmed that AI predictions significantly increased one-box-type choices relative to control (𝑝< 0.001). Human-expert predictions produced a somewhat larger effect than AI predictions (odds ratio = 1.67, 𝑝= 0.032). Qualitative responses suggested that some participants regarded human experts as more socially acceptable or appropriate sources of prediction than AI (Extended Data Table 2). Nevertheless, both predictive sources shifted choices in the same direction, suggesting that AI can exert a behavioral influence resembling that of established human sources of predictive authority. While the direction of the effect was consistent across scenarios, its magnitude varied (Extended Data Fig. 1).

Causal and evidential reasoning about AI prediction Why did the mere presence of AI predictions make participants forgo guaranteed rewards? The system neither revealed its prediction nor provided any recommendation at the time of choice, ruling out explicit AI persuasion18,19. Moreover, identical payoff structures framed around random processes did not produce comparable behavior, indicating that the effect is not driven by payoff structure alone.

Building on prior theoretical interpretations of Newcomb’s paradox, two forms of reasoning can favor different choices3. Under so-called causal reasoning, individuals treat their action as affecting the payoff outcome independently of a predetermined prediction14. As a result, two-boxing is preferable because it always yields an additional US$1 (Fig. 1A), just as it would in the absence of any predictor.

In contrast, under so-called evidential reasoning, individuals treat whichever action they

4

ultimately take as evidence of what has already been predicted.

Such Reasoning May Be

supported when they believe that a system can indeed predict their future actions (perceived predictiveness) and when they take actions consistent with those they anticipate taking (internal coherence) (Fig. 2A). When both are strong, individuals may associate one-boxing with a full Box B (payoff of US$3) and two-boxing with an empty Box B (payoff of US$1), making one-boxing appear reasonable.

Consistent with the role of perceived predictiveness, post-decision evaluations in Study 2 show that participants in the AI conditions perceived the system’s predictions to be significantly more accurate than chance (non-interactive AI: 62.1%, 𝑝< 0.001; interactive AI: 62.9%, 𝑝< 0.001), whereas participants in the random conditions perceived chance-level accuracy (non-interactive random: 50.4%, 𝑝= 0.322; interactive random: 49.7%, 𝑝= 0.639)(Fig.

2B). Notably, they formed these beliefs despite receiving no information about the system’s ostensible predictive accuracy. These evaluations were collected after choice but before outcome disclosure, making post-hoc justification unlikely.

Perceived predictiveness alone, however, was insufficient to explain one-boxing. Participants who one-boxed and two-boxed held similar beliefs about the AI’s predictive accuracy (non- interactive AI: 𝑝= 0.080; interactive AI: 𝑝= 0.634; Fig. 2B). Individuals who believe that their anticipated action has been predicted may nevertheless choose two-boxing if they treat their actual action as independent of the prediction. This suggests that one-boxing depends not only on perceived predictiveness, but also on internal coherence—the tendency to act consistently with the actions one anticipates taking. AI might, in other words, have this sort of thoroughgoing mental impact. A computational model formalizing behavior as a mixture of causal and evidential reasoning further supports this interpretation, suggesting that AI prediction shifts participants toward evidential reasoning over causal reasoning (see Supplementary Text).

Qualitative responses further suggested that participants who chose one-boxing often de- scribed their decisions in relation to the AI’s prediction, whereas participants who chose two- boxing more often described their choices as independent of the prediction (Extended Data Table 3). Notably, participants rarely explained their choices in terms of preferring one option over the other. These responses are consistent with the interpretation that the observed be-

5

havioral differences primarily reflect differences in reasoning about the prediction rather than differences in preferences over the available options. Finally, exploratory analyses in Studies 1 and 2 examined individual characteristics associ- ated with one-boxing (Extended Data Fig. 2). Sociodemographic variables and beliefs about free will or determinism, as well as affective attitudes toward AI, showed no association with choice in the task. The only individual characteristics associated with one-boxing were higher AI literacy (encompassing awareness, usage, and evaluation of AI systems; 𝑝= 0.011) and also greater risk-taking preference (𝑝= 0.003). This pattern is inconsistent with explanations attributing the effect to limited understanding of AI.

Ai Prediction Becomes Self-Fulfilling

Following the standard formulation of Newcomb’s paradox, Studies 1–3 required participants to make a choice before the prediction was revealed (Fig. 1A). In real-world settings, however, people often interact repeatedly with predictive AI systems and observe whether their predictions prove correct. We therefore conducted Study 4 (𝑁= 201) to examine how repeated interaction with AI prediction influences behavior over time.

Participants completed the same choice task over five consecutive rounds, receiving feedback after each round about the AI’s prediction, their own choice, and the resulting outcome. Partic- ipants were randomly assigned to one of two conditions: in one condition the AI consistently predicted one-boxing, whereas in the other it consistently predicted two-boxing. Participants were not informed of the prediction policy, and the system did not update its predictions during the task.

Behavior diverged depending on the AI’s prediction policy (interaction 𝑝= 0.002; Extended Data Table 4; Fig. 3A). When the AI consistently predicted one-boxing, the proportion of one- boxing remained stable across rounds (slope 𝑝= 0.763). In contrast, when the AI consistently predicted two-boxing, the proportion of one-boxing declined significantly over time (slope 𝑝< 0.001). Notably, even after AI’s five consecutive predictive failures in the two-boxing prediction condition, the proportion of one-boxing in the final round (30.6%) remained significantly higher than in the random condition of Study 2 (15.3%; 𝑝= 0.003). These results indicate that AI

6

prediction can influence people’s reasoning and behavior even when they experience its repeated failures. This behavioral adaptation, in turn, altered the accuracy of the AI’s predictions as a by- product (Fig. 3B). Accuracy significantly increased when the AI consistently predicted two- boxing, while remaining relatively stable when the AI consistently predicted one-boxing. As a consequence, overall prediction accuracy increased from 50.7% to 59.2% (slope 𝑝= 0.002; Extended Data Table 5), even though the AI’s prediction policy remained fixed and participants’ beliefs about its prediction accuracy remained stable at 57.6%, on average (slope 𝑝= 0.513; Extended Data Table 6).

This pattern is consistent with a self-fulfilling prophecy, in which even an initially arbitrary prediction brings about the very outcome it predicts12,20. Although this increase in accuracy resulted from human behavioral adaptation, observers who focus on prediction outcomes may nevertheless conclude that the AI is predictive, potentially reinforcing broader beliefs about AI capability.

Ongoing interactions around AI predictions may therefore make AI appear increasingly capable, even in the absence of computational improvements.

Discussion

People appear to think differently about the future in the presence of an AI system, and some might perceive it as having preternatural authority. AI prediction does not merely forecast human behavior, but its existence alone can also shape the reasoning people use to make decisions in the first place. The mere presence of AI prediction can lead people to forgo a guaranteed reward, even in the absence of strategic incentives or social pressure.

In our experiments, this pattern generalized across different AI presentations and across both economically incentivized and vignette-based decision contexts. The effect also remained detectable under repeated interactions, even when participants repeatedly observed mismatches between the system’s prediction and their own choices. Together, these results suggest that AI prediction can influence behavior by shaping how people reason about their own future actions.

We do not interpret this pattern as irrationality per se. Forgoing guaranteed rewards can be rational even from an expected utility perspective under so-called evidential reasoning13,14,21.

7

That is, predictive authority binds people’s anticipated actions to what the system predicts, narrowing the futures they themselves regard as plausible. People may therefore align their choices with the future implied by the system’s prediction11, even when this conflicts with immediate economic incentives. The mechanism observed here depends on the reflexive nature of human cognition22 and is not limited to the one-shot temporal structure of Newcomb’s paradox. Once a prediction becomes part of how people reason about their future actions, it can continue to shape subsequent behavior and even generate self-reinforcing dynamics12,23.

This shift toward evidential reasoning has important implications for human agency24. In our experiments, because the prediction was generated in advance, participants were objectively free to choose the higher-payoff option. Nevertheless, the prior prediction may already have shaped the decision-making process, even without conscious awareness. The experience of agency may therefore come to reflect not only one’s own intentions, but also the anticipated prediction25,26.

The influence of prediction on decision-making is not unique to AI. Study 3 supports this point, showing that predictions from a human expert produced an even stronger effect than AI predictions.

Human behavior has long been organized around anticipated expectations within families, communities, and other social institutions20,27,28. Among these, some of the most influential sources of predictive authority have historically included oracles, prophets, and diviners29,30. Our findings suggest that, even without universal trust in or acceptance of AI, predictive AI systems may become a new source of predictive authority31.

AI is particularly well positioned to occupy this role because it is increasingly perceived as a capable predictor of human behavior across many domains of everyday life1,32,33. In our experiments, despite receiving neither information about, nor experience with, AI’s predictive capability, some participants nevertheless inferred predictive authority and incorporated it into their reasoning about their own future actions.

This Suggests That Predictive Binding May

arise naturally in interactions with AI. Moreover, AI systems can plausibly be perceived as maintaining stable expectations about people’s future behavior, allowing predictive authority to persist and reinforcing self-fulfilling dynamics4.

This possibility bears directly on collective action34. Social dilemmas require individuals

8

to forgo immediate gains in anticipation of what others—or a larger system—will do15,35,36. Machine predictors may therefore influence collective outcomes not only by informing strategic decisions, but by shaping how people reason about the future behavior of themselves and others27,37,38. In this sense, AI could either support or undermine human cooperation, depending on which behaviors become anticipated and reinforced within a given environment39,40. In hybrid human–machine social systems2,41,42, AI will influence collective behavior through the interpersonal expectations and behavioral adaptations the systems create43.

Our study intentionally isolates a minimal decision environment, allowing us to identify a basic behavioral mechanism under tightly controlled conditions. This design necessarily omits many features of real-world choice, including strategic interaction, social norms, and cultural and historical context15,28,38. At the same time, this simplicity is a strength: it shows that the behavioral influence of prediction can emerge even in the absence of social enforcement, institutional rules, or actual predictive capabilities.

In our experiments, the AI systems did not perform any algorithmic prediction and so the observed effects reflect human psychology alone. Comparable effects across multiple AI pre- sentations—a virtual agent, an LLM-based assistant, and the label “AI” alone—suggest that the phenomenon depends less on technical sophistication than on perceived predictive authority.

Future work should examine how such psychological and cognitive effects interact with pre- dictive systems that demonstrably anticipate behavior. Real-world AI systems with verifiable predictive accuracy may elicit stronger or more persistent beliefs, potentially amplifying the effects observed here.

As AI becomes commonplace, its influence may lie not only in what it predicts, but in how prediction reshapes decision-making itself. A culture with machines may affect the culture of humans. Our results highlight AI as a source of behavioral influence that can operate without coercion, recommendation, or incentive change. Addressing the societal impact of AI will therefore require attention not only to what AI predicts or does, but also to how the perceived authority of AI reshapes the decisions people experience as their own.

Ai Adaptive Learning

This project focuses on ai adaptive learning using modern AI and machine learning techniques. The content below is adapted from research literature and practical implementation notes.

We propose a novel high-performance and interpretable canon-

addition, unlike tree learning, DNNs enable gradient descent- ical deep tabular data learning architecture, TabNet. TabNet based end-to-end learning for tabular data which can have a uses sequential attention to choose which features to reason multitude of benefits: (i) efficiently encoding multiple data from at each decision step, enabling interpretability and more types like images along with tabular data; (ii) alleviating the efficient learning as the learning capacity is used for the most need for feature engineering, which is currently a key aspect

salient features. We demonstrate that TabNet outperforms in tree-based tabular data learning methods; (iii) learning other variants on a wide range of non-performance-saturated from streaming data and perhaps most importantly (iv) end- tabular datasets and yields interpretable feature attributions to-end models allow representation learning which enables plus insights into its global behavior. Finally, we demonstrate many valuable application scenarios including data-efficient

self-supervised learning for tabular data, significantly improv- domain adaptation (Goodfellow, Bengio, and Courville 2016), ing performance when unlabeled data is abundant. generative modeling (Radford, Metz, and Chintala 2015) and

Introduction We propose a new canonical DNN architecture for tabular

Deep neural networks (DNNs) have shown notable success data, TabNet. The main contributions are summarized as: efficiently encode the raw data into meaningful representa- enabling flexible integration into end-to-end learning. tions, fuel the rapid progress. One data type that has yet to 2. TabNet uses sequential attention to choose which fea- see such success with a canonical architecture is tabular data. tures to reason from at each decision step, enabling in-

Despite being the most common data type in real-world AI terpretability and better learning as the learning capacity (as it is comprised of any categorical and numerical features), is used for the most salient features (see Fig. 1). This under-explored, with variants of ensemble decision trees for each input, and unlike other instance-wise feature se- Why? First, because DT-based approaches have certain bene- and van der Schaar 2019), TabNet employs a single deep

fits: (i) they are representionally efficient for decision mani- learning architecture for feature selection and reasoning. folds with approximately hyperplane boundaries which are 3. Above design choices lead to two valuable properties: (i) common in tabular data; and (ii) they are highly interpretable TabNet outperforms or is on par with other tabular learn- in their basic form (e.g. by tracking decision nodes) and there ing models on various datasets for classification and re-

are popular post-hoc explainability methods for their ensem- gression problems from different domains; and (ii) TabNet ble form, e.g. (Lundberg, Erion, and Lee 2018) – this is an enables two kinds of interpretability: local interpretability important concern in many real-world applications; (iii) they that visualizes the importance of features and how they are fast to train. Second, because previously-proposed DNN are combined, and global interpretability which quantifies

architectures are not well-suited for tabular data: e.g. stacked the contribution of each feature to the trained model. convolutional layers or multi-layer perceptrons (MLPs) are 4. Finally, for the first time for tabular data, we show signif- vastly overparametrized – the lack of appropriate inductive icant performance improvements by using unsupervised bias often causes them to fail to find optimal solutions for tab- pre-training to predict masked features (see Fig. 2).

ular decision manifolds (Goodfellow, Bengio, and Courville

Why is deep learning worth exploring for tabular data?

One obvious motivation is expected performance improve- Feature selection: Feature selection broadly refers to judi- Copyright © 2021, Association for the Advancement of Artificial ciously picking a subset of features based on their useful-

Professional occupation related Investment related

Feedback from Feedback to

Feature selection Input processing Feature selection Input processing

previous step next step … …

Predicted output (whether the income level >$50k)

selection enables interpretability and better learning as the capacity is used for the most salient features. TabNet employs multiple decision blocks that focus on processing a subset of input features for reasoning. Two decision blocks shown as examples process features that are related to professional occupation and investments, respectively, in order to predict the income level.

Unsupervised pre-training Supervised fine-tuning

Age Cap. gain Education Occupation Gender Relationship Age Cap. gain Education Occupation Gender Relationship 5 2000 ? Exec-managerial F Wife 6 2000 Bachelors Exec-managerial M Husband 1 0 ? Farming-fishing M ? 2 0 High-school Farming-fishing M Unmarried

? 50 Doctorate Prof-specialty M Husband 4 50 Doctorate Prof-specialty M Husband 2 ? ? Handlers-cleaners F Wife 2 0 High-school Handlers-cleaners F Wife 5 3000 Bachelors ? ? Husband 5 3000 Bachelors Exec-managerial M Husband

3 0 Bachelors ? F ? 3 100 Bachelors Prof-specialty F Wife ? 0 High-school Armed-Forces ? Husband 2 0 High-school Armed-Forces M Husband

TabNet decoder Decision making

Age Cap. gain Education Occupation Gender Relationship Income > $50k

3 M False

level can be guessed from the occupation, or the gender can be guessed from the relationship. Unsupervised representation learning by masked self-supervised learning results in an improved encoder model for the supervised learning task.

ward selection and Lasso regularization (Guyon and Elisseeff performance with compact representations. 2003) attribute feature importance based on the entire training Tree-based learning: DTs are commonly-used for tabular data, and are referred as global methods. Instance-wise fea- data learning. Their prominent strength is efficient picking ture selection refers to picking features individually for each of global features with the most statistical information gain

to maximize the mutual information between the selected mance of standard DTs, one common approach is ensembling features and the response variable, and in (Yoon, Jordon, and to reduce variance. Among ensembling methods, random van der Schaar 2019) by using an actor-critic framework to forests (Ho 1998) use random subsets of data with randomly mimic a baseline while optimizing the selection. Unlike these, selected features to grow many trees. XGBoost (Chen and

sity in end-to-end learning – a single model jointly performs recent ensemble DT approaches that dominate most of the feature selection and output mapping, resulting in superior recent data science competitions. Our experimental results

!# + Softmax !" < % !" > % !# > & !# > &

ReLU ReLU &

$" !" − $" % −1 −$" !" + $" % −1 −1 $# !# − $# & % −1 −$# !# + $# & !"

FC FC

W: [$" , - $" , 0, 0] W: [0, 0, $# , - $# ] !" < % b: [-a $" , a $" , -1, -1] b: [-1, -1, -d $# , d $# ] !# < & !" > % !# < & [!" ] [!# ]

M: [1, 0] M: [0, 1]

(right). Relevant features are selected by using multiplicative sparse masks on inputs. The selected features are linearly transformed, and after a bias addition (to represent boundaries) ReLU performs region selection by zeroing the regions. Aggregation of multiple regions is based on addition. As C and C get larger, the decision boundary gets sharper.

for various datasets show that tree-based models can be out- constructs a sequential multi-step architecture, where each performed when the representation capacity is improved with step contributes to a portion of the decision based on the deep learning while retaining their feature selecting property. selected features; (iii) improves the learning capacity via non- Integration of DNNs into DTs: Representing DTs with linear processing of the selected features; and (iv) mimics

DNN building blocks as in (Humbird, Peterson, and McClar- ensembling via higher dimensions and more steps. ren 2018) yields redundancy in representation and ineffi- cient learning. Soft (neural) DTs (Wang, Aggarwal, and Liu Fig. 4 shows the TabNet architecture for encoding tabu- functions, instead of non-differentiable axis-aligned splits. mapping of categorical features with trainable embeddings.

However, losing automatic feature selection often degrades We do not consider any global feature normalization, but performance. In (Yang, Morillo, and Hospedales 2018), a soft merely apply batch normalization (BN). We pass the same D- binning function is proposed to simulate DTs in DNNs, by dimensional features f ∈ <B×D to each decision step, where 2019) proposes a DNN architecture by explicitly leveraging multi-step processing with Nsteps decision steps. The ith

expressive feature combinations, however, learning is based step inputs the processed information from the (i − 1)th step on transferring knowledge from gradient-boosted DT. (Tanno to decide which features to use and outputs the processed ing from primitive blocks while representation learning into sion. The idea of top-down attention in the sequential form edges, routing functions and leaf nodes. TabNet differs from is inspired by its applications in processing visual and text

these as it embeds soft feature selection with controllable data (Hudson and Manning 2018) and reinforcement learn- Self-supervised learning: Unsupervised representation relevant information in high dimensional input. learning improves supervised learning especially in small Feature selection: We employ a learnable mask M[i] ∈ has shown significant advances – driven by the judicious capacity of a decision step is not wasted on irrelevant

choice of the unsupervised learning objective (masked input ones, and thus the model becomes more parameter effi- prediction) and attention-based deep learning. cient. The masking is multiplicative, M[i] · f . We use an attentive transformer (see Fig. 4) to obtain the masks us- TabNet for Tabular Learning ing the processed features from the preceding step, a[i − 1]:

M[i] = sparsemax(P[i − 1] · hi (a[i − 1])). Sparsemax nor-

DTs are successful for learning from real-world tabular malization (Martins and Astudillo 2016) encourages sparsity datasets. With a specific design, conventional DNN building by mapping the Euclidean projection onto the probabilistic blocks can be used to implement DT-like output manifold, simplex, which is observed to be superior in performance and e.g. see Fig. 3). In such a design, individual feature selec- aligned with the goal of sparse feature selection for explain-

tion is key to obtain decision boundaries in hyperplane form, PD which can be generalized to a linear combination of features ability. Note that j=1 M[i]b,j = 1. hi is a trainable func- where coefficients determine the proportion of each feature. tion, shown in Fig. 4 using a FC layer, followed by BN. P[i] TabNet is based on such functionality and it outperforms DTs is the prior scale term, denoting how much a particular feature

Qi while reaping their benefits by careful design which: (i) uses has been used previously: P[i] = j=1 (γ − M[j]), where γ sparse instance-wise feature selection learned from data; (ii) is a relaxation parameter – when γ = 1, a feature is enforced

+ Softmax

Feature Feature …

transformer transformer

x Nsteps Features

+ Softmax

Feature …

transformer transformer Feature Feature Feature Feature transformer

Encoded representation

transformer transformer Attentive transformer … Mask transformer …

Step 2 Decision step dependent

transformer transformer

BN Feature Feature

FC BN transformer transformer

+ 0.5 0.5 0.5 Agg. Agg. Features Features FC FC + +

Reconstructed + … Feature attributes + … features

(a) TabNet encoder architecture (b) TabNet decoder architecture Feature transformer Feature Attentive transformer Shared across decision steps Decision step dependent transformer GLU

Decision step dependent Prior scales

+ 0.5 0.5 0.5

0.5 0.5 0.5

+ Attentive transformer (c) (d)

Prior scales

divides the processed representation to be used by the attentive transformer of the subsequent step as well as for the overall Attentive BN FC

output. For each step, the feature selection mask provides interpretable information about the model’s functionality, and the +

masks can be aggregated to obtain global feature transformer important attribution. (b) TabNet decoder, composed of a feature transformer block at each step. (c) A feature transformer block example – 4-layer network is shown, where 2 are shared across all decision

Prior scales

steps and 2 are decision step-dependent. Each layer is composed of a fully-connected (FC) layer, BN and GLU nonlinearity. (d) +

An attentive transformer block example – a single layer mapping is modulated with a prior scale information which aggregates Sparsemax

how much each feature has been used before the current decision step. sparsemax (Martins and Astudillo 2016) is used for BN FC

normalization of the coefficients, resulting in sparse selection of the salient features. +

to be used only at one decision step and as γ increases, more propose the aggregate.feature importance mask, Magg−b,j = flexibility is provided to use a feature at multiple decision PNsteps ηb [i]Mb,j [i]

PD PNsteps

ηb [i]Mb,j [i].2 i=1 i=1 steps. P is initialized as all ones, 1B×D , without any prior j=1

on the masked features. If some features are unused (as in self- Tabular self-supervised learning: We propose a decoder supervised learning), corresponding P entries are made 0 architecture to reconstruct tabular features from the Tab- to help model’s learning. To further control the sparsity of the Net encoded representations. The decoder is composed of selected features, we propose sparsity regularization in the feature transformers, followed by FC layers at each deci-

form of entropy (Grandvalet and Bengio 2004), Lsparse = sion step. The outputs are summed to obtain the recon-

PNsteps PB PD −Mb,j [i] log(Mb,j [i]+)

i=1 b=1 j=1 Nsteps ·B , where  is a structed features. We propose the task of prediction of miss- small number for numerical stability. We add the sparsity reg- ing feature columns from the others. Consider a binary mask ularization to the overall loss, with a coefficient λsparse . Spar- S ∈ {0, 1}B×D . The TabNet encoder inputs (1 − S) · f̂ sity provides a favorable inductive bias for datasets where and the decoder outputs the reconstructed features, S · f̂ . We

most features are redundant. initialize P = (1 − S) in encoder so that the model em- Feature processing: We process the filtered features using phasizes merely on the known features, and the decoder’s last a feature transformer (see Fig. 4) and then split for the FC layer is multiplied with S to output the unknown features. decision step output and information for the subsequent We consider the reconstruction loss in self-supervised phase:

step, [d[i], a[i]] = fi (M[i] · f ), where d[i] ∈ <B×Nd and 2

PB PD (f̂b,j −fb,j )·Sb,j

a[i] ∈ <B×Na . For parameter-efficient and robust learning b=1 j=1

√ PB PB 2

. Normalization b=1 (fb,j −1/B b=1 fb,j ) with high capacity, a feature transformer should comprise layers that are shared across all decision steps (as the same with the population standard deviation of the ground truth features are input across different decision steps), as well as is beneficial, as the features may have different ranges. We decision step-dependent layers. Fig. 4 shows the implementa- sample Sb,j independently from a Bernoulli distribution with

tion as concatenation of two shared layers and two decision parameter ps , at each iteration. step-dependent layers. Each FC layer is followed by BN and eventually connected to a normalized residual √ connection We study TabNet in wide range of problems, that contain with normalization. Normalization with 0.5 helps to sta- regression or classification tasks, particularly with published bilize learning by ensuring that the variance throughout the benchmarks. For all datasets, categorical inputs are mapped

For faster training, we use large batch sizes with BN. Thus, bedding and numerical columns are input without and pre- except the one applied to the input features, we use ghost BN processing.4 We use standard classification (softmax cross (Hoffer, Hubara, and Soudry 2017) form, using a virtual batch entropy) and regression (mean squared error) loss functions size BV and momentum mB . For the input features, we ob- and we train until convergence. Hyperparameters of the Tab-

serve the benefit of low-variance averaging and hence avoid Net models are optimized on a validation set and listed in ghost BN. Finally, inspired by decision-tree like aggregation Appendix. TabNet performance is not very sensitive to most as in Fig. 3, we construct the overall decision embedding hyperparameters as shown with ablation studies in Appendix. as dout = i=1 PNsteps ReLU(d[i]). We apply a linear mapping In Appendix, we also present ablation studies on various de-

Wfinal dout to get the output mapping.1 sign and guidelines on selection of the key hyperparameters. Interpretability: TabNet’s feature selection masks can shed For all experiments we cite, we use the same training, val- light on the selected features at each step. If Mb,j [i] = 0, idation and testing data split with the original work. Adam optimization algorithm (Kingma and Ba 2014) and Glorot then j th feature of the bth sample should have no contribution uniform initialization are used for training of all models.5

to the decision. If fi were a linear function, the coefficient

Mb,j [i] would correspond to the feature importance of fb,j . Instance-wise feature selection

Although each decision step employs non-linear processing, their outputs are combined later in a linear way. We aim Selection of the salient features is crucial for high perfor- to quantify an aggregate feature importance in addition to mance, especially for small datasets. We consider 6 tabular requires a coefficient that can weigh the relative importance samples). The datasets are constructed in such a way that of each step in the decision. We simply propose ηb [i] = only a subset of the features determine the output. For Syn1-

PNd Syn3, salient features are same for all instances (e.g., the

c=1 ReLU(db,c [i]) to denote the aggregate decision con- tribution at ith decision step for the bth sample. Intuitively, if 2

Normalization is used to ensure D

P j=1 Magg−b,j = 1. db,c [i] < 0, then all features at ith decision step should have 3

0 contribution to the overall decision. As its value increases, prove the performance, but interpretation of individual dimensions

it plays a higher role in the overall linear combination. Scal- may become challenging. ing the decision mask at each decision step with ηb [i], we Specially-designed feature engineering, e.g. logarithmic trans- formation of variables highly-skewed distributions, may further

For discrete outputs, we additionally employ softmax during

training (and argmax during inference). An open-source implementation will be released.

Global: using only globally-salient features, Tree Ensembles (Geurts, Ernst, and Wehenkel 2006), Lasso-regularized model, L2X

Syn Syn Syn Syn Syn Syn

No selection .5 ± .0 .7 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .6 ± .0 Tree .5 ± .1 .8 ± .0 .8 ± .0 .6 ± .0 .7 ± .0 .7 ± .0 Lasso-regularized .4 ± .0 .5 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .7 ± .0

INVASE .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0

Global .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0 TabNet .6 ± .0 .8 ± .0 .8 ± .0 .7 ± .0 .7 ± .0 .8 ± .0

output of Syn depends on features X -X ), and global fea- Table 3: Performance for Poker Hand induction dataset. ture selection, as if the salient features were known, would give high performance. For Syn4-Syn6, salient features are Model Test accuracy (%) instance dependent (e.g., for Syn4, the output depends on ei- DT 50.0 ther X -X or X -X depending on the value of X ), which MLP 50.0

makes global feature selection suboptimal. Table 1 shows that Deep neural DT 65.1

TabNet outperforms others (Tree Ensembles (Geurts, Ernst, XGBoost 71.1

and Wehenkel 2006), LASSO regularization, L2X (Chen LightGBM 70.0 van der Schaar 2019). For Syn1-Syn3, TabNet performance TabNet 99.2 is close to global feature selection - it can figure out what Rule-based 100.0 features are globally important. For Syn4-Syn6, eliminating instance-wise redundant features, TabNet improves global feature selection. All other methods utilize a predictive model Poker Hand (Dua and Graff 2017): The task is classifica-

with 43k parameters, and the total number of parameters is tion of the poker hand from the raw suit and rank attributes of 101k for INVASE due to the two other models in the actor- the cards. The input-output relationship is deterministic and critic framework. TabNet is a single architecture, and its size hand-crafted rules can get 100% accuracy. Yet, conventional is 26k for Syn1-Syn and 31k for Syn4-Syn6. The compact DNNs, DTs, and even their hybrid variant of deep neural DTs

representation is one of TabNet’s valuable properties. (Yang, Morillo, and Hospedales 2018) severely suffer from the imbalanced data and cannot learn the required sorting and Performance on real-world datasets ranking operations (Yang, Morillo, and Hospedales 2018).

Tuned XGBoost, CatBoost, and LightGBM show very slight

as it can perform highly-nonlinear processing with its depth, Model Test accuracy (%) without overfitting thanks to instance-wise feature selection.

CatBoost 85.1 Table 4: Performance on Sarcos dataset. Three TabNet mod-

AutoML Tables 94.9 els of different sizes are considered.

Forest Cover Type (Dua and Graff 2017): The task is clas- MLP 2.1 0.14M

sification of forest cover type from cartographic variables. Adaptive neural tree 1.2 0.60M approaches that are known to achieve solid performance (AutoML 2019), an automated search framework based on TabNet-M 0.2 0.59M ensemble of models including DNN, gradient boosted DT, TabNet-L 0.1 1.75M with very thorough hyperparameter search. A single TabNet without fine-grained hyperparameter search outperforms it. Sarcos (Vijayakumar and Schaal 2000): The task is re-

gressing inverse dynamics of an anthropomorphic robot arm.

very small model is possible with a random forest. In the very and TabNet merely focuses on the relevant ones. For Syn4, small model size regime, TabNet’s performance is on par the output depends on either X -X or X -X depending parameters. When the model size is not constrained, TabNet feature selection – it allocates a mask to focus on the indi- achieves almost an order of magnitude lower test MSE. cator X , and assigns almost all-zero weights to irrelevant

features (the ones other than two feature groups). models are denoted with -S and -M. Real-world datasets: We first consider the simple task of mushroom edibility prediction (Dua and Graff 2017). Tab- Model Test acc. (%) Model size Net achieves 100% test accuracy on this dataset. It is indeed Sparse evolutionary MLP 78.4 81K known (Dua and Graff 2017) that “Odor” is the most discrim-

What is this project about?

This project covers practical implementation and research aspects of the topic using AI/ML techniques.