Enquire Now
Medical Computer Vision · Clinical Diagnostics · PyTorch / TensorFlow · GPU Optimized · 2026

Deep Learning Capstone Project Ideas Github

Tensor Pipeline · Custom Loss Formulations · Model Quantization · Accelerated Inference — A rigorous deep learning engineering project focused on automated pathological lesion segmentation and radiological disease classification. Architected for thesis defense viva presentations, IEEE reproduction, and high-throughput production deployment.

PyTorch
Core Framework
AMP FP16
Mixed Precision
TensorRT
Quantized Serving

Learn to Accumulate Evidence from All Training Samples: Theory and Practice

Computationally Efficient Way To Turn A Determin-

istic neural network uncertainty-aware. The resul-

Uncertainty Using The Learned Evidence. To En-

sure theoretically sound evidential models, the ev-

Special Activation Functions For Model Training

and inference. This constraint often leads to infe-

Them To Many Large-Scale Datasets. To Unveil The

real cause of this undesired behavior, we theoreti- cally investigate evidential models and identify a

Into Such Regions. A Deeper Analysis Of Eviden-

tial activation functions based on our theoretical

Underpinning Inspires The Design Of A Novel Regu-

larizer that effectively alleviates this fundamental

Lenging Real-World Datasets And Settings Confirm

our theoretical findings and demonstrate the effec- tiveness of our proposed approach.

1. Introduction

Deep Learning (DL) models have found great success in many real-world applications such as speech recognition (Kamath et al., 2019), machine translation (Singh et al., 2017), and computer vision (Voulodimos et al., 2018). How- ever, these highly expressive models may easily fit the noise in the training data, which leads to overconfident predictions (Nguyen et al., 2015). The challenge is further compounded when learning from limited labeled data, which is common 1Rochester Institute of Technology. Correspondence to: Qi Yu Proceedings of the 40 th International Conference on Machine Learning, Honolulu, Hawaii, USA. PMLR 202, 2023. Copyright 2023 by the author(s).

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

for applications from specialized domain (e.g., medicine, public safety, and military operations) where data collec- tion and annotation is highly costly. Accurate uncertainty quantification is essential for successful application of DL models in these domains. To this end, DL models have been augmented to become uncertainty-aware (Gal & Ghahra- mani, 2016; Blundell et al., 2015; Pearce et al., 2020). How- ever, commonly used extensions require expensive sampling operations (Gal & Ghahramani, 2016; Blundell et al., 2015), which significantly increase the computational costs (Lak- shminarayanan et al., 2017).

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

The recently developed evidential models bring together evidential theory (Shafer, 1976; Jøsang, 2016) and deep neural architectures that turn a deterministic neural network uncertainty-aware. By leveraging the learned evidence, evi- dential models are capable of quantifying fine-grained un- certainty that helps to identify the sources of ‘unknowns’.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

Furthermore, since only lightweight modifications are intro- duced to existing DL architectures, additional computational costs remain minimum. Such evidential models have been successfully extended to classification (Sensoy et al., 2018), regression (Amini et al., 2020), meta-learning (Pandey & Yu, 2022a), and open-set recognition (Bao et al., 2021) settings.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

Architectures In Rela-

tively simple learning problems. They suffer from a sig- nificant performance drop when facing large datasets with more complex features even in the common classification setting. As shown in Figure 1, an evidential model using ReLU activation and an evidential MSE loss (Sensoy et al., 2018) only achieves 36% test accuracy on Cifar100, which is almost 40% lower than a standard model trained using softmax. Additionally, most evidential models can easily break down with minor architecture changes and/or have a much stronger dependency on hyperparameter tuning to achieve reasonable predictive performance. The experiment section provides more details on these failure cases.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

Arxiv:2306.11113V2 [Cs.Lg] 24 Jun 2023

Learn to Accumulate Evidence from All Training Samples: Theory and Practice Figure 2. Visualization of zero-evidence region for evidential mod- els with ReLU activation in a binary classification setting. Existing models fail to learn from samples that are mapped to such zero- evidence region (shared area at the bottom left quadrant).

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

To train uncertainty-aware evidential models that can also predict well, we perform a novel theoretical analysis with a focus on the standard classification setting to unveil the underlying cause of the performance gap. Our theoreti- cal results show that existing evidential models learn sub- optimally compared to corresponding softmax counterparts.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

Such sub-optimal training is mainly attributed to the inher- ent learning deficiency of evidential models that prevents them from learning across all training samples. More specif- ically, they are incapable to acquire new knowledge from training samples mapped to “zero-evidence regions” in the evidence space, where the predicted evidence reduces to zero. The sub-optimal learning phenomenon is illustrated in Figure 2 (detailed discussion is presented in Section 4.2).

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

We analyze different variants of evidential models present in the existing literature and observe this limitation across all the models and settings. Our theoretical results inspire the design of a novel Regularized Evidential model (RED) that includes positive evidence regularization in its train- ing objective to battle the learning deficiency. Our major

Contributions Can Be Summarized As Follows:

• We identify a fundamental limitation of evidential models, i.e., lack the capability to learn from any data samples that lie in the “zero-evidence” region in the evidence space.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

• We theoretically show the superiority of evidential models with exp activation over other activation functions. • We conduct novel evidence regularization that enables evidential models to avoid the “zero-evidence” region so that they can effectively learn from all training samples.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

• We carry out experiments over multiple challenging real- world datasets to empirically validate the presented theory, and show the effectiveness of our proposed ideas.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

2. Related Works

Uncertainty Quantification in Deep Learning.

Accu-

rate quantification of predictive uncertainty is essential for development of trustworthy Deep Learning (DL) models. Deep ensemble techniques (Pearce et al., 2020; Lakshmi- narayanan et al., 2017) have been developed for uncer- tainty quantification. An ensemble of neural networks is constructed and the agreement/disagreement across the en- semble components is used to quantify different uncertain- ties. Ensemble-based methods significantly increase the number of model parameters, which are computationally expensive at both training and test times. Alternatively, Bayesian neural networks (Gal & Ghahramani, 2016)(Blun- dell et al., 2015)(Mobiny et al., 2021) have been devel- oped that consider a Bayesian formalism to quantify dif- ferent uncertainties. For instance, (Blundell et al., 2015) use Bayes-by-backdrop to learn a distribution over neural network parameters, whereas (Gal & Ghahramani, 2016) enable dropout during inference phase to obtain predictive uncertainty. Bayesian methods resort to some form of ap- proximation to address the intractability issue in marginal- ization of latent variables. Moreover, these methods are also computationally expensive as they require sampling for uncertainty quantification.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

Evidential Deep Learning.

Evidential Models Introduce A

conjugate higher-order evidential prior for the likelihood dis- tribution that enables the model to capture the fine-grained uncertainties. For instance, Dirichlet prior is introduced over the multinomial likelihood for evidential classification (Bao et al., 2021; Zhao et al., 2020), and NIG prior is in- troduced over the Gaussian likelihood (Amini et al., 2020; Pandey & Yu, 2022b) for the evidential regression models.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

Adversarial robustness (Kopetzki et al., 2021) and calibra- tion (Tomani & Buettner, 2021) of evidential models have also been well studied. Usually, these models are trained with evidential losses in conjunction with heuristic evidence regularization to guide the uncertainty behavior (Pandey & Yu, 2022a; Shi et al., 2020) in addition to reasonable gen- eralization performance. Some evidential models assume access to out-of-distribution data during training (Malinin & Gales, 2019; 2018) and use the OOD data to guide the un- certainty behavior. A recent survey (Ulmer, 2021) provides a thorough review of the evidential deep learning field.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

In this work, we focus on evidential classification models and consider settings where no OOD data is used during model training to make the proposed approach more broadly applicable to practical real-world situations.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

3.1. Preliminaries And Problem Setup

Standard classification models use a softmax transformation on the output from the neural network FΘ for input x to ob- tain the class probabilities in K-class classification problem.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

Such models are trained with the cross-entropy based loss.

2

Learn to Accumulate Evidence from All Training Samples: Theory and Practice For a given training sample (x, y), the loss is given by

(1)

where smk is the softmax output.

These Models Have

achieved state-of-the-art performance on many benchmark problems. A detailed gradient analysis shows that they can effectively learn from all training data samples (see Ap- pendix A). Nevertheless, these models lack a systematic mechanism to quantify different sources of uncertainty, a highly desired property in many real-world problems.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

Figure 3. Graphical model for Evidential Deep Learning Evidential classification models formulate training as an evidence acquisition process and consider a higher-order Dirichlet prior Dir(p|α) over the predictive Multino- mial distribution Mult(y|p). Different from a standard Bayesian formulation which optimizes Type II Maximum Likelihood to learn the Dirichlet hyperparameter (Bishop & Nasrabadi, 2006), evidential models directly predict α using data features x and then generate the prediction y by marginalizing the Multinomial parameter p. Figure 3 de- scribes this generative process. Such higher-order prior en- ables the model to systematically quantify different sources of uncertainty. In evidential models, the softmax layer of the standard neural networks is replaced by a non-negative

∀X ∈[−∞, ∞],

such that for input x, the neural network model FΘ with parameters Θ can output evidence e for different classes. Dirichlet prior α is evaluated as α = e+1 to ensure α ≥1.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

The trained evidential model outputs Dirichlet parameters α for input x that can quantify fine-grained uncertainties in addition to the prediction y. Mathematically, for K−class

(4)

The activation function A(·) assumes three common forms to transform the neural network output into evidence: (1)

Relu(·) = Max(0, ·), (2) Softplus(·) = Log(1 +

exp(·)), and (3) exp(·). Evidential models assign input sample to that class for which the output evidence is greatest. Moreover, they quantify the confidence in the prediction for K class classification prob- lem through vacuity ν (i.e., measure of lack of confidence

(5)

For any training sample (x, y), the evidential models aim to maximize the evidence for the correct class, minimize the evidence for the incorrect classes, and output accurate confi- dence. To this end, three variants of evidential loss functions have been proposed (Sensoy et al., 2018): 1) Bayes risk with sum of squares loss, 2) Bayes risk with cross-entropy loss, and 3) Type II Maximum Likelihood loss. Please refer to equations (21), (22), and (23) in the Appendix for the spe- cific forms of these losses. Additionally, incorrect evidence regularization terms are introduced to guide the model to output low evidence for classes other than the ground truth class (See Appendix C for discussion on the regularization).

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

With evidential training, accurate evidential deep learning models are expected to output high evidence for the correct class, low evidence for all other classes, and output very high vacuity for unseen/out-of-distribution samples.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

3.2. Theoretical Analysis of Learning Deficiency in

Evidential Learning

To identify the underlying reason that causes the perfor- mance gap of evidential models as described earlier, we consider a K class classification problem and a represen- tative evidential model trained using Bayes risk with sum of squares loss given in (21). We first provide an important definition that is critical for our theoretical analysis.

deep learning capstone project ideas github Diagram
Figure: System Model & Simulation Flow for Deep Learning Capstone Project Ideas Github

Definition 1 (Zero-Evidence Region). A Zero-evidence sample is a data sample for which the model outputs zero evidence for all classes. A region in the evidence space that contains zero-evidence samples is a zero-evidence region.

For a reasonable evidential model, novel data samples not yet seen during training, difficult data samples, and out-of- distribution samples should become zero-evidence samples.

Theorem 1. Given a training sample (x, y), if an evidential neural network outputs zero evidence e, then the gradients of the evidential loss evaluated on this training sample over the network parameters reduce to zero.

Proof. Consider an input x with one-hot ground truth label y. Let the ground truth class index be gt, i.e., ygt = 1, with corresponding Dirichlet parameter αgt, and y̸=gt = 0. Moreover, let o, e, and α represent the neural network output vector before applying the activation A, the evidence vector, and the Dirichlet parameters respectively.

3

Learn to Accumulate Evidence from All Training Samples: Theory and Practice Now, the gradient of the loss with respect to the neural network output can be computed using the chain rule:

(7)

Based on the actual form of A, we have three cases:

(8)

For a zero-evidence sample, the logits ok satisfy the rela-

= 0

Case II: SoftPlus(·) to transform logits to evidence

Sigmoid(Ok) →0 & ∂Ek

∂ok →0.

∂Ek

∂ok →0. Moreover, there is no term in the first part of the loss gradient in (7) to counterbalance these zero-approaching gradients.

So, for zero-evidence training samples, for any node k,

(11)

Since the gradient of the loss with respect to all the nodes is zero, there is no update to the model from such samples. This implies that the evidential models fail to learn from a zero-evidence data sample.

For completeness, we present the analysis of standard clas- sification models in Appendix A, detailed proof of the evi- dential models trained using Bayes risk with sum of squares error along with other evidential lossses in Appendix B, and impact of incorrect evidence regularization in Appendix C.

Evidential Models Can Not Learn From A Train-

ing sample that the model has never seen and for which the model accurately outputs “I don’t know”, i.e., ek =

0

∀k ∈[1, K]. Such samples are expected and likely to be present during model training. However, the supervised in- formation in such training data points is completely missed by evidential models so they fail to acquire any new knowl- edge from all such training data samples (i.e., data samples in zero-evidence region of the evidence space).

Corollary 1. Incorrect evidence regularization can not help evidential models learn from zero-evidence samples. Intuitively, the incorrect evidence regularization encourages the model to output zero evidence for all classes other than the ground truth class and the regularization does not have any impact on the evidence for the ground truth class. So, the regularization updates the model parameters such that the model is likely to map input samples closer to zero- evidence region in the evidence space. Thus, the regular- ization does not address the failure of evidential models to learn from zero evidence samples.

Theorem 2. For a data sample x, if an evidential model outputs logits ok ≤0 ∀k ∈[0, K], the exponential acti- vation function leads to a larger gradident update on the model parameters than softplus and ReLu.

Limited by space, we present the proof of Theorem 2 along with additional analysis in the Appendix D. The proof fol- lows the gradient analysis of the exponential, Softplus, and ReLU based models. It implies that the the training of evidential models is most effective with the exponential activation function. Intuitively, the ReLU based activation completely destroys all the information in the negative logits, and has largest region in evidence space in which training data have zero evidence. Softplus activation improves over the ReLU, and compared to ReLU, has smaller region in evidence space where training data have zero evidence.

However, Softplus based evidential models fail to cor- rect the acquired knowledge when the model has strong wrong evidence. Moreover, these models are likely to suf- fer from vanishing gradients problem when the number of classes increases (i.e., classification problem becomes more challenging). Finally, exponential activation has the smallest zero-evidence region in the evidence space without suffering from the issues of SoftPlus based evidential models.

Correct Evidence Regularization

We now consider an evidential model with exponential func- tion to transform the logits into evidence. We propose a novel vacuity-guided correct evidence regularization term

S Represents The Regularization Term

whose value is given by the magnitude of the vacuity output by the evidential model and αgt −1 represents the predicted evidence for the ground truth class. The regularization term λcor determines the relative importance of the correct

4

Learn to Accumulate Evidence from All Training Samples: Theory and Practice evidence regularization term compared to the evidential loss and incorrect evidence regularization and is treated as constant during model parameter update.

Theorem 3. Correct evidence regularization Lcor(x, y) can address the issue of learning from zero-evidence train- ing samples.

Proof. The proposed regularization term Lcor(x, y) does not contain any evidence terms other than the evidence for the ground truth node. So, the gradient of the regularization for nodes other than the ground truth node will be 0 i.e.

K̸=Gt = 0 And There Will Be No Update On These

nodes. For the ground truth node gt, ygt = 1, the gradient

(15)

The gradient value equals the magnitude of the vacuity. The vacuity is bounded in the range [0, 1], and zero-evidence sample, the vacuity is maximum, leading to the greatest

= −1. In Other Words, The Reg-

ularization encourages the model to update the parameters such that the correct evidence αgt −1 increases. As the model evidence increases, the vacuity decreases, and the contribution of the regularization Lcor(x, y) is minimized.

Thus, the proposed regularization enables the evidential model to learn from zero-evidence samples.

4.1. Evidential Model Training

We formulate an overall objective used to train the pro- posed Regularized evidential model (RED). Essentially, the evidential model is trained to maximize the correct evi- dence, minimize the incorrect evidence, and avoid the zero- evidence region during training. The overall loss is

(16)

where Levid(x, y) is the loss based on the evidential framework given by (21), (23), or (22) (See Appendix B), Linc(x, y) represents the incorrect evidence regularization (See Appendix Section C), Lcor(x, y) represents the pro- posed novel correct evidence regularization term in (12), and η1 = λ1 × min(1.0, epoch index/10) controls the impact of incorrect evidence regularization to the overall model training. In this work, we consider the forward-KL based incorrect evidence regularization given in (42) based on (Sensoy et al., 2018).

4.2. Evidence Space Visualization

Figure 4. Evidence space visualization to demonstrate the effec- tiveness of the proposed method. Figure 2 visualizes the evidence space in ReLU-based ev- idential models by considering the pre-ReLU output in a binary classification setting. Ideally, all samples that belong to Class 1 should be mapped to the blue region (region of high evidence for Class 1, low evidence for all other classes), all samples that belong to Class 2 should be mapped to the red region, and all out-of distribution samples should be mapped to the zero-evidence region (no evidence for all classes). To realize this goal, the models are trained using the evidential loss Levid with incorrect evidence regular- ization Linc. However, there is no update to the evidential model from such samples of zero-evidence region. Model’s prior belief of “I don’t know” for such samples does not get updated even after being exposed to the true label. For the samples with high incorrect evidence and low correct evidence, evidential model aims to correct itself. However, many such samples are likely to get mapped to the zero- evidence region (as shown by blue and orange arrows in Figure 2) after which there is no update to the model. Such fundamental limitation holds true for all evidential models.

The evidence space visualization for RED is shown in Figure 4 to illustrate how it addresses the above limitation. Cor- rect evidence regularization (indicated by green arrows) is weighted by the magnitude of the vacuity and is maximum in the zero-evidence region. In this problematic region, the proposed regularization fully dominates the model update as there is no update to the model from the two loss com- ponents (Levid and Linc) in (16). As the sample gets far away from the zero evidence region, the vacuity decreases proportionally, the impact of the proposed regularization to model update becomes insignificant, and the evidential losses (Levid & Linc) guide the model training. In this way, RED can effectively learn from all training samples irrespective of the model’s existing evidence.

5

Learn to Accumulate Evidence from All Training Samples: Theory and Practice

5. Experiments

Datasets and setup.

We Consider The Standard Supervised

classification problem with MNIST (LeCun, 1998), Ci- far10, and Cifar100 datasets (Krizhevsky et al., 2009), and few-shot classification with mini-ImageNet dataset (Vinyals et al., 2016). We employ the LeNet model for MNIST, ResNet18 model (He et al., 2016) for Cifar10/Cifar100, and ResNet12 model (He et al., 2016) for mini-ImageNet.

We first conduct experiments to demonstrate the learning deficiency of existing evidential models to confirm our the- oretical findings. We then evaluate the proposed correct evidence regularization to show its effectiveness. We finally conduct ablation studies to investigate the impact of evi- dential losses on model generalization and the uncertainty quantification of the proposed evidential model. Limited by space, additional clarifications, experiment results includ- ing few-shot classification experiments, experiments over challenging tiny-Imagenet datasett with Swin Transformer, hyperparameter details, and discussions are presented in the Appendix.

5.1. Learning Deficiency Of Evidential Models

Sensitivity to the change of the architecture.

We First

consider a toy illustrative experiment with two frameworks: 1) standard softmax, 2) evidential learning, and experiment with the LeNet (LeCun et al., 1999) model considered in EDL (Sensoy et al., 2018) with a minor modification to the architecture: no dropout in the model. To construct the toy dataset, we randomly select 4 labeled data points from the MNIST training dataset as shown in the Figure 5. For the evidential model, we use ReLU to transform the network outputs to evidence, and train the model with MSE-based evidential loss (Sensoy et al., 2018) given in (21) without incorrect evidence regularization. We train both models using only these 4 training data points.

Figure 6 compares the training accuracy and training loss trends of the evidential model with the standard softmax model (trained with the cross-entropy loss). Before any training, both models have 0% accuracy and the loss is high as expected. For the evidential model, in the first few iter- ations, the model learns from the training dataset, and the model’s accuracy increases to 50%. Afterward, the eviden- tial model fails to learn as the evidential model maps two of the training data samples to the zero-evidence region. Even in such a trivial setting, the evidential model fails to fit the 4 training data points showing their learning deficiency that empirically verifies the conclusion in Theorem 1. It is also worth noting that the range of the evidential model’s loss is significantly smaller than the standard model. This is mainly due to the bounded nature of the evidential MSE loss(i.e., it is bounded in the range [0, 2]) (a detailed theoretical analy- sis of the evidential losses is provided in the Appendix). In contrast, the standard model trained with cross-entropy loss

: 6

Figure 5. Toy dataset with 4 data points.

(B) Training Loss Trend

Figure 6. Training of standard and evidential models easily fits the trivial dataset, obtains near 0 loss, and perfect accuracy of 100% after a few iterations of training.

6

Figure 7. Zero-evidence trend during model training Additionally, we visualize the zero-evidence data samples for the toy dataset setting. We plot the total evidence for each training sample as training progresses for the first 100 iterations. The total evidence trend as training progresses for the first 100 iterations is shown in Figure 7. The ev- idential model’s predictions are correct for data samples with ground truth labels of 3 and 6, and incorrect for the remaining two data samples. After few iterations of training, the remaining two samples have zero total evidence (i.e.

samples are mapped to zero evidence region), the model never learns from them, and the model only achieves overall 50% training accuracy even after 100 iterations. Clearly, the evidential model continues to output zero evidence for two of the training examples and fails to learn from them.

Such learning deficiency of evidential models limits their extension to challenging settings. In contrast, the standard model easily overfits the 4 training examples and achieves 100% accuracy.

Sensitivity to hyperparameter tuning.

In This Experi-

ment, evidential models are trained using evidential losses given in (21), (22), or (23) with incorrect evidence regular- ization to guide the model for accurate uncertainty quan-

6

Learn to Accumulate Evidence from All Training Samples: Theory and Practice Figure 8. Impact of different incorrect evidence regularization strengths to the test set accuracy on Cifar100 dataset tification. We study the impact of the incorrect evidence regularization λ1 to the evidential model’s performance using Cifar100. The result shows that the generalization performance of evidential models is highly sensitive to λ1 values. To illustrate, we consider the Type II Maximum Likelihood loss in (23) with different λ1 to control KL reg- ularization (results on other loss functions are presented in the Appendix). As shown in Figure 8, when some regular- ization is introduced, evidential model’s test performance improves slightly. However, when strong regularization is used, the model focuses strongly on minimizing the incor- rect evidence. Such regularization causes the model to push many training samples into or close to the zero-evidence regions, which hurts the model’s learning capabilities. In contrast, the proposed model can continue to learn from samples in zero-evidence regions, which shows its robust- ness to incorrect evidence regularization. Moreover, our model has stable performance across all hyperparameter settings as it can effectively learn from all training samples.

Challenging datasets and settings.

We Next Consider

standard classification models for the Cifar100 dataset and 1-shot classification with the mini-ImageNet dataset. We develop evidential extensions of the classification models using Type II Maximum Likelihood loss given in (23) with- out any incorrect evidence regularization and use ReLU to transform logits to evidence. As shown in Figure 10, com- pared to the standard classification model, the evidential model’s predictive performance is sub-optimal (almost 20% lower for both classification problems). This is mainly due to the fact that evidential model maps many of the training data points to zero-evidence region, which is equivalent to the model saying “I don’t know to which class this sample belongs” and stopping to learn from them. Consequently, the model fails to acquire new knowledge (i.e., update itself), even after being exposed to correct supervision (the label information). In these cases, instead of learning, the eviden- tial model chooses to ignore the training data on which it does not have any evidence and remains to be ignorant.

Visualization of zero-evidence samples.

We Next Show

the 2-dimensional visualization of the latent representation for the randomly selected 500 training examples based on

(B) 1-Shot Results

Figure 10. Learning trends in complex classification problems the tSNE plot for ReLU based evidential model trained on the Cifar100 dataset with λ1 = 0.1. Figure 9 plot visualizes the latent embedding of zero evidence (Zero E) training sam- ples with non-zero evidence (Non-Zero E) training samples.

As can be seen, both zero and non-zero evidence samples ap- pear to be dispersed, overlap at different regions, and cover a large area in the embedding space. This further confirms the challenge of effectively learning from these samples

5.2. Effectiveness Of The Red

Evidential activation function.

We First Experiment With

different activation functions for the evidential models to show the superior predictive performance and generalization capability of exp activation validating our Theorem 2. We consider evidential models trained with evidential log loss given by (23) in Table 1 (Additional results along with hy- perparameter details are presented in Appendix Section F).

As can be seen, exp activation to transform network outputs into evidence leads to superior performance compared to ReLU and Softplus based transformations. Furthermore, our proposed model with correct evidence regularization further improves over the exp-based evidential models as it enables the evidential model to continue learning from zero-evidence samples.

76.43±0.21

We next present the test set performance change as training

7

Learn to Accumulate Evidence from All Training Samples: Theory and Practice progresses with MNIST dataset and two different evidential losses in Figure 11 where we observe similar results. The exp activation shows superior performance, as it has small- est zero-evidence region, and does not suffer from many learning issues present in other activation functions.

(B) Evidential Log Loss

Figure 11. Impact of evidential activation functions to the Test

Accuracy

Correct evidence regularization.

We Now Study The Im-

pact of the proposed correct evidence regularization using the MNIST and Cifar100 classification problems. We con- sider the evidential baseline model that uses exp activation to acquire evidence, and is trained with Type II Maximum Likelihood based loss with different incorrect evidence reg- ularization strengths. We introduce the proposed novel cor- rect evidence regularization to the model. As can be seen in Figure 12, the model with correct-evidence regularization has superior generalization performance compared to the baseline evidential model. This is mainly due to the fact that with proposed correct evidence regularization, the evi- dential model can also learn from the zero-evidence training samples to acquire new knowledge instead of ignoring them.

Our proposed model considers knowledge from all the train- ing data and aims to acquire new knowledge to improve its generalization instead of ignoring the samples on which it has no knowledge. Finally, even though strong incorrect evidence regularization hurts the model’s generalization, the proposed model is robust and generalizes better, empirically validating our Theorem 3. Limited by space, we present additional results in Appendix F.3.2.

Zero-evidence Sample Anaysis.

Similar To The Toy

MNIST zero-evidence analysis, we consider the Cifar100 dataset, and carry out the analysis for this complex dataset/setting. Instead of focusing on a few training ex- amples, we present the average statistics of the evidence (E) for the 50,000 training samples in the 100 class classi- fication problem for a model trained for 200 epochs using a log-based evidential loss in (23) with λ1 = 1.0. For ref- erence, the samples with less than 0.01 average evidence (i.e., E ≤0.01) are samples on which the model is not confident (i.e., having a high vacuity of ν ≥0.99), and are close to the ideal zero-evidence region. Our proposed RED model effectively avoids such zero evidence regions, and has the lowest number of samples (i.e. only 0.06% of total training dataset compared to 58.96% of SoftPlus based,

(D) Trend For Λ1 = 1.0

Figure 12. Impact of correct evidence regularization to test accu- racy: (a), (b) - MNIST Results; (c), (d) - Cifar100 Results and 100% of ReLU based evidential models) in very low evidence regions.

Table 2. Zero-Evidence Analysis for Complex Dataset-Setting

5.3. Ablation Study

Impact of loss function.

We Next Study The Impact Of

the evidential loss function on the model’s performance using MNIST and CIFAR100 classification problems. We consider all three activations: ReLU, SoftPlus, and exp to transform neural network outputs to evidence and carry out experiments over CIFAR100 with identical model and settings. As seen in Table 3, the generalization performance of evidential model is consistently sub-optimal when trained with evidential MSE loss given by (21) compared to the two other evidential losses (22) & (23). This is consistent across all three evidence activation functions. This is mainly due to the bounded nature of the evidential MSE loss (21): for all training samples, evidential MSE loss is bounded in the range of [0, 2]. Type II Maximum Likelihood loss given in (23) and cross-entropy based evidential loss given in (22) show comparable empirical results.

Next, we consider exp activation and conduct experiments over the MNIST dataset for incorrect evidence regulariza- tion strengths of λ1 = 0&1. We again observe similar results where the training with the Evidential MSE loss in (21) leads to sub-optimal test performance. Additional re- sults, along with theoretical analysis are presented in the Appendix. In the subsequent experiments, we consider the Type II Maximum Likelihood loss (23) for evidential model training due to its simplicity and some theoretical advan-

8

Learn to Accumulate Evidence from All Training Samples: Theory and Practice tages (see Appendix E). We leave a thorough investigation of these two evidential losses ((22) & (23)) as future work.

Table 3. Impact of evidential losses on classification performance

(B) Trend For Λ1 = 1.0

Figure 13. Impact of evidential losses on test set accuracy

Figure 14. Accuracy-Vacuity Curve

Study of uncertainty information.

We Now Investigate

the uncertainty behavior of the proposed evidential model with Cifar100 experiments.

We Present The Accuracy-

Vacuity curve for different incorrect evidence regulariza- tion strengths (λ1) in Figure 14. Vacuity reflects the lack of confidence in the predictions, and the accuracy of effec- tive evidential model should increase with lower vacuity threshold. Without any incorrect evidence regularization (i.e., λ1 = 0), the evidential model is highly confident on its predictions and all test samples are concentrated on the low vacuity region. As the incorrect evidence regularization strength is increased, the model outputs more accurate confi- dence in the predictions. Strong incorrect evidence regular- ization hurts the generalization over the test set as indicated by low accuracy when all test samples are considered (i.e., vacuity threshold of 1.0). In all cases, the evidential model shows reasonable uncertainty behavior: the model’s test set accuracy increases as the vacuity threshold is decreased.

Next, we look at the accuracy of the evidential models on their top-K % most confident predictions over the test set. Table 4 shows the accuracy trend of Top-K (%) confident samples. Consider the most confident 20% samples (cor- responding to 2000 test samples of Cifar100 dataset). The proposed model leads to highest accuracy (of 99.35%) com- pared to all the models. Similar trend is seen for different K values where the proposed model shows comparable to superior results demonstrating its accurate uncertainty quantification capability.

76.43

We next consider out-of-distribution (OOD) detection ex- periments for the Cifar100-trained evidential model using SVHN dataset (as OOD) (Netzer et al., 2011). As seen in Table 5, the evidential models, on average, output very high vacuity for the OOD samples, showing the potential for OOD detection.

0.7552

We present the AUROC score for Cifar100 trained models with SVHN dataset test set as the OOD samples in Table 6. In AUROC calculation, we use the maximum softmax score for the standard model, and predicted vacuity score for all the evidential models. As can be seen, the exp-based model outperforms all other activation functions, and the proposed model RED can learn from all the training samples that leads to the best performance.

6. Conclusion

In this paper, we theoretically investigate the evidential mod- els to identify their learning deficiency, which makes them fail to learn from zero-evidence regions. We then show the superiority of the evidential model with exp evidential activation over the ReLU and SoftPlus based models.

We further analyze the evidential losses, and introduce a novel correct evidence regularization over the exp-based ev- idential model. The proposed model effectively pushes the training samples out of the zero-evidence regions, leading to superior learning capabilities. We conduct extensive experi- ments that empirically validate all theoretical claims while demonstrating the effectiveness of the proposed approach.

Acknowledgements

This research was supported in part by an NSF IIS award IIS-1814450 and an ONR award N00014-18-1-2875. The views and conclusions contained in this paper are those of the authors and should not be interpreted as representing any funding agency.

9

Learn to Accumulate Evidence from All Training Samples: Theory and Practice

Ai Adaptive Learning

This project focuses on ai adaptive learning using modern AI and machine learning techniques. The content below is adapted from research literature and practical implementation notes.

We propose a novel high-performance and interpretable canon-

addition, unlike tree learning, DNNs enable gradient descent- ical deep tabular data learning architecture, TabNet. TabNet based end-to-end learning for tabular data which can have a uses sequential attention to choose which features to reason multitude of benefits: (i) efficiently encoding multiple data from at each decision step, enabling interpretability and more types like images along with tabular data; (ii) alleviating the efficient learning as the learning capacity is used for the most need for feature engineering, which is currently a key aspect

salient features. We demonstrate that TabNet outperforms in tree-based tabular data learning methods; (iii) learning other variants on a wide range of non-performance-saturated from streaming data and perhaps most importantly (iv) end- tabular datasets and yields interpretable feature attributions to-end models allow representation learning which enables plus insights into its global behavior. Finally, we demonstrate many valuable application scenarios including data-efficient

self-supervised learning for tabular data, significantly improv- domain adaptation (Goodfellow, Bengio, and Courville 2016), ing performance when unlabeled data is abundant. generative modeling (Radford, Metz, and Chintala 2015) and

Introduction We propose a new canonical DNN architecture for tabular

Deep neural networks (DNNs) have shown notable success data, TabNet. The main contributions are summarized as: efficiently encode the raw data into meaningful representa- enabling flexible integration into end-to-end learning. tions, fuel the rapid progress. One data type that has yet to 2. TabNet uses sequential attention to choose which fea- see such success with a canonical architecture is tabular data. tures to reason from at each decision step, enabling in-

Despite being the most common data type in real-world AI terpretability and better learning as the learning capacity (as it is comprised of any categorical and numerical features), is used for the most salient features (see Fig. 1). This under-explored, with variants of ensemble decision trees for each input, and unlike other instance-wise feature se- Why? First, because DT-based approaches have certain bene- and van der Schaar 2019), TabNet employs a single deep

fits: (i) they are representionally efficient for decision mani- learning architecture for feature selection and reasoning. folds with approximately hyperplane boundaries which are 3. Above design choices lead to two valuable properties: (i) common in tabular data; and (ii) they are highly interpretable TabNet outperforms or is on par with other tabular learn- in their basic form (e.g. by tracking decision nodes) and there ing models on various datasets for classification and re-

are popular post-hoc explainability methods for their ensem- gression problems from different domains; and (ii) TabNet ble form, e.g. (Lundberg, Erion, and Lee 2018) – this is an enables two kinds of interpretability: local interpretability important concern in many real-world applications; (iii) they that visualizes the importance of features and how they are fast to train. Second, because previously-proposed DNN are combined, and global interpretability which quantifies

architectures are not well-suited for tabular data: e.g. stacked the contribution of each feature to the trained model. convolutional layers or multi-layer perceptrons (MLPs) are 4. Finally, for the first time for tabular data, we show signif- vastly overparametrized – the lack of appropriate inductive icant performance improvements by using unsupervised bias often causes them to fail to find optimal solutions for tab- pre-training to predict masked features (see Fig. 2).

ular decision manifolds (Goodfellow, Bengio, and Courville

Why is deep learning worth exploring for tabular data?

One obvious motivation is expected performance improve- Feature selection: Feature selection broadly refers to judi- Copyright © 2021, Association for the Advancement of Artificial ciously picking a subset of features based on their useful-

Professional occupation related Investment related

Feedback from Feedback to

Feature selection Input processing Feature selection Input processing

previous step next step … …

Predicted output (whether the income level >$50k)

selection enables interpretability and better learning as the capacity is used for the most salient features. TabNet employs multiple decision blocks that focus on processing a subset of input features for reasoning. Two decision blocks shown as examples process features that are related to professional occupation and investments, respectively, in order to predict the income level.

Unsupervised pre-training Supervised fine-tuning

Age Cap. gain Education Occupation Gender Relationship Age Cap. gain Education Occupation Gender Relationship 5 2000 ? Exec-managerial F Wife 6 2000 Bachelors Exec-managerial M Husband 1 0 ? Farming-fishing M ? 2 0 High-school Farming-fishing M Unmarried

? 50 Doctorate Prof-specialty M Husband 4 50 Doctorate Prof-specialty M Husband 2 ? ? Handlers-cleaners F Wife 2 0 High-school Handlers-cleaners F Wife 5 3000 Bachelors ? ? Husband 5 3000 Bachelors Exec-managerial M Husband

3 0 Bachelors ? F ? 3 100 Bachelors Prof-specialty F Wife ? 0 High-school Armed-Forces ? Husband 2 0 High-school Armed-Forces M Husband

TabNet decoder Decision making

Age Cap. gain Education Occupation Gender Relationship Income > $50k

3 M False

level can be guessed from the occupation, or the gender can be guessed from the relationship. Unsupervised representation learning by masked self-supervised learning results in an improved encoder model for the supervised learning task.

ward selection and Lasso regularization (Guyon and Elisseeff performance with compact representations. 2003) attribute feature importance based on the entire training Tree-based learning: DTs are commonly-used for tabular data, and are referred as global methods. Instance-wise fea- data learning. Their prominent strength is efficient picking ture selection refers to picking features individually for each of global features with the most statistical information gain

to maximize the mutual information between the selected mance of standard DTs, one common approach is ensembling features and the response variable, and in (Yoon, Jordon, and to reduce variance. Among ensembling methods, random van der Schaar 2019) by using an actor-critic framework to forests (Ho 1998) use random subsets of data with randomly mimic a baseline while optimizing the selection. Unlike these, selected features to grow many trees. XGBoost (Chen and

sity in end-to-end learning – a single model jointly performs recent ensemble DT approaches that dominate most of the feature selection and output mapping, resulting in superior recent data science competitions. Our experimental results

!# + Softmax !" < % !" > % !# > & !# > &

ReLU ReLU &

$" !" − $" % −1 −$" !" + $" % −1 −1 $# !# − $# & % −1 −$# !# + $# & !"

FC FC

W: [$" , - $" , 0, 0] W: [0, 0, $# , - $# ] !" < % b: [-a $" , a $" , -1, -1] b: [-1, -1, -d $# , d $# ] !# < & !" > % !# < & [!" ] [!# ]

M: [1, 0] M: [0, 1]

(right). Relevant features are selected by using multiplicative sparse masks on inputs. The selected features are linearly transformed, and after a bias addition (to represent boundaries) ReLU performs region selection by zeroing the regions. Aggregation of multiple regions is based on addition. As C and C get larger, the decision boundary gets sharper.

for various datasets show that tree-based models can be out- constructs a sequential multi-step architecture, where each performed when the representation capacity is improved with step contributes to a portion of the decision based on the deep learning while retaining their feature selecting property. selected features; (iii) improves the learning capacity via non- Integration of DNNs into DTs: Representing DTs with linear processing of the selected features; and (iv) mimics

DNN building blocks as in (Humbird, Peterson, and McClar- ensembling via higher dimensions and more steps. ren 2018) yields redundancy in representation and ineffi- cient learning. Soft (neural) DTs (Wang, Aggarwal, and Liu Fig. 4 shows the TabNet architecture for encoding tabu- functions, instead of non-differentiable axis-aligned splits. mapping of categorical features with trainable embeddings.

However, losing automatic feature selection often degrades We do not consider any global feature normalization, but performance. In (Yang, Morillo, and Hospedales 2018), a soft merely apply batch normalization (BN). We pass the same D- binning function is proposed to simulate DTs in DNNs, by dimensional features f ∈ <B×D to each decision step, where 2019) proposes a DNN architecture by explicitly leveraging multi-step processing with Nsteps decision steps. The ith

expressive feature combinations, however, learning is based step inputs the processed information from the (i − 1)th step on transferring knowledge from gradient-boosted DT. (Tanno to decide which features to use and outputs the processed ing from primitive blocks while representation learning into sion. The idea of top-down attention in the sequential form edges, routing functions and leaf nodes. TabNet differs from is inspired by its applications in processing visual and text

these as it embeds soft feature selection with controllable data (Hudson and Manning 2018) and reinforcement learn- Self-supervised learning: Unsupervised representation relevant information in high dimensional input. learning improves supervised learning especially in small Feature selection: We employ a learnable mask M[i] ∈ has shown significant advances – driven by the judicious capacity of a decision step is not wasted on irrelevant

choice of the unsupervised learning objective (masked input ones, and thus the model becomes more parameter effi- prediction) and attention-based deep learning. cient. The masking is multiplicative, M[i] · f . We use an attentive transformer (see Fig. 4) to obtain the masks us- TabNet for Tabular Learning ing the processed features from the preceding step, a[i − 1]:

M[i] = sparsemax(P[i − 1] · hi (a[i − 1])). Sparsemax nor-

DTs are successful for learning from real-world tabular malization (Martins and Astudillo 2016) encourages sparsity datasets. With a specific design, conventional DNN building by mapping the Euclidean projection onto the probabilistic blocks can be used to implement DT-like output manifold, simplex, which is observed to be superior in performance and e.g. see Fig. 3). In such a design, individual feature selec- aligned with the goal of sparse feature selection for explain-

tion is key to obtain decision boundaries in hyperplane form, PD which can be generalized to a linear combination of features ability. Note that j=1 M[i]b,j = 1. hi is a trainable func- where coefficients determine the proportion of each feature. tion, shown in Fig. 4 using a FC layer, followed by BN. P[i] TabNet is based on such functionality and it outperforms DTs is the prior scale term, denoting how much a particular feature

Qi while reaping their benefits by careful design which: (i) uses has been used previously: P[i] = j=1 (γ − M[j]), where γ sparse instance-wise feature selection learned from data; (ii) is a relaxation parameter – when γ = 1, a feature is enforced

+ Softmax

Feature Feature …

transformer transformer

x Nsteps Features

+ Softmax

Feature …

transformer transformer Feature Feature Feature Feature transformer

Encoded representation

transformer transformer Attentive transformer … Mask transformer …

Step 2 Decision step dependent

transformer transformer

BN Feature Feature

FC BN transformer transformer

+ 0.5 0.5 0.5 Agg. Agg. Features Features FC FC + +

Reconstructed + … Feature attributes + … features

(a) TabNet encoder architecture (b) TabNet decoder architecture Feature transformer Feature Attentive transformer Shared across decision steps Decision step dependent transformer GLU

Decision step dependent Prior scales

+ 0.5 0.5 0.5

0.5 0.5 0.5

+ Attentive transformer (c) (d)

Prior scales

divides the processed representation to be used by the attentive transformer of the subsequent step as well as for the overall Attentive BN FC

output. For each step, the feature selection mask provides interpretable information about the model’s functionality, and the +

masks can be aggregated to obtain global feature transformer important attribution. (b) TabNet decoder, composed of a feature transformer block at each step. (c) A feature transformer block example – 4-layer network is shown, where 2 are shared across all decision

Prior scales

steps and 2 are decision step-dependent. Each layer is composed of a fully-connected (FC) layer, BN and GLU nonlinearity. (d) +

An attentive transformer block example – a single layer mapping is modulated with a prior scale information which aggregates Sparsemax

how much each feature has been used before the current decision step. sparsemax (Martins and Astudillo 2016) is used for BN FC

normalization of the coefficients, resulting in sparse selection of the salient features. +

to be used only at one decision step and as γ increases, more propose the aggregate.feature importance mask, Magg−b,j = flexibility is provided to use a feature at multiple decision PNsteps ηb [i]Mb,j [i]

PD PNsteps

ηb [i]Mb,j [i].2 i=1 i=1 steps. P is initialized as all ones, 1B×D , without any prior j=1

on the masked features. If some features are unused (as in self- Tabular self-supervised learning: We propose a decoder supervised learning), corresponding P entries are made 0 architecture to reconstruct tabular features from the Tab- to help model’s learning. To further control the sparsity of the Net encoded representations. The decoder is composed of selected features, we propose sparsity regularization in the feature transformers, followed by FC layers at each deci-

form of entropy (Grandvalet and Bengio 2004), Lsparse = sion step. The outputs are summed to obtain the recon-

PNsteps PB PD −Mb,j [i] log(Mb,j [i]+)

i=1 b=1 j=1 Nsteps ·B , where  is a structed features. We propose the task of prediction of miss- small number for numerical stability. We add the sparsity reg- ing feature columns from the others. Consider a binary mask ularization to the overall loss, with a coefficient λsparse . Spar- S ∈ {0, 1}B×D . The TabNet encoder inputs (1 − S) · f̂ sity provides a favorable inductive bias for datasets where and the decoder outputs the reconstructed features, S · f̂ . We

most features are redundant. initialize P = (1 − S) in encoder so that the model em- Feature processing: We process the filtered features using phasizes merely on the known features, and the decoder’s last a feature transformer (see Fig. 4) and then split for the FC layer is multiplied with S to output the unknown features. decision step output and information for the subsequent We consider the reconstruction loss in self-supervised phase:

step, [d[i], a[i]] = fi (M[i] · f ), where d[i] ∈ <B×Nd and 2

PB PD (f̂b,j −fb,j )·Sb,j

a[i] ∈ <B×Na . For parameter-efficient and robust learning b=1 j=1

√ PB PB 2

. Normalization b=1 (fb,j −1/B b=1 fb,j ) with high capacity, a feature transformer should comprise layers that are shared across all decision steps (as the same with the population standard deviation of the ground truth features are input across different decision steps), as well as is beneficial, as the features may have different ranges. We decision step-dependent layers. Fig. 4 shows the implementa- sample Sb,j independently from a Bernoulli distribution with

tion as concatenation of two shared layers and two decision parameter ps , at each iteration. step-dependent layers. Each FC layer is followed by BN and eventually connected to a normalized residual √ connection We study TabNet in wide range of problems, that contain with normalization. Normalization with 0.5 helps to sta- regression or classification tasks, particularly with published bilize learning by ensuring that the variance throughout the benchmarks. For all datasets, categorical inputs are mapped

For faster training, we use large batch sizes with BN. Thus, bedding and numerical columns are input without and pre- except the one applied to the input features, we use ghost BN processing.4 We use standard classification (softmax cross (Hoffer, Hubara, and Soudry 2017) form, using a virtual batch entropy) and regression (mean squared error) loss functions size BV and momentum mB . For the input features, we ob- and we train until convergence. Hyperparameters of the Tab-

serve the benefit of low-variance averaging and hence avoid Net models are optimized on a validation set and listed in ghost BN. Finally, inspired by decision-tree like aggregation Appendix. TabNet performance is not very sensitive to most as in Fig. 3, we construct the overall decision embedding hyperparameters as shown with ablation studies in Appendix. as dout = i=1 PNsteps ReLU(d[i]). We apply a linear mapping In Appendix, we also present ablation studies on various de-

Wfinal dout to get the output mapping.1 sign and guidelines on selection of the key hyperparameters. Interpretability: TabNet’s feature selection masks can shed For all experiments we cite, we use the same training, val- light on the selected features at each step. If Mb,j [i] = 0, idation and testing data split with the original work. Adam optimization algorithm (Kingma and Ba 2014) and Glorot then j th feature of the bth sample should have no contribution uniform initialization are used for training of all models.5

to the decision. If fi were a linear function, the coefficient

Mb,j [i] would correspond to the feature importance of fb,j . Instance-wise feature selection

Although each decision step employs non-linear processing, their outputs are combined later in a linear way. We aim Selection of the salient features is crucial for high perfor- to quantify an aggregate feature importance in addition to mance, especially for small datasets. We consider 6 tabular requires a coefficient that can weigh the relative importance samples). The datasets are constructed in such a way that of each step in the decision. We simply propose ηb [i] = only a subset of the features determine the output. For Syn1-

PNd Syn3, salient features are same for all instances (e.g., the

c=1 ReLU(db,c [i]) to denote the aggregate decision con- tribution at ith decision step for the bth sample. Intuitively, if 2

Normalization is used to ensure D

P j=1 Magg−b,j = 1. db,c [i] < 0, then all features at ith decision step should have 3

0 contribution to the overall decision. As its value increases, prove the performance, but interpretation of individual dimensions

it plays a higher role in the overall linear combination. Scal- may become challenging. ing the decision mask at each decision step with ηb [i], we Specially-designed feature engineering, e.g. logarithmic trans- formation of variables highly-skewed distributions, may further

For discrete outputs, we additionally employ softmax during

training (and argmax during inference). An open-source implementation will be released.

Global: using only globally-salient features, Tree Ensembles (Geurts, Ernst, and Wehenkel 2006), Lasso-regularized model, L2X

Syn Syn Syn Syn Syn Syn

No selection .5 ± .0 .7 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .6 ± .0 Tree .5 ± .1 .8 ± .0 .8 ± .0 .6 ± .0 .7 ± .0 .7 ± .0 Lasso-regularized .4 ± .0 .5 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .7 ± .0

INVASE .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0

Global .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0 TabNet .6 ± .0 .8 ± .0 .8 ± .0 .7 ± .0 .7 ± .0 .8 ± .0

output of Syn depends on features X -X ), and global fea- Table 3: Performance for Poker Hand induction dataset. ture selection, as if the salient features were known, would give high performance. For Syn4-Syn6, salient features are Model Test accuracy (%) instance dependent (e.g., for Syn4, the output depends on ei- DT 50.0 ther X -X or X -X depending on the value of X ), which MLP 50.0

makes global feature selection suboptimal. Table 1 shows that Deep neural DT 65.1

TabNet outperforms others (Tree Ensembles (Geurts, Ernst, XGBoost 71.1

and Wehenkel 2006), LASSO regularization, L2X (Chen LightGBM 70.0 van der Schaar 2019). For Syn1-Syn3, TabNet performance TabNet 99.2 is close to global feature selection - it can figure out what Rule-based 100.0 features are globally important. For Syn4-Syn6, eliminating instance-wise redundant features, TabNet improves global feature selection. All other methods utilize a predictive model Poker Hand (Dua and Graff 2017): The task is classifica-

with 43k parameters, and the total number of parameters is tion of the poker hand from the raw suit and rank attributes of 101k for INVASE due to the two other models in the actor- the cards. The input-output relationship is deterministic and critic framework. TabNet is a single architecture, and its size hand-crafted rules can get 100% accuracy. Yet, conventional is 26k for Syn1-Syn and 31k for Syn4-Syn6. The compact DNNs, DTs, and even their hybrid variant of deep neural DTs

representation is one of TabNet’s valuable properties. (Yang, Morillo, and Hospedales 2018) severely suffer from the imbalanced data and cannot learn the required sorting and Performance on real-world datasets ranking operations (Yang, Morillo, and Hospedales 2018).

Tuned XGBoost, CatBoost, and LightGBM show very slight

as it can perform highly-nonlinear processing with its depth, Model Test accuracy (%) without overfitting thanks to instance-wise feature selection.

CatBoost 85.1 Table 4: Performance on Sarcos dataset. Three TabNet mod-

AutoML Tables 94.9 els of different sizes are considered.

Forest Cover Type (Dua and Graff 2017): The task is clas- MLP 2.1 0.14M

sification of forest cover type from cartographic variables. Adaptive neural tree 1.2 0.60M approaches that are known to achieve solid performance (AutoML 2019), an automated search framework based on TabNet-M 0.2 0.59M ensemble of models including DNN, gradient boosted DT, TabNet-L 0.1 1.75M with very thorough hyperparameter search. A single TabNet without fine-grained hyperparameter search outperforms it. Sarcos (Vijayakumar and Schaal 2000): The task is re-

gressing inverse dynamics of an anthropomorphic robot arm.

very small model is possible with a random forest. In the very and TabNet merely focuses on the relevant ones. For Syn4, small model size regime, TabNet’s performance is on par the output depends on either X -X or X -X depending parameters. When the model size is not constrained, TabNet feature selection – it allocates a mask to focus on the indi- achieves almost an order of magnitude lower test MSE. cator X , and assigns almost all-zero weights to irrelevant

features (the ones other than two feature groups). models are denoted with -S and -M. Real-world datasets: We first consider the simple task of mushroom edibility prediction (Dua and Graff 2017). Tab- Model Test acc. (%) Model size Net achieves 100% test accuracy on this dataset. It is indeed Sparse evolutionary MLP 78.4 81K known (Dua and Graff 2017) that “Odor” is the most discrim-

What is this project about?

This project covers practical implementation and research aspects of the topic using AI/ML techniques.