Enquire Now
Medical Computer Vision · Clinical Diagnostics · PyTorch / TensorFlow · GPU Optimized · 2026

Credit Card Fraud Detection Machine Learning Capstone

Tensor Pipeline · Custom Loss Formulations · Model Quantization · Accelerated Inference — A rigorous deep learning engineering project focused on automated pathological lesion segmentation and radiological disease classification. Architected for thesis defense viva presentations, IEEE reproduction, and high-throughput production deployment.

PyTorch
Core Framework
AMP FP16
Mixed Precision
TensorRT
Quantized Serving

Credit Card Fraud Detection in the Nigerian Financial Sector: A Comparison of Unsupervised TensorFlow-Based Anomaly Detection

Abstract

Credit card fraud is a major cause of national concern in the Nigerian financial sector, affecting hundreds of transactions per second and impacting international e-commerce negatively. Despite the rapid spread and adoption of online marketing, millions of Nigerians are prevented from transacting in several countries with local credit cards due to bans and policies directed at restricting credit card fraud. Presently, a myriad of technologies exist to detect fraudulent transactions, a few of which are adopted by Nigerian financial institutions to proactively manage the situation. Fraud detection allows institutions to restrict offenders from networks and with a centralized banking identity management system, such as the Bank Verification Number used by the Central Bank of Nigeria, offenders who may have stolen other people’s identities can be back-traced and their bank accounts frozen. This paper aims to compare the effectiveness of two fraud detection technologies that are projected to work fully independent of human intervention to possibly predict and detect fraudulent credit card transactions. Autoencoders as an Unsupervised Tensorflow-Based Anomaly Detection Technique generally offers greater performance in dimensionality reduction than the Principal Component Analysis, and this theory was tested out on Nigerian credit card transaction data. Results demonstrate that autoencoders are better suited to analyzing complex and extensive datasets and offer more reliable results with minimal mislabeling than the PCA algorithm.

credit card fraud detection machine learning capstone Diagram
Figure: System Model & Simulation Flow for Credit Card Fraud Detection Machine Learning Capstone

1. Introduction

In recent years, the Central Bank of Nigeria has continuously enacted a wide range of proactive and highly restrictive policies to curb the rise of credit card fraud. From limiting international transaction amounts to $100 monthly and further reducing the limit to $20 per card, the country has struggled to mitigate the effects of one-click e-payment channels on transactional integrity.

credit card fraud detection machine learning capstone Diagram
Figure: System Model & Simulation Flow for Credit Card Fraud Detection Machine Learning Capstone

Credit card fraud is a relatively broad terms that describes a situation when unauthorized users gain access to an individual's credit card information in order to make purchases, other transactions, or move funds to another destination . Essentially, an individual commits credit fraud when they use another's identity and creditworthiness to obtain credit or purchase goods and services without the intention of repaying the debt . The unprecedented rise of financial technology platforms, otherwise known as fintech apps, is a leading cause of credit card fraud in the country as there are now more informal access channels with minimal connection to the country’s banking verification systems . Another significant enabler of the situation in the country is the use of Point-of-Sale machines, otherwise known as POS machines, where card transactions are performed with hand-held machines in the most unregulated environment areas with requirements for machine acquisition. Over 12.2 billion Naira was lost to fraudulent credit card activity in 2023 .

credit card fraud detection machine learning capstone Diagram
Figure: System Model & Simulation Flow for Credit Card Fraud Detection Machine Learning Capstone

Fig 1: Increasing number of POS terminals in Nigeria over a five-year timeline Research over the year show that credit card fraud is a global problem creating vacuums in the financial sectors of dozens of countries around the world [5, 6, 7]. Consequently, credit card fraud detection technologies have progressively constituted an essential research area of high interest to many scholars and experts. With the rapid spread and growing relevance of machine learning and big data analysis, more technologies, algorithms and detection methods are uncovered and fine-tuned by researchers across various levels of complexity.

The biggest challenge to research in credit card fraud detection in Nigeria is access to actionable, reliable, and uncompromised datasets for study. This data is most reliable when accessed from traditional financial institutions in the country. However, preliminary assessments show that while policies are enacted and technologies are deployed by this institution is variably efficient attempts to curb credit card fraud, limited efforts have been invested in documenting the instances or at the very least, making the collected data accessible to academic researchers for studies.

Another challenge is the potential for imbalance of data where the number of legitimate transactions is always much larger than the number of fraudulent transactions . This results in poor reflection of the problematic situation upon visualization. With machine learning algorithms, this imbalance also leads to a problem called the “curse of dimensionality”, characterized by a situation where the number of records are insufficient for the number of features . Over-fitting and extremely poor generalization are often the results of training with an inadequate or highly imbalanced dataset [10, 11]. Other times, a higher dimensionality to secure a larger number of fraudulent records can cause elongated training periods and ultimate burnout of the system in limited research budgets. Therefore, dimensionality reduction techniques are well-recommended for consolidating stable middle grounds in the iterations and visualizing reliable data.

This paper aims to explore the efficiency of unsupervised machine learning algorithms using Nigerian credit card fraud dataset. It explores autoencoders with code construction from first principles and the Principal Component Analysis (PCA) algorithm with an objective to establish autoencoders as the superior and more reliable fraud detection technique. This paper shows that PCA performs a learning process on a linear transformation which extends or transports the data into another spatial dimension, where vectors of projections are defined by variance of the data . Through a sub-process of restricting the dimensionality to a usually lower number of components that account for most of the variance of the data set, dimensionality reduction is eventually achieved. Autoencoders are neural networks trained using back propagation to reconstruct the input that can be used to reduce the data into a low dimensional latent space by stacking multiple non-linear transformations or layers . They are constructed an encoder- decoder architecture. The encoder maps the input to latent space and decoder reconstructs the input.

The main contributions of this paper are as follows:

-

To focus the ongoing research in the credit card fraud detection scene in the cumulative West African financial setting, specifically Nigeria.

-

To realign credit card fraud detection as an anomaly detection situation.

-

To build an autoencoders system from first principles using Python Programming Language and train it to accurately predict fraudulent transactions. Construct a PCA code base to label fraudulent truncations in a dataset.

-

To show how experimental results on a GitHub Dataset with a test set and training show that autoencoders outperforms the PCA algorithm and produces more reliable output for accurate fraud detection.

-

Python. Frameworks – TensorFlow, Keras, SciKitLearn

2. Literature Review Of Related Works

With the rise of global adaptation of credit cards, much so to the apogee of cashless economies thriving solely on the use of credit cards and the accompanying RFID chips, the risk of fraud and malignant truncations are also on the rise . Hundreds of scholars around the world have delved into in-depth research over the years; developing algorithms, technologies, and proposing mechanisms for efficient detection of credit card fraud in national, local, and sub-level financial systems.

Algorithms such as logistic Regression, Genetic Algorithm, Artificial Immune System Model, Naive Bayes, Random Forest, K Nearest Neighbor, Gradient Boosting, Support Vector Machines, neural network algorithm, and the PCA have been explored over the years by many researchers aiming to develop technologies and improve novel algorithms for credit card fraud detection [15, 16, 17]. In 2013, Patel et al proposed a system for credit card fraud detection using Genetic Algorithm (GA) . The genetic algorithm is an optimization technique that mimics the natural evolution processes to select the best characteristics for adaptation a higher likelihood for survival and reproduction. The algorithm used in this study was aimed at obtaining better solutions as time progresses. It operated on the principle of limit dependency, noting that when a card is copied or stolen or lost and captured by fraudsters it is usually used until its available limit is depleted. Thus, rather than the number of correctly classified transactions, a solution which minimizes the total available limit on cards subject to fraud is more prominent.

The main objective was to significantly mitigate the occurrence of false alerts using genetic algorithm where a set of interval valued parameters are optimized. In 2014, Halaviee et al developed a novel model for credit card fraud detection based on the translating the AIS processes into machine learning, simulating natural immune system functionality to protect and prevent anomalies from occurring in systems . This machine learning algorithm worked off the principle of immune system adaptability in humans, where the AIS addresses detecting non-self-cells, imitating the functions of human body which occurs during generating detector cells, detecting non-self-cells, and cleaning the body from non-self while learning its pattern. The algorithm aimed to develop a system where unwanted signals are immediately detected upon entrance or infiltration of the system. The final developed system attempted to improve a previously introduced algorithm in several categories to attain higher precision. It finally proposed a novel implementation model for the method in order to reduce training time.

In 2019, Fiore et al explored the use of generative adversarial networks (GAN) to improvise effectiveness in the classification process of credit card fraud detection . A GAN consists of two feed-forward neural networks, a Generator G and a Discriminator D working in an adversarial setting with each other, with G producing new elements and its opponent D instantly evaluating their usefulness and authenticity. This study aimed to generate a large number of reliable and useful examples of the minority class that can be used to re-balance the training sets used by the binary classifier. A careful experimental evaluation showed that a classifier trained on the augmented set largely outperforms the same classifier trained on the original data, especially as far as sensitivity is concerned, resulting in a really effective credit card fraud detection solution. While the framework was developed in the context of credit card fraud detection, the study noted that it could be useful when extended to other application domains.

In 2020, Madhav et al proposed a system based on the PCA and K-Means for the analysis of credit card fraud data . PCA is used for extracting multiple features simultaneously. Firstly, PCA is applied on the fraud dataset. Then PCA creates new features from original features. By making use of these new features, the relation between multiple attributes or associated attributes can be represented in the form of graphs for analysis purpose. The KMeans algorithm is applied on the features which is used to identify the similarity index of associated attributes.

The Random Forest algorithm is one of the most widely researched fraud detection algorithms in the past two decades . It is also the basis of many patented technologies and commercialized products for large-scale detection. In many cases, it is supported by other algorithms for increased functionality and faster runtimes in enterprise-level operations. In 2022, Devi at al Combined RFA technology with the PCA algorithm to develop a web-based technology for inputting a dataset and detecting fraudulent transactions . In 2023, Afiriye et al presented a study which compared three classification and prediction techniques, Decision Tree, Logistic Regression, and Random Forest, conclusively describing the Random Forest Algorithm as the most suitable supervised learning technique for credit card fraud detection . The researchers balanced the dataset prior to generating the models using the under-sampling technique, to ensure that the model does not favor solely the majority class and prevent over fitting the model to the data. With an AUC value of 98.9% and an accuracy value of 96.0%, the Random Forest model performed better than the other two models, making it the most suitable model for predicting fraudulent transactions.

In 2023, Thorat et al experimented the efficiency of the Support Vector Machine (SVM) algorithm in comparison to other machine learning models in fraud detection . Their proposed system with the SVM model of real databases acquired a maximum accuracy of 99.9%.

The Artificial Neural Network (ANN) comes with 97.32% accuracy while Hidden Markov Model (HMM) has 94.7% accuracy. ANN has a high processing time and excessive training for large neural networks, difficult to set up and run. Also, Bayesian Networks need excessive training and have 96.52% accuracy.

3. Methodology

In this paper, fraud detection in Nigerian credit card transaction data is treated as an anomaly detection problem. The major objectives of the methods applied is to determine the most reliable and well-fitted unsupervised detection technique. It is a layered comparison of Autoencoders built from scratch and the Principal Component Analysis algorithm. The visualization formats of both techniques are vastly different. However, the precision of data represented contributes to the advantages of one method over the other, as is discussed in the following sections.

3.1 Autoencoders

Autoencoders are a type of neural network architecture designed and primarily used for unsupervised learning and data compression . They are categorized as a type of unsupervised learning technique that using back propagation to learn the compressed information in raw data.

The main objective of autoencoders is to perform dimensionality reduction on a dataset, a technique used to reduce the number of features in a dataset while retaining as much of the important information as possible, preventing over-fitting and elongated training periods .

Autoencoders yield a recreation of the input. The autoencoders comprises of two sub-systems: an encoder and a decoder. During training, the encoder learns a set of training, a process identified as a latent representation, from the data it receives. At the same time, the decoder undergoes a similar training process to reconstruct the input. The autoencoders can at that point be connected to perform predictive operations on future input data. Autoencoders are exceptionally generalizable and can be utilized on distinctive information sorts, including picture data, time series, and text.

Fig 2: How Autoeconders Work

The principle of operation of autoencoders is similar to principal component analysis (PCA). The major difference lies in the non-linearity of the activation function used in autoecnoders. If the AF was linear within each processing layer of the network, the latent variable of the autoencodcers would be exactly the same as the major functions in the PCA algorithm.

Generally, the activation function used in autoencoders is non-linear, typical activation functions are ReLU (Rectified Linear Unit) and sigmoid. Autoencoders can be segmented into three parts that can be represented mathematically as:

Ø, Ѱ = Arg Min || X – (Ѱ O Ø) X||2

Encoding: The encoder function, denoted by ϕ, connects the input data X, to latent space F. The decoder is represented by ψ which maps the latent space F to the output. In the autoencoders architecture, the output is the same as the input, denoting an attempt to recreate the original input following a process of generalized non-linear compression.

The encoding network can be represented by the standard neural network function passed through an activation function, where z is the latent dimension;

Z = Σ(Wx + B)

Decoding: The decoding network is represented similarly but with different a weight, bias, and potentially different activation functions.

X´ = Σ´(W´Z + B´)

The loss function is expressed in terms of the functions above, and it is this loss function that is utilized in training the neural network through the standard back propagation procedure. L(X,X´) = ||x-x´||2 = ||x - σ´(W´(σ(Wx + b)) + b´)||2 Another common loss function used for autoencoders is the mean squared error (MSE) loss,

L = 1/N * Sum((X - X')2)

A reduction of the loss function through the gradient descent enables the autoencoders system to learn to encode and decode the input data efficiently. The compressed representation is then used in anomaly detection and in our case, credit card fraud detection. One major objective of the autoencoders system is to select the encoder and decoder functions in such a way that minimal information is required to encode the dataset such that it be can regenerated on the other side.

Ai Adaptive Learning

This project focuses on ai adaptive learning using modern AI and machine learning techniques. The content below is adapted from research literature and practical implementation notes.

We propose a novel high-performance and interpretable canon-

addition, unlike tree learning, DNNs enable gradient descent- ical deep tabular data learning architecture, TabNet. TabNet based end-to-end learning for tabular data which can have a uses sequential attention to choose which features to reason multitude of benefits: (i) efficiently encoding multiple data from at each decision step, enabling interpretability and more types like images along with tabular data; (ii) alleviating the efficient learning as the learning capacity is used for the most need for feature engineering, which is currently a key aspect

salient features. We demonstrate that TabNet outperforms in tree-based tabular data learning methods; (iii) learning other variants on a wide range of non-performance-saturated from streaming data and perhaps most importantly (iv) end- tabular datasets and yields interpretable feature attributions to-end models allow representation learning which enables plus insights into its global behavior. Finally, we demonstrate many valuable application scenarios including data-efficient

self-supervised learning for tabular data, significantly improv- domain adaptation (Goodfellow, Bengio, and Courville 2016), ing performance when unlabeled data is abundant. generative modeling (Radford, Metz, and Chintala 2015) and

Introduction We propose a new canonical DNN architecture for tabular

Deep neural networks (DNNs) have shown notable success data, TabNet. The main contributions are summarized as: efficiently encode the raw data into meaningful representa- enabling flexible integration into end-to-end learning. tions, fuel the rapid progress. One data type that has yet to 2. TabNet uses sequential attention to choose which fea- see such success with a canonical architecture is tabular data. tures to reason from at each decision step, enabling in-

Despite being the most common data type in real-world AI terpretability and better learning as the learning capacity (as it is comprised of any categorical and numerical features), is used for the most salient features (see Fig. 1). This under-explored, with variants of ensemble decision trees for each input, and unlike other instance-wise feature se- Why? First, because DT-based approaches have certain bene- and van der Schaar 2019), TabNet employs a single deep

fits: (i) they are representionally efficient for decision mani- learning architecture for feature selection and reasoning. folds with approximately hyperplane boundaries which are 3. Above design choices lead to two valuable properties: (i) common in tabular data; and (ii) they are highly interpretable TabNet outperforms or is on par with other tabular learn- in their basic form (e.g. by tracking decision nodes) and there ing models on various datasets for classification and re-

are popular post-hoc explainability methods for their ensem- gression problems from different domains; and (ii) TabNet ble form, e.g. (Lundberg, Erion, and Lee 2018) – this is an enables two kinds of interpretability: local interpretability important concern in many real-world applications; (iii) they that visualizes the importance of features and how they are fast to train. Second, because previously-proposed DNN are combined, and global interpretability which quantifies

architectures are not well-suited for tabular data: e.g. stacked the contribution of each feature to the trained model. convolutional layers or multi-layer perceptrons (MLPs) are 4. Finally, for the first time for tabular data, we show signif- vastly overparametrized – the lack of appropriate inductive icant performance improvements by using unsupervised bias often causes them to fail to find optimal solutions for tab- pre-training to predict masked features (see Fig. 2).

ular decision manifolds (Goodfellow, Bengio, and Courville

Why is deep learning worth exploring for tabular data?

One obvious motivation is expected performance improve- Feature selection: Feature selection broadly refers to judi- Copyright © 2021, Association for the Advancement of Artificial ciously picking a subset of features based on their useful-

Professional occupation related Investment related

Feedback from Feedback to

Feature selection Input processing Feature selection Input processing

previous step next step … …

Predicted output (whether the income level >$50k)

selection enables interpretability and better learning as the capacity is used for the most salient features. TabNet employs multiple decision blocks that focus on processing a subset of input features for reasoning. Two decision blocks shown as examples process features that are related to professional occupation and investments, respectively, in order to predict the income level.

Unsupervised pre-training Supervised fine-tuning

Age Cap. gain Education Occupation Gender Relationship Age Cap. gain Education Occupation Gender Relationship 5 2000 ? Exec-managerial F Wife 6 2000 Bachelors Exec-managerial M Husband 1 0 ? Farming-fishing M ? 2 0 High-school Farming-fishing M Unmarried

? 50 Doctorate Prof-specialty M Husband 4 50 Doctorate Prof-specialty M Husband 2 ? ? Handlers-cleaners F Wife 2 0 High-school Handlers-cleaners F Wife 5 3000 Bachelors ? ? Husband 5 3000 Bachelors Exec-managerial M Husband

3 0 Bachelors ? F ? 3 100 Bachelors Prof-specialty F Wife ? 0 High-school Armed-Forces ? Husband 2 0 High-school Armed-Forces M Husband

TabNet decoder Decision making

Age Cap. gain Education Occupation Gender Relationship Income > $50k

3 M False

level can be guessed from the occupation, or the gender can be guessed from the relationship. Unsupervised representation learning by masked self-supervised learning results in an improved encoder model for the supervised learning task.

ward selection and Lasso regularization (Guyon and Elisseeff performance with compact representations. 2003) attribute feature importance based on the entire training Tree-based learning: DTs are commonly-used for tabular data, and are referred as global methods. Instance-wise fea- data learning. Their prominent strength is efficient picking ture selection refers to picking features individually for each of global features with the most statistical information gain

to maximize the mutual information between the selected mance of standard DTs, one common approach is ensembling features and the response variable, and in (Yoon, Jordon, and to reduce variance. Among ensembling methods, random van der Schaar 2019) by using an actor-critic framework to forests (Ho 1998) use random subsets of data with randomly mimic a baseline while optimizing the selection. Unlike these, selected features to grow many trees. XGBoost (Chen and

sity in end-to-end learning – a single model jointly performs recent ensemble DT approaches that dominate most of the feature selection and output mapping, resulting in superior recent data science competitions. Our experimental results

!# + Softmax !" < % !" > % !# > & !# > &

ReLU ReLU &

$" !" − $" % −1 −$" !" + $" % −1 −1 $# !# − $# & % −1 −$# !# + $# & !"

FC FC

W: [$" , - $" , 0, 0] W: [0, 0, $# , - $# ] !" < % b: [-a $" , a $" , -1, -1] b: [-1, -1, -d $# , d $# ] !# < & !" > % !# < & [!" ] [!# ]

M: [1, 0] M: [0, 1]

(right). Relevant features are selected by using multiplicative sparse masks on inputs. The selected features are linearly transformed, and after a bias addition (to represent boundaries) ReLU performs region selection by zeroing the regions. Aggregation of multiple regions is based on addition. As C and C get larger, the decision boundary gets sharper.

for various datasets show that tree-based models can be out- constructs a sequential multi-step architecture, where each performed when the representation capacity is improved with step contributes to a portion of the decision based on the deep learning while retaining their feature selecting property. selected features; (iii) improves the learning capacity via non- Integration of DNNs into DTs: Representing DTs with linear processing of the selected features; and (iv) mimics

DNN building blocks as in (Humbird, Peterson, and McClar- ensembling via higher dimensions and more steps. ren 2018) yields redundancy in representation and ineffi- cient learning. Soft (neural) DTs (Wang, Aggarwal, and Liu Fig. 4 shows the TabNet architecture for encoding tabu- functions, instead of non-differentiable axis-aligned splits. mapping of categorical features with trainable embeddings.

However, losing automatic feature selection often degrades We do not consider any global feature normalization, but performance. In (Yang, Morillo, and Hospedales 2018), a soft merely apply batch normalization (BN). We pass the same D- binning function is proposed to simulate DTs in DNNs, by dimensional features f ∈ <B×D to each decision step, where 2019) proposes a DNN architecture by explicitly leveraging multi-step processing with Nsteps decision steps. The ith

expressive feature combinations, however, learning is based step inputs the processed information from the (i − 1)th step on transferring knowledge from gradient-boosted DT. (Tanno to decide which features to use and outputs the processed ing from primitive blocks while representation learning into sion. The idea of top-down attention in the sequential form edges, routing functions and leaf nodes. TabNet differs from is inspired by its applications in processing visual and text

these as it embeds soft feature selection with controllable data (Hudson and Manning 2018) and reinforcement learn- Self-supervised learning: Unsupervised representation relevant information in high dimensional input. learning improves supervised learning especially in small Feature selection: We employ a learnable mask M[i] ∈ has shown significant advances – driven by the judicious capacity of a decision step is not wasted on irrelevant

choice of the unsupervised learning objective (masked input ones, and thus the model becomes more parameter effi- prediction) and attention-based deep learning. cient. The masking is multiplicative, M[i] · f . We use an attentive transformer (see Fig. 4) to obtain the masks us- TabNet for Tabular Learning ing the processed features from the preceding step, a[i − 1]:

M[i] = sparsemax(P[i − 1] · hi (a[i − 1])). Sparsemax nor-

DTs are successful for learning from real-world tabular malization (Martins and Astudillo 2016) encourages sparsity datasets. With a specific design, conventional DNN building by mapping the Euclidean projection onto the probabilistic blocks can be used to implement DT-like output manifold, simplex, which is observed to be superior in performance and e.g. see Fig. 3). In such a design, individual feature selec- aligned with the goal of sparse feature selection for explain-

tion is key to obtain decision boundaries in hyperplane form, PD which can be generalized to a linear combination of features ability. Note that j=1 M[i]b,j = 1. hi is a trainable func- where coefficients determine the proportion of each feature. tion, shown in Fig. 4 using a FC layer, followed by BN. P[i] TabNet is based on such functionality and it outperforms DTs is the prior scale term, denoting how much a particular feature

Qi while reaping their benefits by careful design which: (i) uses has been used previously: P[i] = j=1 (γ − M[j]), where γ sparse instance-wise feature selection learned from data; (ii) is a relaxation parameter – when γ = 1, a feature is enforced

+ Softmax

Feature Feature …

transformer transformer

x Nsteps Features

+ Softmax

Feature …

transformer transformer Feature Feature Feature Feature transformer

Encoded representation

transformer transformer Attentive transformer … Mask transformer …

Step 2 Decision step dependent

transformer transformer

BN Feature Feature

FC BN transformer transformer

+ 0.5 0.5 0.5 Agg. Agg. Features Features FC FC + +

Reconstructed + … Feature attributes + … features

(a) TabNet encoder architecture (b) TabNet decoder architecture Feature transformer Feature Attentive transformer Shared across decision steps Decision step dependent transformer GLU

Decision step dependent Prior scales

+ 0.5 0.5 0.5

0.5 0.5 0.5

+ Attentive transformer (c) (d)

Prior scales

divides the processed representation to be used by the attentive transformer of the subsequent step as well as for the overall Attentive BN FC

output. For each step, the feature selection mask provides interpretable information about the model’s functionality, and the +

masks can be aggregated to obtain global feature transformer important attribution. (b) TabNet decoder, composed of a feature transformer block at each step. (c) A feature transformer block example – 4-layer network is shown, where 2 are shared across all decision

Prior scales

steps and 2 are decision step-dependent. Each layer is composed of a fully-connected (FC) layer, BN and GLU nonlinearity. (d) +

An attentive transformer block example – a single layer mapping is modulated with a prior scale information which aggregates Sparsemax

how much each feature has been used before the current decision step. sparsemax (Martins and Astudillo 2016) is used for BN FC

normalization of the coefficients, resulting in sparse selection of the salient features. +

to be used only at one decision step and as γ increases, more propose the aggregate.feature importance mask, Magg−b,j = flexibility is provided to use a feature at multiple decision PNsteps ηb [i]Mb,j [i]

PD PNsteps

ηb [i]Mb,j [i].2 i=1 i=1 steps. P is initialized as all ones, 1B×D , without any prior j=1

on the masked features. If some features are unused (as in self- Tabular self-supervised learning: We propose a decoder supervised learning), corresponding P entries are made 0 architecture to reconstruct tabular features from the Tab- to help model’s learning. To further control the sparsity of the Net encoded representations. The decoder is composed of selected features, we propose sparsity regularization in the feature transformers, followed by FC layers at each deci-

form of entropy (Grandvalet and Bengio 2004), Lsparse = sion step. The outputs are summed to obtain the recon-

PNsteps PB PD −Mb,j [i] log(Mb,j [i]+)

i=1 b=1 j=1 Nsteps ·B , where  is a structed features. We propose the task of prediction of miss- small number for numerical stability. We add the sparsity reg- ing feature columns from the others. Consider a binary mask ularization to the overall loss, with a coefficient λsparse . Spar- S ∈ {0, 1}B×D . The TabNet encoder inputs (1 − S) · f̂ sity provides a favorable inductive bias for datasets where and the decoder outputs the reconstructed features, S · f̂ . We

most features are redundant. initialize P = (1 − S) in encoder so that the model em- Feature processing: We process the filtered features using phasizes merely on the known features, and the decoder’s last a feature transformer (see Fig. 4) and then split for the FC layer is multiplied with S to output the unknown features. decision step output and information for the subsequent We consider the reconstruction loss in self-supervised phase:

step, [d[i], a[i]] = fi (M[i] · f ), where d[i] ∈ <B×Nd and 2

PB PD (f̂b,j −fb,j )·Sb,j

a[i] ∈ <B×Na . For parameter-efficient and robust learning b=1 j=1

√ PB PB 2

. Normalization b=1 (fb,j −1/B b=1 fb,j ) with high capacity, a feature transformer should comprise layers that are shared across all decision steps (as the same with the population standard deviation of the ground truth features are input across different decision steps), as well as is beneficial, as the features may have different ranges. We decision step-dependent layers. Fig. 4 shows the implementa- sample Sb,j independently from a Bernoulli distribution with

tion as concatenation of two shared layers and two decision parameter ps , at each iteration. step-dependent layers. Each FC layer is followed by BN and eventually connected to a normalized residual √ connection We study TabNet in wide range of problems, that contain with normalization. Normalization with 0.5 helps to sta- regression or classification tasks, particularly with published bilize learning by ensuring that the variance throughout the benchmarks. For all datasets, categorical inputs are mapped

For faster training, we use large batch sizes with BN. Thus, bedding and numerical columns are input without and pre- except the one applied to the input features, we use ghost BN processing.4 We use standard classification (softmax cross (Hoffer, Hubara, and Soudry 2017) form, using a virtual batch entropy) and regression (mean squared error) loss functions size BV and momentum mB . For the input features, we ob- and we train until convergence. Hyperparameters of the Tab-

serve the benefit of low-variance averaging and hence avoid Net models are optimized on a validation set and listed in ghost BN. Finally, inspired by decision-tree like aggregation Appendix. TabNet performance is not very sensitive to most as in Fig. 3, we construct the overall decision embedding hyperparameters as shown with ablation studies in Appendix. as dout = i=1 PNsteps ReLU(d[i]). We apply a linear mapping In Appendix, we also present ablation studies on various de-

Wfinal dout to get the output mapping.1 sign and guidelines on selection of the key hyperparameters. Interpretability: TabNet’s feature selection masks can shed For all experiments we cite, we use the same training, val- light on the selected features at each step. If Mb,j [i] = 0, idation and testing data split with the original work. Adam optimization algorithm (Kingma and Ba 2014) and Glorot then j th feature of the bth sample should have no contribution uniform initialization are used for training of all models.5

to the decision. If fi were a linear function, the coefficient

Mb,j [i] would correspond to the feature importance of fb,j . Instance-wise feature selection

Although each decision step employs non-linear processing, their outputs are combined later in a linear way. We aim Selection of the salient features is crucial for high perfor- to quantify an aggregate feature importance in addition to mance, especially for small datasets. We consider 6 tabular requires a coefficient that can weigh the relative importance samples). The datasets are constructed in such a way that of each step in the decision. We simply propose ηb [i] = only a subset of the features determine the output. For Syn1-

PNd Syn3, salient features are same for all instances (e.g., the

c=1 ReLU(db,c [i]) to denote the aggregate decision con- tribution at ith decision step for the bth sample. Intuitively, if 2

Normalization is used to ensure D

P j=1 Magg−b,j = 1. db,c [i] < 0, then all features at ith decision step should have 3

0 contribution to the overall decision. As its value increases, prove the performance, but interpretation of individual dimensions

it plays a higher role in the overall linear combination. Scal- may become challenging. ing the decision mask at each decision step with ηb [i], we Specially-designed feature engineering, e.g. logarithmic trans- formation of variables highly-skewed distributions, may further

For discrete outputs, we additionally employ softmax during

training (and argmax during inference). An open-source implementation will be released.

Global: using only globally-salient features, Tree Ensembles (Geurts, Ernst, and Wehenkel 2006), Lasso-regularized model, L2X

Syn Syn Syn Syn Syn Syn

No selection .5 ± .0 .7 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .6 ± .0 Tree .5 ± .1 .8 ± .0 .8 ± .0 .6 ± .0 .7 ± .0 .7 ± .0 Lasso-regularized .4 ± .0 .5 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .7 ± .0

INVASE .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0

Global .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0 TabNet .6 ± .0 .8 ± .0 .8 ± .0 .7 ± .0 .7 ± .0 .8 ± .0

output of Syn depends on features X -X ), and global fea- Table 3: Performance for Poker Hand induction dataset. ture selection, as if the salient features were known, would give high performance. For Syn4-Syn6, salient features are Model Test accuracy (%) instance dependent (e.g., for Syn4, the output depends on ei- DT 50.0 ther X -X or X -X depending on the value of X ), which MLP 50.0

makes global feature selection suboptimal. Table 1 shows that Deep neural DT 65.1

TabNet outperforms others (Tree Ensembles (Geurts, Ernst, XGBoost 71.1

and Wehenkel 2006), LASSO regularization, L2X (Chen LightGBM 70.0 van der Schaar 2019). For Syn1-Syn3, TabNet performance TabNet 99.2 is close to global feature selection - it can figure out what Rule-based 100.0 features are globally important. For Syn4-Syn6, eliminating instance-wise redundant features, TabNet improves global feature selection. All other methods utilize a predictive model Poker Hand (Dua and Graff 2017): The task is classifica-

with 43k parameters, and the total number of parameters is tion of the poker hand from the raw suit and rank attributes of 101k for INVASE due to the two other models in the actor- the cards. The input-output relationship is deterministic and critic framework. TabNet is a single architecture, and its size hand-crafted rules can get 100% accuracy. Yet, conventional is 26k for Syn1-Syn and 31k for Syn4-Syn6. The compact DNNs, DTs, and even their hybrid variant of deep neural DTs

representation is one of TabNet’s valuable properties. (Yang, Morillo, and Hospedales 2018) severely suffer from the imbalanced data and cannot learn the required sorting and Performance on real-world datasets ranking operations (Yang, Morillo, and Hospedales 2018).

Tuned XGBoost, CatBoost, and LightGBM show very slight

as it can perform highly-nonlinear processing with its depth, Model Test accuracy (%) without overfitting thanks to instance-wise feature selection.

CatBoost 85.1 Table 4: Performance on Sarcos dataset. Three TabNet mod-

AutoML Tables 94.9 els of different sizes are considered.

Forest Cover Type (Dua and Graff 2017): The task is clas- MLP 2.1 0.14M

sification of forest cover type from cartographic variables. Adaptive neural tree 1.2 0.60M approaches that are known to achieve solid performance (AutoML 2019), an automated search framework based on TabNet-M 0.2 0.59M ensemble of models including DNN, gradient boosted DT, TabNet-L 0.1 1.75M with very thorough hyperparameter search. A single TabNet without fine-grained hyperparameter search outperforms it. Sarcos (Vijayakumar and Schaal 2000): The task is re-

gressing inverse dynamics of an anthropomorphic robot arm.

very small model is possible with a random forest. In the very and TabNet merely focuses on the relevant ones. For Syn4, small model size regime, TabNet’s performance is on par the output depends on either X -X or X -X depending parameters. When the model size is not constrained, TabNet feature selection – it allocates a mask to focus on the indi- achieves almost an order of magnitude lower test MSE. cator X , and assigns almost all-zero weights to irrelevant

features (the ones other than two feature groups). models are denoted with -S and -M. Real-world datasets: We first consider the simple task of mushroom edibility prediction (Dua and Graff 2017). Tab- Model Test acc. (%) Model size Net achieves 100% test accuracy on this dataset. It is indeed Sparse evolutionary MLP 78.4 81K known (Dua and Graff 2017) that “Odor” is the most discrim-

What is this project about?

This project covers practical implementation and research aspects of the topic using AI/ML techniques.