1
Abstract— In the domain of Biometrics, recognition systems based on iris, fingerprint or palm print scans etc. are often considered more dependable due to extremely low variance in the properties of these entities with respect to time. However, over the last decade data processing capability of computers has increased manifold, which has made real-time video content analysis possible. This shows that the need of the hour is a robust and highly automated Face Detection and Recognition algorithm with credible accuracy rate. The proposed Face Detection and Recognition system using Discrete Wavelet Transform (DWT) accepts face frames as input from a database containing images from low cost devices such as VGA cameras, webcams or even CCTV’s, where image quality is inferior. Face region is then detected using properties of L*a*b* color space and only Frontal Face is extracted such that all additional background is eliminated. Further, this extracted image is converted to grayscale and its dimensions are resized to 128 x 128 pixels. DWT is then applied to entire image to obtain the coefficients. Recognition is carried out by comparison of the DWT coefficients belonging to the test image with those of the registered reference image. On comparison, Euclidean distance classifier is deployed to validate the test image from the database.
Accuracy for various levels of DWT Decomposition is obtained and hence, compared. Keywords— discrete wavelet transform, face detection, face recognition, person identification.
I. Introduction
face recognition system is essentially an application intended to identify or verify a person either from a digital image or a video frame obtained from a video source. Although other reliable methods of biometric personal identification exist, for e.g., fingerprint analysis or iris scans, these methods inherently rely on the cooperation of the participants, whereas a personal identification system based on analysis of frontal or profile images of the face is often effective without the
Or
intervention.
Automatic
identification or verification may be achieved by comparing selected facial features from the image and a facial database. This technique is typically used in security systems. Given a Manuscript received September 11, 2011.
1 D. J. Rajdev is with the Thadomal Shahani Engineering College, Mumbai, 2 A. R. Chadha is with the Thadomal Shahani Engineering College,
E-Mail:
3 P. P. Vaidya is with the Thadomal Shahani Engineering College, Mumbai, 4 M. M. Roja is an Associate Professor in the Electronics and large database of images and a photograph, the problem is to select from the database a small set of records such that one of the image records matched the photograph. The success of the method could be measured in terms of the ratio of the answer list to the number of records in the database. The recognition problem is made difficult by the great variability in head rotation and tilt, lighting intensity and angle, facial expression, aging, etc. A robust facial recognition system must be able to cope with the above factors and yet provide satisfactory accuracy levels. A general statement of the problem of machine recognition of faces can be formulated as: given a still or video image of a scene, identify or verify one or more persons in the scene using a stored database of faces. The solution to the problem involves segmentation of faces, feature extraction from face regions, recognition, or verification. In identification problems, the input to the system is an unknown face, and the system reports back the determined identity from a database of known individuals, whereas in verification problems, the system needs to confirm or reject the claimed identity of the input face.
Some of the various applications of face recognition include driving licenses, immigration, national ID, passport, voter registration, security application, medical records, personal
Human-Robot-Interaction,
human-computer-interaction, smart cards etc. Face recognition is such a challenging yet interesting problem that it has attracted researchers who have different backgrounds: pattern recognition, neural networks, computer vision, and computer graphics, hence the literature is vast and diverse. The usage of a mixture of techniques makes it difficult to classify these systems based on what types of techniques they use for feature representation or classification. To have clear categorization, the proposed paper follows the holistic approach .
Specifically, the following techniques are employed for facial
Feature Extraction And Recognition:
1) Holistic matching methods: These methods use the whole face region as a raw input to the recognition system. One of the most widely used representations of the face region is Eigenpictures, which is inherently based on principal component analysis.
2) Feature-based matching methods: Generally, in these methods, local features such as the eyes, nose and mouth are first extracted and their locations and local statistics are fed as inputs into a classifier.
3) Hybrid methods: It uses both local features and whole face region to recognize a face. This method could potentially offer the better of the two types of methods.
And Face Recognition
Divya Jyoti1, Aman Chadha2, Pallavi Vaidya3, and M. Mani Roja4
2
Most electronic imaging applications often desire and require high resolution images. ‘High resolution’ basically means that pixel density within an image is high, and therefore a HR image can offer more details and subtle transitions that may be critical in various applications . For instance, high resolution medical images could be very helpful for a doctor to make an accurate diagnosis. It may be easy to distinguish an object from similar ones using high resolution satellite images, and the performance of pattern recognition in computer vision can easily be improved if such images are provided. Over the past few decades, charge-coupled device (CCD) and CMOS image sensors have been widely used to capture digital images.
Although these sensors are suitable for most imaging applications, the current resolution level and consumer price will not satisfy the future demand .
Past studies by researches and scientists that have investigated the challenging task of face detection and recognition have therefore, typically used high resolution images. Moreover, most standard face databases such as the MIT-CBCL Face Recognition Database , CMU Multi-PIE , The Yale Face Database etc., that are basically used as a standard test data set by researchers to benchmark their results, also employ high quality images.
Results obtained by solutions proposed by researchers are therefore, relevant for theoretical understanding of face detection and identification in most cases. Practical conditions being rarely optimal, a number of factors play an important role in hampering system performance. Image degradation, i.e., loss of resolution caused mainly by large viewing distances as demonstrated in , and lack of specialized high resolution image capturing equipment such as commercial cameras are the underlying factors for poor performance of face detection and recognition systems in practical situations. There are two paradigms to alleviate this problem, but both have clear disadvantages. One option is to use super-resolution algorithms to enhance the image as proposed in , but as resolution decreases, super-resolution becomes more vulnerable to environmental variations, and it introduces distortions that affect recognition performance. A detailed analysis of super-resolution constraints has been presented in . On the other hand, it is also possible to match in the low-resolution domain by downsampling the training set, but this is undesirable because features important for recognition depend on high frequency details that are erased by downsampling.
These features are permanently lost upon performing downsampling and cannot be recovered with upsampling . The proposed system has been designed keeping in view these critical factors and to address such bottlenecks.
Ii. Idea Of The Proposed Solution
The database consists of a set of face samples of 50 people. There are 5 test images and 5 training or reference images. Frontal face images are detected and hence, extracted. DWT is applied to the entire image so as to obtain the global features which include approximate coefficients (low frequency
Frequency
coefficients). The approximate coefficients thus obtained, are stored and the detail coefficients are discarded. Various levels of DWT are realized and their corresponding accuracy rates are determined.
A. Frontal Face Image Detection And Extraction
The face detection problem can be defined as, given an input an arbitrary image, which could be a digitized video signal or a scanned photograph, determine whether or not there are any human faces in the image and if there are, then return a code corresponding to their location. Face detection as a computer vision task has many applications. It has direct relevance to the face recognition problem, because the first and foremost important step of an automatic human face recognition system is usually identifying and locating the faces in an unknown image .
For our purpose, face detection is actually a face localization problem in which the image position of single face has to be determined . The goal of our facial feature detection is to detect the presence of features, such as eyes, nose, nostrils, eyebrow, mouth, lips, ears, etc., with the assumption that there is only one face in an image . The system should also be robust against human affective states of like happy, sad, disgusted etc. The difficulties associated with face detection systems due to the variations in image appearance such as pose, scale, image rotation and orientation, illumination and facial expression make face detection a difficult pattern recognition problem. Hence, for face detection following problems need to
Be Taken Into Account :
1) Size: A face detector should be able to detect faces in different sizes. Thus, the scaling factor between the reference and the face image under test, needs to be given due consideration.
2) Expressions: The Appearance Of A Face Changes
considerably for different facial expressions and thus, makes face detection more difficult. 3) Pose variation: Face images vary due to relative camera-face pose and some facial features such as an eye or the nose may become partially or wholly occluded.
Another source of variation is the distance of the face from the camera, changes in which can result in perspective distortion.
4) Lighting and texture variation: Changes in the light source in particular can change a face’s appearance can also cause a change in its apparent texture.
5) Presence or absence of structural components: Facial features such as beards, moustaches and glasses may or may not be present. And also there may be variability among these components including shape, colour and size.
The proposed system employs global feature matching for face recognition. However, all computation takes place only on the frontal face, by eliminating the hair and background as these may vary from one image to another. All systems therefore need frontal face extraction. One approach to achieve the aforementioned is by manually cropping the test image for required region or by precisely aligning the user's face with the camera before the test sample is clicked. Both these methods
3
may introduce a high degree of human error and so they have been avoided. Instead, automated Frontal Face Detection and Extraction is put to use. Therefore, a robust automatic face recognition system should be capable of handling the above problem with no need for human intervention. Thus, it is practical and advantageous to realize automatic face detection in a functional face recognition system. Commonly used methods for skin detection include Gabor filters, neural networks and template matching.
It has been proved that Gabor filters give optimum output for a wide range of variations in the test image with respect to user image, but it is the most time intensive procedure .
Moreover, it is unlikely that the test image would be severely out of sync for on the spot face recognition, so this method is not used. Most neural-network based algorithms , are tedious and require training samples for different skin types which add to the already vast reference image database; hence, even this does not fit the program's requirements. Even template matching has severe drawbacks, including high computational cost and fails to work as expected when the user's face is positioned at an angle in the test image.
After considering all the above factors, a classical appearance based methodology is applied to extract Frontal face. The default sRGB colour space is transformed to L*a*b* gamut, because L*a*b* separates intensity from a and b colour components . L*a*b* colour is designed to approximate human vision in contrast to the RGB and CMYK colour models. It aspires to perceptual uniformity, and its L component closely matches human perception of lightness. It can thus be used to make accurate colour balance corrections by modifying output curves in the a and b components, or to adjust the lightness contrast using the L component. In RGB or CMYK spaces, which model the output of physical devices rather than human visual perception, these transformations can only be done with the help of appropriate blend modes in the editing application . This distinction makes L*a*b* space more
Perceptually Uniform As Compared To Srgb And Thus
identifying tones, and not just a single colour, can be accomplished using L*a*b* space. Fig. 1 shows the RGB colour model (B) relating to the CMYK model (C). The larger gamut of L*a*b* (A) gives more of a spectrum to work with, thus making the gamut of the device the only limitation.
Tone identification is applied to detect skin by calculating the gray threshold of a and b colour components and then converting the image to pure black and white (BW) using the obtained threshold. Thus, RGB colour space can separate out only specific pigments, but L*a*b* space can separate out tones. Fig. 2 shows skin color differentiation in the form of white color using L*a*b* space whereas Fig. 4 shows no such differentiation using RGB. The extracted frontal face has been shown in Fig. 3.
There is higher probability of skin surface being the lighter part of the image as compared to the gray threshold , This may happen due to illumination and natural skin colour (in most cultures), so, pure white regions in the black and white image correspond to skin. It is assumed that face will have at least one hole, i.e., a small patch of absolute black due to eyes, chin, dimples etc. and on the basis of presence of holes frontal face is separated from other skin surfaces like hands. A bounding box is created around the Frontal Face and after cropping the excess area, frontal face extraction is complete.
The above technique has been tested extensively on images obtained from standalone VGA cameras, webcams and camera equipped mobile devices having a resolution of 640 × 480.
Even for resolutions as low as 320 × 200, where the test image is poorly illuminated or extremely grainy, the algorithm was able to successfully extract frontal face from test images. Thus, the proposed system is robust enough to achieve desired result even when low cost equipment like CCTV's and low resolution webcams are used. Also, since equipment with inferior picture quality like CCTV's and low resolution webcams are used, the
Fig. 1. Rgb, Cmyk And L*A*B* Colour Model
Fig. 2. Reference image (left), image in black and white a plane (middle) and image in black and white b plane (right) Fig. 3. Image in L*a*b* color space (left), frontal face extracted image (middle) and grayscale resized image (right) Fig. 4. B&W image of red color space (left), B&W image of green color space (middle) and B&W image of blue color space (right)
Ai Adaptive Learning
This project focuses on ai adaptive learning using modern AI and machine learning techniques. The content below is adapted from research literature and practical implementation notes.
We propose a novel high-performance and interpretable canon-
addition, unlike tree learning, DNNs enable gradient descent- ical deep tabular data learning architecture, TabNet. TabNet based end-to-end learning for tabular data which can have a uses sequential attention to choose which features to reason multitude of benefits: (i) efficiently encoding multiple data from at each decision step, enabling interpretability and more types like images along with tabular data; (ii) alleviating the efficient learning as the learning capacity is used for the most need for feature engineering, which is currently a key aspect
salient features. We demonstrate that TabNet outperforms in tree-based tabular data learning methods; (iii) learning other variants on a wide range of non-performance-saturated from streaming data and perhaps most importantly (iv) end- tabular datasets and yields interpretable feature attributions to-end models allow representation learning which enables plus insights into its global behavior. Finally, we demonstrate many valuable application scenarios including data-efficient
self-supervised learning for tabular data, significantly improv- domain adaptation (Goodfellow, Bengio, and Courville 2016), ing performance when unlabeled data is abundant. generative modeling (Radford, Metz, and Chintala 2015) and
Introduction We propose a new canonical DNN architecture for tabular
Deep neural networks (DNNs) have shown notable success data, TabNet. The main contributions are summarized as: efficiently encode the raw data into meaningful representa- enabling flexible integration into end-to-end learning. tions, fuel the rapid progress. One data type that has yet to 2. TabNet uses sequential attention to choose which fea- see such success with a canonical architecture is tabular data. tures to reason from at each decision step, enabling in-
Despite being the most common data type in real-world AI terpretability and better learning as the learning capacity (as it is comprised of any categorical and numerical features), is used for the most salient features (see Fig. 1). This under-explored, with variants of ensemble decision trees for each input, and unlike other instance-wise feature se- Why? First, because DT-based approaches have certain bene- and van der Schaar 2019), TabNet employs a single deep
fits: (i) they are representionally efficient for decision mani- learning architecture for feature selection and reasoning. folds with approximately hyperplane boundaries which are 3. Above design choices lead to two valuable properties: (i) common in tabular data; and (ii) they are highly interpretable TabNet outperforms or is on par with other tabular learn- in their basic form (e.g. by tracking decision nodes) and there ing models on various datasets for classification and re-
are popular post-hoc explainability methods for their ensem- gression problems from different domains; and (ii) TabNet ble form, e.g. (Lundberg, Erion, and Lee 2018) – this is an enables two kinds of interpretability: local interpretability important concern in many real-world applications; (iii) they that visualizes the importance of features and how they are fast to train. Second, because previously-proposed DNN are combined, and global interpretability which quantifies
architectures are not well-suited for tabular data: e.g. stacked the contribution of each feature to the trained model. convolutional layers or multi-layer perceptrons (MLPs) are 4. Finally, for the first time for tabular data, we show signif- vastly overparametrized – the lack of appropriate inductive icant performance improvements by using unsupervised bias often causes them to fail to find optimal solutions for tab- pre-training to predict masked features (see Fig. 2).
ular decision manifolds (Goodfellow, Bengio, and Courville
Why is deep learning worth exploring for tabular data?
One obvious motivation is expected performance improve- Feature selection: Feature selection broadly refers to judi- Copyright © 2021, Association for the Advancement of Artificial ciously picking a subset of features based on their useful-
Professional occupation related Investment related
Feedback from Feedback to
Feature selection Input processing Feature selection Input processing
previous step next step … …
Predicted output (whether the income level >$50k)
selection enables interpretability and better learning as the capacity is used for the most salient features. TabNet employs multiple decision blocks that focus on processing a subset of input features for reasoning. Two decision blocks shown as examples process features that are related to professional occupation and investments, respectively, in order to predict the income level.
Unsupervised pre-training Supervised fine-tuning
Age Cap. gain Education Occupation Gender Relationship Age Cap. gain Education Occupation Gender Relationship 5 2000 ? Exec-managerial F Wife 6 2000 Bachelors Exec-managerial M Husband 1 0 ? Farming-fishing M ? 2 0 High-school Farming-fishing M Unmarried
? 50 Doctorate Prof-specialty M Husband 4 50 Doctorate Prof-specialty M Husband 2 ? ? Handlers-cleaners F Wife 2 0 High-school Handlers-cleaners F Wife 5 3000 Bachelors ? ? Husband 5 3000 Bachelors Exec-managerial M Husband
3 0 Bachelors ? F ? 3 100 Bachelors Prof-specialty F Wife ? 0 High-school Armed-Forces ? Husband 2 0 High-school Armed-Forces M Husband
TabNet decoder Decision making
Age Cap. gain Education Occupation Gender Relationship Income > $50k
3 M False
level can be guessed from the occupation, or the gender can be guessed from the relationship. Unsupervised representation learning by masked self-supervised learning results in an improved encoder model for the supervised learning task.
ward selection and Lasso regularization (Guyon and Elisseeff performance with compact representations. 2003) attribute feature importance based on the entire training Tree-based learning: DTs are commonly-used for tabular data, and are referred as global methods. Instance-wise fea- data learning. Their prominent strength is efficient picking ture selection refers to picking features individually for each of global features with the most statistical information gain
to maximize the mutual information between the selected mance of standard DTs, one common approach is ensembling features and the response variable, and in (Yoon, Jordon, and to reduce variance. Among ensembling methods, random van der Schaar 2019) by using an actor-critic framework to forests (Ho 1998) use random subsets of data with randomly mimic a baseline while optimizing the selection. Unlike these, selected features to grow many trees. XGBoost (Chen and
sity in end-to-end learning – a single model jointly performs recent ensemble DT approaches that dominate most of the feature selection and output mapping, resulting in superior recent data science competitions. Our experimental results
!# + Softmax !" < % !" > % !# > & !# > &
ReLU ReLU &
$" !" − $" % −1 −$" !" + $" % −1 −1 $# !# − $# & % −1 −$# !# + $# & !"
FC FC
W: [$" , - $" , 0, 0] W: [0, 0, $# , - $# ] !" < % b: [-a $" , a $" , -1, -1] b: [-1, -1, -d $# , d $# ] !# < & !" > % !# < & [!" ] [!# ]
M: [1, 0] M: [0, 1]
(right). Relevant features are selected by using multiplicative sparse masks on inputs. The selected features are linearly transformed, and after a bias addition (to represent boundaries) ReLU performs region selection by zeroing the regions. Aggregation of multiple regions is based on addition. As C and C get larger, the decision boundary gets sharper.
for various datasets show that tree-based models can be out- constructs a sequential multi-step architecture, where each performed when the representation capacity is improved with step contributes to a portion of the decision based on the deep learning while retaining their feature selecting property. selected features; (iii) improves the learning capacity via non- Integration of DNNs into DTs: Representing DTs with linear processing of the selected features; and (iv) mimics
DNN building blocks as in (Humbird, Peterson, and McClar- ensembling via higher dimensions and more steps. ren 2018) yields redundancy in representation and ineffi- cient learning. Soft (neural) DTs (Wang, Aggarwal, and Liu Fig. 4 shows the TabNet architecture for encoding tabu- functions, instead of non-differentiable axis-aligned splits. mapping of categorical features with trainable embeddings.
However, losing automatic feature selection often degrades We do not consider any global feature normalization, but performance. In (Yang, Morillo, and Hospedales 2018), a soft merely apply batch normalization (BN). We pass the same D- binning function is proposed to simulate DTs in DNNs, by dimensional features f ∈ <B×D to each decision step, where 2019) proposes a DNN architecture by explicitly leveraging multi-step processing with Nsteps decision steps. The ith
expressive feature combinations, however, learning is based step inputs the processed information from the (i − 1)th step on transferring knowledge from gradient-boosted DT. (Tanno to decide which features to use and outputs the processed ing from primitive blocks while representation learning into sion. The idea of top-down attention in the sequential form edges, routing functions and leaf nodes. TabNet differs from is inspired by its applications in processing visual and text
these as it embeds soft feature selection with controllable data (Hudson and Manning 2018) and reinforcement learn- Self-supervised learning: Unsupervised representation relevant information in high dimensional input. learning improves supervised learning especially in small Feature selection: We employ a learnable mask M[i] ∈ has shown significant advances – driven by the judicious capacity of a decision step is not wasted on irrelevant
choice of the unsupervised learning objective (masked input ones, and thus the model becomes more parameter effi- prediction) and attention-based deep learning. cient. The masking is multiplicative, M[i] · f . We use an attentive transformer (see Fig. 4) to obtain the masks us- TabNet for Tabular Learning ing the processed features from the preceding step, a[i − 1]:
M[i] = sparsemax(P[i − 1] · hi (a[i − 1])). Sparsemax nor-
DTs are successful for learning from real-world tabular malization (Martins and Astudillo 2016) encourages sparsity datasets. With a specific design, conventional DNN building by mapping the Euclidean projection onto the probabilistic blocks can be used to implement DT-like output manifold, simplex, which is observed to be superior in performance and e.g. see Fig. 3). In such a design, individual feature selec- aligned with the goal of sparse feature selection for explain-
tion is key to obtain decision boundaries in hyperplane form, PD which can be generalized to a linear combination of features ability. Note that j=1 M[i]b,j = 1. hi is a trainable func- where coefficients determine the proportion of each feature. tion, shown in Fig. 4 using a FC layer, followed by BN. P[i] TabNet is based on such functionality and it outperforms DTs is the prior scale term, denoting how much a particular feature
Qi while reaping their benefits by careful design which: (i) uses has been used previously: P[i] = j=1 (γ − M[j]), where γ sparse instance-wise feature selection learned from data; (ii) is a relaxation parameter – when γ = 1, a feature is enforced
+ Softmax
Feature Feature …
transformer transformer
x Nsteps Features
+ Softmax
Feature …
transformer transformer Feature Feature Feature Feature transformer
Encoded representation
transformer transformer Attentive transformer … Mask transformer …
Step 2 Decision step dependent
transformer transformer
BN Feature Feature
FC BN transformer transformer
+ 0.5 0.5 0.5 Agg. Agg. Features Features FC FC + +
Reconstructed + … Feature attributes + … features
(a) TabNet encoder architecture (b) TabNet decoder architecture Feature transformer Feature Attentive transformer Shared across decision steps Decision step dependent transformer GLU
Decision step dependent Prior scales
+ 0.5 0.5 0.5
0.5 0.5 0.5
+ Attentive transformer (c) (d)
Prior scales
divides the processed representation to be used by the attentive transformer of the subsequent step as well as for the overall Attentive BN FC
output. For each step, the feature selection mask provides interpretable information about the model’s functionality, and the +
masks can be aggregated to obtain global feature transformer important attribution. (b) TabNet decoder, composed of a feature transformer block at each step. (c) A feature transformer block example – 4-layer network is shown, where 2 are shared across all decision
Prior scales
steps and 2 are decision step-dependent. Each layer is composed of a fully-connected (FC) layer, BN and GLU nonlinearity. (d) +
An attentive transformer block example – a single layer mapping is modulated with a prior scale information which aggregates Sparsemax
how much each feature has been used before the current decision step. sparsemax (Martins and Astudillo 2016) is used for BN FC
normalization of the coefficients, resulting in sparse selection of the salient features. +
to be used only at one decision step and as γ increases, more propose the aggregate.feature importance mask, Magg−b,j = flexibility is provided to use a feature at multiple decision PNsteps ηb [i]Mb,j [i]
PD PNsteps
ηb [i]Mb,j [i].2 i=1 i=1 steps. P is initialized as all ones, 1B×D , without any prior j=1
on the masked features. If some features are unused (as in self- Tabular self-supervised learning: We propose a decoder supervised learning), corresponding P entries are made 0 architecture to reconstruct tabular features from the Tab- to help model’s learning. To further control the sparsity of the Net encoded representations. The decoder is composed of selected features, we propose sparsity regularization in the feature transformers, followed by FC layers at each deci-
form of entropy (Grandvalet and Bengio 2004), Lsparse = sion step. The outputs are summed to obtain the recon-
PNsteps PB PD −Mb,j [i] log(Mb,j [i]+)
i=1 b=1 j=1 Nsteps ·B , where is a structed features. We propose the task of prediction of miss- small number for numerical stability. We add the sparsity reg- ing feature columns from the others. Consider a binary mask ularization to the overall loss, with a coefficient λsparse . Spar- S ∈ {0, 1}B×D . The TabNet encoder inputs (1 − S) · f̂ sity provides a favorable inductive bias for datasets where and the decoder outputs the reconstructed features, S · f̂ . We
most features are redundant. initialize P = (1 − S) in encoder so that the model em- Feature processing: We process the filtered features using phasizes merely on the known features, and the decoder’s last a feature transformer (see Fig. 4) and then split for the FC layer is multiplied with S to output the unknown features. decision step output and information for the subsequent We consider the reconstruction loss in self-supervised phase:
step, [d[i], a[i]] = fi (M[i] · f ), where d[i] ∈ <B×Nd and 2
PB PD (f̂b,j −fb,j )·Sb,j
a[i] ∈ <B×Na . For parameter-efficient and robust learning b=1 j=1
√ PB PB 2
. Normalization b=1 (fb,j −1/B b=1 fb,j ) with high capacity, a feature transformer should comprise layers that are shared across all decision steps (as the same with the population standard deviation of the ground truth features are input across different decision steps), as well as is beneficial, as the features may have different ranges. We decision step-dependent layers. Fig. 4 shows the implementa- sample Sb,j independently from a Bernoulli distribution with
tion as concatenation of two shared layers and two decision parameter ps , at each iteration. step-dependent layers. Each FC layer is followed by BN and eventually connected to a normalized residual √ connection We study TabNet in wide range of problems, that contain with normalization. Normalization with 0.5 helps to sta- regression or classification tasks, particularly with published bilize learning by ensuring that the variance throughout the benchmarks. For all datasets, categorical inputs are mapped
For faster training, we use large batch sizes with BN. Thus, bedding and numerical columns are input without and pre- except the one applied to the input features, we use ghost BN processing.4 We use standard classification (softmax cross (Hoffer, Hubara, and Soudry 2017) form, using a virtual batch entropy) and regression (mean squared error) loss functions size BV and momentum mB . For the input features, we ob- and we train until convergence. Hyperparameters of the Tab-
serve the benefit of low-variance averaging and hence avoid Net models are optimized on a validation set and listed in ghost BN. Finally, inspired by decision-tree like aggregation Appendix. TabNet performance is not very sensitive to most as in Fig. 3, we construct the overall decision embedding hyperparameters as shown with ablation studies in Appendix. as dout = i=1 PNsteps ReLU(d[i]). We apply a linear mapping In Appendix, we also present ablation studies on various de-
Wfinal dout to get the output mapping.1 sign and guidelines on selection of the key hyperparameters. Interpretability: TabNet’s feature selection masks can shed For all experiments we cite, we use the same training, val- light on the selected features at each step. If Mb,j [i] = 0, idation and testing data split with the original work. Adam optimization algorithm (Kingma and Ba 2014) and Glorot then j th feature of the bth sample should have no contribution uniform initialization are used for training of all models.5
to the decision. If fi were a linear function, the coefficient
Mb,j [i] would correspond to the feature importance of fb,j . Instance-wise feature selection
Although each decision step employs non-linear processing, their outputs are combined later in a linear way. We aim Selection of the salient features is crucial for high perfor- to quantify an aggregate feature importance in addition to mance, especially for small datasets. We consider 6 tabular requires a coefficient that can weigh the relative importance samples). The datasets are constructed in such a way that of each step in the decision. We simply propose ηb [i] = only a subset of the features determine the output. For Syn1-
PNd Syn3, salient features are same for all instances (e.g., the
c=1 ReLU(db,c [i]) to denote the aggregate decision con- tribution at ith decision step for the bth sample. Intuitively, if 2
Normalization is used to ensure D
P j=1 Magg−b,j = 1. db,c [i] < 0, then all features at ith decision step should have 3
0 contribution to the overall decision. As its value increases, prove the performance, but interpretation of individual dimensions
it plays a higher role in the overall linear combination. Scal- may become challenging. ing the decision mask at each decision step with ηb [i], we Specially-designed feature engineering, e.g. logarithmic trans- formation of variables highly-skewed distributions, may further
For discrete outputs, we additionally employ softmax during
training (and argmax during inference). An open-source implementation will be released.
Global: using only globally-salient features, Tree Ensembles (Geurts, Ernst, and Wehenkel 2006), Lasso-regularized model, L2X
Syn Syn Syn Syn Syn Syn
No selection .5 ± .0 .7 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .6 ± .0 Tree .5 ± .1 .8 ± .0 .8 ± .0 .6 ± .0 .7 ± .0 .7 ± .0 Lasso-regularized .4 ± .0 .5 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .7 ± .0
INVASE .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0
Global .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0 TabNet .6 ± .0 .8 ± .0 .8 ± .0 .7 ± .0 .7 ± .0 .8 ± .0
output of Syn depends on features X -X ), and global fea- Table 3: Performance for Poker Hand induction dataset. ture selection, as if the salient features were known, would give high performance. For Syn4-Syn6, salient features are Model Test accuracy (%) instance dependent (e.g., for Syn4, the output depends on ei- DT 50.0 ther X -X or X -X depending on the value of X ), which MLP 50.0
makes global feature selection suboptimal. Table 1 shows that Deep neural DT 65.1
TabNet outperforms others (Tree Ensembles (Geurts, Ernst, XGBoost 71.1
and Wehenkel 2006), LASSO regularization, L2X (Chen LightGBM 70.0 van der Schaar 2019). For Syn1-Syn3, TabNet performance TabNet 99.2 is close to global feature selection - it can figure out what Rule-based 100.0 features are globally important. For Syn4-Syn6, eliminating instance-wise redundant features, TabNet improves global feature selection. All other methods utilize a predictive model Poker Hand (Dua and Graff 2017): The task is classifica-
with 43k parameters, and the total number of parameters is tion of the poker hand from the raw suit and rank attributes of 101k for INVASE due to the two other models in the actor- the cards. The input-output relationship is deterministic and critic framework. TabNet is a single architecture, and its size hand-crafted rules can get 100% accuracy. Yet, conventional is 26k for Syn1-Syn and 31k for Syn4-Syn6. The compact DNNs, DTs, and even their hybrid variant of deep neural DTs
representation is one of TabNet’s valuable properties. (Yang, Morillo, and Hospedales 2018) severely suffer from the imbalanced data and cannot learn the required sorting and Performance on real-world datasets ranking operations (Yang, Morillo, and Hospedales 2018).
Tuned XGBoost, CatBoost, and LightGBM show very slight
as it can perform highly-nonlinear processing with its depth, Model Test accuracy (%) without overfitting thanks to instance-wise feature selection.
CatBoost 85.1 Table 4: Performance on Sarcos dataset. Three TabNet mod-
AutoML Tables 94.9 els of different sizes are considered.
Forest Cover Type (Dua and Graff 2017): The task is clas- MLP 2.1 0.14M
sification of forest cover type from cartographic variables. Adaptive neural tree 1.2 0.60M approaches that are known to achieve solid performance (AutoML 2019), an automated search framework based on TabNet-M 0.2 0.59M ensemble of models including DNN, gradient boosted DT, TabNet-L 0.1 1.75M with very thorough hyperparameter search. A single TabNet without fine-grained hyperparameter search outperforms it. Sarcos (Vijayakumar and Schaal 2000): The task is re-
gressing inverse dynamics of an anthropomorphic robot arm.
very small model is possible with a random forest. In the very and TabNet merely focuses on the relevant ones. For Syn4, small model size regime, TabNet’s performance is on par the output depends on either X -X or X -X depending parameters. When the model size is not constrained, TabNet feature selection – it allocates a mask to focus on the indi- achieves almost an order of magnitude lower test MSE. cator X , and assigns almost all-zero weights to irrelevant
features (the ones other than two feature groups). models are denoted with -S and -M. Real-world datasets: We first consider the simple task of mushroom edibility prediction (Dua and Graff 2017). Tab- Model Test acc. (%) Model size Net achieves 100% test accuracy on this dataset. It is indeed Sparse evolutionary MLP 78.4 81K known (Dua and Graff 2017) that “Odor” is the most discrim-
What is this project about?
This project covers practical implementation and research aspects of the topic using AI/ML techniques.