Enquire Now
Medical Computer Vision · Clinical Diagnostics · PyTorch / TensorFlow · GPU Optimized · 2026

Fake News Detection Machine Learning Capstone Project

Tensor Pipeline · Custom Loss Formulations · Model Quantization · Accelerated Inference — A rigorous deep learning engineering project focused on automated pathological lesion segmentation and radiological disease classification. Architected for thesis defense viva presentations, IEEE reproduction, and high-throughput production deployment.

PyTorch
Core Framework
AMP FP16
Mixed Precision
TensorRT
Quantized Serving

A Benchmark Study of Machine Learning Models for Online Fake

News Detection

Junaed Younus Khan*1, Md. Tawkat Islam Khondaker*1, Sadia Afroz2, Gias Uddin3 and Anindya

Abstract

The proliferation of fake news and its propagation on social media has become a major concern due to its ability to create devastating impacts. Different machine learning approaches have been suggested to detect fake news. However, most of those focused on a specific type of news (such as political) which leads us to the question of dataset-bias of the models used. In this research, we conducted a benchmark study to assess the performance of different applicable machine learning approaches on three different datasets where we accumulated the largest and most diversified one.

fake news detection machine learning capstone project Diagram
Figure: System Model & Simulation Flow for Fake News Detection Machine Learning Capstone Project

We explored a number of advanced pre-trained language models for fake news detection along with the traditional and deep learning ones and compared their performances from different aspects for the first time to the best of our knowledge. We find that BERT and similar pre-trained models perform the best for fake news detection, especially with very small dataset. Hence, these models are significantly better option for languages with limited electronic contents, i.e., training data. We also carried out several analysis based on the models’ performance, article’s topic, article’s length, and discussed different lessons learned from them. We believe that this benchmark study will help the research community to explore further and news sites/blogs to select the most appropriate fake news detection method.

fake news detection machine learning capstone project Diagram
Figure: System Model & Simulation Flow for Fake News Detection Machine Learning Capstone Project

Introduction

Fake news can be defined as a type of yellow journalism or propaganda that consists of deliberate misinformation or hoaxes spread via traditional print and broadcast news media or online social media . With the growth of online news portals, social-networking sites, and other online media, online fake news has become a major concern nowadays.

But people are often unable to spend enough time to cross-check references and be sure of the credibility of news. Hence, considering the scale of the users and contributors to the online media, automated detection of fake news is probably the only way to take remedial measures, and therefore currently receiving huge attention from the research community.

Several research works have been carried out on automated fake news detection using both traditional machine learning and deep learning methods over the years [14, 30, 52, 56, 60, 63, 67]. However, most of them focused on de- tecting news of particular types (such as political). Accordingly, they developed their models and designed features for specific datasets that match their topic of interest. These approaches might suffer from dataset bias and perform poorly on news of another topic. Hence, it is important to study if these are sufficient for different types of news pub- lished in online media by evaluating various models on different diverse datasets and comparing their performances.

However, the existing comparative studies on fake news detection methods also focused on a specific type of dataset or explored a limited number of models. For example, Wang built a benchmark dataset namely, Liar, and experimented *The authors contribute equally to this paper. Names are sorted in alphabetical order.

Arxiv:1905.04749V2 [Cs.Cl] 26 Mar 2021

some existing models on it . However, the length of this dataset is not sufficient for neural network based advanced models, and some models were found to suffer from overfitting. Gilda explored a few machine learning approaches but did not evaluate any neural network-based model . Recently, Gravanis et al. evaluated a number of machine learning models on different datasets to address the issue of dataset-bias . However, they also did not explore any deep learning based models in their study. Moreover, very few works have been done to explore advanced pre-trained language models (e.g., BERT, ELECTRA, ELMo) for fake news detection [29, 32] in spite of their state-of-the-art performances in various natural language processing and text classification tasks .

Our study fills this gap by evaluating a wide range of machine learning approaches that include both traditional (e.g., SVM, LR, Decision Tree, Naive Bayes, k-NN) and deep learning (e.g., CNN, LSTM, Bi-LSTM, C-LSTM, HAN, Conv-HAN) models on three different datasets. We have prepared a new combined dataset containing 80k news of a great variety of topics (e.g., politics, economy, investigation, health-care, sports, entertainment) collecting from various sources. To the best of our knowledge, this is the largest dataset used for fake news detection study. We also explored a variety of pre-trained models, e.g., BERT , RoBERTa , DistilBERT , ELECTRA , ELMo in our comparative analysis. To the best of our knowledge, no previous study has incorporated such advanced pre-trained models to compare their performance with other machine learning models on fake news detection task. In particular, we answer the following research questions.

RQ1: How accurate are the traditional machine learning vs deep learning models to detect fake news? We find that deep learning models generally outperform the traditional machine learning models, Among the traditional learning models, Na¨ıve Bayes which achieves 93% accuracy on combined corpus. Among the deep learning models, Bi-LSTM and C-LSTM show great promise with 95% accuracy on combined corpus.

RQ2: Can the advanced pre-trained language models outperform the traditional and deep learning models? We investigated pre-trained models like BERT, DistilBERT, RoBERTa, ELECTRA, and ELMo. Overall, these models outperform traditional and deep learning ones. For example, the pre-trained RoBERTa shows 96% accuracy on combined corpus, which is more than the traditional and deep learning models. We also find that BERT and similar transformer-based models (BERT, DistilBERT, RoBERTa, ELECTRA) perform better than ELMo.

RQ3: Which model performs best with small training data? The superior performance of deep learning and pre-trained models we observed in our datasets could be due to large dataset sizes. However, the construction of a large dataset may not always be possible. We, therefore, attempted to understand whether smaller datasets can still be used to train the models without a considerable reduction in accuracy.

We see that the pre-trained models can achieve high performance with very small training dataset compared to tradi- tional or deep learning models. For example, RoBERTa achieved over 90% accuracy with only 500 training data used for fine-tuning while traditional and deep learning models fail to achieve even 80% accuracy with such small dataset (see Figure 3). In contrast, the best performing traditional learning model Na¨ıve Bayes only achieved 65% accuracy with a sample size of 500 training set. Therefore, our finding can be useful for electronic-resource-limited languages where fake news dataset collections are likely to be small in size. In such cases, based on our observations, pre-trained models are the best option to achieve quality performance for these languages. Note that different languages such as Dutch, Italian, Arabic, Bangla, etc. have pre-trained BERT models [4, 15, 47] that can be fine-tuned with small fake news dataset to develop detection tool.

Replication Package with code and data is shared online at https://github.com/JunaedYounusKhan51/ FakeNewsDetection. Paper Organizations. The rest of this paper is structured as follows. In Section 2, we compare related research works. In Section 3, we describe our study setup by introducing the datasets, the features, and the models we used in our experiments. Section 4 presents the performance of different models on three datasets and answer three research questions. Section 5 compares the performance and analyzes the misclassified cases. We conclude in Section 6.

Related Work

Related work can broadly be divided into the following categories: (1) Exploratory analysis of the characteristics of fake news, (2) Traditional machine learning based detection, (3) Deep learning based detection, (4) Advanced language model based detection, and (5) Benchmark studies.

2

Exploratory analysis of the characteristics of fake news Several research works have been done over the years on the characteristics of fake news and its’ detection. Conroy et al. mentioned three types of fake news: Serious Fabrications, Large-Scale Hoaxes, and Humorous Fakes .

They have termed fake news as a news article that is intentionally and verifiably false and could mislead readers . This narrow definition is useful in the sense that it can eliminate the ambiguity between fake news and other related concepts, e.g., hoaxes, and satires.

Traditional Machine Learning Based Detection

Different traditional machine learning based approaches have been proposed for the automatic detection of fake news. In , the authors proposed to use linguistic-based features such as total words, characters per word, frequencies of large words, frequencies of phrases, i.e., “n-grams” and bag-of-words approaches , parts-of-speech (POS) tagging for fake news detection.

Conroy et al. argued that simple content-related n-grams and part-of-speech (POS) tagging had been proven insuf- ficient for the classification task . Rather, they suggested Deep Syntax analysis using Probabilistic Context-Free Grammars (PCFG) following another work by Feng et al. to distinguish rule categories (i.e., lexicalized, non- lexicalized, parent nodes, etc.) for deception detection with 85-91% accuracy. However, Shlok Gilda reported that while bi-gram TF-IDF yielded highly effective models for detecting fake news, the PCFG features had little to add to the models’ efficacy .

Many research works also suggested the use of sentiment analysis for deception detection as some correlation might be found between the sentiment of the news article and its type. Reference proposed expanding the possi- bilities of word-level analysis by measuring the utility of features like part of speech frequency, and semantic categories such as generalizing terms, positive and negative polarity (sentiment analysis).

Cliche described the detection of sarcasm on twitter using n-grams, words learned from tweets specifically tagged as sarcastic . His work also included the use of sentiment analysis as well as identification of topics (words that are often grouped together in tweets) to improve prediction accuracy.

Deep Learning Based Detection

Several research works used deep learning models to detect fake news. Wang et al. built a hybrid convolutional neural network model that outperforms other traditional machine learning models . Rashkin et al. performed an extensive analysis of linguistic features and showed promising result with LSTM . Singhania et al. proposed a three-level hierarchical attention network, one each for words, sentences, and the headline of a news article . Ruchansky et al.

created the CSI model where they have captured text, the response of an article, and the source characteristics based on users’ behaviour . Among the recent works, Shu et al. argued that a critical aspect of fake news detection is the explainability of such detection in . The authors developed a sentence-comment co-attention sub-network to exploit both news contents and user comments. In this way, the authors focused on jointly capturing explainable check-worthy sentences and user comments for fake news detection. In the work , the authors developed a multimodal variational auto-encoder by using a bi-modal variational auto-encoder coupled with a binary classifier for the task of fake news detection.

The authors claimed that this end-to-end network utilizes the multimodal representations obtained from the bi-modal variational auto-encoder to classify posts as fake or not. Zhou et al. focused on studying the patterns of spreading of fake news in social networks, and the relationships among the spreaders . Hamdi et al. proposed a hybrid approach to detect misinformation in Twitter . The authors extracted user characteristics using node2vec to verify the credibility of the contents.

Advanced Language Model Based Detection

Currently, Advanced pre-trained language models (i.e., BERT, ELECTRA, ELMo) are receiving great attention for several natural language tasks including text classification . However, only a few studies have explored them for fake news detection. For example, Jwa et al. detected fake news by analyzing the relationship between the headline and the body text of news . The authors claimed that the deep-contextualizing nature of BERT improves F-score by 0.14 over the previous state-of-the-art models. Kula et al. presented a hybrid architecture

3

Table 1: Comparison between our benchmark study and prior benchmark studies

Literature

.

Did Not Report Any

results.

Analyzed Their

performances.

Detection Methods

.

Detection Approaches

.

Dataset Namely, Liar

.

To Suffer From

overfitting.

Different And Diverse

datasets.

Detection

.

Any Deep Learning

based model.

Models Along With

traditional ones.

Different Datasets

.

On Different Datasets

.

Sports, Entertainment

connecting BERT with RNN to tackle the impact of fake news . Lee et al. worked on hyperpartisan dataset and leveraged BERT on semi-supervised pseudo-label dataset .

Benchmark Studies

While most of the existing researches have focused on defining the types of fake news and suggesting different ap- proaches to detect them, very few studies are carried out to compare such approaches independently on different datasets. Among the categories, the benchmark-based studies are the most similar to our study. Table 1 compares our work with the previous benchmark-based studies along three themes: (1) experimental setup and results, (2) dataset length and diversity, and (3) range of models explored. We discuss the related work below.

Wang et al. compared the performance of SVM, LR, Bi-LSTM, and CNN models on their proposed dataset “LIAR” . Oshikawa et al. compared various machine learning models (e.g., SVM, CNN, LSTM) for fake news detection on different datasets . Gravanis et al. compared several traditional machine learning models (i.e., k-NN, Decision Tree, Naive Bayes, SVM, AdaBoost, Bagging) for fake news detection on different datasets . Dwivedi et al. presented a literature survey on various fake news detection methods . Zhang et al. presented a comprehensive overview of the existing datasets and approaches proposed for fake news detection in previous literature .

In summary, these few existing comparative studies lack in terms of the range of evaluated models and the diversity of the used datasets. Moreover, a complete exploration of the advanced pre-trained language models for fake news detection and comparison among them and with other models (i.e., traditional and deep learning) were missing in previous works. The benchmark study presented in this paper is focused on dealing with the above issues. We extend the state-of-the-art research in fake news detection by offering a comprehensive an in-depth study of 19 models (eight traditional shallow learning models, six traditional deep learning models, and five advanced pre-trained language models).

Study Setup

In this section, we first introduce the datasets used in our study and discuss how we preprocess those (Section 3). Then we discuss different features that we used in our models in Section 3. Finally, we discuss the traditional learning, deep learning and pre-trained models that we investigated in our study (Section 3). Finally, we discuss the performance metrics we used to evaluate the models and the train and test data settings in Section 3.

Studied Datasets

In this comparative study, we make use of three following datasets. Table 2 shows the detailed statistics of them. We describe the datasets below.

Liar

Liar1 is a publicly available dataset that has been used in . It includes 12.8K human-labeled short statements from POLITIFACT.COM. It comprises six labels of truthfulness ratings: pants-fire, false, barely-true, half-true, mostly-true, and true. In our work, we try to differentiate real news from all types of hoax, propaganda, satire, and misleading news.

Hence, we mainly focus on classifying news as real and fake. For the binary classification of news, we transform these labels into two labels. Pants-fire, false, barely-true are contemplated as fake and half-true, mostly-true, and true are as true. Our converted dataset contains 56% true and 44% fake statements. This dataset mostly deals with political issues that include statements of democrats and republicans, as well as a significant amount of posts from online social media.

The dataset provides some additional meta-data like the subject, speaker, job, state, party, context, history. However, in the real-life scenario, we may not have this meta-data always available. Therefore, we experiment on the texts of the dataset using textual features.

Fake Or Real News

Fake or real news dataset is developed by George McIntire. The fake news portion of this dataset was collected from Kaggle fake news dataset2 comprising news of the 2016 USA election cycle. The real news portion was collected from media organizations such as the New York Times, WSJ, Bloomberg, NPR, and the Guardian for the duration of 2015 or 2016. The GitHub repository of the dataset includes around 6.3k news with an equal allocation of fake and real news, and half of the corpus comes from political news.

Combined Corpus

Apart from the other two datasets, we have built a combined corpus that contains around 80k news among which 51% are real, and 49% are fake. One important property of this corpus is that it incorporates a wide range of topics including national and international politics, economy, investigation, health-care, sports, entertainment, and others. To demonstrate the topic diversity, we show the inter-topic distances3 of our combined corpus using LDA-based (Latent Dirichlet Allocation) topic modeling in Figure 1. Based on the empirical analysis of inter-topic distances, we divided the dataset into ten clusters (circles) where each cluster represents a topic. The coordinates of each topic cluster (circle) were measured following the MDS (Multidimensional Scaling) algorithm . X-axis (PC1) and Y-axis (PC2) maintained an aspect ratio to 1 to preserve the MDS distances. We used Jensen-Shannon divergence to compute distances between topics. The area of a cluster was calculated by the portion of tokens that respective topic generated compared to the total tokens in the corpus. We named the topic of a cluster based on the most relevant terms representing that cluster. The most relevant terms were determined on the basis of frequency. For example, the most relevant (i.e., most frequent) terms for cluster-7 are ‘Trump’, ‘Clinton’, ‘Election’, ‘Campaign’, etc (Figure 1). Hence, the news of this cluster represents the 2016 US election. On the other hand, the most relevant terms for cluster-3 are ’Bank’, ’Job’, ’Financial’, ’Tax’, ’Market’, etc. Thus, this cluster is related to the Economy. Additionally, overlapping of clusters (e.g., Economy and Politics) indicates shared relevant words (e.g., ‘Government’, ‘People’) between them.

We have collected news from several sources of the same time domain mostly from 2015 to 2017 4,5,6. Multiple types of fake news such as hoax, satire, and propaganda have come from The Onion, Borowitz Report, Clickhole, American News, DC Gazette, Natural News, and Activist Report. We have collected the real news from the trusted sources like the New York Times, Breitbart, CNN, Business Insider, the Atlantic, Fox News, Talking Points Memo, Buzzfeed News, National Review, New York Post, the Guardian, NPR, Gigaword News, Reuters, Vox, and the Washington Post.

Data Preprocessing

Before feeding into the models, raw texts of news required some preprocessing. We first eliminated unnecessary IP and URL addresses from our texts. The next step was to remove stop words. After that, we cleaned our corpus by correcting the spelling of words. We split every text by white-space and remove suffices from words by stemming 1https://www.cs.ucsb.edu/˜william/data/liar_dataset.zip

2Https://Www.Kaggle.Com/Mrisdal/Fake-News

3generated using pyLDAvis: https://pyldavis.readthedocs.io/ 4https://homes.cs.washington.edu/˜hrashkin/factcheck.html 5https://github.com/suryattheja/Fake-news-detection

6

Figure 1: Inter-topic distance map of Combined Corpus. them. Finally, we rejoined the word tokens by white-space to present our clean text corpus which had been tokenized later for feeding into the models.

Studied Features

We used lexical and sentiment features, n-gram, and Empath generated features for traditional machine learning mod- els, and pre-trained word embedding for deep learning models.

Lexical And Sentiment Features

Several studies have proposed to use lexical and sentiment features for fake news detection [50, 52, 57]. For lexical features, we used word count, average word length, article length, count of numbers, count of parts of speech, and count of exclamation mark. We calculated the sentiment (i.e., positive and negative polarity) of every article and used them as sentiment features.

N-Gram Feature

Word-based n-gram was used to represent the context of the document and generate features to classify the document as fake and real . We used both uni-gram and bi-gram features in this benchmark and evaluated their effectiveness.

Empath Generated Features

Empath is a tool that can generate lexical categories from a given text using a small set of seed terms . Using Empath, we calculated these categories (e.g., violence, crime, pride, sympathy, deception, war) for every news data and used them as features to identify key information in a news article. Since it has been used in literature for understanding deception in review systems , we feel motivated to investigate their contribution in this context.

Pre-Trained Word Embedding

For neural network models, word embeddings were initialized with 100-dimensional pre-trained embeddings from GloVe . GloVe is an unsupervised learning algorithm for obtaining vector representations for words. It was trained on a dataset of one billion tokens (words) with a vocabulary of 400 thousand words.

Studied Models

We experimented various traditional, deep learning and pre-trained language models in this work. Here, we describe all the models that we studied.

Traditional Machine Learning Models

We built our first three models using SVM (Support Vector Machine), LR (Logistic Regression), and Decision Tree with the lexical and sentiment features. Among the four main variants of the SVM kernel, we used the linear one. We also evaluated ensemble learning method like AdaBoost combining 30 decision trees with lexical and sentiment features. Next, we explored the Multinomial Naive Bayes classifier with the n-gram features. We used the Empath generated features with k-NN (k-Nearest Neighbors) classifier. We use the square-root of the total training data size as k as suggested by Lall and Sharma . Hence, the value of k was chosen to be 70, 90, and 250 for Liar, Fake or Real, and Combined Corpus respectively.

Deep Learning Models

In this study, we have evaluated six deep learning models for fake news detection including CNN, LSTM, Bi-LSTM, C-LSTM, HAN, and Convolutional HAN. The models are described below with their experimental setups. (1) CNN: One dimensional convolutional neural network can extract features and classify texts after transforming words in the sentence corpus into vectors .The one-dimensional convolutional model was initialized with 100- dimensional pre-trained GloVe embeddings. It contained 128 filters of filter size 3 and a max pooling layer of pool size 2 is selected. A dropout probability of 0.8 was preserved which was expunged for Combined Corpus. The model was compiled with ADAM optimizer with a learning rate of 0.001 to minimize binary cross-entropy loss. A sigmoid activation function was used for the final output layer. A batch size of 64 and 512 was used for training the datasets over 10 epochs.

(2) LSTM: Our LSTM model was pre-trained with 100-dimensional GloVe embeddings. The output dimension and time steps were set to 300. ADAM optimizer with learning rate 0.001 was applied to minimize binary cross-entropy loss. Sigmoid was the activation function for the final output layer. The model was trained over 10 epochs with batch size 64 and 512.

(3) Bi-LSTM: Usually, news that is deemed as fake is not fully comprised of false information, rather it is blended with true information. To detect the anomaly in a certain part of the news, we need to examine it both with previous and next events of action. We constructed a Bi-LSTM model to perform this task. Bi-LSTM was initialized with 100- dimensional pre-trained GloVe embeddings. The output dimension of 100 and time steps of 300 was applied. ADAM optimizer with a learning rate of 0.001 was used to minimize binary cross-entropy loss. The training batch size was set to 128 and loss over each epoch was observed with a callback. The learning rate was reduced by a factor of 0.1. We also used an early stop to monitor validation accuracy to check whether the accuracy was deteriorating for 5 epochs.

The loss of the binary cross-entropy of the model was minimized by ADAM with a learning rate of 0.0001. (4) C-LSTM: The C-LSTM based model contained one convolutional layer and one LSTM layer. We used 128 filters with filter size 3 on top of which a max pooling layer of pool size 2 was set. We fed it to our LSTM architecture with 100 output dimensions and dropout 0.2. Finally, we used sigmoid as the activation function of our output layer.

(5) HAN: We used a hierarchical attention network consisting of two attention mechanisms for word-level and sentence-level encoding. Before training, we set the maximum number of sentences in a news article as 20 and the maximum number of words in a sentence as 100. In both level encoding, a bidirectional GRU with output dimension 100 was fed to our customized attention layer. We used word encoder as input to our sentence encoder time-distributed layer. We optimized our model with ADAM that learned at a rate of 0.001.

(6) Convolutional HAN: In order to extract high-level features of the input, we incorporated a one-dimensional convolutional layer before each bidirectional GRU layer in HAN. This layer selected features of each tri-gram from the news article before feeding it to the attention layer.

8

Figure 2: Fine-tuning of pre-trained language models.

Advanced Language Models

Here, we first discuss the advanced language models that we used in this study and then describe their experimental setup. (1) BERT: BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained model which was designed to learn contextual word representations of unlabeled texts . Among the two versions of BERT (i.e., BERT-Base and BERT-Large) proposed originally, we used BERT-Base for this study considering the huge time and memory requirements of the BERT-Large model. The BERT-Base model has 12 layers (transformer blocks) with 12 attention heads and 110 million parameters.

(2) RoBERTa: RoBERTa (Robustly optimized BERT approach), originally suggested in , is the second pre- trained model that we experimented. It achieves better performance than original BERT models by using larger mini- batch sizes to train the model for a longer time over more data. It also removes the NSP loss in BERT and trains on longer sequences. Moreover, it dynamically changes the masking pattern applied to the training data.

(3) DistilBERT: DistilBERT is a smaller, faster, cheaper, and lighter version of original BERT which has 40% fewer parameters than the BERT-Base model. Though original BERT models perform better, DistilBERT is more appropriate for production-level usage due to its’ low resource requirements. Considering potential users of non-profit blogs and online media, we think low-resource models have a good appeal. Hence, this is worth investigating.

(4) ELECTRA: ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately) is a transformer model for self-supervised language representation learning. This model pre-trained with the use of another (small) masked language model. First, a language model takes an input text and randomly masked the text with generated input token. Then, ELECTRA models are trained to distinguish ”real” input tokens vs ”fake” input tokens generated by the former language model. At small scale, ELECTRA can achieve strong results even when trained on a single GPU.

(5) ELMo: ELMo (Embeddings from Language Models) is a contextualized word representation learned from a deep bidirectional language model that is trained on a large text corpus . We used the original pre-trained ELMo model proposed by the authors that has 2 bi-LSTM layers and 93.6 million parameters.

Experimental Setup Of Advanced Language Models:

We appended a classification head composed of a single linear layer on the top of the pre-trained advanced language models. The architecture of the classifier head is kept simple to focus on what information can readily be extracted from these pre-trained models. We used the respective pre-trained embeddings of the corresponding models (e.g., BERT embeddings, ELECTRA embeddings, ELMo embeddings) as the input of the classification heads and fine-tuned them for the fake news detection task (Figure 2). We trained them on all the datasets for 10 epochs with a mini-batch size of 32. We applied early stop to prevent our models from overfitting . Validation loss was considered as the metric of the early stopping while delta is set to zero . We set the maximum sequence length of the input data to 300.

For the Combined Corpus dataset, we configured the gradient accumulation steps as 2 due to the large dataset size. We used AdamW optimizer with the learning rate set to 4e-5, ß1 to 0.9, ß2 to 0.999, and epsilon to 1e-8 . Finally, we used binary cross-entropy to calculate the loss . We performed the experiments on NVIDIA Tesla T4 GPU provided by Google Colab.

Evaluation Metrics

We created a standard training and test set for each of the three datasets by splitting it in an 80:20 ratio so that different models can be evaluated on the same ground. For the first two datasets (i.e., Liar, Fake or Real), we did the split

9

randomly as they only contain one type of news. On the other hand, as the Combined Corpus covers a wide variety of topics, we took 80% (20%) data from each topic and include them in train (test) set to maintain a balanced distribution of every topic in training and test data.

We report the performance of each model in terms of accuracy, precision, recall, and F1-score. For precision, recall, and F1-score, we considered the macro-average of both class. In our experiment, we considered real news as ‘positive class’, and fake news as ‘negative class’. Hence, True Positive (TP) means the news is actually real, and also predicted as real while False Positive (FP) indicates that the news is actually false, but predicted as real. True Negative (TN) and False Negative (FN) imply accordingly. Accuracy is the number of correctly predicted instances out of all instances.

(1)

Precision is the ratio between the number of correctly predicted instances and all the predicted instances for a given class. For real and fake classes, we presented this metric as P(R) and P(F) respectively. Hence, the macro-average precision, P will be the average of P(R) and P(F).

(2)

Recall represents the ratio of the number of correctly predicted instances and all instances belonging to a given class. For real and fake classes, we presented this metric as R(R) and R(F) respectively. Hence, the macro-average recall, R will be the average of R(R) and R(F).

(3)

F1-score is the harmonic mean of the precision and recall.

Study Results

In this section, we answer three research questions: RQ1. How accurate are the traditional and deep learning models to detect fake news in our datasets? RQ2. Can the advanced pre-trained language models outperform the traditional and deep learning models? RQ3. Which model performs best with small training data? Previous studies on fake news detection mainly focused on traditional machine learning models. Therefore, it is im- portant to compare their performance with the deep learning models. We address this concern in RQ1. In particular, the goal of RQ1 is to compare the performance of different traditional machine learning models (e.g., SVM, Naive Bayes, Decision Tree) and deep learning models (e.g., CNN, LSTM, Bi-LSTM) on fake news detection. Consider- ing the great success of pre-trained advanced language models on various text classification tasks, it is important to investigate how these models perform on fake news detection compared to the traditional and deep learning models.

The answers to RQ2 will offer insights into whether and how the pre-trained advanced language models are useful to detect fake news. A common issue for any supervised learning problem is the limitation of labeled data. Intuitively, the more performance we can get with less amount of labeled data, the easier it would be to investigate and develop machine learning models to facilitate fake news detection. Therefore, as part of RQ3, we investigate the performance of the models we used in our study on smaller samples of our datasets.

10

Table 3: Performance of Traditional Machine Learning Models

.70

.70

.70

How accurate are the traditional and deep learning models to detect fake news in our datasets?

(Rq1)

In Table 3, we report the performances of various traditional machine learning models in detecting fake news. We observe that among the traditional machine learning models, Naive Bayes with n-gram features performs the best with 93% accuracy on our Combined Corpus. We also find that the addition of sentiment features with lexical features does not improve the performance considerably. For lexical and sentiment features, SVM and LR models perform better than other traditional machine learning models as suggested by most of the prior studies [10, 52, 60, 63, 64]. On the other hand, Empath generated features do not show promising performance for fake news detection, although they had been used earlier for understanding deception in review systems .

In Table 4, we report the performances of different deep learning models. The baseline CNN model is considered as the best model for Liar in , but we find it to be the second-best among all the models. LSTM-based models are most vulnerable to overfitting for this dataset which is reflected by its performance. Although Bi-LSTM is also a vic- tim of overfitting on the Liar dataset as mentioned in , we find it to be the third-best neural network-based model according to its performance on the dataset. The models successfully used for text classification like C-LSTM, HAN hardly surmount the overfitting problem for the Liar dataset. Our hybrid Conv-HAN model exhibits the best perfor- mance among the neural models for the Liar dataset with 0.59 accuracy and 0.59 F1-score. LSTM-based models show an improvement on the Fake or Real dataset whereas CNN and Conv-HAN continue their impressive performance.

LSTM-based models exhibit their best performance on our Combined Corpus where both Bi-LSTM and C-LSTM achieve 0.95 accuracy and 0.95 F1-score. CNN and all hierarchical attention models including Conv-HAN maintain a decent performance on this dataset with more than 0.90 accuracy and F1-score. This result indicates that, although neural network-based models may suffer from overfitting for a small dataset (LIAR), they show high accuracy and F1-score on a moderately large dataset (Combined Corpus).

We find that the traditional machine learning models are generally outperformed by the deep learning models in fake news detection, i.e., the overall accuracy of the traditional models is much lower than the deep learning ones (Table

Ai Adaptive Learning

This project focuses on ai adaptive learning using modern AI and machine learning techniques. The content below is adapted from research literature and practical implementation notes.

We propose a novel high-performance and interpretable canon-

addition, unlike tree learning, DNNs enable gradient descent- ical deep tabular data learning architecture, TabNet. TabNet based end-to-end learning for tabular data which can have a uses sequential attention to choose which features to reason multitude of benefits: (i) efficiently encoding multiple data from at each decision step, enabling interpretability and more types like images along with tabular data; (ii) alleviating the efficient learning as the learning capacity is used for the most need for feature engineering, which is currently a key aspect

salient features. We demonstrate that TabNet outperforms in tree-based tabular data learning methods; (iii) learning other variants on a wide range of non-performance-saturated from streaming data and perhaps most importantly (iv) end- tabular datasets and yields interpretable feature attributions to-end models allow representation learning which enables plus insights into its global behavior. Finally, we demonstrate many valuable application scenarios including data-efficient

self-supervised learning for tabular data, significantly improv- domain adaptation (Goodfellow, Bengio, and Courville 2016), ing performance when unlabeled data is abundant. generative modeling (Radford, Metz, and Chintala 2015) and

Introduction We propose a new canonical DNN architecture for tabular

Deep neural networks (DNNs) have shown notable success data, TabNet. The main contributions are summarized as: efficiently encode the raw data into meaningful representa- enabling flexible integration into end-to-end learning. tions, fuel the rapid progress. One data type that has yet to 2. TabNet uses sequential attention to choose which fea- see such success with a canonical architecture is tabular data. tures to reason from at each decision step, enabling in-

Despite being the most common data type in real-world AI terpretability and better learning as the learning capacity (as it is comprised of any categorical and numerical features), is used for the most salient features (see Fig. 1). This under-explored, with variants of ensemble decision trees for each input, and unlike other instance-wise feature se- Why? First, because DT-based approaches have certain bene- and van der Schaar 2019), TabNet employs a single deep

fits: (i) they are representionally efficient for decision mani- learning architecture for feature selection and reasoning. folds with approximately hyperplane boundaries which are 3. Above design choices lead to two valuable properties: (i) common in tabular data; and (ii) they are highly interpretable TabNet outperforms or is on par with other tabular learn- in their basic form (e.g. by tracking decision nodes) and there ing models on various datasets for classification and re-

are popular post-hoc explainability methods for their ensem- gression problems from different domains; and (ii) TabNet ble form, e.g. (Lundberg, Erion, and Lee 2018) – this is an enables two kinds of interpretability: local interpretability important concern in many real-world applications; (iii) they that visualizes the importance of features and how they are fast to train. Second, because previously-proposed DNN are combined, and global interpretability which quantifies

architectures are not well-suited for tabular data: e.g. stacked the contribution of each feature to the trained model. convolutional layers or multi-layer perceptrons (MLPs) are 4. Finally, for the first time for tabular data, we show signif- vastly overparametrized – the lack of appropriate inductive icant performance improvements by using unsupervised bias often causes them to fail to find optimal solutions for tab- pre-training to predict masked features (see Fig. 2).

ular decision manifolds (Goodfellow, Bengio, and Courville

Why is deep learning worth exploring for tabular data?

One obvious motivation is expected performance improve- Feature selection: Feature selection broadly refers to judi- Copyright © 2021, Association for the Advancement of Artificial ciously picking a subset of features based on their useful-

Professional occupation related Investment related

Feedback from Feedback to

Feature selection Input processing Feature selection Input processing

previous step next step … …

Predicted output (whether the income level >$50k)

selection enables interpretability and better learning as the capacity is used for the most salient features. TabNet employs multiple decision blocks that focus on processing a subset of input features for reasoning. Two decision blocks shown as examples process features that are related to professional occupation and investments, respectively, in order to predict the income level.

Unsupervised pre-training Supervised fine-tuning

Age Cap. gain Education Occupation Gender Relationship Age Cap. gain Education Occupation Gender Relationship 5 2000 ? Exec-managerial F Wife 6 2000 Bachelors Exec-managerial M Husband 1 0 ? Farming-fishing M ? 2 0 High-school Farming-fishing M Unmarried

? 50 Doctorate Prof-specialty M Husband 4 50 Doctorate Prof-specialty M Husband 2 ? ? Handlers-cleaners F Wife 2 0 High-school Handlers-cleaners F Wife 5 3000 Bachelors ? ? Husband 5 3000 Bachelors Exec-managerial M Husband

3 0 Bachelors ? F ? 3 100 Bachelors Prof-specialty F Wife ? 0 High-school Armed-Forces ? Husband 2 0 High-school Armed-Forces M Husband

TabNet decoder Decision making

Age Cap. gain Education Occupation Gender Relationship Income > $50k

3 M False

level can be guessed from the occupation, or the gender can be guessed from the relationship. Unsupervised representation learning by masked self-supervised learning results in an improved encoder model for the supervised learning task.

ward selection and Lasso regularization (Guyon and Elisseeff performance with compact representations. 2003) attribute feature importance based on the entire training Tree-based learning: DTs are commonly-used for tabular data, and are referred as global methods. Instance-wise fea- data learning. Their prominent strength is efficient picking ture selection refers to picking features individually for each of global features with the most statistical information gain

to maximize the mutual information between the selected mance of standard DTs, one common approach is ensembling features and the response variable, and in (Yoon, Jordon, and to reduce variance. Among ensembling methods, random van der Schaar 2019) by using an actor-critic framework to forests (Ho 1998) use random subsets of data with randomly mimic a baseline while optimizing the selection. Unlike these, selected features to grow many trees. XGBoost (Chen and

sity in end-to-end learning – a single model jointly performs recent ensemble DT approaches that dominate most of the feature selection and output mapping, resulting in superior recent data science competitions. Our experimental results

!# + Softmax !" < % !" > % !# > & !# > &

ReLU ReLU &

$" !" − $" % −1 −$" !" + $" % −1 −1 $# !# − $# & % −1 −$# !# + $# & !"

FC FC

W: [$" , - $" , 0, 0] W: [0, 0, $# , - $# ] !" < % b: [-a $" , a $" , -1, -1] b: [-1, -1, -d $# , d $# ] !# < & !" > % !# < & [!" ] [!# ]

M: [1, 0] M: [0, 1]

(right). Relevant features are selected by using multiplicative sparse masks on inputs. The selected features are linearly transformed, and after a bias addition (to represent boundaries) ReLU performs region selection by zeroing the regions. Aggregation of multiple regions is based on addition. As C and C get larger, the decision boundary gets sharper.

for various datasets show that tree-based models can be out- constructs a sequential multi-step architecture, where each performed when the representation capacity is improved with step contributes to a portion of the decision based on the deep learning while retaining their feature selecting property. selected features; (iii) improves the learning capacity via non- Integration of DNNs into DTs: Representing DTs with linear processing of the selected features; and (iv) mimics

DNN building blocks as in (Humbird, Peterson, and McClar- ensembling via higher dimensions and more steps. ren 2018) yields redundancy in representation and ineffi- cient learning. Soft (neural) DTs (Wang, Aggarwal, and Liu Fig. 4 shows the TabNet architecture for encoding tabu- functions, instead of non-differentiable axis-aligned splits. mapping of categorical features with trainable embeddings.

However, losing automatic feature selection often degrades We do not consider any global feature normalization, but performance. In (Yang, Morillo, and Hospedales 2018), a soft merely apply batch normalization (BN). We pass the same D- binning function is proposed to simulate DTs in DNNs, by dimensional features f ∈ <B×D to each decision step, where 2019) proposes a DNN architecture by explicitly leveraging multi-step processing with Nsteps decision steps. The ith

expressive feature combinations, however, learning is based step inputs the processed information from the (i − 1)th step on transferring knowledge from gradient-boosted DT. (Tanno to decide which features to use and outputs the processed ing from primitive blocks while representation learning into sion. The idea of top-down attention in the sequential form edges, routing functions and leaf nodes. TabNet differs from is inspired by its applications in processing visual and text

these as it embeds soft feature selection with controllable data (Hudson and Manning 2018) and reinforcement learn- Self-supervised learning: Unsupervised representation relevant information in high dimensional input. learning improves supervised learning especially in small Feature selection: We employ a learnable mask M[i] ∈ has shown significant advances – driven by the judicious capacity of a decision step is not wasted on irrelevant

choice of the unsupervised learning objective (masked input ones, and thus the model becomes more parameter effi- prediction) and attention-based deep learning. cient. The masking is multiplicative, M[i] · f . We use an attentive transformer (see Fig. 4) to obtain the masks us- TabNet for Tabular Learning ing the processed features from the preceding step, a[i − 1]:

M[i] = sparsemax(P[i − 1] · hi (a[i − 1])). Sparsemax nor-

DTs are successful for learning from real-world tabular malization (Martins and Astudillo 2016) encourages sparsity datasets. With a specific design, conventional DNN building by mapping the Euclidean projection onto the probabilistic blocks can be used to implement DT-like output manifold, simplex, which is observed to be superior in performance and e.g. see Fig. 3). In such a design, individual feature selec- aligned with the goal of sparse feature selection for explain-

tion is key to obtain decision boundaries in hyperplane form, PD which can be generalized to a linear combination of features ability. Note that j=1 M[i]b,j = 1. hi is a trainable func- where coefficients determine the proportion of each feature. tion, shown in Fig. 4 using a FC layer, followed by BN. P[i] TabNet is based on such functionality and it outperforms DTs is the prior scale term, denoting how much a particular feature

Qi while reaping their benefits by careful design which: (i) uses has been used previously: P[i] = j=1 (γ − M[j]), where γ sparse instance-wise feature selection learned from data; (ii) is a relaxation parameter – when γ = 1, a feature is enforced

+ Softmax

Feature Feature …

transformer transformer

x Nsteps Features

+ Softmax

Feature …

transformer transformer Feature Feature Feature Feature transformer

Encoded representation

transformer transformer Attentive transformer … Mask transformer …

Step 2 Decision step dependent

transformer transformer

BN Feature Feature

FC BN transformer transformer

+ 0.5 0.5 0.5 Agg. Agg. Features Features FC FC + +

Reconstructed + … Feature attributes + … features

(a) TabNet encoder architecture (b) TabNet decoder architecture Feature transformer Feature Attentive transformer Shared across decision steps Decision step dependent transformer GLU

Decision step dependent Prior scales

+ 0.5 0.5 0.5

0.5 0.5 0.5

+ Attentive transformer (c) (d)

Prior scales

divides the processed representation to be used by the attentive transformer of the subsequent step as well as for the overall Attentive BN FC

output. For each step, the feature selection mask provides interpretable information about the model’s functionality, and the +

masks can be aggregated to obtain global feature transformer important attribution. (b) TabNet decoder, composed of a feature transformer block at each step. (c) A feature transformer block example – 4-layer network is shown, where 2 are shared across all decision

Prior scales

steps and 2 are decision step-dependent. Each layer is composed of a fully-connected (FC) layer, BN and GLU nonlinearity. (d) +

An attentive transformer block example – a single layer mapping is modulated with a prior scale information which aggregates Sparsemax

how much each feature has been used before the current decision step. sparsemax (Martins and Astudillo 2016) is used for BN FC

normalization of the coefficients, resulting in sparse selection of the salient features. +

to be used only at one decision step and as γ increases, more propose the aggregate.feature importance mask, Magg−b,j = flexibility is provided to use a feature at multiple decision PNsteps ηb [i]Mb,j [i]

PD PNsteps

ηb [i]Mb,j [i].2 i=1 i=1 steps. P is initialized as all ones, 1B×D , without any prior j=1

on the masked features. If some features are unused (as in self- Tabular self-supervised learning: We propose a decoder supervised learning), corresponding P entries are made 0 architecture to reconstruct tabular features from the Tab- to help model’s learning. To further control the sparsity of the Net encoded representations. The decoder is composed of selected features, we propose sparsity regularization in the feature transformers, followed by FC layers at each deci-

form of entropy (Grandvalet and Bengio 2004), Lsparse = sion step. The outputs are summed to obtain the recon-

PNsteps PB PD −Mb,j [i] log(Mb,j [i]+)

i=1 b=1 j=1 Nsteps ·B , where  is a structed features. We propose the task of prediction of miss- small number for numerical stability. We add the sparsity reg- ing feature columns from the others. Consider a binary mask ularization to the overall loss, with a coefficient λsparse . Spar- S ∈ {0, 1}B×D . The TabNet encoder inputs (1 − S) · f̂ sity provides a favorable inductive bias for datasets where and the decoder outputs the reconstructed features, S · f̂ . We

most features are redundant. initialize P = (1 − S) in encoder so that the model em- Feature processing: We process the filtered features using phasizes merely on the known features, and the decoder’s last a feature transformer (see Fig. 4) and then split for the FC layer is multiplied with S to output the unknown features. decision step output and information for the subsequent We consider the reconstruction loss in self-supervised phase:

step, [d[i], a[i]] = fi (M[i] · f ), where d[i] ∈ <B×Nd and 2

PB PD (f̂b,j −fb,j )·Sb,j

a[i] ∈ <B×Na . For parameter-efficient and robust learning b=1 j=1

√ PB PB 2

. Normalization b=1 (fb,j −1/B b=1 fb,j ) with high capacity, a feature transformer should comprise layers that are shared across all decision steps (as the same with the population standard deviation of the ground truth features are input across different decision steps), as well as is beneficial, as the features may have different ranges. We decision step-dependent layers. Fig. 4 shows the implementa- sample Sb,j independently from a Bernoulli distribution with

tion as concatenation of two shared layers and two decision parameter ps , at each iteration. step-dependent layers. Each FC layer is followed by BN and eventually connected to a normalized residual √ connection We study TabNet in wide range of problems, that contain with normalization. Normalization with 0.5 helps to sta- regression or classification tasks, particularly with published bilize learning by ensuring that the variance throughout the benchmarks. For all datasets, categorical inputs are mapped

For faster training, we use large batch sizes with BN. Thus, bedding and numerical columns are input without and pre- except the one applied to the input features, we use ghost BN processing.4 We use standard classification (softmax cross (Hoffer, Hubara, and Soudry 2017) form, using a virtual batch entropy) and regression (mean squared error) loss functions size BV and momentum mB . For the input features, we ob- and we train until convergence. Hyperparameters of the Tab-

serve the benefit of low-variance averaging and hence avoid Net models are optimized on a validation set and listed in ghost BN. Finally, inspired by decision-tree like aggregation Appendix. TabNet performance is not very sensitive to most as in Fig. 3, we construct the overall decision embedding hyperparameters as shown with ablation studies in Appendix. as dout = i=1 PNsteps ReLU(d[i]). We apply a linear mapping In Appendix, we also present ablation studies on various de-

Wfinal dout to get the output mapping.1 sign and guidelines on selection of the key hyperparameters. Interpretability: TabNet’s feature selection masks can shed For all experiments we cite, we use the same training, val- light on the selected features at each step. If Mb,j [i] = 0, idation and testing data split with the original work. Adam optimization algorithm (Kingma and Ba 2014) and Glorot then j th feature of the bth sample should have no contribution uniform initialization are used for training of all models.5

to the decision. If fi were a linear function, the coefficient

Mb,j [i] would correspond to the feature importance of fb,j . Instance-wise feature selection

Although each decision step employs non-linear processing, their outputs are combined later in a linear way. We aim Selection of the salient features is crucial for high perfor- to quantify an aggregate feature importance in addition to mance, especially for small datasets. We consider 6 tabular requires a coefficient that can weigh the relative importance samples). The datasets are constructed in such a way that of each step in the decision. We simply propose ηb [i] = only a subset of the features determine the output. For Syn1-

PNd Syn3, salient features are same for all instances (e.g., the

c=1 ReLU(db,c [i]) to denote the aggregate decision con- tribution at ith decision step for the bth sample. Intuitively, if 2

Normalization is used to ensure D

P j=1 Magg−b,j = 1. db,c [i] < 0, then all features at ith decision step should have 3

0 contribution to the overall decision. As its value increases, prove the performance, but interpretation of individual dimensions

it plays a higher role in the overall linear combination. Scal- may become challenging. ing the decision mask at each decision step with ηb [i], we Specially-designed feature engineering, e.g. logarithmic trans- formation of variables highly-skewed distributions, may further

For discrete outputs, we additionally employ softmax during

training (and argmax during inference). An open-source implementation will be released.

Global: using only globally-salient features, Tree Ensembles (Geurts, Ernst, and Wehenkel 2006), Lasso-regularized model, L2X

Syn Syn Syn Syn Syn Syn

No selection .5 ± .0 .7 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .6 ± .0 Tree .5 ± .1 .8 ± .0 .8 ± .0 .6 ± .0 .7 ± .0 .7 ± .0 Lasso-regularized .4 ± .0 .5 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .7 ± .0

INVASE .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0

Global .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0 TabNet .6 ± .0 .8 ± .0 .8 ± .0 .7 ± .0 .7 ± .0 .8 ± .0

output of Syn depends on features X -X ), and global fea- Table 3: Performance for Poker Hand induction dataset. ture selection, as if the salient features were known, would give high performance. For Syn4-Syn6, salient features are Model Test accuracy (%) instance dependent (e.g., for Syn4, the output depends on ei- DT 50.0 ther X -X or X -X depending on the value of X ), which MLP 50.0

makes global feature selection suboptimal. Table 1 shows that Deep neural DT 65.1

TabNet outperforms others (Tree Ensembles (Geurts, Ernst, XGBoost 71.1

and Wehenkel 2006), LASSO regularization, L2X (Chen LightGBM 70.0 van der Schaar 2019). For Syn1-Syn3, TabNet performance TabNet 99.2 is close to global feature selection - it can figure out what Rule-based 100.0 features are globally important. For Syn4-Syn6, eliminating instance-wise redundant features, TabNet improves global feature selection. All other methods utilize a predictive model Poker Hand (Dua and Graff 2017): The task is classifica-

with 43k parameters, and the total number of parameters is tion of the poker hand from the raw suit and rank attributes of 101k for INVASE due to the two other models in the actor- the cards. The input-output relationship is deterministic and critic framework. TabNet is a single architecture, and its size hand-crafted rules can get 100% accuracy. Yet, conventional is 26k for Syn1-Syn and 31k for Syn4-Syn6. The compact DNNs, DTs, and even their hybrid variant of deep neural DTs

representation is one of TabNet’s valuable properties. (Yang, Morillo, and Hospedales 2018) severely suffer from the imbalanced data and cannot learn the required sorting and Performance on real-world datasets ranking operations (Yang, Morillo, and Hospedales 2018).

Tuned XGBoost, CatBoost, and LightGBM show very slight

as it can perform highly-nonlinear processing with its depth, Model Test accuracy (%) without overfitting thanks to instance-wise feature selection.

CatBoost 85.1 Table 4: Performance on Sarcos dataset. Three TabNet mod-

AutoML Tables 94.9 els of different sizes are considered.

Forest Cover Type (Dua and Graff 2017): The task is clas- MLP 2.1 0.14M

sification of forest cover type from cartographic variables. Adaptive neural tree 1.2 0.60M approaches that are known to achieve solid performance (AutoML 2019), an automated search framework based on TabNet-M 0.2 0.59M ensemble of models including DNN, gradient boosted DT, TabNet-L 0.1 1.75M with very thorough hyperparameter search. A single TabNet without fine-grained hyperparameter search outperforms it. Sarcos (Vijayakumar and Schaal 2000): The task is re-

gressing inverse dynamics of an anthropomorphic robot arm.

very small model is possible with a random forest. In the very and TabNet merely focuses on the relevant ones. For Syn4, small model size regime, TabNet’s performance is on par the output depends on either X -X or X -X depending parameters. When the model size is not constrained, TabNet feature selection – it allocates a mask to focus on the indi- achieves almost an order of magnitude lower test MSE. cator X , and assigns almost all-zero weights to irrelevant

features (the ones other than two feature groups). models are denoted with -S and -M. Real-world datasets: We first consider the simple task of mushroom edibility prediction (Dua and Graff 2017). Tab- Model Test acc. (%) Model size Net achieves 100% test accuracy on this dataset. It is indeed Sparse evolutionary MLP 78.4 81K known (Dua and Graff 2017) that “Odor” is the most discrim-

What is this project about?

This project covers practical implementation and research aspects of the topic using AI/ML techniques.