Enquire Now
Medical Computer Vision · Clinical Diagnostics · PyTorch / TensorFlow · GPU Optimized · 2026

Speaker Verification Xvector

Tensor Pipeline · Custom Loss Formulations · Model Quantization · Accelerated Inference — A rigorous deep learning engineering project focused on automated pathological lesion segmentation and radiological disease classification. Architected for thesis defense viva presentations, IEEE reproduction, and high-throughput production deployment.

PyTorch
Core Framework
AMP FP16
Mixed Precision
TensorRT
Quantized Serving

Ai Activity Recognition

This project focuses on ai activity recognition using modern AI and machine learning techniques. The content below is adapted from research literature and practical implementation notes.

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

Defying forgetting in classification tasks

Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia,

Aleš Leonardis, Gregory Slabaugh, Tinne Tuytelaars

Abstract—Artificial neural networks thrive in solving the classification problem for a particular rigid task, acquiring knowledge through

generalized learning behaviour from a distinct training phase. The resulting network resembles a static entity of knowledge, with endeavours to extend this knowledge without targeting the original task resulting in a catastrophic forgetting. Continual learning shifts

this paradigm towards networks that can continually accumulate knowledge over different tasks without the need to retrain from scratch. We focus on task incremental classification, where tasks arrive sequentially and are delineated by clear boundaries. Our main contributions concern (1) a taxonomy and extensive overview of the state-of-the-art; (2) a novel framework to continually determine the stability-plasticity trade-off of the continual learner; (3) a comprehensive experimental comparison of 1 state-of-the-art continual

learning methods and 4 baselines. We empirically scrutinize method strengths and weaknesses on three benchmarks, considering Tiny Imagenet and large-scale unbalanced iNaturalist and a sequence of recognition datasets. We study the influence of model capacity, weight decay and dropout regularization, and the order in which the tasks are presented, and qualitatively compare methods in terms of required memory, computation time and storage.

Index Terms—Continual Learning, lifelong learning, task incremental learning, catastrophic forgetting, classification, neural networks

1 I NTRODUCTION

In recent years, machine learning models have been Data can stem from changing input domains (e.g. varying reported to exhibit or even surpass human level perfor- imaging conditions) or can be associated with different mance on individual tasks, such as Atari games or object tasks (e.g. fine-grained classification problems). Continual recognition . While these results are impressive, they are learning is also referred to as lifelong learning , , ,

obtained with static models incapable of adapting their be- , , , sequential learning , , or incremen- havior over time. As such, this requires restarting the train- tal learning , , , , , , . The main ing process each time new data becomes available. In our criterion is the sequential nature of the learning process, dynamic world, this practice quickly becomes intractable for with only a small portion of input data from one or few

data streams or may only be available temporarily due to tasks available at once. The major challenge is to learn with- storage constraints or privacy issues. This calls for systems out catastrophic forgetting: performance on a previously that adapt continually and keep on learning over time. learned task or domain should not significantly degrade Human cognition exemplifies such systems, with a ten- over time as new tasks or domains are added. This is a direct

dency to learn concepts sequentially. Revisiting old concepts result of a more general problem in neural networks, namely by observing examples may occur, but is not essential to the stability-plasticity dilemma , with plasticity referring preserve this knowledge, and while humans may gradually to the ability of integrating new knowledge, and stability forget old information, a complete loss of previous knowl- retaining previous knowledge while encoding it. Albeit a

edge is rarely attested . challenging problem, progress in continual learning has led By contrast, artificial neural networks cannot learn in to real-world applications starting to emerge , , . this manner: they suffer from catastrophic forgetting of old concepts as new ones are learned . To circumvent this Scope. To keep focus, we limit the scope of our study in two

problem, research on artificial neural networks has focused ways. First, we only consider the task incremental setting, mostly on static tasks, with usually shuffled data to ensure where data arrives sequentially in batches and one batch i.i.d. conditions, and vast performance increase by revisiting corresponds to one task, such as a new set of categories training data over multiple epochs. to be learned. In other words, we assume for a given task,

Continual Learning studies the problem of learning from all data becomes available simultaneously for offline train- an infinite stream of data, with the goal of gradually extend- ing. This enables learning for multiple epochs over all its ing acquired knowledge and using it for future learning . training data, repeatedly shuffled to ensure i.i.d. conditions.

Importantly, data from previous or future tasks cannot be

• M. De Lange, R. Aljundi and T. Tuytelaars are with the Center for accessed. Optimizing for a new task in this setting will result Processing Speech and Images, Department Electrical Engineering, KU in catastrophic forgetting, with significant drops in perfor- • M. Masana is with Computer Vision Center, UAB. mance for old tasks, unless special measures are taken. The • S. Parisot, X. Jia, A. Leonardis, G. Slabaugh are with Huawei. efficacy of those measures, under different circumstances,

Manuscript received 6 Sep. 2019; revised 2 Jul. 2020. is exactly what this paper is about. Furthermore, task Log number TPAMI-2020-02-0259. incremental learning confines the scope to a multi-head con-

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

figuration, with an exclusive output layer or head for each of the common practice to use one dataset with different task. This is in contrast to the even more challenging class pixel permutations (typically permuted MNIST ). In incremental setup with all tasks sharing a single head. This addition, they stress that task incremental learning under introduces additional interference in learning and increases multi-head settings hides the true difficulty of the problem.

the number of output nodes to choose from. Instead, we While we agree with this statement, we still opted for a assume it is known which task a given sample belongs to. multi-head based survey, as it allows comparing existing Second, we focus on classification problems only, as methods without major modifications. They also propose classification is arguably one of the most established tasks a couple of desiderata for evaluating continual learning

for artificial neural networks, with good performance using methods. However, their study is limited to the fairly easy relatively simple, standard and well understood network MNIST and Fashion-MNIST datasets with far from realistic architectures. The setup is described in more detail in data encountered in practical applications. None of these Section 2, with Section 7 discussing open issues towards works systematically addresses the questions raised above.

tackling a more general setting. Paper overview. First, we describe in Section 2 the widely Motivation. There is limited consensus in the literature adopted task incremental setting for this paper. Next, on experimental setups and datasets to be used. Although Section 3 surveys different approaches towards continual papers provide evidence for at least one specific setting of learning, structuring them into three main groups: replay,

model architecture, combination of tasks and hyperparam- regularization-based and parameter isolation methods. An eters under which the proposed method reduces forgetting important issue when aiming for a fair method comparison and outperforms alternative approaches, there is no com- is the selection of hyperparameters, with in particular the prehensive experimental comparison performed to date. learning rate and the stability-plasticity trade-off. Hence,

Findings. To establish a fair comparison, the following Section 4 introduces a novel framework to tackle this prob- question needs to be considered: “How can the trade-off be- lem without requiring access to data from previous or future tween stability and plasticity be set in a consistent manner, tasks. We provide details on the methods selected for our using only data from the current task?” For this purpose, experimental evaluation in Section 5, and describe the actual

we propose a generalizing framework that dynamically experiments, main findings, and qualitative comparison in determines stability and plasticity for continual learning Section 6. We look further ahead in Section 7, highlighting methods. Within this principled framework, we scrutinize additional challenges in the field, moving beyond the task 1 representative approaches in the spectrum of continual incremental setting towards true continual learning. Then,

learning methods, and find method-specific preferences for we emphasize the relation with other fields in Section 8, and configurations regarding model capacity and combinations conclude in Section 9. Finally, implementation details and with typical regularization schemes such as weight decay or further experimental results are provided in supplemental dropout. All representative methods generalize well to the material. Code is publicly available for reproducibility 1 .

nicely balanced small-scale Tiny Imagenet setup. However, this no longer holds when transferring to large-scale real-

2 T HE TASK I NCREMENTAL L EARNING S ETTING

world datasets, such as the profoundly unbalanced iNatu- ralist, and especially in the unevenly distributed RecogSeq Due to the general difficulty and variety of challenges in sequence with highly dissimilar recognition tasks. Never- continual learning, many methods relax the general setting theless, rigorous methods based on isolation of parameters to an easier task incremental one. The task incremental set-

seem to withstand generalizing to these challenging set- ting considers a sequence of tasks, receiving training data of tings. Additionally, we find the ordering in which tasks are just one task at a time to perform training until convergence. presented to have insignificant impact on performance. A Data (X (t) , Y (t) ) is randomly drawn from distribution D(t) , summary of our main findings can be found in Section 6.8. with X (t) a set of data samples for task t, and Y (t) the

Related work. Continual learning has been the subject of corresponding ground truth labels. The goal is to control describe a wide range of methods and approaches, yet to data (X (t) , Y (t) ) from previous tasks t < T : without an empirical evaluation or comparison. In the same T X

E(X (t) ,Y (t) ) [ℓ(ft (X (t) ; θ), Y (t) )] (1)

continual learning, but with an emphasis on dynamic en- t=1 vironments for robotics. Pfülb and Gepperth perform an empirical study on catastrophic forgetting and develop with loss function ℓ, parameters θ, T the number of tasks a protocol for setting hyperparameters and method evalu- seen so far, and ft representing the network function for ation. However, they only consider two methods, namely task t. The definition can be extended to zero-shot learning

Elastic Weight Consolidation (EWC) and Incremental for t > T , but mainly remains future work in the field as Moment Matching (IMM) . Also, their evaluation is lim- discussed in Section 7. For the current task T , the statistical only three methods, EWC , PathNet and Gepp- NT Net . The experiments are performed on a simple fully 1 X (T ) (T )

ℓ(f (xi ; θ), yi ) . (2) connected network, including only 3 datasets: MNIST, CUB- NT i=1

2 and AudioSet. Farquhar and Gal survey continual

learning evaluation strategies and highlight shortcomings 1. Code available at: HTTPS :// GITHUB . COM / MATTDL /CL SURVEY

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

The main research focus in task incremental learning aims to Continual Learning Methods determine optimal parameters θ∗ by optimizing (1). How-

Replay Regularization-based Parameter isolation

ever, data is no longer available for old tasks, impeding methods methods methods evaluation of statistical risk for the new parameter values. Rehearsal Pseudo Constrained Prior-focused Data-focused Fixed Dynamic The concept ’task’ refers to an isolated training phase Rehearsal Network Architectures

iCaRL GEM EWC LwF with a new batch of data, belonging to a new group ER DGR A-GEM IMM LFL PackNet PNN of classes, a new domain, or a different output space. SER PR

TEM CCLUGM

GSS SI

R-EWC

EBLL PathNet Expert Gate

DMC Piggyback RCL

Similar to , a finer subdivision into three cate- CoPE LGM MAS

HAT DAN

gories emerges based on the marginal output and in- Walk

put distributions P (Y (t) ) and P (X (t) ) of a task t, with P (X (t) ) 6= P (X (t+1) ). First, class incremental learn- learning families of methods and the different branches ing defines an output space for all observed class labels within each family. The leaves enlist example methods. {Y (t) } ⊂ {Y (t+1) } with P (Y (t) ) 6= P (Y (t+1) ). Task incremen- tal learning defines {Y (t) } 6= {Y (t+1) }, which additionally requires task label t to indicate the isolated output nodes

Y (t) for current task t. Incremental domain learning defines for rehearsal, or to constrain optimization of the new task {Y (t) } = {Y (t+1) } with P (Y (t) ) = P (Y (t+1) ). Additionally, loss to prevent previous task interference. data incremental learning defines a more general contin- Rehearsal methods , , , explicitly retrain on ual learning setting for any data stream without notion of a limited subset of stored samples while training on new

task, class or domain. Our experiments focus on the task tasks. The performance of these methods is upper bounded incremental learning setting for a sequence of classification by joint training on previous and current tasks. Most no- tasks. table is class incremental learner iCaRL , storing a subset of exemplars per class, selected to best approximate class

means in the learned feature space. At test time, the class 3 C ONTINUAL L EARNING A PPROACHES means are calculated for nearest-mean classification based Early works observed the catastrophic interference prob- on all exemplars. In data incremental learning, Rolnick et lem when sequentially learning examples of different input al. suggest reservoir sampling to limit the number of

patterns. Several directions have been explored, such as stored samples to a fixed budget assuming an i.i.d. data reducing representation overlap , , , , , stream. Continual Prototype Evolution (CoPE) com- replaying samples or virtual samples from the past , bines the nearest-mean classifier approach with an efficient or introducing dual architectures , , . However, reservoir-based sampling scheme. Additional experiments

due to resource restrictions at the time, these works mainly on rehearsal for class incremental learning are provided considered few examples (in the order of tens) and were in . based on specific shallow architectures. While rehearsal might be prone to overfitting the subset With the recent increased interest in neural networks, of stored samples and seems to be bounded by joint training,

continual learning and catastrophic forgetting also received constrained optimization is an alternative solution leaving more attention. The influence of dropout and different more leeway for backward/forward transfer. As proposed activation functions on forgetting have been studied empir- in GEM under the task incremental setting, the key ically for two sequential tasks , and several works stud- idea is to only constrain new task updates to not interfere

ied task incremental learning from a theoretical perspective with previous tasks. This is achieved through projecting the , , , . estimated gradient direction on the feasible region outlined More recent works have addressed continual learning by previous task gradients through first order Taylor series with longer task sequences and a larger number of exam- approximation. A-GEM relaxes the problem to project on

ples. In the following, we will review the most important one direction estimated by randomly selected samples from specific information is stored and used throughout the se- solution to a pure online continual learning setting without quential learning process: task boundaries, proposing to select sample subsets that • Replay methods maximally approximate the feasible region of historical data.

• Regularization-based methods In the absence of previous samples, pseudo rehearsal is • Parameter isolation methods an alternative strategy used in early works with shallow neural networks. The output of previous model(s) given

Note that our categorization overlaps to some extent with

random inputs are used to approximate previous task sam- that introduced in previous work , . However, we ples . With deep networks and large input vectors (e.g. believe it offers a more general overview and covers most full resolution images), random input cannot cover the input existing works. A summary can be found in Figure 1. space . Recently, generative models have shown the ability to generate high quality images , , opening up 3.1 Replay methods possibilities to model the data generating distribution and

This line of work stores samples in raw format or generates retrain on generated examples . However, this also adds pseudo-samples with a generative model. These previous complexity in training generative models continually, with task samples are replayed while learning a new task to extra care to balance retrieved examples and avoid mode alleviate forgetting. They are either reused as model inputs collapse.

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

3.2 Regularization-based methods 4 C ONTINUAL H YPERPARAMETER F RAMEWORK

This line of works avoids storing raw inputs, prioritizing Methods tackling the continual learning problem typically privacy, and alleviating memory requirements. Instead, an involve extra hyperparameters to balance the stability- extra regularization term is introduced in the loss function, plasticity trade-off. These hyperparameters are in many consolidating previous knowledge when learning on new cases found via a grid search, using held-out validation

data. We can further divide these methods into data-focused data from all tasks. However, this inherently violates the and prior-focused methods. main assumption in continual learning, namely no access to previous task data. This may lead to overoptimistic results, that cannot be reproduced in a true continual learning 3.2.1 Data-focused methods setting. As it is of our concern in this survey to provide

a comprehensive and fair study on the different continual The basic building block in data-focused methods is knowl- learning methods, we need to establish a principled frame- edge distillation from a previous model (trained on a previ- work to set hyperparameters without violating the continual ous task) to the model being trained on the new data. Silver learning setting. Besides a fair comparison over the existing

given new task input images, mainly for improving new the stability-plasticity trade-off, and therefore extends its task performance. It has been re-introduced by LwF to value to real-world continual learners. mitigate forgetting and transfer knowledge, using the pre- First, our proposed framework assumes only access vious model output as soft labels for previous tasks. Other to new task data to comply with the continual learning

works , have been introduced with related ideas, paradigm. Next, the framework aims at the best trade-off however, it has been shown that this strategy is vulnerable in stability and plasticity in a dynamic fashion. For this to domain shift between tasks . In an attempt to overcome purpose, the set of hyperparameters H related to forgetting of shallow autoencoders to constrain task features in their we do for all compared methods in Section 5. The hyperpa-

corresponding learned low dimensional space. rameters are initialized to ensure minimal forgetting of pre- vious tasks. Then, if a predefined threshold on performance 3.2.2 Prior-focused methods is not achieved, hyperparameter values are decayed until reaching the desired performance on the new task validation To mitigate forgetting, prior-focused methods estimate a data. Algorithm 1 illustrates the two main phases of our

distribution over the model parameters, used as prior when framework, to be repeated for each new task: learning from new data. Typically, importance of all neural Maximal Plasticity Search first finetunes model copy ′ network parameters is estimated, with parameters assumed θt on new task data, given previous model parameters θt . independent to ensure feasibility. During training of later The learning rate η ∗ is obtained via a coarse grid search,

tasks, changes to important parameters are penalized. Elas- which aims for highest accuracy A∗ on a held-out validation tic weight consolidation (EWC) was the first to estab- set from the new task. The accuracy A∗ represents the lish this approach. Variational Continual Learning (VCL) best accuracy that can be achieved while disregarding the introduced a variational framework for this family , previous tasks.

sprouting a body of Bayesian based works , . Zenke Stability Decay. In the second phase we train θt with tance estimation, allowing increased flexibility and online to their highest values to ensure minimum forgetting. We user adaptation as in . Further work extends this to task define threshold p indicating tolerated drop in new task per- free settings . formance compared to finetuning A∗ . When this threshold is

not met, we decrease hyperparameter values in H by scalar multiplication with decay factor α, and subsequently repeat 3.3 Parameter isolation methods this phase. This corresponds to increasing model plasticity in order to reach the desired performance threshold. To This family dedicates different model parameters to each avoid redundant stability decay iterations, decayed hyper- task, to prevent any possible forgetting. When no constraints parameters are propagated to later tasks.

apply to architecture size, one can grow new branches for new tasks, while freezing previous task parameters ,

5 C OMPARED M ETHODS

, or dedicate a model copy to each task . Alternatively, the architecture remains static, with fixed parts allocated In Section 6, we carry out a comprehensive comparison to each task. Previous task parts are masked out during between representative methods from each of the three new task training, either imposed at parameters level , families of continual learning approaches introduced in Sec- , or unit level . These works typically require a tion 3. For clarity, we first provide a brief description of the

task oracle, activating corresponding masks or task branch selected methods, highlighting their main characteristics. during prediction. Therefore, they are restrained to a multi- head setup, incapable to cope with a shared head between 5.1 Replay methods tasks. Expert Gate avoids this problem through learning iCaRL was the first replay method, focused on learn- an auto-encoder gate. ing in a class-incremental way. Assuming fixed allocated

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

Algorithm 1 Continual Hyperparameter Selection Frame- in a gradual error build-up to the prior tasks as more work dissimilar tasks are added , . This error build-up input H hyperparameter set, α ∈ [0, 1] decaying factor, also applies in a class-incremental setup, as shown in . p ∈ [0, 1] accuracy drop margin, Dt+1 new task data, Another drawback resides in the additional overhead to

Ψ coarse learning rate grid forward all new task data, and storing the outputs. LwF require θt previous task model parameters is specifically designed for classification, but has also been require CLM continual learning method applied to other problems, such as object detection . // Maximal Plasticity Search Encoder Based Lifelong Learning (EBLL) extends

1: A∗ = 0 LwF by preserving important low dimensional feature rep- 2: for η ∈ Ψ do resentations of previous tasks. For each task, an under-

A ← Finetune Dt+1 , η; θt ⊲ Finetuning accuracy

 3: complete autoencoder is optimized end-to-end, projecting 4: if A > A∗ then features on a lower dimensional manifold. During train- 5: A∗ , η ∗ ← A, η ⊲ Update best values ing, an additional regularization term impedes the current // Stability Decay feature projections to deviate from previous task optimal

6: do ones. Although required memory grows linearly with the A ← CLM Dt+1 , η ∗ ; θt number of tasks, autoencoder size constitutes only a small  7: 8: if A < (1 − p)A∗ then fraction of the backbone network. The main computational 9: H←α·H ⊲ Hyperparameter decay overhead occurs in autoencoder training, and collecting

10: while A < (1 − p)A∗ feature projections for the samples in each optimization step.

Elastic Weight Consolidation (EWC) introduces

network parameter uncertainty in the Bayesian frame- memory, it selects and stores samples (exemplars) closest work . Following sequential Bayesian estimation, the to the feature mean of each class. During training, the old posterior of previous tasks t < T constitutes the prior estimated loss on new classes is minimized, along with for new task T , founding a mechanism to propagate old the distillation loss between targets obtained from previous task importance weights. The true posterior is intractable,

model predictions and current model predictions on the and is therefore estimated using a Laplace approximation previously learned classes. As the distillation loss strength with precision determined by the Fisher Information Matrix correlates with preservation of the previous knowledge, this (FIM). Near a minimum, this FIM shows equivalence to hyperparameter is optimized in our proposed framework. the positive semi-definite second order derivative of the

In our study, we consider the task incremental setting, and

loss , and is in practice typically approximated by the therefore implement iCarl also in a multi-head fashion to empirical FIM to avoid additional backward passes : perform a fair comparison with other methods. "  #

GEM exploits exemplars to solve a constrained op- T δL 2

timization problem, projecting the current task gradient in Ωk = E(x,y)∼DT , (3) δθkT a feasible area, outlined by the previous task gradients. The authors observe increased backward transfer by altering the with importance weight ΩTk calculated after training task gradient projection with a small constant γ ≥ 0, constituting T . Although originally requiring one FIM per task, this can

H in our framework. be resolved by propagating a single penalty . Further, The major drawback of replay methods is limited scal- the FIM is approximated after optimizing the task, inducing ability over the number of classes, requiring additional gradients close to zero, and hence very little regularization. computation and storage of raw input samples. Although This is inherently coped with in our framework, as the

fixing memory limits memory consumption, this also dete- regularization strength is initially very high and lowered riorates the ability of exemplar sets to represent the original only to decay stability. Variants of EWC are proposed to distribution. Additionally, storing these raw input samples address these issues in , and . may also lead to privacy issues. Synaptic Intelligence (SI) breaks the EWC

paradigm of determining the new task importance weights 5.2 Regularization-based methods ΩTk in a separate phase after training. Instead, they maintain an online estimate ω T during training to eventually attain

This survey strongly focuses on regularization-based meth-

T ods, comparing seven methods in this family. The regular- X ωkt ization strength correlates to the amount of knowledge re- ΩTk = , (4) t=1 (∆θkt )2 + ξ tention, and therefore constitutes H in our hyperparameter framework. with ∆θkt = θkt − θkt−1 the task-specific parameter distance,

Learning without Forgetting (LwF) retains knowl- and damping parameter ξ avoiding division by zero. Note edge of preceding tasks by means of knowledge distilla- that the accumulated importance weights ΩTk are still only tion . Before training the new task, network outputs for updated after training task T as in EWC. Further, the the new task data are recorded, and are subsequently used knife cuts both ways for efficient online calculation. First,

during training to distill prior task knowledge. However, stochastic gradient descent incurs noise in the approxi- the success of this method depends heavily on the new task mated gradient during training, and therefore the authors data and how strong it is related to prior tasks. Distribution state importance weights tend to be overestimated. Second, shifts with respect to the previously learned tasks can result catastrophic forgetting in a pretrained network becomes

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

inevitable, as importance weights can’t be retrieved. In decay, the capacity used for the new task can be minimized another work, Riemannian Walk combines the SI path instead to avoid capacity saturation and ensuring stable integral with an online version of EWC to measure parame- learning for future tasks. ter importance. PackNet iteratively assigns parameter subsets to

Memory Aware Synapses (MAS) redefines the pa- consecutive tasks by constituting binary masks. For this rameter importance measure to an unsupervised setting. purpose, new tasks establish two training phases. First, the Instead of calculating gradients of the loss function L as network is trained without altering previous task parameter in (3), the authors obtain gradients of the squared L -norm subsets. Subsequently, a portion of unimportant free param-

of the learned network output function fT : eters are pruned, measured by lowest magnitude. Then, the second training round retrains this remaining subset δ kfT (x; θ)k of important parameters. The pruning mask preserves task

ΩTk = Ex∼DT [ ]. (5)

δθkT performance as it ensures to fix the task parameter subset for Previously discussed methods require supervised data for future tasks. PackNet allows explicit allocation of network the loss-based importance weight estimations, and are there- capacity per task, and therefore inherently limits the total fore confined to available training data. By contrast, MAS number of tasks. H contains the per-layer pruning fraction.

enables importance weight estimation on an unsupervised HAT requires only one training phase, incorporating held-out dataset, hence capable of user-specific data adap- task-specific embeddings for attention masking. The per- tation. layer embeddings are gated through a Sigmoid to attain

Incremental Moment Matching (IMM) estimates unit-based attention masks in the forward pass. The Sig-

Gaussian posteriors for task parameters, in the same vein moid slope is rectified through training of each epoch, as EWC, but inherently differs in its use of model merging. initially allowing mask modifications and culminating into In the merging step, the mixture of Gaussian posteriors near-binary masks. To facilitate capacity for further tasks, a is approximated by a single Gaussian distribution, i.e. a regularization term imposes sparsity on the new task atten-

new set of merged parameters θ1:T and corresponding tion mask. Core to this method is constraining parameter covariances Σ1:T . Although the merging strategy implies updates between two units deemed important for previous a single merged model for deployment, it requires storing tasks, based on the attention masks. Both regularization models during training for each learned task. In their work, strength and Sigmoid slope are considered in H. We only

two methods for the merge step are proposed: mean-IMM report results on Tiny Imagenet as difficulties regarding and mode-IMM. In the former, weights θk of task-specific asymmetric capacity and hyperparameter sensitivity pre- networks are averaged following the weighted sum: vent using HAT for more difficult setups . We refer to Appendix B.5 and B.6 for a detailed analysis on capacity XT

θk1:T = αtk θkt , (6) usage of parameter isolation methods. t PT with αtk the mixing ratio of task t, subject to αtk = 1. t

Alternatively, the second merging method mode-IMM aims

6 E XPERIMENTS

for the mode of the Gaussian mixture, with importance In this section we first discuss the experimental setup in weights obtained from (3): Section 6.1, followed by a comparison of all the methods on a common BASE model in Section 6.2. The effects of changing

1 XT t t t XT

θk1:T = 1:T αk Ωk θk , Ω1:T k = αtk Ωtk . (7) the capacity of this model are discussed in Section 6.3. Next,

Ωk t t

in Section 6.4 we look at the effect of two popular methods When two models converge to a different local minimum for regularization. We continue in Section 6.5, scrutinizing due to independent initialization, simply averaging the the behaviour of continual learning methods in a real- models might result in increased loss, as there are no guar- world setup, abandoning the artificially imposed balance

antees for a flat or convex loss surface between the two between tasks. In addition, we investigate the effect of the points in parameter space . Therefore, IMM suggests task ordering in both the balanced and unbalanced setup three transfer techniques aiming for an optimal solution in Section 6.6, and elucidate a qualitative comparison in in the interpolation of the task-specific models: i) Weight- Table 9 of Section 6.7. Finally, Section 6.8 summarizes our

Transfer initializes the new task network with previous task main findings in Table 10. parameters; ii) Drop-Transfer is a variant of dropout with previous task parameters as zero point; iii) L2-transfer 6.1 Experimental Setup is a variant of L2-regularization, again with previous task parameters redefining the zero point. In this study, we com- Datasets. We conduct image classification experiments on pare mean-IMM and mode-IMM with both weight-transfer three datasets, the main characteristics of which are summa-

and L2-transfer. As mean-IMM is consistently outperformed rized in Table 1. First, we use the Tiny Imagenet dataset . by mode-IMM, we refer for its full results to appendix. This is a subset of 2 classes from ImageNet , rescaled to image size 6 × 64. Each class contains 5 samples subdivided into training (80%) and validation (20%), and 5.3 Parameter isolation methods 5 samples for evaluation. In order to construct a balanced

This family of methods isolates parameters for specific tasks dataset, we assign an equal amount of 2 randomly chosen and can guarantee maximal stability by fixing the parameter classes to each task in a sequence of 1 consecutive tasks. subsets of previous tasks. Although this refrains stability This task incremental setting allows using an oracle at test

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

iNaturalist and RecogSeq dataset characteristics. layers. The models are based on a VGG configuration , but with less parameters due to the small image size. We

Tiny Imagenet iNaturalist RecogSeq

reduce the feature extractor to comprise 4 max-pooling

Tasks 1 1 8 layers, each preceded by a stack of identical convolutional

Classes/task 2 5 to 3 1 to 2

Train data/task 8k 0.6k to 66k 2k to 73k

layers with a consistent 3 × 3 receptive field. The first max- Val. data/task 1k 0.1k to 9k 0.6k to 13k pooling layer is preceded by one conv. layer with 6 filters. Task selection random class supercategory dataset Depending on the model, we increase subsequent conv. layer stacks with a multiple of factor 2. The models have that VOC Actions represents the human action classification 1 units for the SMALL model and 5 units for the other

subset of the VOC challenge 20 . three models. The multi-head setting imposes a separate fully connected layer with softmax for each task, with the

Task Classes Samples number of outputs defined by the classes in that task. The

Train Val Test detailed description can be found in Appendix A.

Oxford Flowers 1 20 30 30 The size of iNaturalist and RecogSeq impose arduous

MIT Scenes 6 53 6 6 learning. Therefore, we conduct the experiments solely for

Caltech-UCSD Birds 2 59 28 28

Stanford Cars 1 81 40 40 AlexNet , pretrained on ImageNet.

FGVC-Aircraft 1 66 16 16 Evaluation Metrics. To measure performance in the con-

VOC Actions 1 31 15 15

Letters 5 68 5 5 tinual learning setup, we evaluate accuracy and forgetting

SVHN 1 732 130 130 per task, after training each task. We define this measure of forgetting as the difference between the expectation of acquired knowledge of a task, i.e. the accuracy when first time for our evaluation per task, ensuring all tasks are learning a task, and the accuracy obtained after training one roughly similar in terms of difficulty, size, and distribution, or more additional tasks. In the figures, we focus on evolu-

making the interpretation of the results easier. tion of accuracy for each task as more tasks are added. In the The second dataset is based on iNaturalist , which tables, we report average accuracy and average forgetting aims for a more real-world setting with a large number of on the final model, obtained by evaluating each task after fine-grained categories and highly imbalanced classes. On learning the entire task sequence.

top, we impose task imbalance and domain shifts between Baselines. The discussed continual learning methods in tasks by assigning 1 super-categories of species as separate Section 5 are compared against several baselines: tasks. We selected the most balanced 1 super-categories from the total of 1 and only retained categories with at least 1) Finetuning starts form the previous task model 1 samples. More details on the statistics for each of these to optimize current task parameters. This baseline

tasks can be found in Table 8. We only utilize the training greedily trains each task without considering pre- data, subdivided in training (70%), validation (20%) and vious task performance, hence introducing catas- evaluation (10%) sets, with all images measuring 8 × 600. trophic forgetting, and representing the minimum Thirdly, we adopt a sequence of 8 highly diverse recog- desired performance.

2) Joint training considers all data in the task sequence nition tasks (RecogSeq) as in and , which also targets simultaneously, hence violating the continual learn- distributing an imbalanced number of classes, and reaches ing setup (indicated with appended ’∗’ in reported beyond object recognition with scene and action recogni-

results). This baseline provides a target reference

tion. This sequence is composed of 8 consecutive datasets, from fine-grained to coarse classification of objects, actions performance. and scenes, going from flowers, scenes, birds and cars, to For the replay methods we consider two additional fine- aircrafts, actions, letters and digits. Details are provided in tuning baselines, extending baseline (1) with the benefit of

In this survey we scrutinize the effects of different task

orderings for both Tiny Imagenet and iNaturalist in Sec- 3) Basic rehearsal with Full Memory (R-FM) fully tion 6.6. Apart from that section, discussed results on both exploits total available exemplar memory R, incre- datasets are performed on a random ordering of tasks, and mentally dividing equal capacity over all previous the fixed dataset ordering for RecogSeq. tasks. This is a baseline for replay methods defining

memory management policies to exploit all memory Models. We summarize in Table 3 the models used for (e.g. iCaRL). experiments in this work. Due to the limited size of Tiny Im- 4) Basic rehearsal with Partial Memory (R-PM) pre- agenet we can easily run experiments with different models. allocates fixed exemplar memory R/T over all This allows to analyze the influence of model capacity (Sec- tasks, assuming the amount of tasks T is known be-

tion 6.3) and regularization for each model configuration forehand. This is used by methods lacking memory (Section 6.4). This is important, as the effect of model size management policies (e.g. GEM). and architecture on the performance of different continual learning methods has not received much attention so far. We Replay Buffers. Replay methods (GEM, iCaRL) and cor- configure a BASE model, two models with less (SMALL) and responding baselines (R-PM, R-FM) can be configured with

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

Model Feature Extractor Classifier (w/o head) Total Parameters Pretrained Multi-head

Conv. Layers MaxPool Parameters FC layers Parameters

Tiny Imagenet SMALL 6 4 334k 2 279k 613k ✗ X

BASE 6 4 1.15m 2 2.36k 3.51m ✗ X

WIDE 6 4 4.5m 2 4.46m 8.95m ✗ X

DEEP 2 4 4.28m 2 2.36k 6.64m ✗ X

iNaturalist/ RecogSeq AlexNet 5 3 2.47m 2 (with Dropout) 54.5m 57.0m X(Imagenet) X

arbitrary replay buffer size. Too large buffers would result in 6.2.1 Discussion unfair comparison to regularization and parameter isolation methods, with in its limit even holding all data for the General observations. As a reference, it is relevant to highlight previous tasks, as in joint learning, which is not compli- the soft upper bound obtained when training all tasks ant with the continual learning setup. Therefore, we use jointly. This is indicated by the star symbol in each of the

the memory required to store the BASE model as a basic subpanels. For the Tiny Imagenet dataset, all tasks have reference for the amount of exemplars. This corresponds the same number of classes, same amount of training data to the additional memory needed to propagate importance and similar level of difficulty, resulting in similar accuracies weights in the prior-focused methods (EWC, SI, MAS). For for all tasks under the joint training scheme. The average

Tiny Imagenet, this gives a total replay buffer capacity of accuracy for joint training on this dataset is 55.70%, while 4.5k exemplars. We also experiment with a buffer of 9k random guessing would result in 5%. exemplars to examine the influence of increased buffer ca- Further, as reported ample times in literature, the fine- pacity. Note that the BASE reference implies replay methods tuning baseline suffers severely from catastrophic forget-

to use a significantly different amount of memory than the ting: initially good results are obtained when learning a non-replay based methods; more memory for the SMALL task, but as soon as a new task is added, performance model), and less memory for the WIDE and DEEP models. drops, resulting in poor average accuracy of only 21.30% Comparisons between replay and non-replay based meth- and 26.90% average forgetting.

ods should thus only be done for the BASE model. In the With 49.13% average accuracy, PackNet shows highest following, notation of the methods without replay buffer overall performance after training all tasks. When learning size refers to the default 4.5k buffer. new tasks, due to compression PackNet only has a fraction of the total model capacity available. Therefore, it typically

Learning details. During training the models are optimized

performs worse on new tasks compared to other methods. with stochastic gradient descent with a momentum of 0.9.

However, by fixing task parameters through masking, it

Training lasts for 7 epochs unless preempted by early

allows complete knowledge retention until the end of the stopping. Standard AlexNet is configured with dropout task sequence (no forgetting yields flat curves in the figure), for iNaturalist and RecogSeq. For Tiny Imagenet we do not resulting in highest accumulation of knowledge. This holds use any form of regularization by default, except for Sec- at least when working with long sequences, where forget- tion 6.4 explicitly scrutinizing regularization influence. The ting errors gradually build up for other methods.

framework we proposed in Section 4 can only be exploited in the case of forgetting-related hyperparameters, in which MAS and iCaRL show competitive results w.r.t. PackNet case we set p = 0.2 and α = 0.5. All methods discussed in (resp. 46.90% and 47.27%). iCaRL starts with a significantly Section 5 satisfy this requirement, except for IMM with L lower accuracy for new tasks, due to its nearest-neighbour

transfer for which we could not identify a specific hyper- based classification, but improves over time, especially for parameter related to forgetting. For specific implementation the first few tasks. Further, doubling replay-buffer size to 9k details about the continual learning methods we refer to enhances iCaRL performance to 48.76%. Appendix A. Regularization-based methods. In prior experiments where we

did a grid search over a range of hyperparameters over the whole sequence (rather than using the framework in- troduced in Section 4), we observed MAS to be remarkably 6.2 Comparing Methods on the B ASE Network more robust to the choice of hyperparameter values com- pared to the two related methods EWC and SI. Switching Tiny Imagenet. We start evaluation of continual learning to the continual hyperparameter selection framework de-

methods discussed in Section 5 with a comparison using the scribed in Section 4, this robustness leads to superior results BASE network, on the Tiny Imagenet dataset, with random for MAS with 46.90% compared to the other two (42.43% ordering of tasks. This provides a balanced setup, making and 33.93%). Especially improved forgetting stands out interpretation of results easier. (1.58% vs 7.51% and 15.77%). Further, SI underperforms

ing test accuracy evolution for a specific task (e.g. Task 1 the BASE model (see results Appendix B.2). SI inherently for the leftmost panel) as more tasks are added for training. constrains parameter importance estimation to the training Since the n-th task is added for training only after n steps, data only, which is opposed to EWC and MAS able to curves get shorter as we move to subpanels on the right. determine parameter importance both on validation and

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

training data in a separate phase after task training. Further, regularization, which does not guarantee uniform capacity mode-IMM catastrophically forgets in evaluation on the first over layers and tasks. task, but shows transitory recovering through backward transfer in all subsequent tasks. Overall for Tiny Imagenet, 6.3 Effects of Model Capacity

IMM does not seem competitive with the other strategies for

continual learning - especially if one takes into account that Tiny Imagenet. A crucial design choice in continual learning they need to store all previous task models, making them concerns network capacity. Aiming to learn a long task much more expensive in terms of storage. sequence, a high capacity model seems preferable. However, The data-driven methods LwF and EBLL obtain similar learning the first task using such model, with only data

results as EWC (41.91% and 45.34% vs 42.43%). EBLL from a single task, holds the risk of overfitting, jeopardiz- improves over LwF by residing closer to the optimal task ing generalization performance. So far, we compared all representations, lowering average forgetting and improving methods on the BASE model. Next, we study effects of accuracy. Apart from the first few tasks, curves for these extending or reducing capacity in this architecture. Details

methods are quite flat, indicating low levels of forgetting. of the models are given in Section 6.1 and Appendix A.

In the following, we discuss results of model choice, again

Replay methods. iCaRL starts from the same model as GEM using the random ordering of Tiny Imagenet. These results and the regularization-based methods, but uses a feature are summarized in the top part of Table 4 for parameter classifier based on a nearest-neighbor scheme. As indicated isolation and regularization-based methods and Table 5 for earlier, this results in lower accuracies on the first task after replay methods .

training. Remarkably, for about half of the tasks, the iCaRL accuracy increases when learning additional tasks, resulting 6.3.1 Discussion in a salient negative average forgetting of −1.11%. Such

Overfitting. We observed overfitting for several models and

level of backward transfer is unusual. After training all ten methods, and not only for the WIDE and DEEP models. tasks, a competitive result of 47.27% average accuracy can

Comparing different methods, SI seems quite vulnerable

be reported. Comparing iCaRL to its baseline R-FM shows to overfitting issues, while PackNet prevents overfitting to significant improvements over basic rehearsal (47.27% vs some extent, thanks to the network compression phase. 37.31%). Doubling the size of the replay buffer (iCaRL 9k) increases performance even more, pushing iCarl closer to General Observations. Selecting highest accuracy disregard- PackNet with 48.76% average accuracy. ing which model is used, iCaRL 9k and PackNet remain on

The results of GEM are significantly lower than those the lead with 49.94% and 49.13%. Further, LwF shows to be of iCaRL (45.13% vs 47.27%). GEM is originally designed competitive with MAS (resp. 46.79% and 46.90%). for an online learning setup, while in this comparison each Baselines. Taking a closer look at the finetuning baseline method can exploit multiple epochs for a task. Additional results (see purple box in Table 4), we observe that it does

experiments in Appendix B.3 compare GEM sensitivity to not reach the same level of accuracy with the SMALL model the amount of epochs with iCaRL, from which we procure as with the other models. In particular, the initial accuracies a GEM setup with 5 epochs for all the experiments in are similar, yet the level of forgetting is much more severe, this work. Furthermore, the lack of memory management due to the limited model capacity to learn new tasks.

policy in GEM gives iCaRL a compelling advantage w.r.t. the Joint training results on the DEEP network (blue box) are amount of exemplars for the first tasks, e.g. training Task 2 inferior compared to shallower networks, implying the deep comprises a replay buffer of Task 1 with factor 1 (number of architecture to be less suited to accumulate all knowledge tasks) more exemplars. Surprisingly, GEM 9k with twice as from the task sequence. As performance of joint learning

much exemplar capacity doesn’t perform better than GEM serves as a soft upper bound for the continual learning 4.5k. This unexpected behavior may be due to the random methods, this already serves as an indication that a deep exemplar selection yielding a less representative subset. model may not be optimal. Nevertheless, GEM convincingly improves accuracy of the In accordance, replay baselines R-PM and R-FM show

basic replay baseline R-PM with the same partial memory to be quite model agnostic with all but the DEEP model scheme (45.13% vs 36.09%). performing very similar. The baselines experience both less Comparing the two basic rehearsal baselines that use the forgetting and increased average accuracy when doubling memory in different ways, we observe the scheme exploit- replay buffer size.

ing full memory from the start (R-FM) giving significantly S MALL model. Finetuning and SI suffer from severe forget- better results than R-PM for tasks 1 to 7, but not for the ting (> 20%) imposed by the decreased capacity of the last three. As more tasks are seen, both baselines converge network (see red underlining), making these combinations to the same memory scheme, where for the final task R-FM worthless in practice. Other methods experience alleviated

allocates an equal portion of memory to each task in the forgetting, with EWC most saliently benefitting from a small sequence, hence equivalent to R-PM. network (see green underlinings). Parameter isolation methods. Both HAT and PackNet exhibit W IDE model. Remarkable in WIDE model results is SI con- the virtue of masking, reporting zero forgetting (see the sistently outperforming EWC, in contrast with other model

two flat curves). HAT performs better for the first task, but builds up a salient discrepancy with PackNet for later 2. Note that, for all of our observations, we also checked the results obtained for the two other task orders (reported in the middle and tasks. PackNet rigorously assigns layer-wise portions of bottom part of Table 4 and Table 5). The reader can check that the network capacity per task, while HAT imposes sparsity by observations are quite consistent.

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

finetuning: 21.3 (26.90) PackNet: 49.1 (0.00) SI: 33.9 (15.77) MAS: 46.9 (1.58) LwF: 41.9 (3.08) joint*: 55.7 (n/a) HAT: 43.5 (0.00) EWC: 42.4 (7.51) mode-IMM: 36.8 (0.98) EBLL: 45.3 (1.44)

Evaluation on Task

T T T T T T T T T T

Accuracy %

0 T T T T T T T T T T

Training Sequence Per Task

finetuning: 21.3 (26.90) R-PM 4.5k: 36.0 (10.96) R-FM 4.5k: 37.3 (9.21) GEM 4.5k: 45.1 (4.96) iCaRL 4.5k: 47.2 (-1.11) joint*: 55.7 (n/a) R-PM 9k: 38.6 (7.23) R-FM 9k: 42.3 (3.94) GEM 9k: 41.7 (5.18) iCaRL 9k: 48.7 (-1.76)

Evaluation on Task

T T T T T T T T T T

Accuracy %

T T T T T T T T T T

Training Sequence Per Task

BASE model with random ordering, reporting average accuracy (forgetting) in the legend.

results (see orange box; even more clear in middle and this in more detail in Appendix B.5.

bottom tables in Table 4, which will be discussed later). EWC performs worse for the high capacity W IDE and DEEP mod- Model agnostic. Over all orderings which will be discussed in els and, as previously discussed, attains best performance Section 6.6, some methods don’t exhibit a preference for any for the SMALL model. LwF, EBLL and PackNet mainly reach model, except for the common aversion for the detrimental

their top performance when using the WIDE model, with SI DEEP model. This is in general most salient for all replay performing most stable on both the WIDE and BASE models. methods in Table 5, but also for PackNet, MAS, LwF and IMM also shows increased performance when using the EBLL in Table 4. BASE and WIDE model.

D EEP model. Over the whole line of methods (yellow box), Conclusion. We can conclude that (too) deep model archi- extending the BASE model with additional convolutional tectures do not provide a good match with the continual layers results in lower performance. As we already observed learning setup. For the same amount of feature extractor overfitting on the BASE model, additional layers may intro- parameters, WIDE models obtain significant better results

duce extra unnecessary layers of abstraction. For the D EEP (on average 11% better over all methods in Table 4 and model iCaRL outperforms all continual learning methods Table 5). Also too small models should be avoided, as the with both memory sizes. HAT deteriorates significantly for limited available capacity can cause forgetting. On average the DEEP model, especially in the random ordering, for we observe a modest 0.89% more forgetting for the SMALL

which we observe issues regarding asymmetric allocation model compared to the BASE model. At the same time, for of network capacity. This makes the method sensitive to some models, poor results may be due to overfitting, which ordering of tasks, as early fixed parameters determine per- can possibly be overcome using regularization, as we will formance for later tasks at saturated capacity. We discuss study next.

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

random (top), easy to hard (middle) and hard to easy (bottom) ordering of tasks, reporting average accuracy (forgetting).

Model finetuning joint* PackNet HAT SI EWC MAS LwF EBLL mode-IMM SMALL 16.2 (34.84) 57.0 (n/a) 46.6 (0.00) 44.1 (0.00) 23.9 (23.26) 45.1 (0.86) 40.5 (0.78) 44.0 (-0.44) 44.1 (-0.53) 29.6 (3.06) BASE 21.3 (26.90) 55.7 (n/a) 49.1 (0.00) 43.5 (0.00) 33.9 (15.77) 42.4 (7.51) 46.9 (1.58) 41.9 (3.08) 45.3 (1.44) 36.8 (0.98)

WIDE 25.2 (24.15) 57.2 (n/a) 47.6 (0.00) 43.7 (0.50) 33.8 (15.16) 31.1 (17.07) 45.0 (2.58) 46.7 (1.19) 46.2 (1.72) 36.4 (1.66) DEEP 20.8 (20.60) 51.0 (n/a) 35.5 (0.00) 8.0 (3.21) 24.5 (12.15) 29.1 (7.92) 33.5 (0.91) 32.2 (2.58) 27.7 (3.14) 27.5 (0.47)

Model finetuning joint* PackNet HAT SI EWC MAS LwF EBLL mode-IMM SMALL 16.0 (35.40) 57.0 (n/a) 49.2 (0.00) 43.8 (0.00) 35.9 (13.02) 40.1 (7.88) 44.2 (1.77) 45.0 (1.89) 42.0 (1.73) 26.1 (3.03) BASE 23.2 (24.85) 55.7 (n/a) 50.5 (0.00) 43.9 (-0.02) 33.3 (14.28) 34.0 (13.25) 44.0 (1.30) 43.4 (2.53) 43.6 (1.71) 36.8 (-1.17)

WIDE 19.6 (29.95) 57.2 (n/a) 47.3 (0.00) 41.6 (0.25) 34.8 (13.14) 28.3 (21.16) 45.5 (1.52) 42.6 (1.07) 44.1 (3.46) 38.6 (-1.09) DEEP 23.9 (17.91) 51.0 (n/a) 34.5 (0.00) 28.6 (0.87) 24.5 (14.41) 24.4 (15.22) 35.2 (2.37) 27.6 (3.70) 29.7 (1.95) 25.8 (-2.09)

Model finetuning joint* PackNet HAT SI EWC MAS LwF EBLL mode-IMM SMALL 18.6 (28.68) 57.0 (n/a) 44.4 (0.00) 33.5 (1.55) 40.3 (4.50) 41.6 (3.54) 40.9 (1.35) 42.3 (0.63) 43.6 (-0.05) 24.9 (1.66) BASE 21.1 (22.73) 55.7 (n/a) 43.1 (0.00) 40.2 (0.18) 40.7 (2.58) 41.8 (1.51) 41.9 (0.14) 41.5 (1.40) 41.5 (0.82) 34.5 (0.23)

WIDE 25.2 (22.69) 57.2 (n/a) 45.5 (0.00) 41.7 (0.07) 37.9 (8.05) 29.9 (15.77) 43.5 (-0.29) 43.8 (1.26) 42.4 (0.85) 35.2 (-0.82) DEEP 15.3 (20.66) 51.0 (n/a) 30.7 (0.00) 27.2 (3.21) 22.9 (13.23) 22.3 (13.86) 32.9 (2.19) 30.7 (2.51) 30.1 (2.67) 26.3 (1.20)

to easy (bottom) ordering of tasks, reporting average accuracy (forgetting).

Model finetuning joint* R-PM 4.5k R-PM 9k R-FM 4.5k R-FM 9k GEM 4.5k GEM 9k iCaRL 4.5k iCaRL 9k SMALL 16.2 (34.84) 57.0 (n/a) 36.9 (12.21) 40.1 (9.23) 39.6 (9.57) 41.3 (6.33) 39.4 (7.38) 41.5 (4.17) 43.2 (-1.39) 46.3 (-1.07) BASE 21.3 (26.90) 55.7 (n/a) 36.0 (10.96) 38.6 (7.23) 37.3 (9.21) 42.3 (3.94) 45.1 (4.96) 41.7 (5.18) 47.2 (-1.11) 48.7 (-1.76)

WIDE 25.2 (24.15) 57.2 (n/a) 36.4 (12.45) 41.5 (5.15) 39.2 (9.26) 41.5 (6.02) 40.3 (6.97) 44.2 (3.94) 44.2 (-1.43) 49.9 (-2.80) DEEP 20.8 (20.60) 51.0 (n/a) 27.6 (7.14) 28.9 (6.27) 32.2 (3.33) 33.1 (4.70) 29.6 (6.57) 23.7 (6.93) 36.1 (-0.93) 37.1 (-1.64)

Model finetuning joint* R-PM 4.5k R-PM 9k R-FM 4.5k R-FM 9k GEM 4.5k GEM 9k iCaRL 4.5k iCaRL 9k SMALL 16.0 (35.40) 57.0 (n/a) 37.0 (11.64) 40.8 (8.08) 38.4 (9.62) 40.8 (7.73) 39.0 (9.45) 42.1 (6.93) 48.5 (-1.17) 46.6 (-2.03) BASE 23.2 (24.85) 55.7 (n/a) 36.7 (8.88) 39.8 (6.59) 38.2 (8.35) 40.9 (6.21) 30.7 (10.39) 40.2 (7.41) 46.3 (-1.53) 47.4 (-2.22)

WIDE 19.6 (29.95) 57.2 (n/a) 39.4 (9.10) 42.5 (5.32) 40.6 (7.52) 43.2 (4.09) 44.4 (4.66) 40.7 (8.14) 48.1 (-2.01) 44.5 (-2.44) DEEP 23.9 (17.91) 51.0 (n/a) 30.9 (6.32) 32.0 (4.16) 29.6 (6.85) 33.9 (3.56) 25.2 (12.79) 29.0 (6.30) 32.4 (-0.35) 34.6 (-1.25)

Model finetuning joint* R-PM 4.5k R-PM 9k R-FM 4.5k R-FM 9k GEM 4.5k GEM 9k iCaRL 4.5k iCaRL 9k SMALL 18.6 (28.68) 57.0 (n/a) 33.6 (11.06) 37.6 (6.34) 37.1 (7.13) 38.4 (4.88) 38.8 (6.78) 39.2 (7.44) 46.8 (-0.92) 46.9 (-1.61) BASE 21.1 (22.73) 55.7 (n/a) 32.9 (8.98) 34.3 (7.50) 33.5 (9.08) 36.8 (5.43) 35.1 (7.44) 35.9 (6.68) 43.2 (-0.47) 44.5 (-1.71)

WIDE 25.2 (22.69) 57.2 (n/a) 34.8 (7.94) 40.2 (6.67) 36.7 (6.63) 37.4 (5.62) 37.2 (7.86) 37.9 (6.94) 41.4 (-1.49) 49.6 (-2.90) DEEP 15.3 (20.66) 51.0 (n/a) 24.2 (6.67) 23.4 (5.93) 25.8 (6.06) 29.9 (2.91) 27.0 (6.47) 31.2 (5.52) 30.9 (0.85) 37.9 (-1.35)

6.4 Effects of Regularization Finetuning. Goodfellow et al. observe reduced catastrophic forgetting in a transfer learning setup with finetuning when

Tiny Imagenet. In the previous subsection we mentioned

using dropout . Extended to learning a sequence of the problem of overfitting in continual learning. Although

1 consecutive tasks in our experiments, finetuning consis-

an evident solution would be to apply regularization, this tently benefits from dropout regularization. This is opposed might interfere with the continual learning methods. There- to weight decay, resulting in increased forgetting and a fore, we investigate the effects of two popular regularization lower performance on the final model. In spite of good methods, namely dropout and weight decay, for the param-

results for dropout, we regularly observe an increase in the

eter isolation and regularization-based methods in Table 6, level of forgetting, which is compensated for by starting and for the replay methods in Table 7. For dropout we from a better initial model, due to a reduction in overfitting. set the probability of retaining the units to p = 0.5, and

Dropout leading to more forgetting is something we also

weight decay applies a regularization strength of λ = 10−4 . observe for many other methods (see blue boxes), and

Any form of regularization is applied in both phases of our

exacerbates as the task sequence grows in length. framework.

Joint training and PackNet. mainly relish higher accura-

6.4.1 Discussion cies with both regularization setups. By construction, they General observations. In Table 6 and Table 7, negative results benefit from the regularization without interference issues. (with regularization hurting performance) are underlined in PackNet even reaches a top performance of almost 56%, that red. This occurs repeatedly, especially in combination with is 7% higher than closest competitor MAS.

the SMALL model or with weight decay. Over all methods HAT mainly favours weight decay of the model parameters, and models dropout mainly shows to be fruitful. This is excluding embedding parameters. The embeddings are al- consistent with earlier observations . There are, however, ready prone to sparsity regularization, for which we found a few salient exceptions (discussed below). Weight decay troublesome learning when included for weight decay as

over the whole line mainly improves the wide network well. Dropout results in worse results, probably due to in- accuracies. In the following, we will describe the most terfering with the unit-based masking. However, difficulties notable observations and exceptions to the main tendencies. regarding asymmetric capacity in the DEEP model seem to

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

Model finetuning joint* PackNet HAT SI EWC MAS LwF EBLL mode-IMM SMALL Plain 16.2 (34.84) 57.0 (n/a) 49.0 (0.00) 44.1 (0.00) 23.9 (23.26) 45.1 (0.86) 40.5 (0.78) 44.0 (-0.44) 44.1 (-0.53) 29.6 (3.06) Dropout 19.5 (32.62) 55.9 (n/a) 50.7 (0.00) 26.8 (2.46) 38.3 (9.24) 40.0 (7.57) 40.2 (7.63) 31.5 (18.38) 34.1 (13.20) 29.3 (0.90)

Weight Decay 15.0 (34.61) 56.9 (n/a) 49.9 (0.00) 42.4 (0.00) 37.9 (6.43) 41.2 (1.52) 37.3 (4.43) 41.6 (-0.28) 42.6 (-0.63) 30.5 (0.93) BASE Plain 21.3 (26.90) 55.7 (n/a) 47.6 (0.00) 43.5 (0.00) 33.9 (15.77) 42.4 (7.51) 46.9 (1.58) 41.9 (3.08) 45.3 (1.44) 36.8 (0.98) Dropout 29.2 (26.44) 61.4 (n/a) 54.2 (0.00) 38.1 (0.36) 43.1 (10.83) 42.0 (12.54) 48.9 (0.87) 41.4 (8.72) 44.6 (7.66) 34.2 (1.44)

Weight Decay 19.1 (29.31) 57.1 (n/a) 48.2 (0.00) 44.9 (0.19) 39.6 (8.11) 44.3 (3.51) 44.2 (1.15) 40.9 (1.29) 41.2 (0.82) 37.4 (0.46) WIDE Plain 25.2 (24.15) 57.2 (n/a) 48.3 (0.00) 43.7 (0.50) 33.8 (15.16) 31.1 (17.07) 45.0 (2.58) 46.7 (1.19) 46.2 (1.72) 36.4 (1.66) Dropout 30.7 (26.11) 62.2 (n/a) 55.9 (0.00) 34.9 (0.81) 43.7 (8.80) 33.9 (19.73) 47.9 (1.37) 45.0 (6.85) 46.1 (5.31) 42.4 (-0.93)

Weight Decay 22.7 (27.19) 59.6 (n/a) 47.7 (0.00) 44.1 (0.06) 42.4 (8.36) 37.4 (13.47) 47.2 (1.60) 48.1 (0.62) 48.1 (0.82) 39.1 (1.41) DEEP Plain 20.8 (20.60) 51.0 (n/a) 34.7 (0.00) 8.0 (3.21) 24.5 (12.15) 29.1 (7.92) 33.5 (0.91) 32.2 (2.58) 27.7 (3.14) 27.5 (0.47) Dropout 23.0 (27.30) 59.5 (n/a) 46.2 (0.00) 27.5 (2.79) 32.7 (15.09) 31.1 (17.06) 39.0 (5.02) 37.8 (7.78) 36.8 (4.87) 33.6 (-0.6

Weight Decay 19.4 (21.27) 54.6 (n/a) 36.9 (0.00) 30.2 (2.02) 26.0 (9.51) 22.4 (13.19) 19.3 (13.62) 33.1 (1.16) 31.7 (1.39) 25.6 (-0.15)

be alleviated using either type of regularization. TABLE 7: Replay methods: dropout (p = 0.5) and weight decay (λ = 10−4 ) regularization for Tiny Imagenet models.

SI most saliently thrives with regularization (with increases

in performance of around 10%, see green box), which we Model R-PM 4.5k R-FM 4.5k GEM 4.5k iCaRL 4.5k previously found to be sensitive to overfitting. The regular- SMALL Plain 36.9 (12.21) 39.6 (9.57) 39.4 (7.38) 43.2 (-1.39) ization aims to find a solution to a well-posed problem, sta- Dropout 35.5 (12.02) 35.7 (9.87) 33.8 (6.25) 44.8 (4.83)

Weight Decay 35.5 (13.38) 39.1 (9.49) 37.6 (5.11) 44.9 (-0.81)

bilizing the path in parameter space. Therefore, this might

BASE Plain 36.0 (10.96) 37.3 (9.21) 45.1 (4.96) 47.2 (-1.11)

provide a better importance weight estimation along the Dropout 43.3 (10.59) 45.7 (7.38) 36.0 (12.13) 48.4 (2.68) path. Nonetheless, SI remains noncompetitive with leading Weight Decay 36.1 (10.13) 37.7 (8.01) 38.0 (8.74) 45.9 (-2.32)

average accuracies of other methods. WIDE Plain 36.4 (12.45) 39.2 (9.26) 40.3 (6.97) 44.2 (-1.43)

Dropout 42.2 (12.31) 45.5 (8.85) 42.7 (6.33) 45.5 (2.74)

EWC and MAS suffer from interference with additional Weight Decay 39.7 (8.28) 38.7 (10.01) 45.2 (5.92) 46.5 (-1.38)

regularization. For dropout, more redundancy in the model DEEP Plain 27.6 (7.14) 32.2 (3.33) 29.6 (6.57) 36.1 (-0.93)

Dropout 34.4 (9.57) 37.2 (7.65) 32.7 (8.15) 41.7 (3.58)

means less unimportant parameters left for learning new Weight Decay 26.7 (9.01) 31.7 (4.86) 27.2 (5.94) 33.7 (-0.47) tasks. For weight decay, parameters deemed important for previous tasks are also decayed in each iteration, affecting performance on older tasks. A similar effect was noticed both in terms of classes and available data per task. We in . However, in some cases, the effect of this interference train AlexNet pretrained on Imagenet, and track task per-

is again compensated for by having better initial models. formance for RecogSeq and iNaturalist in Figure 3 and Fig- EWC only benefits from dropout on the WIDE and DEEP ure 4a. The experiments exclude both noncompetitive HAT models, and similar to MAS prefers no regularization on results due to problems of asymmetric capacity allocation the SMALL model. discussed in Appendix B.5, and replay methods as they lack

LwF and EBLL suffer from using dropout, with only the DEEP a policy to cope with unbalanced data, which would make model significantly improving performance. For weight de- comparison highly biased to our implementation. cay the methods follow the general trend, enhancing the

WIDE net accuracies. 6.5.1 Discussion

IMM with dropout exhibits higher performance only for the iNaturalist. Overall PackNet shows the highest average

WIDE and DEEP model, coming closer to the performance

accuracy, tailgated by mode-IMM with superior forgetting obtained by the other methods. by exhibiting positive backward transfer. Both PackNet and iCaRL and GEM. Similar to the observations in Table 6, there mode-IMM attain accuracies very close or even surpassing is a striking increased forgetting for dropout for all methods joint training, such as for Task 1 and Task 8.

in Table 7. Especially iCaRL shows inreased average forget-

Performance drop. Evaluation on the first 4 tasks shows

ting, albeit consistently accompanied with higher average salient dips when learning Task 5 for prior-based methods accuracy. Except for the SMALL model, all models for the EWC, SI and MAS. In the random ordering Task 5 is an replay baselines benefit from dropout. For GEM this benefit extremely easy task of supercategory ’Fungi’ which contains is only notable for the WIDE and DEEP models. Weight decay only 5 classes and few data. Using expert gates to measure

doesn’t improve average accuracy for the replay methods, relatedness, the first four tasks show no particularly salient nor forgetting, except mainly for the WIDE model. In gen- relatedness peaks or dips for Task 5. Instead, forgetting eral, the influence of regularization seems limited for replay might rather be caused by the limited amount of training methods in Table 7 compared to non-replay based methods data, with only a few target classes, enforcing the network

in Table 6. to overly fit to this task.

SI. For iNaturalist we did not observe overfitting, which

6.5 Effects of a Real-world Setting might be the cause for stable SI behaviour in comparison to iNaturalist and RecogSeq. Up to this point all experiments Tiny Imagenet. conducted on Tiny Imagenet are nicely balanced, with an LwF and EBLL. LwF catastrophically forgets on the unbal- equal amount of data and classes per task. In further ex- anced dataset with 13.77% forgetting, which is in high

periments we scrutinize highly unbalanced task sequences, contrast with the results acquired on Tiny Imagenet. The

This is the author’s version of an article that has been published in this journal. Changes were made to this version by the publisher prior to publication.

The final version of record is available at http://dx.doi.org/10.1109/TPAMI.2021.30574

supercategories in iNaturalist constitute a completely dif- TABLE 8: The unbalanced iNaturalist task sequence details ferent task, imposing severe distribution shifts between the for random, related, and unrelated orderings. tasks. In the contrary, Tiny Imagenet is constructed from

Task Classes Samples Ordering

a subset of randomly collected classes, implying similar levels of homogeneity between task distributions. On top, Train Val Test Rand. Rel. Unrel.

Amphibia 2 53 7 15 1 4 1

EBLL constraints the new task features to reside closely to Animalia 1 13 1 3 2 5 9 the optimal presentation for previous task features, which Arachnida 9 11 1 3 3 8 7 for the random ordering enforces forgetting to nearly halve Aves 3 658 93 188 4 1 1

Fungi 5 6 9 1 5 6 2

from 13.77% to 7.51%, resulting in a striking 7.91% increase Insecta 1 260 36 74 6 9 3 in average accuracy over LwF (from 45.39% to 53.30%). Mammalia 4 86 12 24 7 2 8

Mollusca 1 17 2 4 8 7 4

RecogSeq shows similar findings to iNaturalist for LwF Plantae 2 351 49 100 9 1 5 and EBLL, exacerbating in performance as they are subject Reptilia 5 101 14 29 1 3 6 to severe distribution shifts between tasks. This is most notable for the fierce forgetting when learning the last

SVHN task. Besides the data-focused methods, we also 6.6.1 Discussion

observe deterioration for the prior-focused regularization- based methods. This shows increasing difficulty of these Tiny Imagenet. The main observations on the random or- methods to cope with highly different recognition tasks. dering for the BASE model in Section 6.2 and model capacity By contrast, PackNet thrives in this setup, taking the lead in Section 6.3 remain valid for the two additional orderings,

by high margin and approaching joint performance. In its with PackNet and iCaRL competing for highest average advantage, PackNet freezes the previous task parameters, to accuracy, subsequently followed by MAS and LwF. In the impose zero forgetting regardless of the task dissimilarities. following, we will instead focus on general observations For all continual learning methods, Appendix B.4 reports between the three different orderings.

reproduced results similar to the original RecogSeq setup Task Ordering Hypothesis. Starting from our curriculum learn- , , showcasing hyperparameter sensitivity and an urge ing hypothesis we would expect the easy-to-hard ordering for hyperparameter robustness. to enhance performance w.r.t. the random ordering, and the opposite effect for the hard-to-easy ordering. However, es-

pecially SI and EWC show unexpected better results for the 6.6 Effects of Task Order hard-to-easy ordering than for the easy-to-hard ordering.

For PackNet and MAS, we see a systematic improvement

In this experiment we scrutinize how changing the task when switching from hard-to-easy to easy-to-hard ordering. order affects both the balanced (Tiny Imagenet) and unbal- The gain is, however, relatively small. Overall, the impact of anced (iNaturalist) task sequences. Our hypothesis resem- the task order seems insignificant. bles curriculum learning , implying knowledge is better

Replay Buffer. Introducing the two other orderings, we now

captured starting with the general easier tasks followed by observe that iCaRL doesn’t always improve when increasing harder specific tasks. Instead, experimental results on both the replay buffer size for the easy-to-hard ordering (see datasets exhibit ordering agnostic behaviour, corresponding

SMALL and WIDE model). More exemplars induce more

to similar previous findings . samples to distill knowledge from previous tasks, but might Tiny Imagenet. The previously discussed results of param- deteriorate stochasticity in the estimated feature means in eter isolation and regularization-based methods on top of the nearest-neighbour classifier. GEM does not as consis- random ordering. We define a new order based on task which could also root in the reduced stochasticity of the

difficulty, by measuring the accuracy over the 4 models constraining gradients from previous tasks. obtained on held-out datasets for each of the tasks. This results in an ordering from hard to easy consisting of task iNaturalist. The general observations for the random or- sequence [5, 7, 10, 2, 9, 8, 6, 4, 3, 1] from the random dering remain consistent for the other orderings as well, ordering, wit

Frequently Asked Questions

What is this project about?

This project covers practical implementation and research aspects of the topic using AI/ML techniques.