3D Object Detection for Autonomous Driving: A Comprehensive Survey Jiageng Mao 1 · Shaoshuai Shi 3 · Xiaogang Wang 1,2 · Hongsheng Li 1,2
Received: 7 February 2023
Abstract Autonomous driving, in recent years, has been receiv- ing increasing attention for its potential to relieve drivers’ bur- dens and improve the safety of driving. In modern autonomous driving pipelines, the perception system is an indispensable com- ponent, aiming to accurately estimate the status of surrounding environments and provide reliable observations for prediction and planning. 3D object detection, which aims to predict the locations, sizes, and categories of the 3D objects near an au- tonomous vehicle, is an important part of a perception system.
This paper reviews the advances in 3D object detection for au- tonomous driving. First, we introduce the background of 3D ob- ject detection and discuss the challenges in this task. Second, we conduct a comprehensive survey of the progress in 3D ob- ject detection from the aspects of models and sensory inputs, in- cluding LiDAR-based, camera-based, and multi-modal detection approaches. We also provide an in-depth analysis of the poten- tials and challenges in each category of methods. Additionally, we systematically investigate the applications of 3D object de- tection in driving systems. Finally, we conduct a performance analysis of the 3D object detection approaches, and we further summarize the research trends over the years and prospect the future directions of this area.
Keywords 3D object detection · perception · autonomous driving · deep learning · computer vision · robotics
1 Introduction
Autonomous driving, which aims to enable vehicles to perceive the surrounding environments intelligently and move safely with little or no human effort, has attained rapid progress in recent years. Autonomous driving techniques have been broadly ap- plied in many scenarios, including self-driving trucks, robotaxis, delivery robots, etc., and are capable of reducing human error and enhancing road safety. As a core component of autonomous driving systems, automotive perception helps autonomous vehi- cles understand the surrounding environments with sensory in- 2 Centre for Perceptual and Interactive Intelligence
3 Max Planck Institute For Informatics, Germany
put. Perception systems generally take multi-modality data (im- ages from cameras, point clouds from LiDAR scanners, high- definition maps etc.) as input, and predict the geometric and se- mantic information of critical elements on a road. High-quality perception results serve as reliable observations for the follow- ing steps such as object tracking, trajectory prediction, and path planning.
To obtain a comprehensive understanding of driving environ- ments, many vision tasks can be involved in a perception system, e.g. object detection and tracking, lane detection, and seman- tic and instance segmentation. Among these perception tasks, 3D object detection is one of the most indispensable tasks in an automotive perception system. 3D object detection aims to pre- dict the locations, sizes, and classes of critical objects, e.g. cars, pedestrians, cyclists, in the 3D space. In contrast to 2D object detection which only generates 2D bounding boxes on images and ignores the actual distance information of objects from the ego-vehicle, 3D object detection focuses on the localization and recognition of objects in the real-world 3D coordinate system.
The geometric information predicted by 3D object detection in real-world coordinates can be directly utilized to measure the distances between the ego-vehicle and critical objects, and to further help plan driving routes and avoid collisions.
3D object detection methods have evolved rapidly with the advances of deep learning techniques in computer vision and robotics. These methods have been trying to address the 3D ob- ject detection problem from a particular aspect, e.g. detection from a particular sensory type or data representation, and lack a systematic comparison with the methods of other categories.
Hence a comprehensive analysis of the strengths and weaknesses of all types of 3D object detection methods is desirable and can provide some valuable findings to the research community.
To this end, we propose to comprehensively review the 3D object detection methods for autonomous driving applications and provide in-depth analysis and a systematic comparison on different categories of approaches. Compared to the existing sur- veys [5, 151, 232], our paper broadly covers the recent advances in this area, e.g. 3D object detection from range images, self- /semi-/weakly-supervised 3D object detection, 3D detection in end-to-end driving systems. In contrast to the previous surveys that only focus on detection from point cloud [90, 75, 360], from monocular images [317, 179], and from multi-modal in- puts , our paper systematically investigate the 3D object detection methods from all sensory types and in most application
2
Jiageng Mao et al.
Self-Supervised 3D Object Detection
3D Object Detection in Driving Systems (Section 9)
Transformer Applications In 3D Object Detection
Fig. 1: Hierarchically-structured taxonomy of 3D object detection for autonomous driving. scenarios. The major contributions of this work can be summa-
Rized As Follows:
– We provide a comprehensive review of the 3D object detec- tion methods from different perspectives, including detec- tion from different sensory inputs (LiDAR-based, camera- based, and multi-modal detection), detection from temporal sequences, label-efficient detection, as well as the applica- tions of 3D object detection in driving systems.
– We summarize 3D object detection approaches structurally and hierarchically, conduct a systematic analysis of these methods, and provide valuable insights for the potentials and challenges of different categories of methods.
– We conduct a comprehensive performance and speed anal- ysis on the 3D object detection approaches, identify the re- search trends over years, and provide insightful views on the future directions of 3D object detection.
The structure of this paper is organized as follows. First, we introduce the problem definition, datasets, and evaluation met- rics of 3D object detection in Section 2. Then, we review and analyze the 3D object detection methods based on LiDAR sen- sors (Section 3), cameras (Section 4), multi-sensor fusion (Sec- tion 5), and Transformer-based architectures (Section 6). Next, we introduce the detection methods that leverage temporal data in Section 7 and utilize fewer labels in Section 8. We subse- quently discuss some critical problems of 3D object detection in driving systems in Section 9. Finally, we conduct a speed and performance analysis, investigate the research trends, and prospect the future directions of 3D object detection in Sec- tion 10. A hierarchically-structured taxonomy is shown in Fig- ure 1. We also provide a constantly updated project page here.
2.1 What Is 3D Object Detection?
Problem definition. 3D object detection aims to predict bound- ing boxes of 3D objects in driving scenarios from sensory inputs. A general formula of 3D object detection can be represented as
(1)
where B = {B1, · · · , BN} is a set of N 3D objects in a scene, fdet is a 3D object detection model, and Isensor is one or more sensory inputs. How to represent a 3D object Bi is a crucial prob- lem in this task, since it determines what 3D information should be provided for the following prediction and planning steps. In most cases, a 3D object is represented as a 3D cuboid that in-
(2)
where (xc, yc, zc) is the 3D center coordinate of a cuboid, l, w, h is the length, width, and height of a cuboid respectively, θ is the heading angle, i.e. the yaw angle, of a cuboid on the ground plane, and class denotes the category of a 3D object, e.g. cars, trucks, pedestrians, cyclists. In , additional parameters vx and vy that describe the speed of a 3D object along x and y axes on the ground are employed.
Sensory inputs. There are many types of sensors that can pro- vide raw data for 3D object detection. Among the sensors, radars, cameras, and LiDAR (Light Detection And Ranging) sensors are the three most widely adopted sensory types. Radars have long detection range and are robust to different weather conditions.
Due to the Doppler effect, radars could provide additional ve- locity measurements. Cameras are cheap and easily accessible, and can be crucial for understanding semantics, e.g. the type of traffic sign. Cameras produce images Icam ∈RW ×H×3 for 3D object detection, where W, H are the width and height of an im- age, and each pixel has 3 RGB channels. Albeit cheap, cameras have intrinsic limitations to be utilized for 3D object detection.
First, cameras only capture appearance information, and are not capable of directly obtaining 3D structural information about a scene. On the other hand, 3D object detection normally requires accurate localization in the 3D space, while the 3D information, e.g. depth, estimated from images normally has large errors. In addition, detection from images is generally vulnerable to ex- treme weather and time conditions. Detecting objects from im- ages at night or on foggy days is much harder than detection on sunny days, which leads to the challenge of attaining sufficient robustness for autonomous driving.
As an alternative solution, LiDAR sensors can obtain fine- grained 3D structures of a scene by emitting laser beams and then measuring their reflective information. A LiDAR sensor that emits m beams and conducts measurements for n times in one scan cycle can produce a range image Irange ∈Rm×n×3, where each pixel of a range image contains range r, azimuth α, and inclination φ in the spherical coordinate system as well as 3D Object Detection for Autonomous Driving: A Comprehensive Survey
3D Object Detection From Point Cloud
Fig. 2: An illustration of 3D object detection in autonomous driving scenarios. the reflective intensity. Range images are the raw data format ob- tained by LiDAR sensors, and can be further converted into point clouds by transforming spherical coordinates into Cartesian co- ordinates. A point cloud can be represented as Ipoint ∈RN×3, where N denotes the number of points in a scene, and each point has 3 channels of xyz coordinates. Both range images and point clouds contain accurate 3D information directly acquired by Li- DAR sensors. Hence in contrast to cameras, LiDAR sensors are more suitable for detecting objects in the 3D space, and LiDAR sensors are also less vulnerable to time and weather changes.
However, LiDAR sensors are much more expensive than cam- eras, which may limit the applications in driving scenarios. An illustration of 3D object detection is shown in Figure 2.
Analysis: comparisons with 2D object detection. 2D object detection, which aims to generate 2D bounding boxes on images, is a fundamental problem in computer vision. 3D object detec- tion methods have borrowed many design paradigms from the 2D counterparts: proposals generation and refinement, anchors, non maximum suppression, etc. However, from many aspects, 3D object detection is not a naive adaptation of 2D object detec- tion methods to the 3D space. (1) 3D object detection methods have to deal with heterogeneous data representations. Detection from point clouds requires novel operators and networks to han- dle irregular point data, and detection from both point clouds and images needs special fusion mechanisms. (2) 3D object detection methods normally leverage distinct projected views to generate object predictions. As opposed to 2D object detection methods that detect objects from the perspective view, 3D methods have to consider different views to detect 3D objects, e.g. from the bird’s-eye view, point view, and cylindrical view. (3) 3D object detection has a high demand for accurate localization of objects in the 3D space. A decimeter-level localization error can lead to a detection failure of small objects such as pedestrians and cy- clists, while in 2D object detection, a localization error of several pixels may still maintain a high Intersection over Union (IoU) between predicted and ground truth bounding boxes. Hence ac- curate 3D geometric information is indispensable for 3D object detection from either point clouds or images.
Analysis: comparisons with indoor 3D object detection. There is also a branch of works [226, 227, 228, 169] on 3D object de- tection in indoor scenarios. Indoor datasets, e.g. ScanNet , SUN RGB-D , provide 3D structures of rooms reconstructed from RGB-D sensors and 3D annotations including doors, win- dows, beds, chairs, etc. 3D object detection in indoor scenes is also based on point clouds or images. However, compared to in- door 3D object detection, there are unique challenges of detec- tion in driving scenarios. (1) Point cloud distributions from Li- DAR and RGB-D sensors are different. In indoor scenes, points are relatively uniformly distributed on the scanned surfaces and most 3D objects receive a sufficient number of points on their surfaces. However, in driving scenes most points fall in a near neighborhood of the LiDAR sensor, and those 3D objects that are far away from the sensor will receive only a few points. Thus methods in driving scenarios are specially required to handle various point cloud densities of 3D objects and accurately de- tect those faraway and sparse objects. (2) Detection in driving scenarios has a special demand for inference latency. Perception in driving scenes has to be real-time to avoid accidents. Hence those methods are required to be computationally efficient, oth- erwise they will not be applied in real-world applications.
2.2 Datasets
A large number of driving datasets have been built to provide multi-modal sensory data and 3D annotations for 3D object de- tection. Table 1 lists the datasets that collect data in driving sce- narios and provide 3D cuboid annotations. KITTI is a pio- neering work that proposes a standard data collection and an- notation paradigm: equipping a vehicle with cameras and Li- DAR sensors, driving the vehicle on roads for data collection, and annotating 3D objects from the collected data. The follow- ing works made improvements mainly from the 4 aspects. (1) Increasing the scale of data. Compared to , the recent large- scale datasets [268, 15, 186] have more than 10x point clouds, images and annotations. (2) Improving the diversity of data.
only contains driving data obtained in the daytime and in good weather, while recent datasets [51, 30, 218, 15, 268, 321, 186, 315] provide data captured at night or in rainy days. (3) Provid- ing more annotated categories. Some datasets [154, 321, 84, 315, 15] can provide more fine-grained object classes, including ani- mals, barriers, traffic cones, etc. They also provide fine-grained sub-categories of existing classes, e.g. the adult and child cate- gory of the existing pedestrian class in . (4) Providing data of more modalities. In addition to images and point clouds, re- cent datasets provide more data types, including high-definition maps [114, 30, 268, 315], radar data , long-range LiDAR data [313, 307], thermal images .
4
Jiageng Mao et al. Table 1: Datasets for 3D object detection in driving scenarios.
-
Analysis: future prospects of driving datasets. The research community has witnessed an explosion of datasets for 3D object detection in autonomous driving scenarios. A subsequent ques- tion may be asked: what will the next-generation autonomous driving datasets look like? Considering the fact that 3D object detection is not an independent task but a component in driving systems, we propose that future datasets will include all impor- tant tasks in autonomous driving: perception, prediction, plan- ning, and mapping, as a whole and in an end-to-end manner, so that the development and evaluation of 3D object detection methods will be considered from an overall and systematic view.
There are some datasets [268, 15, 352] working towards this goal.
2.3 Evaluation Metrics
Various evaluation metrics have been proposed to measure the performance of 3D object detection methods. Those evaluation metrics can be divided into two categories. The first category tries to extend the Average Precision (AP) metric in 2D
(3)
where p(r) is the precision-recall curve same as . The ma- jor difference with the 2D AP metric lies in the matching cri- terion between ground truths and predictions when calculating precision and recall. KITTI proposes two widely-used AP metrics: AP3D and APBEV , where AP3D matches the predicted objects to the respective ground truths if the 3D Intersection over Union (3D IoU) of two cuboids is above a certain threshold, and APBEV is based on the IoU of two cuboids from the bird’s-eye view (BEV IoU). NuScenes proposes APcenter where a pre- dicted object is matched to a ground truth object if the distance of their center locations is below a certain threshold, and NuScenes Detection Score (NDS) is further proposed to take both APcenter and the error of other parameters, i.e. size, heading, velocity, into consideration. Waymo proposes APhungarian that applies the Hungarian algorithm to match the ground truths and predic- tions, and AP weighted by Heading (APH) is proposed to in- corporate heading errors as a coefficient into the AP calculation.
The other category of approaches tries to resolve the eval- uation problem from a more practical perspective. The idea is that the quality of 3D object detection should be relevant to the downstream task, i.e. motion planning, so that the best detection methods should be most helpful to planners to ensure the safety of driving in practical applications. Toward this goal, PKL measures the detection quality using the KL-divergence of the ego vehicle’s future planned states based on the predicted and ground truth detections respectively. SDE leverages the min- imal distance from the object boundary to the ego vehicle as the support distance and measures the support distance error.
Analysis: pros and cons of different evaluation metrics. AP- based evaluation metrics [82, 15, 268] can naturally inherit the advantages from 2D detection. However, those metrics overlook the influence of detection on safety issues, which are also criti- cal in real-world applications. For instance, a misdetection of an object near the ego vehicle and far away from the ego vehicle may receive a similar level of punishment in AP calculation, but a misdetection of nearby objects is substantially more dangerous than a misdetection of faraway objects in practical applications.
Thus AP-based metrics may not be the optimal solution from the perspective of safe driving. PKL and SDE partly resolve the problem by considering the effects of detection in downstream tasks, but additional challenges will be introduced when modeling those effects. PKL requires a pre-trained motion planner for evaluating the detection performance, but a pre-trained planner also has innate errors that could make the evaluation process inaccurate. SDE requires reconstructing object boundaries which is generally complicated and challeng- ing.
3 Lidar-Based 3D Object Detection
In this section, we introduce the 3D object detection methods based on LiDAR data, i.e. point clouds or range images. In Sec- tion 3.1, we review and analyze the LiDAR-based 3D object de- tection models based on different data representations, includ- ing the point-based, grid-based, point-voxel based, and range- based methods. In Section 3.2, we investigate the learning ob- jectives for 3D object detectors, including the anchor-based and anchor-free frameworks, as well as the auxiliary tasks adopted 3D Object Detection for Autonomous Driving: A Comprehensive Survey
(Hu Et Al.)
Fig. 3: Chronological overview of the LiDAR-based 3D object detection methods.
Points & Features
Fig. 4: An illustration of point-based 3D object detection methods. in LiDAR-based 3D object detection. A chronological overview of the LiDAR-based 3D detection methods is shown in Figure 3.
3.1 Data Representations For 3D Object Detection
Problem and Challenge. In contrast to images where pixels are regularly distributed on an image plane, point cloud is a sparse and irregular 3D representation that requires specially designed models for feature extraction. Range image is a dense and com- pact representation, but range pixels contain 3D information in- stead of RGB values. Hence directly applying conventional con- volutional networks on range images may not be an optimal solu- tion. On the other hand, detection in autonomous driving scenar- ios generally has a requirement for real-time inference. There- fore, how to develop a model that could effectively handle point cloud or range image data while maintaining a high efficiency remains an open challenge to the research community.
3.1.1 Point-Based 3D Object Detection
General Framework. Point-based 3D object detection methods generally inherit the success of deep learning techniques on point cloud [224, 225, 300, 184] and propose diverse architectures to detect 3D objects directly from raw points. Point clouds are first passed through a point-based backbone network, in which the points are gradually sampled and features are learned by point cloud operators. 3D bounding boxes are then predicted based on the downsampled points and features. A general point-based Table 2: A taxonomy of point-based detection methods based on point cloud sampling and feature learning.
Ball Query
Seg.
Transformer
detection framework is shown in Figure 4 and a taxonomy of point-based detectors is in Table 2. There are two basic compo- nents of a point-based 3D object detector: point cloud sampling and feature learning.
Point Cloud Sampling. Farthest Point Sampling (FPS) in Point- Net++ has been broadly adopted in point-based detectors, in which the farthest points are sequentially selected from the original point set. PointRCNN is a pioneering work that adopts FPS to progressively downsample input point cloud and generate 3D proposals from the downsampled points. Similar de- sign paradigm has also been adopted in many following works with improvements like segmentation guided filtering , fea- ture space sampling , random sampling , voxel-based sampling , and coordinate refinement .
Point Cloud Feature Learning. A series of works [252, 340, 379, 290] leverage set abstraction in to learn features from point cloud. Specifically, context points are first collected within a pre-defined radius by ball query. Then, the context points and
6
Jiageng Mao et al.
Detection Head
Fig. 5: An illustration of grid-based 3D object detection methods. features are aggregated through multi-layer perceptrons and max- pooling to obtain the new features. There are also other works resorting to different point cloud operators, including graph op- erators [256, 361, 203, 74, 97], attentional operators , and Transformer .
Analysis: potentials and challenges on point cloud feature learning and sampling. The representation power of point-based detectors is mainly restricted by two factors: the number of con- text points and the context radius adopted in feature learning. In- creasing the number of context points will gain more representa- tion power but at the cost of increasing much memory consump- tion. Suitable context radius in ball query is also an important factor: the context information may be insufficient if the radius is too small and the fine-grained 3D information may lose if the radius is too large. These two factors have to be determined care- fully to balance the efficacy and efficiency of detection models.
Point cloud sampling is a bottleneck in inference time for most point-based methods. Random uniform sampling can be conducted in parallel with high efficiency. However, consider- ing points in LiDAR sweeps are not uniformly distributed, ran- dom uniform sampling may tend to over-sample those regions of high point cloud density while under-sample those sparse re- gions, which normally leads to poor performance compared to farthest point sampling. Farthest point sampling and its variants can attain a more uniform sampling result by sequentially select- ing the farthest point from the existing point set. Nevertheless, farthest point sampling is intrinsically a sequential algorithm and can not become highly parallel. Thus farthest point sampling is normally time-consuming and not ready for real-time detection.
3.1.2 Grid-Based 3D Object Detection
General Framework. Grid-based 3D object detectors first ras- terize point clouds into discrete grid representations, i.e. voxels, pillars, and bird’s-eye view (BEV) feature maps. Then they ap- ply conventional 2D convolutional neural networks or 3D sparse neural networks to extract features from the grids. Finally, 3D objects can be detected from the BEV grid cells. An illustration of grid-based 3D object detection is shown in Figure 5 and a tax- onomy of grid-based detectors is in Table 3. There are two basic components in grid-based detectors: grid-based representations and grid-based neural networks.
Grid-based representations. There are 3 major types of grid representations: voxels, pillars, and BEV feature maps. Voxels. If we rasterize the detection space into a regular 3D grid, voxels are the grid cells. A voxel can be non-empty if point clouds fall into this grid cell. Since point clouds are sparsely dis- tributed, most voxel cells in the 3D space are empty and contain no point. In practical applications, only those non-empty voxels are stored and utilized for feature extraction. VoxelNet is a pioneering work that utilizes sparse voxel grids and proposes a novel voxel feature encoding (VFE) layer to extract features from the points inside a voxel cell. A similar voxel encoding strategy has been adopted by a series of following works [385, 81, 350, 297, 375, 129, 57, 254]. In addition, there are two cate- gories of approaches trying to improve the voxel representation for 3D object detection: (1) Multi-view voxels. Some methods propose a dynamic voxelization and fusion scheme from diverse views, e.g. from both the bird’s-eye view and the perspective view , from the cylindrical and spherical view , from the range view . (2) Multi-scale voxels. Some papers gen- erate voxels of different scales or use reconfigurable vox- els .
Pillars. Pillars can be viewed as special voxels in which the voxel size is unlimited in the vertical direction. Pillar features can be aggregated from points through a PointNet and then scattered back to construct a 2D BEV image for feature extrac- tion. PointPillars is a seminal work that introduces the pil- lar representation and is followed by [302, 70].
BEV feature maps. Bird’s-eye view feature map is a dense 2D representation, where each pixel corresponds to a specific region and encodes the points information in this region. BEV feature maps can be obtained from voxels and pillars by projecting the 3D features into the bird’s-eye view, or they can be directly ob- tained from raw point clouds by summarizing points statistics within the pixel region. The commonly-used statistics include binary occupancy [335, 334, 2] and the height and density of local point cloud [41, 10, 364, 3, 263, 368, 8, 126].
Grid-based neural networks. There are 2 major types of grid- based networks: 2D convolutional neural networks for BEV fea- ture maps and pillars, and 3D sparse neural networks for voxels.
2D convolutional neural networks. Conventional 2D convo- lutional neural networks can be applied to the BEV feature map to detect 3D objects from the bird’s-eye view. In most works, the 2D network architectures are generally adapted from those suc- cessful designs in 2D object detection, e.g. ResNet adopted
In , Region Proposal Network (Rpn) And Feature
3D Object Detection for Autonomous Driving: A Comprehensive Survey
7
Table 3: A taxonomy of grid-based detection methods based on models and data representations.
✓
✓
✓
✓
Pyramid Network (FPN) in [10, 8, 124, 119, 263, 129], and spatial attention in [130, 167, 346]. 3D sparse neural networks. 3D sparse convolutional neural networks are based on two specialized 3D convolutional oper- ators: sparse convolutions and submanifold convolutions , which can efficiently conduct 3D convolutions only on those non-empty voxels. Compared to [283, 68, 201, 125] that per- form standard 3D convolutions on the whole voxel space, sparse convolutional operators are highly efficient and can obtain a real- time inference speed. SECOND is a seminal work that implements these two sparse operators with GPU-based hash tables and builds a sparse convolutional network to extract 3D voxel features. This network architecture has been applied in nu- merous works [385, 350, 347, 297, 81, 36, 388, 333, 57, 375] and becomes the most widely-used backbone network in voxel- based detectors. There is also a series of works trying to improve the sparse operators , extend into a two-stage detec- tor [254, 57], and introduce the Transformer architecture into voxel-based detection [187, 70].
Analysis: pros and cons of different grid representations. In contrast to the 2D representations like BEV feature maps and pillars, voxels contain more structured 3D information. In addi- tion, deep voxel features can be learned through a 3D sparse net- work. However, a 3D neural network brings additional time and memory costs. BEV feature map is the most efficient grid repre- sentation that directly projects point cloud into a 2D pseudo im- age without specialized 3D operators like sparse convolutions or pillar encoding. 2D detection techniques can also be seamlessly applied to BEV feature maps without much modification. BEV- based detection methods generally can obtain high efficiency and a real-time inference speed. However, simply summarizing points statistics inside pixel regions loses too much 3D informa- tion, which leads to less accurate detection results compared to voxel-based detection. Pillar-based detection approaches lever- age PointNet to encode 3D points information inside a pillar cell, and the features are then scattered back into a 2D pseudo image for efficient detection, which balances the effectiveness and effi- ciency of 3D object detection.
Analysis: challenges of the grid-based detection methods. A critical problem that all grid-based methods have to face is choos- ing the proper size of grid cells. Grid representations are essen- tially discrete formats of point clouds by converting the con- tinuous point coordinates into discrete grid indices. The quan- tization process inevitably loses some 3D information and its efficacy largely depends on the size of grid cells: smaller grid size yields high resolution grids, and hence maintains more fine- grained details that are crucial to accurate 3D object detection.
Nevertheless, reducing the size of grid cells leads to a quadratic increase in memory consumption for the 2D grid representations like BEV feature maps or pillars. As for the 3D grid represen- tation like voxels, the problem can become more severe. There- fore, how to balance the efficacy brought by smaller grid sizes and the efficiency influenced by the memory increase remains an open challenge to all grid-based 3D object detection methods.
3.1.3 Point-Voxel Based 3D Object Detection
Point-voxel based approaches resort to a hybrid architecture that leverages both points and voxels for 3D object detection. Those methods can be divided into two categories: the single-stage and two-stage detection frameworks. An illustration of the two cate- gories is shown in Figure 6 and a taxonomy is in Table 4.
Single-stage point-voxel detection frameworks. Single-stage point-voxel based 3D object detectors try to bridge the features of points and voxels with the point-to-voxel and voxel-to-point transform in the backbone networks. Points contain fine-grained geometric information and voxels are efficient for computation, and combining them together in the feature extraction stage natu- rally benefits from both two representations. The idea that lever- ages point-voxel feature fusion in backbones has been explored by many works, with the contributions like point-voxel convo- lutions [165, 273], auxiliary point-based networks [94, 131, 58], and multi-scale feature fusion [195, 204, 88].
Two-stage point-voxel detection frameworks. Two-stage point- voxel based 3D object detectors resort to different data repre- sentations for different detection stages. Specifically, at the first stage, they employ a voxel-based detection framework to gener- ate a set of 3D object proposals. In the second stage, keypoints are first sampled from the input point cloud, and then the 3D pro- posals are further refined from the keypoints through novel point
Operators. Pv-Rcnn Is A Seminal Work That Adopts
as the first-stage detector, and the RoI-grid pooling operator is proposed for the second-stage refinement. The following works try to improve the second-stage head with novel modules and op-
8
Jiageng Mao et al.
Voxels
Stage-I : Generating 3D Proposals with Voxel-based Detector
Box Refinement
Fig. 6: An illustration of point-voxel based 3D object detection methods. Table 4: A taxonomy of point-voxel based detection methods.
Channel-Wise Transformer
erators, e.g. RefinerNet , VectorPool , point-wise atten-
Tion , Scale-Aware Pooling , Roi-Grid Attention ,
channel-wise Transformer , and point density-aware refine- ment module . Analysis: potentials and challenges of the point-voxel based methods. The point-voxel based methods can naturally benefit from both the fine-grained 3D shape and structure information obtained from points and the computational efficiency brought by voxels. However, some challenges still exist in these methods.
For the hybrid point-voxel backbones, the fusion of point and voxel features generally relies on the voxel-to-point and point- to-voxel transform mechanisms, which can bring non-negligible time costs. For the two-stage point-voxel detection frameworks, a critical challenge is how to efficiently aggregate point fea- tures for 3D proposals, as the existing modules and operators are generally time-consuming. In conclusion, compared to the pure voxel-based detection approaches, the point-voxel based detec- tion methods can obtain a better detection accuracy while at the cost of increasing the inference time.
3.1.4 Range-Based 3D Object Detection
Range image is a dense and compact 2D representation in which each pixel contains 3D distance information instead of RGB val- ues. Range-based methods address the detection problem from two aspects: designing new models and operators that are tai- lored for range images, and selecting suitable views for detec- tion. An illustration of the range-based 3D object detection meth- ods is shown in Figure 7 and a taxonomy is in Table 5.
Range-based detection models. Since range images are 2D rep- resentations like RGB images, range-based 3D object detectors can naturally borrow the models in 2D object detection to handle range images. LaserNet is a seminal work that leverages the deep layer aggregation network (DLA-Net) to obtain multi-scale features and detect 3D objects from range images.
Some papers also adopt other 2D object detection architectures, e.g. U-Net is applied in [193, 152, 269], RPN and
R-Cnn Are Employed In [152, 11], Fcn Is Used
in , and FPN is leveraged in . Range-based operators. Pixels of range images contain 3D dis- tance information instead of color values, so the standard convo- lutional operator in conventional 2D network architectures is not optimal for range-based detection, as the pixels in a sliding win- dow may be far away from each other in the 3D space. Some works resort to novel operators to effectively extract features from range pixels, including range dilated convolutions , graph operators , and meta-kernel convolutions .
Views for range-based detection. Range images are captured from the range view (RV), and ideally, the range view is a spher- ical projection of a point cloud. It has been a natural solution for many range-based approaches [192, 11, 69, 27] to detect 3D objects directly from the range view. Nevertheless, detection from the range view will inevitably suffer from the occlusion and scale-variation issues brought by the spherical projection.
To circumvent these issues, many methods have been working on leveraging other views for predicting 3D objects, e.g. the cylindrical view (CYV) leveraged in , a combination of 3D Object Detection for Autonomous Driving: A Comprehensive Survey
9
Table 5: A taxonomy of range-based detection methods based on views, models, and operators.
-
Rapoport-Lavie et al.
Transform Into Point View (Pv)
Fig. 7: An illustration of range-based 3D object detection. the range-view, bird’s-eye view (BEV), and/or point-view (PV) adopted in [153, 269, 193, 152].
Analysis: potentials and challenges of the range-based meth- ods. Range image is a dense and compact 2D representation, so the conventional or specialized 2D convolutions can be seam- lessly applied on range images, which makes the feature extrac- tion process quite efficient. Nevertheless, compared to bird’s-eye view detection, detection from the range view is vulnerable to occlusion and scale variation. Hence, feature extraction from the range view and object detection from the bird’s eye view be- comes the most practical solution to range-based 3D object de- tection.
3.2 Learning Objectives For 3D Object Detection
Problem and Challenge. Learning objectives are critical in ob- ject detection. Since 3D objects are quite small relative to the whole detection range, special mechanisms to enhance the lo- calization of small objects are strongly required in 3D detection.
On the other hand, considering point cloud is sparse and objects normally have incomplete shapes, accurately estimating the cen- ters and sizes of 3D objects is a long-standing challenge.
3.2.1 Anchor-Based 3D Object Detection
Anchors are pre-defined cuboids with fixed shapes that can be placed in the 3D space. 3D objects can be predicted based on the positive anchors that have a high intersection over union (IoU) with ground truth. We will introduce the anchor-based 3D ob- ject detection methods from the aspect of anchor configurations and loss functions. An illustration of anchor-based learning ob- jectives is shown in Figure 8 and a taxonomy is in Table 6.
Prerequisites. The ground truth 3D objects can be represented as [xg, yg, zg, lg, wg, hg, θg] with the class clsg. The anchors [xa, ya, za, la, wa, ha, θa] are used to generate predicted 3D ob- jects [x, y, z, l, w, h, θ] with a predicted class probability p.
Anchor configurations. Anchor-based 3D object detection ap- proaches generally detect 3D objects from the bird’s-eye view, in which 3D anchor boxes are placed at each grid cell of a BEV feature map. 3D anchors normally have a fixed size for each cat- egory, since objects of the same category have similar sizes.
Loss functions. The anchor-based methods employ the classifi- cation loss Lcls to learn the positive and negative anchors, and the regression loss Lreg is utilized to learn the size and loca- tion of an object based on a positive anchor. Additionally, Lθ is applied to learn the object’s heading angle. The loss function is Ldet = Lcls + Lreg + Lθ.
(4)
VoxelNet is a seminal work that leverages the anchors that have a high IoU with the ground truth 3D objects as posi- tive anchors, and the other anchors are treated as negatives. To accurately classify those positive and negative anchors, for each category, the binary cross entropy loss can be applied to each anchor on the BEV feature map, which can be formulated as
(5)
where p is the predicted probability for each anchor and the tar- get q is 1 if the anchor is positive and 0 otherwise. In addition to the binary cross entropy loss, the focal loss [157, 358] has also been employed to enhance the localization ability:
(6)
where α = 0.25 and γ = 2 are adopted in most works. The regression targets can be further applied to those positive anchors to learn the sizes and locations of 3D objects:
(La)2 + (Wa)2 Is The Diagonal Length Of An An-
chor from the bird’s-eye view. Then the SmoothL1 loss is adopted to regress the targets, which is represented as
V∈{∆X,∆Y,∆Z,∆L,∆W,∆H}
SmoothL1(u −v).
10
Jiageng Mao et al.
Bin-Based Heading Estimation
Fig. 8: An illustration of anchor-based learning objectives. To learn the heading angle θ, the radian orientation offset can
∆Θ = Θg −Θa,
Lθ = SmoothL1(θ −∆θ).
(9)
However, directly regressing the radian offset is normally hard due to the large regression range. Alternatively, the bin-based heading estimation is a better solution to learn the heading angle, in which the angle space is first divided into bins, and bin- based classification Ldir and residual regression are employed:
(10)
where ∆θ′ is the residual offset within a bin. The sine function
(11)
and Lθ can be computed following Eqn. 9 or Eqn. 10. In addition to the loss functions that learn the objects’ sizes, locations, and orientations separately, the intersection over union (IoU) loss that considers all object parameters as a whole
(12)
where bg and b are the ground truth and predicted 3D bound- ing boxes, and IoU(·) calculates the 3D IoU in a differential manner. Apart from the IoU loss, the corner loss is also introduced to minimize the distances between the eight corners
Where Cg
i and ci are the ith corner of the ground truth and pre- dicted cuboid respectively. Analysis: potentials and challenges of the anchor-based ap- proaches. The anchor-based methods can benefit from the prior knowledge that 3D objects of the same category should have similar shapes, so they can generate accurate object predictions with the help of 3D anchors. However, since 3D objects are rel- atively small with respect to the detection range, a large number of anchors are required to ensure complete coverage of the whole detection range, e.g. around 70k anchors are utilized in on the KITTI dataset. Furthermore, for those extremely small objects such as pedestrians and cyclists, applying anchor-based assignments can be quite challenging. Considering the fact that anchors are generally placed at the center of each grid cell, if the grid cell is large and objects in the cell are small, the anchor of this cell may have a low IoU with the small objects, which may hamper the training process.
Table 6: A taxonomy of anchor-based methods based on loss functions.
3.2.2 Anchor-Free 3D Object Detection
Anchor-free approaches eliminate the complicated anchor de- signs and can be flexibly applied to diverse views, e.g. the bird’s- eye view, point view, and range view. An illustration of anchor- free learning objectives is shown in Figure 9 and a taxonomy is in Table 7. The major difference between the anchor-based and anchor-free methods lies in the selection of positive and negative samples. We will introduce the anchor-free methods from the perspective of positive assignments, including grid-based, point- based, range-based, and set-to-set assignments. We still adopt the notations in Section 3.2.1 for simplicity.
Grid-based assignment. In contrast to the anchor-based meth- ods that rely on the IoUs with anchors to determine the positive and negative samples, the anchor-free methods leverage various grid-based assignment strategies for BEV grid cells, pillars, and voxels. PIXOR is a pioneering work that leverages the grid cells inside the ground truth 3D objects as positives, and the others as negatives. This inside-object assignment strategy is adopted in , and further improved in [81, 103, 36] by select- ing the grid cells nearest to the object center. CenterPoint utilizes a Gaussian kernel at each object center to assign posi- tive labels. These methods can still use Eqn. 5 or Eqn. 6 as the
Classification Loss, And The Regression Target Is
∆= [dx, dy, zg, log(lg), log(wg), log(hg), sin(θg), cos(θg)],
(14)
where dx and dy are the offsets between positive grid cells and object centers. The SmoothL1 loss is leveraged to regress ∆. Point-based assignment. Most point-based detection approaches resort to the anchor-free and point-based assignment strategy, in which the points are first segmented and those foreground points inside or near 3D objects are selected as positive samples, and 3D bounding boxes are finally learned from those foreground points. This foreground point segmentation strategy has been adopted in most point-based detectors [252, 342, 339, 208], with improvements such as adding centerness scores , etc.
Range-based assignment. Anchor-free assignments can also be employed on range images. A common solution is to select the range pixels inside 3D objects as positive samples, which has been adopted in [192, 69]. Different from other methods where the regression targets are based on the global 3D coordinate sys- tem, the range-based methods resort to an object-centric coordi- 3D Object Detection for Autonomous Driving: A Comprehensive Survey
Anchor-Free Assignments
Fig. 9: An illustration of anchor-free learning objectives. Table 7: A taxonomy of anchor-free detection methods based on the sample types for prediction and the assignment strategies.
Inside Objects
nate system for regression. Eqn. 14 can still be applied in these methods with an additional coordinate transform. Set-to-set assignment. DETR is an influential 2D detection method that introduces a set-to-set assignment strategy to auto- matically assign the predictions to the respective ground truths
(15)
where M is a one-to-one mapping from each positive sample to a 3D object. The set-to-set assignments have also been explored in 3D object detection approaches [196, 297, 332], and fur- ther introduces a novel cost function for the Hungarian matching.
Analysis: potentials and challenges of the anchor-free ap- proaches. The anchor-free detection methods abandon the com- plicated anchor design and exhibit stronger flexibility in terms of the assignment strategies. With the anchor-free assignments, 3D objects can be predicted directly on various representations, in- cluding points, range pixels, voxels, pillars, and BEV grid cells.
The learning process is also greatly simplified without introduc- ing additional shape priors. Among those anchor-free methods, the center-based methods have shown great potential in detecting small objects and have outperformed the anchor-based detection methods on the widely used benchmarks [15, 268].
Despite these merits, a general challenge to the anchor-free methods is to properly select positive samples to generate 3D object predictions. In contrast to the anchor-based methods that only select those high IoU samples, the anchor-free methods may possibly select some bad positive samples that yield inaccurate object predictions. Hence, careful design to filter out those bad positives is important in most anchor-free methods.
3.2.3 3D Object Detection With Auxiliary Tasks
Numerous approaches resort to auxiliary tasks to enhance the spatial features and provide implicit guidance for accurate 3D object detection. The commonly used auxiliary tasks include se- mantic segmentation, intersection over union prediction, object shape completion, and object part estimation.
Table 8: A taxonomy of detection methods based on auxiliary tasks.
[36, 254]
Semantic segmentation. Semantic segmentation can help 3D object detection in 3 aspects: (1) Foreground segmentation could provide implicit information on objects’ locations. Point-wise foreground segmentation has been broadly adopted in most point- based 3D object detectors [252, 379, 340, 94] for proposal gen- eration. (2) Spatial features can be enhanced by segmentation.
In , a semantic context encoder is leveraged to enhance spa- tial features with semantic knowledge. (3) Semantic segmenta- tion can be utilized as a pre-processing step to filter out back- ground samples and make 3D object detection more efficient.
and leverage semantic segmentation to remove those redundant points to speed up the subsequent detection model. IoU prediction. Intersection over union (IoU) can serve as a useful supervisory signal to rectify the object confidence scores.
proposes an auxiliary branch to predict an IoU score SIoU for each detected 3D object. During inference, the original con- fidence scores Sconf = Scls from the conventional classification branch are further rectified by the IoU scores SIoU:
(16)
where the hyper-parameter β controls the degrees of suppressing the low-IoU predictions and enhancing the high-IoU predictions. With the IoU rectification, the high-quality 3D objects are easier to be selected as the final predictions. Similar designs have also been adopted in [376, 153, 81, 103].
Object shape completion. Due to the nature of LiDAR sen- sors, faraway objects generally receive only a few points on their surfaces, so 3D objects are generally sparse and incomplete. A straightforward way of boosting the detection performance is to complete object shapes from sparse point clouds. Complete shapes could provide more useful information for accurate and robust detection. Many shape completion techniques have been proposed in 3D detection, including a shape decoder , shape signatures , and a probabilistic occupancy grid [329, 328].
Object part estimation. Identifying the part information inside objects is helpful in 3D object detection, as it reveals more fine- grained 3D structure information of an object. Object part esti- mation has been explored in some works [36, 254].
Analysis: future prospects of multitask learning for 3D ob- ject detection. 3D object detection is innately correlated with many other 3D perception and generation tasks. Multitask learn- ing of 3D detection and segmentation is more beneficial com- pared to training 3D object detectors independently, and shape completion can also help 3D object detection. There are also other tasks that can help boost the performance of 3D object de- tectors. For instance, scene flow estimation could identify static and moving objects, and tracking the same 3D object in a point cloud sequence yields a more accurate estimation of this object.
Hence, it will be promising to integrate more perception tasks into the existing 3D object detection pipeline.
12
Jiageng Mao et al.
(Chen Et Al.)
Fig. 10: Chronological overview of the camera-based 3D object detection methods.
4 Camera-Based 3D Object Detection
In this section, we introduce camera-based 3D object detection methods. In Section 4.1, we review and analyze the monocular 3D object detection methods, which can be further divided into the image-only, depth-assisted, and prior-guided approaches. In Section 4.2, we investigate the 3D object detection methods based on stereo images. In Section 4.3, we introduce the 3D object de- tection methods with multiple cameras. A chronological overview of the camera-based 3D object detection methods is shown in Figure 10.
4.1 Monocular 3D Object Detection
Problem and Challenge. Detecting objects in the 3D space from monocular images is an ill-posed problem since a single image cannot provide sufficient depth information. Accurately predict- ing the 3D locations of objects is the major challenge in monoc- ular 3D object detection. Many endeavors have been made to tackle the object localization problem, e.g. inferring depth from images, leveraging geometric constraints and shape priors. Nev- ertheless, the problem is far from being solved. Monocular 3D detection methods still perform much worse than the LiDAR- based methods due the poor 3D localization ability, which leaves an open challenge to the research community.
4.1.1 Image-Only Monocular 3D Object Detection
Inspired by the 2D detection approaches, a straightforward so- lution to monocular 3D object detection is to directly regress the 3D box parameters from images via a convolutional neu- ral network. The direct-regression methods naturally borrow de- signs from the 2D detection network architectures, and can be trained in an end-to-end manner. These approaches can be di- vided into the single-stage/two-stage, or anchor-based/anchor- free methods. An illustration of image-only 3D object detection is shown in Figure 11 and a taxonomy is in Table 9.
Single-stage anchor-based methods. Anchor-based monocu- lar detection approaches rely on a set of 2D-3D anchor boxes placed at each image pixel, and use a 2D convolutional neural network to regress object parameters from the anchors. Specifi- cally, for each pixel [u, v] on the image plane, a set of 3D anchors [wa, ha, la, θa]3D, 2D anchors [wa, ha]2D, and depth anchors da are pre-defined. An image is passed through a convolutional net- work to predict the 2D box offsets δ2D = [δx, δy, δw, δh]2D and the 3D box offsets δ3D = [δx, δy, δd, δw, δh, δl, δθ]3D based on each anchor. Then, the 2D bounding boxes b2D = [x, y, w, h]2D Table 9: A taxonomy of image-only monocular detection meth- ods based on frameworks.
(17)
and the 3D bounding boxes b3D = [x, y, z, l, w, h, θ]3D can be
(18)
where [uc, vc] is the projected object center on the image plane. Finally, the projected center [uc, vc] and its depth dc are trans-
(19)
where K and T are the camera intrinsics and extrinsics. M3D-RPN is a seminal paper that proposes the anchor- based framework, and many papers have tried to improve this framework, e.g. extending it into video-based 3D detection , introducing differential non-maximum suppression , design- ing an asymmetric attention module .
Single-stage anchor-free methods. Anchor-free monocular de- tection approaches predict the attributes of 3D objects from im- ages without the aid of anchors. Specifically, an image is passed through a 2D convolutional neural network and then multiple heads are applied to predict the object attributes separately. The prediction heads generally include a category head to predict the object’s category, a keypoint head to predict the coarse object center [u, v], an offset head to predict the center offset [δx, δy] based on [u, v], a depth head to predict the depth offset δd, a size head to predict the object size [w, h, l], and an orientation head to predict the observation angle α. The 3D object center [x, y, z] can be converted from the projected center [uc, vc] and depth dc: 3D Object Detection for Autonomous Driving: A Comprehensive Survey
Refine
Single-Stage Anchor-based Monocular 3D Object Detection Single-Stage Anchor-Free Monocular 3D Object Detection
To The 3D Space
Fig. 11: An illustration of image-only monocular 3D object de- tection methods. Image samples are from .
(20)
where σ is the sigmoid function. The yaw angle θ of an object can be converted from the observation angle α using
Θ = Α + Arctan(X
z ).
(21)
CenterNet first introduces the single-stage anchor-free framework for monocular 3D object detection. Many follow- ing papers work on improving this framework, including novel depth estimation schemes [166, 294, 369], an FCOS -like architecture , a new IoU-based loss function , key- points , pair-wise relationships , camera extrinsics pre- diction , and view transforms [241, 237, 262].
Two-stage methods. Two-stage monocular detection approaches generally extend the conventional two-stage 2D detection archi- tectures to 3D object detection. Specifically, they utilize a 2D detector in the first stage to generate 2D bounding boxes from an input image. Then in the second stage, the 2D boxes are lifted up to the 3D space by predicting the 3D object parameters from the 2D RoIs. ROI-10D extends the conventional Faster R- CNN architecture with a novel head to predict the parame- ters of 3D objects in the second stage. A similar design paradigm has been adopted in many works with improvements like disen- tangling the 2D and 3D detection loss , predicting heading angles in the first stage , learning more accurate depth in- formation [233, 258, 172].
Table 10: A taxonomy of depth-assisted monocular detection methods based on data representations and detection networks (2D: convolutional networks; 3D: point cloud networks).
Lidar
Coord.
✓
✓
✓
✓
Weng et al.
✓
✓
Analysis: potentials and challenges of the image-only meth- ods. The image-only methods aim to directly regress the 3D box parameters from images via a modified 2D object detection framework. Since these methods take inspiration from the 2D detection methods, they can naturally benefit from the advances in 2D object detection and image-based network architectures.
Most methods can be trained end-to-end without pre-training or post-processing, which is quite simple and efficient. A critical challenge of the image-only methods is to accu- rately predict depth dc for each 3D object. As shown in , simply replacing the predicted depth with ground truth yields more than 20% car AP gain on the KITTI dataset, while replacing other parameters only results in an incremental gain.
This observation indicates that the depth error dominates the to- tal errors and becomes the most critical factor hampering accu- rate monocular detection. Nevertheless, depth estimation from monocular images is an ill-posed problem, and the problem be- comes severer with only box-level supervisory signals.
4.1.2 Depth-assisted monocular 3D object detection Depth estimation is critical in monocular 3D object detection. To achieve more accurate monocular detection results, many pa- pers resort to pre-training an auxiliary depth estimation network.
Specifically, a monocular image is first passed through a pre- trained depth estimator, e.g. MonoDepth or DORN , to generate a depth image. Then, there are mainly two categories of methods to deal with depth images and monocular images.
The depth-image based methods fuse images and depth maps with a specialized neural network to generate depth-aware fea- tures that could enhance the detection performance. The pseudo- LiDAR based methods convert a depth image into a pseudo- LiDAR point cloud, and LiDAR-based detectors can then be ap- plied to the point cloud to predict 3D objects. An illustration of depth-assisted monocular 3D object detection is shown in Fig- ure 12 and a taxonomy of these methods is in Table 10.
Depth-image based methods. Most depth-image based meth- ods leverage two backbone networks for RGB and depth images respectively. They obtain depth-aware image features by fusing the information from the two backbones with specialized opera- tors. More accurate 3D bounding boxes can be learned from the depth-ware features and can be further refined with depth im- ages. MultiFusion is a pioneering work that introduces the depth-image based detection framework. Following papers adopt similar design paradigms with improvements in network archi- tectures, operators, and training strategies, e.g. a point-based at- tentional network , depth-guided convolutions , depth-
14
Jiageng Mao et al.
Coordinate Map
Fig. 12: An illustration of depth-assisted monocular 3D object detection methods. Image and depth samples are from . conditioned message passing , disentangling appearance and localization features , and a novel depth pre-training framework .
Pseudo-LiDAR based methods. Pseudo-LiDAR based methods transform a depth image into a pseudo-LiDAR point cloud, and LiDAR-based detectors can then be employed to detect 3D ob- jects from the point cloud. Pseudo-LiDAR point cloud is a data representation first introduced in , where they convert a depth map D ∈RH×W into a pseudo point cloud P ∈RHW ×3.
Specifically, for each pixel [u, v] and its depth value d in a depth image, the corresponding 3D point coordinate [x, y, z] in the
(22)
where [cu, cv] is the camera principal point, and fu and fv are the focal lengths along the horizontal and vertical axis respec- tively. Thus P can be obtained by back-projecting each pixel in D into the 3D space. P is referred as the pseudo-LiDAR rep- resentation: it is essentially a 3D point cloud but is extracted from a depth image instead of a real LiDAR sensor. Finally, LiDAR-based 3D object detectors can be directly applied on the pseudo-LiDAR point cloud P to predict 3D objects. Many papers have worked on improving the pseudo-LiDAR detection framework, including augmenting pseudo point cloud with color information , introducing instance segmentation , de- signing a progressive coordinate transform scheme , im- proving pixel-wise depth estimation with separate foreground and background prediction , domain adaptation from real LiDAR point cloud , and a new physical sensor design .
PatchNet challenges the conventional idea of leverag- ing the pseudo-LiDAR representation P ∈RHW ×3 for monoc- ular 3D object detection. They conduct an in-depth investigation and provide an insightful observation that the power of pseudo- LiDAR representation comes from the coordinate transformation instead of the point cloud representation. Hence, a coordinate map M ∈RH×W ×3 where each pixel encodes a 3D coordi- nate can attain a comparable monocular detection result with the pseudo-LiDAR point cloud representation. This observation en- ables us to directly apply a 2D neural network on the coordinate map to predict 3D objects, eliminating the need of leveraging the time-consuming LiDAR-based detectors on point clouds.
Analysis: potentials and challenges of the depth-assisted ap- proaches. The depth-assisted approaches pursue more accurate depth estimation by leveraging a pre-trained depth prediction network. Both the depth image representation and the pseudo- LiDAR presentation could significantly boost the monocular de- tection performance. Nevertheless, compared to the image-only methods that only require 3D box annotations, pre-training a depth prediction network requires expensive ground truth depth maps, and it also hampers the end-to-end training of the whole framework. Furthermore, pre-trained depth estimation networks suffer from poor generalization ability. Pretrained depth maps are usually not well calibrated on the target dataset and typi- cally the scale needs to be adapted to the target dataset. Thus there remains a non-negligible domain gap between the source domain leveraged for depth pre-training and the target domain for monocular detection. Given the fact that driving scenarios are normally diverse and complex, pre-training depth networks on a restricted domain may not work well in real-world applications.
4.1.3 Prior-Guided Monocular 3D Object Detection
Numerous approaches try to tackle the ill-posed monocular 3D object detection problem by leveraging the hidden prior knowl- edge of object shapes and scene geometry from images. The prior knowledge can be learned by introducing pre-trained sub- networks or auxiliary tasks, and they can provide extra informa- tion or constraints to help accurately localize 3D objects. The broadly adopted prior knowledge includes object shapes, geom- etry consistency, temporal constraints, and segmentation infor- mation. An illustration of the prior types is shown in Figure 13.
Object shapes. Many methods resort to shape reconstruction of 3D objects directly from images. The reconstructed shapes can be further leveraged to determine the locations and poses of 3D Object Detection for Autonomous Driving: A Comprehensive Survey
Temporal Constraint
Temporal Constraint
Fig. 13: An illustration of the prior types in monocular 3D object detection methods. Samples are from [39, 98, 319, 9, 213]. the 3D objects. There are 5 types of reconstructed representa- tions: computer-aided design (CAD) models, wireframe models, signed distance function (SDF), points, and voxels.
Some papers [362, 25, 98] learn morphable wireframe mod- els to represent 3D objects. Other works [122, 182, 359, 9] lever- age DeepSDF to learn implicit signed distance functions or low-dimensional shape parameters from CAD models, and they further propose a render-and-compare approach to learn the parameters of 3D objects. Some works [319, 320] utilize voxel patterns to represent 3D objects. Other papers [118, 31] resort to point cloud reconstruction from images and estimate the loca- tions of 3D objects with 2D-3D correspondences.
Geometric consistency. Given the extrinsics matrix T ∈SE(3) that transforms a 3D coordinate in the object frame to the camera frame, and the camera intrinsics matrix K that project the 3D coordinate onto the image plane, the projection of a 3D point [x, y, z] in the object frame into the image pixel coordinate [u, v]
(23)
where d is the depth of transformed 3D coordinate in the camera frame. Eqn. 23 provides a geometric relationship between 3D points and 2D image pixel coordinates, which can be leveraged in various ways to encourage consistency between the predicted 3D objects and the 2D objects on images. There are mainly 5 types of geometric constraints in monocular detection: 2D-3D boxes consistency, keypoints, object’s height-depth relationship, inter-objects relationship, and ground plane constraints.
Some works [197, 13, 112, 200] propose to encourage the consistency between 2D and 3D boxes by minimizing reprojec- tion errors. These methods introduce a post-processing step to optimize the 3D object parameters by gradually fitting the pro- jected 3D boxes to 2D bounding boxes on images. There is also a branch of papers [135, 257, 133] that predict the object key- points from images, and the keypoints can be leveraged to cali- brate the sizes of locations of 3D objects. Object’s height-depth relationship can also serve as a strong geometric prior. Specifi- cally, given the physical height of an object H in the 3D space, the visual height h on images, and the corresponding depth of the object d, there exists a geometric constraint: d = f ·H/h, where f is the camera focal length. This constraint can be leveraged to Table 11: A taxonomy of prior-guided monocular detection methods based on prior types.
[319, 9, 39, 99]
obtain more accurate depth estimation and has been broadly ap- plied in a lot of works [17, 369, 172, 258]. There are also some papers [381, 47] trying to model the inter-objects relationships by exploiting new geometric relations among objects. Other pa- pers [39, 13, 159, 161] leverage the assumption that 3D objects are generally on the ground plane to better localize those objects.
Temporal constraints. Temporal association of 3D objects can be leveraged as strong prior knowledge. The temporal object relationships have been exploited as depth-ordering and multi-frame object fusion with a 3D Kalman filter .
Segmentation Image segmentation helps monocular 3D object detection mainly in two aspects. First, object segmentation masks are crucial for instance shape reconstruction in some works [319, 9]. Second, segmentation indicates whether an image pixel is in- side a 3D object from the perspective view, and this information has been utilized in [39, 99] to help localize 3D objects.
Analysis: potentials and challenges of leveraging prior knowl- edge in monocular 3D detection. With shape reconstruction, we could obtain more detailed object shape information from images, which is beneficial to 3D object detection. We can also attain more accurate detection results through the projection or render-and-compare loss. However, there exist two challenges for shape reconstruction applied in monocular 3D object detec- tion. First, shape reconstruction normally requires an additional step of pre-training a reconstruction network, which hampers end-to-end training of the monocular detection pipeline. Second, object shapes are generally learned from CAD models instead of real-world instances, which imposes the challenge of generaliz- ing the reconstructed objects to real-world scenarios.
Geometric consistencies are broadly adopted and can help improve detection accuracy. Nevertheless, some methods formu- late the geometric consistency as an optimization problem and optimize object parameters in post-processing, which is quite time-consuming and hampers end-to-end training.
Image segmentation is useful information in monocular 3D detection. However, training segmentation networks requires ex- pensive pixel annotations. Pre-training segmentation models on external datasets will suffer from the generalization problem.
4.2 Stereo-Based 3D Object Detection
Problem and Challenge. Stereo-based 3D object detection aims to detect 3D objects from a pair of images. Compared to monoc- ular images, paired stereo images provide additional geometric constraints that can be utilized to infer more accurate depth in- formation. Hence, the stereo-based methods generally obtain a
16
Jiageng Mao et al.
Volume-Based Stereo 3D Object Detection
Fig. 14: An illustration of stereo-based 3D object detection methods. Image and disparity samples are from . better detection performance than the monocular-based meth- ods. Nevertheless, stereo cameras typically require very accurate calibration and synchronization, which are normally difficult to achieve in real applications. An illustration of stereo-based 3D object detection approaches is shown in Figure 14 and a taxon- omy is in Table 12.
Stereo matching and depth estimation. A stereo camera can produce a pair of images, i.e. the left image IL and the right im- age IR, in one shot. With the stereo matching techniques [188, 29], a disparity map can be estimated from the paired stereo im- ages leveraging multi-view geometry . Ideally, for each pixel on the left image IL(u, v), there exists a pixel on the right image IR(u, v + p) with the disparity value p so that the two pixels picture the same 3D location. Finally, the disparity map can be transformed into a depth image with the following formula:
(24)
where d is the depth value, f is the focal length, and b is the baseline length of the stereo camera. The pixel-wise disparity constraints from stereo images enable more accurate depth esti- mation compared to monocular depth prediction.
2D-detection based methods. Conventional 2D object detection frameworks can be modified to resolve the stereo detection prob- lem. Specifically, paired stereo images are passed through an image-based detector with Siamese backbone networks to gen- erate left and right regions of interest (RoIs) for the left and right images respectively. Then in the second stage, the left and right RoIs are fused to estimate the parameters of 3D objects. Stereo R-CNN first proposes to extend 2D detection frameworks to stereo 3D detection. This design paradigm has been adopted in numerous papers. proposes a novel stereo triangulation Table 12: A taxonomy of stereo-based detection methods based on auxiliary tasks and data representations.
2D
Det.
2D
Seg.
✓
✓
Qian et al.
✓
✓
learning sub-network at the second stage; [331, 223, 267, 32] learn instance-level disparity by object-centric stereo matching and instance segmentation; proposes adaptive instance dis- parity estimation; [160, 217] introduce single-stage stereo detec- tion frameworks; [38, 40] propose an energy-based framework for stereo-based 3D object detection.
Pseudo-LiDAR based methods. The disparity map predicted from stereo images can be transformed into the depth image and then converted into the pseudo-LiDAR point cloud. Hence, sim- ilar to the monocular detection methods, the pseudo-LiDAR rep- resentation can also be employed in stereo-based 3D object de- tection methods. Those methods try to improve the disparity es- 3D Object Detection for Autonomous Driving: A Comprehensive Survey
Query-Based Multi-View 3D Object Detection
Fig. 15: An illustration of multi-view 3D object detection methods. Figures are from and . timation in stereo matching for more accurate depth prediction. introduces a depth cost volume in stereo matching net- works; proposes an end-to-end stereo matching and detec- tion framework; [116, 128] leverage semantic segmentation and predict disparity for foreground and background regions sepa- rately; proposes a Wasserstein loss for disparity estimation.
Volume-based methods. There exists a category of methods that skip the pseudo-LiDAR representation and perform 3D object detection directly on 3D stereo volumes. DSGN proposes a 3D geometric volume derived from stereo matching networks and applies a grid-based 3D detector on the volume to detect 3D
Objects. And Improve By Leveraging Knowledge
distillation and 3D feature volumes respectively. Potentials and challenges of the stereo-based methods. Com- pared to the monocular detection methods, the stereo-based meth- ods can obtain more accurate depth and disparity estimation with stereo matching techniques, which brings a stronger object lo- calization ability and significantly boosts the 3D object detection performance. Nevertheless, an auxiliary stereo matching network brings additional time and memory consumption. Compared to LiDAR-based 3D object detection, detection from stereo images can serve as a much cheaper solution for 3D perception in au- tonomous driving scenarios. However, there still exists a non- negligible performance gap between the stereo-based and the LiDAR-based 3D object detection approaches.
4.3 Multi-View 3D Object Detection
Problem and Challenge. Autonomous vehicles are generally equipped with multiple cameras to obtain complete environmen- tal information from multiple viewpoints. Recently, multi-view 3D object detection has evolved rapidly. Some multi-view 3D detection approaches try to construct a unified BEV space by projecting multi-view images into the bird’s-eye view, and then employ a BEV-based detector on top of the unified BEV fea- ture map to detect 3D objects. The transformation from cam- era views to the bird’s-eye view is ambiguous without accurate depth information, so image pixels and their BEV locations are not perfectly aligned. How to build reliable transformations from camera views to the bird’s-eye view is a major challenge in these methods. Other methods resort to 3D object queries that are gen- erated from the bird’s-eye view and Transformers where cross- view attention is applied to object queries and multi-view image features. The major challenge is how to properly generate 3D object queries and design more effective attention mechanisms in Transformers.
BEV-based multi-view 3D object detection. LSS is a pi- oneering work that proposes a lift-splat-shoot paradigm to solve the problem of BEV perception from multi-view cameras. There are three steps in LSS. Lift: bin-based depth prediction is con- ducted on image pixels and multi-view image features are lifted to 3D frustums with depth bins. Splat: 3D frustums are splat- ted into a unified bird’s-eye view plane and image features are transformed into BEV features in an end-to-end manner. Shoot: downstream perception tasks are performed on top of the BEV feature map. This paradigm has been successfully adopted by many following works. BEVDet [106, 105] improves LSS with a four-step multi-view detection pipeline, where the image view encoder encodes features from multi-view images, the view transformer transforms image features from camera views to the bird’s-eye view, the BEV encoder further encodes the BEV fea- tures, and the detection head is employed on top of the BEV features for 3D detection. The major bottleneck in [106, 219] is depth prediction, as it is normally inaccurate and will result in inaccurate feature transforms from camera views to the bird’s- eye view. To obtain more accurate depth information, many pa- pers resort to mining additional information from multi-view im- ages and past frames, e.g. leverages explicit depth super- vision, introduces surround-view temporal stereo, uses dynamic temporal stereo, combines both short-term and long-term temporal stereo for depth prediction. In addition, there are also some papers [323, 104] that completely abandon the design of depth bins and categorical depth prediction. They simply assume that the depth distribution along the ray is uni- form, so the camera-to-BEV transformation can be conducted with higher efficiency.
Query-based multi-view 3D object detection. In addition to the BEV-based approaches, there is also a category of methods where object queries are generated from the bird’s-eye view and interact with camera view features. Inspired by the advances in Transformers for object detection , DETR3D intro- duces a sparse set of 3D object queries, and each query cor- responds to a 3D reference point. The 3D reference points can collect image features by projecting their 3D locations onto the multi-view image planes and then object queries interact with image features through Transformer layers. Finally, each object query will decode a 3D bounding box. Many following papers try to improve this design paradigm, such as introducing spatially- aware cross-view attention and adding 3D positional em- beddings on top of image features . BEVFormer in- troduces dense grid-based BEV queries and each query corre- sponds to a pillar that contains a set of 3D reference points.
Spatial cross-attention is applied to object queries and sparse image features to obtain spatial information, and temporal self- attention is applied to object queries and past BEV queries to fuse temporal information.
5 Multi-Modal 3D Object Detection
In this section, we introduce the multi-modal 3D object detection approaches that fuse multiple sensory inputs. According to the sensor types, the approaches can be divided into three categories:
18
Jiageng Mao et al.
(Fang Et Al.)
Camera-LiDAR Early-Fusion based 3D Object Detector Camera-LiDAR Intermediate-Fusion based 3D Object Detector
(Chen Et Al.)
Fig. 16: Chronological overview of the multi-modal 3D object detection methods. Table 13: A taxonomy of multi-sensor fusion-based detection methods based on fused stages, representations, and operators.
Output
Camera rep. LiDAR rep.
Box Consistency
LiDAR-camera, radar, and map fusion-based methods. In Sec- tion 5.1, we review and analyze the multi-modal detection ap- proaches with LiDAR-camera fusion, including the early-fusion based, the intermediate-fusion based, and the late-fusion based methods. In Section 5.2, we investigate the multi-modal detec- tion approaches with radar signals. In Section 5.3, we introduce the multi-modal 3D detection approaches with high-definition maps. A chronological overview of the multi-modal 3D object detection approaches is shown in Figure 16.
5.1 Multi-modal detection with LiDAR-camera fusion Problem and Challenge. Camera and LiDAR are two comple- mentary sensor types for 3D object detection. Cameras provide color information from which rich semantic features can be ex- tracted, while LiDAR sensors specialize in 3D localization and provide rich information about 3D structures. Many endeavors have been made to fuse the information from cameras and Li- DARs for accurate 3D object detection. Since LiDAR-based de- tection methods perform much better than camera-based meth- ods, the state-of-the-art approaches are mainly based on LiDAR- based 3D object detectors and try to incorporate image informa- tion into different stages of a LiDAR detection pipeline. In view of the complexity of LiDAR-based and camera-based detection systems, combining the two modalities together inevitably brings additional computational overhead and inference time latency.
Therefore, how to efficiently fuse the multi-modal information remains an open challenge. A taxonomy of multi-modal 3D ob- ject detection methods is in Table 13.
5.1.1 Early-Fusion Based 3D Object Detection
Early-fusion based methods aim to incorporate the knowledge from images into point cloud before they are fed into a LiDAR- based detection pipeline. Hence the early-fusion frameworks are generally built in a sequential manner: 2D detection or segmen- tation networks are firstly employed to extract knowledge from images, and then the image knowledge is passed to point cloud, and finally the enhanced point cloud is fed to a LiDAR-based 3D object detector. Based on the fusion types, the early-fusion methods can be divided into two categories: region-level knowl- edge fusion and point-level knowledge fusion. An illustration of the early-fusion based approaches is shown in Figure 17.
Region-level knowledge fusion. Region-level fusion methods aim to leverage knowledge from images to narrow down the ob- ject candidate regions in 3D point cloud. Specifically, an image is first passed through a 2D object detector to generate 2D bound- ing boxes, and then the 2D boxes are extruded into 3D viewing frustums. The 3D viewing frustums are applied on LiDAR point cloud to reduce the searching space. Finally, only the selected point cloud regions are fed into a LiDAR detector for 3D ob- ject detection. F-PointNet first proposes this fusion mecha- nism, and many endeavors have been made to improve the fusion framework. divides a viewing frustum into grid cells and applies a convolutional network on the grid cells for 3D detec- tion; proposes a novel geometric agreement search; exploits the pillar representation; introduces a model fitting algorithm to find the object point cloud inside each frustum.
Point-level knowledge fusion. Point-level fusion methods aim to augment input point cloud with image features. The augmented point cloud is then fed into a LiDAR detector to attain a bet- ter detection result. PointPainting is a seminal work that leverages image-based semantic segmentation to augment point clouds. Specifically, an image is passed through a segmentation network to obtain pixel-wise semantic labels, and then the se- mantic labels are attached to the 3D points by point-to-pixel projection. Finally, the points with semantic labels are fed into a LiDAR-based 3D object detector. This design paradigm has been 3D Object Detection for Autonomous Driving: A Comprehensive Survey
3D Detection Results
3D Detection Results
Fig. 17: An illustration of early-fusion based 3D object detection methods.
Roi Feature Fusion
Feature Fusion in Proposal Generation and RoI head
Roi Head
Fig. 18: An illustration of intermediate-fusion based 3D object detection methods. followed by a lot of papers [330, 260, 191]. Apart from semantic segmentation, there also exist some works trying to exploit other information from images, e.g. depth image completion .
Analysis: potentials and challenges of the early-fusion meth- ods. The early-fusion based methods focus on augmenting point clouds with image information before they are passed through a LiDAR 3D object detection pipeline. Most methods are com- patible with a wide range of LiDAR-based 3D object detectors and can serve as a quite effective pre-processing step to boost detection performance. Nevertheless, the early-fusion methods generally perform multi-modal fusion and 3D object detection in a sequential manner, which brings additional inference la- tency. Given the fact that the fusion step generally requires a complicated 2D object detection or semantic segmentation net- work, the time cost brought by multi-modal fusion is normally non-negligible. Hence, how to perform multi-modal fusion effi- ciently at the early stage has become a critical challenge.
5.1.2 Intermediate-fusion based 3D object detection Intermediate-fusion based methods try to fuse image and LiDAR features at the intermediate stages of a LiDAR-based 3D object detector, e.g. in backbone networks, at the proposal generation stage, or at the RoI refinement stage. These methods can also be classified according to the fusion stages. An illustration of intermediate-fusion based approaches is shown in Figure 18.
Fusion in backbone networks. Many endeavors have been made to progressively fuse image and LiDAR features in the backbone networks. In those methods, point-to-pixel correspondences are firstly established by LiDAR-to-camera transform, and then with the point-to-pixel correspondences, features from a LiDAR back- bone can be fused with features from an image backbone through different fusion operators. The multi-modal fusion can be con- ducted in the intermediate layers of a grid-based detection back- bone, with novel fusion operators such as continuous convo-
Ai Adaptive Learning
This project focuses on ai adaptive learning using modern AI and machine learning techniques. The content below is adapted from research literature and practical implementation notes.
We propose a novel high-performance and interpretable canon-
addition, unlike tree learning, DNNs enable gradient descent- ical deep tabular data learning architecture, TabNet. TabNet based end-to-end learning for tabular data which can have a uses sequential attention to choose which features to reason multitude of benefits: (i) efficiently encoding multiple data from at each decision step, enabling interpretability and more types like images along with tabular data; (ii) alleviating the efficient learning as the learning capacity is used for the most need for feature engineering, which is currently a key aspect
salient features. We demonstrate that TabNet outperforms in tree-based tabular data learning methods; (iii) learning other variants on a wide range of non-performance-saturated from streaming data and perhaps most importantly (iv) end- tabular datasets and yields interpretable feature attributions to-end models allow representation learning which enables plus insights into its global behavior. Finally, we demonstrate many valuable application scenarios including data-efficient
self-supervised learning for tabular data, significantly improv- domain adaptation (Goodfellow, Bengio, and Courville 2016), ing performance when unlabeled data is abundant. generative modeling (Radford, Metz, and Chintala 2015) and
Introduction We propose a new canonical DNN architecture for tabular
Deep neural networks (DNNs) have shown notable success data, TabNet. The main contributions are summarized as: efficiently encode the raw data into meaningful representa- enabling flexible integration into end-to-end learning. tions, fuel the rapid progress. One data type that has yet to 2. TabNet uses sequential attention to choose which fea- see such success with a canonical architecture is tabular data. tures to reason from at each decision step, enabling in-
Despite being the most common data type in real-world AI terpretability and better learning as the learning capacity (as it is comprised of any categorical and numerical features), is used for the most salient features (see Fig. 1). This under-explored, with variants of ensemble decision trees for each input, and unlike other instance-wise feature se- Why? First, because DT-based approaches have certain bene- and van der Schaar 2019), TabNet employs a single deep
fits: (i) they are representionally efficient for decision mani- learning architecture for feature selection and reasoning. folds with approximately hyperplane boundaries which are 3. Above design choices lead to two valuable properties: (i) common in tabular data; and (ii) they are highly interpretable TabNet outperforms or is on par with other tabular learn- in their basic form (e.g. by tracking decision nodes) and there ing models on various datasets for classification and re-
are popular post-hoc explainability methods for their ensem- gression problems from different domains; and (ii) TabNet ble form, e.g. (Lundberg, Erion, and Lee 2018) – this is an enables two kinds of interpretability: local interpretability important concern in many real-world applications; (iii) they that visualizes the importance of features and how they are fast to train. Second, because previously-proposed DNN are combined, and global interpretability which quantifies
architectures are not well-suited for tabular data: e.g. stacked the contribution of each feature to the trained model. convolutional layers or multi-layer perceptrons (MLPs) are 4. Finally, for the first time for tabular data, we show signif- vastly overparametrized – the lack of appropriate inductive icant performance improvements by using unsupervised bias often causes them to fail to find optimal solutions for tab- pre-training to predict masked features (see Fig. 2).
ular decision manifolds (Goodfellow, Bengio, and Courville
Why is deep learning worth exploring for tabular data?
One obvious motivation is expected performance improve- Feature selection: Feature selection broadly refers to judi- Copyright © 2021, Association for the Advancement of Artificial ciously picking a subset of features based on their useful-
Professional occupation related Investment related
Feedback from Feedback to
Feature selection Input processing Feature selection Input processing
previous step next step … …
Predicted output (whether the income level >$50k)
selection enables interpretability and better learning as the capacity is used for the most salient features. TabNet employs multiple decision blocks that focus on processing a subset of input features for reasoning. Two decision blocks shown as examples process features that are related to professional occupation and investments, respectively, in order to predict the income level.
Unsupervised pre-training Supervised fine-tuning
Age Cap. gain Education Occupation Gender Relationship Age Cap. gain Education Occupation Gender Relationship 5 2000 ? Exec-managerial F Wife 6 2000 Bachelors Exec-managerial M Husband 1 0 ? Farming-fishing M ? 2 0 High-school Farming-fishing M Unmarried
? 50 Doctorate Prof-specialty M Husband 4 50 Doctorate Prof-specialty M Husband 2 ? ? Handlers-cleaners F Wife 2 0 High-school Handlers-cleaners F Wife 5 3000 Bachelors ? ? Husband 5 3000 Bachelors Exec-managerial M Husband
3 0 Bachelors ? F ? 3 100 Bachelors Prof-specialty F Wife ? 0 High-school Armed-Forces ? Husband 2 0 High-school Armed-Forces M Husband
TabNet decoder Decision making
Age Cap. gain Education Occupation Gender Relationship Income > $50k
3 M False
level can be guessed from the occupation, or the gender can be guessed from the relationship. Unsupervised representation learning by masked self-supervised learning results in an improved encoder model for the supervised learning task.
ward selection and Lasso regularization (Guyon and Elisseeff performance with compact representations. 2003) attribute feature importance based on the entire training Tree-based learning: DTs are commonly-used for tabular data, and are referred as global methods. Instance-wise fea- data learning. Their prominent strength is efficient picking ture selection refers to picking features individually for each of global features with the most statistical information gain
to maximize the mutual information between the selected mance of standard DTs, one common approach is ensembling features and the response variable, and in (Yoon, Jordon, and to reduce variance. Among ensembling methods, random van der Schaar 2019) by using an actor-critic framework to forests (Ho 1998) use random subsets of data with randomly mimic a baseline while optimizing the selection. Unlike these, selected features to grow many trees. XGBoost (Chen and
sity in end-to-end learning – a single model jointly performs recent ensemble DT approaches that dominate most of the feature selection and output mapping, resulting in superior recent data science competitions. Our experimental results
!# + Softmax !" < % !" > % !# > & !# > &
ReLU ReLU &
$" !" − $" % −1 −$" !" + $" % −1 −1 $# !# − $# & % −1 −$# !# + $# & !"
FC FC
W: [$" , - $" , 0, 0] W: [0, 0, $# , - $# ] !" < % b: [-a $" , a $" , -1, -1] b: [-1, -1, -d $# , d $# ] !# < & !" > % !# < & [!" ] [!# ]
M: [1, 0] M: [0, 1]
(right). Relevant features are selected by using multiplicative sparse masks on inputs. The selected features are linearly transformed, and after a bias addition (to represent boundaries) ReLU performs region selection by zeroing the regions. Aggregation of multiple regions is based on addition. As C and C get larger, the decision boundary gets sharper.
for various datasets show that tree-based models can be out- constructs a sequential multi-step architecture, where each performed when the representation capacity is improved with step contributes to a portion of the decision based on the deep learning while retaining their feature selecting property. selected features; (iii) improves the learning capacity via non- Integration of DNNs into DTs: Representing DTs with linear processing of the selected features; and (iv) mimics
DNN building blocks as in (Humbird, Peterson, and McClar- ensembling via higher dimensions and more steps. ren 2018) yields redundancy in representation and ineffi- cient learning. Soft (neural) DTs (Wang, Aggarwal, and Liu Fig. 4 shows the TabNet architecture for encoding tabu- functions, instead of non-differentiable axis-aligned splits. mapping of categorical features with trainable embeddings.
However, losing automatic feature selection often degrades We do not consider any global feature normalization, but performance. In (Yang, Morillo, and Hospedales 2018), a soft merely apply batch normalization (BN). We pass the same D- binning function is proposed to simulate DTs in DNNs, by dimensional features f ∈ <B×D to each decision step, where 2019) proposes a DNN architecture by explicitly leveraging multi-step processing with Nsteps decision steps. The ith
expressive feature combinations, however, learning is based step inputs the processed information from the (i − 1)th step on transferring knowledge from gradient-boosted DT. (Tanno to decide which features to use and outputs the processed ing from primitive blocks while representation learning into sion. The idea of top-down attention in the sequential form edges, routing functions and leaf nodes. TabNet differs from is inspired by its applications in processing visual and text
these as it embeds soft feature selection with controllable data (Hudson and Manning 2018) and reinforcement learn- Self-supervised learning: Unsupervised representation relevant information in high dimensional input. learning improves supervised learning especially in small Feature selection: We employ a learnable mask M[i] ∈ has shown significant advances – driven by the judicious capacity of a decision step is not wasted on irrelevant
choice of the unsupervised learning objective (masked input ones, and thus the model becomes more parameter effi- prediction) and attention-based deep learning. cient. The masking is multiplicative, M[i] · f . We use an attentive transformer (see Fig. 4) to obtain the masks us- TabNet for Tabular Learning ing the processed features from the preceding step, a[i − 1]:
M[i] = sparsemax(P[i − 1] · hi (a[i − 1])). Sparsemax nor-
DTs are successful for learning from real-world tabular malization (Martins and Astudillo 2016) encourages sparsity datasets. With a specific design, conventional DNN building by mapping the Euclidean projection onto the probabilistic blocks can be used to implement DT-like output manifold, simplex, which is observed to be superior in performance and e.g. see Fig. 3). In such a design, individual feature selec- aligned with the goal of sparse feature selection for explain-
tion is key to obtain decision boundaries in hyperplane form, PD which can be generalized to a linear combination of features ability. Note that j=1 M[i]b,j = 1. hi is a trainable func- where coefficients determine the proportion of each feature. tion, shown in Fig. 4 using a FC layer, followed by BN. P[i] TabNet is based on such functionality and it outperforms DTs is the prior scale term, denoting how much a particular feature
Qi while reaping their benefits by careful design which: (i) uses has been used previously: P[i] = j=1 (γ − M[j]), where γ sparse instance-wise feature selection learned from data; (ii) is a relaxation parameter – when γ = 1, a feature is enforced
+ Softmax
Feature Feature …
transformer transformer
x Nsteps Features
+ Softmax
Feature …
transformer transformer Feature Feature Feature Feature transformer
Encoded representation
transformer transformer Attentive transformer … Mask transformer …
Step 2 Decision step dependent
transformer transformer
BN Feature Feature
FC BN transformer transformer
+ 0.5 0.5 0.5 Agg. Agg. Features Features FC FC + +
Reconstructed + … Feature attributes + … features
(a) TabNet encoder architecture (b) TabNet decoder architecture Feature transformer Feature Attentive transformer Shared across decision steps Decision step dependent transformer GLU
Decision step dependent Prior scales
+ 0.5 0.5 0.5
0.5 0.5 0.5
+ Attentive transformer (c) (d)
Prior scales
divides the processed representation to be used by the attentive transformer of the subsequent step as well as for the overall Attentive BN FC
output. For each step, the feature selection mask provides interpretable information about the model’s functionality, and the +
masks can be aggregated to obtain global feature transformer important attribution. (b) TabNet decoder, composed of a feature transformer block at each step. (c) A feature transformer block example – 4-layer network is shown, where 2 are shared across all decision
Prior scales
steps and 2 are decision step-dependent. Each layer is composed of a fully-connected (FC) layer, BN and GLU nonlinearity. (d) +
An attentive transformer block example – a single layer mapping is modulated with a prior scale information which aggregates Sparsemax
how much each feature has been used before the current decision step. sparsemax (Martins and Astudillo 2016) is used for BN FC
normalization of the coefficients, resulting in sparse selection of the salient features. +
to be used only at one decision step and as γ increases, more propose the aggregate.feature importance mask, Magg−b,j = flexibility is provided to use a feature at multiple decision PNsteps ηb [i]Mb,j [i]
PD PNsteps
ηb [i]Mb,j [i].2 i=1 i=1 steps. P is initialized as all ones, 1B×D , without any prior j=1
on the masked features. If some features are unused (as in self- Tabular self-supervised learning: We propose a decoder supervised learning), corresponding P entries are made 0 architecture to reconstruct tabular features from the Tab- to help model’s learning. To further control the sparsity of the Net encoded representations. The decoder is composed of selected features, we propose sparsity regularization in the feature transformers, followed by FC layers at each deci-
form of entropy (Grandvalet and Bengio 2004), Lsparse = sion step. The outputs are summed to obtain the recon-
PNsteps PB PD −Mb,j [i] log(Mb,j [i]+)
i=1 b=1 j=1 Nsteps ·B , where is a structed features. We propose the task of prediction of miss- small number for numerical stability. We add the sparsity reg- ing feature columns from the others. Consider a binary mask ularization to the overall loss, with a coefficient λsparse . Spar- S ∈ {0, 1}B×D . The TabNet encoder inputs (1 − S) · f̂ sity provides a favorable inductive bias for datasets where and the decoder outputs the reconstructed features, S · f̂ . We
most features are redundant. initialize P = (1 − S) in encoder so that the model em- Feature processing: We process the filtered features using phasizes merely on the known features, and the decoder’s last a feature transformer (see Fig. 4) and then split for the FC layer is multiplied with S to output the unknown features. decision step output and information for the subsequent We consider the reconstruction loss in self-supervised phase:
step, [d[i], a[i]] = fi (M[i] · f ), where d[i] ∈ <B×Nd and 2
PB PD (f̂b,j −fb,j )·Sb,j
a[i] ∈ <B×Na . For parameter-efficient and robust learning b=1 j=1
√ PB PB 2
. Normalization b=1 (fb,j −1/B b=1 fb,j ) with high capacity, a feature transformer should comprise layers that are shared across all decision steps (as the same with the population standard deviation of the ground truth features are input across different decision steps), as well as is beneficial, as the features may have different ranges. We decision step-dependent layers. Fig. 4 shows the implementa- sample Sb,j independently from a Bernoulli distribution with
tion as concatenation of two shared layers and two decision parameter ps , at each iteration. step-dependent layers. Each FC layer is followed by BN and eventually connected to a normalized residual √ connection We study TabNet in wide range of problems, that contain with normalization. Normalization with 0.5 helps to sta- regression or classification tasks, particularly with published bilize learning by ensuring that the variance throughout the benchmarks. For all datasets, categorical inputs are mapped
For faster training, we use large batch sizes with BN. Thus, bedding and numerical columns are input without and pre- except the one applied to the input features, we use ghost BN processing.4 We use standard classification (softmax cross (Hoffer, Hubara, and Soudry 2017) form, using a virtual batch entropy) and regression (mean squared error) loss functions size BV and momentum mB . For the input features, we ob- and we train until convergence. Hyperparameters of the Tab-
serve the benefit of low-variance averaging and hence avoid Net models are optimized on a validation set and listed in ghost BN. Finally, inspired by decision-tree like aggregation Appendix. TabNet performance is not very sensitive to most as in Fig. 3, we construct the overall decision embedding hyperparameters as shown with ablation studies in Appendix. as dout = i=1 PNsteps ReLU(d[i]). We apply a linear mapping In Appendix, we also present ablation studies on various de-
Wfinal dout to get the output mapping.1 sign and guidelines on selection of the key hyperparameters. Interpretability: TabNet’s feature selection masks can shed For all experiments we cite, we use the same training, val- light on the selected features at each step. If Mb,j [i] = 0, idation and testing data split with the original work. Adam optimization algorithm (Kingma and Ba 2014) and Glorot then j th feature of the bth sample should have no contribution uniform initialization are used for training of all models.5
to the decision. If fi were a linear function, the coefficient
Mb,j [i] would correspond to the feature importance of fb,j . Instance-wise feature selection
Although each decision step employs non-linear processing, their outputs are combined later in a linear way. We aim Selection of the salient features is crucial for high perfor- to quantify an aggregate feature importance in addition to mance, especially for small datasets. We consider 6 tabular requires a coefficient that can weigh the relative importance samples). The datasets are constructed in such a way that of each step in the decision. We simply propose ηb [i] = only a subset of the features determine the output. For Syn1-
PNd Syn3, salient features are same for all instances (e.g., the
c=1 ReLU(db,c [i]) to denote the aggregate decision con- tribution at ith decision step for the bth sample. Intuitively, if 2
Normalization is used to ensure D
P j=1 Magg−b,j = 1. db,c [i] < 0, then all features at ith decision step should have 3
0 contribution to the overall decision. As its value increases, prove the performance, but interpretation of individual dimensions
it plays a higher role in the overall linear combination. Scal- may become challenging. ing the decision mask at each decision step with ηb [i], we Specially-designed feature engineering, e.g. logarithmic trans- formation of variables highly-skewed distributions, may further
For discrete outputs, we additionally employ softmax during
training (and argmax during inference). An open-source implementation will be released.
Global: using only globally-salient features, Tree Ensembles (Geurts, Ernst, and Wehenkel 2006), Lasso-regularized model, L2X
Syn Syn Syn Syn Syn Syn
No selection .5 ± .0 .7 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .6 ± .0 Tree .5 ± .1 .8 ± .0 .8 ± .0 .6 ± .0 .7 ± .0 .7 ± .0 Lasso-regularized .4 ± .0 .5 ± .0 .8 ± .0 .5 ± .0 .6 ± .0 .7 ± .0
INVASE .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0
Global .6 ± .0 .8 ± .0 .9 ± .0 .7 ± .0 .7 ± .0 .8 ± .0 TabNet .6 ± .0 .8 ± .0 .8 ± .0 .7 ± .0 .7 ± .0 .8 ± .0
output of Syn depends on features X -X ), and global fea- Table 3: Performance for Poker Hand induction dataset. ture selection, as if the salient features were known, would give high performance. For Syn4-Syn6, salient features are Model Test accuracy (%) instance dependent (e.g., for Syn4, the output depends on ei- DT 50.0 ther X -X or X -X depending on the value of X ), which MLP 50.0
makes global feature selection suboptimal. Table 1 shows that Deep neural DT 65.1
TabNet outperforms others (Tree Ensembles (Geurts, Ernst, XGBoost 71.1
and Wehenkel 2006), LASSO regularization, L2X (Chen LightGBM 70.0 van der Schaar 2019). For Syn1-Syn3, TabNet performance TabNet 99.2 is close to global feature selection - it can figure out what Rule-based 100.0 features are globally important. For Syn4-Syn6, eliminating instance-wise redundant features, TabNet improves global feature selection. All other methods utilize a predictive model Poker Hand (Dua and Graff 2017): The task is classifica-
with 43k parameters, and the total number of parameters is tion of the poker hand from the raw suit and rank attributes of 101k for INVASE due to the two other models in the actor- the cards. The input-output relationship is deterministic and critic framework. TabNet is a single architecture, and its size hand-crafted rules can get 100% accuracy. Yet, conventional is 26k for Syn1-Syn and 31k for Syn4-Syn6. The compact DNNs, DTs, and even their hybrid variant of deep neural DTs
representation is one of TabNet’s valuable properties. (Yang, Morillo, and Hospedales 2018) severely suffer from the imbalanced data and cannot learn the required sorting and Performance on real-world datasets ranking operations (Yang, Morillo, and Hospedales 2018).
Tuned XGBoost, CatBoost, and LightGBM show very slight
as it can perform highly-nonlinear processing with its depth, Model Test accuracy (%) without overfitting thanks to instance-wise feature selection.
CatBoost 85.1 Table 4: Performance on Sarcos dataset. Three TabNet mod-
AutoML Tables 94.9 els of different sizes are considered.
Forest Cover Type (Dua and Graff 2017): The task is clas- MLP 2.1 0.14M
sification of forest cover type from cartographic variables. Adaptive neural tree 1.2 0.60M approaches that are known to achieve solid performance (AutoML 2019), an automated search framework based on TabNet-M 0.2 0.59M ensemble of models including DNN, gradient boosted DT, TabNet-L 0.1 1.75M with very thorough hyperparameter search. A single TabNet without fine-grained hyperparameter search outperforms it. Sarcos (Vijayakumar and Schaal 2000): The task is re-
gressing inverse dynamics of an anthropomorphic robot arm.
very small model is possible with a random forest. In the very and TabNet merely focuses on the relevant ones. For Syn4, small model size regime, TabNet’s performance is on par the output depends on either X -X or X -X depending parameters. When the model size is not constrained, TabNet feature selection – it allocates a mask to focus on the indi- achieves almost an order of magnitude lower test MSE. cator X , and assigns almost all-zero weights to irrelevant
features (the ones other than two feature groups). models are denoted with -S and -M. Real-world datasets: We first consider the simple task of mushroom edibility prediction (Dua and Graff 2017). Tab- Model Test acc. (%) Model size Net achieves 100% test accuracy on this dataset. It is indeed Sparse evolutionary MLP 78.4 81K known (Dua and Graff 2017) that “Odor” is the most discrim-
What is this project about?
This project covers practical implementation and research aspects of the topic using AI/ML techniques.