Ai Voice Bot Projects
This project focuses on ai voice bot projects using modern AI and machine learning techniques. The content below is adapted from research literature and practical implementation notes.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE 1
A Survey on Multimodal Large Language Models
Shukang Yin*, Chaoyou Fu*†, Sirui Zhao*, Ke Li,
Xing Sun, Tong Xu, and Enhong Chen, Fellow, IEEE
Abstract—Recently, Multimodal Large Language Model (MLLM) represented by GPT-4V has been a new rising research hotspot,
which uses powerful Large Language Models (LLMs) as a brain to perform multimodal tasks. The surprising emergent capabilities of MLLM, such as writing stories based on images and OCR-free math reasoning, are rare in traditional multimodal methods, suggesting a potential path to artificial general intelligence. To this end, both academia and industry have endeavored to develop MLLMs that can compete with or even better than GPT-4V, pushing the limit of research at a surprising speed. In this paper, we aim to trace and
summarize the recent progress of MLLMs. First of all, we present the basic formulation of MLLM and delineate its related concepts,
including architecture, training strategy and data, as well as evaluation. Then, we introduce research topics about how MLLMs can be extended to support more granularity, modalities, languages, and scenarios. We continue with multimodal hallucination and extended techniques, including Multimodal ICL (M-ICL), Multimodal CoT (M-CoT), and LLM-Aided Visual Reasoning (LAVR). To conclude the paper, we discuss existing challenges and point out promising research directions. In light of the fact that the era of MLLM has only just
begun, we will keep updating this survey and hope it can inspire more research. An associated GitHub link collecting the latest papers is available at https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models.
Index Terms—Multimodal Large Language Model, Vision Language Model, Large Language Model.
1 I NTRODUCTION manifests two representative traits compared with the tradi- tional counterparts: (1) MLLM is based on LLM with billion-
R ECENT years have seen the remarkable progress of
LLMs , , , , . By scaling up data size and
model size, these LLMs raise extraordinary emergent abil- scale parameters, which is not available in previous models. (2) MLLM uses new training paradigms to unleash its full ities, typically including instruction following , , In- potential, such as using multimodal instruction tuning , Context Learning (ICL) , and Chain of Thought (CoT) . to encourage the model to follow new instructions. Although LLMs have demonstrated surprising zero/few- Armed with the two traits, MLLM exhibits new capabilities,
shot reasoning performance on most Natural Language such as writing website code based on images , under- Processing (NLP) tasks, they are inherently “blind” to vision standing the deep meaning of a meme , and OCR-free since they can only understand discrete text. Concurrently, math reasoning . Large Vision Models (LVMs) can see clearly , , , Ever since the release of GPT-4 , there has been a
, but commonly lag in reasoning. research frenzy over MLLMs because of the amazing mul- In light of this complementarity, LLM and LVM run timodal examples it shows. Rapid development is fueled towards each other, leading to the new field of Multimodal by efforts from both academia and industry. Preliminary Large Language Model (MLLM). Formally, it refers to the research on MLLMs focuses on text content generation
LLM-based model with the ability to receive, reason, and grounded in text prompts and image , /video , output with multimodal information. Prior to MLLM, there /audio . Subsequent works have expanded the capa- have been a lot of works devoted to multimodality, which bilities or the usage scenarios, including: (1) Better granular- can be divided into discriminative , , and gen- ity support. Finer control on user prompts is developed to
erative , , paradigms. CLIP , as a represen- support specific regions through boxes or a certain ob- tative of the former, projects visual and textual information ject through a click . (2) Enhanced support on input and into a unified representation space, building a bridge for output modalities , , such as image, video, audio, downstream multimodal tasks. In contrast, OFA is a and point cloud. Besides input, projects like NExT-GPT
representative of the latter, which unifies multimodal tasks further support output in different modalities. (3) Improved in a sequence-to-sequence manner. MLLM can be classified language support. Efforts have been made to extend the as the latter according to the sequence operation, but it success of MLLMs to other languages (e.g. Chinese) with relatively limited training corpus , . (4) Extension • †Chaoyou Fu is the project leader.
to more realms and usage scenarios. Some studies transfer • *Shukang Yin, Chaoyou Fu, and Sirui Zhao contribute equally. the strong capabilities of MLLMs to other domains such as • Shukang Yin, Sirui Zhao, Tong Xu, and Enhong Chen are with the medical image understanding , , and document China, No.96, JinZhai Road Baohe District, Hefei, Anhui, 230026, China. • Chaoyou Fu, Ke Li, and Xing Sun are with the Tencent YouTu Lab, agents , and GUI agents , , . An MLLM
Corresponding author: Chaoyou Fu, Sirui Zhao, and Enhong Chen. In view of such rapid progress and the promising results
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE 2
Fuyu-8B V* SPHINX MoE-LLaVA Qwen-VL-Max
Available/Unavailable MobileVLM Vary Monkey TextMonkey Mobile-Agent
MMICL Xcomposer NExT-GPT
10-12 2024 Gemini AnyMAL Woodpecker
Video-LLaMA 3D-LLM GPT-4V Qwen-VL
8-9 Video-LLaVA
Kosmos-2 Lynx GPT4RoI PointLLM ASM
VisCPM LISA
Pengi Chameleon LanguageBIND
DetGPT VisionLLM Otter LLaVA-Med LLaVA-1.5
LaVIN MultiModal-GPT Shikra MotionGPT CogVLM
Emu Ferret
PaLM-E LLaMA-Adapter
4-5 GLaMM
Kosmos-1 LLaVA VideoChat
VIMA Flamingo MiniGPT-4 InstructBLIP
BLIP-2 HuggingGPT LTU EmbodiedGPT
2022 MM-REACT ViperGPT GPT4Tools mPLUG-Owl
released GitHub page, which is updated daily.
of this field, we write this survey to provide researchers to humans, modality encoders such as image/audio en- with a grasp of the basic idea, main method, and current coders are human eyes/ears that receive and pre-process progress of MLLMs. Note that we mainly focus on visual optical/acoustic signals, while LLMs are like human brains and language modalities, but also include works involving that understand and reason with the processed signals. In
other modalities like video and audio. Specifically, we cover between, the modality interface serves to align different the most important aspects of MLLMs with corresponding modalities. Some MLLMs also include a generator to output summaries and open a GitHub page that would be updated other modalities apart from text. A diagram of the architec- in real time. To the best of our knowledge, this is the first ture is plotted in Fig. 2. In this section, we introduce each
survey on MLLM. module in sequence.
The following parts of the survey are structured as
such: the survey starts with a comprehensive review of the essential aspects of MLLMs, including (1) Mainstream archi- 2.1 Modality encoder tectures (§2); (2) A full recipe of training strategy and data (§3); (3) Common practices of performance evaluation (§4). The encoders compress raw information, such as images Then, we delve into a deeper discussion on some important or audio, into a more compact representation. Rather than
topics about MLLMs, each focusing on a main problem: (1) training from scratch, a common approach is to use a pre- What aspects can be further improved or extended (§5)? trained encoder that has been aligned to other modalities. (2) How to relieve the multimodal hallucination issue (§6)? For example, CLIP incorporates a visual encoder se- The survey continues with the introduction of three key mantically aligned with the text through large-scale pre-
techniques (§7), each specialized in a specific scenario: M- training on image-text pairs. Therefore, it is easier to use ICL (§7.1) is an effective technique commonly used at the such initially pre-aligned encoders to align with LLMs inference stage to boost few-shot performance. Another im- through alignment pre-training (see §3.1). portant technique is M-CoT (§7.2), which is typically used in The series of commonly used image encoders are sum-
complex reasoning tasks. Afterward, we delineate a general marized in Table 1. Apart from vanilla CLIP image en- idea to develop LLM-based systems to solve composite coders , some works also explore using other variants. reasoning tasks or to address common user queries (§7.3). For example, MiniGPT-4 adopts an EVA-CLIP , Finally, we finish our survey with a summary and potential (ViT-G/14) encoder, which is trained with improved
research directions. training techniques. In contrast, Osprey introduces a convolution-based ConvNext-L encoder to utilize higher resolution and multi-level features. Some works also
2 A RCHITECTURE explore encoder-free architecture. For instance, the image
A typical MLLM can be abstracted into three modules, i.e. patches of Fuyu-8b are directly projected before sending a pre-trained modality encoder, a pre-trained LLM, and a to LLMs. Thus, the model naturally supports flexible image modality interface to connect them. Drawing an analogy resolution input.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE 3
Variants Pretraining Corpus Resolution Samples (B) Parameter Size (M)
OpenCLIP-ConvNext-L LAION-2B 320 29 197.4
CLIP-ViT-L/14 OpenAI’s WIT 224/336 13 304.0
EVA-CLIP-ViT-G/14 LAION-2B,COYO-700M 224 11 1000.0
OpenCLIP-ViT-G/14 LAION-2B 224 34 1012.7
OpenCLIP-ViT-bigG/14 LAION-2B 224 34 1844.9
Text training data composition are of less importance compared with input resolution, found by empirical studies .
Text Similar encoders are also available for other modali-
Audio ties. For example, Pengi uses CLAP model as
LLM the audio encoder. ImageBind-LLM uses the Image-
Video Bind encoder, which supports encoding image, text, Modality
Connector Generator
audio, depth, thermal, and Inertial Measurement Unit (IMU)
Encoder data. Equipped with the strong encoder, ImageBind-LLM
… … can respond to the input of multiple modalities.
2.2 Pre-trained LLM
Q Q K V
Q-Former K
Instead of training an LLM from scratch, it is more effi-
V V cient and practical to start with a pre-trained one. Through
tremendous pre-training on web corpus, LLMs have been embedded with rich world knowledge, and demonstrate Learnable Queries strong generalization and reasoning capabilities.
We summarize the commonly used and publicly avail-
able LLMs in Table 2. Notably, most LLMs fall in the causal includes an encoder, a connector, and a LLM. An optional decoder category, following GPT-3 . Among them, Flan- generator can be attached to the LLM to generate more
T5 series are relatively early LLMs used in works
modalities besides text. The encoder takes in images, audios like BLIP-2 and InstructBLIP . LLaMA series , or videos and outputs features, which are processed by the and Vicuna family are representative open-sourced connector so that the LLM can better understand. There are
LLMs that have attracted much academic attention. Since
broadly three types of connectors: projection-based, query- the two LLMs are predominantly pre-trained on English based, and fusion-based connectors. The former two types corpus, they are limited in multi-language support, such adopt token-level fusion, processing features into tokens to as Chinese. In contrast, Qwen is a bilingual LLM that be sent along with text tokens, while the last type enables a supports Chinese and English well. feature-level fusion inside the LLM.
It should be noted that scaling up the parameter size
of LLMs also brings additional gains, similar to the case When choosing encoders, one often considers factors find that simply scaling up LLM from 7B to 13B brings like resolution, parameter size, and pretraining corpus. comprehensive improvement on various benchmarks. Fur- Notably, many works have empirically verified that us- thermore, when using a 34B LLM, the model shows emer- ing higher resolution can achieve remarkable performance gent zero-shot Chinese capability, given that only English
input resolution can be categorized into direct scaling and a similar phenomenon by scaling up LLMs from 13B to 35B patch-division methods. The direct scaling way inputs im- and 65B/70B, where the larger model size brings consistent ages of higher resolutions to the encoder, which often gains on benchmarks specifically designed for MLLMs. involves further tuning the encoder or replacing a There are also works that use smaller LLMs to facilitate
pre-trained encoder with higher resolution . Similarly, deployment on mobile devices. For example, MobileVLM
CogAgent uses a dual-encoder mechanism, where two series , use downscaled LLaMA (termed as
encoders process high and low-resolution images, respec- MobileLLaMA 1.4B/2.7B), enabling efficient inference on tively. High-resolution features are injected into the low- mobile processors. resolution branch through cross-attention. Patch-division Recently, explorations of Mixture of Experts (MoE) archi- methods cut a high-resolution image into patches and reuse tecture for LLMs have garnered rising attention , ,
the low-resolution encoder. For example, Monkey and . Compared with dense models, the sparse architecture SPHINX divide a large image into smaller patches enables scaling up total parameter size without increasing and send sub-images together with a downsampled high- computational cost, by selective activation of the parame- resolution image to the image encoder, where the sub- ters. Empirically, MM1 and MoE-LLaVA find that
global features, respectively. In contrast, parameter size and dense counterpart on almost all the benchmarks.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE 4
Model Release Date Pretrain Data Scale Parameter Size (B) Language Support Architecture
Flan-T5-XL/XXL Oct-2022 - 3/ 11 en, fr, de Encoder-Decoder
LLaMA Feb-2023 1.4T tokens 7/ 13/ 33/ 65 en Causal Decoder
Vicuna Mar-2023 1.4T tokens 7/ 13/ 33 en Causal Decoder
LLaMA-2 Jul-2023 2T tokens 7/ 13/ 70 en Causal Decoder
Qwen Sep-2023 3T tokens 1.8 / 7/ 14/ 72 en, zh Causal Decoder
2.3 Modality interface are first embedded with visual knowledge and then con-
catenated with text features as prefixes.
Since LLMs can only perceive text, bridging the gap be-
In terms of parameter size, learnable interfaces generally
tween natural language and other modalities is necessary. comprise a small portion compared with encoders and
However, it would be costly to train a large multimodal
LLMs. Take Qwen-VL as an example, the parameter
model in an end-to-end manner. A more practical way is size of the Q-Former is about 0.08B, accounting for less to introduce a learnable connector between the pre-trained than 1% of the whole parameters, while the encoder and visual encoder and LLM. The other approach is to translate the LLM account for about 19.8% (1.9B) and 80.2% (7.7B), respectively. then send the language to LLM.
Expert Model. Apart from the learnable interface, using
Learnable Connector. It is responsible for bridging the
expert models, such as an image captioning model, is also gap between different modalities. Specifically, the module a feasible way to bridge the modality gap , , , projects information into the space that LLM can understand . The basic idea is to convert multimodal inputs into efficiently. Based on how multimodal information is fused, languages without training. In this way, LLMs can under- there are broadly two ways to implement such interfaces, stand multimodality by the converted languages. For ex-
i.e. token-level and feature-level fusion. ample, VideoChat-Text uses pre-trained vision models For token-level fusion, features output from encoders are to extract visual information such as actions and enriches transformed into tokens and concatenated with text tokens the descriptions using a speech recognition model. Though before being sent into LLMs. A common and feasible solu- using expert models is straightforward, it may not be as tion is to leverage a group of learnable query tokens to ex- flexible as adopting a learnable interface. The conversion of
tract information in a query-based manner , which first foreign modalities into text would cause information loss. has been implemented in BLIP-2 , and subsequently in- For example, transforming videos into textual descriptions herited by a variety of work , , . Such Q-Former- distorts spatial-temporal relationships . style approaches compress visual tokens into a smaller num- ber of representation vectors. In contrast, some methods simply use a MLP-based interface to bridge the modality 3 T RAINING S TRATEGY AND DATA
gap , , , . For example, LLaVA series adopts A full-fledged MLLM undergoes three stages of training, one/two linear MLP , to project visual tokens and i.e. pre-training, instruction-tuning, and alignment tuning. align the feature dimension with word embeddings. Each phase of training requires different types of data and On a related note, MM1 has ablated on design fulfills different objectives. In this section, we discuss train- choices on the connector and found that for token-level ing objectives, as well as data collection and characteristics
fusion, the type of modality adapter is far less important for each training stage. than the number of visual tokens and input resolution.
3.1 Pre-training
of token and feature-level fusion, and empirically reveal that the token-level fusion variant performs better in terms 3.1.1 Training Detail of VQA benchmarks. Regarding the performance gap, the As the first training stage, pre-training mainly aims to align authors suggest that cross-attention models might require different modalities and learn multimodal world knowl- a more complicated hyper-parameter searching process to edge. Pre-training stage generally entails large-scale text-
achieve comparable performance. paired data, e.g. caption data. Typically, the caption pairs de- As another line, feature-level fusion inserts extra mod- scribe images/audio/videos in natural language sentences. ules that enable deep interaction and fusion between text Here, we consider a common scenario where MLLMs are features and visual features. For example, Flamingo trained to align vision with text. As illustrated in Table 3,
inserts extra cross-attention layers between frozen Trans- given an image, the model is trained to predict autore- former layers of LLMs, thereby augmenting language fea- gressively the caption of the image, following a standard tures with external visual cues. Similarly, CogVLM cross-entropy loss. A common approach for pre-training plugs in a visual expert module in each Transformer layer is to keep pre-trained modules (e.g. visual encoders and
to enable dual interaction and fusion between vision and LLMs) frozen and train a learnable interface , , . language features. For better performance, the QKV weight The idea is to align different modalities without losing matrix of the introduced module is initialized from the pre-trained knowledge. Some methods , , also pre-trained LLM. Similarly, LLaMA-Adapter introduces unfreeze more modules (e.g. visual encoder) to enable more learnable prompts into Transformer layers. These prompts trainable parameters for alignment. It should be noted that
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE 5
Input: <image>
Response: {caption} Dataset Samples Date
Coarse-grained Image-Text
data. {<image>} is the placeholder for the visual tokens, and CC-12M 12.4M 2020
SBU Captions 1M 2011
{caption} is the caption for the image. Note that only the LAION-5B 5.9B Mar-2022 part marked in red is used for loss calculation. LAION-2B 2.3B Mar-2022
LAION-COCO 600M Sep-2022
COYO-700M 747M Aug-2022
the training scheme is closely related to the data quality. Fine-grained Image-Text For short and noisy caption data, a lower resolution (e.g.
ShareGPT4V-PT 1.2M Nov-2023
224) can be adopted to speed up the training process, while LVIS-Instruct4V 111K Nov-2023 for longer and cleaner data, it is better to utilize higher ALLaVA 709K Feb-2024 resolutions (e.g. 448 or higher) to mitigate hallucinations. Be- Video-Text sides, ShareGPT4V finds that with high-quality caption
MSR-VTT 200K 2016
data in the pretraining stage, unlocking the vision encode promotes better alignment. Audio-Text
WavCaps 24K Mar-2023
3.1.2 Data
Pretraining data mainly serve two purposes, i.e. (1) aligning
different modalities and (2) providing world knowledge. LAION. This series are large web-scale datasets, with im- The pretraining corpora can be divided into coarse-grained ages scrawled from the internet and associated alt-text as and fine-grained data according to granularities, which we captions. To filter the image-text pairs, the following steps will introduce sequentially. We summarize commonly used are performed: (1) Text with short lengths or images with too
pretraining datasets in Table 4. small or too big sizes are dropped. (2) Image deduplication Coarse-grained caption data share some typical traits in based on URL. (3) Extract CLIP embeddings for images common: (1) The data volume is large since samples are and text, and use the embeddings to drop possibly illegal generally sourced from the internet. (2) Because of the web- content and image-text pairs with low cosine similarity
scrawled nature, the captions are usually short and noisy between embeddings. Here we offer a brief summary of since they originate from the alt-text of the web images. some typical variants: These data can be cleaned and filtered via automatic tools, • LAION-5B : It is a research-purpose dataset of 5.85B for example, using CLIP model to filter out image- image-text pairs. The dataset is multilingual with a 2B text pairs whose similarities are lower than a pre-defined English subset.
threshold. In what follows, we introduce some representa- • LAION-COCO : It contains 600M images extracted tive coarse-grained datasets. from the English subset of LAION-5B. The captions are CC. CC-3M is a web-scale caption dataset of 3.3M synthetic, using BLIP to generate various image cap- from alt-text associated with images. The authors design a COYO-700M . It contains 747M image-text pairs, which complicated pipeline to clean data: (1) For images, those are extracted from CommonCrawl. For data filtering, the
with inappropriate content or aspect ratio are filtered. (2) authors design the following strategies: (1) For images, For text, NLP tools are used to obtain text annotations, with those with inappropriate size, content, format, or aspect samples filtered according to the designed heuristics. (3) For ratio are filtered. Moreover, the images are filtered based If text annotations do not overlap with image labels, the public datasets such as ImageNet and MS-COCO. (2) For
corresponding samples are dropped. text, only English text with satisfactory length, noun forms, CC-12M is a following work of CC-3M and contains and appropriate words are saved. Whitespace before and 12.4M image-caption pairs. Compared with the previous after the sentence will be removed, and consecutive whites- work, CC-12M relaxes and simplifies the data-collection pace characters will be replaced with a single whitespace.
pipeline, thus collecting more data. Moreover, text appearing more than 10 times (e.g. “image SBU Captions . It is a captioned photo dataset con- for”) will be dropped. (3) For image-text pairs, duplicated taining 1M image-text pairs, with images and descriptions samples are removed based on (image pHash, text) tuple. sourced from Flickr. Specifically, an initial set of images is Recently, more works , , have explored acquired by querying the Flickr website with a large number generating high-quality fine-grained data through prompt-
of query terms. The descriptions attached to the images ing strong MLLMs (e.g. GPT-4V). Compared with coarse- thus serve as captions. Then, to ensure that descriptions grained data, these data generally contain longer and more are relevant to the images, the retained images fulfill these accurate descriptions of the images, thus enabling finer- requirements: (1) Descriptions of the images are of satisfac- grained alignment between image and text modalities.
tory length, decided by observation. (2) Descriptions of the However, since this approach generally requires calling and a propositional word (e.g. “on”, “under”) that generally ume is relatively smaller. Notably, ShareGPT4V strikes a suggests spatial relationships. balance by first training a captioner with GPT-4V-generated
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE 6
(A) Pretrain–finetune (BERT, T5) like the image caption task . The output is the answer
Pretrained Finetune on Inference
LM task A on task A (C) Instruction tuning (FLAN) to the instruction conditioned on the input. The instruction • Typically requires many task-specific examples Pretrained LM
Instruction-tune on
many tasks: Inference on task A template is flexible and subject to manual designs , , • One specialized model B, C, D, … for each task
Model learns to perform Inference on
, as exemplified in Table 5. Note that the instruction many tasks via natural unseen task (B) Prompting (GPT-3) language instructions template can also be generalized to the case of multi-round
Improve performance
Pretrained via few-shot prompting or prompt engineering Inference conversations , , , .
LM on task A
Formally, a multimodal instruction sample can be de-
truth response, respectively. The MLLM predicts an answer given the instruction and the multimodal input: 100K data, then scaling up the data volume to 1.2M using the pre-trained captioner. A = f (I, M; θ) (1)
Here, A denotes the predicted answer, and θ are the pa-
3.2 Instruction-tuning rameters of the model. The training objective is typically the 3.2.1 Introduction original auto-regressive objective used to train LLMs , Instruction refers to the description of tasks. Intuitively, , , , based on which the MLLM is encouraged to
instruction tuning aims to teach models to better under- predict the next token of the response. The objective can be stand the instructions from users and fulfill the demanded expressed as: tasks. Tuning in this way, LLMs can generalize to unseen N tasks by following new instructions, thus boosting zero-shot X
L(θ) = − log p(Ri |I, R<i ; θ) (2)
performance. This simple yet effective idea has sparked the i=1 success of subsequent NLP works, such as ChatGPT , InstructGPT , FLAN , , and OPT-IML . where N is the length of the ground-truth response.
The comparisons between instruction tuning and related
typical learning paradigms are illustrated in Fig. 3. The 3.2.3 Data Collection supervised fine-tuning approach usually requires a large Since instruction data are more flexible in formats and amount of task-specific data to train a task-specific model. varied in task formulations, it is usually trickier and more The prompting approach reduces the reliance on large-scale costly to collect data samples. In this section, we summarize
data and can fulfill a specialized task via prompt engi- three typical ways to harvest instruction data at scale, i.e. neering. In such a case, though the few-shot performance data adaptation, self-instruction, and data mixture. has been improved, the zero-shot performance is still quite Data Adaptation. Task-specific datasets are rich sources of
average . Differently, instruction tuning learns how to high-quality data. Hence, abundant works , , , generalize to unseen tasks rather than fitting specific tasks , , , , have utilized existing high- like the two counterparts. Moreover, instruction tuning is quality datasets to construct instruction-formatted datasets. highly related to multi-task prompting . Take the transformation of VQA datasets for an example,
In this section, we delineate the format of instruction the original sample is an input-out pair where the input samples, the training objectives, typical ways to gather in- comprises an image and a natural language question, and struction data, and corresponding commonly used datasets. the output is the textual answer to the question conditioned
on the image. The input-output pairs of these datasets could 3.2.2 Training Detail naturally comprise the multimodal input and response of A multimodal instruction sample often includes an optional the instruction sample (see §3.2.2). The instructions, i.e. the instruction and an input-output pair. The instruction is descriptions of the tasks, can either derive from manual
typically a natural language sentence describing the task, design or from semi-automatic generation aided by GPT. such as, “Describe the image in detail.” The input can be an Specifically, some works , , , , , of them during training. We offer an example of instruction templates for the VQA datasets as shown in Table 6. The Below is an instruction that describes a task. Write a response other works manually design some seed instructions and
that appropriately completes the request use these to prompt GPT to generate more , , . Instruction: <instruction> Note that since the answers of existing VQA and caption Input: {<image>, <text>} datasets are usually concise, directly using these datasets for
Response: <output>
instruction tuning may limit the output length of MLLMs.
There are two common strategies to tackle this problem. The
instruction data. <instruction> is a textual description of the ChatBridge explicitly declares short and brief for short- task. {<image>, <text>} and <output> are input and output answer data, as well as a sentence and single sentence for from the data sample. Note that <text> in the input may be conventional coarse-grained caption data. The second one is
missed for some datasets, such as image caption datasets to extend the length of existing answers . For example, merely have <image>. The example is adapted from . M3 IT proposes to rephrase the original answer by
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE 7
• <Image> {Question} • <Image> Question: {Question} • <Image> {Question} A short answer to the question is • <Image> Q: {Question} A: • <Image> Question: {Question} Short answer: • <Image> Given the image, answer the following question with no more than three words. {Question} • <Image> Based on the image, respond to this question with a short answer: {Question}. Answer: • <Image> Use the provided image to answer the question: {Question} Provide your answer as short as possible:
• <Image> What is the answer to the following question? "{Question}" • <Image> The question "{Question}" can be answered using the image. A short answer is
in the original VQA datasets, respectively.
Video, A: Audio. For data composition, M-T and S-T denote multi-turn and single-turn, respectively.
Dataset Sample Modality Source Composition
LLaVA-Instruct 158K I+T→T MS-COCO 23K caption + 58K M-T QA + 77K reasoning
LVIS-Instruct 220K I+T→T LVIS 110K caption + 110K M-T QA
ALLaVA 1.4M I+T→T VFlan, LAION 709K caption + 709K S-T QA
Video-ChatGPT 100K V+T→T ActivityNet 7K description + 4K M-T QA
VideoChat 11K V+T → T WebVid description + summarization + creation
Clotho-Detail 3.9K A+T→T Clotho caption
prompting ChatGPT with the original question, answer, and and randomly shuffle) and sequential instruction tuning contextual information of the image (e.g. caption and OCR). (text data followed by multimodal data).
Self-Instruction. Although existing multi-task datasets can
3.2.4 Data Quality
contribute a rich source of data, they usually do not meet human needs well in real-world scenarios, such as multiple Recent research has revealed that the data quality of rounds of conversations. To tackle this issue, some works instruction-tuning samples is no less important than quan- collect samples through self-instruction , which utilizes tity. Lynx finds that models pre-trained on large-scale LLMs to generate textual instruction-following data using a but noisy image-text pairs do not perform as well as mod-
few hand-annotated samples. Specifically, some instruction- els pre-trained with smaller but cleaner datasets. Similarly, ter which ChatGPT/GPT-4 is prompted to generate more higher quality can achieve better performance. For data instruction samples with the demonstrations as guidance. filtering, the work proposes some metrics to evaluate data LLaVA extends the approach to the multimodal field quality and, correspondingly, a method to automatically
by translating images into text of captions and bound- filter out inferior vision-language data. Here we discuss two ing boxes, and prompting text-only GPT-4 to generate important aspects regarding data quality. new data with the guidance of requirements and demon- Prompt Diversity. The diversity of instructions has been strations. In this way, a multimodal instruction dataset found to be critical for model performance. Lynx em- is constructed, called LLaVA-Instruct-150k. Following this pirically verifies that diverse prompts help improve model
idea, subsequent works such as MiniGPT-4 , Chat- performance and generalization ability. Bridge , GPT4Tools , and DetGPT develop Task Coverage. In terms of tasks involved in training data, the release of the more powerful multimodal model GPT- the visual reasoning task is superior to captioning and 4V, many works have adopted GPT-4V to generate data of QA tasks for boosting model performance. Moreover, the higher quality, as exemplified by LVIS-Instruct4V and study suggests that enhancing the complexity of instruc-
ALLaVA . We summarize the popular datasets gener- tions might be more beneficial than increasing task diversity ated through self-instruction in Table 7. and incorporating fine-grained spatial annotations.
Data Mixture. Apart from the multimodal instruction
3.3 Alignment tuning
data, language-only user-assistant conversation data can also be used to improve conversational proficiencies 3.3.1 Introduction and instruction-following abilities , , , . Alignment tuning is more often used in scenarios where LaVIN directly constructs a minibatch by randomly models need to be aligned with specific human preferences, sampling from both language-only and multimodal data. e.g. response with fewer hallucinations (see §6). Currently,
MultiInstruct probes different strategies for training Reinforcement Learning with Human Feedback (RLHF) and
with a fusion of single modal and multimodal data, includ- Direct Preference Optimization (DPO) are two main tech- ing mixed instruction tuning (combine both types of data niques for alignment tuning. In this section, we introduce
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE 8
the main ideas of the two techniques in sequence and TABLE 8: A summary of datasets for alignment-tuning. For offer some examples of how they are utilized in addressing input/output modalities, I: Image, T: Text. practical problems, and finally, give a compilation of the Dataset Sample Modality Source related datasets. LLaVA-RLHF 10K I+T→T Human
RLHF-V 5.7K I+T→T Human
VLFeedback 380K I+T→T GPT-4V
3.3.2 Training Detail
RLHF , . This technique aims to utilize rein-
forcement learning algorithms to align LLMs with human response and uses the obtained data to perform dense DPO. preferences, with human annotations as supervision in the Silkie instead collects preference data via prompting training loop. As exemplified in InstructGPT , RLHF GPT-4V and distills the preference supervision into an incorporates three key steps: instruction-tuned model through DPO. 1) Supervised fine-tuning. This step aims to fine-tune a
pre-trained model to present the preliminary desired 3.3.3 Data output behavior. The fine-tuned model in the RLHF The gist of data collection for alignment-tuning is to collect setting is called a policy model. Note that this step might feedback for model responses, i.e. to decide which response be skipped since the supervised policy model π SFT can be is better. It is generally more expensive to collect such data, initialized from an instruction-tuned model (see §3.2). and the amount of data used for this phase is typically
2) Reward modeling. A reward model is trained using pref- even less than that used in previous stages. In this part, we erence pairs in this step. Given a multimodal prompt introduce some datasets and summarize them in Table 8. (e.g. image and text) x and a response pair (yw , yl ), the LLaVA-RLHF . It contains 10K preference pairs col- reward model rθ learns to give a higher reward to the lected from human feedback in terms of honesty and help-
preferred response yw , and vice versa for yl , according to fulness. The dataset mainly serves to reduce hallucinations the following objective: in model responses. L(θ) = −E(x,yw ,yl )∼D [log(σ(rθ (x, yw ) − rθ (x, yl )] (3) RLHF-V . It has 5.7K fine-grained human feedback data collected by segment-level hallucination corrections. where D = {(x, yw , yl )} is the comparison dataset VLFeedback . It utilizes AI to provide feedback on
labeled by human annotators. In practice, the reward model responses. The dataset contains more than 380K model rθ shares a similar structure with the policy model. comparison pairs scored by GPT-4V in terms of helpfulness, 3) Reinforcement learning. In this step, the Proximal Policy faithfulness, and ethical concerns.
Optimization (PPO) algorithm is adopted to optimize the
RL policy model πϕRL . A per-token KL penalty is often
added to the training objective to avoid deviating too far
4 E VALUATION
from the original policy , resulting in the objective: Evaluation is an essential part of developing MLLMs since h it provides feedback for model optimization and helps to
L(ϕ) = −Ex∼D,y∼πϕRL (y|x) rθ (x, y) compare the performance of different models. Compared
i (4) with evaluation methods of traditional multimodal mod- − β · DKL πϕRL (y|x)||π REF (y|x) els, the evaluation of MLLMs exhibits several new traits: (1) Since MLLMs are generally versatile, it is important where β is the coefficient for the KL penalty term. Typ- to evaluate MLLMs comprehensively. (2) MLLMs exhibit ically, both the RL policy πϕRL and the reference model many emergent capabilities that require special attention
π REF are initialized from the supervised model π SFT . (e.g. OCR-free math reasoning) and thus require new eval- The obtained RL policy model is expected to align with uation schemes. The evaluation of MLLMs can be broadly human preferences through this tuning process. categorized into two types according to the question genres, Researchers have explored using the RLHF techniques including closed-set and open-set. for better multimodal alignment. For example, LLaVA-
RLHF collects human preference data and tunes a
4.1 Closed-set
model with fewer hallucinations based on LLaVA . DPO . It learns from human preference labels utilizing Closed-set questions refer to a type of question where the a simple binary classification loss. Compared with the PPO- possible answer options are predefined and limited to a based RLHF algorithm, DPO is exempt from learning an finite set. The evaluation is usually performed on task- explicit reward model, thus simplifying the whole pipeline specific datasets. In this case, the responses can be naturally
to two steps, i.e. human preference data collection and judged by benchmark metrics , , , , , preference learning. The learning objective is as follows: , , . For example, InstructBLIP reports the accuracy on ScienceQA , as well as the CIDEr h πϕRL (yw |x) score on NoCaps and Flickr30K . The evalu-
L(ϕ) = −E(x,yw ,yl )∼D log σ β log REF
π (yw |x) ation settings are typically zero-shot , , , RL (5) or finetuning , , , , , , , . The πϕ (yl |x) i − β log REF first setting often selects a wide range of datasets covering π (yl |x) different general tasks and splits them into held-in and
RLHF-V collects fine-grained (segment-level) prefer- held-out datasets. After tuning on the former, zero-shot
ence data pairs by correcting hallucinations in the model performance is evaluated on the latter with unseen datasets
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE 9
or even unseen tasks. In contrast, the second setting is often exploit a more advanced GPT-4V model to assess observed in the evaluation of domain-specific tasks. For the performance of MLLMs. For example, Woodpecker example, LLaVA and LLaMA-Adapter report fine- adopts GPT-4V to judge the response quality of model tuned performance on ScienceQA . LLaVA-Med answers based on the image. The evaluation is expected to reports results on biomedical VQA , , . be more accurate than using text-only GPT-4 since GPT-4V
The above evaluation methods are usually limited to a has direct access to the image. small range of selected tasks or datasets, lacking a compre- A supplementary approach is to compare the different hensive quantitative comparison. To this end, some efforts capabilities of MLLMs through case studies. For instance, have endeavored to develop new benchmarks specially some studies evaluate two typical advanced commercial-use
sive evaluation benchmark MME that includes a total of samples across various domains and tasks, spanning from
14 perception and cognition tasks. All instruction-answer preliminary skills, such as caption and object counting, to
pairs in MME are manually designed to avoid data leakage. complex tasks that require world knowledge and reasoning, MMBench is a benchmark specifically designed for such as joke understanding and indoor navigation as an ChatGPT to match open responses with pre-defined choices. uation of GPT-4V by designing samples targeting automatic domains and propose specialized benchmarks as well as evaluation on Gemini-Pro by comparing the model against
evaluation tools for assessment. There are also evaluation GPT-4V. The results suggest that GPT-4V and Gemini exhibit strategies designed to evaluate a specific aspect of the comparable visual reasoning abilities in spite of different model , as exemplified by POPE for assessment response styles. of hallucination degree.
5 E XTENSIONS
4.2 Open-set Recent studies have made significant strides in extending
In contrast to the closed-set questions, the responses to the capabilities of MLLMs, spanning from more potent open-set questions can be more flexible, where MLLMs foundational abilities to broader coverage of scenarios. We usually play a chatbot role. Because the content of the trace the principal development of MLLMs in this regard. chat can be arbitrary, it would be trickier to judge than Granularity Support. To facilitate better interaction between
the closed-ended output. The criterion can be classified agents and users, researchers have developed MLLMs with into manual scoring, GPT scoring, and case study. Manual finer support of granularities in terms of model inputs and scoring requires humans to assess the generated responses. outputs. On the input side, models that support finer control This kind of approach often involves hand-crafted ques- from user prompts are developed progressively, evolving
tions that are designed to assess specific dimensions. For from image to region , , and even pixels , example, mPLUG-Owl collects a visually related eval- , . Specifically, Shikra supports region-level uation set to judge capabilities like natural image under- input and understanding. Users may interact with the assis- standing, diagram, and flowchart understanding. Similarly, tant more flexibly by referring to specific regions, which are
GPT4Tools builds two sets for the finetuning and zero- represented in bounding boxes of natural language forms. shot performance, respectively, and evaluates the responses Ferret takes a step further and supports more flexible in terms of thought, action, arguments, and the whole. referring by devising a hybrid representation scheme. The Since manual assessment is labor intensive, some re- model supports different forms of prompts, including point,
searchers have explored rating with GPT, namely GPT scor- box, and sketch. Similarly, Osprey supports point input ing. This approach is often used to evaluate performance by utilizing a segmentation model . Aided by the excep- on multimodal dialogue. LLaVA proposes to score the tional capabilities of the pre-trained segmentation model, responses via text-only GPT-4 in terms of different aspects, Osprey enables specifying a single entity or part of it with a
such as helpfulness and accuracy. Specifically, 30 images single click. On the output side, grounding capabilities are are sampled from the COCO validation set, each improved in line with the development of input support. associated with a short question, a detailed question, and Shikra supports response grounded in the image with a complex reasoning question via self-instruction on GPT-4. box annotations, resulting in higher precision and finer
The answers generated by both the model and GPT-4 are referring experience. LISA further supports mask- sent to GPT-4 for comparison. Subsequent works follow this level understanding and reasoning, which makes pixel-level idea and prompt ChatGPT or GPT-4 , , , grounding possible. , to rate results , , , , or judge Modality Support. Increased support for modalities is a which one is better . tendency for MLLM studies. On the one hand, researchers
A main issue of applying text-only GPT-4 as an evaluator have explored adapting MLLMs to support the input of is that the judge is only based on image-related text content, more multimodal content, such as 3D point cloud , such as captions or bounding box coordinates, without , , . On the other hand, MLLMs are also accessing the image . Thus, it may be questionable to set extended to generate responses of more modalities, such GPT-4 as the performance upper bound in this case. With as image , , , , audio , , ,
the release of the vision interface of GPT, some works , , and video , . For example, NExT-GPT
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE 10
proposes a framework that supports inputs and outputs of the image content . As a fundamental and important mixed modalities, specifically, combinations of text, image, problem, the issue has received increased attention. In this audio, and video, with the help of diffusion models , section, we briefly introduce some related concepts and attached to the MLLM. The framework applies an research development. encoder-decoder architecture and puts LLM as a pivot for
Language Support. Current models are predominantly 6.1 Preliminaries
unilingual, probably due to the fact that high-quality non- Current research on multimodal hallucinations can be fur- English training corpus is scarce. Some works have been de- ther categorized into three types : voted to developing multilingual models so that a broader 1) Existence Hallucination is the most basic form, meaning range of users can be covered. VisCPM transfers model that models incorrectly claim the existence of certain capabilities to the multilingual setting by designing a multi-
objects in the image. stage training scheme. Specifically, the scheme takes English 2) Attribute Hallucination means describing the attributes as a pivotal language, with abundant training corpus. Uti- of certain objects in a wrong way, e.g. failure to identify lizing a pre-trained bilingual LLM, the multimodal capa- a dog’s color correctly. It is typically associated with ex- bilities are transferred to Chinese by adding some trans- istence hallucination since descriptions of the attributes
lated samples during instruction tuning. Taking a similar should be grounded in objects present in the image. approach, Qwen-VL is developed from the bilingual 3) Relationship Hallucination is a more complex type and is LLM Qwen and supports both Chinese and English. also based on the existence of objects. It refers to false
During pre-training, Chinese data is mixed into the training
descriptions of relationships between objects, such as corpus to preserve the bilingual capabilities of the model, relative positions and interactions. taking up 22.7% of the whole data volume. Scenario/Task Extension. Apart from developing common In what follows, we first introduce some specific eval- general-purpose assistants, some studies have focused on uation methods (§6.2), which are useful to gauge the per- more specific scenarios where practical cond
Frequently Asked Questions
What is this project about?
This project covers practical implementation and research aspects of the topic using AI/ML techniques.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.
This section provides additional detailed analysis and supporting information derived from the research paper content to ensure comprehensive coverage of the topic with expanded discussion on key concepts, methods, and findings.