Glossary term
Large Multimodal Models
What are Large Multimodal Models?
Large Multimodal Models (LMMs), also known as Multimodal Large Language Models (MLLMs), are advanced AI systems that can process and generate information across multiple data modalities, such as text, images, audio, and video. Unlike traditional AI models that are typically limited to a single type of data, LMMs can understand and synthesize information from various sources, providing a more comprehensive understanding of complex inputs.
These models are considered a significant innovation for businesses and other applications because they enable enhanced decision-making through more accurate predictions and insights derived from diverse data types. They also streamline workflows by automating complex processes and improve customer experiences by providing personalized interactions.
LMMs are a step towards artificial general intelligence, as they exhibit emergent capabilities like writing stories based on images or performing OCR-free math reasoning. They use techniques such as Multimodal Instruction Tuning (M-IT), Multimodal In-Context Learning (M-ICL), Multimodal Chain of Thought (M-CoT), and LLM-Aided Visual Reasoning (LAVR) to achieve their functionality.
The development of LMMs is driven by the integration of additional modalities into Large Language Models (LLMs), which allows them to understand the world through multiple senses, leading to more sophisticated reasoning abilities and continuous learning capabilities.
Researchers and companies are actively exploring LMMs, with models like OpenAI's GPT-4 and DALL-E 3 being notable examples. These models are pushing the boundaries of AI by not only processing multimodal inputs but also generating multimodal outputs, which could include text, images, and potentially even animations or audio.
Current context: In 2023 and early 2024, GPT-4V and DALL-E 3 were prominent examples of multimodal AI. They illustrated systems that could process multimodal inputs or generate outputs such as text and images, while later model generations expanded the available modalities and capabilities.
In 2023 and early 2024, models such as OpenAI's GPT-4V and DALL-E 3 were prominent examples of multimodal AI. They illustrated how systems could process multimodal inputs or generate outputs such as text and images, while later model generations expanded the available modalities and capabilities.
The field of LMMs is rapidly evolving, with new ideas and techniques being developed to improve their performance and capabilities. As these models become more advanced, they are expected to play a central role in the next generation of AI applications, offering more natural and intuitive ways for humans to interact with technology.
Examples of Large Language Models
Large Multimodal Models (LMMs) are a recent development in AI that combine different types of data, such as text and images, to generate more comprehensive outputs. Here are examples of three such models:
-
GPT-4-V — This model, developed by OpenAI, allows users to upload an image as an input and ask a question about the image, a task known as visual question answering (VQA). It can analyze image inputs provided by the user and generate a response. However, it has been noted that GPT-4-V can sometimes make mistakes in its responses and may miss mathematical symbols.
-
Google Gemini — Gemini is a large language model developed by Google that works across various Google products, including search and ads. It uses a new code-generating system called AlphaCode 2 and is designed to be more efficient, being both faster and cheaper to run than previous models. Gemini comes in different versions, including Gemini Nano for running natively and offline on Android, Gemini Pro for powering Google AI services, and Gemini Ultra, the most powerful version.
-
LLaVA — LLaVA (Large Language-and-Vision Assistant) is an end-to-end trained large multimodal model that connects a vision encoder and a large language model for general-purpose applications. It is an open-source chatbot trained by fine-tuning LLaMA/Vicuna on GPT-generated multimodal instruction-following data. LLaVA achieves 85.1% relative score compared with GPT-4, indicating the effectiveness of the proposed self-instruct method in multimodal settings.
These models represent the cutting edge of multimodal AI, each with their unique strengths and applications. They are pushing the boundaries of how we use AI by integrating different types of data to generate more comprehensive and useful outputs.
Current example context: These were influential 2023 and early 2024 examples rather than a current frontier list. Gemini launched as a multimodal family in Nano, Pro, and Ultra tiers and has since advanced through newer generations. In its original paper, LLaVA v1 reported its 85.1% relative score on a synthetic multimodal instruction-following benchmark; that result is protocol-specific rather than a general percentage of GPT-4 capability.
Examples of Large Multimodal Models
Dated and protocol-qualified example context
Large Multimodal Models (LMMs) combine different types of data, such as text and images, to generate more comprehensive outputs. The following are three influential examples from 2023 and early 2024:
Google Gemini: Gemini is a family of multimodal models developed by Google that works across various Google products, including Search and Ads. At its initial December 2023 launch, Gemini shipped in three tiers — Nano for running natively and offline on Android, Pro for powering Google AI services, and Ultra, the most capable version at the time — alongside a new code-generating system called AlphaCode 2. Google has continued to release newer Gemini generations since then under updated naming.
LLaVA: LLaVA (Large Language-and-Vision Assistant) is an end-to-end trained large multimodal model that connects a vision encoder and a large language model for general-purpose applications. It is an open-source chatbot trained by fine-tuning LLaMA/Vicuna on GPT-generated multimodal instruction-following data. In its original paper, LLaVA (v1) reported an 85.1% relative score compared with GPT-4 on a synthetic multimodal instruction-following benchmark, indicating the effectiveness of the proposed self-instruct method in multimodal settings.
-
Google Gemini — Gemini is a family of multimodal models developed by Google that works across various Google products, including Search and Ads. At its initial December 2023 launch, Gemini shipped in three tiers — Nano for running natively and offline on Android, Pro for powering Google AI services, and Ultra, the most capable version at the time — alongside a new code-generating system called AlphaCode 2. Google has continued to release newer Gemini generations since then under updated naming.
-
LLaVA — LLaVA (Large Language-and-Vision Assistant) is an end-to-end trained large multimodal model that connects a vision encoder and a large language model for general-purpose applications. It is an open-source chatbot trained by fine-tuning LLaMA/Vicuna on GPT-generated multimodal instruction-following data. In its original paper, LLaVA (v1) reported an 85.1% relative score compared with GPT-4 on a synthetic multimodal instruction-following benchmark, indicating the effectiveness of the proposed self-instruct method in multimodal settings.
These models were important early examples of multimodal AI, each with distinct strengths and applications. Their releases demonstrated how integrating different data types could produce more comprehensive and useful outputs.
What are the differences between GPT-4V and Gemini?
GPT-4-V and Gemini are both large multimodal models, but they have distinct characteristics and capabilities.
| Feature | GPT-4-V | Gemini |
|---|---|---|
| Data Processing | Handles text and images | Extends to audio and video for a comprehensive multimodal experience |
| Response Style | Precision and succinctness in responses | Detailed, expansive answers with imagery and links |
| Performance | Varies by benchmark; excels in text benchmarks | Outperforms GPT-4-V in HumanEval and Natural2Code benchmarks |
| Speed | Incredibly quick responses | Can slow down due to beta deployment |
| Integration and Extensions | Wide array of third-party plugins and extensions | Integrated with Google Bard and tailored versions for different platforms |
| Fact-checking | Provides source links at the end of responses | Offers a button to perform a Google search for information |
| Availability | Available through ChatGPT and API | Gemini Pro is freely available through Bard |
These differences highlight the unique strengths of each model and their suitability for different tasks and applications.
How did GPT-4V and Gemini 1.0 compare at launch?
GPT-4V and Gemini 1.0 were both important launch-era multimodal systems, but model capability and product-interface behavior were separate concerns. Google's published code evaluation also used GPT-4 as its comparator, not GPT-4V.
Historical launch-era matrix (archived January 5, 2024)
The original page archived a seven-dimension comparison on January 5, 2024. It did not cite row-level sources or define a common test protocol, prompt set, latency sample, product version, or exact model snapshot. The matrix below restores every dimension while separating model capabilities from ChatGPT, Bard, API, plugin, and search-interface behavior. Unsupported subjective claims from the original matrix are identified rather than repeated as facts.
| Dimension | GPT-4V launch-era scope | Gemini 1.0 launch-era scope | Limitation |
|---|---|---|---|
| Data processing | Accepted text and image input and generated text | The model family was developed across text, image, audio, and video; supported inputs varied by tier and product surface | Training modalities, model capability, and modalities exposed in a product are different claims |
| Response style | The archived page described responses as precise and succinct, but provided no evaluation | The archived page described responses as detailed and expansive, but provided no evaluation | Prompting and product UI affect style; no model-level comparison is verified |
| Performance | No GPT-4V code score was reported in the cited Gemini technical report | Gemini Ultra's HumanEval and Natural2Code results are preserved in the separate table below | The report compared Gemini Ultra with GPT-4, so it does not establish that Gemini outperformed GPT-4V |
| Speed | No comparable latency measurement was archived | No comparable latency measurement was archived | The original quick-versus-slow claim lacked hardware, region, load, sampling, and product-version controls |
| Integration and extensions | ChatGPT plugins and third-party integrations were product-platform features, not capabilities of GPT-4V | Gemini models were integrated into Google products, including the launch-era Bard experience | Integration breadth describes surrounding products and distribution, not the underlying models |
| Fact-checking | Source links were interface or application behavior rather than an inherent GPT-4V fact-checking system | Bard's Google-search affordance was a product-interface feature rather than an inherent Gemini fact-checking capability | Neither UI feature proves factual accuracy |
| Availability | Access depended on the ChatGPT or API product, account, and rollout | Gemini Nano, Pro, and Ultra had different launch-era deployment targets; Gemini Pro was surfaced through Bard | Availability is time-sensitive product policy, not a stable capability comparison |
This matrix preserves launch-era distinctions across modalities, style, performance, speed, integrations, fact-checking interfaces, and availability. It is historical qualitative evidence, not a current feature checklist or a controlled benchmark, and it cannot be combined with the quantitative code scores below.
Separate technical-report code scores
In the Gemini 1.0 technical report, Gemini Ultra and GPT-4 received the following code-generation scores:
| Benchmark | Gemini Ultra | GPT-4 |
|---|---|---|
| HumanEval | 74.4% | 67.0% |
| Natural2Code | 74.9% | 73.9% |
These figures reflect the report's December 2023 evaluation setup rather than current model performance, and they should not be presented as a direct benchmark comparison between Gemini Ultra and GPT-4V.
More terms
Continue exploring the glossary.
Glossary term
Unsupervised Learning
It's time to build
Collaborate with your team on reliable Generative AI features.
Want expert guidance? Book a 1:1 onboarding session from your dashboard.