What is Google Gemini?

Stephen M. Walker II · Co-Founder / CEO

What is Google Gemini?

Klu Google Gemini Brand

Google Gemini is a family of multimodal AI models developed by Google DeepMind, first announced in December 2023. It was trained on text, code, images, audio, and video, making it "natively multimodal" from the outset rather than combining separate single-modality models. Some of the key features and capabilities of the original Gemini release include:

  • Multimodal Reasoning — Gemini is designed to reason seamlessly across text, images, video, audio, and code, making it a versatile tool for various applications.
  • Massive Multitask Language Understanding (MMLU) — At launch, Gemini Ultra was the first model reported to outperform human experts on MMLU, a widely used benchmark for testing AI models.
  • Code Generation — Gemini can generate code based on different inputs, showcasing its potential in various software applications.
  • Visual Reasoning — The model can reason visually across languages and recognize images quickly, even for connect-the-dots pictures.
Klu Google Gemini Brand

Google has continued to develop and expand Gemini well beyond this initial release — see Gemini's later versions below — integrating it into products such as Search, Ads, and the Chrome browser.

Google Gemini Model Details

Google Gemini, developed by Google, is a multimodal generative AI model. It's capable of understanding text, images, videos, and audio, making it Google's most versatile AI model to date.

It comes in three variants: Gemini Ultra, Gemini Pro, and Gemini Nano, each optimized for specific tasks and platforms.

Klu Google Gemini Sizes
  • Gemini Ultra was designed for highly complex tasks and launched to developers and enterprise customers in February 2024, following its initial testing phase. At launch, Google reported it exceeded "current state-of-the-art results on 30 of the 32 widely-used academic benchmarks used in large language model (LLM) research and development".
  • Gemini Pro was designed for a wide range of tasks and was integrated across various Google products and platforms.
  • Gemini Nano was designed for efficient on-device tasks and was made available on Google products like the Pixel 8 phone and Bard chatbot.

At launch (December 2023), Google reported Gemini Ultra surpassing human experts in Massive Multitask Language Understanding (MMLU) with a score of 90.0%, while scoring 87.8% on the HellaSwag commonsense reasoning benchmark, which was behind GPT-4 at the time. These figures reflect the original Gemini Ultra release and have since been superseded by newer models from both Google and OpenAI.

Training Data

The specific details about the training data used for the original Gemini models were not disclosed by Google.

Evaluation Data

Gemini's performance was evaluated using a variety of benchmarks. At launch, Google reported that Gemini Ultra outperformed GPT-4 on 30 of 32 benchmarks, while scoring 87.8% on the HellaSwag commonsense reasoning test, behind GPT-4 at the time.

What are Google Gemini's key features?

Google Gemini is a cutting-edge AI model with several key features that make it stand out from other large language models (LLMs). Some of its most notable features include:

  • Multimodal capabilities — Gemini is designed to seamlessly reason across text, images, video, audio, and code, making it a versatile and powerful AI system.
  • Massive Multitask Language Understanding (MMLU) — Gemini is the first model to surpass human experts in MMLU, demonstrating its extensive knowledge and problem-solving capabilities.
  • Understanding and generating high-quality code — Gemini is particularly adept at understanding and generating code in various programming languages, making it a valuable tool for developers and software engineers.
  • Integration with Google products — Gemini is now available on Google products in its Nano and Pro sizes, such as the Pixel 8 phone and Bard chat assistant, with plans to integrate it into Google's Search, Ads, Chrome, and other services.
  • Gemini Ultra — The most capable version at launch, reported to exceed "current state-of-the-art results on 30 of the 32 widely-used academic benchmarks used in the field".

Overall, Google Gemini is a highly advanced AI model with diverse capabilities that span across various domains, making it a significant leap forward in artificial intelligence and a potential challenge to other LLMs like OpenAI's ChatGPT.

How does Google Gemini work?

Google Gemini is a multimodal large language model (LLM) that can understand and process not just text but also images, videos, and audio. It was developed by Google DeepMind and serves as the successor to LaMDA and PaLM 2. Gemini is designed to be flexible and capable of running on various platforms, from Google's Tensor Processing Units (TPUs) to other computing systems.

At launch, Gemini was accessible through integrations into Google Bard and the Google Pixel 8, expanding to Google Vertex developers on December 13, 2023.

Klu Google Gemini Ultra vs GPT-4V

Gemini's architecture is built on top of Transformer decoders, with separate text and vision encoders. It has been trained on a wide range of data, including text, code, audio, image, and video, making it a versatile model capable of completing complex tasks in various domains, such as math and physics. Google claims that Gemini largely outperforms OpenAI's GPT-4 model in most benchmark tests.

Some of the key features and capabilities of Google Gemini include:

  1. Multimodal Inputs — Gemini can take textual input and a wide variety of audio and visual inputs, such as natural language text, images, audio, video, 3D models, and graphs.
  2. Versatility — Gemini is designed to be adaptable and efficient, capable of running on different platforms and handling various tasks.
  3. Improved Code Generation — Google claims that Gemini's AlphaCode 2 system performs better than 85% up from 50% for the original AlphaCode, making it a significant improvement in code generation capabilities.

Following its initial release, Google continued to fine-tune Gemini and expand its safety testing, integrating it into services such as Search, Ads, and Bard (later rebranded Gemini).

What are the benefits of using Google Gemini?

Google Gemini is a powerful and versatile AI model that offers numerous benefits across various applications and industries. Some of the key advantages of using Google Gemini include:

  • Multimodal capabilities — Gemini is trained on Tensor Processing Units (TPUs) across image, audio, video, and text data, making it highly capable of understanding and reasoning in multi-modal tasks for different domains.

  • Strong generalist capabilities — Gemini can process a wide variety of inputs, such as natural images, charts, screenshots, PDFs, and videos, and produce text and image outputs, making it excellent at understanding and reasoning in various tasks.

  • Efficiency — Gemini is designed with cost- and latency-optimization in mind, making it a more efficient model for deploying at scale.

  • Safety and responsible deployment — Google has employed best-in-class adversarial testing techniques to identify safety issues and has built dedicated safety classifiers to help its model with toxicity and other potential problems.

  • Integration with Google products — Gemini is poised to be integrated into various Google products, including the search engine, ad products, and the Chrome browser, marking the beginning of a new era in AI development.

  • Enhanced user experiences — Gemini promises advanced reasoning, planning, and understanding capabilities, particularly in Google's products like the Bard chatbot and Search Generative Experience.

  • Availability for developers — Google Cloud will make Gemini Ultra available in an early access program for developers, rolling out more broadly in early 2024. Gemini Pro will be available starting December 13 in Google Cloud's Vertex AI and AI Studio, while a version called Gemini Nano for on-device applications will be available on Google Pixel phones.

Klu Google Gemini Sizes

Overall, Google Gemini is a significant advancement in AI technology, with the potential to revolutionize various industries and applications, from search and chatbots to enterprise solutions and on-device tasks.

What are the limitations of Google Gemini?

Google Gemini is a multimodal AI model that showcases impressive capabilities in various modalities, but it also has some limitations. Some of the key limitations of Google Gemini include:

  • English-only interactions — The current version of Gemini Pro is available only in English, which hinders its global accessibility.

  • Integration within Bard — The integration of Gemini Pro within Google's Bard chatbot is limited, with future enhancements expected from Google.

  • Inconsistencies in factual accuracy — Gemini Pro has faced criticism for inconsistencies in factual accuracy and translation errors.

  • Coding limitations — Developers have reported limitations in coding when using Gemini Pro.

  • Multimodal capabilities — Despite being touted as "natively multimodal," Gemini's multimodal capabilities are not yet fully available, and it may struggle to handle visual information as effectively as it claims to.

  • Comparison to other AI models — While Gemini may outperform some AI models, such as GPT-4, it has not disclosed full details of its architecture, training data, or size, making it difficult to determine its true capabilities.

  • Geographical constraints — The availability of Gemini Pro is restricted in certain regions, such as the European Union.

These limitations reflect the initial Gemini Pro and Ultra releases; subsequent Gemini versions have since addressed many of these concerns (see below).

What was the controversy around Google Gemini's launch demo?

Klu Google Gemini MML Benchmark

At its December 2023 launch, Google published a demonstration video showcasing Gemini's multimodal capabilities. The video drew criticism after Google confirmed it was edited for presentation rather than showing a live, real-time interaction. Key points about the episode:

  • The video was produced to illustrate Gemini's capabilities, but Google later clarified that the AI did not respond in real time to voice or video prompts as the editing implied.

  • Critics argued the video overstated how the demonstrated capabilities compared to existing models like GPT-4, and that the AI did not autonomously devise the game featured in the video.

  • The controversy did not affect the underlying launch of Gemini Nano, Gemini Pro, and Gemini Ultra, with Google continuing the rollout as planned; Gemini Ultra was reported to outperform GPT-4 on MMLU at the time.

The episode became a widely cited example of the gap between AI demo videos and live product behavior, and Google has since been more explicit in labeling demonstration content.

What came after the original Gemini release?

The details above describe Gemini as it launched in December 2023. Google has continued to release newer Gemini generations since, each expanding on context length, multimodal reasoning, and availability:

  • Gemini 1.5 introduced a much longer context window and improved reasoning over the original 1.0 generation.
  • Gemini 2.0 and Gemini 2.5 added further gains in reasoning, coding, and agentic capabilities, and became the default models across Google's consumer and developer products.
  • Later Gemini 3.x releases continued this progression with additional improvements to reasoning and multimodal performance.

For current model details, benchmarks, and availability, see Google's official Gemini documentation rather than treating the figures above as current.

More terms

Continue exploring the glossary.

Learn how teams define, measure, and improve LLM systems.

Glossary term

What is Deep Reinforcement Learning?

Deep Reinforcement Learning combines neural networks with a reinforcement learning architecture that enables software-defined agents to learn the best actions possible in virtual environment scenarios to maximize the notion of cumulative reward. It has driven advancements in AI, notably powering AlphaGo's 2016 victory over Lee Sedol, and continues to underpin autonomous vehicles and sophisticated recommendation systems.
Read term

Glossary term

Transformer Architecture

A Transformer is a type of deep learning model that was first proposed in 2017. It's a neural network that learns context and meaning by tracking relationships in sequential data, such as words in a sentence or frames in a video. The Transformer model is particularly notable for its use of an attention mechanism, which allows it to focus on different parts of the input sequence when making predictions.
Read term

It's time to build

Collaborate with your team on reliable Generative AI features.
Want expert guidance? Book a 1:1 onboarding session from your dashboard.

Talk to sales