February 15, 2024

Google Gemini Pro 1.5

Stephen M. Walker II · Co-Founder / CEO

Google Gemini Pro 1.5

Google's Gemini Pro 1.5, along with its specialized variant Gemini 1.5 Flash, was announced in February 2024, with Gemini 1.5 Flash following in May 2024. It succeeded the original Gemini 1.0 family and was later succeeded itself by newer Gemini generations. This page describes Gemini 1.5 as it was released; treat pricing, availability, and benchmark details below as historical rather than current.

Key Advancements and Features

Gemini Pro 1.5 introduced advancements across model architecture, multimodal processing, and long-context handling relative to Gemini 1.0, including a substantially larger context window and broader multimodal input support.

Long Context Performance

To further assess Gemini 1.5 Flash's capabilities in handling long contexts, a "needle in a haystack" evaluation was conducted. This test involves inserting a small, crucial piece of information (the "needle") within a large amount of irrelevant text (the "haystack") and evaluating the model's ability to locate and utilize this information accurately.

To evaluate Gemini 1.5 Flash's long-context capabilities, we created a specialized needle-haystack dataset. We used a recently published book, "Nuclear War" by Annie Jacobsen, replacing all proper nouns with science fiction alternatives generated by an LLM. The haystack consisted of random 10kb text chunks, with critical 4-sentence "needles" inserted at various points.

Gemini 1.5 Flash Needle in Haystack Evaluation (1M tokens)

The Gemini 1.5 Flash model's performance with a 1M context window shows significant variability across different context lengths and depths. High scores (10) are frequently achieved at various depths, especially at shorter context lengths (20,000 to 265,000) and again at the longest length (1,000,000). However, there are notable performance drops, particularly at certain depths within specific context lengths. After several attempts we witnessed a range of both low and high scores, making it difficult to draw a complete conclusion. Next, we look at the first 20k context.

Gemini Needle in a Haystack Evaluation for 20k Tokens

The Gemini 1.5 Flash model exhibits unexpected performance inconsistencies when handling relatively short context lengths between 0 and 20,000 tokens. This observation is particularly noteworthy as shorter contexts are typically easier to manage, yet the model's accuracy fluctuates considerably across different depths within this range. These issues could potentially impact the model's reliability in tasks involving shorter documents or information snippets, warranting further investigation into its behavior with brief inputs.

10 Question Q&A

Our tests revealed that while the model performed well in many scenarios, it exhibited significant performance drops at certain context lengths and depths.

Notably, we observed catastrophic forgetfulness when dealing with 120k context length and multiple inserted facts, highlighting potential limitations in the model's long-context processing abilities.

Enhanced Model Architecture and Performance

Gemini Pro 1.5 uses a Mixture-of-Experts (MoE) transformer architecture, a change from the dense architecture of Gemini 1.0. According to Google's technical report, this contributed to gains over Gemini 1.0 on a range of benchmarks covering language understanding, reasoning, coding, and multimodal tasks.

Breakthrough in Long-Context Understanding

One of the most notable advancements is the expansion of the context window to 1 million tokens, with successful tests up to 10 million tokens. This enables the model to process, analyze, and summarize vast amounts of content within a given prompt, including:

  • Up to 1 hour of video
  • 11 hours of audio
  • Codebases with over 30,000 lines of code
  • Over 700,000 words of text

While these advancements are promising, we maintain a cautious perspective on Google Gemini 1.5's real-world deployment and performance until further independent verification and widespread adoption provide more conclusive evidence of its capabilities.

Advanced Multimodal Capabilities

Gemini Pro 1.5 excels in seamlessly analyzing and reasoning across different modalities, including text, video, audio, and code. It can effectively reason about conversations, events, and details found in extensive documents or analyze complex multimedia content.

Ethical AI and Multilingual Improvements

Enhanced safeguards against biases and improved alignment with human values have been implemented. Additionally, the model now offers expanded support for low-resource languages and improved translation quality across language pairs.

Gemini 1.5 Flash: Speed and Efficiency

Alongside the main release, Google introduced Gemini 1.5 Flash, a specialized variant designed for:

  • Ultra-fast inference, reducing latency in real-time applications
  • Optimized performance on less powerful hardware, suitable for edge computing and mobile devices
  • Seamless API integration with existing software ecosystems
  • More flexible fine-tuning options for specific use cases

Availability and Pricing at Release

At release, Gemini Pro 1.5 became available to developers and enterprise customers through AI Studio and Vertex AI, with pricing tiers starting at a 128,000-token context window and scaling up to 1 million tokens. The figures below reflect pricing at the time of Gemini 1.5's release and do not reflect current Google pricing.

Gemini 1.5 Pro Pricing

At general availability, Gemini 1.5 Pro used tiered pricing based on context window size, from 128,000 to 1 million tokens. Input costs ranged from $3.50 to $7.00 per million tokens, while output cost $10.50 to $21.00 per million tokens. Context caching incurred additional fees. The free tier was limited to 2 requests per minute, 32,000 tokens per minute, and 50 requests per day. The paid tier allowed 360 requests per minute, 2 million tokens per minute, and 10,000 requests per day. Regional restrictions applied to free tier usage in the EEA, UK, and Switzerland.

FeatureFree of chargePay-as-you-go (prices in USD)
Rate Limits2 RPM (requests per minute)
32,000 TPM (tokens per minute)
50 RPD (requests per day)
360 RPM (requests per minute)
2 million TPM (tokens per minute)
10,000 RPD (requests per day)
Price (input)Free of charge$3.50 / 1 million tokens (for prompts up to 128K tokens)
$7.00 / 1 million tokens (for prompts longer than 128K)
Context cachingNot applicable$0.875 / 1 million tokens (for prompts up to 128K tokens)
$1.75 / 1 million tokens (for prompts longer than 128K)
$4.50 / 1 million tokens per hour (storage)
Price (output)Free of charge$10.50 / 1 million tokens (for prompts up to 128K tokens)
$21.00 / 1 million tokens (for prompts longer than 128K)
Prompts/responses used to improve our productsYesNo

Gemini 1.5 Flash Pricing

At general availability, Gemini 1.5 Flash's free tier offered 15 requests per minute (RPM), 1 million tokens per minute (TPM), and 1,500 requests per day (RPD). Pay-as-you-go users had increased limits of 1000 RPM and 2 million TPM. Input pricing started at $0.35 per million tokens for prompts up to 128K, increasing to $0.70 for longer prompts. Output cost $1.05 per million tokens (up to 128K) and $2.10 for longer outputs. Context caching was available at $0.0875 per million tokens (up to 128K) and $1.00 per million tokens per hour for storage.

FeatureFree of chargePay-as-you-go (prices in USD)
Rate Limits15 RPM (requests per minute)
1 million TPM (tokens per minute)
1,500 RPD (requests per day)
1000 RPM (requests per minute)
2 million TPM (tokens per minute)
Price (input)Free of charge$0.35 / 1 million tokens (for prompts up to 128K tokens)
$0.70 / 1 million tokens (for prompts longer than 128K)
Context cachingNot applicable$0.0875 / 1 million tokens (for prompts up to 128K tokens)
$0.175 / 1 million tokens (for prompts longer than 128K)
$1.00 / 1 million tokens per hour (storage)
Price (output)Free of charge$1.05 / 1 million tokens (for prompts up to 128K tokens)
$2.10 / 1 million tokens (for prompts longer than 128K)
Prompts/responses used to improve our productsYesNo

Conclusion

Gemini Pro 1.5 and Gemini 1.5 Flash extended Gemini 1.0 with a context window of up to 1 million tokens (tested by Google to 10 million) and multimodal processing across text, audio, video, and code, with Gemini 1.5 Flash tuned for lower latency and cost. These capabilities supported use cases such as large-scale document analysis, codebase review, and multimedia content interpretation, and informed the design of subsequent Gemini model generations.

More terms

Continue exploring the glossary.

Learn how teams define, measure, and improve LLM systems.

Glossary term

What is linear regression?

Linear regression is a statistical model used to estimate the relationship between a dependent variable and one or more independent variables. It's a fundamental tool in statistics, data science, and machine learning for predictive analysis.
Read term

Glossary term

What is the Google 'No Moat' Memo?

The "no moat" memo is a leaked document from a Google researcher, which suggests that Google and OpenAI lack a competitive edge or "moat" in the AI industry. The memo argues that open-source AI models are outperforming these tech giants, being faster, more customizable, more private, and more capable overall.
Read term

It's time to build

Collaborate with your team on reliable Generative AI features.
Want expert guidance? Book a 1:1 onboarding session from your dashboard.

Talk to sales