May 13, 2024

GPT-4 Getting Worse

Stephen M. Walker II · Co-Founder / CEO

Top tip

Automated Evaluations are now available, enabling comprehensive assessment of GPT-4's performance for your specific use cases.

Is GPT-4 Getting Worse?

The question of whether GPT-4 is getting worse is complex and multifaceted.

But, yes, it is.

Historical Klu comparison (archived July 24, 2024)

The following table preserves the benchmark, retrieval, and composite-score results reported by this page's historical Klu evaluation.

MODELRELEASEBENCHMARKSRETRIEVALCOMPSCORE
GPT-403143.503.943.59
GPT-4 32k03143.503.943.58
GPT-406133.503.903.58
GPT-4 32k06133.503.903.56
GPT-4 Turbo2024-04-093.503.983.48
GPT-4o2024-05-133.423.963.44
GPT-3.5 Turbo03013.023.893.42
GPT-3.5 Turbo06133.013.883.33

GPT-4's speed, vision capabilities, and ability to follow complex ideas are improving. However, when examining raw capabilities across all known benchmarks, the original GPT-4 release maintains its position.

Provenance and aggregate-score caveat

The repository archived this comparison on July 24, 2024. The original page provided no source, dataset, benchmark definitions, prompts, evaluator, sample size, harness, aggregation formula, or complete model-snapshot provenance for these values. The meanings and scales of BENCHMARKS, RETRIEVAL, and COMPSCORE therefore cannot be independently verified. The table is retained as evidence of what this page previously reported, not as current or validated performance data.

The historical interpretation above contrasted improvements in speed, vision, and instruction following with a claimed decline in the composite score. The general point that model generations can trade off capabilities remains useful, but this unsupported table cannot establish that the March 2023 GPT-4 release was universally better or that later snapshots regressed. Available evidence shows behavior drift between model snapshots on some tasks, but it does not establish that GPT-4 became uniformly worse.

While some users have reported perceived declines in performance, particularly in specific tasks or contexts, it's important to consider several factors.

Model updates, changes in training data, and evolving user expectations can all influence perceptions of performance.

Additionally, OpenAI continuously works on improving and fine-tuning their models, as evidenced by the release of GPT-4 Omni (GPT-4o) on May 13, 2024.

This latest model showcases advanced multimodal capabilities, processing and generating text, audio, image, and video inputs and outputs, which may address some of the concerns raised by users.

Current multimodal correction

At its May 2024 launch, GPT-4o accepted text, audio, image, and video inputs and generated text, audio, and image outputs. It did not generate video output.

Therefore, while there may be instances where GPT-4's performance appears to fluctuate, ongoing advancements and updates aim to enhance its overall capabilities and user experience.

Reported Issues

Many users have reported experiencing issues with GPT-4 that weren't present in earlier versions:

  • Decreased reasoning capabilities and logical errors
  • Difficulty maintaining context and following instructions
  • Reduced accuracy in specialized tasks like coding
  • Less nuanced understanding of prompts

Quantitative Study

A study by researchers from Stanford and UC Berkeley has attempted to quantify these changes:

These changes are likely due to two reasons: performance improvements via model pruning, quantization and other techniques, and additional RLHF to handle edge cases, safety issues, and minimize extra token output.

Study interpretation caveat

The paper itself attributed much of the direct-execution drop to June responses adding Markdown code fences and other non-code text. After the authors removed non-code text, GPT-4's code correctness increased from 52% in March to 70% in June. The study therefore documented behavior and instruction-following drift on specific tasks, not a uniform loss of underlying capability.

Possible Explanations

Several factors might contribute to these perceived and measured changes:

  1. Increased user base straining the system
  2. Modifications to improve speed at the cost of accuracy
  3. Changes in content moderation and safety measures
  4. Ongoing model updates and fine-tuning

Differing Opinions

It's important to note that experiences vary, and not all users report a decline in quality. Some find GPT-4 to be faster and more human-like in its responses, even if less accurate in certain areas.

Lack of Official Communication

OpenAI has not provided detailed explanations for these changes, leading to speculation and frustration among users. This lack of transparency has made it difficult for users to understand the reasons behind the perceived changes in performance.

Implications

The reported decline in quality has several implications:

  • Users may need to verify GPT-4's outputs more carefully
  • Businesses relying on GPT-4 APIs may need to reassess their strategies
  • There's an increased interest in open-source alternatives

While GPT-4 remains a powerful tool, users should be aware of its limitations and potential inconsistencies. As with any AI technology, it's crucial to approach its outputs critically and verify important information from authoritative sources.

More terms

Continue exploring the glossary.

Learn how teams define, measure, and improve LLM systems.

Glossary term

What is natural language processing (NLP)?

Natural language processing (NLP) is a subfield of artificial intelligence (AI) that deals with the interaction between computers and human (natural) languages.
Read term

Glossary term

What is the situated approach in AI?

The situated approach in AI refers to the development of agents that are designed to operate effectively within their environment. This approach emphasizes the importance of creating AI systems "from the bottom-up," focusing on basic perceptual and motor skills necessary for an agent to function and survive in its environment. It de-emphasizes abstract reasoning and problem-solving skills that are not directly tied to interaction with the environment.
Read term

It's time to build

Collaborate with your team on reliable Generative AI features.
Want expert guidance? Book a 1:1 onboarding session from your dashboard.

Talk to sales