Glossary term
What is the Nvidia H100?
What is the Nvidia H100?
The NVIDIA H100 is a Tensor Core GPU designed for data center and cloud-based AI and high-performance computing (HPC) workloads. It is built on the NVIDIA Hopper architecture, which introduced several innovations over the prior A100 generation, including a dedicated Transformer Engine for accelerating large language models.
Key features of the NVIDIA H100 include:
- Architecture — The H100 is built on a 4 nm process and features 14,592 CUDA cores, 456 tensor cores, and 80 GB of HBM2e memory.
- Performance — The GPU operates at a frequency of 1,095 MHz, which can be boosted up to 1,755 MHz, with memory running at 1,593 MHz.
- Connectivity — The H100 supports PCIe Gen 5 and NVLink for high-bandwidth, low-latency communication with other GPUs and devices.
- Multi-Instance GPU (MIG) — The H100 can be partitioned into up to seven right-sized instances, allowing multiple workloads to share a single GPU.
- AI Acceleration — A fourth-generation Transformer Engine and NVIDIA AI Enterprise software optimize the development and deployment of accelerated AI workflows.
When it launched in 2022, NVIDIA reported that the H100 delivered up to 5x faster AI training and 30x faster AI inference on large language models compared to the A100. At launch, list pricing for the H100 was reported at roughly $30,000 per GPU, though actual market prices varied with supply and demand. The H100 has since been followed by newer NVIDIA architectures, including the H200 and Blackwell-generation GPUs, which supersede it for new large-scale deployments; the H100 remains widely deployed in existing data centers.
Key Features and Specifications
The H100 GPU has a maximum thermal design power (TDP) of up to 700W, configurable down to 300–350W depending on the form factor. It supports up to seven MIG partitions of 10GB each, and its fourth-generation Tensor Cores handle matrix computations faster and more efficiently than the prior generation.
With Hopper Confidential Computing, the H100's MIG partitions can also secure sensitive applications running on shared data center infrastructure.
Performance
At its release, the H100 set records across the MLPerf training benchmarks of the time, including tests for large language models, recommenders, computer vision, medical imaging, and speech recognition.
Optimizations across the full technology stack also enabled near-linear performance scaling on the demanding LLM test as submissions scaled from hundreds to thousands of H100 GPUs.
Use Cases
The H100 GPU is suitable for a wide range of use cases. It is ideal for applications that require high-performance computing, such as complex AI models and scientific research. It is also a perfect match for PCIe expansions and GPU servers.
The H100 GPU is particularly effective for generative AI and large language models (LLMs). It has been used to set new records in the MLPerf training benchmarks, demonstrating its superior performance in these areas.
Conclusion
The NVIDIA H100 Tensor Core GPU marked a significant step forward in GPU technology, combining fourth-generation Tensor Cores with a dedicated Transformer Engine for large language model training and inference. While newer NVIDIA architectures have since surpassed it, the H100 remains a capable choice for organizations running established high-performance computing and AI workloads.
More terms
Continue exploring the glossary.
Glossary term
What is AI Governance?
It's time to build
Collaborate with your team on reliable Generative AI features.
Want expert guidance? Book a 1:1 onboarding session from your dashboard.