computers

Top 10 Fastest GPUs by Compute Performance

The ten fastest graphics processing units ranked by FP32 floating-point compute performance, from NVIDIA's flagship data-center Rubin Ultra to the consumer GeForce RTX 5090.

Updated August 1, 2026 10 ranked 4 sources

Quick answer: 1. NVIDIA Rubin Ultra GPU, 2. AMD Radeon Instinct MI455X, 3. NVIDIA GeForce RTX 6090, 4. NVIDIA Rubin GPU, 5. NVIDIA RTX PRO 6000 Blackwell, 6. NVIDIA RTX PRO 6000D Blackwell, 7. NVIDIA RTX PRO 6000D Blackwell Max-Q, 8. NVIDIA RTX PRO 6000 Blackwell Max-Q, 9. NVIDIA GeForce RTX 5090, 10. NVIDIA GeForce RTX 5090 D.

Graphics processing units (GPUs) have evolved far beyond gaming graphics to become the workhorses of artificial intelligence, scientific computing, and data center workloads. The fastest GPUs today are measured in hundreds of teraflops (TFLOPS) of single-precision (FP32) compute performance, with NVIDIA's latest Rubin Ultra architecture leading the pack at 260 TFLOPS. The GPU market is dominated by NVIDIA, which holds the top positions across both data-center and consumer segments, with AMD competing in the data center with its Instinct series. This list ranks the top ten GPUs by FP32 (single-precision floating-point) compute performance, based on manufacturer specifications and benchmark data from TopCPU and GPU benchmark databases. Both data-center and consumer GPUs are included, reflecting the full spectrum of the fastest graphics processors available.

FP32 compute (TFLOPS)

FP32 compute (TFLOPS)
NVIDIA Rubin Ultra GPU
260 TFLOPS
AMD Radeon Instinct MI455X
157.3 TFLOPS
NVIDIA GeForce RTX 6090
132.7 TFLOPS
NVIDIA Rubin GPU
130 TFLOPS
NVIDIA RTX PRO 6000 Blackwell
126 TFLOPS
NVIDIA RTX PRO 6000D Blackwell
126 TFLOPS
NVIDIA RTX PRO 6000D Blackwell Max-Q
110.1 TFLOPS
NVIDIA RTX PRO 6000 Blackwell Max-Q
109.7 TFLOPS
NVIDIA GeForce RTX 5090
104.8 TFLOPS
NVIDIA GeForce RTX 5090 D
104.8 TFLOPS

The ranking

1

NVIDIA Rubin Ultra GPU

The NVIDIA Rubin Ultra GPU is the most powerful GPU ever announced, delivering 260 TFLOPS of FP32 compute performance. Part of NVIDIA's next-generation Rubin architecture (successor to Blackwell), the Rubin Ultra is designed for the largest AI training and inference workloads in hyperscale data centers. The GPU features advanced packaging with multiple dies interconnected through NVIDIA's NVLink technology. Rubin Ultra is expected to power the next generation of large language models and scientific simulations, with availability targeted for late 2026.

Manufacturer: NVIDIAArchitecture: Rubin (Ultra variant)Market segment: Data center / AIMemory: HBM4 (expected)Status: Announced
2

AMD Radeon Instinct MI455X

The AMD Radeon Instinct MI455X is AMD's flagship data-center GPU, delivering 157.3 TFLOPS of FP32 compute performance. Designed to compete with NVIDIA's Blackwell and Rubin architectures in the AI training and inference market, the MI455X features AMD's CDNA 5 architecture with advanced chiplet design. The GPU includes high-bandwidth memory and supports AMD's Infinity Architecture for multi-GPU scaling. AMD has positioned the MI455X as a strong alternative for open-source AI frameworks and ROCm ecosystem users.

Manufacturer: AMDArchitecture: CDNA 5 (Instinct)Market segment: Data center / AIMemory: HBM3eStatus: Available
3

NVIDIA GeForce RTX 6090

The NVIDIA GeForce RTX 6090 is the flagship consumer graphics card of the next-generation GeForce lineup, delivering 132.7 TFLOPS of FP32 compute performance. Built on the Rubin architecture, the RTX 6090 is designed for the most demanding gaming, content creation, and AI workloads at the enthusiast level. It features the latest generation of RT cores for ray tracing, Tensor cores for AI acceleration, and GDDR7 memory. The RTX 6090 represents a massive leap in consumer GPU performance over the previous generation.

Manufacturer: NVIDIAArchitecture: RubinMarket segment: Consumer / EnthusiastMemory: GDDR7Status: Announced
4

NVIDIA Rubin GPU

The standard NVIDIA Rubin GPU delivers 130 TFLOPS of FP32 compute performance, positioned as the base data-center GPU in the Rubin architecture family. Unlike the Rubin Ultra which uses advanced multi-die packaging, the standard Rubin GPU is a single monolithic die design optimized for a broader range of data center workloads. It offers a balance of performance, power efficiency, and cost that makes it suitable for widespread deployment in cloud data centers and enterprise AI infrastructure.

Manufacturer: NVIDIAArchitecture: RubinMarket segment: Data center / AIMemory: HBM4 (expected)Status: Announced
5

NVIDIA RTX PRO 6000 Blackwell

The NVIDIA RTX PRO 6000 Blackwell is the flagship workstation GPU based on the Blackwell architecture, delivering 126 TFLOPS of FP32 compute performance. Designed for professional visualization, rendering, scientific computing, and AI development, the RTX PRO 6000 replaces the RTX 6000 Ada Generation. It features fourth-generation RT cores, fifth-generation Tensor cores, and a massive memory pool for handling the largest professional workloads, from 3D rendering to medical imaging and scientific simulation.

Manufacturer: NVIDIAArchitecture: BlackwellMarket segment: Workstation / ProfessionalMemory: GDDR7 (expected 96 GB)Status: Available
6

NVIDIA RTX PRO 6000D Blackwell

The NVIDIA RTX PRO 6000D Blackwell is a variant of the RTX PRO 6000 for the Chinese market, delivering 126 TFLOPS of FP32 compute performance — identical to the standard RTX PRO 6000. The 'D' designation indicates compliance with US export regulations for the Chinese market. The card maintains the same Blackwell architecture, RT cores, and Tensor cores as the global version but may have reduced NVLink interconnect capabilities or other feature adjustments to meet regulatory requirements.

Manufacturer: NVIDIAArchitecture: BlackwellMarket segment: Workstation / China marketMemory: GDDR7 (expected 96 GB)Status: Available
7

NVIDIA RTX PRO 6000D Blackwell Max-Q

The NVIDIA RTX PRO 6000D Blackwell Max-Q delivers 110.1 TFLOPS of FP32 compute performance, using NVIDIA's Max-Q technology to optimize power efficiency and thermal performance. The Max-Q variant of the RTX PRO 6000D is designed for workstation laptops and form-factor-constrained systems where power and cooling are limited. Despite the lower peak performance compared to the desktop version, it remains one of the most powerful mobile workstation GPUs ever built, capable of running demanding professional applications.

Manufacturer: NVIDIAArchitecture: BlackwellMarket segment: Mobile workstationTechnology: Max-Q designStatus: Available
8

NVIDIA RTX PRO 6000 Blackwell Max-Q

The NVIDIA RTX PRO 6000 Blackwell Max-Q delivers 109.7 TFLOPS of FP32 compute performance, slightly below the Chinese-market D variant Max-Q. This is the standard global Max-Q version of the RTX PRO 6000 Blackwell workstation GPU, optimized for mobile workstations. It balances compute performance with power efficiency, making it suitable for high-end mobile workstations used in CAD, video editing, 3D rendering, and scientific computing on the go.

Manufacturer: NVIDIAArchitecture: BlackwellMarket segment: Mobile workstationTechnology: Max-Q designStatus: Available
9

NVIDIA GeForce RTX 5090

The NVIDIA GeForce RTX 5090 is the flagship consumer GPU of the current generation, delivering 104.8 TFLOPS of FP32 compute performance. Built on the Blackwell architecture, the RTX 5090 features 21,760 CUDA cores, fourth-generation RT cores, and fifth-generation Tensor cores with 32 GB of GDDR7 memory on a 512-bit memory bus. The RTX 5090 is the most powerful consumer graphics card available, capable of 4K and 8K gaming at high frame rates with full ray tracing enabled, and serves as a powerful platform for AI development and content creation.

Manufacturer: NVIDIAArchitecture: BlackwellMarket segment: Consumer / EnthusiastMemory: 32 GB GDDR7CUDA cores: 21,760
10

NVIDIA GeForce RTX 5090 D

The NVIDIA GeForce RTX 5090 D is the variant of the RTX 5090 designed for the Chinese market, also delivering 104.8 TFLOPS of FP32 compute performance. Like the RTX PRO 6000D, the 'D' designation indicates compliance with US export regulations. The RTX 5090 D maintains identical FP32 performance to the standard RTX 5090, making it one of the most powerful consumer GPUs available in China. It features the same Blackwell architecture, CUDA core count, and GDDR7 memory as the global version.

Manufacturer: NVIDIAArchitecture: BlackwellMarket segment: Consumer / China marketMemory: 32 GB GDDR7CUDA cores: 21,760

Full comparison

# Name ManufacturerArchitectureMarket segmentMemoryStatusTechnologyCUDA cores
1 NVIDIA Rubin Ultra GPU NVIDIARubin (Ultra variant)Data center / AIHBM4 (expected)Announced
2 AMD Radeon Instinct MI455X AMDCDNA 5 (Instinct)Data center / AIHBM3eAvailable
3 NVIDIA GeForce RTX 6090 NVIDIARubinConsumer / EnthusiastGDDR7Announced
4 NVIDIA Rubin GPU NVIDIARubinData center / AIHBM4 (expected)Announced
5 NVIDIA RTX PRO 6000 Blackwell NVIDIABlackwellWorkstation / ProfessionalGDDR7 (expected 96 GB)Available
6 NVIDIA RTX PRO 6000D Blackwell NVIDIABlackwellWorkstation / China marketGDDR7 (expected 96 GB)Available
7 NVIDIA RTX PRO 6000D Blackwell Max-Q NVIDIABlackwellMobile workstationAvailableMax-Q design
8 NVIDIA RTX PRO 6000 Blackwell Max-Q NVIDIABlackwellMobile workstationAvailableMax-Q design
9 NVIDIA GeForce RTX 5090 NVIDIABlackwellConsumer / Enthusiast32 GB GDDR721,760
10 NVIDIA GeForce RTX 5090 D NVIDIABlackwellConsumer / China market32 GB GDDR721,760

How we ranked this

GPUs are ranked by their FP32 (single-precision floating-point) compute performance, measured in teraflops (TFLOPS, trillions of floating-point operations per second). FP32 performance is the most widely used metric for comparing raw GPU compute capability across different architectures and use cases. Figures are based on manufacturer-provided specifications and verified against GPU benchmark databases. The list includes both data-center/enterprise GPUs and consumer graphics cards. Where multiple variants of the same GPU exist (e.g., different power limits or form factors), the highest-performance variant is listed.

FAQ

What is the fastest GPU in the world?

The NVIDIA Rubin Ultra GPU is the fastest GPU ever announced, with 260 TFLOPS of FP32 compute performance. However, it is a data-center GPU not yet available for consumers. The fastest available consumer GPU is the NVIDIA GeForce RTX 5090 with 104.8 TFLOPS.

What does TFLOPS mean in GPUs?

TFLOPS stands for trillion floating-point operations per second. It measures how many trillions of calculations a GPU can perform per second and is the standard metric for comparing raw compute performance across different GPUs.

Is AMD competitive with NVIDIA in GPU performance?

AMD's Radeon Instinct MI455X is the second-fastest GPU overall with 157.3 TFLOPS, competing directly with NVIDIA's data-center offerings. However, NVIDIA dominates the consumer GPU market and holds the vast majority of positions in the overall performance rankings.

Sources