New GPUThor attack defeats NVIDIA ECC protection for root access

New GPUThor attack defeats NVIDIA ECC protection for root access

New GPUThor attack defeats NVIDIA ECC protection for root access

A newly discovered Rowhammer attack called GPUThor can bypass error-correcting code (ECC) protection on NVIDIA GPUs, enabling denial-of-service (DoS) and root-level privilege escalation.

In a paper published by the University of Toronto, researchers say GPUThor achieves far more practical bit flip rates than previous concepts such as their own GPUHammer or GPUBreach, which became irrelevant after the introduction of ECC.

The attack was demonstrated against NVIDIA Ampere-class workstation GPUs with GDDR6 memory, including RTX A4000, RTX A4500, RTX A5000 and RTX A6000, all of which are widely used in AI and cloud infrastructure.

Picture

Improving the GPUThor attack

Rowhammer is the name given to a class of attacks that involve repeatedly accessing (“hammering”) rows of memory, increasing the likelihood that bits in adjacent memory areas will be flipped and change state from one to zero or vice versa.

This can lead to data corruption and security risks. Since training AI models relies heavily on GPU performance, a successful Rowhammer attack could have devastating effects on the model’s accuracy.

To mitigate the risks of this type of attack, NVIDIA uses mitigations such as SECDED ECC to correct single-bit errors and detect double-bit errors in monitored memory blocks.

However, researchers at the University of Toronto have adapted GPUThor Therefore, the hammering follows an uneven pattern at a rate that avoids activation of GDDR6’s Target Row Refresh (TRR) remedies.

Specific hammering pattern used by GPUThor
Specific hammering pattern used by GPUThor
Source: University of Toronto

They achieved this by considering two undocumented GPU behaviors: how repeated memory requests are merged and how frequently TRR is activated.

The researchers say the adjustment results in 6.6 times more aggressor row activations compared to previous attack concepts, achieving between 72,000 and 377,000 flips per GB on the tested GPUs without ECC protection.

Bit flip rates for each attack method
Bit flip rates for each attack method
Source: University of Toronto

These results are between 4,548 and 23,597 times higher than GPUHammer, the researchers’ previous attack, and approach the bit flip rates achieved by powerful CPU Rowhammer attacks like Blacksmith.

At GPUThor bit flip rates, finding an exploitable bit flip is possible within about 1.1 minutes, as opposed to 21.9 hours for GPUHammer.

The researchers explain that with ECC enabled, GPUThor generated 387 double-bit errors that ECC detects but cannot correct, as well as two triple-bit errors that ECC incorrectly repaired, resulting in data corruption.

DoS and privilege escalation

Researchers at the University of Toronto have shown that GPUThor can trigger a DoS condition on an ECC-enabled RTX A6000, causing the GPU to reset every two hours and terminating all workloads.

After a repeated attack on the same card, the device will eventually mark itself as requiring replacement.

The more interesting attack is privilege escalation to the root level, which the researchers believe is possible by corrupting GPU page tables, granting arbitrary memory access to an unprivileged CUDA program, and opening a root shell on the host system.

How to defend yourself against the attack

Aside from the four models proven to be vulnerable to GPUThor, the researchers say that despite restrictions that improve resilience to denial-of-service (DoS) conditions, privilege escalation can still work on Ampere server-class (A100) GPUs because they still rely on SECDED-level ECC.

On some Blackwell GPUs, the RAS repair resiliency feature makes a GPUThor attack more time consuming, but does not prevent it.

According to the researchers GPUThor paper Published yesterday, even HBM3/e and GDDR7 GPUs with on-die ECC could be vulnerable when multi-bit flips are triggered.

The researchers reported their findings to NVIDIA on April 29 and to the company on August 21 has published a guide give instructions.

NVIDIA recommends enabling both SYS-ECC and IOMMU/DMA isolation, monitoring GPU error telemetry, and restricting sharing or execution of untrusted workloads.

The company says the risk varies depending on the DRAM device, memory technology, platform design, in-DRAM defenses, and system configuration, and notes that no bit flips were observed on tested GDDR6X or HBM2e GPUs with the same attack patterns.

The researchers recommend avoiding sharing cross-tenant GPUs when possible, monitoring ECC error counters, and restricting untrusted CUDA workloads. They added that full protection will likely require stronger multi-bit ECC and hardware protections in future GPUs.


Item image

Overall prevention scores can hide what happens after the first access. Once attackers use valid credentials, prevention drops sharply.

The 2026 Blue Report measures defense technology for technology in 338 million simulations conducted in customer production environments.

Get the report

Leave a Reply

Your email address will not be published. Required fields are marked *