A Modder’s RTX 4080 Was Enough To Play AAA Games, But Not For Running LLMs, So He Integrated NVIDIA’s Tesla V100 At A Throwaway Price To Run 27B AI Models

Omar Sohail
An NVIDIA Tesla V100 gets modded to a gaming PC to run LLMs and AAA games
16GB VRAM can handle some AI models, but not the beefy ones / Image credits - Tymscar

Running higher-quality AI models means your existing GPU will have to be equipped with a ton of VRAM to make the experience enjoyable. For one RTX 4080 owner, playing the most visually taxing and graphically demanding games might be a walk in the park, but for running LLMs, it’s a Herculean task. Wanting to accomplish both feats in a single gaming PC, a modder successfully ran NVIDIA’s Tesla V100 in his system, but encountered a few challenges along the way.

Adding the Tesla V100 grants the modder 32GB of usable VRAM, sufficient for running significantly improved AI models like Qwen 3.6 at 32 tokens/second

Before you ask, it’s impossible to attach the Tesla V100 to a desktop motherboard like a “plug and play” job, requiring Tymscar to purchase an SXM2-to-PCIe adapter. Sourcing the GPUs with 16GB of HBM2 memory and the accessory cost him £200, which translates into roughly $266. On eBay, you can grab these for only $100 apiece. With the Tesla V100’s 5,120 CUDA cores and 4,096-bit bus width that delivers 900GB/s of bandwidth, the GPU still has some computing juice remaining.

Related Story NVIDIA Trims Vera Rubin Memory as HBM4 Prices Threaten to Eat 29% of Every Rack’s Cost – Report

Successfully sourcing an SXM2-to-PCIe adapter wasn’t the most difficult of challenges, but it isn’t simple either, especially when you find out later that the Tesla V100 doesn’t have a PCIe slot, display inputs, or PCIe power connectors. As you can see in the image below, attaching the GPU to the vapor chamber-style cooler might be perfectly fine for those who don’t mind the excessive noise, but at 82dB, it’ll make anyone uncomfortable.

With a little tweaking here and there to reduce the fan noise by using a 9V battery and a PWM jumper, the V100’s modded cooler was now operating at 10 percent of the original maximum RPM. With this problem out of the way, the modder had successfully found a way to get 32GB of usable VRAM into his system. Now, no game will ever require a whopping 32GB of video memory, unless you decide to run a newer title at 16K resolution, so the best use case would be to fire up Qwen3.6 27B.

The modder was running Qwen3.6-27B-MTP quantized at Q5_K_M, which comes in at 19GB, and with a context size of 128K tokens, there was sufficient VRAM to run the LLM at 32 tokens per second. Prompt processing was between 133 and 160 tokens per second, making it decent performance if you happen to stumble across previous-generation AI GPUs for home-computing purposes. Best of all, for less than $300, you can have your very own small to medium-sized AI models running at home, free of cost and without any internet connection.

News Source: Tymscar

Omar Sohail Photo

About the author: Omar Sohail is a reporter and analyst for Wccftech's mobile section, specializing in the technology and business of the mobile industry. His expertise lies in the intricate hardware supply chain, covering developments in semiconductor manufacturing, chip lithography, and camera sensor technology.

Follow Wccftech on Google to get more of our news coverage in your feeds.

Button