Hobbyists Run Modern AI on Decade-Old 'E-Waste' Servers
16 Jul 2026
The old-hardware AI experiment
Two separate hobbyist projects are testing a question with real implications for cash-strapped founders: how far can genuinely old, decommissioned hardware go with today's AI workloads?
The first is a nearly year-long benchmarking effort covering 15 decommissioned NVIDIA Tesla-series enterprise GPUs, including the K80 ($60, 24GB GDDR5 VRAM), the P100-16GB (~$75, 16GB VRAM), the V100-16GB (under $200, 16GB VRAM), and the M60 ($50). The benchmarking tool built for this project has been published on GitHub, making it available to homelab users and builders of low-cost GPU nodes.
A companion build pairs the GPU testing with a used E5-2690 Xeon CPU ($40, 56 threads, 3.50GHz boost) mounted on a Supermicro X10DRG-Q dual-socket motherboard ($200, 7 PCIe slots) — the kind of setup increasingly associated with cheap, DIY inference rigs.
Running a modern MoE model on 13-year-old gear
Separately, a builder ran Google's Gemma 4, a 26-billion-parameter mixture-of-experts open-weights model (30 layers, 8 active experts per token, 262,000-token vocabulary), on a 13-year-old HP StoreVirtual storage box — dual Ivy Bridge Xeons, no GPU at all. The model generated at roughly 5 tokens per second.
Getting there wasn't plug-and-play. The ik_llama.cpp fork used for inference assumed AVX2 support that pre-2013 Ivy Bridge Xeons don't have, causing a compatibility break. The author used Claude to diagnose and fix the issue, then submitted a patch to the ikawrakow/ik_llama.cpp GitHub repo (issue #2138), which is awaiting maintainer review as of publication.
The caveats
None of the benchmarked Tesla GPUs will receive further CUDA compatibility updates or new drivers — they're frozen in time from a software support standpoint. Older GPU and CPU hardware also consumes more energy per token generated, though no specific power figures were provided in this report. And the AVX2 compatibility gap on pre-2013 Xeons is a reminder that legacy instruction sets can quietly break modern inference stacks.
Several details remain unclear: no per-GPU throughput numbers were given for the 15 Tesla cards, there's no combined system build cost, it's not specified which AI workloads were tested across the GPU lineup, and the quantization level or memory footprint used to fit the 26B model onto the old server wasn't disclosed.
Why founders should care
For early-stage teams watching every dollar of infrastructure spend, these projects suggest — though don't prove at production scale — that some AI inference workloads may be less hardware-constrained than commonly assumed. A model running at 5 tokens/sec on a 13-year-old, GPU-less storage box is unlikely to power a customer-facing product, but it's a plausible signal that prototyping, testing, or low-throughput internal tooling could be shifted onto cheap secondhand hardware rather than cloud GPU instances.
That said, the tradeoffs are real. Decommissioned Tesla GPUs likely represent one of the last remaining low-cost sources of idle VRAM, but the total absence of future driver or CUDA support means founders betting on this path should weigh short-term savings against the near-certainty of eventual obsolescence and maintenance risk. Teams with in-house hardware skills — or access to AI tools like Claude for troubleshooting compatibility issues, as used here — are best positioned to make this tradeoff work; teams without that capacity may find the debugging overhead erases the cost advantage.
Most notably, the publication of the benchmarking tool on GitHub could let other resource-constrained builders replicate or extend this testing without redoing the groundwork, which may modestly lower the barrier for founders considering a scrappy, hardware-first approach to early AI product development.