Distributed LLM Ran on Six Consumer Devices
A user successfully executed a 120B parameter model across a heterogeneous network, reaching 1.1 tokens per second.
Updated on Sept. 28, 2026 in Artificial Intelligence

Live Poll
Is it worth the effort to repurpose consumer hardware for running high-performance AI models?
A user has successfully run a 120B parameter large language model by networking six consumer devices into a single distributed cluster. This research-stage experiment achieved a processing speed of 1.1 tokens per second using a 4-bit quantized version of the gpt-oss-120b model.
Why it matters
This experiment demonstrates that massive large language models can be executed on fragmented consumer hardware rather than centralized enterprise servers. By distributing a model's memory requirements, the work highlights a path toward running state-of-the-art AI without high-end data center equipment.
The 120B parameter model required 47GB of memory across the cluster, well below the 240GB required for an unquantized version. The setup utilized a 4-bit quantized version of gpt-oss-120b, which typically requires a 60GB minimum memory footprint to run.
The players
Tiny
The primary computer used in the cluster, equipped with 28.7GB of system memory and an RTX 3060 graphics processor.
The details
The experiment used pipeline parallelism—a technique where sequential layers of a neural network are partitioned and assigned to different hardware nodes to balance memory load. The AI model was sliced layer-by-layer across the network, which included a Galaxy S24+, an RTX 3060 mini PC, a Windows notebook, and three Mac computers. The primary node, named Tiny, anchored the operation with 28.7GB of system memory and an RTX 3060 graphics processing unit.
Timeline
September 28, 2026: The experiment findings were documented and published.
The Tech Race
Running a 120B parameter model on consumer hardware challenges the current reliance on specialized data center clusters for large-scale AI inference. The experiment confirms that distributed memory pooling can effectively bypass the extreme VRAM constraints that typically gate access to high-parameter LLMs.
This development suggests that enthusiasts may eventually be able to host massive models locally by pooling existing household devices. It requires significant technical configuration of pipeline parallelism, but indicates a future where model access is not strictly limited by local GPU capacity.
The takeaway
This experiment proves that massive model parameter counts can be circumvented by clever layer-wise distribution across heterogeneous hardware. Watch for future open-source tools that automate the orchestration of these distributed node clusters.
Further reading
For more on the current technical capabilities of neural network deployment, see Artificial Intelligence.
Source note: This article includes information reported by Wccftech.
Live Poll
Is it worth the effort to repurpose consumer hardware for running high-performance AI models?






