Distributed LLM Ran on Six Consumer Devices

A user successfully executed a 120B parameter model across a heterogeneous network, reaching 1.1 tokens per second.

Updated on Sept. 28, 2026 in Artificial Intelligence

Isometric editorial illustration of six interconnected hardware blocks arranged in a symmetric grid to represent distributed computing.
A researcher has successfully executed a 120B parameter large language model by networking six consumer-grade devices, achieving 1.1 tokens per second. AI Illustration. Upload story photo >

Live Poll

Is it worth the effort to repurpose consumer hardware for running high-performance AI models?

A user has successfully run a 120B parameter large language model by networking six consumer devices into a single distributed cluster. This research-stage experiment achieved a processing speed of 1.1 tokens per second using a 4-bit quantized version of the gpt-oss-120b model.

Why it matters

This experiment demonstrates that massive large language models can be executed on fragmented consumer hardware rather than centralized enterprise servers. By distributing a model's memory requirements, the work highlights a path toward running state-of-the-art AI without high-end data center equipment.

The 120B parameter model required 47GB of memory across the cluster, well below the 240GB required for an unquantized version. The setup utilized a 4-bit quantized version of gpt-oss-120b, which typically requires a 60GB minimum memory footprint to run.

The players

Tiny

The primary computer used in the cluster, equipped with 28.7GB of system memory and an RTX 3060 graphics processor.

The details

The experiment used pipeline parallelism—a technique where sequential layers of a neural network are partitioned and assigned to different hardware nodes to balance memory load. The AI model was sliced layer-by-layer across the network, which included a Galaxy S24+, an RTX 3060 mini PC, a Windows notebook, and three Mac computers. The primary node, named Tiny, anchored the operation with 28.7GB of system memory and an RTX 3060 graphics processing unit.

Timeline

  1. September 28, 2026: The experiment findings were documented and published.

The Tech Race

Running a 120B parameter model on consumer hardware challenges the current reliance on specialized data center clusters for large-scale AI inference. The experiment confirms that distributed memory pooling can effectively bypass the extreme VRAM constraints that typically gate access to high-parameter LLMs.

This development suggests that enthusiasts may eventually be able to host massive models locally by pooling existing household devices. It requires significant technical configuration of pipeline parallelism, but indicates a future where model access is not strictly limited by local GPU capacity.

The takeaway

This experiment proves that massive model parameter counts can be circumvented by clever layer-wise distribution across heterogeneous hardware. Watch for future open-source tools that automate the orchestration of these distributed node clusters.

Further reading

For more on the current technical capabilities of neural network deployment, see Artificial Intelligence.

Source note: This article includes information reported by Wccftech.

Live Poll

Is it worth the effort to repurpose consumer hardware for running high-performance AI models?

Distributed LLM Ran on Six Consumer Devices