Forlinx Released M.2 AI Accelerator Using Rockchip Silicon

The new hardware integrates up to 5GB of DRAM to support offline inference for models with up to 7 billion parameters.

Updated on Sept. 29, 2026 in Semiconductors

Isometric editorial illustration of a stack of three silicon wafers with central processor cores, representing high-performance AI hardware.
Forlinx Embedded released a new M.2 AI accelerator card utilizing 3D-stacked memory and Rockchip processors to facilitate offline inference for large language models. AI Illustration. Upload story photo >

Live Poll

Would you prefer to run AI models on your own hardware instead of using cloud services?

Forlinx Embedded has released a new M.2 AI accelerator card powered by Rockchip RK1820 and RK1828 processors. This hardware is now available for deployment in Linux and Android systems to execute local offline inference.

Why it matters

By integrating DRAM directly into the processor package, this hardware reduces memory bandwidth bottlenecks that typically throttle local language-model performance. This approach enables the execution of 3B to 7B parameter models without relying on host system memory.

The module achieves 20 TOPS of INT8 computing performance using a 3D stacked architecture that places up to 5GB of DRAM directly alongside the processor. This design significantly outperforms standard host-memory configurations for the 3B to 7B parameter models supported.

The players

Forlinx Embedded

A developer of industrial-grade embedded hardware and computing modules focused on localized AI processing.

Rockchip

A semiconductor company specializing in system-on-chip solutions for mobile, embedded, and AI-driven edge devices.

The details

The accelerator utilizes a 3D stacked architecture where memory is placed within the same package as the processor, shortening the path data must travel to reach the compute cores. It supports PCIe cascading, a technique for connecting multiple expansion cards to distribute individual Transformer layers—the underlying mathematical building blocks of AI models—across several chips. The system is managed via the Rockchip RKNN3 software development kit.

Timeline

  1. September 29, 2026: Official publication of the release details.

The Tech Race

The development follows the industry trend of adapting the M.2 2280 form factor to house specialized AI compute modules for edge environments. It sits in competition with existing edge AI accelerators by prioritizing integrated memory density to enable models previously confined to server-grade hardware.

Developers and engineers can use these modules to deploy offline AI models like Qwen2.5 on embedded systems with limited host memory. Initial deployment requires integration with the Rockchip RKNN3 SDK on Linux or Android-compatible platforms.

The takeaway

This card provides a practical path for running 7B-parameter models in power-constrained environments. Watch for benchmark results comparing this stacked-memory performance against standalone NPU modules in upcoming developer community testing.

Further reading

For more on the latest hardware trends, explore the latest updates in our Semiconductors section.

Source note: This article includes information reported by LinuxGizmos.

Live Poll

Would you prefer to run AI models on your own hardware instead of using cloud services?

Forlinx Released M.2 AI Accelerator Using Rockchip Silicon | Highwise Tech