WHAT'S NEW?
News

GEEKOM Builds Four-Node A9 Mega AI Cluster for Local DeepSeek V4 Flash Inference

Grant Soriano
·
September 2, 2026  ·  2 min read
Add as Preferred Source on Google
Geekom A9 Mega Mini PC Cluster

GEEKOM has demonstrated a four-node AI cluster built from its A9 Mega Mini PCs, running DeepSeek V4 Flash as a distributed platform. Instead of using a traditional data-center server, the setup connects four A9 Mega systems through USB4 to handle AI workloads locally. The configuration highlights how compact PCs can be combined to provide additional compute capacity while keeping the underlying systems independently usable.

Ryzen AI Max+ 395 Powers Each Node

Each GEEKOM A9 Mega is powered by the AMD Ryzen AI Max+ 395, which combines 16 Zen 5 CPU cores, Radeon 8060S graphics, and unified memory. Ubuntu and ROCm are used alongside DwarfStar to distribute an optimized DeepSeek V4 Flash model across the four systems. Applications and AI agents can then connect through an OpenAI-compatible API, while USB4 provides the link between the nodes without requiring a proprietary high-speed switch or server rack.

Designed for Local AI Workloads

The cluster is intended for organizations that need to keep AI processing and sensitive information within their own infrastructure. GEEKOM says prompts, documents, source code, credentials, and intermediate results can remain local instead of being sent to a public cloud.

The setup can also be used for private knowledge assistants that work with internal policies, manuals, contracts, and reports. Other supported workloads include document review, source-code analysis, local retrieval-augmented generation (RAG), multi-source research, and controlled workflow automation.

Up to 250K Tokens for Agent Workloads

The system can also be used with agent platforms such as Hermes Agent, allowing tools, policies, memory, logs, code, and retrieved information to be processed as part of multi-step tasks. According to GEEKOM, the technical path has operated with contexts of up to 250K tokens, which the company says makes the setup suitable for long documents, large codebases, and more complex jobs.

In testing, the four-node system delivered approximately 14.61 tokens per second at single concurrency, with a P95 time to first token of around 0.42 seconds in the 32- and 128-token tests. GEEKOM notes that the configuration is intended to provide additional capacity and acceleration for long prompts rather than simply increasing text-generation speed.

Flexible Four-Node Expansion

The setup does not require all four A9 Mega systems to be deployed at once. Organizations can start with one or two systems and add additional nodes as their AI workloads grow, with each A9 Mega capable of operating independently or contributing to distributed inference.

GEEKOM positions the configuration for businesses, laboratories, classrooms, and edge environments looking to deploy local AI without moving directly to a traditional server-based infrastructure. Pricing and availability for the four-node A9 Mega AI cluster were not provided in the announcement.

Grant Soriano
PC Review Lead and Sales Officer at Gadget Pilipinas | Website

Grant is a Financial Management graduate from UST. His passion for gadgets and tech crossed him over in the industry where he could apply his knowledge as an enthusiast and in-depth analytic skills as a Finance Major. His passion allows him to earn at the same time help Gadget Pilipinas' readers in making smart, value-based decisions and purchases with his reviews and guides.

Grant Soriano

347 articles published

Grant is a Financial Management graduate from UST. His passion for gadgets and tech crossed him over in the industry where he could apply his knowledge as an enthusiast and in-depth analytic skills as a Finance Major. His passion allows him to earn at the same time help Gadget Pilipinas' readers in making smart, value-based decisions and purchases with his reviews and guides.

View all articles →

Related News