The Mac mini popularized the small form factor for running AI agents, thanks to powerful chipsets and adequate unified memory. Adopting this approach, Singaporean startup Acrab believes it has the perfect competitor in the form of an AI chip called Gelix 1. Sporting a 20-core ARM CPU and fabricated on the 5nm process, the company claims that it can beat Apple’s older M4 Pro in AI inference and run up to 100-billion-parameter models.
Acrab’s own benchmarks reveal that the Gelix 1 achieved a pre-fill rate of 1,416 tokens per second when running Gemma 26B
The AI inference capabilities of the Gelix 1 are further fueled by its multi-core neural processing unit, and its unified memory approach allows Acrab’s AI chip to obtain a bandwidth of 273GB/s. Interestingly, this figure matches the M4 Pro’s bandwidth, and, like the Mac mini, the Gelix 1 is also designed to operate in a compact enclosure without performance or thermal limitations.
On the benchmark side of things, Acrab claims that the Gelix 1 achieved a pre-fill rate of 1,416 tokens per second running a Gemma 26B model with a 40,000-token context window. In comparison, the M4 Pro Mac mini achieved only 188 tokens per second in the same test. DigiTimes has reported that independent verifications have yet to be done, so it’s imperative that readers treat this data with a pinch of salt for now.
There are obviously several glaring questions, such as Acrab’s claim of being able to run 100B AI models. The M4 Pro Mac mini is currently limited to 48GB of unified RAM, which should breeze through 35B LLMs. However, a 100B model will require 128GB of unified memory, which might be the maximum configuration of the Gelix 1. The 20-core ARM CPU sounds oddly familiar to NVIDIA’s DGX Spark and RTX Spark, so we’re more than excited to see if it houses the same processor from MediaTek.
Consumers and businesses with privacy and latency concerns can drive the Gelix 1, enabling them to run larger AI models while occupying a negligible amount of desk space. Sadly, pricing and availability details are currently unknown, but if anyone out there plans on running 100B models locally, their desired workflow won’t lead to an affordable configuration.
Follow Wccftech on Google to get more of our news coverage in your feeds.
