The deal strengthens AMD’s push to build a full-stack AI platform and gives it access to Taalas’ specialized inference technology, which is designed to reduce the memory and computing bottlenecks that can limit AI performance.
Taalas takes a fundamentally different approach to AI-chip design by embedding model weights directly into the silicon rather than repeatedly moving them between processors and external memory. Its first-generation HC1 chip, built on TSMC’s 6nm process, is optimized for Meta’s Llama 3.1 8B model and combines memory and computing on a single chip, eliminating the need for high-bandwidth memory (HBM). This can reduce memory bottlenecks while avoiding costly advanced packaging technologies such as CoWoS and 3D memory stacks.
The company claims significant inference performance advantages, citing 16,960 tokens per second per user for Llama 3.1 8B, versus 594 for Nvidia and Groq and 1,981 for Cerebras in its comparisons. These figures are Taalas’ own claims.
The approach essentially treats “the model as the computer,” with model weights permanently embedded using mask-programmed read-only memory. Only a small number of layers need to change when adapting the design to a new model, allowing Taalas to produce customized chips relatively quickly.











