AMD is buying Canadian startup Taalas, which hard-codes model weights directly into inference chips. That makes them extremely fast but locks each chip to a single model. A demo chip hit over 16,000 tokens per second per user running Llama 3.1-8B. Google is reportedly working on a similar approach for Gemini.

Financial terms or expected timeline for completion were not disclosed

Early tech demos show model-specific integrated circuits churning out up to 17,000 tokens a second