AMD acquires Taalas to build 10x faster AI chips
AMD has agreed to acquire Toronto-based startup Taalas to integrate its model-locked silicon technology, which runs specific AI models up to ten times faster than traditional GPUs.

AMD has agreed to acquire Toronto-based startup Taalas in a transaction expected to close in Q4 2026, subject to regulatory approval. While financial terms were not disclosed, this acquisition represents AMD's second recent Canadian inference deal, following its integration of the Untether AI team in 2025 and its long-standing Canadian engineering presence established by buying ATI Technologies in 2006. Taalas specializes in what it terms 'Hardcore Models,' an architecture that physically wires a specific AI model's weights and dataflow directly into the silicon. This model-locked design removes general programmability entirely, meaning each chip runs exactly one model, and switching models requires fabricating a new chip.
To make this rigid approach commercially viable, Taalas utilizes a structured-ASIC manufacturing method where TSMC customizes only the final two metal layers of a nearly complete processor. This allows the foundry to deliver a finished, customized part in just two months, compared to the six months required to manufacture general-purpose chips like Nvidia's Blackwell from scratch. Taalas, which emerged from stealth in February, demonstrated over 16,000 tokens per second per user on Llama 3.1-8B, with claims of reaching 17,000 tokens per second per user. This represents a performance increase of roughly 10x over current GPU setups, alongside a 20x lower build cost and 10x lower power consumption. However, this initial silicon relies on aggressive quantization down to a custom 3-bit format, which reduces output quality compared to GPU baselines, though second-generation chips will adopt standard 4-bit floating-point formats.
For AI practitioners, this acquisition signals a major shift toward disaggregated inference. Mirroring Nvidia's $20 billion Groq deal seven months prior, AMD aims to pair these specialized Taalas chips with its Instinct GPUs. For enterprise developers running static, high-volume models at scale, the trade-off of zero programmability becomes highly attractive. Instead of relying on expensive, power-hungry GPU clusters, practitioners can deploy highly efficient, single-purpose chips that Taalas bets will be a thousand times more efficient than software-based alternatives. This setup allows developers to offload stable, high-throughput workloads to dedicated silicon while reserving flexible GPUs for training and experimental architectures.
This is our own summary of reporting by AlphaSignal



