AI

AMD bets that etching AI weights into silicon forever will beat Nvidia at inference

Susan Hill

AMD is acquiring Taalas, a Toronto startup that builds chips which embed AI model weights permanently into silicon rather than loading them from memory during inference. The company’s HC1 chip generates 16,960 tokens per second on Meta’s Llama 3.1 8B — a figure AMD claims is 48 times faster than Nvidia‘s GPUs and 8.5 times faster than Cerebras, the inference chip specialist.

The technology behind Taalas exploits a structural inefficiency in how current AI inference works. Standard GPUs are flexible: they can run any model, switch between tasks on demand, and be reprogrammed when a model is updated. That flexibility is also the bottleneck. Loading billions of model weights from high-bandwidth memory into compute units on every inference request creates latency that adding more GPUs cannot eliminate.

Taalas embeds those weights directly into the chip’s silicon at the time of manufacture. The model lives in the hardware, not in memory that gets fetched per request. The throughput numbers AMD is advertising reflect that architectural choice — but at a price the company does not downplay: once a chip is manufactured with a specific model baked in, updating it requires a new silicon tape-out. Taalas says only two metal layers need to change for a model refresh, which shortens the cycle, but these chips cannot adapt on the fly to a newer model version.

Under AMD’s integration plan, Taalas chips will handle token generation — the most time-consuming phase of inference — while AMD’s Instinct GPUs handle prefill and attention computation at the start of each request. The combined system runs on AMD’s Helios rack-scale platform. For cloud operators running fixed-model workloads — customer service bots, document processing pipelines, real-time translation — the promise is throughput per rack watt that flexible GPU clusters cannot match at equivalent cost.

Whether that calculus survives the industry’s model refresh pace is the central question. Major labs now update their production models roughly every six months. A chip locked to Llama 3.1 8B today may need to be retired when a successor becomes the standard deployment target. Taalas’s two-metal-layer refresh claim helps, but a silicon tape-out still takes months to complete — a meaningful lag when models turn over faster than hardware cycles.

The acquisition comes seven months after Nvidia bought Groq — a different approach to the same inference speed problem — for a reported $20 billion. Groq’s tensor streaming processors also optimize for throughput but retain the ability to load different models. AMD is paying an undisclosed but presumably smaller sum for a startup that made the opposite bet.

For buyers, the Taalas acquisition belongs to the roadmap rather than the catalog. The HC2 chip, targeting models with up to 20 billion parameters, is still in development. The 48x benchmark applies to HC1 under specific conditions — real-world deployments with mixed traffic and varying batch sizes will produce different results. Deal close is expected in the fourth quarter of 2026, pending regulatory review. HC1 shipped in February on 6nm; HC2 ship date and target process node remain unconfirmed.

Tags: , , , ,

Discussion

There are 0 comments.