AMD Acquires Taalas to Build Chips That Run One AI Model 10x Faster

AMD acquires Toronto startup Taalas, which hardwires AI models directly into silicon, promising 17,000 tokens/sec at 20x lower cost than GPUs

·
·
AMD Acquires Taalas to Build Chips That Run One AI Model 10x Faster
AuthorTaalas Inc.
Read3 min
TopicGpus · Business
SubtopicAccelerators
  • AMD acquires Taalas: AMD has agreed to buy Toronto-based Taalas, a startup that hardwires AI models directly into custom silicon for ultra-fast inference.
  • Extreme performance claims: Taalas's first chip runs Llama 3.1-8B at 17,000 tokens/sec per user -- roughly 10x faster than GPU setups -- at 20x lower cost and 10x less power.
  • Model-locked silicon: Each Taalas chip runs exactly one model; switching models requires a new chip, but TSMC can produce a customized part in just two months.
  • Mirrors Nvidia-Groq: The deal comes seven months after Nvidia's $20B Groq deal, and positions AMD to pair Taalas chips with Instinct GPUs for disaggregated inference, just as Nvidia does with Groq.
  • AMD's second Canadian inference acquisition: AMD also acquired the Untether AI team in 2025, and has had a Canadian engineering presence since buying ATI Technologies in 2006.
  • Deal terms and timeline: Financial terms were not disclosed; the deal is expected to close in Q4 2026, subject to regulatory approval.

AMD has agreed to acquire Taalas, a Toronto-based startup that takes a radically different approach to AI inference: instead of building a general-purpose chip and running a model on it, Taalas builds the chip around the model. The result is silicon that is physically wired to run one specific AI model, and nothing else. Financial terms were not disclosed, but AMD confirmed this is a full acquisition, not an acquihire.

What Taalas Actually Built

The core idea behind Taalas is deceptively simple. Taalas builds what it calls "Hardcore Models": processors tailored to a single model's weights, produced by finalizing a small number of a chip's metal layers once the model is fixed. Think of it like the difference between a programmable calculator and a circuit board wired to do only one calculation -- the latter is vastly faster and cheaper for that one task.

The approach eliminates the biggest bottlenecks in modern AI serving. The startup essentially hardwires a model's dataflow between compute elements and burns in the weights. Inference runs very fast, but almost all programmability is removed, so only one model can be run. The tradeoff is real: switching to a new model means designing a new chip.

The manufacturing model is what makes this commercially viable. Taalas came out of stealth mode in February with a demo chip that could achieve more than 16,000 tokens per second per user on Llama 3.1-8B. The startup has claimed it can launch new chips in just two months, well faster than the industry standard. That speed comes from a structured-ASIC approach: Taalas assembles a nearly complete chip and only customizes the final two metal layers per model, so TSMC needs roughly two months to finish a chip, versus six months to fabricate a chip like Nvidia's Blackwell from scratch.

The performance numbers Taalas has put forward are striking, though they come with caveats:

  • 17,000 tokens per second per user on Llama 3.1-8B -- roughly 10x faster than current state-of-the-art GPU setups, per the company's own data
  • 20x lower build cost and 10x lower power consumption versus comparable GPU inference
  • First-generation chip uses aggressive quantization to a custom 3-bit data type, which degrades output quality relative to GPU baselines
  • Second-generation silicon moves to standard 4-bit floating-point formats
  • The startup is betting it can produce models that are a thousand times more efficient than their software counterparts, with single chips that could outperform small GPU data centers

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves