AMD Acquires Taalas to Build Chips That Run One AI Model 10x Faster

AMD acquires Toronto startup Taalas, which hardwires AI models directly into silicon, promising 17,000 tokens/sec at 20x lower cost than GPUs

·
·
AMD Acquires Taalas to Build Chips That Run One AI Model 10x Faster
Read6 min
TypeNews
TopicGpus · Business
SubtopicAccelerators
  • AMD acquires Taalas: AMD has agreed to buy Toronto-based Taalas, a startup that hardwires AI models directly into custom silicon for ultra-fast inference.
  • Extreme performance claims: Taalas's first chip runs Llama 3.1-8B at 17,000 tokens/sec per user -- roughly 10x faster than GPU setups -- at 20x lower cost and 10x less power.
  • Model-locked silicon: Each Taalas chip runs exactly one model; switching models requires a new chip, but TSMC can produce a customized part in just two months.
  • Mirrors Nvidia-Groq: The deal comes seven months after Nvidia's $20B Groq deal, and positions AMD to pair Taalas chips with Instinct GPUs for disaggregated inference, just as Nvidia does with Groq.
  • AMD's second Canadian inference acquisition: AMD also acquired the Untether AI team in 2025, and has had a Canadian engineering presence since buying ATI Technologies in 2006.
  • Deal terms and timeline: Financial terms were not disclosed; the deal is expected to close in Q4 2026, subject to regulatory approval.

AMD has agreed to acquire Taalas, a Toronto-based startup that takes a fundamentally different approach to AI inference: instead of building a general-purpose chip and running a model on it, Taalas builds the chip around the model. The result is silicon physically wired to run one specific AI model. Financial terms were not disclosed, but AMD confirmed this is a full acquisition.

How Taalas Chips Work

Taalas builds what it calls "Hardcore Models": processors tailored to a single model's weights, produced by finalizing a small number of a chip's metal layers once the model is fixed. Think of it like the difference between a programmable calculator and a circuit board wired to do only one calculation. The latter is vastly faster and cheaper for that one task.

The approach eliminates the biggest bottlenecks in modern AI serving by hardwiring a model's dataflow between compute elements and burning in the weights. Inference runs extremely fast, but switching to a new model means designing a new chip.

That constraint is less punishing than it sounds, because of how Taalas manufactures. The startup assembles a nearly complete chip and customizes only the final two metal layers per model, so TSMC needs roughly two months to finish a chip versus six months to fabricate something like Nvidia's Blackwell from scratch. Taalas emerged from stealth in February with a demo chip achieving more than 16,000 tokens per second per user on Llama 3.1-8B, built by a team of 24 on just $30 million.

The performance claims are striking, though they come from the company's own data:

  • 17,000 tokens per second per user on Llama 3.1-8B, roughly 10x faster than current GPU setups
  • 20x lower build cost and 10x lower power consumption versus comparable GPU inference
  • First-generation chip uses aggressive quantization to a custom 3-bit data type, which degrades output quality relative to GPU baselines
  • Second-generation silicon moves to standard 4-bit floating-point formats
  • Single chips that could outperform small GPU data centers on targeted workloads

There is a hard ceiling on model size. Taalas's first chip only runs Llama 3.1-8B. Bigger models require more chips: running something like DeepSeek-671B would need around 30 separate tape-outs. The practical sweet spot, at least for now, is models up to around 8 billion parameters.

The Team Behind It

Taalas was founded in 2023 and is led by CEO Ljubisa Bajic, a former AMD executive and ex-CEO, CTO, and president of Tenstorrent; COO Lejla Bajic; and CTO Drago Ignjatovic. Taalas's announcement noted that many of the team "grew up at AMD Canada," giving the acquisition an unusual cultural continuity. Vamsi Boppana, senior vice president of AMD's Artificial Intelligence Group, will oversee integration.

The company raised $219 million in total from investors including Quiet Capital, Fidelity, and chip-industry veteran Pierre Lamond.

Why AMD Wanted This

AMD is counting on its Instinct GPUs to drive data center growth as cloud companies snap up advanced AI chips. But training a model and serving it at scale are fundamentally different workloads, and the economics of inference favor specialization. Running the same model millions of times rewards hardware built for exactly that job.

AMD has been assembling its inference stack piece by piece. In July it announced a partnership with Cerebras to integrate Cerebras chips into its systems. The Taalas deal fits that pattern: AMD plans to integrate Taalas technology into its accelerator roadmap and pair it with Instinct GPUs in Helios rack systems.

This is also AMD's second acquisition of a Canadian AI chip firm in just over a year. In 2025, AMD acquired the team behind Toronto startup Untether AI, which had been developing energy-efficient inference chips. AMD's Canadian presence dates to its 2006 purchase of Markham-based ATI Technologies, which became a long-running engineering hub.

The Nvidia Parallel

The deal mirrors what Nvidia did last December: a $20 billion licensing arrangement with Groq, framed around high-performance inference for AI agents like code assistants. Nvidia plans to use Groq chips alongside its GPUs, relying on Groq specifically for the decode stage of inference. That disaggregated setup frees Groq chips from holding the entire model and KV cache while maintaining token-generation speed.

AMD is now positioned to do the same with Taalas. AMD already announced a disaggregated inference arrangement with Cerebras, where AMD GPUs handle the prefill portion of the workload and Cerebras accelerates decode. A Taalas chip could fill that decode role instead, with the entire system controlled and supplied by AMD.

The broader pattern is consistent across the industry. Qualcomm closed its acquisition of compiler startup Modular in July 2026. Anthropic is building an in-house silicon team. Every major chip company is racing to own the inference stack end-to-end, and GPU-only inference is giving way to specialized, model-matched hardware.

Winners and Losers

The clearest beneficiaries:

  • AMD gets a differentiated inference story paired with Instinct GPUs in Helios racks, making it harder for customers to default to Nvidia
  • Taalas employees gain AMD's manufacturing relationships, scale, and global sales force
  • Cloud operators and enterprises running high-volume inference on stable models could access dramatically cheaper serving costs
  • Canada's AI ecosystem gets a concrete vote of confidence, with AMD committing to retain and grow Canadian talent

The costs are more diffuse. Nvidia's lead in general-purpose AI compute is unthreatened, but its claim to inference efficiency just got harder to defend. Cerebras, which AMD was partnering with for disaggregated inference, may find itself competing with an in-house AMD solution as Taalas matures.

What Changes at Scale

The most consequential second-order effect is what this does to the economics of deploying small, stable models at massive scale. If Taalas's claims hold under independent scrutiny, a company running millions of Llama-class inference calls per day could replace a GPU cluster with a rack of Taalas chips at a fraction of the cost and power draw.

There is also an edge and physical AI angle. Power constraints and cost sensitivity make GPU-based inference impractical in many physical AI deployments, and a two-month tape-out cycle is fast enough to track the pace at which frontier small models are being released. AMD could use Taalas chips for the entire inference workload on small models in those environments, not just as a decode accelerator alongside GPUs.

Will It Close?

Subject to regulatory approval, the deal is expected to close in the fourth quarter. Given that Taalas is a small Canadian startup with no dominant market position, significant antitrust hurdles seem unlikely. The bigger risk is integration: Taalas's value is almost entirely in its engineering team and its proprietary tool flow for rapid chip customization. Keeping that team intact inside a large public company is the real execution challenge.

AMD's Canadian track record is relevant here. The ATI acquisition in 2006 produced a thriving engineering hub that has persisted for two decades. The Untether AI team acquisition in 2025 is still too recent to evaluate. Taalas will be the third data point, and given that Bajic and much of his team are AMD alumni returning to the company, the cultural fit is probably as strong as it gets for a deal of this kind.

Trending
  • No trending articles

Comments

avatar

Next Reads