
Groq Inc
AI inference chip company competing directly with Nvidia. Builds the LPU, a Language Processing Unit based on a Tensor Streaming Processor architecture with 230MB on-chip SRAM and no HBM, using static compile-time scheduling to eliminate GPU non-determinism. Delivers Llama 3 70B at roughly 800 tokens per second via GroqCloud's API.