LATESTEDITORIALLEADERBOARDTOPICS
LATESTEDITORIALLEADERBOARDTOPICS
AboutAdvertisePrivacy
vLLM profile image

vLLM

Open-source LLM inference and serving engine, originated at UC Berkeley's Sky Computing Lab. Built around PagedAttention for efficient KV cache memory management, with continuous batching, tensor and pipeline parallelism, and quantization support (FP8, GPTQ, AWQ). Supports 200+ Hugging Face model architectures with an OpenAI-compatible API.
Last 12 months

Focus areas

INFRA•OPEN_SOURCE•LLMS•INFERENCE OPTIMIZATION•MODEL SERVING•KERNELS

Links

Website•X

vLLM's Hybrid HiSparse Triples Concurrent Requests on Million-Token Contexts

vLLMvLLM·Infra
·12 hrs ago·

vLLM v0.28.0 Ships Sparse Attention and 60% Faster Speculative Decoding

vLLMvLLM·Infra
·Aug 27·

NovaSky's IsoExec Fixes the Hidden Math Bug Corrupting AI Training Runs

vLLMvLLM·Post Training
·Aug 21·

vLLM's DSpark Beats Every Fixed-Length Config on DeepSeek-V4 at Any Load

vLLMvLLM·Infra
·Aug 15·

vLLM v0.26.0 Ships Day-0 Support for Inkling's 1T-Parameter Multimodal Model

vLLMvLLM·Infra
·Jul 26·

vLLM's AFD Plugin Cuts DeepSeek-V3.2 Response Time by 47%

vLLMvLLM·Infra
·Jul 24·

vLLM's TileRT Integration Hits 618 tok/s by Splitting AI's Two Hardest Jobs

vLLMvLLM·Infra
·Jul 15·

Ant Group Pushes Qwen3-Omni to 5.4x Faster With 0.6s Audio Response

vLLMvLLM·Infra
·Jul 03·

vLLM-Omni Squeezes 172% More Audio Out of Four Speech Models

vLLMvLLM·Audio
·Jun 29·

Baidu's Unlimited-OCR Parses 200-Page PDFs With Flat Memory Usage

vLLMvLLM·Image
·Jun 28·

AI moves faster than any field in history. AlphaSignal tracks every paper, repo, model, and release in real time—organized, searchable, and tuned to your work.

Product

LatestEditorialLeaderboardTopicsArchive

Resources

NewsletterAdvertiseGet ListedEnterprise

Company

AboutEditorial TeamTermsPrivacy

© 2026 AlphaSignal