vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72 #AI #Hardware #Performance
DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0 #AI #LLM #Performance
Taking vLLM Apart: A Practical Guide to Disaggregated Serving #AI #LLM #Tech
Watermarking in vLLM #AI #LLM #Watermarking
Announcing vllm-metal: Concurrent Serving on Apple Silicon #AI #Apple #OpenSource
PD Serving of Qwen3.8-2.4T #AI #LLM #Performance
Scaling Multi-GPU Video Captioning with PyNvVideoCodec and vLLM #AI #Video #Performance
vLLM x Novita AI: Chord, Faster INT4 MoE for Kimi K2.x. Up to 1.3x on H200, 2.15x on Untuned B300 #AI #LLM #Performance
How we trained the fastest DSpark for Kimi-K3 using GB300 NVL72 #AI #LLM #Performance
vime × RL-Kernel × AMD: Bitwise Train-Rollout Consistency on ROCm #AI #MachineLearning #OpenSource
Kimi K3 Performance Optimizations in vLLM: The Road to 2.8× Throughput #AI #Performance #LLM
Tiered KV Cache Offloading in vLLM #AI #MachineLearning #OpenSource
Following the Bottleneck: Optimizing MiniMax M3 on AMD Instinct MI355X #AI #LLM #Hardware
vLLM x AgentX: Optimizing for Real-World Agentic Serving #AI #LLM #OpenSource
GLM 5.3 Optimizations, Part 1: Hybrid HiSparse Offloading in vLLM #AI #LLM #Performance
Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin #AI #Hardware #OpenSource
MiniMax H3 on vLLM-Omni: From System-Wide Optimization to Real-Time Serving with FastVideo's FastH3 #AI #LLM #Performance
Exploring Speculative Decoding in vLLM on AMD GPUs #AI #OpenSource #Hardware
Large-Scale Sharded Weight Transfer with Ray Direct Transport (RDT) in vLLM #AI #MachineLearning #OpenSource
IsoExec: Unified Execution to Eliminate Trainer-Inference Mismatch in SkyRL #AI #MachineLearning #RL
VeRL-Omni v0.2.0: Faster Diffusion RL and Stable Omni Training #AI #MachineLearning #OpenSource
Distributed Layerwise Offload: Scaling Toward 200B+ DiT Models Efficiently in vLLM-Omni #AI #LLM #DistributedComputing
Adaptive Verification in vLLM: DSpark confidence-scheduled verification #AI #MachineLearning #LLM
Day 0 Support for Qwen3.8-2.4T-A95B on vLLM #AI #LLM #OpenSource
Announcing Day-0 Support for NVIDIA Nemotron 3.5 Lightning on vLLM #AI #LLM #OpenSource
Efficient Decode Context Parallelism with vLLM for Long Context Workloads #AI #LLM #Performance
vLLM Reaches 25K Total TPS/GPU on Qwen3.5 #AI #LLM #Performance
Optimizing vLLM on Arm CPUs #AI #ARM #Performance
Parallel All the Way Down: Beyond Single-Token Generation with Speculative Decoding #AI #MachineLearning #LLM
Kimi K3 Is Here: Efficient Day-0 Support on vLLM #AI #LLM #Tech
Announcing vLLM AFD Plugin: Disaggregating Attention and FFN for Flexible MoE Serving #AI #LLM #DeepLearning
From Day 0 to Production SLAs: Serving GLM-5.2 on 24 NVIDIA B300 GPUs with vLLM #AI #LLM #Performance
A Preview of Production-Scale Kimi K3 Support on vLLM #AI #LLM #OpenSource
Beyond a Single Model: Building Mixture-of-Models Systems with vLLM Semantic Router #AI #LLM #OpenSource
Keeping vLLM Production Quality: A Look Inside CI, Benchmarking, and the Release Process #AI #OpenSource #DevOps
TML Inkling on vLLM: Day-0 Support with Optimized Performance #AI #MachineLearning #LLM
vLLM x TileRT: Specialized Decode for Latency-Critical Serving #AI #LLM #Performance
EAGLE3 Speculative Decoding on AMD Instinct GPUs: Training and Serving with vLLM and AMD Quark #AI #GPU #OpenSource
vime + ROCm: End-to-End RL Post-Training on AMD Instinct™ GPUs #AI #MachineLearning #OpenSource
vLLM × HPC-Ops: High-Performance Attention and MoE Backends from Tencent Hunyuan #AI #MachineLearning #HPC