Skip to content
Change the repository type filter

All

    Repositories list

    • vLLM plugin for attention-ffn disaggregation support
      Python
      Apache License 2.0
      1765176Updated Aug 3, 2026Aug 3, 2026
    • recipes

      Public
      Common recipes to run vLLM
      JavaScript
      Apache License 2.0
      35394735105Updated Aug 3, 2026Aug 3, 2026
    • vllm

      Public
      A high-throughput and memory-efficient inference and serving engine for LLMs
      Python
      Apache License 2.0
      20k88k2k4.2kUpdated Aug 3, 2026Aug 3, 2026
    • guidellm

      Public
      Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
      Python
      Apache License 2.0
      2021.5k5019Updated Aug 3, 2026Aug 3, 2026
    • TPU inference for vLLM, with unified JAX and PyTorch support.
      Python
      Apache License 2.0
      27439772327Updated Aug 3, 2026Aug 3, 2026
    • Community maintained hardware plugin for vLLM on Ascend
      C++
      Apache License 2.0
      1.9k2.5k1.5k1.1kUpdated Aug 3, 2026Aug 3, 2026
    • Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference
      Go
      Apache License 2.0
      7945.1k19789Updated Aug 3, 2026Aug 3, 2026
    • FlashKDA

      Public
      Cuda
      MIT License
      01000Updated Aug 3, 2026Aug 3, 2026
    • aibrix

      Public
      Cost-efficient and pluggable Infrastructure components for GenAI inference
      Go
      Apache License 2.0
      6395k31238Updated Aug 2, 2026Aug 2, 2026
    • Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
      Python
      Apache License 2.0
      6003.6k4686Updated Aug 2, 2026Aug 2, 2026
    • A safetensors extension to efficiently store sparse quantized tensors on disk
      Python
      Apache License 2.0
      108308744Updated Aug 2, 2026Aug 2, 2026
    • Stateful API logic for agentic applications using vLLM
      Rust
      Apache License 2.0
      23603111Updated Aug 2, 2026Aug 2, 2026
    • vllm-omni

      Public
      A framework for efficient model inference with omni-modality models
      Python
      Apache License 2.0
      1.4k5.8k609715Updated Aug 2, 2026Aug 2, 2026
    • Community maintained hardware plugin for vLLM on Apple Silicon
      Python
      Apache License 2.0
      2011.5k84Updated Aug 2, 2026Aug 2, 2026
    • vvm

      Public
      Manage multiple vLLM installations with isolated Python virtual environments. Switch between releases, commits, branches, and PRs instantly.
      Rust
      Apache License 2.0
      0900Updated Aug 2, 2026Aug 2, 2026
    • HTML
      12459416Updated Jul 31, 2026Jul 31, 2026
    • A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
      Python
      Apache License 2.0
      1766843759Updated Jul 31, 2026Jul 31, 2026
    • vLLM Quantization plugin for bitsandbytes
      Python
      Apache License 2.0
      1201Updated Jul 31, 2026Jul 31, 2026
    • Community maintained hardware plugin for vLLM on Intel Gaudi
      Python
      Apache License 2.0
      14650459Updated Jul 31, 2026Jul 31, 2026
    • vLLM Quantization plugin for GGUF
      Python
      Apache License 2.0
      292949Updated Jul 31, 2026Jul 31, 2026
    • perf-eval

      Public
      Performance benchmark & accuracy evaluation for vLLM
      Python
      1512023Updated Jul 31, 2026Jul 31, 2026
    • Fast and memory-efficient exact attention
      Python
      BSD 3-Clause "New" or "Revised" License
      3k133041Updated Jul 31, 2026Jul 31, 2026
    • ci-infra

      Public
      This repo hosts code for vLLM CI & Performance Benchmark infrastructure.
      Python
      Apache License 2.0
      7545058Updated Jul 30, 2026Jul 30, 2026
    • MSA

      Public
      Python
      MIT License
      49000Updated Jul 30, 2026Jul 30, 2026
    • vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization
      Python
      Apache License 2.0
      4572.5k9983Updated Jul 30, 2026Jul 30, 2026
    • TypeScript
      91012Updated Jul 30, 2026Jul 30, 2026
    • vLLM Daily Summarization of Merged PRs
      55200Updated Jul 30, 2026Jul 30, 2026
    • The vLLM XPU kernels for Intel GPU
      C++
      Apache License 2.0
      92571253Updated Jul 30, 2026Jul 30, 2026
    • DeepGEMM

      Public
      DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
      Cuda
      MIT License
      1.1k000Updated Jul 30, 2026Jul 30, 2026
    • High-performance Rust benchmark client for vLLM serving endpoints.
      Rust
      Apache License 2.0
      115113Updated Jul 28, 2026Jul 28, 2026
    ProTip! When viewing an organization's repositories, you can use the props. filter to filter by custom property.