AI-Native & Cloud-Native FS: A high-performance file semantic layer for cloud object storage, integrated with high-speed cache. CNCF Sandbox Project.
-
Updated
Jul 25, 2026 - Rust
AI-Native & Cloud-Native FS: A high-performance file semantic layer for cloud object storage, integrated with high-speed cache. CNCF Sandbox Project.
Production inference for encoder models - ColBERT, GLiNER, ColPali, embeddings etc. - as vLLM plugins for online and in-process deployment
Prefix-Aware Attention for LLM Decoding
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
A curated list of plugins built on top of vLLM
独立、可单独安装的 vLLM KV-cache 池 + 空闲队列可视化插件。 与具体缓存方案解耦,自动适配: 三区 (ThreePhaseBlockQueue, vllm-kv-cache-plugin) — Cold / Warm / Hot 双区 (TwoPhaseBlockQueue, kv_cache_affinity) — Aged / Fresh 原生 vLLM 队列兜底 — 单 Free 区
A manager to load vllm plugins without rebuilding image for each new plugin.
vLLM Plugins for additional features like decoding strategies, monitoring, models etc
Add a description, image, and links to the vllm-plugins topic page so that developers can more easily learn about it.
To associate your repository with the vllm-plugins topic, visit your repo's landing page and select "manage topics."