DEV Community

#vllm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
KV Cache Quantization in LLM Serving: FP8 and INT8 Tradeoffs, the Silent config.json Trap, and How to Measure It Fairly

KV Cache Quantization in LLM Serving: FP8 and INT8 Tradeoffs, the Silent config.json Trap, and How to Measure It Fairly

Comments
7 min read
What Does It Actually Cost to Self-Host an LLM? The Batching Math Nobody Shows You

What Does It Actually Cost to Self-Host an LLM? The Batching Math Nobody Shows You

Comments
6 min read
Five Ways Your LLM Serving Benchmark Is Lying to You (and How to Catch Each One)

Five Ways Your LLM Serving Benchmark Is Lying to You (and How to Catch Each One)

Comments 1
15 min read
vLLM's weight cache can serve another checkpoint's weights when the tensor layout matches

vLLM's weight cache can serve another checkpoint's weights when the tensor layout matches

Comments
4 min read
vLLM ignores LoRA rank_pattern and alpha_pattern and serves the adapter at the wrong scale

vLLM ignores LoRA rank_pattern and alpha_pattern and serves the adapter at the wrong scale

Comments
5 min read
Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default

Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default

Comments
4 min read
Evaluating Speculative Decoding in vLLM on AMD MI300X GPUs

Evaluating Speculative Decoding in vLLM on AMD MI300X GPUs

Comments
1 min read
Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4

Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4

1
Comments
11 min read
Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price

Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price

1
Comments
8 min read
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Comments
9 min read
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

1
Comments
9 min read
Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Comments
13 min read
KV Cache on 16 GB GPUs: Making Long Context Actually Fit

KV Cache on 16 GB GPUs: Making Long Context Actually Fit

Comments 1
22 min read
Inside vLLM: Following One Request from the API to GPU Execution

Inside vLLM: Following One Request from the API to GPU Execution

1
Comments 2
24 min read
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

7
Comments
9 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.