Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
vllm
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
KV Cache Quantization in LLM Serving: FP8 and INT8 Tradeoffs, the Silent config.json Trap, and How to Measure It Fairly
AI Tech News
AI Tech News
AI Tech News
Follow
Oct 6
KV Cache Quantization in LLM Serving: FP8 and INT8 Tradeoffs, the Silent config.json Trap, and How to Measure It Fairly
#
llm
#
vllm
#
performance
#
devops
Comments
Add Comment
7 min read
What Does It Actually Cost to Self-Host an LLM? The Batching Math Nobody Shows You
AI Tech News
AI Tech News
AI Tech News
Follow
Oct 6
What Does It Actually Cost to Self-Host an LLM? The Batching Math Nobody Shows You
#
llm
#
vllm
#
performance
#
devops
Comments
Add Comment
6 min read
Five Ways Your LLM Serving Benchmark Is Lying to You (and How to Catch Each One)
AI Tech News
AI Tech News
AI Tech News
Follow
Oct 5
Five Ways Your LLM Serving Benchmark Is Lying to You (and How to Catch Each One)
#
llm
#
vllm
#
performance
#
devops
Comments
1
 comment
15 min read
vLLM's weight cache can serve another checkpoint's weights when the tensor layout matches
The Homelab Postmortem
The Homelab Postmortem
The Homelab Postmortem
Follow
Oct 3
vLLM's weight cache can serve another checkpoint's weights when the tensor layout matches
#
llm
#
vllm
#
selfhosted
#
debugging
Comments
Add Comment
4 min read
vLLM ignores LoRA rank_pattern and alpha_pattern and serves the adapter at the wrong scale
The Homelab Postmortem
The Homelab Postmortem
The Homelab Postmortem
Follow
Oct 3
vLLM ignores LoRA rank_pattern and alpha_pattern and serves the adapter at the wrong scale
#
llm
#
vllm
#
lora
#
selfhosted
Comments
Add Comment
5 min read
Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default
Reno Lu
Reno Lu
Reno Lu
Follow
Oct 2
Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default
#
vllm
#
localllm
#
speculativedecoding
#
quantization
Comments
Add Comment
4 min read
Evaluating Speculative Decoding in vLLM on AMD MI300X GPUs
ram.mehta.1899@gmail.com
ram.mehta.1899@gmail.com
ram.mehta.1899@gmail.com
Follow
Oct 1
Evaluating Speculative Decoding in vLLM on AMD MI300X GPUs
#
vllm
#
amd
#
mi300x
#
rocm
Comments
Add Comment
1 min read
Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4
xbill
xbill
xbill
Follow
for
AWS Community Builders
Sep 30
Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4
#
aws
#
sagemaker
#
gemma
#
vllm
1
 reaction
Comments
Add Comment
11 min read
Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price
xbill
xbill
xbill
Follow
for
AWS Community Builders
Sep 30
Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price
#
aws
#
sagemaker
#
gemma
#
vllm
1
 reaction
Comments
Add Comment
8 min read
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4
xbill
xbill
xbill
Follow
for
AWS Community Builders
Sep 25
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4
#
aws
#
sagemaker
#
gemma
#
vllm
Comments
Add Comment
9 min read
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers
xbill
xbill
xbill
Follow
for
AWS Community Builders
Sep 30
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers
#
aws
#
sagemaker
#
gemma
#
vllm
1
 reaction
Comments
Add Comment
9 min read
Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 14
Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU
#
ai
#
localllm
#
vllm
#
kubernetes
Comments
Add Comment
13 min read
KV Cache on 16 GB GPUs: Making Long Context Actually Fit
Rost
Rost
Rost
Follow
Sep 11
KV Cache on 16 GB GPUs: Making Long Context Actually Fit
#
llm
#
llamacpp
#
vllm
#
ollama
Comments
1
 comment
22 min read
Inside vLLM: Following One Request from the API to GPU Execution
yuan lei
yuan lei
yuan lei
Follow
Sep 7
Inside vLLM: Following One Request from the API to GPU Execution
#
vllm
#
llm
#
inference
#
python
1
 reaction
Comments
2
 comments
24 min read
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 26
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4
#
aws
#
sagemaker
#
gemma
#
vllm
7
 reactions
Comments
Add Comment
9 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account