DEV Community

#vllm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16

Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16

1
Comments
9 min read
Serving Gemma 4 on an AMD MI300X: What $1.99 an Hour Buys

Serving Gemma 4 on an AMD MI300X: What $1.99 an Hour Buys

11
Comments 5
9 min read
Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Comments
13 min read
KV Cache on 16 GB GPUs: Making Long Context Actually Fit

KV Cache on 16 GB GPUs: Making Long Context Actually Fit

Comments 1
22 min read
From adapter to deployment: merging LoRA weights and serving with vLLM or a Space

From adapter to deployment: merging LoRA weights and serving with vLLM or a Space

Comments
3 min read
Inside vLLM: Following One Request from the API to GPU Execution

Inside vLLM: Following One Request from the API to GPU Execution

1
Comments 2
24 min read
Deploying the 600GB Inkling-NVFP4 Model on Spot A3: A GKE and vLLM Deep Dive

Deploying the 600GB Inkling-NVFP4 Model on Spot A3: A GKE and vLLM Deep Dive

4
Comments
61 min read
vLLM v0.28.0: the breaking change small GPU users must read

vLLM v0.28.0: the breaking change small GPU users must read

Comments
5 min read
Installing Rust for vLLM on Graviton: a G5g walk-through 🦀

Installing Rust for vLLM on Graviton: a G5g walk-through 🦀

Comments
10 min read
FastMCP Is Now MCPServer: Migrating a Python MCP Server to the MCP SDK 2.x

Dependency upgrades break shared environments

FastMCP Is Now MCPServer: Migrating a Python MCP Server to the MCP SDK 2.x

13
Comments 10
12 min read
Ollama to vLLM: When to Migrate Your Local LLM Server

Ollama to vLLM: When to Migrate Your Local LLM Server

Comments
15 min read
Deploying Inference Using NVIDIA Dynamo and vLLM

Deploying Inference Using NVIDIA Dynamo and vLLM

12
Comments
8 min read
The unofficial TPU migration guide: Cloud TPU API to Compute Engine

The unofficial TPU migration guide: Cloud TPU API to Compute Engine

6
Comments 2
17 min read
The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

3
Comments
11 min read
Self-Hosted Gemma 4 on TPU v6e: Deployment & SRE with Antigravity

Self-Hosted Gemma 4 on TPU v6e: Deployment & SRE with Antigravity

Comments
8 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.