DEV Community

#llamacpp

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
llama.cpp vs Ollama in 2026: Which Runtime Should You Run?

llama.cpp vs Ollama in 2026: Which Runtime Should You Run?

Comments
18 min read
How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM

How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM

Comments
5 min read
llama.cpp vs Ollama: Which Should You Run in 2026?

llama.cpp vs Ollama: Which Should You Run in 2026?

Comments
5 min read
Gemma 4 on an 2021 4 GB Laptop GPU: QAT Takes It From 9.5 GiB to 1.6

Gemma 4 on an 2021 4 GB Laptop GPU: QAT Takes It From 9.5 GiB to 1.6

10
Comments 3
13 min read
ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide

ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide

Comments 2
20 min read
KV Cache on 16 GB GPUs: Making Long Context Actually Fit

KV Cache on 16 GB GPUs: Making Long Context Actually Fit

Comments 1
22 min read
Running llama.cpp on a 32 GB MacBook Air: A Direct Comparison with Ollama

Running llama.cpp on a 32 GB MacBook Air: A Direct Comparison with Ollama

Comments 1
9 min read
Fix Local LLM Quality: Context Stacking & Rope Freq Tweaks

Fix Local LLM Quality: Context Stacking & Rope Freq Tweaks

Comments
8 min read
Two Rust Clients for Gemma 4: Calling the Endpoint vs. Calling the MCP Server 🦀

Two Rust Clients for Gemma 4: Calling the Endpoint vs. Calling the MCP Server 🦀

8
Comments 1
17 min read
"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window

"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window

Comments
10 min read
Nine ways to talk to a local model

Nine ways to talk to a local model

Comments
9 min read
Can Qwen 3.8 running on your laptop really replace Claude Opus for Agentic coding?

Can Qwen 3.8 running on your laptop really replace Claude Opus for Agentic coding?

7
Comments 3
20 min read
My inference server decided my second GPU no longer exists. Here is how I got it back without upgrading a driver.

My inference server decided my second GPU no longer exists. Here is how I got it back without upgrading a driver.

Comments 1
3 min read
GnLOLot Releases MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF for Enhanced Local AI Development

GnLOLot Releases MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF for Enhanced Local AI Development

Comments
3 min read
Running a 1.5B-Parameter LLM Entirely On-Device for Mental Health — The NilaMind Architecture

Running a 1.5B-Parameter LLM Entirely On-Device for Mental Health — The NilaMind Architecture

Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.