Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
llamacpp
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
llama.cpp vs Ollama in 2026: Which Runtime Should You Run?
Rost
Rost
Rost
Follow
Sep 14
llama.cpp vs Ollama in 2026: Which Runtime Should You Run?
#
llamacpp
#
ollama
#
llm
#
selfhosting
Comments
Add Comment
18 min read
How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM
Mr Say Nothing
Mr Say Nothing
Mr Say Nothing
Follow
Sep 18
How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM
#
ai
#
gguf
#
ollama
#
llamacpp
Comments
Add Comment
5 min read
llama.cpp vs Ollama: Which Should You Run in 2026?
Mr Say Nothing
Mr Say Nothing
Mr Say Nothing
Follow
Sep 18
llama.cpp vs Ollama: Which Should You Run in 2026?
#
ai
#
llamacpp
#
ollama
#
localllm
Comments
Add Comment
5 min read
Gemma 4 on an 2021 4 GB Laptop GPU: QAT Takes It From 9.5 GiB to 1.6
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 10
Gemma 4 on an 2021 4 GB Laptop GPU: QAT Takes It From 9.5 GiB to 1.6
#
gemma
#
llamacpp
#
mcp
#
cuda
10
 reactions
Comments
3
 comments
13 min read
ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide
Rost
Rost
Rost
Follow
Sep 12
ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide
#
llm
#
selfhosting
#
llamacpp
#
ollama
Comments
2
 comments
20 min read
KV Cache on 16 GB GPUs: Making Long Context Actually Fit
Rost
Rost
Rost
Follow
Sep 11
KV Cache on 16 GB GPUs: Making Long Context Actually Fit
#
llm
#
llamacpp
#
vllm
#
ollama
Comments
1
 comment
22 min read
Running llama.cpp on a 32 GB MacBook Air: A Direct Comparison with Ollama
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 7
Running llama.cpp on a 32 GB MacBook Air: A Direct Comparison with Ollama
#
ai
#
localllm
#
llamacpp
#
ollama
Comments
1
 comment
9 min read
Fix Local LLM Quality: Context Stacking & Rope Freq Tweaks
Umair Bilal
Umair Bilal
Umair Bilal
Follow
Aug 23
Fix Local LLM Quality: Context Stacking & Rope Freq Tweaks
#
localllms
#
aiagents
#
ollama
#
llamacpp
Comments
Add Comment
8 min read
Two Rust Clients for Gemma 4: Calling the Endpoint vs. Calling the MCP Server 🦀
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 11
Two Rust Clients for Gemma 4: Calling the Endpoint vs. Calling the MCP Server 🦀
#
rust
#
mcp
#
gemma
#
llamacpp
8
 reactions
Comments
1
 comment
17 min read
"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window
Jasur Yuldoshev
Jasur Yuldoshev
Jasur Yuldoshev
Follow
Aug 21
"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window
#
llamacpp
#
llm
#
performance
#
debugging
Comments
Add Comment
10 min read
Nine ways to talk to a local model
the kilted dev
the kilted dev
the kilted dev
Follow
Aug 11
Nine ways to talk to a local model
#
localllm
#
ollama
#
llamacpp
#
buildinpublic
Comments
Add Comment
9 min read
Can Qwen 3.8 running on your laptop really replace Claude Opus for Agentic coding?
Deepu K Sasidharan
Deepu K Sasidharan
Deepu K Sasidharan
Follow
Sep 11
Can Qwen 3.8 running on your laptop really replace Claude Opus for Agentic coding?
#
ai
#
llamacpp
#
localllm
#
qwen38
7
 reactions
Comments
3
 comments
20 min read
My inference server decided my second GPU no longer exists. Here is how I got it back without upgrading a driver.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Aug 24
My inference server decided my second GPU no longer exists. Here is how I got it back without upgrading a driver.
#
gpu
#
llamacpp
#
ollama
#
homelab
Comments
1
 comment
3 min read
GnLOLot Releases MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF for Enhanced Local AI Development
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 18
GnLOLot Releases MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF for Enhanced Local AI Development
#
gguf
#
llamacpp
#
quantized
#
minicpm5
Comments
Add Comment
3 min read
Running a 1.5B-Parameter LLM Entirely On-Device for Mental Health — The NilaMind Architecture
sampathmannam
sampathmannam
sampathmannam
Follow
Jul 14
Running a 1.5B-Parameter LLM Entirely On-Device for Mental Health — The NilaMind Architecture
#
llamacpp
#
opensource
#
android
Comments
Add Comment
5 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account