Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
localllm
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
How I upgrade xiaoai speaker local llm: Sub-200ms AI
Umair Bilal
Umair Bilal
Umair Bilal
Follow
Sep 14
How I upgrade xiaoai speaker local llm: Sub-200ms AI
#
aiagents
#
localllm
#
iot
#
node
Comments
1
 comment
10 min read
Best Local LLM for Coding: 8GB to 24GB VRAM Picks
Mr Say Nothing
Mr Say Nothing
Mr Say Nothing
Follow
Sep 18
Best Local LLM for Coding: 8GB to 24GB VRAM Picks
#
ai
#
localllm
#
coding
#
ollama
Comments
Add Comment
4 min read
llama.cpp vs Ollama: Which Should You Run in 2026?
Mr Say Nothing
Mr Say Nothing
Mr Say Nothing
Follow
Sep 18
llama.cpp vs Ollama: Which Should You Run in 2026?
#
ai
#
llamacpp
#
ollama
#
localllm
Comments
Add Comment
5 min read
Ollama vs LM Studio: Which Local LLM Tool Should You Use?
Mr Say Nothing
Mr Say Nothing
Mr Say Nothing
Follow
Sep 18
Ollama vs LM Studio: Which Local LLM Tool Should You Use?
#
ai
#
ollama
#
lmstudio
#
localllm
Comments
Add Comment
4 min read
What Happens When You Ask an LLM a Question
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 14
What Happens When You Ask an LLM a Question
#
ai
#
localllm
#
llmbasics
#
gpu
Comments
Add Comment
8 min read
Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 14
Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU
#
ai
#
localllm
#
vllm
#
kubernetes
Comments
Add Comment
13 min read
I Ran DeepSeek V4 Flash Across Two DGX Sparks Over Ethernet
Kevin Tang
Kevin Tang
Kevin Tang
Follow
Sep 6
I Ran DeepSeek V4 Flash Across Two DGX Sparks Over Ethernet
#
ai
#
hardware
#
localllm
#
networking
Comments
Add Comment
11 min read
VRAM and RAM for local LLMs — honest planning bands, not a GPU tier list
Tech-Gurunomics
Tech-Gurunomics
Tech-Gurunomics
Follow
Sep 6
VRAM and RAM for local LLMs — honest planning bands, not a GPU tier list
#
localllm
#
llm
#
ollama
#
hardware
Comments
Add Comment
4 min read
Running Ollama on a 32 GB MacBook Air: A Practical First Setup
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 7
Running Ollama on a 32 GB MacBook Air: A Practical First Setup
#
ai
#
localllm
#
ollama
#
applesilicon
Comments
Add Comment
6 min read
Running llama.cpp on a 32 GB MacBook Air: A Direct Comparison with Ollama
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 7
Running llama.cpp on a 32 GB MacBook Air: A Direct Comparison with Ollama
#
ai
#
localllm
#
llamacpp
#
ollama
Comments
1
 comment
9 min read
Running a 35B MoE Model on an 8 GB Laptop GPU: Testing FreeToken
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 7
Running a 35B MoE Model on an 8 GB Laptop GPU: Testing FreeToken
#
ai
#
localllm
#
freetoken
#
gpu
Comments
3
 comments
7 min read
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.
#
llm
#
benchmarks
#
localllm
#
reproducibility
1
 reaction
Comments
1
 comment
3 min read
I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.
#
llm
#
prompting
#
debugging
#
localllm
1
 reaction
Comments
Add Comment
3 min read
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.
#
llm
#
benchmarks
#
localllm
#
privateai
2
 reactions
Comments
Add Comment
4 min read
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number
#
llm
#
benchmarks
#
localllm
#
agents
1
 reaction
Comments
Add Comment
4 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account