Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
inference
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Laya is a 421M open-weights answer to Jev
techaiwire
techaiwire
techaiwire
Follow
Sep 19
Laya is a 421M open-weights answer to Jev
#
llm
#
openweights
#
inference
#
benchmarks
5
 reactions
Comments
Add Comment
4 min read
Fujitsu MONAKA ships 144 Armv9 cores in November
techaiwire
techaiwire
techaiwire
Follow
Sep 18
Fujitsu MONAKA ships 144 Armv9 cores in November
#
hardware
#
linux
#
inference
#
datacenters
5
 reactions
Comments
Add Comment
4 min read
Inside vLLM: Following One Request from the API to GPU Execution
yuan lei
yuan lei
yuan lei
Follow
Sep 7
Inside vLLM: Following One Request from the API to GPU Execution
#
vllm
#
llm
#
inference
#
python
1
 reaction
Comments
2
 comments
24 min read
DGX Spark (GB10) memory sizing for LLM serving: the numbers
Jahn
Jahn
Jahn
Follow
Sep 2
DGX Spark (GB10) memory sizing for LLM serving: the numbers
#
nvidia
#
llm
#
inference
#
gpu
Comments
Add Comment
7 min read
On-Device AI in Kotlin
pielouNW
pielouNW
pielouNW
Follow
Sep 2
On-Device AI in Kotlin
#
ai
#
kotlin
#
llm
#
inference
Comments
Add Comment
7 min read
Name the Blackwell serving cell you are actually in
Jahn
Jahn
Jahn
Follow
Sep 9
Name the Blackwell serving cell you are actually in
#
nvidia
#
gpu
#
llm
#
inference
1
 reaction
Comments
1
 comment
5 min read
Training vs Inference: Why Building Costs Millions and Asking Costs Cents
Internals Decoded
Internals Decoded
Internals Decoded
Follow
Aug 25
Training vs Inference: Why Building Costs Millions and Asking Costs Cents
#
ai
#
training
#
inference
#
compute
Comments
Add Comment
9 min read
I Got 28 TPS Out of Free Kaggle GPUs. Here's What It Took.
Aditya Raut
Aditya Raut
Aditya Raut
Follow
Aug 23
I Got 28 TPS Out of Free Kaggle GPUs. Here's What It Took.
#
ai
#
gpu
#
inference
#
python
Comments
Add Comment
5 min read
Speculative Decoding and MTP: Why Guessing Is Free
Jessie Jia
Jessie Jia
Jessie Jia
Follow
Aug 20
Speculative Decoding and MTP: Why Guessing Is Free
#
ai
#
technical
#
inference
#
mtp
Comments
Add Comment
6 min read
what a turn actually costs me
Saltorious
Saltorious
Saltorious
Follow
Aug 13
what a turn actually costs me
#
engineering
#
inference
#
localmodels
Comments
Add Comment
2 min read
AMD's Move on Weight Storage: The Taalas Bet
Peremptory
Peremptory
Peremptory
Follow
Aug 12
AMD's Move on Weight Storage: The Taalas Bet
#
amd
#
compute
#
inference
#
aiinfrastructure
Comments
Add Comment
2 min read
A 4B model on a 6GB laptop beat Claude Opus on our 440K-token corpus. The fix was giving the model less to do.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Aug 24
A 4B model on a 6GB laptop beat Claude Opus on our 440K-token corpus. The fix was giving the model less to do.
#
privateai
#
llm
#
inference
#
localllm
1
 reaction
Comments
1
 comment
4 min read
AMD Bets Weight Storage Is the Real Bottleneck
Peremptory
Peremptory
Peremptory
Follow
Aug 11
AMD Bets Weight Storage Is the Real Bottleneck
#
amd
#
hardware
#
aiinfrastructure
#
inference
Comments
Add Comment
2 min read
vLLM reinvented the operating system, and nobody told you
Yathiskumar
Yathiskumar
Yathiskumar
Follow
Aug 11
vLLM reinvented the operating system, and nobody told you
#
systemdesign
#
llm
#
inference
#
operatingsystems
1
 reaction
Comments
Add Comment
12 min read
Inverse Problems: Why Predicting Backward Is Harder Than It Looks
zeromathai
zeromathai
zeromathai
Follow
Sep 6
Inverse Problems: Why Predicting Backward Is Harder Than It Looks
#
machinelearning
#
generativeai
#
inference
#
deeplearning
Comments
Add Comment
6 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account