Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
gpu
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Same nvJPEG2000, different numbers: timer boundaries and frames in flight
Fyodor Serzhenko
Fyodor Serzhenko
Fyodor Serzhenko
Follow
Sep 18
Same nvJPEG2000, different numbers: timer boundaries and frames in flight
#
cuda
#
gpu
#
performance
#
cpp
Comments
Add Comment
13 min read
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture
ComradePenguin
ComradePenguin
ComradePenguin
Follow
Sep 18
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture
#
cuda
#
nvidia
#
gpu
1
 reaction
Comments
Add Comment
10 min read
DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)
Sho Tanaka (tsho)
Sho Tanaka (tsho)
Sho Tanaka (tsho)
Follow
Sep 17
DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)
#
machinelearning
#
pytorch
#
deepspeed
#
gpu
Comments
Add Comment
10 min read
Half of my FP8 convolutions were silently running in f32 – a one-line XLA fix
Yehor Cherednichenko
Yehor Cherednichenko
Yehor Cherednichenko
Follow
Sep 17
Half of my FP8 convolutions were silently running in f32 – a one-line XLA fix
#
machinelearning
#
performance
#
compilers
#
gpu
2
 reactions
Comments
1
 comment
6 min read
What Happens When You Ask an LLM a Question
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 14
What Happens When You Ask an LLM a Question
#
ai
#
localllm
#
llmbasics
#
gpu
Comments
Add Comment
8 min read
CUDA Cores vs Tensor Cores Explained
Sumukh Shenoy
Sumukh Shenoy
Sumukh Shenoy
Follow
Sep 14
CUDA Cores vs Tensor Cores Explained
#
gpu
#
machinelearning
#
deeplearning
#
beginners
Comments
1
 comment
2 min read
Inside SFPU Overflow Bugs: How a 40-Year-Old Rounding Trick Breaks on Modern AI Accelerators
stmanst
stmanst
stmanst
Follow
Sep 9
Inside SFPU Overflow Bugs: How a 40-Year-Old Rounding Trick Breaks on Modern AI Accelerators
#
ai
#
machinelearning
#
gpu
#
debugging
Comments
Add Comment
3 min read
I Put a Paid AI Video Generator on My Own Gaming GPU — No Cloud Bill
erniou86
erniou86
erniou86
Follow
Sep 9
I Put a Paid AI Video Generator on My Own Gaming GPU — No Cloud Bill
#
ai
#
gpu
#
selfhosting
#
machinelearning
Comments
1
 comment
3 min read
Reverse-Engineering NVIDIA: Modifying a CUDA binary
Stjepan
Stjepan
Stjepan
Follow
Sep 6
Reverse-Engineering NVIDIA: Modifying a CUDA binary
#
nvidia
#
gpu
#
cuda
#
hex
Comments
Add Comment
4 min read
From API to GPU, Week 6 (Part 2): Watching a Neural Network Learn
Dinesh Kumar Ramasamy
Dinesh Kumar Ramasamy
Dinesh Kumar Ramasamy
Follow
Sep 5
From API to GPU, Week 6 (Part 2): Watching a Neural Network Learn
#
ai
#
llm
#
gpu
#
machinelearning
Comments
Add Comment
17 min read
From API to GPU, Week 6 (Part 1): A Model That Predicts, and How Wrong It Is
Dinesh Kumar Ramasamy
Dinesh Kumar Ramasamy
Dinesh Kumar Ramasamy
Follow
Sep 5
From API to GPU, Week 6 (Part 1): A Model That Predicts, and How Wrong It Is
#
ai
#
llm
#
gpu
#
machinelearning
Comments
Add Comment
13 min read
Scale Before the Spike: Predictive Autoscaling for GPU Workloads on Kubernetes
Ramkumar Nagaraj
Ramkumar Nagaraj
Ramkumar Nagaraj
Follow
Sep 4
Scale Before the Spike: Predictive Autoscaling for GPU Workloads on Kubernetes
#
kubernetes
#
autoscaling
#
gpu
#
devops
Comments
Add Comment
6 min read
Ollama Not Using GPU? Fix It on Linux, Windows and WSL
Mr Say Nothing
Mr Say Nothing
Mr Say Nothing
Follow
Sep 16
Ollama Not Using GPU? Fix It on Linux, Windows and WSL
#
ollama
#
gpu
#
linux
#
windows
Comments
Add Comment
6 min read
DGX Spark (GB10) memory sizing for LLM serving: the numbers
Jahn
Jahn
Jahn
Follow
Sep 2
DGX Spark (GB10) memory sizing for LLM serving: the numbers
#
nvidia
#
llm
#
inference
#
gpu
Comments
Add Comment
7 min read
A 4 GB Laptop GPU Beats a 12-Core CPU by 4.3x on Gemma 4
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 16
A 4 GB Laptop GPU Beats a 12-Core CPU by 4.3x on Gemma 4
#
machinelearning
#
gpu
#
benchmarking
#
python
15
 reactions
Comments
6
 comments
6 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account