DEV Community

#gpu

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Same nvJPEG2000, different numbers: timer boundaries and frames in flight

Same nvJPEG2000, different numbers: timer boundaries and frames in flight

Comments
13 min read
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture

Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture

1
Comments
10 min read
DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)

DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)

Comments
10 min read
Half of my FP8 convolutions were silently running in f32 – a one-line XLA fix

Half of my FP8 convolutions were silently running in f32 – a one-line XLA fix

2
Comments 1
6 min read
What Happens When You Ask an LLM a Question

What Happens When You Ask an LLM a Question

Comments
8 min read
CUDA Cores vs Tensor Cores Explained

CUDA Cores vs Tensor Cores Explained

Comments 1
2 min read
Inside SFPU Overflow Bugs: How a 40-Year-Old Rounding Trick Breaks on Modern AI Accelerators

Inside SFPU Overflow Bugs: How a 40-Year-Old Rounding Trick Breaks on Modern AI Accelerators

Comments
3 min read
I Put a Paid AI Video Generator on My Own Gaming GPU — No Cloud Bill

I Put a Paid AI Video Generator on My Own Gaming GPU — No Cloud Bill

Comments 1
3 min read
Reverse-Engineering NVIDIA: Modifying a CUDA binary

Reverse-Engineering NVIDIA: Modifying a CUDA binary

Comments
4 min read
From API to GPU, Week 6 (Part 2): Watching a Neural Network Learn

From API to GPU, Week 6 (Part 2): Watching a Neural Network Learn

Comments
17 min read
From API to GPU, Week 6 (Part 1): A Model That Predicts, and How Wrong It Is

From API to GPU, Week 6 (Part 1): A Model That Predicts, and How Wrong It Is

Comments
13 min read
Scale Before the Spike: Predictive Autoscaling for GPU Workloads on Kubernetes

Scale Before the Spike: Predictive Autoscaling for GPU Workloads on Kubernetes

Comments
6 min read
Ollama Not Using GPU? Fix It on Linux, Windows and WSL

Ollama Not Using GPU? Fix It on Linux, Windows and WSL

Comments
6 min read
DGX Spark (GB10) memory sizing for LLM serving: the numbers

DGX Spark (GB10) memory sizing for LLM serving: the numbers

Comments
7 min read
A 4 GB Laptop GPU Beats a 12-Core CPU by 4.3x on Gemma 4

A 4 GB Laptop GPU Beats a 12-Core CPU by 4.3x on Gemma 4

15
Comments 6
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.