DEV Community

#pytorch

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
I Trained a 6.4M-Parameter Transformer From Scratch to Talk About Recipes

I Trained a 6.4M-Parameter Transformer From Scratch to Talk About Recipes

1
Comments
5 min read
Why My Medical AI Took 6.4 Seconds Per Scan and How I Got It to 3.1.

Why My Medical AI Took 6.4 Seconds Per Scan and How I Got It to 3.1.

Comments 2
6 min read
RoPE: How 2D Rotations Solved Transformer Long-Context

RoPE: How 2D Rotations Solved Transformer Long-Context

1
Comments
4 min read
Running Qwen3 Through the ExecuTorch MLX Delegate: Up to 4.52x Faster on M1 Max

Running Qwen3 Through the ExecuTorch MLX Delegate: Up to 4.52x Faster on M1 Max

Comments
7 min read
Testing PyTorch 2.13 MPS FlexAttention on M1 Max: Up to 7.83x Faster for Sparse Attention

Testing PyTorch 2.13 MPS FlexAttention on M1 Max: Up to 7.83x Faster for Sparse Attention

Comments
7 min read
What Does `unsqueeze` Do in PyTorch? (And Why Your Model Keeps Asking For It)

What Does `unsqueeze` Do in PyTorch? (And Why Your Model Keeps Asking For It)

Comments
7 min read
PyTorch Broadcasting Explained: The 3 Rules (and the Silent Bug That Bites Everyone)

PyTorch Broadcasting Explained: The 3 Rules (and the Silent Bug That Bites Everyone)

1
Comments
4 min read
Debugging a Python "Memory Leak" That Was Actually a Measurement Bug (ru_maxrss vs VmRSS)

Debugging a Python "Memory Leak" That Was Actually a Measurement Bug (ru_maxrss vs VmRSS)

Comments
4 min read
Classifier-free guidance above 7.5 oversaturated our product renders

Classifier-free guidance above 7.5 oversaturated our product renders

1
Comments
4 min read
Using the channels-last memory format reduced the latency of our conversation backbone by 22%

Using the channels-last memory format reduced the latency of our conversation backbone by 22%

1
Comments
4 min read
The SDXL VAE overflow that decoded black images in fp16

The SDXL VAE overflow that decoded black images in fp16

1
Comments
4 min read
Data Science Workload: Giới hạn RAM trên Dell Pro Max 14 MC14250

Data Science Workload: Giới hạn RAM trên Dell Pro Max 14 MC14250

Comments
3 min read
The seam our tiled upscaler left on every 4K product render

The seam our tiled upscaler left on every 4K product render

Comments
4 min read
Perplexity held flat after INT4. Task accuracy dropped 7 points.

Perplexity held flat after INT4. Task accuracy dropped 7 points.

Comments
4 min read
Developer Take On: A High-Resolution Neural Cellular Automata

Developer Take On: A High-Resolution Neural Cellular Automata

Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.