DEV Community

#transformers

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Detectar volatilidad anómala en los mercados con un Temporal Fusion Transformer

Detectar volatilidad anómala en los mercados con un Temporal Fusion Transformer

Comments
2 min read
Puzzle Solution Revealed - Transformer: Need for Position Embedding

Puzzle Solution Revealed - Transformer: Need for Position Embedding

Comments
13 min read
Masked Self-Attention, Explained Through Avengers: Endgame

Masked Self-Attention, Explained Through Avengers: Endgame

1
Comments
6 min read
Depth-Attention: Opening an Attention Channel Between Transformer Layers

Depth-Attention: Opening an Attention Channel Between Transformer Layers

Comments
11 min read
Top 5 Dev Tools Released in Late July 2026

Top 5 Dev Tools Released in Late July 2026

Comments
3 min read
How a Baseten Engineer Traced 7 Years of Attention Mechanism Evolution -- From GPT-2 to Kimi K3, in Runable PyTorch

How a Baseten Engineer Traced 7 Years of Attention Mechanism Evolution -- From GPT-2 to Kimi K3, in Runable PyTorch

2
Comments 4
7 min read
Attention Is All You Need: The Translation Problem That Led to ChatGPT

Attention Is All You Need: The Translation Problem That Led to ChatGPT

Comments 2
8 min read
The KV Cache Is the Bottleneck: A 2026 Field Guide to Attention Variants

The KV Cache Is the Bottleneck: A 2026 Field Guide to Attention Variants

1
Comments 1
4 min read
Self-attention, explained without the heavy math

Self-attention, explained without the heavy math

3
Comments
3 min read
How a Transformer Plays Tic-Tac-Toe

How a Transformer Plays Tic-Tac-Toe

Comments
1 min read
How Modern Transformer Blocks Work — From RMSNorm to MoE

How Modern Transformer Blocks Work — From RMSNorm to MoE

Comments
5 min read
Why Positional Embeddings Matter — APE, RPE, and RoPE Explained for Developers

Why Positional Embeddings Matter — APE, RPE, and RoPE Explained for Developers

Comments
5 min read
🧠 人工智能发展方向:当前是否到头?

🧠 人工智能发展方向:当前是否到头?

Comments
1 min read
Why KV Cache Matters — How MQA, GQA, and MLA Make LLM Inference Faster

Why KV Cache Matters — How MQA, GQA, and MLA Make LLM Inference Faster

Comments
5 min read
Why Attention Becomes the Bottleneck — And How Efficient Attention Fixes It

Why Attention Becomes the Bottleneck — And How Efficient Attention Fixes It

Comments
3 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.