Skip to content
Mermaid templates

Transformer Block

Pre-norm transformer block with multi-head attention + MLP and residuals.

Template previewMermaid
Rendering…

Make it your own.

flowchart TD
X[Input embeddings] --> N1[LayerNorm]
N1 --> Q[Q proj]
N1 --> K[K proj]
N1 --> V[V proj]
Q --> Attn[Multi-head\nself-attention]
K --> Attn
V --> Attn
Attn --> Add1((+))
X --> Add1
Add1 --> N2[LayerNorm]
N2 --> FC1[Linear up]
FC1 --> GELU[GELU]
GELU --> FC2[Linear down]
FC2 --> Add2((+))
Add1 --> Add2
Add2 --> Y[Output]
classDef op fill:#dbeafe,stroke:#1e3a8a;
classDef act fill:#fce7f3,stroke:#9d174d;
class Attn,FC1,FC2,Q,K,V op
class GELU act