FlashAttention: 把 注意力 的 瓶颈 从算力挪 回访存
注意力慢在访存,不在算力。FlashAttention 把 Q/K/V 分块搬进 SRAM,用在线 softmax 保证分块结果与整行 softmax 逐位相等,N×N 中间矩阵从不落地显存。
Inference optimisation6 min · 2.5K characters
Topic index
1 post
注意力慢在访存,不在算力。FlashAttention 把 Q/K/V 分块搬进 SRAM,用在线 softmax 保证分块结果与整行 softmax 逐位相等,N×N 中间矩阵从不落地显存。