š Day 1 of #OpenSourceWeek: FlashMLA Honored to share FlashMLA - our efficient MLA decoding kernel for Hopper GPUs, optimized for variable-length sequences and now in production. ā BF16 support ā Paged KV cache (block size 64) ā” 3000 GB/s memory-bound & 580 TFLOPS compute-bound on H800 š Explore on GitHub: t.co/4JvJTn5HX2
- Views
- 1,714,030
- Likes
- 10,057
- Replies
- 547
- Reposts
- 1,308
- Quotes
- 332
- Bookmarks
- 1,637
š Day 1 of : FlashMLA
The post introduces FlashMLA as an efficient decoding kernel released during OpenSourceWeek. It highlights supported features, production status, benchmark results, and a GitHub link for further exploration.
- Views
- 1.7M
- Likes
- 10.1K
- Replies
- 547
- Reposts
- 1.3K
- Quotes
- 332
- Bookmarks
- 1.6K
- Topics
Analysis
The post introduces FlashMLA as an efficient decoding kernel released during OpenSourceWeek. It highlights supported features, production status, benchmark results, and a GitHub link for further exploration.
Sign in to see the full analysis.
Sign inš Day 1 of #OpenSourceWeek: FlashMLA Honored to share FlashMLA - our efficient MLA decoding kernel for Hopper GPUs, optimized for variable-length sequences and now in production. ā BF16 support ā Paged KV cache (block size 64) ā” 3000 GB/s memory-bound & 580 TFLOPS compute-bound on H800 š Explore on GitHub: https://t.co/4JvJTn5HX2
š Day 1 of #OpenSourceWeek: FlashMLA
- 01The Event Hook
Open with the first day of an event and name the featured item.
- 02The Announcement
State what is being shared and give its main benefit, use case, or current status.
- 03The Proof Points
List key features and add measurable performance or outcome data.
- 04The Call To Action
Invite readers to explore the full resource through a direct link.