Introducing SubQ - a major breakthrough in LLM intelligence. It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA), And the first frontier model with a 12 million token context window which is: - 52x faster than FlashAttention at 1MM tokens - Less than 5% the cost of Opus Transformer-based LLMs waste compute by processing every possible relationship between words (standard attention). Only a small fraction actually matter. @subquadratic finds and focuses only on the ones that do. That's nearly 1,000x less compute and a new way for LLMs to scale.

- Views
- 13,488,733
- Likes
- 22,471
- Replies
- 1,484
- Reposts
- 2,731
- Quotes
- 1,970
- Bookmarks
- 19,058
Introducing SubQ - a major breakthrough in LLM intelligence.
SubQ’s co-founder announces an LLM built on fully sub-quadratic sparse attention, claiming a 12-million-token context window, faster processing, and much lower costs than competing models. The video also introduces SubQ Code, an AI coding agent for working across large codebases and document sets.
- Views
- 13.5M
- Likes
- 22.5K
- Replies
- 1.5K
- Reposts
- 2.7K
- Quotes
- 2K
- Bookmarks
- 19.1K
- Topics
Analysis
SubQ’s co-founder announces an LLM built on fully sub-quadratic sparse attention, claiming a 12-million-token context window, faster processing, and much lower costs than competing models. The video also introduces SubQ Code, an AI coding agent for working across large codebases and document sets.
Sign in to see the full analysis.
Sign inVisual hook
Introducing SubQ - a major breakthrough in LLM intelligence.
- 01The Announcement
Introduce the subject and frame it as a major breakthrough.
- 02The Proof Points
Position the subject as a first-of-its-kind solution, then give clear performance comparisons.
- 03The Simple Explanation
Explain the old problem, show that only part of it matters, and reveal the new approach and its payoff.
