Untitled source document
No live link available for this source.
What we drew from this source
The claims Via News extracted from this document. We point to the source; we don't replace it.
SubQ dynamically selects which token relationships are important on the fly, differently for each piece of text, rather than using fixed patterns as prior sparse-attention mechanisms have done.
60% confidenceSubQ is either the biggest breakthrough since the Transformer or it's AI Theranos.
60% confidenceIn hindsight, releasing third-party benchmarks alongside the initial announcement would have preempted the skepticism.
60% confidenceThe Appen evaluation validated Subquadratic's architecture and suggests SubQ could be a game changer given models' struggles with speed and inefficiency.
60% confidenceAchieving competitive sparse attention is extremely difficult — akin to running a four-minute mile — and pretty much every approach under the sun has already been attempted.
60% confidenceSparse attention is justified because not all word relationships in a document are important.
60% confidenceSubQ is faster, cheaper, and uses significantly less energy than any other LLM on the market.
60% confidenceSubquadratic hopes to kick off a new age of LLM efficiency and believes nobody will be building on transformers in a few years.
60% confidenceSubQ matches the performance of the best models from Google DeepMind, OpenAI, and Anthropic on key tasks like coding.
60% confidenceIt costs $2,600 to run Anthropic's Claude Opus 4.6 through the RULER 128 benchmark, versus $8 for SubQ.
60% confidenceTens of thousands of potential users have signed up for early access to SubQ, including more than 500 enterprise customers.
60% confidenceSubQ scored 98% on needle-in-a-haystack with context windows of 6 million and 12 million tokens, sustaining near-perfect long-context retrieval at scales few models are tested at.
60% confidenceSubQ is the first sparse-attention LLM that rivals mainstream dense-attention models in performance.
60% confidenceSubquadratic may have built something real and useful, but the public evidence does not yet justify the stronger claim that they have solved the quadratic attention bottleneck.
60% confidenceSubQ continues to provide frontier-level performance in coding.
60% confidenceSubQ can process up to 12 times as much text at once as most other models, enabling analysis of hundreds of documents or entire codebases.
60% confidence
