DeepSeek faster, cheaper: innovation speeds processing of long text 10 times, paper says
Chinese AI start-up DeepSeek has developed a new "native sparse attention" (NSA) method that significantly speeds up and reduces the cost of processing long texts for next-generation language models. The technology, detailed in a paper by CEO Liang Wenfeng and his team, achieves up to an 11-fold increase in processing speed by training AI to focus on key information rather than analyzing every word.
The NSA method combines algorithmic innovations with hardware improvements to enhance efficiency without compromising performance. This advancement could greatly improve AI's capabilities in solving complex problems, writing extensive programs, and maintaining context in lengthy conversations. DeepSeek's announcement on X followed the release of xAI's Grok 3 model, highlighting the ongoing advancements in AI technology.
via SCMP Full Text Feed
Chinese AI start-up DeepSeek has developed a new "native sparse attention" (NSA) method that significantly speeds up and reduces the cost of processing long texts for next-generation language models. The technology, detailed in a paper by CEO Liang Wenfeng and his team, achieves up to an 11-fold increase in processing speed by training AI to focus on key information rather than analyzing every word.
The NSA method combines algorithmic innovations with hardware improvements to enhance efficiency without compromising performance. This advancement could greatly improve AI's capabilities in solving complex problems, writing extensive programs, and maintaining context in lengthy conversations. DeepSeek's announcement on X followed the release of xAI's Grok 3 model, highlighting the ongoing advancements in AI technology.
via SCMP Full Text Feed