熊出没之夺宝熊兵
英伟达新一代算力平台补齐拼图 结合DeepSeek模型Token成本可降35倍_我的网站

一 | 英伟达硬件演进正加速驶向下一站,为智能体(Agent)打造新一代算力平台。 当地时间8月24日,英伟达宣布,推理加速器Groq 3 LPX机架已进入全面量产阶段。英伟达高级总监Dion Harris表示,Groq机架将与Vera中央处理器和Rubin图形处理器一同部署在新型云服务商Nebius的数据中心,并将于今年晚些时候上线。 New Delhi, Oct, 20 (UNI) A Delhi Special Court fixed October, 31 for hearing arguments on the bail application in the matter registered under the provisions of Prevention Money Laundering Act against Delhi's Health Minister Satyendar.
Special Judge Vikash Dhull after hearing arguments in the bail application on behalf of Vaibhav Jain and Ankush Jain said, " Put up the matter on 27.10.2022 for hearing arguments on the bail application of Satyendar Jain."
Counsel for Jain submitted before the Court that main counsel Sh. N. Hariharan who is to argue the bail application on behalf of Jain has to go for some medical checkup and will not be available before 31.10.2022.
After hearing the request on behalf of Jain the Court said, "Put up on 31.10.2022 for arguments on the bail application of accused Satyendar Kumar Jain. Superintendent, Tihar Jail is directed to produce accused Satyendar Kumar Jain, Vaibhav Jain and Ankush Jain through video conferencing on 31.10.2022"
The Principal District & Sessions Court of Rouse Avenue, New Delhi on September, 22 transfer the matter registered under the provisions of Prevention Money Laundering Act against Delhi's Health Minister Satyendar Jain from the Court of Special Judge Geetanjali Goel's to Special Judge Vikash Dhull. Jain challenged the order of transfer in High Court and Apex Court and later on withdraw the same from Apex Court. Now the matter is at the stage of hearing arguments on bail application of Jain.
Government Counsel for ED opposed the bail application by submitting that Jain served as a Health Minister in Delhi may managed to get forged documents and can influence the doctors and Jail official.
ED has earlier opposed the bail application of Satyendar Jain by submitting that if bail granted he may influence the co-accused, witnesses and other documents related to the case.
Counsel for Jain opposed the submission of ED by submitting that this is a malafide application to derail the trial and to prolong the custody of Jain. He contented that Jain is neither a Health Minister nor Jail Minister but is in jail and because the hospital where Jain was admitted under the control of Delhi Government is no ground to alleged a bias in the case.
Investigating Agency has arrested Satyendar Jain in this Matter on May, 30 under the provisions of PMLA and he is now in the judicial custody at Tihar.
Central Bureau of Investigation (CBI) in its chargesheet has alleged amassing assets to the tune of Rs. 1.47 crore in Disproportionate Assets case while Enforcement Directorate has submitted the attachment of Rs. 4.81 crore in connection with money laundering investigation.
UNI XC GNK。 Groq 3 LPX机架是英伟达专为Vera Rubin平台设计的推理加速扩展组件,需与 Vera Rubin NVL72 GPU主机架搭配部署——NVL72负责更广泛的训练、推理以及上下文处理,LPX专门加速输出AI生成结果(Token),主打超低延迟Token生成,两者异构协同,可大幅提升万亿参数、超长上下文智能体推理的每瓦吞吐与交互响应速度。 英伟达此前将Vera Rubin平台定义为专为智能体AI与科学计算时代设计的机架级超级计算平台。英伟达此次将Groq 3 LPX正式推向全面生产,核心目的并不是用专用推理芯片取代GPU,而是为Vera Rubin补充传统GPU架构在低延迟Token生成方面的能力。

二 | 相比Blackwell,Vera Rubin在智能体工作上的效率得到了极大的提升,英伟达释出多个最新测试数据(本次测试使用的基准为SemiAnalysis AgentX,该基准针对Kimi K3、MiniMax M3、GLM5.3、Qwen3.5、DeepSeek V4 Pro等各类模型)—— 英伟达称,在Artificial Analysis测试中,搭载Gemma 4 31B模型、100K Token上下文时,Groq 3 LPX实现每秒3400个输出Token,为该模型创下最高纪录;针对智能体编程等低延迟工作负载,其响应速度最高达到最近竞争平台的4倍。英伟达称,这种速度可以让智能体更快完成代码编写、测试、工具调用和结果验证等连续任务。 与此同时,英伟达公布Vera Rubin NVL72在真实智能体工作负载下的新测试结果:基于SemiAnalysis的AgentX工作负载、采用DeepSeek V4 Pro模型,Vera Rubin每兆瓦(MW)吞吐量最高达到上一代GB300 NVL72的30倍,每Token成本最高降低35倍。对于运营商而言,这意味着在固定的电力和基础设施预算下,可以支持更高的交互式代理容量,或者以更低的运营成本提供相同的容量。
英伟达强调,智能体任务不同于传统聊天对话,任务上下文会随着数百个步骤不断累积,甚至达到数十万Token。

三 | 英伟达援引OpenRouter数据称,智能体AI工作负载消耗的Token数量可以达到简单聊天请求的15倍。原因在于,一个智能体可能先查询数据库,再搜索新闻和文件,调用子智能体进行分析,运行代码并反复验证,整个过程中上下文不断累积。 近期腾讯、阿里、DeepSeek等头部厂商相继推进AI智能体产品迭代与生态布局,AI应用端发展聚焦终端场景落地与生态壁垒搭建。天风国际证券最新研报称,当前Harness加速开放迭代,前沿AI竞争正向Agent执行框架与商业化兑现深化。 英伟达正紧随这一趋势调整硬件优化方向,当AI迈入智能体时代,AI基础设施的评价体系或将随之变化,基础设施服务商的硬件迭代重心将从追求训练更大模型,转向单位能耗有效Token产出、单位成本任务处理能力,以及智能体复杂任务的执行时延。
(文章来源:财联社)。
Current article:http://kpd.guizhuanguchuanluanzhoudi.shop/017519l/os4v.html
Published on:19:20:17
