跳到正文
今天10月9日周五1 条
  1. a16z66

    AWS CEO Matt Garman 与 a16z 的 Raghu Raghuram 对谈称,AWS 本可把全部 GPU 卖给 frontier 实验室,但刻意为创业公司保留算力以维持生态健康,约 60% 的资源请求会以某种形式获得批准。

    AI 生成摘要 · 以原文为准

    引用a16z@a16z

    AWS CEO Matt Garman 与 a16z 的 Raghu Raghuram 对谈:全球最大云提供商如何为智能体时代重构: 2026 年 2200 亿美元资本支出,订购 200 万块 NVIDIA GPU,AWS 自研 AI 芯片明年之前售罄,还推出无需信用卡、30 秒即可开通的 AWS 账户,让智能体立即动起来。 0:55 营收 1700 亿美元,增长 37% 2:40 30-40% 的营收最初来自初创公司 4:55 初创公司想要面向智能体的云 9:10 30 秒开通 AWS 账户 12:25 智能体不需要坚不可摧的数据库 16:30 为什么 AWS 不会把所有 GPU 卖给前沿实验室 22:00 为什么 Matt 不担心泡沫 24:45 为什么 AWS 自建电厂 27:45 为下一次供应短缺做规划 30:25 为一个县每位居民节省 5000 美元 33:10 AWS 如何打造自研芯片 40:05 是什么阻碍了企业级智能体 45:35 为什么 Bedrock 从不把你的数据发给模型 49:30 AWS 的新 AI 安全工具 51:50 为什么 AWS 的 10 人团队如今只需 3-4 人 YouTube: https://www.youtube.com/watch?v=rn_afJaPldg @mattsgarman @awscloud @RaghuRaghuram

    同一新闻,精选展示《AWS CEO Matt Garman 对谈 a16z:2200 亿美元 CapEx 与为智能体时代重构 AWS》

10月8日周四
  1. NVIDIA Newsroom62

    NVIDIA 承诺未来五年投入 10 亿美元推动美国科学研究

    NVIDIA 宣布未来五年承诺投入价值 10 亿美元的资源,支持美国在量子计算、医疗健康和能源安全等领域的超级智能研发,并在华盛顿 Science: A New Golden Age 活动上公布。资金将用于支持美国高等教育研究机构、推进量子计算领导力,以及支持云服务商满足美国政府任务需求。

    AI 生成摘要 · 以原文为准

    关注理由:NVIDIA 承诺五年内投入 10 亿美元支持美国科研,重点覆盖量子计算、医疗与能源安全,明确了其对国家级科研算力与量子方向的投入布局。

  2. EE Times30

    NVIDIA Jetson 迁移至 SiMa.ai Modalix MLSoC 指南

    SiMa.ai 发布技术指南,介绍如何将机器学习模型和应用从 NVIDIA/CUDA 环境迁移到其 Modalix MLSoC 与 Palette 软件栈。模型可导出为 ONNX 后由 Model Compiler 完成量化与编译,无需重新训练;应用侧使用 Palette Neat 适配,从预处理到 DeepStream。

    AI 生成摘要 · 以原文为准

10月7日周三
  1. AMD26

    为 AI 时代而生。 我们很自豪能以第六代 AMD EPYC 处理器,为新一代 @HPE ProLiant 服务器提供动力。

    AI 生成摘要 · 以原文为准

    引用HPE@HPE

    Future-ready servers are here. Meet HPE ProLiant Gen13 Servers, powered by 6th Gen @AMD EPYC™ Server processors, with automatic security, intelligent ops, and breakthrough performance. Generation after generation, HPE and AMD are helping enterprises modernize for what’s next.

10月5日周一
10月4日周日
10月3日周六
  1. AMD49

    周六早晨,我们一边享受咖啡☕、翻翻新闻📰,一边看着我们 AMD 向 vLLM 提交的代码。 开源万岁。

    AI 生成摘要 · 以原文为准

    引用Ramine Roane@roaner

    #1 company contributing code to @vllm_project right now? @AMD: 502 commits to core vLLM in 90 days. 13% of all org contributions, almost 4x Nvidia. Also grateful to Red Hat, IBM, Embedded LLM & Inferact for building it with us. Open source wins when hardware has a choice.

  2. Fabricated Knowledge33

    笑死,这渲染搞得我火大——兄弟,我们根本不是在 PCIE 上用显卡的

    AI 生成摘要 · 以原文为准

    引用Bearly AI@bearlyai

    Someone used Claude Opus 5.5 to make a 3D animation of Nvidia Blackwell GPUs (from server down to an atom). Took 1 hour and full animation “runs from single HTML file, directly in browser. No vid editor. No pre-rendered 3D sequence. Just browser-based experience.” Very cool.

10月2日周五
  1. NVIDIA Newsroom64

    OpenAI 发布 GPT-6 Astra Ultrafast,NVIDIA Blackwell GPU 加速下 token 生成速度最高提升 8 倍

    GPT-6 Astra Ultrafast 现已在 OpenAI API 及面向符合条件的 ChatGPT Work 与 Codex 用户开放,运行在 NVIDIA Blackwell GPU 上。

    AI 生成摘要 · 以原文为准

    关注理由:OpenAI 基于 NVIDIA Blackwell 推出 Ultrafast 推理模式,展示了模型厂商通过自研推理优化榨取 GPU 硬件性能的路径,对算力平台软件生态有意义。

10月1日周四
9月30日周三
9月29日周二
  1. AMD Press Releases85

    AMD宣布以约82亿美元全股票收购李飞飞创立的World Labs

    AMD于9月28日宣布达成最终协议,将以约82亿美元全股票收购李飞飞(Dr. Fei-Fei Li)创立的AI模型研究公司World Labs,交易预计于2026年底前完成,尚待监管审批等常规交割条件。World Labs总部位于旧金山,开发从文本、图像和视频输入生成、重建并模拟交互式3D环境的空间智能模型,以及机器人学习与模拟技术。

    AI 生成摘要 · 以原文为准

    关注理由:AMD以约82亿美元全股票收购李飞飞创办的World Labs,把空间智能模型研究能力纳入自身路线图,直接影响其AI硬件与软件的模型导向开发方向。

  2. SemiAnalysis72

    SemiAnalysis解析GLM-5.3稀疏注意力如何影响HBM内存占用

    SemiAnalysis撰文分析GLM-5.3的稀疏注意力与DSA架构对HBM内存和推理成本的影响。文章指出top-k选择仍需完整上下文驻留HBM,稀疏注意力并不消除内存容量瓶颈,SGLang的HiSparse通过将KV cache按LRU策略在HBM与主存DRAM间分层换入换出并跨层重叠加载,在高并发长上下文下维持95%以上缓存命中率。

    AI 生成摘要 · 以原文为准

9月28日周一
9月27日周日
  1. Dylan Patel53

    SemiAnalysis 的 Dylan Patel 发推称 Rubin HBM 降规格(Despec)等于 Nvidia 减料,并引用一篇讽刺文:GB200 NVL72 约 400 万美元、GB300 约 500 万美元(重约 1580 kg),Vera Rubin NVL72 已被报至最高 880 万美元,机柜重量受数据中心地板承重限制而价格不受限,$/kg 每代约 1.75 倍增长,外推 2030 年与可卡因批发价出现交叉。内容为戏谑观点而非测算结论。

    AI 生成摘要 · 以原文为准

    引用Dylan Patel@dylan522p

    Fentanyl Grade Compute: Why a GB300 Rack Will Out-Price Blow by Weight in 2030 A GB300 NVL72 weighs roughly 1,580 kg fully populated and recent purchase orders put it at $5M per rack. That is $3,165/kg. Strip out the 1.5 tons of busbar, manifold and coolant and the GPU packages alone are well into gold territory, but we're pricing the rack, because that's what you actually take delivery of. Where that sits on the illicit commodity curve today ($/kg): $2,400 Cannabis flower $3,165 GB300 NVL72 $3,500 Fentanyl $28,000 Cocaine $65,000 Heroin $138,000 Gold So today NVIDIA ships a product that is denser in value than weed, roughly fentanyl-grade, and still an order of magnitude short of cocaine. For now... GB200 NVL72 was $4M, GB300 is ~$5M, and Vera Rubin NVL72 is already quoted at up to $8.8M at essentially the same rack mass. Rack weight is constrained by the datacenter floor loading; rack price is constrained by nothing. That's ~1.75x $/kg per generation with HBM, CoWoS and power all supply-limited through the decade. Cocaine: Colombia's new president has pledged a hard-line security crackdown with $1B in US aid behind it Every prior crackdown has consolidated the industry into fewer, better-capitalized operators with better logistics and higher yields per hectare. Consolidation is deflationary. We model US wholesale drifting from ~$28k/kg toward ~$12k/kg by 2030 as the supply chain professionalizes. **Crossover: 2030.** Our extrapolated NVL72-class rack hits ~$17k/kg while wholesale cocaine falls to ~$12k/kg. At that point the rational cartel pivots to smuggling racks, except a rack has a fixed 500+ kW power draw, a 3,300 lb forklift requirement and an export-control regime that is actually enforced. Cocaine has none of these problems, which is why it will remain the superior product for anyone without a substation. Key risks to the thesis: US retail coke is still ~$60–200/g, i.e. $60k–200k/kg so the crossover only holds at wholesale. Retail compute (H100 hour on a neocloud) is also marked up, so we consider this an apples-to-apples wholesale comparison. Full model available to Coke Research subscribers.

9月25日周五
9月24日周四
  1. SemiAnalysis76

    SemiAnalysis 发布 ClusterMAX 3.0 GPU 云评级:Nebius 升入铂金级

    SemiAnalysis 发布 ClusterMAX 3.0 GPU 云评级,覆盖 77 家受测供应商与 323 家市场观察对象,较 2.0 的 209 家继续扩大,仅 19 家获得奖牌级评级。Nebius 加入 CoreWeave 升入铂金级,Google Cloud 进入金级,Azure 降至银级,Fluidstack 降为不可用,Crusoe 降至铜级。

    AI 生成摘要 · 以原文为准

    关注理由:SemiAnalysis 第三版 GPU 云评测覆盖 323 家供应商并给出分层结果,同时梳理融资、Blackwell 部署与可靠性实践,为评估 neocloud 服务能力与采购谈判提供参考框架。

9月22日周二
9月21日周一
9月18日周五
  1. NVIDIA54

    NVIDIA 发文展示 World Labs 的 Atlas 模型:仅用 32 张园区照片作为三维空间上下文生成新视角,可实时漫游 NVIDIA 的 Voyager 园区并自由选择路线。该模型在 NVIDIA Blackwell GPU 上训练,基于 NVIDIA 开放平台构建。

    AI 生成摘要 · 以原文为准

    引用World Labs@theworldlabs

    From 32 input images to real-time flight through @nvidia's Voyager headquarters. Trained on NVIDIA Blackwell GPUs, Atlas uses these images as 3D spatial context to generate new views, letting you explore with pixel-perfect camera control. Take a look around.