资讯日报 · DAILY BRIEF
星期二
15:00 更新
36 条苹果 41 页诉状点名三人却放过伊夫:古尔曼拆解这份"留白"背后的三重算计 ↗
一份41页的诉状,把苹果与OpenAI的暗战摆上了法庭,但最耐人寻味的,是它点了谁的名、又故意略过了谁。据多家媒体援引彭博社马克·古尔曼的分析,苹果自7月12日起诉OpenAI及其硬件子公司io Products盗用商业秘密,诉状中明确点名了Tang Tan、Chang Liu、Yu-Ting Alyssa Peng三人,却对前苹果设计师乔纳森·伊夫只字未提。 古尔曼认为,这份留白并非疏忽,而是经过权衡的结果,背后至少叠着三层考量。 第一 层是关联性的强弱——在苹果的判断里,伊夫与这起商业秘密争议的实质牵连有限,不足以把他推上被告席。 第二层则牵出人际与资本的网络:伊夫与乔布斯遗孀Laurene Powell Jobs私交甚好,而后者在相关投资与支持上产生的影响,让直接点名伊夫的代价和连带效应都变得复杂。第三层是舆论层面的避险——若把这位苹果设计灵魂人物卷进诉讼,势必将引发外界对苹果设计话语权是否易主的无尽争论,苹果显然不愿在打官司的同时,再给自己添一场关于设计地位变动的公共危机。 当诉状用41页的篇幅细数OpenAI与io Products的"偷师"指控,却在对伊夫的处理上按下静音键,这份刻意留出的空白,本身就成了观察苹果诉讼策略的一扇窗。
三星电子成立RX机器人事业部,加速机器人业务商业化 ↗
7月21日,三星电子正式成立机器人事业部RX(Robotics eXperience,机器人体验),该部门将直接向公司首席执行官汇报,并被纳入三星未来增长战略的重要布局。 据了解,RX事业部将负责统筹三星机器人领域的中长期发展规划,推动核心机器人技术研发与产品业务落地,同时加强韩国及海外研发体系建设,以提升机器人技术商业化能力。 此次组织调整意味着三星正进一步加大对机器人领域的投入,推动相关业务从技术研发阶段向规模化应用阶段转型。作为全球领先的消费电子和智能硬件企业,三星希望通过机器人技术拓展新的增长空间,并强化其在人工智能、智能家居和未来服务机器人领域的竞争力。 近年来,随着人工智能、大模型和智能制造技术快速发展,机器人产业正成为科技企业竞争的新方向。包括三星、英伟达、谷歌、特斯拉等企业均在加快机器人技术布局,推动AI能力与实体设备深度融合。 三星此次成立独立机器人事业部,显示其正在通过组织架构调整整合研发资源,加快机器人产品商业化进程。未来,RX事业部能否推出具备市场竞争力的机器人产品,将成为观察三星AI与智能硬件战略落地的重要指标。
智谱AI落地1GW国产AI算力中心,并收购中科加禾强化AI Infra能力 ↗
据外媒报道,智谱AI(Z.ai)已完成1GW级国产AI算力数据中心建设,整体采用国产AI芯片。同时,公司于近日完成对国产AI异构算力软件企业中科加禾(XCore Sigma)的收购,进一步完善从算力资源到基础软件的AI基础设施布局。 据了解,中科加禾源自中国科学院计算技术研究所编译实验室,长期专注于异构算力软件栈、编译优化、运行时系统以及AI推理基础设施建设,被业内视为国内领先的AI Infra技术团队之一。此次收购将增强智谱在国产芯片适配、算力调度和模型部署优化方面的能力。 业内认为,智谱此次两项布局分别对应AI产业链中的“算力供给”和“算力释放”环节。其中,1GW级国产算力中心能够为大规模模型训练提供稳定计算资源;中科加禾的软件技术则有助于提升不同类型AI芯片的计算效率,通过编译器、Runtime和推理引擎优化,提高硬件利用率,降低模型推理成本。 随着大模型竞争进入基础设施和工程能力比拼阶段,AI企业正在从单纯追求模型规模,转向构建覆盖芯片、算力、软件栈和应用部署的完整技术体系。近期,市场对智谱下一代基础模型持续关注。有观点认为,在大规模国产算力资源、成熟AI Infra体系以及长期积累的后训练(Post-training)能力支持下,智谱未来模型有望继续向更大参数规模、更高智能水平方向发展,同时提升推理效率和实际部署能力。 此次算力中心建设与基础软件能力整合,标志着国产大模型企业正加速构建自主可控的AI基础设施生态,也反映出AI产业竞争正在从模型能力扩展至全栈技术体系的竞争。
唐杰亲自剧透GLM-5.5:一句"史诗级plus",把国产大模型的军备赛又挑高了一截 ↗
国产大模型的密集爆发,正把和美国旗舰闭源AI的身位拉近到肉眼难辨。快科技7月20日报道,在千问、Kimi接连秀出肌肉之后,轮到智谱出手了,传闻中的GLM-5.5被描述为一轮史诗级提升。而这次不是媒体猜测,是智谱创始人唐杰亲自下了场。 有网友在评论里提到,千问和Kimi都完成了史诗级进化,顺口问了句GLM还有没有戏。唐杰的回应当场把期待值拉满——他用了一个词:史诗级plus。这意味着,在智谱自己的判断里,这一代的提升幅度比网友提到的那些还要大。 具体的数字现在当然无从谈起,但唐杰显然希望外界对GLM保持信心。他的底气来自上一代的战绩:在Kimi K3登场之前,GLM-5.2是国产大模型中 第一 个在AI编程能力上追平Opus级别性能的选手。更值得注意的是它的体量——GLM-5.2的参数规模只有7400多亿,差不多是Kimi K3的四分之一,大概也只有美国 顶级 AI的五分之一甚至更少,却硬是靠后训练把能力托到了那个水位,堪称小参数撬动大能力的样本。 唐杰口中的史诗级plus新模型,指向的应该是下一代GLM-5.5。此前圈内传过GLM-5.3的名字,但现在看GLM-5.5的可能性更大,当然也不排除发布时直接冠上GLM-6的名号。一个可以确定的趋势是,下一代大模型的参数量至少会迈过1万亿的门槛。考虑到智谱历史上曾使用过DeepSeek开源的基模,外界猜测GLM-5.5或许会做到与V4Pro 同级 的1.6万亿参数;再叠加智谱公认颇强的后训练功力,其性能跃升确实值得押注,很可能又是一个对标Fable5量级的国产大模型。 不过在国产阵营的等待名单里,眼下最让人坐立难安的,还是DeepSeek V4的正式版。报道发出时,7月中旬的最后一刻——当天零点——正悬在头顶,那条小鲸鱼究竟是打算半夜甩出大招,还是又一次跳了票,谁也说不准。当智谱用一句史诗级plus把话题点燃,国产大模型这场接力赛,显然还远没到撞线的时候。
Adobe Project Indigo引入AI摄影助手,利用大模型优化拍摄与编辑体验 ↗
Adobe正在为旗下实验性iOS相机应用Project Indigo加入一系列AI功能,通过大型语言模型(LLM)分析照片内容,为用户提供摄影点评、编辑建议以及智能优化方案。该应用此前已支持专业拍摄控制、多帧超分辨率和多种摄影模式,此次更新进一步强化了AI辅助创作能力。 新功能包括照片点评、拍摄与编辑建议、 高级 物体移除、AI景深生成以及照片风格迁移等。其中,照片点评功能能够围绕构图、光线、色彩和情感表达等维度提供类似专业摄影师的分析;拍摄建议功能则会针对重新构图、曝光调整以及画面元素优化提出具体指导,帮助用户提升拍摄技巧。 Project Indigo开发负责人Marc Levoy曾参与谷歌Pixel相机技术研发。他表示,当前多数生成式AI图像工具依赖用户输入提示词,但普通用户往往难以准确描述想要的调整效果。因此,Project Indigo更多采用按钮化和确定性操作,让AI直接提供可执行的编辑方案。 在智能编辑方面,Project Indigo新增的对象移除功能可以自动识别并删除背景人物、垃圾桶、电线杆、车辆等干扰元素,同时支持用户自定义描述需要移除的对象。相比传统需要手动圈选目标的方式,该功能能够减少操作成本,并提升处理效果。 此外,应用还支持通过AI生成景深效果,为普通照片模拟背景虚化,并提供水彩、钢笔画、彩色墨线、单色、逆光等风格迁移效果。不过,风格转换并非全新技术,其实际应用价值仍有待进一步验证。 目前,Project Indigo的新AI能力由基于谷歌Gemini的Nano Banana模型提供支持,Adobe也表示未来可能测试包括自研Adobe Firefly模型在内的其他AI模型。这些功能目前集中在“AI Playground”测试区域,仅向部分用户开放,是否最终全面推广仍未确定。 随着AI图像工具逐渐从简单生成和自动修图向智能摄影助手方向发展,Adobe正在探索如何利用大模型帮助用户理解摄影、优化创作流程,而不仅仅是生成视觉效果。该方向也代表了移动影像领域AI应用从“自动处理”向“辅助决策”演进的趋势。
2 万亿参数隔空交锋:Kimi K3 登顶后,月之暗面喊话马斯克"欢迎入伙" ↗
一场发生在中文微博与英文社交平台之间的隔空喊话,把大模型参数竞赛的火药味推到了台前。7月20日,月之暗面发布的新一代开源大模型Kimi K3登顶全球榜单,引发广泛关注,而就在几天前,埃隆·马斯克的一句话,让这场竞争从榜单延伸到了口水仗。 7月18日上午,马斯克在社交平台发文称,旗下那款2万亿参数的模型在各方面都优于当前1.5万亿参数版本,将在下周完成初步训练,且有可能超越Kimi;同时他放话,新模型的推理速度和Token效率将接近现有的1.5T模型,也就是Grok4.5。这番表态,等于在Kimi K3刚刚登顶的节点上,直接把参数规模的对标线拉到了2万亿以上。 月之暗面的回应干脆又带刺。它在微博上公开发文@马斯克,只扔下一句话:欢迎加入2万亿+俱乐部。一句看似轻松的邀约,实则是趁着登顶之势,把自家的开源成绩摆成了俱乐部的门槛。 不过,这场隔空对话里还留着不少未解之谜。目前xAI尚未公布Grok4.6的正式发布时间、训练细节与完整技术指标,马斯克的2万亿模型究竟是代号4.6还是另一款新作,外界仍只能等待。而把时间拨回不久前的7月9日,xAI已经发布了Grok4.5——这是该公司首个专门面向编程与智能体任务训练的模型,由xAI与Cursor联合完成训练,在提供前沿智能水平的同时兼顾领先的速度与成本效率,马斯克本人将其称为Opus级模型。当2万亿参数的新王尚未露面,Kimi K3已经先一步把开源榜单的椅子坐热,这场俱乐部之争的下一回合,恐怕要等马斯克的新模型真正走出训练阶段才算开场。
马斯克把Grok塞进Excel:选中一片数据就能问涨跌原因,图表直接插进表格 ↗
马斯克旗下的xAI,把自家大模型Grok直接装进了微软Office的心脏地带。据多家财经媒体 7 月 21 日消息,xAI为微软Excel推出了一款名为Grok For Excel的免费Microsoft365 插件,目前已上线Microsoft Marketplace,同时兼容Word与PowerPoint。这意味着Grok不再只活在聊天框里,而是成了办公套件里随手可调用的数据分析员。 这款插件的使用逻辑极其贴近真实办公场景:用户在Excel里选中一片数据区域,就能直接向Grok发问,询问数据的变动情况、背后的原因以及其中的亮点,而Grok的回答会明确引用对应的数据单元格,不是凭空给结论。它还能理解公式语法,完成平均值计算、排序、填充等常规编辑操作,生成的图表则可以一键插入工作表。换句话说,过去要写函数、拖公式、手动画图的活儿,现在对着数据说句话就能推进。 更关键的是,Grok并不只盯着表格里那一小块被选中的数据。借助连接器,它能从电子邮件、SharePoint或Google Drive中提取上下文信息,把散落在各处的资料拼进当前分析。这一能力让Excel里的问答不再是孤岛,而是能结合邮件往来、团队文档和外部文件做出更完整的判断。当Grok以免费插件的姿态嵌入Word、PowerPoint和Excel三件套,办公软件与对话式AI的边界,又被磨掉了一层。
苹果 41 页诉状怒撕OpenAI,为何唯独放过了传奇设计师伊夫? ↗
苹果公司近日正式向法院提起诉讼,指控OpenAI及其硬件子公司通过挖角员工和获取内部文件等手段盗用商业机密。在长达 41 页的诉状中,多名前高管及核心工程师被点名控告,但令人意外的是,前首席设计师乔纳森·伊夫却并未出现在名单中。 三重考量避开核心人物 知名科技记者马克·古尔曼对此进行了深度剖析,指出苹果未点名伊夫的首要原因是其关联性有限。伊夫虽然通过旗下设计公司参与了相关硬件项目,但他并不负责实际的招聘与日常运营工作。 此外,深层的资金与人脉关系也是苹果选择回避的关键因素。伊夫与乔布斯遗孀关系密切,后者不仅投资了涉事硬件公司,其家族姓氏在苹果内部依然拥有举足轻重的特殊影响力。 顾及形象避免舆论反噬 除了人际考量,法庭上的舆论风险同样是苹果必须衡量的因素。在后乔布斯时代,苹果的战略重心逐渐向运营效率与成本控制倾斜,设计的核心地位已大不如前。 如果将伊夫卷入这场诉讼,辩方律师极可能在法庭质询环节借机发难,迫使伊夫公开谈论苹果设计理念的转变与退化。这种潜在的舆论风暴不仅无法帮助苹果赢得官司,反而可能对公司的品牌形象造成二次伤害。
经典版 Outlook 迎来重大更新:微软宣布今年内引入 AI 自动起草邮件功能 ↗
微软正在积极推进旗下办公软件的智能升级。根据 最新 规划,微软计划在年底前向 Windows 10 和 Windows 11 平台上的经典版 Outlook 用户推送 Copilot 功能更新。 AI 赋能邮件撰写 这次更新最核心的亮点是加入了 AI 起草电子邮件功能。持有相关许可证的用户届时可以直接在写信框中一键唤醒 AI,让系统协助撰写新邮件或润色已有草稿。 微软此举旨在让老用户也能享受智能化带来的便捷体验。不过也有分析指出,这一变动可能打破部分传统用户的习惯,甚至增加日常使用的复杂性。 推动软件升级迭代 其实微软一直以来都在引导用户向新版 Outlook 迁移。团队正不断丰富新版应用的功能生态,试图以此吸引更多用户主动完成版本过渡。 尽管开发重心早已偏向新版,微软依然选择为经典版追加 AI 功能。这一策略能否兼顾老用户的习惯与智能化趋势,仍有待市场后续检验。
谷歌密造"Frozen v2"专用芯片:把Gemini烙进硅片,单位功耗token能力翻 6 到 10 倍 ↗
把模型直接焊进芯片里,让硅片替算法省下算力和数据传输——谷歌正沿着这条路,悄悄打磨一款代号Frozen v2 的新型服务器芯片。据多家媒体报道,这款芯片的目标是更高效地运行Gemini模型,报道称它将把Gemini的部分架构 永久 嵌入芯片之中,从而减少回答用户问题时所需的计算量与数据搬运量。谷歌工程师预计,与其 最新 一代通用AI芯片TPU(张量处理器)相比,Frozen芯片在单位功耗下可提供 6 至 10 倍的Token处理能力。 这款芯片的登场时间被定在 2028 年,但谷歌明确它不会取代通用型TPU,而是成为定制芯片产品线中更专一、更具针对性的专用分支。消息公布后,资本市场立刻给出了反应:谷歌A类股(GOOGL)美股早盘一度涨近3.7%,C类股(GOOG)盘中 最高 涨幅约3.9%。 面对媒体追问,谷歌在声明中回应称,公司团队持续研究并试验各类创新技术,旨在为用户和客户提供 最佳 性能与 最高 效率,并强调虽然并非所有项目都会进入量产,但这种严格的探索正是其全栈战略的核心组成部分。谷歌还表示,通过从底层协同设计硬件与软件,整个系统得以深度集成,并针对真实工作负载高度优化。不过谷歌眼下把Frozen v2 定位为一次试验性项目,暂无像TPU那样大规模生产的计划。 分析人士认为,这个项目的背后,是谷歌内部日益吃紧的算力困境。此前算力不足不仅加剧了内部资源争夺,据称还一度迫使谷歌云拒绝部分外部客户业务。就在上月,谷歌还同意每月向SpaceX支付近 10 亿美元,以填补算力缺口、兑现对企业客户的算力承诺。Frozen v2 显然是为缓解这道缺口而落子的一步。 把架构固化进芯片,代价是灵活性的收缩。报道指出,只有当谷歌未来的Gemini模型继续沿用相同底层架构时,这款芯片才能兼容后续版本——一旦模型底层大改,烙进去的那部分设计便可能失效。而比芯片更紧迫的,是谷歌AI业务当下正面临的连环挑战:有消息称下一代Gemini Pro的发布时间已经推迟,同时多名 资深 AI研究人员流向了竞争对手。当专用芯片还在 2028 年的路线图里,谷歌要先稳住的是眼下的模型节奏与人才防线。
15 亿美元和解获法官点头:Anthropic版权案落地,作家集体诉讼画上句号 ↗
一桩牵动整个AI行业的版权拉锯战,终于等来了法官落槌。据界面新闻援引报道,当地时间 7 月 20 日,美国加州北区联邦法官Araceli Martínez-Olguín正式批准了人工智能公司Anthropic涉及 15 亿美元的和解协议,并驳回了针对赔偿金额过低的异议。这份协议要解决的,是由作家团体提起的集体诉讼——这群作家早在 2024 年便指控Anthropic未经许可,将他们的著作拿来训练旗下AI聊天机器人Claude。 法官的批准,意味着围绕训练数据版权归属的这场标志性交锋暂告段落,也为AI公司使用受版权保护文本作为训练素材的合法性边界,添上了一个分量十足的判例注脚。
Anthropic获批15亿美元版权和解协议,将向50万部作品支付赔偿 ↗
据路透社报道,美国联邦法院周一正式批准Anthropic公司与作家、出版商达成的15亿美元集体诉讼和解协议。该协议用于解决Anthropic因训练AI模型过程中涉嫌侵犯版权引发的诉讼,公司随后将开始向相关版权方支付赔偿。 美国加州北区地方法院法官Araceli Martinez-Olguin签署了最终批准文件。此前,已退休法官威廉·阿尔苏普曾初步批准该和解方案,并裁定Anthropic曾非法下载并存储数百万本受版权保护的书籍。 根据赔偿方案,每部涉案作品预计可获得3000美元赔偿,涉及作品数量约50万部,赔偿金额将由拥有版权的作者和出版商共同分配。这一金额被认为是美国版权法领域规模 最大 的和解之一。 不过,该协议并未完全解决围绕AI训练数据的法律争议。阿尔苏普法官此前在案件核心问题上支持Anthropic,认为利用受版权保护文本训练AI模型可能属于“合理使用”范畴,被业内视为AI版权判例的重要进展。但与此同时,他指出Anthropic获取部分训练数据的方式存在违法问题。 Anthropic此前主要通过两种方式获取训练数据:一部分来自合法购买并扫描的书籍,另一部分则来自Library Genesis、Pirate Library Mirror等盗版网站下载的内容。法院认为,后者涉及非法获取版权材料。为避免进一步审判以及可能面临的高额赔偿,Anthropic最终选择达成和解。 业内人士指出,虽然此次和解结束了该案件,但并未形成针对整个AI行业的统一法律标准。由于Anthropic选择和解,案件未进入上诉程序,阿尔苏普此前关于AI训练“合理使用”的判断也不会成为具有约束力的司法先例。 目前,围绕AI模型训练数据版权问题的诉讼仍在持续,包括针对谷歌、Meta、Midjourney和OpenAI等公司的相关案件。近期,Hachette、Cengage、Elsevier、作家Scott Turow以及SCRIBE等出版机构和作者还联合起诉谷歌,指控其未经授权使用版权作品训练Gemini人工智能平台。 随着生成式AI快速发展,如何平衡模型训练需求与版权保护,正成为全球AI产业面临的重要法律与商业挑战。Anthropic和解案虽然提供了一种解决路径,但行业关于AI数据合规的长期规则仍待进一步明确。
谷歌研发“Frozen v2”AI芯片,预计2028年发布,能效提升最高10倍 ↗
谷歌母公司Alphabet正在研发一款代号为“Frozen v2”的新型AI服务器芯片,用于提升自主研发Gemini模型的运行效率。据The Information援引匿名消息人士报道,该芯片预计于2028年推出,其单位能耗产生代币数量的效率可能达到谷歌现有AI芯片的6至10倍。 针对相关报道,谷歌向TechCrunch表示,公司持续探索和测试创新技术,以提升系统性能与效率,但并非所有研发项目都会最终实现量产。谷歌强调,通过硬件与软件协同设计的“全栈开发”模式,可以打造更高集成度、针对实际AI工作负载优化的计算系统。 近年来,随着生成式AI应用快速扩张,全球AI算力需求持续增长,科技企业纷纷加大自研芯片投入,以提升模型运行效率并降低对外部供应商的依赖。英伟达凭借GPU技术长期占据AI芯片市场主导地位,推动OpenAI、Anthropic等AI公司也开始探索自主芯片路线。今年6月,OpenAI发布 首款 定制推理芯片“Jalapeño”;近期,Anthropic也被曝正与三星洽谈芯片制造合作。 Alphabet此前宣布,为推动人工智能战略发展,计划投入1800亿至1900亿美元资金。随着AI基础设施投资规模扩大,芯片效率成为衡量投入回报的重要指标。Frozen v2的相关消息公布后,市场信心有所提升,Alphabet股价在报道发布后的周一早盘上涨约3%,投资者关注其AI投入能否转化为长期竞争优势。 随着AI模型规模持续增长,自研芯片正成为科技企业构建算力生态、提升成本控制能力的重要战略方向。谷歌此次布局也反映出AI产业竞争正从模型能力延伸至芯片、基础设施和全栈技术体系的竞争。
FlashRT:引导智能体部署实时多模态应用的代理驾驭工具 ↗
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing serving systems and auto-parallelism compilers commit to limited transformations and fixed workload assumptions, so achieving high performance on a new application requires hand-crafting an efficient implementation. We present FlashRT, an agent harness that guides coding agents to lift simple developer-written reference implementations into optimized multi-GPU deployments that flexibly weigh target metrics like latency and throughput. Using a new chain-of-program paradigm, FlashRT directs a generic coding agent through a multi-pass transformation process where an agent transforms the reference into an intermediate representation (IR) to capture data dependencies and persistent-state scopes, validates this IR via a sequential interpreter, and performs static analyses to identify candidate transformations. Then, the agent iteratively implements, verifies, and benchmarks each candidate under a measurement-gated optimization loop to produce effective deployments that span different hardware budgets. Across various applications, including video world models and multimodal LLMs, FlashRT converts reference implementations into highly efficient deployments, delivering up to ~70x latency reduction and 2.8x throughput improvement on NVIDIA B200 GPUs. On AMD MI355X GPUs, FlashRT matches the peak latency reduction while increasing peak throughput improvement to 3.6x, demonstrating that agent-driven optimization can be more scalable on platforms with less mature expert optimization. In fact, for Qwen3-Omni text-to-audio inference, FlashRT reduces response latency by 65% compared to the expert vLLM-Omni implementation on AMD MI355X.
LLM即教练:面向不可验证任务的体验式学习 ↗
Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.
分布偏移下用于忠实生成的令牌级离线策略学习 ↗
We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens. Experiments on document summarization tasks show that TOPL achieves strong out-of-distribution generalization across 11 datasets against a diverse set of sequence-level and token-level baselines. We further demonstrate that TOPL transfers effectively to machine translation, suggesting that its benefits generalize across different faithful generation tasks. Through ablation studies, we confirm that our token-level learning signal is critical to good performance; sequence-level analogues do not confer similar benefits. Finally, we show that TOPL induces interpretable model updates: the LoRA adapters learned through TOPL function as linear classification heads and steering vectors.
RynnBrain 1.1:迈向更强能力和通用性的具身基础模型 ↗
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.
HOMIE:通过多模态智能增强的以人物为中心的视频个性化 ↗
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page:
SWE-Pruner Pro:编程大语言模型早已知道该裁剪什么 ↗
Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent itself encodes internal representations indicating the relevance of code context when reading tool output. Based on this finding, we propose SWE-Pruner Pro, which prunes tool outputs directly inside the agent. Concretely, a small head turns the agent's own internal representations into a keep-or-prune label for each line, with a length-aware embedding keyed to each tool output's line count. Across two open-weight backbones and four multi-turn benchmarks, SWE-Pruner Pro saves up to 39% of prompt and completion tokens while preserving task quality, with bounded inference overhead. Notably, on MiMo-V2-Flash SWE-Pruner Pro additionally raises the SWE-Bench Verified resolve rate by +3.8% and the long-context Oolong accuracy by +2.2 points.
DiFA:扩散模型的推理时前向过程对齐 ↗
The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of numerical integration. This perspective casts the model as an exact estimator, neglecting the inherent statistical uncertainty of the denoising process. In this work, we propose Forward-Process Aligned Diffusion prediction (DiFA), a training-free framework that reframes inference-time data prediction refinement as a sequential state estimation problem. Rather than reusing past outputs solely for numerical integration, DiFA treats iterative data predictions along the reverse trajectory as correlated observations to build a forward-aligned temporal consensus. Inspired by Kalman filtering, this consensus aggregates historical predictions according to structural consistency and noise-level compatibility. To counteract the over-smoothing tendency of temporal consensus, we introduce a deviation guidance mechanism to adaptively preserve residual details. Empirically, DiFA yields significant improvements on CIFAR-10 and ImageNet across the evaluated metrics, including FID, IS, and FD-DINOv2, demonstrating that aligning inference with the forward statistical structure substantially improves generative fidelity.
TimeLens2:基于多模态大语言模型的通用视频时序定位 ↗
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Existing training strategies are misaligned with this set-valued task: long-video labels often rely on brittle one-pass annotation, while reinforcement-learning rewards either fail to distinguish non-overlapping predictions or require fragile segment matching. TimeLens2 treats temporal evidence as an interval set throughout supervision and optimization. TimeLens2-93K constructs reliable multi-span supervision through caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. Our temporal Wasserstein reward computes exact one-dimensional \(W_1\) between uniform distributions over merged interval supports, providing dense, matching-free feedback under unequal cardinalities and equivalent fragmentation; temporal IoU complements it with precise-overlap feedback. Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters. The 2B, 4B, and 8B variants improve over their Qwen3-VL backbones by 14.2, 13.0, and 18.1 mIoU points, respectively.
EvolvingWorld:面向互动文学世界中角色扮演智能体与世界模型协同演化的开放模式框架 ↗
This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or isolated scene generation, failing to capture how characters and worlds evolve together over time. To address this, EvolvingWorld models literary simulation as a long-horizon process where characters interact, scenes progress, and character and world states are persistently updated. Unlike prior systems relying on fixed schemas, EvolvingWorld adopts an open-schema framework to support simulation across diverse literary worlds. The framework consists of two coupled modules: a Character Agent for multi-character role-play and persistent profile evolution, and an LLM-based World Model for global and location/entity-level state maintenance and scene progression. Based on this architecture, we formulate 7 trainable tasks for scene initialization, interaction generation, and state update. We construct a dataset from 57 books, producing 138,596 supervised training samples and 222 snapshots for testing. Furthermore, we introduce a trajectory-level LLM-as-Judge evaluation protocol spanning 10 dimensions and 20 metrics. Experiments show that EvolvingWorld can improve long-horizon simulation by effectively maintaining persistent, coherent character and world development.
面向大语言模型后训练的蒸馏强化学习 ↗
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restricting OPD to within-family distillation. We propose Distilled Reinforcement Learning (Distilled RL), which integrates teacher supervision into the RL objective to provide fine-grained guidance, selectively transfer new knowledge and avoid unconditional imitation. Distilled RL contains three components: reverse importance sampling with clipping, negative sample reset, and sequence-level geometric normalization. Through a concise and interpretable case study, we demonstrate that Distilled RL can effectively transfer previously unavailable knowledge from a teacher model to a student model. Extensive experiments across both within-family and cross-family distillation settings show that Distilled RL substantially outperforms standard RL and OPD in terms of both pass@1 and pass@k. Our code is available at
10:00 更新
5 条Reverse-engineering is cheap now ↗
I keep hearing anecdotes from people who used coding agents to reverse-engineer and automate devices in their homes. I think this is an interesting illustration of the impact of the reduced cost of writing code. Prior to agents, it was entirely possible to reverse-engineer home devices. The problem was the ROI - was it really worth all of that effort? More importantly, any experienced programmer knows that undocumented, unstable APIs like that may well change or break in the future. Is that initial work worth the effort if you're committing yourself to a frustrating cycle of maintenance in the future? Coding agents change that equation entirely. The effort to get a simple automation working has dropped, as has the cost of trying and failing to get it to work. Since the code is so cheap, the idea of having to maintain it in the future - or throw it away and start again - carries way less psychological baggage. Tags: reverse-engineering, coding-agents, ai-assisted-programming, generative-ai, ai, llms
23:00 更新 · 等待中
本批次发布后,条目会出现在这里。
次日 03:00 补充更新 · 等待中
本批次发布后,条目会出现在这里。
没有符合当前条件的资讯。