AI周报第517期:当AI再无可窃取的内容时,会发生什么?

内容来源:https://aiweekly.co/issues/what-happens-when-ai-runs-out-of-content-to-steal
内容总结:
人工智能遭遇“数据荒”:高质量人类文本成稀缺资源
随着AI生成内容充斥互联网,全球仍存在海量未利用数据,但曾支撑大语言模型(LLM)发展的廉价、洁净且自由获取的文本正面临污染,知识产权争议加剧,获取成本日益攀升。本周,有报道称AI公司开始购买旧书籍,而英伟达则发布了一款可通过视频、动作和合成后果训练机器人的模拟器。
行业动态:开源与闭源路线分化,算力争夺白热化
- 模型市场双向发展:中国公司月之暗面(Moonshot)开源了拥有2.8万亿参数、单次激活约1040亿参数的混合专家模型Kimi K3,权重可免费下载;Anthropic则推出价格减半但性能接近旗舰的Opus 5模型,闭源路线降低门槛。
- 基建军备竞赛升温:Anthropic宣布将部署最高2吉瓦的AMD MI450 GPU,AMD承诺未来最高投资50亿美元;英伟达投资“安全超级智能”(SSI)公司,并称其Vera Rubin系统将提升该实验室算力一个数量级。
- 成本与监管博弈:美国政府推动超200家电力公司、开发商等承诺,由大型数据中心自费承担新增电力基础设施,避免转嫁居民;欧盟则延长部分高风险AI规则期限,同时明确禁止生成非自愿性内容及儿童虐待材料的系统。
数据困境:下一个护城河是“可靠经验”
- 旧书成新“矿藏”:AI公司收购纸质旧书,因其内容可确定早于AI生成时代,纯净人类文本成为有采购价值的战略物资。
- 训练数据配方受保护:加州要求公开训练数据集文档,xAI正就此提出法律挑战,训练数据构成已成为商业机密。
- AI开始“自我编写教材”:一项名为“技能自玩”(Skill Self-Play)的技术让AI自行生成任务并验证结果,训练材料从固定资源转为系统自我制造的产物。
- 环境即数据:英伟达Cosmos-H-Dreams通过外科手术视频和机器人运动学学习,并模拟后续结果,将环境反馈转化为训练数据。
关键结论:模型市场正分裂为“开源下载”与“闭门降价”两派;廉价智能依赖昂贵基建,社会正争论由谁承担外部成本;彻底解决AI数据来源,领域控制权——包括授权的人类档案、可验证的合成实践或专有物理环境——才是可持续优势。出版商、模拟器开发商和机器人运营商正成为模型产业链的核心环节。
其他要闻速览
- 儿童对AI说不:调查发现,许多孩子认为生成式AI“令人毛骨悚然”甚至“很酷但没用”,采纳问题正变成品味与身份认同问题。
- 维基百科反击AI洗稿:推出新流程,允许编辑在特定条件下移除疑似AI生成的贡献,恢复者需对内容及来源负责。
- AI或致人类表达趋同:跨学科研究警告,大规模使用可能强化主流风格,边缘化少数声音,风险不仅是模型自噬,更是人类语言日益同质化。
- AI代理自主编写训练课程:模型系统开始自我生成任务并验证结果,训练材料由系统协助制造,而非有限存量。
- 英伟达将动作转化为数据:新技术从手术视频和机器人运动学中学习并模拟推演,环境本身成为新的“文档”。
- 信任危机频发:Claude的公开分享链接被搜索引擎收录,提示“公开”与“可被检索”是截然不同的产品承诺。
- 值得关注:开源权重模型对闭源商业模式的冲击,Anthropic对强大AI发布的安全边界主张,以及AI资源隐藏债务等议题。
- 趣闻:Spotify不为AI音乐加标签,听众自建数据库识别;一初创公司开放上百台机器人供在线操作绘画、剑术、化学实验;ChatGPT婉拒模仿斯蒂芬·金文风,但提供“小镇阴森感”的类似风格创作。
中文翻译:
世界仍然蕴含着海量未使用的数据。但曾推动第一轮大语言模型繁荣的廉价、清洁且无需许可的文本,正逐渐被AI输出所污染、被其所有者争夺,且替换成本日益高昂。本周,据报道AI公司正在购买旧书,而英伟达发布了一款模拟器,可通过视频、动作和合成后果来训练机器人。
获取更多AI周刊内容
更多信号,更少噪音——选择你的频道。
你正在阅读每周简报。以下是关注报道的其他方式——所有频道均免费,随时可退出。
→ 探索16个深度专题
每周主题通讯:生成式AI、机器学习、AI商业、机器人技术、前沿研究、地缘政治、医疗保健等。浏览全部16个深度专题 →
→ AI突发警报
早晨Espresso之后发生的重要动态,不重复你已读过的内容。通常不会额外发送邮件;最多一条下午更新,外加罕见的紧急例外。获取突发警报 →
→ AI今日新闻(实时)
扫描器发现新闻时实时更新的动态面板:过去48小时的评分报道、每周实体动向,以及覆盖113家AI公司、人物和主题的季度趋势线。打开AI今日新闻 →
野外动态
专家网络发掘的三条链接,尚未在Espresso中发送给你。
- 儿童可能是AI的第一批反炒作群体。《连线》杂志发现孩子们形容生成式AI为“ creepy”、恶心且不酷。普及不仅是一个能力问题;它正在变成一个品味和身份认同的问题。阅读孩子们的说法
- 维基百科正在设计对AI复制的免疫反应。其新流程允许编辑在满足特定条件时移除疑似AI生成的贡献;任何恢复这些内容的人需承担审查内容及其来源的责任。阅读该流程
- 大语言模型可能将人类表达拉向同一中心。一项跨学科综述认为,广泛使用会强化主流风格并边缘化另类声音。风险不仅在于模型自我学习;还在于人们开始听起来更加相似。阅读该综述
快讯速览
开放前沿
Kimi K3的权重可下载。Opus 5展示了另一条更廉价访问的路径。
- Kimi K3使2.8万亿参数模型可下载。月之暗面发布了2.8万亿参数混合专家模型的权重,每个token激活1040亿参数。
- Opus 5使Anthropic的近前沿层级更便宜。Anthropic表示Opus 5以一半价格接近Fable 5的水平。
算力圈地运动
模型只有在某方投入硬件、电力和资本之后,使用成本才会下降。
- Anthropic承诺以吉瓦规模采用AMD芯片。Anthropic计划部署高达两吉瓦的AMD MI450 GPU;AMD还承诺未来最高50亿美元的股权投资。
- 英伟达为SSI带来数量级的算力跃升。英伟达投资了Safe Superintelligence,并表示Vera Rubin系统将使该实验室的算力提升一个数量级。
谁买单
电网和合规成本正成为明确的政治选择。
- 华盛顿试图将数据中心账单从家庭身上移开。白宫表示,超过200家公用事业公司、开发商、合作社和州承诺,大型数据中心应为其所需的新电力基础设施出资。
- 欧洲延长AI期限——并增加一条硬性禁令。欧洲延长了高风险AI时间表的部分内容,同时禁止生成未经同意的性内容或儿童虐待材料的系统。
进入敏感系统
一个共享链接在技术上可以是公开的,而用户仍将其体验为私密的。
- Claude的公开共享链接变得可被搜索。《连线》杂志在搜索结果中发现了用户创建的Claude公开链接——这提醒人们,“共享”和“可被发现”是完全不同的产品承诺。
当AI耗尽干净的人类文本时会发生什么?
上周旧书的故事看起来像一个奇怪的采购细节。把它放在接下来三个链接旁边,它就变成了后爬取时代AI经济的地图。
- 旧书成为AI库存。据报道,AI公司正在购买印刷书籍,因为可以保证它们早于AI生成内容。这不是怀旧;干净的人类文本现在具有采购价值。
- xAI对抗“配料标签”。加州要求提供训练数据集的一般性文档;xAI正在挑战该规定。语料配方现在成了竞争性信息。
- Agent开始编写自己的课程。Skill Self-Play让Agent生成任务并验证结果。训练材料变成了系统协助制造的东西,而不是最终会被耗尽的一堆固定数据。
- 英伟达将运动转化为训练数据。Cosmos-H-Dreams从手术视频和机器人运动学中学习,然后模拟接下来会发生什么。新的“文档”是一个带有后果的环境。
要点:下一个数据护城河不是“更多内容”,而是对可靠经验的控制——授权的人类档案、可验证的合成实践或专有的物理世界环境。出版商、模拟器构建者和机器人运营者成为模型堆栈的一部分。
核心要点
- 模型市场正在一分为二:Kimi K3扩大了可下载的范围,而Opus 5降低了闭源模型访问的价格。
- 廉价的智能建立在昂贵的基础设施之上。AMD和英伟达的承诺使算力竞赛清晰可见;费率支付方的承诺则追问谁吸收其外部成本。
- 信任仍然在默认设置处破裂。Claude用户创建了公开共享链接,但搜索引擎使“公开”的可发现性远超许多人的预期。
- 上游,稀缺资产正在变成可靠经验——而非原始数量。下一代持久的AI优势可能就在那里。
值得阅读
阅读论证本身,而不是又一篇复述:一篇分析解释了为什么Kimi K3给闭源模型经济学带来压力;Anthropic划定了它希望围绕强大发布设定的安全线。
- Kimi K3改变了开放权重的经济学。Nathan Lambert认为,接近前沿的开放权重加速了扩散,同时压缩了资助闭源模型实验室的利润空间。
- Anthropic表示不会全面禁止开放权重——但要求测试。Dario Amodei称非危险的开放权重模型是公共品,拒绝全面禁令,并主张在强大发布前进行能力测试。
值得观看
两条专家分享的新视频,在AI TV上策展。
- 过度优化与RLHF的坏名声——Nathan Lambert
- AI公司隐藏的债务比你想象的更多——The Tech Report
等等,什么?
- Spotify不会标记AI音乐,所以听众自己建立了注册库。Spotify不标记AI生成的曲目,因此SoullessMusic和SlopTracker正在建立独立的注册库作为替代。
- 一家初创公司将100台机器人上线,用于绘画、击剑和混合化学品。Enigma正在开放对100多台机器人的在线访问,这些机器人可以绘画、击剑和执行简单的化学实验。
- ChatGPT拒绝模仿斯蒂芬·金风格——然后提供“小镇恐惧”。被要求模仿斯蒂芬·金的风格时,ChatGPT拒绝了直接模仿,但提供了氛围恐怖和具有类似感受的“小镇恐惧”。
本周投票
哪个来源对AI能力的下一次跃升最为重要?
上周,133位读者参与了投票:
在本周的围堵失败之后,你会把下一笔AI安全资金花在哪里?
哪个来源对AI能力的下一次跃升最为重要?
周五见。
Alexis
英文来源:
The world still contains vast amounts of unused data. But the cheap, clean and permissionless text that powered the first LLM boom is becoming polluted by AI output, contested by its owners and costly to replace. This week, AI companies were reportedly buying old books while Nvidia released a simulator that teaches robots through video, motion and synthetic consequences.
Get more from AI Weekly
More signal, less noise — pick your channels.
You're reading the weekly brief. Below are the other ways to follow the story — every channel free, easy to leave.
→ Explore 16 deep divesWeekly topic-specific newsletters: Generative AI, Machine Learning, AI in Business, Robotics, Frontier Research, Geopolitics, Healthcare, and more.Browse all 16 deep dives →
→ Breaking AI alertsImportant developments that happen after your morning Espresso, without repeating what you already read. Usually no extra email; at most one afternoon update, plus a rare critical exception.Get breaking alerts →
→ AI News Today (live)Live dashboard updated as the scanner finds news: scored stories from the last 48 hours, weekly entity movers, and quarterly trend lines across 113 AI companies, people, and topics.Open AI News Today →
In the Wild
Three links the expert network surfaced that we have not already sent you in Espresso.
- Children may be AI’s first anti-hype demographic. Wired found kids describing generative AI as creepy, disgusting and uncool. Adoption is not only a capability problem; it is becoming a question of taste and identity. Read what the kids said
- Wikipedia is designing an immune response to AI copy. Its new process lets editors remove suspected AI-generated contributions when defined conditions are met; anyone restoring them assumes responsibility for reviewing the content and its sources. Read the process
- LLMs may pull human expression toward the same center. A cross-disciplinary review argues that widespread use can reinforce dominant styles and marginalize alternative voices. The risk is not only models learning from themselves; it is people beginning to sound more alike. Read the review
Quick Hits
The Open Frontier
Kimi K3’s weights are downloadable. Opus 5 shows a separate route to cheaper access. - Kimi K3 makes a 2.8-trillion-parameter model downloadable. Moonshot released the weights for a 2.8-trillion-parameter mixture-of-experts model that activates 104 billion parameters per token.
- Opus 5 makes Anthropic’s near-frontier tier cheaper. Anthropic says Opus 5 comes close to Fable 5 at half the price.
The Compute Land Grab
The models get cheaper to use only after somebody commits the hardware, power and capital. - Anthropic commits to AMD at gigawatt scale. Anthropic plans to deploy up to two gigawatts of AMD MI450 GPUs; AMD also committed to a future equity investment of up to $5 billion.
- Nvidia gives SSI an order-of-magnitude compute jump. Nvidia invested in Safe Superintelligence and says Vera Rubin systems will increase the lab’s compute by an order of magnitude.
Who Pays the Bill
The grid and compliance costs are becoming explicit political choices. - Washington tries to move the data-center bill off households. The White House says more than 200 utilities, developers, cooperatives and states pledged that large data centers should fund the new power infrastructure they require.
- Europe extends AI deadlines—and adds a hard prohibition. Europe extended parts of the high-risk AI timetable while prohibiting systems that generate non-consensual sexual or child-abuse material.
Into Sensitive Systems
A share link can be technically public while users still experience it as private. - Claude's public share links became searchable. Wired found user-created public Claude links in search results—a reminder that “shared” and “discoverable” are very different product promises.
What happens when AI runs out of clean human text?
Last week’s old-books story looked like an odd procurement detail. Put it beside the next three links and it becomes a map of the post-crawl AI economy. - Old books become AI inventory. AI companies are reportedly buying printed books because they are guaranteed to predate AI-generated content. That is not nostalgia; clean human text now has procurement value.
- xAI fights the ingredient label. California asks for general documentation of training datasets; xAI is challenging the rule. The corpus recipe is now competitive information.
- Agents begin writing their own curriculum. Skill Self-Play has agents generate tasks and verify the results. Training material becomes something the system helps manufacture, not a fixed pile it eventually exhausts.
- Nvidia turns movement into training data. Cosmos-H-Dreams learns from surgical video and robot kinematics, then simulates what happens next. The new “document” is an environment with consequences.
The point: the next data moat is not “more content.” It is control over reliable experience—licensed human archives, verifiable synthetic practice or proprietary physical-world environments. Publishers, simulator builders and robot operators become part of the model stack.
Key Takeaways - The model market is splitting in two: Kimi K3 expands what can be downloaded, while Opus 5 lowers the price of closed-model access.
- Cheap intelligence rests on expensive infrastructure. The AMD and Nvidia commitments make the compute race visible; the ratepayer pledge asks who absorbs its external cost.
- Trust still breaks at the default. Claude users created public share links, but search engines made “public” far more discoverable than many expected.
- Upstream, the scarce asset is becoming reliable experience—not raw volume. That is where the next durable AI advantage may sit.
Worth Reading
Read the argument, not another recap: one analysis explains why Kimi K3 pressures closed-model economics; Anthropic draws the safety line it wants around powerful releases. - Kimi K3 changes the economics of open weights. Nathan Lambert argues that near-frontier open weights accelerate diffusion while pressuring the margins that finance closed-model labs.
- Anthropic says no blanket open-weight ban—but demands testing. Dario Amodei calls non-dangerous open-weight models a public good, rejects a categorical ban and argues for capability testing before powerful releases.
Worth Watching
Two fresh expert-shared videos, curated on AI TV. - Over-Optimization and RLHF’s Bad Reputation — Nathan Lambert
- AI Companies Are Hiding More Debt Than You Think — The Tech Report
Wait, What? - Spotify will not label AI music, so listeners built their own registries. Spotify does not label AI-generated tracks, so SoullessMusic and SlopTracker are building independent registries instead.
- A startup put 100 robots online to paint, sword-fight and mix chemicals. Enigma is opening online access to more than 100 robots that can paint, fight with swords and perform simple chemistry experiments.
- ChatGPT refuses Stephen King’s style—then offers “small-town dread”. Asked for Stephen King’s style, ChatGPT refused the exact imitation but offered atmospheric horror and “small-town dread” with a similar feeling.
This week’s poll
Which source will matter most for the next jump in AI capability?
Last week, 133 of you voted:
After this week’s containment failures, where would you spend the next AI-security dollar?
Which source will matter most for the next jump in AI capability?
Back Friday.
Alexis
文章标题:AI周报第517期:当AI再无可窃取的内容时,会发生什么?
文章链接:https://news.qimuai.cn/?post=4686
本站文章均为原创,未经授权请勿用于任何商业用途