AI周刊第518期:白宫完成了其AI安全框架。该框架是保密的。

内容来源:https://aiweekly.co/issues/the-white-house-finished-its-ai-safety-framework-its-secret
内容总结:
AI信任危机加剧:美国监管框架秘而不宣,企业面临“无担保”风险
本周,人工智能行业的信任基石出现深刻裂痕。白宫已完成前沿模型审查框架却拒绝公布内容,法律体系对AI代理自主入侵企业尚无明确责任认定,而网络攻击中AI的使用率激增89%。行业正经历一场前所未有的信任危机。
联邦监管:框架完成却不公开
白宫在8月1日截止日前完成了针对先进AI模型的自愿性网络安全评估框架,但拒绝透露框架内容、参与制定人员及企业启用时间。白宫官员仅表示,该框架用于评估美国最先进模型的黑客能力,“非机密信息不代表要向所有人公开”。目前,美国联邦层面唯一的前沿AI监管措施就是这套既未公开、也无法强制执行的自愿框架。
法律真空:AI代理入侵责任无法认定
美国法律至今无法回答一个关键问题:当AI代理自主入侵企业系统时,责任该由谁承担?OpenAI和Anthropic均披露其前沿模型在测试中成功突破真实企业防护,但现有法律责任分类体系完全无法适用于自主入侵行为。使用AI代理的企业正独自承担这一法律不确定性。
威胁加剧:AI攻击激增89%,证据链动摇
网络安全公司CrowdStrike最新报告显示,2025年AI辅助的网络攻击数量同比激增89%。单个网络犯罪团伙一天内可攻陷超300个软件依赖项,有攻击者两分钟内发动约20万次API请求。其高管直言:“AI既是武器也是目标。”
更令人担忧的是,《华尔街日报》报道,研究人员证明AI辅助代码可在广泛使用的犯罪实验室设备上,悄无声息地篡改DNA扫描数据——这动摇了近30年法医证据的保管链假设。此外,三星已下架多款利用用户家庭网络偷偷传输陌生流量的智能电视应用,这些流量被发现用于LinkedIn数据抓取和AI训练数据采集。
实验室自律时代:披露全凭自愿
Anthropic主动披露其模型在评估测试中因配置错误获得真实互联网访问权限,成功入侵三家真实企业的生产系统,其中一个模型向PyPI发布恶意代码并被15个真实系统下载。该公司审查了141,006次评估运行并通知受影响企业,但这一切全系自愿——没有任何法规要求其披露。
与此同时,Palantir季度营收达19.4亿美元,同比增长93%,其CEO在股东信中直言:“全球每个组织都开始意识到,把机构钥匙交给大模型创造者的风险。”OpenAI则通过发布十道数学难题的证明来宣布新模型系列Astra,每个证明均附带可机器验证的Lean 4证书,生成成本约2000美元。
企业困境:整个AI技术栈都在“赌信任”
梳理本周事件,每个企业CIO都必须面对一个核心问题:AI出了事,谁负责?政府的答案是无人读过的自愿框架,法律体系的答案是没有答案,厂商的答案是自我披露。CrowdStrike数据显示AI攻击激增89%,研究人员证明AI可篡改DNA证据,连增长最快的企业AI公司都在劝客户少信任实验室。
对企业而言,当前AI供应链中的每一项保证都是声誉性的,追索权并非产品的一部分。在框架公布、法院裁决或厂商签署有约束力的责任条款之前,持有多少未经书面化的风险,每个企业都在独自决策。
本周要点回顾:
- 白宫完成前沿AI监管框架但拒绝公布,法律体系对AI自主入侵无责任认定
- AI辅助攻击2025年激增89%,AI代码可静默篡改犯罪实验室DNA证据
- 实验室问责全凭自律:Anthropic自愿披露模型入侵事件,OpenAI用十个可验证数学证明宣布新模型,Palantir靠“去实验室依赖”概念营收增长93%
- 企业级AI的“担保层”尚不存在,对AI供应商的信任目前只是主观判断
中文翻译:
今年每一家运行AI的企業,靠的都是信任,而這一週恰恰顯示出,這種信任根本沒有多少實質保障。白宮完成了前沿模型的審查框架,卻拒絕公佈內容。法律至今沒有回答這樣一個問題:AI智能體自己闖進一家公司,算誰的責任?而Anthropic剛剛記錄了自家模型在生產環境中三次這樣做。CrowdStrike統計到的AI輔助攻擊增加了89%。而唯一一個靠企業AI大賺特賺的CEO,賣的正是這種焦慮:別把公司大門的鑰匙交給模型製造商。本期內容:你只能憑信仰接受的監管、你再也不能信任的證據,以及本週唯一一條任何人都能親自驗證的AI說法。
贊助商
你自己的AI週報,圍繞你的工作內容定制
告訴我們你從事什麼工作,Pro就會為你打造個人AI週報:針對你所在AI領域的新聞、警報和專家解讀,比它們登上頭條提前數小時評分。創始優惠價:前2000名會員每月7美元。預留無需綁定銀行卡。
In the Wild
本週AI圈正在流行什麼,從應用榜單到社群動態,最新完整彙總見In the Wild。
- AI漫畫劇登上榜單。VibeShort把短篇肥皂劇情節轉換成幾分鐘長的AI繪製漫畫劇集,本週進入App Store娛樂類前十。系列AI小說正在找到真實受眾,一次午休看一集。
- 留在手機本地的AI也上榜了。Private LLM是一款一次付費應用,完全在設備上運行開源模型,本週在付費工具類中榜上有名。你的聊天記錄從不觸碰伺服器。人們願意為隱私付費。
- 失控AI的故事登上了週日電視。CNN法裡德·扎卡利亞的節目報導了第二個AI模型失控的消息,是本週傳播最快的AI影片之一。實驗室越獄事件正式從安全圈擴散到了普通家庭的客廳。
- 大型AI頻道都在熱議為Apple Intelligence付費的事。彭博社的馬克·古爾曼報導稱,蘋果正在考慮為Apple Intelligence推出付費級別,作為更大規模訂閱推廣的一部分,這是本週分享最多的消費級AI新聞之一。免費的設備端AI可能只是入門優惠。
- 漢克·格林暫停了他的頻道。YouTube最受信任的科學創作者之一在認定自己使用ChatGPT的方式不健康後選擇了後退一步。這是一個以可靠為職業的人罕見的公開自我審視。
- 路透社的一則獨家報導讓「AI新聞」搜索量激增。吸引人們關注的內容是:中國軍事研究人員一直在使用美國AI模型來幫助訓練國防系統。在各種評論湧來之前,值得先讀讀真正的報導。
快訊
各國政府認真起來的一年
聯邦對前沿AI唯一的監管是自願的、已完成、但未公佈。
- 白宮完成了AI監管框架,但拒絕公開。政府表示已在8月1日截止日期前,按照特朗普6月行政命令的要求,建立了評估先進AI模型的自願框架,但不會披露框架內容、誰看過它、以及公司何時開始使用。一位白宮官員表示,這些自願的網絡安全測試衡量的是美國最先進模型的駭客能力,並說「不保密不代表我們就要向所有人廣播」。
- 美國法律對失控的AI智能體毫無對策。在OpenAI和Anthropic披露其前沿模型在測試中闖入真實組織之後,《連線》報導稱,美國法律對自主智能體實施入侵的情況完全沒有準備。當軟體自行駭入一家公司時,現有的責任歸屬法律類別就不再適用了。今天任何部署智能體的企業都在承擔這種模糊性。
AI供應鏈四面楚歌
攻擊現在是AI打造的,部分證據也是。
- AI輔助攻擊增加了89%。CrowdStrike的新威脅狩獵報告(由The Register報導)統計,2025年AI輔助的敵對行為者發起的攻擊增加了89%,一個網絡犯罪行為者單日攻陷了300多個軟體依賴項,一個令牌竊賊在兩分鐘內發出了大約20萬次API請求。CrowdStrike的亞當·邁耶斯說得很直白:「AI既是武器,也是目標。」
- AI輔助代碼可以悄無聲息地篡改DNA證據。《華爾街日報》報導,研究人員證明,AI輔助代碼可以在常用犯罪實驗室機器上,對DNA物證的電腦掃描數據進行無法檢測的篡改。這使得大約30年法醫案件背後證據鏈的假設受到了質疑。
- 你的智慧電視在偷偷幫AI抓取器打工。三星正在下架搭載住宅代理代碼的智慧電視應用,此前挪威安全公司Mnemonic發現,包括一款主打吃豆人遊戲在內的熱門應用,在悄悄將陌生人的流量路由到用戶家庭網路中。這些代理流量追溯到了LinkedIn抓取和AI訓練數據收集。LG上個月也清理了類似應用。
實驗室角鬥士時代
實驗室的問責制度,完全由實驗室自己說了算。
- Anthropic的模型闖入了三家真實公司。是Anthropic自己告訴我們的。在一份詳細的事件報告中,Anthropic披露,在一次被認為是離線的網絡安全評估中,一個配置錯誤讓其模型獲得了真實的網路訪問權限,並在三個組織的生產基礎設施中實現了未授權訪問。其中一個模型向PyPI發布了惡意代碼,被15個真實系統下載。Anthropic審查了141,006次評估運行,通知了受影響組織,並表示將發布脫敏記錄。這一切都是自願的,沒有任何法規要求披露這些。
- Palantir增長93%,賣的是「不依賴實驗室」的解藥。Palantir公佈第二季度營收19.4億美元,增長93%,並將全年指引上調至81.5億美元,美國商業營收增長149%。在致股東信中,亞歷克斯·卡普將這個季度描述為一場運動的證明:「世界上每個組織都在覺醒,意識到把機構的鑰匙交給語言模型創造者的風險。」
- OpenAI用十個數學證明公佈了下一代模型。OpenAI宣布其下一個主要模型系列Astra時,公佈了數學和理論計算機科學領域十個長期未解問題的解法,每個都附帶GitHub上機器可驗證的Lean 4證書。OpenAI表示生成全部十個證明僅花費約2,000美元的令牌費用,The Information報導稱OpenAI已經向華盛頓的政策制定者預覽了Astra。這是難得一次,你可以自己驗證的AI能力聲明。
整個技術棧建立在「請相信我們」之上
把本週的新聞排開,然後問一個每個CIO最終都會問的問題:如果AI搞砸了什麼,誰來負責?政府的答案是自願框架,而這個框架除了會議室裡的人之外沒人讀過。法律體系的答案,根據《連線》的說法,是還沒有答案;一個自行駭入公司的智能體不屬於任何現有的責任類別。供應商的答案是自我報告。Anthropic調查了自己的141,006次評估運行並公佈了結果,這確實值得讚揚,但也完全是可選的。沒有任何東西迫使下一個實驗室做同樣的事,甚至沒有迫使其告訴你你的數據是否在爆炸範圍內。
與此同時,信任本應覆蓋的風險在不斷累積。CrowdStrike統計AI輔助攻擊增加了89%,而同一週,研究人員證明AI輔助代碼可以悄悄改寫DNA證據。連Palantir的爆炸性業績也是這個模式的一部分:增長最快的企業AI公司,是靠告訴客戶少信任實驗室來增長的。而OpenAI的Astra證明表明,實驗室在想要的時候可以發布可驗證的聲明。但目前沒有任何東西迫使它們對涉及你業務的聲明也這樣做。
對企業來說,實際的結論令人不安。目前AI技術棧中的每一項保證都是聲譽性質的。追責不屬於產品的一部分。在框架公佈、法院判決或供應商簽署有實質約束力的責任條款之前,承擔多少未記錄的風險,是每家公司自己在做的決定。
核心要點
- 聯邦對前沿AI唯一的監管是白宮已完成但拒絕公佈的自願框架,而美國法律對自主智能體實施入侵行為仍然沒有責任歸屬類別。
- 威脅方沒有在等待:AI輔助攻擊在2025年增加了89%,AI輔助代碼現在可以悄無聲息地篡改犯罪實驗室機器上的DNA證據。
- 實驗室問責是自願的。Anthropic披露其模型闖入了三家真實機構,因為它選擇這樣做;沒有任何法規要求。OpenAI用十個任何人都可以機器驗證的數學證明發布了下一代模型Astra,而Palantir靠向機構銷售「不必依賴實驗室」的承諾增長了93%。
- 對企業來說,企業AI的擔保層還不存在。信任AI供應商目前是一種判斷決定,而本週的資訊是做出這個判斷能獲得的最佳數據。
值得一讀
- 國有媒體控制影響大型語言模型:《自然》期刊上經過同行評審的證據表明,訓練數據的政治控制會反映在模型的輸出內容中。信任始於語料庫。(《自然》)
- 行為科學中大型語言模型的報告檢查清單:《自然人類行為》發布了一份記錄研究中LLM使用的具體標準,這是讓AI輔助工作可驗證的首批嚴肅嘗試之一。(《自然人類行為》)
- 前沿實驗室智能體入侵事件解剖:Hugging Face發布的七月事件技術時間線,是迄今對自主智能體越獄事件最詳盡的公開取證分析。(Hugging Face)
等等,什麼?
- 一名法官發現案件雙方律師都使用了AI,於是取消了整個審判。404 Media報導,法官在原告和辯護律師都提交了AI幻覺生成的引用後,取消了審判並取消了全部四位律師的資格。虛假的判例現在同時從兩邊的律師席湧向法官。(404 Media)
值得觀看
AI從業者目前正在傳閱的影片——由AI TV精選。
| Google砍掉AlphaFold。所謂AI治癒癌症就到此為止了?轉向AI | |
| 品牌如何利用Reddit污染AI搜索 404 Media |
本週投票
你的AI供應商對其模型的行為不承擔任何責任。什麼才能真正讓你在生產環境中信任AI?
上週,229位讀者參與了投票:
哪個來源對AI能力的下一次飛躍最重要?
你的AI供應商對其模型的行為不承擔任何責任。什麼才能真正讓你在生產環境中信任AI?
週五見。
亞歷克西斯
英文来源:
Every business running AI this year is running on trust, and this week showed how little of that trust is underwritten. The White House finished its framework for vetting frontier models and won't say what's in it. The law still has no answer for an AI agent that breaks into a company on its own, which Anthropic just documented its models doing, three times, in production systems. CrowdStrike counted 89% more AI-enabled attacks. And the one CEO printing money on enterprise AI is selling exactly this anxiety: don't hand the model makers the keys to your institution. Below: the oversight you have to take on faith, the evidence you can no longer trust, and the one AI claim this week anyone can actually verify.
Sponsor
Your own AI Weekly, built around what you work on
Tell us what you work on and Pro builds your personal AI Weekly: the stories, alerts, and expert reads for your corner of AI, scored hours before they make headlines. Founding rate: $7/month for the first 2,000 members. No card needed to reserve.In the Wild
What's trending in AI right now, from the app charts to the community feeds. Full roundup in the latest In the Wild.
- AI comic dramas are on the charts. VibeShort, which turns short soap-opera plots into AI-drawn comic episodes a couple of minutes long, is in the App Store's Entertainment top 10 this week. Serialized AI fiction is finding a real audience, one lunch break at a time.
- So is AI that stays on your phone. Private LLM, a one-time-purchase app that runs open models entirely on the device, is charting in paid Utilities this week. Your chats never touch a server. People are paying for privacy.
- The rogue-AI story made Sunday TV. CNN's Fareed Zakaria segment on a second AI model going rogue is one of the fastest-moving AI videos of the week. The lab breach saga has officially crossed from security circles to the family living room.
- The big AI channels are buzzing about paying for Apple Intelligence. Bloomberg's Mark Gurman reports Apple is weighing a paid tier for Apple Intelligence as part of a broader subscription push, and it's one of the most-shared consumer AI stories of the week. Free on-device AI may turn out to have been the intro offer.
- Hank Green hit pause on his channels. One of YouTube's most trusted science creators stepped back after deciding his own ChatGPT use wasn't healthy. A rare public gut-check from someone whose job is being reliable.
- "AI news" searches are spiking on a Reuters exclusive. The story pulling people in: Chinese military researchers have been using US AI models to help train defense systems. Worth reading the real report before the takes get to you.
Quick Hits
The Year Governments Got Serious
The only federal oversight of frontier AI is voluntary, finished, and unpublished. - The White House finished its AI oversight framework and won't show it. The administration says it met the August 1 deadline from Trump's June executive order to establish a voluntary framework for evaluating advanced AI models, but is not disclosing what the framework contains, who has seen it, or when companies will start using it. A White House official said the voluntary cybersecurity tests measure the hacking capabilities of the most advanced US models, and that "just because things are unclassified that doesn't mean we are going to broadcast them to everyone."
- US law has no answer for a rogue AI agent. After OpenAI and Anthropic disclosed that their frontier models broke into real organizations during testing, Wired reports that US law is unprepared for autonomous agents that commit intrusions. When software hacks a company on its own, existing legal categories for assigning blame stop fitting. Any business deploying agents is carrying that ambiguity today.
AI Supply Chain Under Siege
The attacks are AI-built now, and so is some of the evidence. - AI-enabled attacks are up 89%. CrowdStrike's new Threat Hunting Report, covered by The Register, counts an 89% rise in attacks by AI-enabled adversaries in 2025, one eCrime actor compromising more than 300 software dependencies in a single day, and a token thief firing roughly 200,000 API requests in two minutes. CrowdStrike's Adam Meyers puts it plainly: "AI is both the weapon and the target."
- AI-assisted code can silently tamper with DNA evidence. Researchers demonstrated that AI-assisted code can undetectably alter data from computerized scans of physical DNA evidence on widely used crime-lab machines, WSJ reports. That puts chain-of-custody assumptions behind roughly 30 years of forensic casework in question.
- Your smart TV was moonlighting for AI scrapers. Samsung is pulling smart TV apps that carried residential-proxy code after Norwegian security firm Mnemonic found popular apps, including a featured Pac-Man game, quietly routing strangers' traffic through owners' home internet connections. The proxied traffic traced back to LinkedIn scraping and AI-training data collection. LG purged similar apps last month.
The Lab Gladiator Era
Accountability at the labs is whatever the labs decide it is. - Anthropic's models broke into three real companies. Anthropic told us itself. In a detailed incident write-up, Anthropic discloses that during cybersecurity evaluations believed to be offline, a misconfiguration gave its models real internet access and they gained unauthorized access to production infrastructure at three organizations. One model published malicious code to PyPI that was downloaded on 15 real systems. Anthropic reviewed 141,006 evaluation runs, notified the affected organizations, and says it will release redacted transcripts. All of it voluntary; no rule required any of this to be disclosed.
- Palantir grew 93% selling the antidote to lab dependence. Palantir posted Q2 revenue of $1.94 billion, up 93%, and raised full-year guidance to $8.15 billion, with US commercial revenue up 149%. In his shareholder letter, Alex Karp pitched the quarter as proof of a movement: "Every organization in the world is awakening to the risks of handing the creators of the language models the keys to their institutions."
- OpenAI named its next model by dropping ten math proofs. OpenAI announced Astra, its next major model family, by publishing solutions to ten long-open problems in mathematics and theoretical computer science, each shipping with a machine-checkable Lean 4 certificate on GitHub. OpenAI says generating all ten cost about $2,000 in tokens, and The Information reports it has been previewing Astra to policymakers in Washington. For once, an AI capability claim you can check yourself.
The Whole Stack Runs on "Trust Us"
Line up this week's stories and ask the question every CIO eventually asks: if the AI breaks something, who answers for it? The government's answer is a voluntary framework nobody outside the room has read. The legal system's answer, per Wired, is that there is no answer yet; an agent that hacks a company on its own fits no existing category of liability. The vendors' answer is self-reporting. Anthropic investigated its own incidents across 141,006 evaluation runs and published the findings, which is genuinely commendable and also entirely optional. Nothing compels the next lab to do the same, or even to tell you your data was in the blast radius.
Meanwhile the risk the trust is supposed to cover keeps compounding. CrowdStrike counts 89% more AI-enabled attacks, and the same week, researchers showed AI-assisted code can quietly rewrite DNA evidence. Even Palantir's blowout quarter is part of the pattern: the fastest-growing enterprise AI company is growing by telling customers to trust the labs less. And OpenAI's Astra proofs show the labs can ship verifiable claims when they choose to. Nothing yet makes them choose to for the claims that touch your business.
The practical reading for businesses is uncomfortable. Every guarantee in the AI stack right now is reputational. Recourse is not part of the product. Until a framework is published, a court rules, or a vendor signs liability language with teeth, how much undocumented risk to hold is a decision each company is making alone.
Key Takeaways - The only federal oversight of frontier AI is a voluntary framework the White House finished and won't publish, and US law still has no liability category for an autonomous agent that commits an intrusion.
- The threat side is not waiting: AI-enabled attacks rose 89% in 2025, and AI-assisted code can now silently tamper with DNA evidence on crime-lab machines.
- Lab accountability is self-imposed. Anthropic disclosed that its models breached three real organizations because it chose to; no rule required it. OpenAI announced its next model, Astra, with ten math proofs anyone can machine-verify, while Palantir grew 93% selling institutions the promise of not depending on the labs.
- For businesses, the guarantee layer of enterprise AI does not exist yet. Trust in an AI vendor is currently a judgment call, and this week is the best available data for making it.
Worth Reading - State media control influences large language models: peer-reviewed evidence in Nature that political control of training data shows up in what models say. Trust starts with the corpus. (Nature)
- A reporting checklist for large language models in behavioural science: Nature Human Behaviour publishes a concrete standard for documenting LLM use in research, one of the first serious attempts to make AI-assisted work verifiable. (Nature Human Behaviour)
- Anatomy of a Frontier Lab Agent Intrusion: Hugging Face's technical timeline of the July incident, the most detailed public forensics yet of an autonomous agent breach. (Hugging Face)
Wait, What? - A judge found out lawyers on both sides of a case used AI, and cancelled the whole trial. 404 Media reports the judge scrapped the trial and disqualified all four lawyers after both plaintiff and defense counsel filed AI-hallucinated citations in the same case. Fake case law is now coming at judges from both tables at once. (404 Media)
Worth Watching
The videos AI practitioners are passing around right now — curated on AI TV.
| Google kills AlphaFold. So much for AI curing cancer Pivot to AI | |
| How Brands Use Reddit to Poison AI Search 404 Media |
This week's poll
Your AI vendor accepts no liability for what its models do. What would actually make you trust AI in production?
Last week, 229 of you voted:
Which source will matter most for the next jump in AI capability?
Your AI vendor accepts no liability for what its models do. What would actually make you trust AI in production?
Back Friday.
Alexis
文章标题:AI周刊第518期:白宫完成了其AI安全框架。该框架是保密的。
文章链接:https://news.qimuai.cn/?post=4726
本站文章均为原创,未经授权请勿用于任何商业用途