我们仍然不清楚人们究竟是如何使用人工智能的。

内容来源:https://www.technologyreview.com/2026/08/18/1142226/how-people-use-ai/
内容总结:
AI使用真相调查:独立研究揭示企业与现实差距
一项最新发布的独立研究显示,人们对人工智能(AI)的实际使用情况,与Anthropic、OpenAI等头部AI企业公开报告所描绘的图景存在显著差异。
该研究项目名为“AI观察站”(AI Observatory),由斯坦福大学、麻省理工学院等机构的学者联合发起。研究人员整合了七个既有数据集,分析了超过2.4万段用户与ChatGPT、Claude、Gemini等主流模型的真实对话,旨在为政策制定者和研究人员提供不受企业影响的独立数据参考。
研究团队指出,Anthropic和OpenAI定期发布的报告通常只展示其希望公众看到的数据,且存在明显的“选择性盲区”。例如,Anthropic的经济指数报告明确聚焦于生产力相关的工作场景。但当研究人员采用Anthropic的筛选方法重新分析数据时发现,高达48%的对话因与工作无关而被排除在外。这些被过滤的对话中,涉及健康、人际关系、成人或非法内容、骚扰仇恨言论以及色情内容的比例,均远高于企业报告所呈现的水平。
值得注意的是,研究还揭示出不同AI模型的使用方式截然不同。例如,Grok和Gemini更多被用于信息检索,其中Grok在新闻和政治领域颇受欢迎,但也是虚假信息集中传播的平台;Claude则主要被用于编程;Gemini更偏向社交和角色扮演;ChatGPT则常见于作业辅导。
此外,用户的使用行为随时间发生了显著变化。数据显示,对话变得更长、更复杂,闲聊内容增多,这暗示AI陪伴功能的兴起,同时AI助手主动承认自己是聊天机器人的次数减少。另一方面,涉及有害或敏感内容的对话频率总体下降,可能反映出各平台的安全防护措施正在加强。
研究还发现,即便是同一模型的不同版本也表现出差异。例如,用户与GPT-3.5驱动的ChatGPT对话较短,而与GPT-4o的对话则更长且更具反复性。
参与该研究的学者表示,这些细微差别在企业自身的报告中往往被忽略,没有任何单一企业的报告能呈现全貌。独立研究对于评估AI的真实利弊至关重要,因为目前许多重大决策都基于极为有限的数据。研究人员呼吁,AI公司应在保护用户隐私的前提下,向独立学术机构开放更多数据接口。否则,决策者将如同“在荒野中摸索”,依据企业叙事做出可能带来深远影响的判断。
中文翻译:
我们尚不清楚人们究竟如何使用AI,但一项新研究表明,工作场景的使用占比远低于AI公司所宣称的水平。
像Anthropic和OpenAI这样的AI公司定期发布关于人们如何使用Claude和ChatGPT等产品的报告,但AI研究人员表示,这些公司只公布他们希望我们看到的数据。
“没有独立的来源可以佐证这些数据,”斯坦福大学可信AI研究实验室的计算机科学博士生安卡·罗伊尔说道。
罗伊尔是名为“AI观察站”的新研究项目的联合负责人,该项目旨在填补这一空白。这是一个公共平台,汇总并分析了用户通过七个现有数据集同意收集的、与Claude和Gemini等流行模型的真实AI对话。其目的是提供独立的信源,帮助研究人员和政策制定者评估人们如何使用生成式AI。罗伊尔表示,目前关于AI收益和风险的重大决策是在非常有限的数据基础上做出的。
AI观察站发现,AI的使用在不同模型之间存在显著差异,并且随时间发生了变化。其研究发现,用户行为中涉及敏感内容的情况远多于主要AI公司报告中所呈现的,这些报告更侧重于工作场景而非个人使用。
Anthropic经济指数是最知名、被引用最广泛的AI使用数据来源之一,但它存在盲区。顾名思义,它聚焦于Claude AI在工作效率相关的用途上,过滤掉了与这些用途无关的对话。
当AI观察站的研究人员将Anthropic的方法应用于他们的数据集时,他们发现近一半的对话(48%)会被过滤掉。这些与工作无关的对话更可能涉及健康和人际关系(44.2%,而Anthropic的分析中为31.2%)、成人或非法话题(7.9%,对比2.1%)、骚扰和仇恨言论(27.5%,对比5.66%)以及性内容(16.7%,对比2.4%)。(OpenAI 2025年关于ChatGPT的报告同样发现,消费者使用中只有30%与工作相关。)
Anthropic已经发布了单独的博客文章,介绍人们如何使用Claude获得支持或陪伴,甚至用于生成儿童性虐待材料,但德克萨斯大学奥斯汀分校助理教授大卫·威德表示,拥有“AI观察站的全局视角分析”,而不是将这些信息“分割到单独的报告中”,有助于研究人员更一致地理解不同的使用方式。威德研究人与AI系统的互动方式,并未参与AI观察站项目。
AI观察站考察的数据集包含了2023年至2025年间的对话,研究发现人们使用AI的方式以及各AI平台的回应都存在差异。
在WildChat(AI观察站研究中包含的最大、最详细的数据集之一)中的对话,随着时间的推移变得越来越长、越来越复杂,表现为提示词令牌、响应令牌和对话轮次的不断增多。
闲聊内容也显著增加。这表明AI陪伴的使用在上升;与此同时,AI助手自我披露(即承认自己是聊天机器人)的次数却在减少。此外,研究人员标记为敏感(即包含潜在有害或受限制内容,包括性骚扰和仇恨言论)的交流变得不那么频繁。这可能表明各平台总体上部署了更有效的防护措施。
AI观察站还发现,话题、互动风格、对话结构以及敏感使用案例的可能性和类型在不同模型之间各不相同。
例如,研究人员发现,人们更频繁地使用Grok和Gemini进行信息检索。Grok尤其适合获取新闻和政治信息,但同时也是错误信息容易集中的地方。(这与另一项研究一致,该研究显示错误信息在Grok上如何容易泛滥。xAI未回应置评请求。)
与此同时,人们更倾向于使用Anthropic进行编程,使用Gemini进行社交和角色扮演,使用ChatGPT协助完成作业。
即使是同一模型的不同版本之间也存在差异。研究人员发现,当ChatGPT由GPT-3.5驱动时,人们的对话较短;而使用GPT-4o时,对话更长且更具迭代性——这很合理,因为该版本以导致情感依赖而闻名。
然而,各公司的报告往往无法捕捉到它们自身模型之间甚至之内的这些细微差别。“没有哪一份公司报告能讲述全貌,”与罗伊尔共同领导这项研究的麻省理工学院媒体实验室近期博士毕业生谢恩·朗普雷说道。
为了创建AI观察站,罗伊尔与来自麻省理工学院、斯坦福大学、数据溯源倡议及其他机构的研究人员汇总了来自七个先前研究收集的真实世界数据集的24,521段对话中的85,633个对话轮次(即用户提示词和相应的AI回复)。这些对话来自5,000名用户在2023年至2025年间与52个不同模型(包括ChatGPT、Gemini、Claude和Grok)的互动。
但与大型实验室自身拥有的数据相比,这些对话不过是九牛一毛。例如,最新的Anthropic经济AI指数基于对100万段Claude对话的分析;OpenAI关于人们如何使用ChatGPT的报告分析了150万段对话。
Anthropic的一位代表表示,该公司发布的研究反映了其研究团队的具体问题和兴趣,并且支持外部独立研究非常重要。OpenAI未回应置评请求。
AI观察站的数据集来源于自愿提供的资料,这意味着它可能低估了敏感用途的使用情况,因为人们可能不太愿意分享这些内容。因此,研究人员提醒说,他们的发现并不能代表所有AI使用情况。
不过,该项目的成果拓宽了研究社区的获取渠道。像罗伊尔和威德这样的独立研究人员表示,AI公司通常不会分享他们的聊天数据供分析,这意味着他们的报告往往侧重于展示自己最好一面的发现。
“当我们想问,例如:Anthropic的通用AI系统……主要是被用于好的方面还是坏的方面……我们没有办法回答这个问题,因为这些信息是专有的,”威德解释道。
AI观察站的数据将对研究人员开放分析,团队希望随着时间的推移扩展其数据集。罗伊尔表示,理想情况下,AI公司应该与独立研究人员共享数据——当然要以保护用户隐私的方式。但就目前情况而言,她说,任何基于AI使用数据做决策的人,都冒着“完全在黑暗中操作,在不知道公司叙事之外实际发生了什么的情况下做出这些重大决策”的风险。
深度专题
人工智能
一家初创公司声称突破了制约大语言模型的瓶颈
Subquadratic现已分享了其新模型的更多细节,但一些人仍持怀疑态度。
一个根本性缺陷使大语言模型极易受到攻击
这使得人们可以轻易诱骗它们做出不该做的事情,比如告诉你如何破坏飞机的导航系统。
Anthropic发现了一个隐藏空间,Claude在那里思索概念
一项新技术让该公司比以往任何时候都更深入地探索了大语言模型的奇异内部运作。
Claude Science是Anthropic最新的旗舰产品
该公司正在加大对AI助力科学研究的投入。
保持联系
获取来自《麻省理工科技评论》的最新资讯
发现特别优惠、热门故事、即将举行的活动等更多内容。
英文来源:
We still don’t know how people are really using AI
But a new study shows that work use cases make up less of the picture than AI companies claim.
AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say.
“There is no independent source to corroborate it,” says Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab.
Reuel is co-lead of a new research project, called the AI Observatory, that aims to fill the gap. It’s a public platform that aggregated and analyzed real AI conversations with popular models like Claude and Gemini that were collected with users’ consent through seven existing datasets. The intent is to provide independent sources of information that can help researchers and policymakers assess how people are using generative AI. Highly consequential decisions about AI’s benefits and risks are currently being made on the basis of very limited data, says Reuel.
The AI Observatory found that AI use differs significantly across models and has changed over time. Its research shows many more sensitive behaviors than are captured in reports from major AI companies, which they say focus more on work than on personal use.
The Anthropic Economic Index is one of the best-known and most widely cited sources of AI usage data, but it has blind spots. As its name suggests, it focuses on work- and productivity-related uses of Claude AI—filtering out conversations that are unrelated to these uses.
When the AI Observatory researchers applied Anthropic’s methods to their dataset, they found that nearly half the conversations—48%—would have been filtered out. Those non-work-related conversations were more likely to involve health and relationships (44.2% versus 31.2% in Anthropic’s analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). (OpenAI’s 2025 report on ChatGPT, similarly, found that only 30% of consumer use was related to work.)
Anthropic has released separate blog posts on how people use Claude for support or companionship, and even to generate CSAM, but “having [the AI Observatory’s] bird’s-eye-view analysis” rather than leaving that information “sectioned off into a separate report” helps researchers understand the different uses more consistently, says David Widder, an assistant professor at the University of Texas at Austin, who researches how people interact with AI systems and is not involved with the AI Observatory.
The datasets the AI Observatory looked at include conversations that took place between 2023 and 2025, and it found differences both in how people were using AI and how various AI platforms responded.
Conversations within WildChat, one of the largest and most detailed datasets included in the AI Observatory’s study, got longer and more elaborate over time, as indicated by growing numbers of prompt tokens, response tokens, and conversation turns.
There was also significantly more small talk over time. That suggests that AI companionship was increasing; meanwhile, the AI assistants’ self-disclosure (i.e., admitting to being a chatbot) decreased.
Additionally, exchanges that the researchers labeled as sensitive—meaning ones with potentially harmful or restricted content, including sexual harassment and hate speech—became less frequent. That might suggest that platforms were generally deploying more effective safeguards.
The AI Observatory also found that topics, interaction styles, conversation structures, and the likelihood and type of sensitive use cases differed from one model to another.
For example, the researchers found that people used Grok and Gemini more frequently for information retrieval. Grok, in particular, was especially popular for information on news and politics, but it was also where misinformation tended to concentrate. (This is consistent with other research that has shown how readily misinformation proliferates on Grok. xAI did not respond to a request for comment.)
Meanwhile, people were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance.
There were even differences between different versions of the same model. Researchers found that people had shorter conversations with ChatGPT when it was powered by GPT-3.5, and longer and more iterative ones with GPT-4o—which makes sense given that that version became known for leading to emotional addiction.
Companies’ reports, however, didn’t tend to capture these nuances between or even within their own models. “No single company report tells the whole story,” says Shayne Longpre, a recent PhD graduate from the MIT Media Lab who co-led the research with Reuel.
To create the AI Observatory, Reuel and researchers from MIT, Stanford, the Data Provenance Initiative, and other institutions aggregated 85,633 conversational turns (that is, the user prompt and corresponding AI response) across 24,521 conversations from seven real-world datasets collected in previous research. These conversations came from 5,000 users interacting with 52 different models, including ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025.
But these conversations are a drop in the proverbial bucket compared with the data that the big labs themselves have access to. The latest Anthropic Economic AI Index, for example, is based on analysis of 1 million Claude conversations; OpenAI’s report on how people are using ChatGPT analyzed 1.5 million conversations.
An Anthropic representative said the company’s published research reflects its research teams’ specific questions and interests and that it’s important to support external independent research. OpenAI did not respond to requests for comment.
The fact that the AI Observatory’s dataset draws from voluntarily provided sources means it’s probably underrepresenting sensitive uses, which people may be less likely to share. Thus, the researchers caution that its findings are not indicative of all AI use.
The project’s work, though, broadens access for the research community. AI companies don’t typically share their chat data for analysis, which means their reports tend to focus on the findings that paint them in the best light, independent researchers like Reuel and Widder say.
“When we want to ask, for example: is Anthropic’s general-purpose AI system … used mostly for good or mostly for bad … we don’t have a way of answering that question because that information is proprietary,” explains Widder.
The AI Observatory’s data will be available to researchers for analysis, and the team hopes to expand its datasets over time. Ideally, Reuel says, the AI companies would share their data with independent researchers—in ways that protect user privacy, of course. But as it currently stands, she says, anyone making decisions based on AI usage data risks “completely operating in the wild and making these really consequential decisions without knowing what’s actually happening beyond those company narratives.”
Deep Dive
Artificial intelligence
A startup claims it broke through a bottleneck that’s holding back LLMs
Subquadratic has now shared more details about its new model. But some are still skeptical.
A fundamental flaw leaves LLMs strikingly vulnerable to attack
It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system.
Anthropic found a hidden space where Claude puzzles over concepts
A new technique has let the company probe deeper than ever into the weird workings of an LLM.
Claude Science is Anthropic’s newest flagship product
The company is doubling down on AI for science.
Stay connected
Get the latest updates from
MIT Technology Review
Discover special offers, top stories, upcoming events, and more.
文章标题:我们仍然不清楚人们究竟是如何使用人工智能的。
文章链接:https://news.qimuai.cn/?post=4838
本站文章均为原创,未经授权请勿用于任何商业用途