亚马逊可使用你的Twitch内容训练其AI——除非你选择退出

qimuai 发布于 阅读:59 一手编译

亚马逊可使用你的Twitch内容训练其AI——除非你选择退出

内容来源:https://www.wired.com/story/amazon-uses-your-twitch-content-to-train-its-ai-how-to-opt-out/

内容总结:

Twitch允许主播选择退出AI训练数据收集,引发广泛关注

近日,流媒体平台Twitch更新了其账户设置,允许主播选择退出将其内容用于训练母公司亚马逊的AI模型。此举虽在一定程度上缓解了部分创作者的担忧,但同时也引发了关于平台何时开始使用用户数据进行AI训练,以及大型科技公司如何处理用户数据的新疑虑。

设置流程简便,但条款细节值得注意

用户只需在Twitch网站或移动应用中,点击头像进入“设置”,在“安全与隐私”部分找到“生成式AI训练”选项,即可关闭该功能。然而,Twitch在相关说明中明确指出,关闭此选项并不会阻止Twitch和亚马逊将频道内容用于隐私政策中描述的其他目的,包括用于辅助主播成长和变现的AI功能,如赞助活动实时协助、观众推荐以及自动审核等社区安全工具。

默认开启引争议,平台高管承认“别无选择”

该新设置旨在承认创作者对其内容使用方式的决定权,但其上线也引发了对过去内容使用情况的追问。在一个相关论坛上,超过1.6万名创作者曾反对其内容在默认情况下被用于训练AI。这一反对声浪源于Twitch社区负责人玛丽·基什在直播中介绍该变更后,她承认此举会引发负面反应。Twitch产品负责人迈克·明顿也表示,默认保持开启是必要的,因为否则“没有人会参与”这一过程。他更坦言,并非只有Twitch这么做,可以合理推测几乎所有公开内容都在以某种方式被用于训练模型,无论是否获得许可,这其中有很多情况已超出平台的直接控制范围。

数据来源与授权边界成谜

明顿的言论引发了更多疑问:Twitch内容何时开始被用于AI训练?亚马逊是唯一的使用方,还是其商业伙伴也参与其中?创作者内容的著作权在多大程度上得到尊重?自2024年3月起生效的Twitch服务条款虽规定用户授予平台及被许可方使用、复制、修改和创作衍生作品等权利,但并未明确规定此类材料可用于训练生成式AI模型。

训练数据困境成为行业普遍难题

此事件折射出AI系统开发者面临的重大挑战——高质量训练数据的日益短缺。尽管OpenAI等公司已与多家出版商达成内容使用协议,但随着技术发展和对数据需求的增加,可用数据源正在减少。为此,各种替代方案应运而生,再次引发了关于训练过程伦理和透明度的争论。此前,Meta被曝使用用户在Facebook和Instagram上分享的内容训练AI,甚至包括其智能眼镜用户截取的画面;谷歌和YouTube也有类似案例。因此,高质量训练数据如同芯片、存储和电力一样,已成为AI头部企业争夺的战略资源。种种迹象表明,这部分需求正大量通过使用用户生成的信息来满足,换取的是用户免费使用日常数字服务的权利。

中文翻译:

Twitch 已更新其账户设置,允许主播选择退出,不让自己的内容被用于帮助训练其母公司亚马逊的人工智能模型。

尽管此举让一些创作者感到安心,但外界仍不清楚该流媒体平台究竟是从何时开始使用用户的帖子、直播和视频来训练亚马逊系统的。这一披露引发了人们对大型科技公司如何处理用户数据的新担忧。

退出流程很简单。从 Twitch 网站或移动应用上,点击你的账户头像,选择“设置”,然后进入菜单中的“安全与隐私”部分。在那里,你会找到“生成式人工智能培训”选项,你可以关闭该选项,以禁止你的内容被用于此目的。

在开关旁边的说明文字中,Twitch 指出,“关闭此选项并不能阻止 Twitch 和亚马逊将你频道的内容用于 Twitch《隐私声明》中描述的其他目的。”这些目的包括旨在促进主播成长和变现的人工智能平台功能,例如赞助活动的实时协助、通过推荐帮助观众发现内容,以及通过 AutoMod 等工具保障社区安全。

这项新设置的目的是承认创作者有权决定其所产内容的使用方式。然而,其推出也引发了关于 Twitch 和亚马逊迄今如何使用这些内容的问题。

在一个专门讨论此话题的论坛上,超过1.6万名创作者表示反对自己的内容被默认用于训练亚马逊的人工智能系统——这一做法直到此次设置更新后才被曝光。

这场风波源于一场直播,Twitch 社区负责人玛丽·基什在直播中解释了这些变更的引入。该高管承认,这一改变会引发负面反应。Twitch 产品负责人迈克·明顿也指出——用他的话说是“坦率的回应”——将使用内容进行人工智能培训的选项保持默认开启是必要的,因为否则“没有人会参与”这一过程。

明顿指出,这些机制并非 Twitch 独有,他认为其他开发人工智能系统的公司很可能也在从 Twitch 和其他服务中提取内容用于训练模型。“我不确定,”他说,“但我认为可以相当合理地推测,几乎所有公开可获取的内容都在以某种方式被用于训练模型,无论是否获得许可。所以我认为我们还需要承认,这里有很多事情甚至超出了我们的直接控制范围。”

这些言论引发了新的问题:Twitch 内容从何时开始被用于训练人工智能模型?亚马逊是唯一使用这些数据的公司吗,还是其商业合作伙伴也有参与?创作者发布内容的著作权在何种程度上以及以何种方式受到尊重?

Twitch 自2024年3月起生效的服务条款规定,用户授予 Twitch 及其再许可方使用、复制、修改、改编、分发其内容以及基于其内容创作衍生作品的权利。然而,到目前为止,条款并未明确说明此类材料可用于训练生成式人工智能模型。

训练数据问题

这一案例凸显了人工智能系统开发者面临的重大挑战之一:用于训练模型的高质量数据日益短缺。

尽管像 OpenAI 这样的公司已与多家出版商(包括 WIRED 的母公司康泰纳仕)达成协议,使用其部分内容用于此目的,但随着技术的进步和数据需求的增加,可用的数据来源正在减少。因此,各种替代方案应运而生,重新引发了关于训练过程伦理和透明度的争论。

一段时间以来,Meta 一直在使用用户在 Facebook 和 Instagram 上分享的帖子和图片来训练其人工智能模型。最近,有消息披露,该公司还在使用员工的活跃数据以及用户通过其智能眼镜拍摄的截图用于同样目的。谷歌和 YouTube 也出现过类似案例。

因此,高质量的训练数据已成为另一种战略资源——与芯片、内存和电力一样——而领先的人工智能模型开发者发现,由于他们难以跟上市场要求的创新速度,这些资源正变得捉襟见肘。种种迹象表明,这部分需求中的一部分将通过用户自身生成的信息来满足,而作为交换,用户可以获得日常数字服务的免费访问。

本报道最初发表于 WIRED 西班牙语版,并从西班牙语翻译而来。

评论

返回顶部

英文来源:

Twitch has updated its account settings to let streamers opt out of having their content used to help train the artificial intelligence models of Twitch’s parent company, Amazon.
Although the move has reassured some creators, it remains unclear exactly when the streaming platform began using the posts, streams, and videos of its users to train Amazon’s systems. The revelation is raising new concerns about how big tech companies handle their users’ data.
The opt-out process is simple. From the Twitch website or mobile app, click on your account avatar, select Settings, and then go to the Security and Privacy section in the menu. There, you’ll find the Generative AI Training option, where you can disable the use of your content for that purpose.
In the language next to the toggle, Twitch notes that “disabling this option does not prevent Twitch and Amazon from using your channel’s content for other purposes described in Twitch’s Privacy Notice.” These include AI-powered platform features designed to facilitate streamers’ growth and monetization, such as real-time assistance for sponsorship campaigns, viewer discovery through recommendations, and community safety via tools like AutoMod.
The new setting aims to recognize creators’ right to decide how the content they produce is used. However, its launch has also raised questions about how Twitch and Amazon have used that content up to now.
In a forum dedicated to the topic, more than 16,000 creators expressed opposition to having their content being used by default to train Amazon’s AI systems—a practice that only came to light following the update to the settings.
The backlash arose after a livestream in which Mary Kish, Twitch’s head of community, explained the introduction of the changes. The executive acknowledged that the change would provoke a negative reaction. Mike Minton, Twitch’s head of product, also noted—in what he described as “a candid response”—that keeping the option to use content for AI training enabled by default was necessary, since otherwise “no one would participate” in the process.
Minton pointed out that these mechanisms are not unique to Twitch and considered it likely that other companies developing AI systems are also extracting content from Twitch and other services to use in training their models. “I don’t know for sure,” he said, “but I think it’s quite reasonable to assume that almost any publicly available content is used to train models in one way or another, with or without permission. So I think we also need to acknowledge that there’s a lot here that’s beyond even our direct control.”
These statements raised new questions: Since when has Twitch content been used to train AI models? Is Amazon the only company using this data, or are its business partners also involved? To what extent and in what ways is the authorship of content published by creators respected?
Twitch’s Terms of Service, in effect since March 2024, have stipulated that users grant Twitch and its sublicensees the right to use, reproduce, modify, adapt, distribute, and create derivative works from their content. However, until now, they have not explicitly stated that such materials could be used to train generative AI models.
The Training Data Problem
This case highlights one of the major challenges facing AI system developers: the growing shortage of high-quality data for training models.
Although companies like OpenAI have reached agreements with various publishers (including WIRED’s corporate parent, Condé Nast) to use some of their content for this purpose, available data sources are dwindling as technology advances and the demand for data increases. As a result, various alternatives have emerged that are reigniting the debate over the ethics and transparency of training processes.
For some time now, Meta has been using posts and images that users share on Facebook and Instagram to train its AI models. Recently, it also came to light that the company was using its employees’ activity and screenshots captured by users of its smart glasses for the same purpose. Similar cases have been documented involving Google and YouTube.
Thus, high-quality training data has become another strategic resource—along with chips, memory, and electricity—that the leading developers of artificial intelligence models have found running short as they struggle to sustain the pace of innovation demanded by the market. All signs point to a portion of that demand being met with information generated by users themselves, in exchange for free access to everyday digital services.
This story originally appeared in WIRED en Español and has been translated from Spanish.
Comments
Back to top

连线杂志AI最前沿

文章目录


    扫描二维码,在手机上阅读