ChatGPT Is a Blurry JPEG of the Web

ChatGPT:互联网的「模糊JPEG」

Ted Chiang · The New Yorker

In 2013, workers at a German construction company noticed something odd about their Xerox photocopier: when they made a copy of the floor plan of a house, the copy differed from the original in a subtle but significant way. In the original floor plan, each of the house’s three rooms was accompanied by a rectangle specifying its area: the rooms were 14.13, 21.11, and 17.42 square metres, respectively. However, in the photocopy, all three rooms were labelled as being 14.13 square metres in size. The company contacted the computer scientist David Kriesel to investigate this seemingly inconceivable result. They needed a computer scientist because a modern Xerox photocopier doesn’t use the physical xerographic process popularized in the nineteen-sixties. Instead, it scans the document digitally, and then prints the resulting image file. Combine that with the fact that virtually every digital image file is compressed to save space, and a solution to the mystery begins to suggest itself.

2013年,一家德国建筑公司的员工发现他们的施乐(Xerox)复印机有些古怪:当他们复印一张房屋平面图时,复印件与原件之间存在一个细微却重大的差异。在原始平面图中,房屋的三个房间各附有一个矩形,标明了各自的面积:房间分别为 14.13、21.11 和 17.42 平方米。然而,在复印件中,所有三个房间都被标记为 14.13 平方米。公司联系了计算机科学家大卫·克里塞尔(David Kriesel)来调查这个看似不可思议的结果。他们需要一位计算机科学家,因为现代施乐复印机不再使用 1960 年代普及的物理静电复印工艺。相反,它以数字方式扫描文档,然后打印生成的图像文件。再加上几乎每个数字图像文件都会为了节省空间而被压缩这一事实,谜底开始浮出水面。

Compressing a file requires two steps: first, the encoding, during which the file is converted into a more compact format, and then the decoding, whereby the process is reversed. If the restored file is identical to the original, then the compression process is described as lossless: no information has been discarded. By contrast, if the restored file is only an approximation of the original, the compression is described as lossy: some information has been discarded and is now unrecoverable. Lossless compression is what’s typically used for text files and computer programs, because those are domains in which even a single incorrect character has the potential to be disastrous. Lossy compression is often used for photos, audio, and video in situations in which absolute accuracy isn’t essential. Most of the time, we don’t notice if a picture, song, or movie isn’t perfectly reproduced. The loss in fidelity becomes more perceptible only as files are squeezed very tightly. In those cases, we notice what are known as compression artifacts: the fuzziness of the smallest JPEG and MPEG images, or the tinny sound of low-bit-rate MP3s.

压缩文件需要两个步骤:首先是编码,在此过程中文件被转换为更紧凑的格式;然后是解码,即该过程的逆过程。如果恢复的文件与原始文件完全相同,则压缩过程被称为无损压缩:没有信息被丢弃。相比之下,如果恢复的文件只是原始文件的近似值,则压缩被称为有损压缩:一些信息已被丢弃且无法恢复。无损压缩通常用于文本文件和计算机程序,因为在这些领域中,即使是一个错误的字符也可能导致灾难性的后果。有损压缩通常用于照片、音频和视频,在这些情况下,绝对的准确性并非必不可少。大多数时候,我们不会注意到图片、歌曲或电影是否被完美再现。只有当文件被极度压缩时,保真度的损失才会变得明显。在这些情况下,我们会注意到所谓的压缩伪影:最小的 JPEG 和 MPEG 图像的模糊感,或低比特率 MP3 的刺耳声音。

Xerox photocopiers use a lossy compression format known as JBIG2, designed for use with black-and-white images. To save space, the copier identifies similar-looking regions in the image and stores a single copy for all of them; when the file is decompressed, it uses that copy repeatedly to reconstruct the image. It turned out that the photocopier had judged the labels specifying the area of the rooms to be similar enough that it needed to store only one of them—14.13—and it reused that one for all three rooms when printing the floor plan.

施乐复印机使用一种称为 JBIG2 的有损压缩格式,专为黑白图像设计。为了节省空间,复印机会识别图像中外观相似的区域,并只存储其中一个副本;当文件解压缩时,它会重复使用该副本以重建图像。事实证明,复印机判定标示房间面积的标签足够相似,只需要存储其中一个——14.13——并在打印平面图时将该标签重复用于所有三个房间。

The fact that Xerox photocopiers use a lossy compression format instead of a lossless one isn’t, in itself, a problem. The problem is that the photocopiers were degrading the image in a subtle way, in which the compression artifacts weren’t immediately recognizable. If the photocopier simply produced blurry printouts, everyone would know that they weren’t accurate reproductions of the originals. What led to problems was the fact that the photocopier was producing numbers that were readable but incorrect; it made the copies seem accurate when they weren’t. (In 2014, Xerox released a patch to correct this issue.)

施乐复印机使用有损压缩格式而不是无损压缩格式这一事实本身并不是问题。问题在于复印机以一种微妙的方式降低了图像质量,其中的压缩伪影并不能立即被识别出来。如果复印机只是生成模糊的打印件,每个人都会知道它们不是原件的精确复制品。导致问题的原因是复印机生成的数字清晰可读但却是错误的;它使复印件看起来很准确,但实际上并非如此。(2014年,施乐发布了一个补丁来修正这个问题。)

I think that this incident with the Xerox photocopier is worth bearing in mind today, as we consider OpenAI’s ChatGPT and other similar programs, which A.I. researchers call large language models. The resemblance between a photocopier and a large language model might not be immediately apparent—but consider the following scenario. Imagine that you’re about to lose your access to the Internet forever. In preparation, you plan to create a compressed copy of all the text on the Web, so that you can store it on a private server. Unfortunately, your private server has only one per cent of the space needed; you can’t use a lossless compression algorithm if you want everything to fit. Instead, you write a lossy algorithm that identifies statistical regularities in the text and stores them in a specialized file format. Because you have virtually unlimited computational power to throw at this task, your algorithm can identify extraordinarily nuanced statistical regularities, and this allows you to achieve the desired compression ratio of a hundred to one.

我认为施乐复印机的这个事件在今天值得铭记,因为我们正在考量 OpenAI 的 ChatGPT 和其他类似的程序,人工智能研究人员称之为大型语言模型。复印机和大型语言模型之间的相似之处可能不会立即显现——但请考虑以下场景。想象一下,你即将永远失去互联网访问权限。为了做准备,你计划创建网络上所有文本的压缩副本,以便将其存储在私人服务器上。不幸的是,你的私人服务器只有所需空间的百分之一;如果你想装下所有东西,就不能使用无损压缩算法。相反,你编写了一种有损算法,识别文本中的统计规律,并将它们存储在一种专门的文件格式中。因为你有几乎无限的计算能力来投入这项任务,你的算法可以识别极其细微的统计规律,这使你能够实现一百比一的预期压缩比。

Now, losing your Internet access isn’t quite so terrible; you’ve got all the information on the Web stored on your server. The only catch is that, because the text has been so highly compressed, you can’t look for information by searching for an exact quote; you’ll never get an exact match, because the words aren’t what’s being stored. To solve this problem, you create an interface that accepts queries in the form of questions and responds with answers that convey the gist of what you have on your server.

现在,失去互联网访问权限并不那么可怕了;你已经将网络上的所有信息存储在你的服务器上了。唯一的限制是,由于文本被高度压缩,你无法通过搜索精确引语来查找信息;你永远无法获得精确匹配,因为存储的不是单词本身。为了解决这个问题,你创建了一个界面,接受问题形式的查询,并以传达你服务器上内容主旨的答案作为回应。

What I’ve described sounds a lot like ChatGPT, or most any other large language model. Think of ChatGPT as a blurry JPEG of all the text on the Web. It retains much of the information on the Web, in the same way that a JPEG retains much of the information of a higher-resolution image, but, if you’re looking for an exact sequence of bits, you won’t find it; all you will ever get is an approximation. But, because the approximation is presented in the form of grammatical text, which ChatGPT excels at creating, it’s usually acceptable. You’re still looking at a blurry JPEG, but the blurriness occurs in a way that doesn’t make the picture as a whole look less sharp.

我所描述的听起来很像 ChatGPT,或者几乎任何其他大型语言模型。把 ChatGPT 想象成网络上所有文本的模糊 JPEG。它保留了网络上的大部分信息,就像 JPEG 保留了高分辨率图像的大部分信息一样,但是,如果你正在寻找精确的比特序列,你是找不到的;你得到的永远只是一个近似值。但是,因为这个近似值是以符合语法的文本形式呈现的——这正是 ChatGPT 擅长生成的——所以它通常是可以接受的。你仍然在看一张模糊的 JPEG,但这种模糊是以一种不会让整张图片看起来不清晰的方式出现的。

This analogy to lossy compression is not just a way to understand ChatGPT’s facility at repackaging information found on the Web by using different words. It’s also a way to understand the “hallucinations,” or nonsensical answers to factual questions, to which large language models such as ChatGPT are all too prone. These hallucinations are compression artifacts, but—like the incorrect labels generated by the Xerox photocopier—they are plausible enough that identifying them requires comparing them against the originals, which in this case means either the Web or our own knowledge of the world. When we think about them this way, such hallucinations are anything but surprising; if a compression algorithm is designed to reconstruct text after ninety-nine per cent of the original has been discarded, we should expect that significant portions of what it generates will be entirely fabricated.

这种与有损压缩的类比不仅是理解 ChatGPT 擅长用不同词语重新包装网络信息的一种方式。它也是理解"幻觉"(hallucinations)的一种方式,即大型语言模型(如 ChatGPT)极易产生的对事实问题的荒谬回答。这些幻觉是压缩伪影,但就像施乐复印机生成的错误标签一样,它们足够可信,以至于识别它们需要将其与原件进行比较,在这种情况下,原件意味着网络或我们要对世界的认知。当我们这样思考时,这种幻觉就不足为奇了;如果一个压缩算法被设计为在丢弃了 99% 的原始数据后重建文本,我们应该预料到它生成的很大一部分内容将完全是捏造的。

This analogy makes even more sense when we remember that a common technique used by lossy compression algorithms is interpolation—that is, estimating what’s missing by looking at what’s on either side of the gap. When an image program is displaying a photo and has to reconstruct a pixel that was lost during the compression process, it looks at the nearby pixels and calculates the average. This is what ChatGPT does when it’s prompted to describe, say, losing a sock in the dryer using the style of the Declaration of Independence: it is taking two points in “lexical space” and generating the text that would occupy the location between them. (“When in the Course of human events, it becomes necessary for one to separate his garments from their mates, in order to maintain the cleanliness and order thereof. . . .”) ChatGPT is so good at this form of interpolation that people find it entertaining: they’ve discovered a “blur” tool for paragraphs instead of photos, and are having a blast playing with it.

当我们记得有损压缩算法使用的一种常用技术是插值(interpolation)——即通过查看间隙两侧的内容来估计缺失的内容时,这个类比就更有意义了。当图像程序显示照片并必须重建在压缩过程中丢失的像素时,它会查看附近的像素并计算平均值。这就是 ChatGPT 在被提示用《独立宣言》的风格描述在烘干机里丢了一只袜子时所做的事情:它在"词汇空间"中选取两点,并生成占据它们之间位置的文本。("在人类事务的发展过程中,当一个人有必要将他的衣物与其配偶分开,以保持其清洁和秩序时……")ChatGPT 非常擅长这种形式的插值,以至于人们觉得它很有趣:他们发现了一个用于段落而不是照片的"模糊"工具,并且玩得不亦乐乎。

Given that large language models like ChatGPT are often extolled as the cutting edge of artificial intelligence, it may sound dismissive—or at least deflating—to describe them as lossy text-compression algorithms. I do think that this perspective offers a useful corrective to the tendency to anthropomorphize large language models, but there is another aspect to the compression analogy that is worth considering. Since 2006, an A.I. researcher named Marcus Hutter has offered a cash reward—known as the Prize for Compressing Human Knowledge, or the Hutter Prize—to anyone who can losslessly compress a specific one-gigabyte snapshot of Wikipedia smaller than the previous prize-winner did. You have probably encountered files compressed using the zip file format. The zip format reduces Hutter’s one-gigabyte file to about three hundred megabytes; the most recent prize-winner has managed to reduce it to a hundred and fifteen megabytes. This isn’t just an exercise in smooshing. Hutter believes that better text compression will be instrumental in the creation of human-level artificial intelligence, in part because the greatest degree of compression can be achieved by understanding the text.

鉴于像 ChatGPT 这样的大型语言模型经常被吹捧为人工智能的前沿,将它们描述为有损文本压缩算法听起来可能有些轻蔑——或者至少是令人泄气的。我确实认为这个视角为拟人化大型语言模型的倾向提供了一个有益的纠正,但压缩类比还有另一个方面值得考虑。自 2006 年以来,一位名叫马库斯·赫特(Marcus Hutter)的人工智能研究员提供了一笔现金奖励——被称为"压缩人类知识奖"(Prize for Compressing Human Knowledge)或赫特奖(Hutter Prize)——奖励给任何能够将维基百科特定 1GB 快照无损压缩得比前一位获奖者更小的人。你可能遇到过使用 zip 文件格式压缩的文件。Zip 格式将赫特的 1GB 文件减少到约 300MB;最近的获奖者已设法将其减少到 115MB。这不仅仅是一个压缩练习。赫特认为,更好的文本压缩将有助于创造人类水平的人工智能,部分原因是最大程度的压缩可以通过理解文本来实现。

To grasp the proposed relationship between compression and understanding, imagine that you have a text file containing a million examples of addition, subtraction, multiplication, and division. Although any compression algorithm could reduce the size of this file, the way to achieve the greatest compression ratio would probably be to derive the principles of arithmetic and then write the code for a calculator program. Using a calculator, you could perfectly reconstruct not just the million examples in the file but any other example of arithmetic that you might encounter in the future. The same logic applies to the problem of compressing a slice of Wikipedia. If a compression program knows that force equals mass times acceleration, it can discard a lot of words when compressing the pages about physics because it will be able to reconstruct them. Likewise, the more the program knows about supply and demand, the more words it can discard when compressing the pages about economics, and so forth.

为了理解压缩与理解之间提出的关系,想象一下你有一个包含一百万个加减乘除算术例子的文本文件。虽然任何压缩算法都可以减小这个文件的大小,但实现最大压缩比的方法可能是推导出算术原理,然后编写一个计算器程序的代码。使用计算器,你不仅可以完美地重建文件中的一百万个例子,还可以重建你将来可能遇到的任何其他算术例子。同样的逻辑也适用于压缩维基百科切片的问题。如果压缩程序知道力等于质量乘以加速度,它在压缩关于物理学的页面时就可以丢弃很多单词,因为它将能够重建它们。同样,程序对供给和需求了解得越多,在压缩关于经济学的页面时可以丢弃的单词就越多,依此类推。

Large language models identify statistical regularities in text. Any analysis of the text of the Web will reveal that phrases like “supply is low” often appear in close proximity to phrases like “prices rise.” A chatbot that incorporates this correlation might, when asked a question about the effect of supply shortages, respond with an answer about prices increasing. If a large language model has compiled a vast number of correlations between economic terms—so many that it can offer plausible responses to a wide variety of questions—should we say that it actually understands economic theory? Models like ChatGPT aren’t eligible for the Hutter Prize for a variety of reasons, one of which is that they don’t reconstruct the original text precisely—i.e., they don’t perform lossless compression. But is it possible that their lossy compression nonetheless indicates real understanding of the sort that A.I. researchers are interested in?

大型语言模型识别文本中的统计规律。对网络文本的任何分析都会显示,像"供应不足"这样的短语经常出现在像"价格上涨"这样的短语附近。一个包含这种相关性的聊天机器人,当被问及供应短缺的影响时,可能会回答关于价格上涨的内容。如果一个大型语言模型汇编了大量经济术语之间的相关性——多到它可以对各种各样的问题提供看似合理的回答——我们应该说它实际上理解经济理论吗?像 ChatGPT 这样的模型没有资格获得赫特奖,原因有很多,其中之一是它们不能精确地重建原始文本——即它们不执行无损压缩。但是,它们的有损压缩是否可能表明了人工智能研究人员感兴趣的那种真正的理解?

Let’s go back to the example of arithmetic. If you ask GPT-3 (the large-language model that ChatGPT was built from) to add or subtract a pair of numbers, it almost always responds with the correct answer when the numbers have only two digits. But its accuracy worsens significantly with larger numbers, falling to ten per cent when the numbers have five digits. Most of the correct answers that GPT-3 gives are not found on the Web—there aren’t many Web pages that contain the text “245 + 821,” for example—so it’s not engaged in simple memorization. But, despite ingesting a vast amount of information, it hasn’t been able to derive the principles of arithmetic, either. A close examination of GPT-3’s incorrect answers suggests that it doesn’t carry the “1” when performing arithmetic. The Web certainly contains explanations of carrying the “1,” but GPT-3 isn’t able to incorporate those explanations. GPT-3’s statistical analysis of examples of arithmetic enables it to produce a superficial approximation of the real thing, but no more than that.

让我们回到算术的例子。如果你让 GPT-3(构建 ChatGPT 的大型语言模型)对一对数字进行加减运算,当数字只有两位数时,它几乎总是能给出正确的答案。但随着数字变大,其准确性显著下降,当数字有五位数时,准确率降至 10%。GPT-3 给出的大多数正确答案在网络上都找不到——例如,包含文本"245 + 821"的网页并不多——所以它并没有进行简单的死记硬背。但是,尽管摄入了大量信息,它也没能推导出算术原理。对 GPT-3 错误答案的仔细检查表明,它在进行算术运算时不进位。网络上肯定包含关于进位的解释,但 GPT-3 无法整合这些解释。GPT-3 对算术示例的统计分析使其能够产生真实事物的表面近似值,但也仅此而已。

Given GPT-3’s failure at a subject taught in elementary school, how can we explain the fact that it sometimes appears to perform well at writing college-level essays? Even though large language models often hallucinate, when they’re lucid they sound like they actually understand subjects like economic theory. Perhaps arithmetic is a special case, one for which large language models are poorly suited. Is it possible that, in areas outside addition and subtraction, statistical regularities in text actually do correspond to genuine knowledge of the real world?

鉴于 GPT-3 在小学教授的科目上的失败,我们如何解释它有时在撰写大学水平论文方面表现良好的事实?即使大型语言模型经常产生幻觉,但当它们清醒时,它们听起来好像真的理解经济理论等学科。也许算术是一个特例,大型语言模型不适合它。是否有可能,在加减法以外的领域,文本中的统计规律实际上确实对应于现实世界的真实知识?

I think there’s a simpler explanation. Imagine what it would look like if ChatGPT were a lossless algorithm. If that were the case, it would always answer questions by providing a verbatim quote from a relevant Web page. We would probably regard the software as only a slight improvement over a conventional search engine, and be less impressed by it. The fact that ChatGPT rephrases material from the Web instead of quoting it word for word makes it seem like a student expressing ideas in her own words, rather than simply regurgitating what she’s read; it creates the illusion that ChatGPT understands the material. In human students, rote memorization isn’t an indicator of genuine learning, so ChatGPT’s inability to produce exact quotes from Web pages is precisely what makes us think that it has learned something. When we’re dealing with sequences of words, lossy compression looks smarter than lossless compression.

我认为有一个更简单的解释。想象一下,如果 ChatGPT 是一个无损算法会是什么样子。如果是那样的话,它总是会通过提供相关网页的逐字引用来回答问题。我们可能会认为该软件只是对传统搜索引擎的轻微改进,并且对它的印象不那么深刻。ChatGPT 对网络材料进行改写而不是逐字引用,这使它看起来像一个学生用自己的话表达想法,而不是简单地反刍她读过的内容;这造成了 ChatGPT 理解材料的错觉。对于人类学生来说,死记硬背并不是真正学习的指标,所以 ChatGPT 无法从网页中生成精确引语正是让我们认为它学到了一些东西的原因。当我们处理单词序列时,有损压缩看起来比无损压缩更聪明。

A lot of uses have been proposed for large language models. Thinking about them as blurry JPEGs offers a way to evaluate what they might or might not be well suited for. Let’s consider a few scenarios.

人们已经为大型语言模型提出了许多用途。将它们视为模糊的 JPEG 提供了一种评估它们可能适合或不适合什么的方法。让我们考虑几种情况。

Can large language models take the place of traditional search engines? For us to have confidence in them, we would need to know that they haven’t been fed propaganda and conspiracy theories—we’d need to know that the JPEG is capturing the right sections of the Web. But, even if a large language model includes only the information we want, there’s still the matter of blurriness. There’s a type of blurriness that is acceptable, which is the re-stating of information in different words. Then there’s the blurriness of outright fabrication, which we consider unacceptable when we’re looking for facts. It’s not clear that it’s technically possible to retain the acceptable kind of blurriness while eliminating the unacceptable kind, but I expect that we’ll find out in the near future.

大型语言模型能取代传统搜索引擎吗?为了让我们对它们有信心,我们需要知道它们没有被灌输宣传和阴谋论——我们需要知道 JPEG 捕捉到了网络的正确部分。但是,即使大型语言模型只包含我们想要的信息,仍然存在模糊性的问题。有一种模糊性是可以接受的,即用不同的词语重述信息。还有一种是完全捏造的模糊性,当我们寻找事实时,我们认为这是不可接受的。目前尚不清楚在技术上是否可能在保留可接受的模糊性的同时消除不可接受的模糊性,但我预计我们在不久的将来就会找到答案。

Even if it is possible to restrict large language models from engaging in fabrication, should we use them to generate Web content? This would make sense only if our goal is to repackage information that’s already available on the Web. Some companies exist to do just that—we usually call them content mills. Perhaps the blurriness of large language models will be useful to them, as a way of avoiding copyright infringement. Generally speaking, though, I’d say that anything that’s good for content mills is not good for people searching for information. The rise of this type of repackaging is what makes it harder for us to find what we’re looking for online right now; the more that text generated by large language models gets published on the Web, the more the Web becomes a blurrier version of itself.

即使有可能限制大型语言模型进行捏造,我们应该使用它们来生成网络内容吗?只有当我们的目标是重新包装网络上已有的信息时,这才有意义。有些公司就是为了做这件事而存在的——我们通常称之为内容农场。也许大型语言模型的模糊性对它们有用,作为一种避免侵犯版权的方式。不过,一般来说,我会说任何对内容农场有利的事情对搜索信息的人来说都不是好事。这种重新包装的兴起正是让我们现在更难在网上找到我们正在寻找的内容的原因;大型语言模型生成的文本在网络上发布的越多,网络就越变成其自身的模糊版本。

There is very little information available about OpenAI’s forthcoming successor to ChatGPT, GPT-4. But I’m going to make a prediction: when assembling the vast amount of text used to train GPT-4, the people at OpenAI will have made every effort to exclude material generated by ChatGPT or any other large language model. If this turns out to be the case, it will serve as unintentional confirmation that the analogy between large language models and lossy compression is useful. Repeatedly resaving a JPEG creates more compression artifacts, because more information is lost every time. It’s the digital equivalent of repeatedly making photocopies of photocopies in the old days. The image quality only gets worse.

关于 OpenAI 即将推出的 ChatGPT 继任者 GPT-4 的信息很少。但我要做一个预测:在收集用于训练 GPT-4 的大量文本时,OpenAI 的人员将尽一切努力排除由 ChatGPT 或任何其他大型语言模型生成的材料。如果事实证明确实如此,这将无意中证实大型语言模型和有损压缩之间的类比是有用的。反复重新保存 JPEG 会产生更多的压缩伪影,因为每次都会丢失更多信息。这相当于数字时代的旧式复印件的复印件。图像质量只会变得更差。

Indeed, a useful criterion for gauging a large language model’s quality might be the willingness of a company to use the text that it generates as training material for a new model. If the output of ChatGPT isn’t good enough for GPT-4, we might take that as an indicator that it’s not good enough for us, either. Conversely, if a model starts generating text so good that it can be used to train new models, then that should give us confidence in the quality of that text. (I suspect that such an outcome would require a major breakthrough in the techniques used to build these models.) If and when we start seeing models producing output that’s as good as their input, then the analogy of lossy compression will no longer be applicable.

事实上,衡量大型语言模型质量的一个有用标准可能是公司是否愿意将其生成的文本用作新模型的训练材料。如果 ChatGPT 的输出对 GPT-4 来说不够好,我们可能会将其视为它对我们来说也不够好的指标。相反,如果一个模型开始生成的文本非常好,以至于可以用来训练新模型,那么这应该让我们对该文本的质量充满信心。(我怀疑这样的结果需要构建这些模型的技术取得重大突破。)如果并且当我们开始看到模型产生的输出与其输入一样好时,那么有损压缩的类比将不再适用。

Can large language models help humans with the creation of original writing? To answer that, we need to be specific about what we mean by that question. There is a genre of art known as Xerox art, or photocopy art, in which artists use the distinctive properties of photocopiers as creative tools. Something along those lines is surely possible with the photocopier that is ChatGPT, so, in that sense, the answer is yes. But I don’t think that anyone would claim that photocopiers have become an essential tool in the creation of art; the vast majority of artists don’t use them in their creative process, and no one argues that they’re putting themselves at a disadvantage with that choice.

大型语言模型能帮助人类创作原创作品吗?要回答这个问题,我们需要明确我们所说的这个问题是什么意思。有一种被称为施乐艺术(Xerox art)或复印艺术的艺术流派,艺术家利用复印机的独特属性作为创作工具。像 ChatGPT 这样的复印机肯定也有可能实现类似的事情,所以,在这个意义上,答案是肯定的。但我不认为有人会声称复印机已成为艺术创作的重要工具;绝大多数艺术家在创作过程中不使用它们,也没有人争辩说他们因为这个选择而处于劣势。

So let’s assume that we’re not talking about a new genre of writing that’s analogous to Xerox art. Given that stipulation, can the text generated by large language models be a useful starting point for writers to build off when writing something original, whether it’s fiction or nonfiction? Will letting a large language model handle the boilerplate allow writers to focus their attention on the really creative parts?

所以让我们假设我们谈论的不是一种类似于施乐艺术的新写作流派。鉴于这一规定,大型语言模型生成的文本能否成为作家在撰写原创作品(无论是小说还是非小说)时的有用起点?让大型语言模型处理样板文件是否会让作家将注意力集中在真正有创意的部分?

Obviously, no one can speak for all writers, but let me make the argument that starting with a blurry copy of unoriginal work isn’t a good way to create original work. If you’re a writer, you will write a lot of unoriginal work before you write something original. And the time and effort expended on that unoriginal work isn’t wasted; on the contrary, I would suggest that it is precisely what enables you to eventually create something original. The hours spent choosing the right word and rearranging sentences to better follow one another are what teach you how meaning is conveyed by prose. Having students write essays isn’t merely a way to test their grasp of the material; it gives them experience in articulating their thoughts. If students never have to write essays that we have all read before, they will never gain the skills needed to write something that we have never read.

显然,没有人能代表所有作家发言,但请允许我提出一个论点:从非原创作品的模糊副本开始并不是创作原创作品的好方法。如果你是一名作家,在写出原创作品之前,你会写很多非原创作品。花在这些非原创作品上的时间和精力并没有浪费;相反,我认为这正是让你最终能够创造出原创作品的原因。花在选择正确的词语和重新排列句子以使其更好地衔接上的时间,教会了你散文是如何传达意义的。让学生写论文不仅仅是测试他们对材料掌握程度的一种方式;它给了他们表达思想的经验。如果学生从来不需要写我们都读过的论文,他们就永远无法获得写出我们从未读过的东西所需的技能。

And it’s not the case that, once you have ceased to be a student, you can safely use the template that a large language model provides. The struggle to express your thoughts doesn’t disappear once you graduate—it can take place every time you start drafting a new piece. Sometimes it’s only in the process of writing that you discover your original ideas. Some might say that the output of large language models doesn’t look all that different from a human writer’s first draft, but, again, I think this is a superficial resemblance. Your first draft isn’t an unoriginal idea expressed clearly; it’s an original idea expressed poorly, and it is accompanied by your amorphous dissatisfaction, your awareness of the distance between what it says and what you want it to say. That’s what directs you during rewriting, and that’s one of the things lacking when you start with text generated by an A.I.

而且情况并非一旦你不再是学生,就可以安全地使用大型语言模型提供的模板。表达思想的挣扎并不会在你毕业后消失——它可能发生在你每次开始起草新作品时。有时只有在写作的过程中,你才会发现你的原创想法。有些人可能会说,大型语言模型的输出看起来与人类作家的初稿并没有太大不同,但是,再一次,我认为这只是表面的相似。你的初稿不是一个表达清晰的非原创想法;它是一个表达得很糟糕的原创想法,伴随着你模糊的不满,你意识到它所说的和你想要它说的之间的距离。这正是指导你重写的动力,也是当你从人工智能生成的文本开始时所缺乏的东西之一。

There’s nothing magical or mystical about writing, but it involves more than placing an existing document on an unreliable photocopier and pressing the Print button. It’s possible that, in the future, we will build an A.I. that is capable of writing good prose based on nothing but its own experience of the world. The day we achieve that will be momentous indeed—but that day lies far beyond our prediction horizon. In the meantime, it’s reasonable to ask, What use is there in having something that rephrases the Web? If we were losing our access to the Internet forever and had to store a copy on a private server with limited space, a large language model like ChatGPT might be a good solution, assuming that it could be kept from fabricating. But we aren’t losing our access to the Internet. So just how much use is a blurry JPEG, when you still have the original?

写作没有什么神奇或神秘的,但它不仅仅是将现有文档放在不可靠的复印机上并按下打印按钮。有可能在未来,我们将构建一个能够完全基于其自身对世界的体验写出优秀散文的人工智能。我们实现那一天的日子确实将是重大的——但那一天远远超出了我们的预测范围。在此期间,有理由问,拥有一个重新措辞网络的东西有什么用?如果我们永远失去互联网访问权限,并且不得不将副本存储在空间有限的私人服务器上,像 ChatGPT 这样的大型语言模型可能是一个很好的解决方案,假设它可以被阻止进行捏造。但我们并没有失去互联网访问权限。那么,当你仍然拥有原件时,一张模糊的 JPEG 到底有多大用处呢?