Notes
← All posts
2026-09-02 ·Experiment·《相声的有限元》 · 2012·CCL2019 · n = 4,046

The formula was real. My ruler was wrong. 公式是真的,不对的是我那把尺子

《相声的有限元》 contains a real formula for how long an audience will laugh, and a rather elegant one: its three optimality criteria compose so that the quality coefficient becomes exactly 1. Its output is seconds of laughter from a live audience; I fed it silent written jokes. Every claim came back null — and then I discovered my measuring instrument had a discrimination of 1.09 over chance, rebuilt it, and watched the answer move — into a place that is worse for the book, not better. 《相声的有限元》里有一个真实存在、而且相当自洽的公式,用来预测观众会笑多久——它的质量系数在三条判据同时满足时,会干干净净地等于 1。这个公式的输出单位是「现场观众笑了几秒」,而我拿 4046 条没有舞台、没有观众的书面笑话去量它,四条主张一条没过。后来我发现自己那把尺子的区分度只有 1.09(随机是 1.00),于是重造了一把——答案跟着变了。

1.000k, exactly

What the quality coefficient becomes when all three criteria hold. (20/7) × 0.35 is exactly 1, and ln(1) is exactly 0. Three optimality conditions that look unrelated on the page compose to normalise the coefficient away, collapsing the whole formula into a single sum of ideal laughter times. It is designed, not fitted — and it is the nicest thing in the book. 三条判据同时满足时,笑果质量系数正好等于 1。 (20/7)×0.35 恰好等于 1,ln(1) 恰好等于 0。三条写出来彼此看不出关系的最优条件,代进去之后正好把系数归一,整个庞大的公式塌缩成一句话:把每个包袱的理想笑果时间加起来。这是我在这本书里看到的最漂亮的一处设计,而它是设计出来的,不是凑出来的。

01 — The question起因

Can you compute whether a joke lands?笑话好不好笑,能算出来吗?

In 2012, two graduate students at Shanghai Jiao Tong University published a 917,000-character book arguing that you can. Li Hongye and Zheng Yu had spent about a decade running the university’s crosstalk club — 相声, the two-person Chinese stage form built out of a straight man and a comic — and they did something that, whatever else you think of it, almost nobody does: they sat down with recordings of their own shows and measured how long the audience laughed, gag by gag, to a tenth of a second.

Out of that came 《相声的有限元》, The Finite Element Method of Crosstalk. The title is a borrowing from engineering: chop a hard problem into small pieces, model each piece, add them up. The book’s central object is the 笑果预期总公式 — the total formula for expected laughter effect.

I wanted to know whether it works. Specifically, whether it could serve as a reward function: could you use it to rank jokes, to pick the best of several drafts, to train a language model to be funnier?

That question turned out to be much harder to ask properly than I expected, and the reason why is most of this post.

2012 年,两个上海交大的研究生出了一本 91.7 万字的书,回答说:能。

李宏烨和郑钰在交大相声协会待了差不多十年。他们做了一件不管你怎么看它、总之几乎没人做过的事:抱着自己历年演出的录音,一个包袱一个包袱地量观众笑了多久,精确到零点一秒。

这本书叫《相声的有限元》。书名是从工程里借的:把一个难题切成有限个小块,每块单独建模,最后加起来。全书的核心是一个东西,叫笑果预期总公式。

我想知道它管不管用。更具体地说,想知道它能不能当奖励函数使:能不能用它给笑话排序,从几个草稿里挑出最好的一个,甚至拿它去训练一个模型,让它写得更好笑。

结果这个问题光是把它问对就比我想的难得多,而这篇文章大部分篇幅讲的就是为什么。

02 — Ground rules先把话说在前头

What this post is not这篇文章不是什么

This is not a book review, and it is not a verdict on anyone’s academic work.

Here is the shape of what I did, stated plainly so nobody has to infer it. Their formula was written for performed two-person crosstalk in front of a live audience, and its output is denominated in seconds of laughter from that audience. What I had was a corpus of short written jokes scraped from the Chinese web — median about ninety characters, no stage, no performer, no audience, no laughter. I applied the formula there anyway, and it did not predict anything.

That is a thermometer problem. If you take a clinical thermometer, hold it in a furnace, and read 42 °C, you have not discovered that the thermometer is broken. You have discovered that you used the wrong instrument.

I am reporting the experiment because a well-documented mismatch is worth something — it tells the next person what not to do, and it produced a few side findings about the datasets that are genuinely useful. But every null below carries that caveat, and I would rather over-state it than let a reader walk away thinking this post demonstrated something it did not.

There is also a separate reason to be careful here. This book became briefly famous in China in 2018 for reasons having nothing to do with its contents, after its author appeared on a television programme and the exchange went badly. That episode is not what I am writing about, I have no view on it, and it has no bearing on whether the mathematics works. The book deserves to be read as a book.

这篇不是书评,也不是对谁的学术成果下判断。

我做的事情,说白了是这样,免得谁还要去猜。他们那个公式是给舞台上、有现场观众的对口相声写的,输出单位是观众笑了几秒。而我手里是一批从中文网上扒下来的书面短笑话,中位数九十来个字,没有舞台,没有演员,没有观众,也没有笑声。我硬把公式搬过去用了,然后它什么也没预测出来。

这是个温度计的问题。你拿一支体温计伸进炉膛,读出来 42 度,你并没有发现体温计坏了,你只是发现自己拿错了工具。

我还是把这次实验写出来,是因为一次记录清楚的错配本身有价值:它至少告诉下一个人别这么干。而且过程里顺手捞到几个关于数据集的发现,是真有用的。但下面每一个”不显著”都带着这个前提,我宁可把话说重一点,也不想让谁读完之后以为这篇文章证明了什么它其实没证明的东西。

另外还有一层需要小心的地方。这本书 2018 年在国内短暂地出过一阵名,起因跟书的内容毫无关系——作者上了一档电视节目,那场对话谈得不太愉快。那件事不是我要写的东西,我对它也没有立场,它跟这套数学成不成立更是半点关系都没有。这本书应该被当成一本书来读。

03 — Before the book在拿到书之前

A pile of guesses, and one thing I refused to guess一堆猜测,和一件我没敢猜的事

The book is out of print and not online. Before a copy turned up, everything had to come from what is public: the publisher’s table of contents, a 2018 long-form interview in Southern People Weekly, and a scattering of press coverage. From those I could establish a handful of things — that the book named a 第一公式, a 第二公式, a 补偿公式 and a total formula; that laughter duration was reported as quality coefficient × effect; that something called 三大判据, “the three criteria”, existed.

So I built a reconstruction. Not the original formula — a clearly-labelled guess at it, in its own config file, with every component tagged by how much public warrant it actually had. The scorer for the real formula was written to raise an exception, with a unit test asserting it never quietly fell back to the reconstruction.

And there was one thing I deliberately did not do. The publisher’s blurb named 三大判据 and never said what they were. I left them empty. The config file records the hole, in as many words: “Named in the publisher’s blurb; described nowhere. Inventing it would be fabrication.”

Hold that thought.

这本书早已绝版,网上也没有。在书到手之前,能用的只有公开信息:出版社网页上的目录、2018 年《南方人物周刊》的一篇长采访、以及零散的媒体报道。从这些里面我能确认几件事:书里点了名有第一公式、第二公式、补偿公式和一个总公式;笑声时长被报道成「效果质量系数 × 效果」;还有一个叫三大判据的东西存在。

于是我做了一个重建版。不是原公式——是一份明确标着”这是猜的”的仿制品,单独放在自己的配置文件里,每个部件都标注了它到底有多少公开依据。而真公式那个 scorer,我写成直接抛异常,还配了一个单元测试,专门守着它绝不会偷偷退回去用重建版顶替。

有一件事我是故意没做的。出版社简介里点了”三大判据”这个名字,但从头到尾没说是哪三条。我把它空在那儿。配置文件里白纸黑字写着:“出版社简介里点了名,但没有任何地方描述过。编一个出来就是伪造。”

这句话待会儿还要用到。

04 — The book arrives书到了

What the formula actually says公式长什么样

Then a copy turned up, and everything changed at once.

Here is 式 8-3, from page 64 — the 笑果预期总公式 itself:

后来书到手了,情况一下子全变了。

下面是式 8-3,出现在第 64 页,就是笑果预期总公式本身:

X = k · X₀ = √(Ĉ/C) / (1 + |ln(D + L_z)|) · Σᵢ yᵢ·Gᵢ / (1 + (2/Wᵢ)|Lᵢ − Sᵢ|)

Two factors. On the left, a laughter-quality coefficient k. On the right, the net laughter time X₀ — a sum over the gags in one unit. X comes out in seconds.

The variables, with the book’s own definitions:

两个因子相乘。左边是笑果质量系数 k,右边是笑果净时间 X₀,对一个单元里的各个单包袱求和。X 的单位是秒。

各个变量,用书里自己的定义:

Symbol符号Name名称Meaning含义Range取值
G搞笑机理standard laughter the mechanism alone produces, in seconds这个机理本身能产生多少笑,单位是秒0.5–2.0 s
y语言加分multiplier for linguistic artistry语言艺术性的加乘因数×1.0–5.0
D对比度share of the unit’s time spent setting up铺垫占单元总时长的比例optimum 0.35最佳值 0.35
L亮度可预料程度 — attention the audience has可预料程度——观众现有的注意力0–1
S深度可猜透程度 — how guessable the punchline is可猜透程度——底儿有多难猜0–1
W温度可信服程度 — how convincing the gag is可信服程度——这个包袱有多让人服气0–1
C单元总时长the unit’s duration on stage, seconds这个单元在台上演了多久,秒—

That triad in the middle — 深度, 亮度, 温度, the “three degrees” of a single gag — is the sharpest thing in the book, and I want to give it its due. Guessability, foreseeability and convincingness are three genuinely different properties, and running them together is exactly the mistake an untrained listener makes. The book works through an example where a gag feels deep but is actually a shallow gag with a temperature problem, and the distinction does real work. Whatever the formula’s fate, that piece of analysis stands on its own.

中间那三个——深度、亮度、温度,书里叫单包袱的三度——是这本书里最见功力的地方,我想专门说一句。可猜透、可预料、可信服,这确实是三件不同的事,而把它们混作一谈,恰恰是外行听相声时最容易犯的错。书里举了个例子,一个包袱听上去”好深”,其实深度很浅,问题出在温度上。这个区分是真能干活的。不管这个公式最后命运如何,这一段分析本身立得住。

05 — The nice part漂亮的地方

Three criteria that compose to exactly one三条判据,乘出来正好是 1

Remember 三大判据 — the three criteria I refused to invent. Here they are, from pages 64 and 65:

还记得三大判据吗,就是我当初没敢编的那三条。它们在第 64、65 页,长这样:

№编号Criterion判据Reads as意思是
式 8-5D = 0.35setup should occupy 35 % of the unit铺垫应该占单元的 35%
式 8-7D + L_z = 1contrast plus the main gag’s brightness sums to one对比度加主包袱亮度等于一
式 8-8Lᵢ − Sᵢ = 0attention available should equal attention required观众有的注意力,正好等于包袱要的注意力

Now watch what happens when all three hold at once. The equivalent-duration term (式 8-4) gives √(Ĉ/C) = √((20/7)·D), and (20/7) × 0.35 is exactly 1. The log term gives ln(D + L_z) = ln(1) = exactly 0. So:

现在看看三条同时满足的时候会发生什么。当量时长那一项(式 8-4)给出 √(Ĉ/C) = √((20/7)·D),而 (20/7) × 0.35 正好是 1。对数那一项给出 ln(D + L_z) = ln(1),正好是 0。于是:

k = 1, exactly. The entire quality coefficient normalises away, and the formula collapses to X = Σ yᵢ·Gᵢ — just the sum of each gag's ideal laughter time. k is bounded to [0, 1] and reaches 1 only at the joint optimum; every criterion you miss costs you a fraction of the laughter you could have had.

k 正好等于 1。整个质量系数干干净净地归一了,公式塌缩成 X = Σ yᵢ·Gᵢ——就是把每个包袱的理想笑果时间加起来。k 被限制在 [0, 1] 之间,只有三条判据同时达标时才取到 1;你差了哪一条,就按比例损失掉本来能拿到的笑声。

I checked this numerically and then wrote a unit test for it, because it is the kind of thing you want to be sure of. It holds exactly.

This is designed, not fitted. Someone chose 20/7 and 20/13 so that the piecewise contrast term would pass through 1 at D = 0.35, and chose a logarithm so that the first formula would zero out there. Three conditions that look unrelated on the page turn out to be one condition viewed from three angles. That is a nice piece of construction, and it is the moment in this project where I sat up straight.

And it retroactively justifies the one call I was most nervous about. Had I invented three plausible-sounding criteria back in section 03 — and I could have, easily, they would have looked fine — they would have been wrong, and no reader could have told. The real three are the formula’s optimality conditions. A guess would have quietly contradicted the mathematics while sitting in the config file with exactly the same confidence as everything else.

我先用数值验了一遍,然后专门写了个单元测试守着它,因为这种事你会想确认一下。它是精确成立的。

**这是设计出来的,不是凑出来的。**有人特意选了 20/7 和 20/13,让那个分段的对比度项在 D = 0.35 处正好穿过 1;又特意用了对数,让第一公式在那里正好归零。三条写在纸上看不出关系的条件,其实是同一件事的三个侧面。这是一手漂亮的搭建,也是我做这个项目全过程中,唯一一次坐直了身子的时刻。

而且它回头替我确认了当初最没底的那个决定。如果第 03 节那会儿我顺手编了三条听着挺像样的判据——我完全编得出来,而且看上去不会有任何毛病——那三条一定是错的,而且没有任何读者能看出来。真正的三条是这个公式的最优条件。一份编出来的东西会安安静静地躺在配置文件里,跟别的部分一样自信,一边悄悄跟数学作对。

06 — The catch关键的问题

Seven parameters. Two of them come from paper.七个参数,只有两个能从纸上算出来

Here is where my whole plan came apart.

G and y are properties of a script. G you get by classifying the gag’s mechanism — imagistic mechanisms are worth 0.5–0.9 s, logical ones 1.0–2.0 s, and there is a combination rule (式 9-2) for gags carrying several. y you get by identifying which of six categories of linguistic bonus the gag uses; a free pun is fixed at exactly 3.5. Both are readable off the page.

C, D, L, S and W are not. And the decisive passage is the analysis workflow in 第33讲, page 281, which lays out how you actually fill in a 笑果分析表:

我整个计划就是在这儿散架的。

G 和 y 是稿子的属性。G 靠判定包袱用的是什么机理拿到——形象搞笑机理值 0.5–0.9 秒,逻辑搞笑机理值 1.0–2.0 秒,一个包袱身兼数种时还有一条组合公式(式 9-2)。y 靠判定它用了六类语言加分里的哪一类拿到,其中自由双关是定值 3.5。这两个都能从纸面上读出来。

C、D、L、S、W 不行。决定性的一段在第 33 讲,第 281 页,那里写的是笑果分析表到底该怎么填:

Step 2. Classify each gag's mechanism and its standard language bonus, and compute the theoretical laughter time. ← from the script alone

Step 3. Record each gag's actual measured laughter time, compare it against the theoretical time, and back-compute the non-standard language bonus and — for gags showing a temperature effect — the brightness L and depth S. ← needs a recording of a live performance

第 2 步。分析每个单包袱的机理和它所含的标准语言加分,算出理论笑果时间。←只看稿子就行

第 3 步。统计每个单包袱的实际笑果时间和温象,跟理论笑果时间比对,倒推出非标准语言加分,以及有温象的单包袱的亮度和深度。←需要一份现场演出的录音

Read that again. Brightness and depth are not estimated from the script — they are back-computed from how long the audience actually laughed. The performance transcripts in the book are annotated accordingly, with entries like C: 8.0 s and C: 0.3 s running down the margin.

So the system is half predictive and half retrodictive, and it says so. Step 2 predicts; step 3 fits the remaining parameters to what happened. That is a perfectly respectable way to build an analytical instrument for a director reviewing a show. It is simply not a thing you can run on a joke you found on the internet.

What is left is step 2’s output, which the book names: 理论笑果时间, the theoretical laughter time, equal to y · G. The book uses it directly in worked examples — one on page 148 reads “理想笑果量为 1.6 s × 3.0 = 4.8 s”. And because k = 1 at the three criteria, X = Σ yᵢ·Gᵢ is the full formula evaluated at its own optimum.

That is what I could test. Not the formula — the formula’s best case.

再读一遍。**亮度和深度不是照着稿子估出来的,是拿观众实际笑了多久倒推出来的。**书里的演出稿边上就密密麻麻标着这些,一行一行的 C:8.0 s、C:0.3 s。

所以这套系统一半是预测,一半是事后归因,而且它自己就是这么写的。第 2 步做预测;第 3 步把剩下的参数拟合到实际发生的事情上。对一个要复盘一场演出的导演来说,这是一件完全站得住脚的分析工具。它只是没法拿去跑一条你从网上翻到的笑话。

剩下能用的,是第 2 步的产出,而书里给了它名字:理论笑果时间,等于 y · G。书里在例子中直接用它,第 148 页那段写的是”理想笑果量为 1.6 s × 3.0 = 4.8 s”。而因为三条判据下 k = 1,X = Σ yᵢ·Gᵢ 就是这个公式在它自己最优点上的取值。

这是我能测的东西。不是这个公式,是这个公式的最好情况。

07 — The results结果

Nothing predicted anything什么也没预测出来

The test set is the held-out split of CCL2019 subtask 2: 4,046 Chinese jokes, each graded by human annotators as weakly / ordinarily / strongly humorous. Splits are by template cluster, not by row, so no near-variant of a training joke sits in test. Correlations are Spearman with bootstrap intervals resampled over topic groups.

测试集是 CCL2019 子任务二的留出集:4046 条中文笑话,每条由人工标注为弱幽默 / 普通幽默 / 强幽默。切分按模板聚类而不是按行随机,所以训练集里某条笑话的近似变体不会跑到测试集里去。相关系数是 Spearman,置信区间用按话题分组重采样的 bootstrap。

Scorer打分方式Spearman ρSpearman ρ95 % CI95% 置信区间
learned model (TF-IDF + ridge)学出来的模型(TF-IDF + 岭回归)0.219[0.188, 0.250]
random随机0.024[−0.005, 0.055]
keyword markers关键词表0.004[−0.027, 0.035]
原始公式 X = Σ yᵢ·Gᵢ (regex labels — superseded in §09)原始公式 X = Σ yᵢ·Gᵢ (正则标签,第 09 节有更正)−0.058[−0.091, −0.025]
my reconstruction我的重建版−0.058[−0.088, −0.029]
its shuffled-coefficient control重建版的系数打乱对照−0.063[−0.095, −0.031]
length only只看长度−0.068[−0.095, −0.038]

The original formula’s text-computable part sits at chance — its interval spans zero. My reconstruction sits at −0.058, which is statistically the same place, and which is also indistinguishable from a control where I randomly permuted its own coefficients.

I also tested the book’s four individually falsifiable claims, one at a time, because that is a sharper test than the whole formula:

原始公式那个能从文字算出来的部分,落在随机水平上——置信区间跨过了零。我的重建版是 −0.058,统计上是同一个位置,而且跟”把它自己的系数随机打乱”的对照组也分不出来。

我还把书里四条可以单独证伪的主张一条一条测了,因为这比测整个公式要锐利得多:

Claim主张Result结果
C1 logical mechanisms beat imagistic onesC1 逻辑搞笑机理强于形象搞笑机理0.444 vs 0.478; one-sided p = 0.950.444 对 0.478;单侧 p = 0.95❌
C2 higher G means funnierC2 G 越高越好笑ρ = −0.032 (p = 0.20)❌
C3 a language bonus helpsC3 有语言加分会更好笑−0.027, CI [−0.047, −0.006]−0.027,CI [−0.047, −0.006]❌
C4 a higher share of logical mechanisms is betterC4 雅俗率越高越好ρ = −0.039 (p = 0.12)❌

C1 is the book’s central ordering claim, and not only does it fail to replicate, the point estimate leans the other way — though its interval includes zero, so the honest statement is “no detectable difference, and certainly not the claimed one”.

Four for four. And now the part that matters.

C1 是这本书最核心的排序主张,而它不但没复现,点估计还朝反方向偏——不过区间是跨零的,所以老实的说法是”看不出差别,而且肯定不是它说的那个方向”。

四条全没过。接下来是要紧的部分。

08 — Why this null is weak为什么这个「不显著」不能当真

Three reasons not to believe my own result三条理由,说明我自己的结果不该被当真

One: my labels are much worse than the book’s. The book has an expert read the script and assign the mechanism. I have regular expressions and word lists. When I validated those detectors against ChinesePun — a corpus of human-annotated Chinese puns — only homophone detection cleared a lift of 1.5 over a control; homograph detection came in at 1.05, which is chance. On the actual test set, 29 % of jokes got no mechanism at all, meaning G = 0. A null measured through bad labels is a weak null, and I cannot separate “the formula does not work” from “my classifier does not work”.

Two: the domain gap, again. The book’s G values are seconds of laughter from a live audience watching two performers. These are ninety-character jokes read silently off a screen. The formula’s own dependent variable does not exist in my data.

Three: the labels I was scoring against are barely a function of the text. More on this in the next section, but the short version is that when the same joke appears twice in this corpus, the two copies get the same funniness grade about 15 % of the time. There is a ceiling on any correlation measurable here, and it is low.

**第一,我的标签比书里的差太远。**书里是让一个懂行的人把稿子读一遍,判定机理。我用的是一堆正则和词表。我拿 ChinesePun——一个人工标注的中文双关语料——去验这些检测器,只有谐音检测的提升倍数过了 1.5;一词多义那个是 1.05,就是随机水平。在真正的测试集上,29% 的笑话我一个机理都没认出来,也就是 G = 0。用一套烂标签量出来的”不显著”,本身就不怎么显著,而我分不清这是”公式不管用”还是”我的分类器不管用”。

**第二,还是那个领域错配。**书里 G 的取值,是现场观众看着两个演员演出时笑了多少秒。我这儿是九十来个字、在屏幕上默读的段子。这个公式的因变量在我的数据里根本不存在。

**第三,我用来对照的那些标签,本身就不怎么取决于文本。**下一节细说,但一句话版本是:同一条笑话在这个语料里出现两次时,两份拷贝拿到同样等级的概率大约是 15%。这里能测出来的相关性有一个天花板,而且这个天花板很低。

So what does the experiment actually establish? That this formula, computed this way, on this corpus, does not rank jokes the way Chinese annotators did. That is a much smaller claim than "the formula does not work", and it is the only one the data supports.那这次实验到底证明了什么?证明了这个公式、按这种算法、在这个语料上,排出来的顺序跟中文标注者的判断对不上。这比"公式不管用"要小得多,而且这是数据唯一支持得起的说法。
09 — Fixing the worst objection把最大的问题修掉

I redid the labels with a model, and the answer moved我把标签用模型重做了一遍,答案变了

The first objection in section 08 — that my regular expressions were too crude to be testing anything — is the only one of the three that money can fix. So I fixed it.

A local 35B model, in a container on the same machine, read the book’s definitions of the ten 搞笑机理 and six 语言加分 categories and labelled 21,765 jokes, zero errors, nothing leaving the box. One rule made this worth doing at all:

第 08 节那三条理由里,只有第一条是花钱能解决的——我那堆正则太糙,糙到根本没在检验什么。那就修。

一个 35B 的模型,跑在同一台机器的容器里,读着书里对十种搞笑机理和六类语言加分的定义,标了 21,765 条笑话,零错误,数据一步都没离开这台机器。有一条规矩让这件事才值得做:

The model classifies. The book supplies every number. The model says which mechanism and which language bonus a joke uses. It is never asked how many seconds, or how funny anything is — the annotation schema contains no numeric field at all, and a test asserts it. Ask a model for the seconds and you have replaced the table you were trying to test.

模型只做分类,秒数全部由书来定。模型只说这条笑话用了哪个机理、哪类语言加分。它从不被问「几秒」或「多好笑」——标注格式里一个数值字段都没有,而且有测试守着这一点。一旦让模型报秒数,你要检验的那张表就被它替换掉了。

The labels really are better. Measured against ChinesePun’s 1,052 human-annotated puns, with a control of CCL2019 jokes that ChinesePun’s own texts have been removed from:

标签确实好了很多。拿 ChinesePun 的 1052 条人工标注双关做召回,拿扣掉重叠之后的 CCL2019 做误报:

labeller标注方式recall on real puns真双关召回fires on control普通笑话误报lift区分度balanced acc平衡准确率
regex正则0.6410.5891.090.526
LLM0.7820.2123.680.785

Lift 1.09 is chance. The regular expressions were not a weak instrument, they were not an instrument. And they left 61.5 % of jokes with no mechanism at all, against 3.6 % for the model.

So the number I published above is wrong, and here is the correction.

区分度 1.09 就是随机。那些正则不是「不太灵」,是根本不是仪器。而且它有 61.5% 的笑话认不出任何机理,模型是 3.6%。

所以我上面写的那个数字是错的,这里更正。

the book’s formula, scored with…书里的公式,用什么标签打分Spearman ρSpearman ρ95 % CI95% 置信区间
regex categories (what section 07 reported)正则类别(第 07 节报的就是这个)−0.058[−0.091, −0.025]
LLM categoriesLLM 类别−0.005[−0.035, +0.025]

The formula does not sit below chance. It sits at chance. My measurement was not merely noisy — it was biased against the book, and I would not have found that without redoing it.

这个公式不是低于随机,它就在随机线上。我之前那次测量不只是有噪声,它是系统性地低估了这本书——不重做一遍我根本发现不了。

The experiment that actually separates two things

The book fixes G and y by table. Take the same LLM categories, throw the book’s seconds away, and fit weights on the training split instead. That distinguishes two very different failures: “these categories say nothing about funniness” versus “these categories say something, and the book’s seconds are not how to read it”.

真正能分开两件事的那个实验

书里的 G 和 y 是查表定死的。现在把同一批 LLM 类别拿过来,把书里的秒数扔掉,改成在训练集上 把权重学出来。这就能分开两种完全不同的失败:「这些类别跟好不好笑没关系」,和「这些类别是有信息的, 只是书里那套秒数不是提取它的正确方式」。

scorer打分方式Spearman ρSpearman ρ95 % CI95% 置信区间
the book’s G·y weights书里的 G·y 权重−0.005[−0.035, +0.025]
same categories, weights fitted同样的类别,权重学出来+0.124[+0.091, +0.155]
reference: length alone*参照:*只看长度−0.068[−0.095, −0.037]
reference: learned TF-IDF model*参照:*学出来的词频模型+0.219[+0.187, +0.251]

The book's seconds are not the weights that extract whatever its categories know. Fitted on the same taxonomy the score reaches +0.124 — clearly above the length confound at −0.068, so the categories are not noise. But "not noise" is a low bar, and the next section is where that number gets put in its place.

书里那套秒数,不是把它的分类里所知道的东西提取出来的正确权重。在同一套分类上把权重学出来能到 +0.124——明显高过「只看长度」的 −0.068,所以这些类别不是噪声。但「不是噪声」是个很低的门槛,下一节就是给这个数字定位置的。

The four claims stay unsupported, but their failure changes character. Under regex labels C1’s point estimate pointed the wrong way and C3 was significantly negative; under LLM labels both are simply flat — C1 differs by −0.001, C3’s interval spans zero. The book is not contradicted here. It is unconfirmed.

Two bugs the rebuild caught in my own work

I leaked. My first prompt used a worked example lifted from CCL2019 — a 0.808-similarity match to a held-out test row, which is showing the model an answer it will later be graded on. Every example is now hand-written, and a check refuses to run the labeller if any of them appears in an evaluated corpus.

My prompt made one label meaningless. In the first version the model put 欲擒故纵 on 67 % of all jokes — nearly every joke’s ending is a surprise, so the label carried almost no information — while finding only two jokes with an imagistic mechanism alone, which made the book’s central logical-versus-imagistic claim untestable. Splitting the mechanism guidance from the pun guidance fixed both.

Neither of those is a fact about the book. They are facts about me, and I would have published the first one without noticing.

What still is not settled

A model reading the book’s taxonomy is not a person who has read the book. On puns there is human ground truth and the table above measures the gap. For the mechanism categories no corpus annotates anything, so accuracy cannot be measured — only stability. Two different models agree on the primary mechanism 0.907 of the time; but three sampled runs all agree only 0.425 of the time. Mechanism assignment is soft. The +0.124 is obtained despite that, which is mildly reassuring and not a substitute for human labels.

And the two objections money cannot fix are untouched: these are still silent written jokes with no audience, and the labels I am scoring against still agree with themselves 15 % of the time.

四条主张仍然全部不成立,但失败的性质变了。用正则标签时,C1 的点估计是朝反方向偏的,C3 是显著为负的;换成 LLM 标签,两条都只是平的——C1 差 −0.001,C3 的区间跨过零。这本书在这里没有被推翻,它只是没有被证实。

重做过程中抓到我自己的两个错

**我泄题了。**第一版 prompt 里的示例,是我从 CCL2019 里随手挑的——跟一条留出测试样本相似度 0.808,等于把答案提前给模型看了。现在所有示例都是手写的,而且加了一道检查:只要有任何一条出现在评测语料里,标注程序就拒绝运行。

我的 prompt 让一个标签失去了意义。第一版里,模型把「欲擒故纵」打在了 67% 的笑话上——几乎每条笑话结尾都有意外,这个标签基本没携带信息——同时只找出两条纯形象机理的笑话,直接导致书里最核心的「逻辑 vs 形象」那条主张没法测。把机理的指引和双关的指引拆开写,两个问题一起解决了。

这两条都不是关于这本书的事实,是关于我的事实。而其中第一条,我原本会毫无察觉地发出去。

还是没解决的

一个模型读着书里的分类体系,不等于一个读过这本书的人。双关这一轴有人工 ground truth,上面那张表量出了差距。但搞笑机理那十几类,任何语料里都没有人工标注,所以准确率测不了,只能测稳定性。两个不同模型在主机理上一致率 0.907;可是三次采样全都一致的比例只有 0.425。**机理判定是软的。**那个 +0.124 是在这种软标签下拿到的,这一点让人稍微放心一些,但它替代不了人工标注。

而那两条花钱解决不了的,一条没动:这些依然是没有观众、默读的书面笑话,我用来对照的标签依然只有 15% 的时候跟自己一致。

10 — Putting +0.124 in its place给 +0.124 定个位置

A bag of characters beats it, twice over一袋字符片段,赢它一倍

The previous section could be read as a rescue: the book’s categories are informative after all, they were just priced wrong. Before anyone reads it that way, here is the comparison that matters.

Take the same training split. Fit three models. One sees only the book’s categories — about twenty binary flags per joke. One sees only character n-grams, the dumbest text feature there is, no theory of humour whatsoever. One sees both.

上一节容易被读成一次挽救:书里的分类其实是有信息的,只是价钱标错了。在有人这么读之前,先把真正要紧的对比摆出来。

同一份训练集,训三个模型。一个只看书里的分类——每条笑话约二十个是非标记。一个只看字符 n-gram,文本特征里最笨的那种,对幽默不含任何理论。还有一个两样都看。

model模型features特征数Spearman ρSpearman ρ95 % CI95% 置信区间
the book’s categories, weights fitted书的分类,权重学出来20+0.1238[+0.0910, +0.1547]
character n-grams alone只看字符 n-gram300,000+0.2470[+0.2176, +0.2782]
both together两样都看300,020+0.2576[+0.2288, +0.2888]

Even after discarding the book's seconds and letting the weights be learned from data, a three-second lexical baseline is twice as good. For the practical job of scoring a written joke, the book's apparatus is beaten — comfortably — by the dumbest thing you could try.

即便把书里的秒数全扔掉、让权重从数据里自己学出来,一个三秒钟就能训完的词频基线还是强了一倍。就「给书面笑话打分」这件实事而言,书里那套东西被你能想到的最笨的做法**明显地**打败了。

Do the categories add anything on top?

That is the question worth asking, and the one I could not answer by reasoning. If the categories see something the raw characters miss, they earn their place even while losing on their own. So: does the combined model beat the text-only model?

Δρ = +0.0106, 95 % confidence interval [−0.0016, +0.0249], with 96.3 % of bootstrap resamples positive.

The interval includes zero, so by the criterion I set in advance this fails. But I am not going to write it up as “adds nothing” — 96 % positive is not an absence. The honest reading is a very small positive effect this sample cannot separate from zero. Either way the ceiling is about +0.025 on a base of +0.247: a few percent, not a change of kind.

The number I did not expect

The two models rank jokes only +0.188 alike.

They are not doing the same thing at all. The book’s twenty flags genuinely see something that three hundred thousand character n-grams do not. And yet putting them together barely moves the needle. Which means:

那把它加到统计模型上面呢?

这才是值得问的问题,而且是我光靠推理答不出来的。如果这些类别看到了原文字面看不到的东西,那它就算单打独斗输了,也仍然有存在价值。所以:两样都看的模型,能不能赢过只看文字的?

Δρ = +0.0106,95% 置信区间 [−0.0016, +0.0249],bootstrap 里 96.3% 为正。

区间跨过了零,按我事先定的标准,没通过。但我不打算把它写成「没用」——96% 为正不是「不存在」。诚实的说法是:很可能有一点点正效应,只是这个样本量分不出它和零的差别。而且不管怎么算,天花板也就是在 +0.247 上再加 0.025:是零头,不是量变。

我没料到的那个数字

两个模型给笑话排的序,只有 +0.188 的相关性。

它俩根本不是在做同一件事。书里那二十个标记,确实看到了三十万个字符片段看不到的东西。可是把两者拼起来,指针几乎没动。这就意味着:

What the book's categories see differently is mostly not what makes a joke land. That is a more precise finding than "it does not work", and a less flattering one than the previous section on its own would suggest.

书里那套分类看到的那些「不一样的东西」,大部分跟笑话好不好笑没关系。这比「它不管用」精确,也比上一节单独读起来要难听一些。

The one thing this does not settle

The comparison is not a fair fight on capacity and was never meant to be. Three hundred thousand features against twenty. Compressing a joke into twenty binary flags and keeping half the ranking power is, read the other way, not a bad showing.

And the two are not competing for the same job. A character-n-gram weight is not something a writer can act on. “This gag’s brightness does not match its depth, and your setup is running long” is. The book was written as an editing instrument for people making crosstalk, not as a scoring function — and a model that wins on ρ while explaining nothing does not replace that.

But on the narrow question of prediction, which is the question I set out to ask: it loses.

这件事没有解决的

这场比较在容量上并不公平,本来也不打算公平。三十万个特征对二十个。反过来读:把一条笑话压缩成二十个是非标记,还能保住一半的排序能力,其实不算难看。

而且这两样东西争的不是同一份工作。一个字符 n-gram 的权重,创作者没法照着改稿子;「这个包袱的亮度配不上它的深度,而且你的铺垫铺长了」——这句话可以。这本书是写给做相声的人当编辑工具的,不是写来当打分函数的。一个赢在 ρ 上、却什么也解释不了的模型,替代不了它。

但就「预测」这个我一开始要问的窄问题而言:它输了。

11 — Side findings顺手捡到的

Two things about the datasets, which may be more useful than anything above两个关于数据集的发现,可能比上面所有内容都有用

CCL2019 subtask 1 is not a humor label. This benchmark has two subtasks and gets cited fairly often. Subtask 2 is humor grading, which is what I used. Subtask 1 is widely described as “humor recognition” — and when I read the official task page, the labels are:

**CCL2019 子任务一不是幽默标签。**这个评测有两个子任务,被引用得还挺多。子任务二是幽默等级划分,也就是我用的那个。子任务一常常被描述成”幽默识别”——而我去翻官方任务说明,标签写的是:

0 = 计算机生成幽默 · 1 = 非生成幽默

That is machine-written versus human-written. It is a provenance label. It has nothing to do with whether anything is funny.

那是机器写的还是人写的。这是个来源标签,跟好不好笑没有半点关系。

I put a hard guard in the code so this label can never reach a funniness computation — the type raises if you try. If you are using this benchmark, check which subtask you actually loaded.

The graded labels disagree with themselves. After normalising away punctuation and whitespace, 125 texts in the corpus appear more than once. Same joke, twice. The two copies should carry the same grade.

They do so 15 % of the time, against a chance rate of 38.5 % from the label marginals.

Below chance is itself informative, and it is worth being careful about what it means. Pure annotator noise pushes agreement toward chance, never below it. Landing well below is the signature of deduplication applied within each label class — if identical strings were removed inside a class but not across classes, then nearly every surviving duplicate has to straddle two classes. That matches the raw files, which contain no exact-string duplicates at all; the repeats I found differ only in whitespace and quotation marks.

So the honest reading is narrow: the same text can and does carry different labels here, no text-only scorer can fit these labels exactly, and the true inter-annotator agreement is not recoverable from this data. The organisers published no annotator count and no agreement figure.

我在代码里加了一道硬拦截,这个标签永远进不了好笑度的计算——类型层面就会抛异常。如果你在用这个评测,检查一下你到底载入的是哪个子任务。

**幽默等级标签跟它自己都对不上。**把标点和空格洗掉之后,语料里有 125 条文本出现了不止一次。同一条笑话,出现两遍。两份拷贝应该拿到同一个等级吧。

只有 15% 的时候是这样,而按标签的边际分布,纯靠瞎撞也有 38.5%。

低于随机本身是有信息量的,但得小心它到底说明了什么。纯粹的标注噪声只会把一致率推向随机水平,不会推到随机以下。掉到随机以下,通常是在每个等级内部单独去重留下的痕迹——如果完全相同的字符串在同一等级内被删掉了、但跨等级没删,那么活下来的重复项几乎必然横跨两个等级。这跟原始文件对得上:原始文件里一条精确重复都没有,我找到的这些只在空格和引号上有差别。

所以老实的读法很窄:同一段文本在这里确实会拿到不同的标签,任何只看文字的打分器都不可能精确拟合这批标签,而真正的标注者一致率没法从这批数据里还原出来。主办方没有公布标注人数,也没有公布一致率。

12 — As a reward function拿它当奖励函数

Two objectives that barely touch each other两个几乎不相交的目标

The original question was whether the formula could train a model. So I ran it: LoRA GRPO on a 1.5B Chinese instruction model, with a frozen learned human-preference model logged alongside as an independent yardstick — loaded in every arm, including the one where it contributes nothing to the reward. That is the point of it.

最初的问题是这个公式能不能拿来训模型。于是我真跑了:在一个 1.5B 的中文指令模型上做 LoRA GRPO,旁边挂一个冻结的、学过真人评分的模型当独立标尺——每个实验组都载入它,包括那个它对奖励毫无贡献的组。挂它就是为了这个。

Arm (lr 5e-5, 300 steps)实验组(lr 5e-5,300 步)Δ formula scoreΔ 公式分Δ frozen human proxyΔ 冻结的人类代理分
optimising the formula优化公式+0.055+0.008
optimising the human proxy优化人类代理+0.014+0.212

Each arm improves its own objective and barely moves the other — 14 % and 7 % transfer. Push the model towards the formula and the human proxy registers almost nothing; push it towards the human proxy and the formula score barely stirs. The two rewards are close to orthogonal, which is exactly what the ρ = −0.026 above predicts.

One control matters a lot here, and it cuts against an easy story. Both arms showed diversity collapse and shortening — in one human-only run, more of it (semantic diversity −0.116, mean length −17 characters). Those are generic pathologies of optimising any dense text reward, not something the formula did. It would have been easy, and wrong, to report them as evidence about 公式相声.

And two identical configurations differed by 23 characters in mean length, so none of these numbers should be read as a precise effect size.

每一组都在改进自己的目标,同时几乎不动另一个——迁移率 14% 和 7%。把模型往公式那边推,人类代理分几乎没反应;往人类代理那边推,公式分也基本不动。这两个奖励接近正交,而这正是上面那个 ρ = −0.026 预示的结果。

这里有一个对照特别重要,而且它砍掉了一个很好讲的故事。两组都出现了多样性坍缩和输出变短——在其中一次只优化人类代理的运行里,还更严重(语义多样性 −0.116,平均长度 −17 字)。这是优化任何稠密文本奖励都会有的通病,不是公式干的。把它们写成关于公式相声的证据,会很顺手,但那是错的。

另外,两次配置完全相同的运行,平均长度差了 23 个字。所以这些数字一个都不该当成精确的效应量来读。

13 — Grading my own guess给自己的猜测打分

The architecture was right. The physics was wrong.架子搭对了,里面的东西全错

The most enjoyable thing I got to do here was mark my own homework. The reconstruction from section 03 was written down precisely enough to be graded against the real formula, which is the whole reason for writing guesses down precisely.

这个项目里我做得最过瘾的一件事,是给自己批卷子。第 03 节那份重建版当初写得足够精确,所以现在可以拿真公式来打分——而把猜测写精确,图的就是这个。

My guess我猜的Reality实际
multiplicative two-factor core乘法的两因子结构式 8-3 is exactly X = k · X₀式 8-3 就是 X = k · X₀✅
contrast optimum at 0.35对比度最优值 0.35式 8-5: D = 0.35式 8-5:D = 0.35✅
language bonus is a multiplier语言加分是乘数y multiplies G, with a stated reasony 乘在 G 上,书里给了理由✅
logical mechanisms outrank imagistic ones逻辑机理排在形象机理之上1.0–2.0 s vs 0.5–0.9 s1.0–2.0 秒 对 0.5–0.9 秒✅
left 三大判据 empty把三大判据空着they are the three optimality conditions它们是三个最优条件✅
contrast term is a Gaussian around 0.35对比度项是绕 0.35 的高斯piecewise square root分段平方根❌
puns are mechanisms双关是机理puns are language bonuses — the other side of the product双关是语言加分——在乘积的另一边❌
my quality coefficient我那个质量系数bears no relation to k跟 k 毫无关系❌
— (never conceived)—(压根没想到)the 三度: brightness, depth, temperature三度:亮度、深度、温度❌

Architecture right, physics wrong. It recovered the multiplicative shape, the 0.35 constant, the multiplicative language bonus and the mechanism ordering — all four real and load-bearing — and it missed the single most distinctive idea in the book entirely.

But here is the finding that should bother anyone who does this kind of work, including me:

架子搭对了,里面的物理全错。它猜中了乘法结构、0.35 这个常数、语言加分是乘数、以及机理的排序——这四条都是真的,而且都是承重的——同时把这本书里最有特色的那个想法整个漏掉了。

但下面这个发现,应该让所有做这类工作的人(包括我自己)不太舒服:

Measured through the regular expressions, the real formula and my partly-wrong guess scored identically — −0.058 and −0.058, statistically the same place. I took that as evidence the test could not tell a correct formula from an incorrect one. It was. Once the labels were redone with a model (section 09), the real formula moved to −0.005 while the reconstruction, which is built on those same regular expressions, stayed where it was. The two had looked identical because the instrument could not separate anything at all.

用那堆正则去量,真公式和我半错的猜测,分数一模一样——−0.058 和 −0.058,统计上同一个位置。我当时把这当成「这个检验分辨不出对错公式」的证据。确实如此。等到第 09 节把标签用模型重做一遍,真公式挪到了 −0.005,而重建版——它本来就架在那堆正则上——原地没动。它俩之前之所以看着一样,是因为那把尺子根本什么都分不出来。

14 — What would settle it怎么才能真正说清楚

The experiment I did not get to run我没能做的那个实验

The honest way to test this theory is not the way I tested it. Two things would do it, in order:

Annotate by the book’s own method. Have someone who knows the book read three hundred jokes and assign the mechanism, the class and the standard language bonus the way 第33讲 prescribes — replacing my regular expressions with the expert judgement the formula was designed around. Then re-run the four claims. If they still fail under expert labels, that is a real result about the theory. If they pass, the failure was mine. This costs a person a couple of days and it is worth more than everything in this post.

Fill in one complete 笑果分析表 for a recorded performance. Script, per-gag laughter timings, brightness and depth back-computed the way the book says. That is the only way to exercise the actual formula rather than its best case — and it is testing it in the domain it was built for, in front of the audience whose laughter it is denominated in.

Until someone does the second one, nobody has tested 公式相声. Including me. Especially me.

老实说,检验这套理论的正确方式不是我用的这种。有两件事能做到,按顺序:

按书里自己的方法标注。找一个读过这本书的人,把三百条笑话过一遍,照第 33 讲的规矩定下机理、类别和标准语言加分——用这个公式当初就指望的专家判断,去替换掉我那堆正则。然后把四条主张重跑一遍。如果在专家标签下还是不成立,那才是关于这套理论的真结论;如果通过了,那说明之前失败的是我。这件事花一个人两三天,价值超过这篇文章里的全部内容。

给一场有录音的演出,完整填一张笑果分析表。演出稿、逐个包袱的笑声时长、按书里的办法倒推出的亮度和深度。这是唯一能跑通真正的公式、而不是它最好情况的路子——而且是在它本来就该待的领域里、在它的单位所指的那批观众面前测它。

在有人做完第二件事之前,没有人检验过公式相声。包括我。尤其是我。

15 — Coda末了

On measuring things that were not built to be measured your way关于用错工具去量一个东西

I set out to find whether a formula could predict whether a joke lands, and what I mostly found was how easy it is to build a test that cannot answer its own question. Three things were wrong with mine: detectors too coarse for the taxonomy, a corpus with no audience — which is what the formula’s units refer to — and labels that agreed with themselves 15 % of the time. Any one of them is enough to make a null uninformative. I had all three.

One of them turned out to be fixable, and fixing it was the most useful thing I did here. It also did not go the way I quietly hoped. The corrected measurement is kinder to the book in one place — the formula sits at chance rather than below it — and markedly less kind in another: strip out its seconds, fit the weights to its own categories, and a bag of character n-grams carrying no theory of humour still wins by a factor of two. The other two problems a GPU cannot touch.

What I did come away with, I am fairly confident about. The formula is real and it is carefully built — the way its three criteria compose to normalise the quality coefficient to exactly 1 is a genuinely elegant piece of construction, and the separation of guessability from foreseeability from convincingness is a sharper piece of analysis than most of what gets written about humor. Somebody sat with a stopwatch and a decade of recordings and tried to turn a craft into a measurable thing. That is a real attempt, and treating it as one is the least it deserves.

Whether the numbers hold up on a stage, I have not shown, and neither has anyone else. That question is still open, and it deserves a better instrument than mine.

我一开始想知道的是,一个公式能不能预测笑话好不好笑。最后主要弄明白的,是设计出一个回答不了自己问题的检验有多容易。我这次的检验有三个毛病:检测器对这套机理体系来说太糙;语料没有观众,而观众恰恰是这个公式单位所指的东西;标签跟自己只有 15% 的时候一致。三条里任何一条,都足以让一个「不显著」变得没有信息量——而我三条全占了。

**其中一条后来被修掉了,而修它是我这次做得最有用的一件事。它也没有往我暗暗希望的方向走。**更正之后的测量,在一处对这本书更宽厚了——公式落在随机线上,而不是低于随机;在另一处则刻薄得多:把它的秒数抽掉、用它自己的分类去学权重,一袋不含任何幽默理论的字符片段,照样赢它一倍。剩下那两个毛病,不是显卡能解决的。

**真正带走的东西,我倒是比较有把握。**这个公式是真实存在的,而且搭得很用心——三条判据乘出来把质量系数归到正好是 1,这是一手确实漂亮的构造;把”可猜透”、“可预料”、“可信服”三件事拆开,也比市面上大多数谈幽默的文字要锐利。有人拿着秒表,对着十年的演出录音,试图把一门手艺变成可以度量的东西。这是一次认真的尝试,把它当成一次认真的尝试来对待,是它最起码应得的。

至于这些数字在舞台上到底站不站得住,我没有证明,别人也没有。这个问题还开着,而且它值得一把比我这把好得多的尺子。