Notes
← All posts
2026-09-01 ·Experiment·talkie-1930-13b · Q4_K_M·Local · RTX 3060 12GB

A model that stopped reading in 1930 says today's Democrats are the Republican party 一个 1930 年就停止阅读的模型说:今天的民主党,是共和党

Take a model that has read nothing written after 1930 and use it as an instrument. Feed it only conduct — who votes for a party, what it does with tariffs, land, unions, churches, foreigners — with every name and every ideological label stripped out, and ask it who this is. It passes its own 1930 controls 32 times out of 34. Then it identifies today's Democratic coalition as the Republican party. 找一个只读过 1931 年之前文本的模型,拿它当一把尺子来量。只告诉它一个机构在做什么事,不给名字,不给年份,不给任何主义标签,让它猜这是谁。它先做对了 34 道 1930 年的常识题里的 32 道。然后它把今天的民主党纲领判成了共和党,把今天的共和党纲领判成了禁酒党。

+6.0nats

How far today's Democratic portrait sits from its own name. On the same ruler, today's Democratic coalition scores closer to the 1930 Republican party than to the Democratic party by six nats — a likelihood ratio of about four hundred to one. Today's Republican platform lands on the Prohibition party, the single-issue party of legislated private morals. 今日民主党的画像,被判成「共和党」时领先的幅度。 同一把尺子量下来,今天的民主党离 1930 年的共和党,比离它自己的名字还近 6 个 nats,差不多是四百倍的差距。而今天的共和党,最像的是那个只为一个私德议题存在的禁酒党。

01 — The idea设想

A witness who does not know how any of it turned out一个不知道结局的证人

talkie-1930-13b is a model trained entirely on text written before 31 December 1930. It has 13 billion parameters, which is not large; it has been instruction-tuned, but it has no reasoning-chain training and no think token in its vocabulary. None of that is the interesting part. The interesting part is that it has never seen a computer, a television, or a jet aircraft — and, more to the point, it has never seen how anything after 1930 turned out.

Think of it as a man from 1930 dropped into 2026. He has lived through the Crash and is in the middle of the Depression. He read about the German election of September 1930, where the Nazis went from 12 seats to 107. He knows the Soviet Union is midway through its first Five-Year Plan. He has been following the civil war in China. He probably has a great many questions for us. We certainly have more for him.

The early twentieth century is where a surprising amount of today’s politics starts. The Keynesian revolution and the idea that governments should manage demand; the Soviet system, whose residue is still with us; the Chinese civil war, already under way; Gandhi, who had just finished the Salt March. So here is the question worth putting to a witness like this. The institutions we think we understand today: how much of each one is still the same thing it was in his lifetime?

The experiment is simple to state. Describe a present-day institution using nothing but what it actually does, with no name, no date and no ideological label, then ask this 1930 model to say what it is. Its answer places today’s conduct on a map of political categories that was drawn before the map got redrawn.

talkie-1930-13b 这个模型的训练语料,全部写于 1930 年 12 月 31 日之前。它只有 13B 参数,不算大。它做过指令微调,但没有推理链训练,词表里连 think 这个标记都没有。这些其实都不重要。重要的是它没见过计算机,没见过电视,没见过喷气式飞机,更没见过 1930 年之后任何一件事情最后是怎么收场的。这情形很像《桃花源记》里那些人:不知有汉,无论魏晋。只不过桃花源里的人是自己躲进山谷,与外界断了几百年音信;它是被语料在 1930 年 12 月 31 日这天硬生生截断的。

你可以把它想象成一个 1930 年的人,昨天刚穿越到 2026 年。他经历过股市崩盘,正身处大萧条当中。他读过 1930 年 9 月的德国国会选举,纳粹的席位从 12 席一下子涨到 107 席。他知道苏联正在搞第一个五年计划。他也一直在关注中国的内战。他大概有一肚子问题想问我们,而我们想问他的只会更多。

二十世纪初本来就是个大变革的年代,今天很多政治格局都能在那儿找到源头。凯恩斯革命带来了宏观经济学,也带来了政府该不该干预经济这个问题;苏联那套体制的遗产到今天还在;中国的国共内战已经打起来了;印度那边,甘地刚刚走完食盐长征。所以有一个问题特别值得问这位证人:我们今天自以为很熟悉的这些机构,跟它们在他那个年代的样子比,还有多少是同一个东西?

实验本身一句话就能讲清楚。只用一个机构实际在做的事去描述它,不给名字,不给年份,不给任何主义标签,然后让这位 1930 年来的模型猜这是谁。它的回答,等于是把今天的行为放回一张 1930 年画好的分类图上,看它落在哪一格里。

这有点像刻舟求剑。那个成语通常是用来笑话人的:船早开出去老远了,他还守着船帮上刻的那道印子找剑。但这一次,刻舟是我们故意干的:我们要的不是找到剑,是量一量船到底漂出去了多远。

02 — The rules出题规则

Judge them by what they do, not by what they call themselves看它做了什么,不看它说了什么

Every description in this experiment is written from conduct, never from self-description. The reason is not squeamishness about propaganda. It is that political organisations advertise themselves in the most beautiful language available, and what they say and what they do routinely come apart.

Economics has a name for the principle: revealed preference. What someone actually chooses tells you what they actually want; what they announce tells you what they want you to believe. Applying that to institutions is the whole method here. Palpatine may declare himself the guardian of democracy, but if you were writing a quiz question about him, the description would have to read dictator of the galaxy.

There is a second rule, and it is easier to break than the first: no leakage. A description must not smuggle the answer in. If a question about a country mentions that its established church is Anglican, you have simply written the answer on the question paper. The same goes for a party whose “doctrine is common ownership”, or one that “rests on three principles of the people”. Those are labels, and labels are exactly what this experiment is trying to see past.

So every item is built from the same set of dimensions, filled in with facts: who votes for it, what it does about land and property, about capital and merchants, about organised labour, about religion and tradition, about how power changes hands, and about foreigners. Section 07 puts this rule on trial by feeding the model the parties’ own self-descriptions as items, and watching where they land.

这个实验里所有的描述,都是照着”这个机构实际干了什么”写的,绝不照着它自己的宣传写。倒不是我对宣传有什么洁癖。原因很简单:政治机构都会挑最漂亮的词把自己包装成仁义礼智信,而它说的和它做的,往往是两回事。

经济学里管这叫显示性偏好。一个人真正选了什么,才说明他真正想要什么;他嘴上宣称什么,只说明他希望你相信什么。把这条规矩套到机构身上,就是这篇文章全部的方法。帕尔帕廷可以自称是民主制度的守护者,但要是拿他出一道题,题面上只能写”银河的独裁者”。

还有第二条规矩,而且比第一条更容易犯:不许泄题。描述里不能偷偷把答案夹带进去。比方说一道猜国家的题,如果里面写了”这个国家的国教是圣公会”,那就等于直接把答案印在了卷子上。同样,说某个党”主张生产资料公有”,或者说它”以三民主义立党”,也是一个道理。这些都是标签,而这个实验想做的,恰恰是绕开标签。

所以每道题都按同一套维度来写,往里填的全是事实:谁投它的票,它拿土地和财产怎么办,拿资本和商人怎么办,拿工会怎么办,拿宗教和传统怎么办,权力怎么交接,对外国人什么态度。到了第 07 节,我会把这条规矩本身送上审判台,办法是直接把两党的自我宣传当成题目喂进去,看它们会落到哪儿。

03 — The instrument实验方法

You cannot simply ask it不能直接问它

The obvious way to run this is a multiple-choice question. It does not work at all. Here is what the model does with “a large four-legged beast used to draw ploughs and carriages”, where A. a horse is the right answer:

最直觉的做法当然是让它做选择题。结果这条路根本走不通。下面这道题是”一种用来拉犁和拉马车的四脚大兽”,正确答案是 A. 马,我们看看它给四个选项各打了多少分:

option选项ABCD
probability概率0.0760.4560.2250.243

The horse comes last. Not near last — last, at a seventh of the probability given to option B.

It is not answering the question. It is answering the format. Its entire strategy amounts to one piece of exam-room folk wisdom: when in doubt, pick B, you probably won’t be wrong. Where the answer sits in the list, and which letters happened to appear in the few-shot examples, swamp everything the question actually says. Any experiment built on this readout would have measured the shape of my prompt and nothing else.

马排最后。不是差一点垫底,是实实在在的最后一名,概率只有 B 选项的七分之一。

它根本不是在答题,它是在答格式。它脑子里只有一条朴素的考场经验:不会就选 B,多半不会错。 答案排在第几位,示例里出现过哪几个字母,这些东西的影响力完全压过了题目本身在说什么。要是拿这种读数去做实验,最后测出来的只是我的提示词长什么样,跟模型知道什么没有半点关系。

Read the probabilities, not the words读概率,不读它说的话

The fix is to stop asking it to produce a letter, and instead score answers it never gets to choose between. The inference engine here is llama.cpp, which supports GBNF grammars: you can force the decoder down one exact string and still read back the probability the model would have assigned to each token on its own.

So for a question like “this country’s head of state is directly elected, but its prime minister answers to parliament”, I do not ask the model to choose. I compute, separately, how likely it finds each of “France”, “Britain”, “the United States”, “China” as the continuation. Highest score wins. No letters, no option order, no format task of any kind.

That is the mechanism. The harder question is whether the reading needs correcting, and how. This model and I are, after all, from different centuries; before trusting anything it says, I had to work out how to talk to it properly. So I wrote a batch of items about 1930 whose answers are already known, tried several ways of communicating, and kept score.

解决办法是别让它吐 A、B、C、D,改成让它给答案打分,而且是它没得挑的答案。这次用的推理引擎是 llama.cpp,它支持 GBNF 语法。所谓语法约束,就是你可以按住模型的手,强迫它把某一个指定的字符串一个字一个字写完,同时你还能读回来:它自己本来会给每个字打多少分。

举个例子。题面是”这个国家的元首由全民直选产生,但总理要对议会负责”。我不让它选,我分别按住它的手,让它写”法国”,写”英国”,写”美国”,写”中国”,看它写哪一个的时候最顺、分数最高。分最高的那个就是它的答案。整个过程里没有字母,没有选项顺序,也没有任何格式任务。

机制就是这么回事。更麻烦的问题在后面:这个读数需不需要修正,怎么修。毕竟我跟这个模型隔了快一百年,在相信它说的任何一句话之前,我得先搞清楚怎么跟它好好说话。于是我准备了一批 1930 年的常识题,答案都是已知的,然后换着法子跟它沟通,一种一种记分。

Four corrections, and which of them actually mattered四层修正,哪几层真的管用

Prompt-variant averaging. The worry here is that a particular phrasing pushes the model around. I used two instructions — “Name the thing described below” and “Read the following account, and say what it is” — crossed with two lead-ins — “This describes ” and “The thing here described is ” — giving four wordings per item, then averaged over them. Measured effect: essentially none.

PMI, or content-free calibration. A candidate’s score is the product of the probabilities of its parts, so intuitively a short name like the Republicans has an advantage over a long one like the Farmer-Labor party of the Northwest for reasons that have nothing to do with the question. And some names are simply commoner in the training text than others, which makes the model readier to say them. The standard fix is to subtract the score the candidate gets under a content-free prompt: in effect log P(candidate | description) − log P(candidate). Measured effect: it made things slightly worse.

Surface-form averaging. The same entity has several names, and the model may prefer one of them on frequency alone. The Kuomintang and the Chinese Nationalist party are the same organisation, but the model may like the first far better because the English press of the day used it more. So each candidate is scored under three synonymous names, the Republican party / the Republicans / the Republican party of the United States, and averaged, which stops its preference for a string from drowning out its judgement about the entity. Measured effect: large. Accuracy on the 1930 control items went from 59% to 76%.

Marginal calibration. The model may also simply have favourites. Imagine a die-hard partisan who likes the sound of one party’s name and reaches for it whenever it appears among the options. To flatten that, I subtract from each candidate’s score the average score that candidate earns across the control items, so what gets compared is not the raw number but how far above its own baseline it landed. Measured effect: the largest of the four. Accuracy went from 76% to 94%.

The scoreboard below is 34 control items whose 1930 answers are known, balanced so that each candidate is the right answer equally often. Some sets are four-way and some six-way, which is why guessing scores 24% rather than 25%.

第一层,提示变体平均。 担心的是问法本身会把模型往某个方向推。我准备了两种指令,一种是”说出下面描述的是什么”,一种是”读一读下面这段记述,说说它是什么”;又准备了两种引导语,一种是”这段描述的是”,一种是”此处所描述之物为”。两两组合,每道题有四种问法,最后取平均。实测下来,这四种问法给出的结果几乎没有差别,这一层基本白加。

第二层,PMI,也叫内容无关校准。 一个候选答案的分数,是它每个组成部分概率的乘积。所以直觉上,the Republicans 这种短名字天生就比 the Farmer-Labor party of the Northwest 这种长名字占便宜,而这跟题目问的是什么毫无关系。另外,有些名字在训练语料里本来就出现得多,模型也就更容易脱口而出。标准的修正办法是先量一下这个候选在”什么内容都没有”的提示下能得多少分,然后减掉,也就是 log P(候选 | 描述) − log P(候选)。实测结果有点尴尬:不但没帮上忙,成绩反而略微掉了一点。

第三层,表面形式平均。 同一个实体往往有好几个名字,而模型可能纯粹因为词频就偏爱其中某一个。the Kuomintang 和 the Chinese Nationalist party 说的是同一个党,但模型很可能明显更认前者,因为当年的英文报纸更爱这么写。所以我给每个候选准备三个同义名,比如共和党就用 the Republican party、the Republicans、the Republican party of the United States,分别打分再取平均,这样它对某个具体字符串的偏好,就压不过它对这个实体本身的判断了。实测效果很大:1930 年常识题的准确率从 59% 提到了 76%。

第四层,边际校准。 模型还可能干脆就有心头好。你可以想象它是某个党的铁杆粉丝,只要选项里出现那个名字,它就想选。为了把这种偏好压平,我从每个候选的分数里减掉它在全部对照题上的平均分。最后拿来比较的不是原始分数,而是”这一题它比自己的平常水平高出多少”。这一层是四层里效果最大的:准确率从 76% 直接提到了 94%。

下面这张成绩单,测的是 34 道 1930 年答案已知的对照题。题目是配匀过的,每个候选当正确答案的次数一样多。其中有的题是四选一,有的是六选一,所以”瞎蒙”的基准线是 24%,而不是 25%。

readout读法correct正确
lettered multiple choice字母选择题12/34 · 35%
two-shot chain-of-thought二样例思维链14/34 · 41%
asked directly, free generation直接问,让它自己说16/34 · 47%
chain-of-thought + self-consistency (5 samples)思维链加自洽投票(采样 5 次)19/34 · 56%
forced likelihood scoring, bare强制打分,不加任何修正20/34 · 59%
+ prompt-variant averaging加上提示变体平均19/34 · 56%
+ PMI baseline加上 PMI 基线18/34 · 53%
+ surface-form averaging加上表面形式平均26/34 · 76%
+ marginal calibration加上边际校准32/34 · 94%

Two rows in that table are more interesting than the 94%, and both deserve a longer look than I can give them here.

The first is PMI. It is the textbook cure for exactly the disease I was worried about, and on this model it did not cure it. What cured it was averaging over synonymous names and then subtracting each name’s own baseline. Whatever is going on underneath, the problem was not that long names get penalised by arithmetic. It was that this model has opinions about particular strings.

The second is chain-of-thought, which is not the lever here: it is worse than not thinking at all. Two-shot reasoning scores 41%, against 47% for simply asking the question. Only by sampling five reasoning chains and taking a majority vote does it claw back to 56%, and even that loses to the crudest version of reading the distribution. Every generation-based readout loses to every scoring-based one.

这张表里有两行比那个 94% 更值得琢磨,而且都够单独写一篇了。

第一行是 PMI。它是教科书上专门治我担心的那个毛病的药,可在这个模型身上它没治好。真正治好的是”同义名取平均”加”扣掉自己的基准线”。不管底下的机制到底是什么,问题都不在于长名字在乘法上吃了亏,而在于这个模型对具体的字符串本身是有偏好的。

第二行是思维链。它在这儿不但不是关键,反而比不动脑子还差:让它按两个示例一步步推理,得分 41%;直接问它,反而有 47%。要采样五条推理链再投票,才勉强拉回 56%,而这个成绩连最粗糙的那版”读概率”都比不过。所有让它自己说话的读法,都输给所有给它打分的读法。

Why it fails is visible in the text it produces. Here it is on the Farmer-Labor party of 1930. It looks at three clues, notices that each one points at a different party, and then picks one anyway:

它为什么会失败,看它自己写出来的字就明白了。下面这道题的正确答案是农工党。它盯着三条线索,发现每一条指向的党都不一样,然后随手挑了一个:

“First, it draws strength from the north-west, which no party but the Republican does. Second, it demands relief for the farmer, which no party but the Democratic does. Third, it demands state ownership of elevators and mills, which no party but the Socialist does. Verdict: the Republican party.""其一,它的力量来自西北,除共和党外无党如此。其二,它要求救济农民,除民主党外无党如此。其三,它要求粮仓与面粉厂归州所有,除社会党外无党如此。结论:共和党。”

The form of the reasoning is impeccable. The conclusion is a coin toss. And the coin is loaded in a familiar way: given six American parties to choose among, the free-generation arm answered “the Socialist party” on half the control items, and the chain-of-thought arm answered “the Republican party” on half. That is the same answer-position bias that destroyed the multiple-choice design, wearing a different costume.

When the association is genuinely strong, though, the reasoning is good, and rather enjoyable to read. Here it is identifying the Russian Communist party of 1930, unprompted and correct:

推理的架子搭得无可挑剔,结论却是抛硬币定的。而且这枚硬币偏心的方式很眼熟:在六个美国政党里选,让它自己开口的时候,一半的对照题都答”社会党”;让它写推理过程的时候,一半都答”共和党”。这跟毁掉选择题的那种答案偏置是同一个东西,只是换了个马甲。

不过话说回来,当关联足够强的时候,它的推理是真的好,读起来还挺有味道。下面是它认 1930 年的俄国共产党,没人给它提示,答对了:

“First, the state holds all, so these are no private persons. Second, the peasants are driven, so they are no lovers of liberty. Third, religion is put down, so they are no Christians. Fourth, they build furiously, so they are no lovers of peace.""其一,国家掌握一切,故此辈非私人。其二,农民被驱赶,故此辈不爱自由。其三,宗教被压制,故此辈非基督徒。其四,他们疯狂建设,故此辈不爱和平。”

The technique that could not be made to run一条没跑起来的路

One approach I had high hopes for had to be abandoned, and it is worth writing down why.

Take a question of the form “how likely is it that this description belongs to the Republican party?” The normal thing to do is what I described above: under this description, compute the probability of “Republican”, “Democratic”, “Congress”, “Kuomintang”. But you can turn it around and compute, given that the answer is the Republican party, the probability of generating this description. That may well be fairer, because the thing being scored is then the same passage every time. In notation: score with logP(description | candidate) instead of logP(candidate | description). Length, word frequency and surface form all cancel exactly, and no calibration is needed at all. It is the cleanest fix available in theory, and on the first few items it worked and was accurate.

In practice I ran into a problem I had not anticipated. Scoring a description means forcing the model through a hundred-odd tokens of prescribed text, and llama.cpp’s grammar-constrained sampler is greedy, with no backtracking. Somewhere in a long forced string it picks the locally best token, from which the required text can no longer be completed, throws Unexpected empty grammar stack, and — this is the expensive part — every subsequent grammar request to that server fails until llama.cpp is restarted.

Put concretely, it is easy to see. Suppose the text I need it to produce contains the word cat. It emits the token ca, and then discovers that t is not among the tokens it may legally emit next. Because of how the sampler is built, there is no way to hand ca back and try again. So it dies there, and takes the rest of the run with it.

Short candidate names never dead-end; the forward arm ran more than five thousand scorings without a single failure. Long forced strings dead-end reliably. Three server restarts established the pattern before I gave up on the arm.

有一条路我本来觉得特别有戏,最后还是被迫放弃了,不过我认为值得记一笔。

设想这样一道题:“读完这段描述,你觉得它讲的是共和党的可能性有多大?“常规做法就是前面说的那个:在这段描述的条件下,分别算共和党、民主党、国大党、国民党的概率。但我们可以把方向反过来:假设答案就是共和党,那么这段描述被写出来的概率是多少?这样也许更公平,因为每次被打分的都是同一段话,长度、用词、名字长短这些干扰因素一次性全抵消了,连校准都不用做。写成式子就是用 logP(描述 | 候选名) 打分,而不是 logP(候选名 | 描述)。理论上这是最干净的方案,头几道题跑下来也确实又准又稳。

可惜实践中我撞上了一个完全没想到的坑。给一段描述打分,意味着要按住模型的手,逼它把一百多个 token 的既定文本一个不差地写完。而 llama.cpp 的语法约束采样是贪心的,不回溯。长串走到某个地方,它挑了一个当下看起来最合适的 token,结果剩下的文字再也拼不出来了,于是抛出 Unexpected empty grammar stack。代价还在后头:从这一刻起,发给这个服务端的所有语法请求全部失败,只能重启 llama.cpp。

举个具体例子就很好懂。假设我要它写出的文字里有 cat 这个词。它先吐了 ca 这个 token,然后发现下一步能合法吐出的 token 里根本没有 t。而按照采样器的设计,我没办法把 ca 收回来重算。于是它就卡死在这儿,顺带把后面整轮实验一起拖下水。

短的候选名永远不会走进死胡同,正着算的那套做法跑了五千多次打分,一次都没失败过。长串则是必然踩雷。我重启了三次服务端才摸清这个规律,然后放弃了这条路。

04 — The rule判定规则

What counts as a finding什么才算数

Every subject in this experiment gets two questions on the same dimensions. The first describes how it actually behaved around 1930. The second describes how it actually behaves now. Then:

这个模型对每一个考察对象都要做两道题。第一道写它 1930 年前后到底在干什么,第二道写它今天到底在干什么。然后:

A drift is recorded only when the 1930 control is answered correctly and the present-day item is not. If the model gets both wrong, the instrument is broken for that subject, and nothing is concluded from it.只有”1930 年那道答对了、今天这道答错了”,才算这个对象在九十多年里真的变了。 如果两道都答错,那说明模型在这个对象上本来就糊涂,整组按失效处理,一个字的结论都不许下。

I am convinced this rule is right, and it made me unhappy more than once, because several results I would have loved to report went into the bin because of it.

It also caught the one genuine embarrassment in the control set: the model swapped Mexico and Portugal with each other in 1930. Before laughing at it, read the two descriptions side by side. Both countries were poor. Both were Catholic. Both were deep in debt. Both had recently had a general raised to power. Both were quarrelling with foreign companies over concessions. This is not a stupid mistake; a reasonably well-read person could make it. But the rule does not care whether a confusion is forgivable. Neither country’s present-day item is admissible, and both were struck from the results.

这条规矩我确信是对的,但它让我难受过好几次,因为有些我特别想拿出来讲的结果,就是被它扔进了垃圾桶。

它还抓到了对照组里唯一一处真正的难堪:模型把 1930 年的墨西哥和葡萄牙互相认错了。不过先别急着笑话它,你把那两段描述并排读一遍就知道了。当年这两个国家都不富裕,都信天主教,都债台高筑,都刚有一位将军上台掌权,都在跟外国公司为特许权吵架。这个错犯得不算蠢,换个读过点书的人也可能栽在这儿。但规矩不管情有可原不情有可原:这两个国家今天那两道题一律不能用,全部从结果里划掉。

05 — The American result美国的结果

Where today’s platforms land in 1930今天的纲领,落在 1930 年的哪一格

The first design I tried was the obvious one: Democrat versus Republican, two options. It produced nothing usable. Of sixteen items, only three gave the same answer under all three calibration schemes I tried, and the headline verdict flipped between two runs that differed only in which items formed the calibration basis. The model was, to put it plainly, going round in circles.

That is a fault in the question, not in the model. In 1930 the American ballot had more than two parties on it, and this model has read about all of them. Confront it with a choice between only Democrats and Republicans and you have deleted most of the categories it actually thinks in. No wonder it flounders.

So I widened the field to six: Republican, Democratic, Socialist, Progressive, Farmer-Labor, Prohibition. The last four are worth a sentence each, because the results below do not read without them.

  • Socialist. A few hundred thousand votes and no state. Its standard-bearer ran for president five times, once from a prison cell. It wanted the mines, the railways and the great trusts in public hands, opposed the late war, and was strongest in the German and Jewish wards of the big cities and in the needle trades.
  • Progressive. An insurgency of western senators against the leadership of their own party. It denounced the power trust and the railroads, and demanded publicly owned water power, direct primaries and the recall of judges. It carried a state at the last presidential contest.
  • Farmer-Labor. A state party of the upper Midwest, where the wheat growers and the iron miners made common cause. It wanted relief for the mortgaged farmer, state-owned grain elevators and flour mills, and cooperative marketing. Its base was Scandinavian settlers.
  • Prohibition. Fifty years of candidates who never won, on one question: should the law forbid a thing many citizens wish to buy. The great parties took up its cause and wrote it into the constitution. Its strength was the Methodist and Baptist congregations of the countryside, and its platform speakers were as often clergymen as politicians.

Each of the six gets one 1930 control item written from its real behaviour at the time, and the model identifies all six correctly. That is a validated instrument, and a more useful one, because it can now say which 1930 party a modern platform resembles instead of merely which of two colours it is.

Then the two present-day coalitions, each written as a full portrait. Here is the Democratic one, in full, exactly as the model received it:

我最先试的是最直觉的设计:民主党对共和党,二选一。结果完全不能用。十六道题里,只有三道在我试过的三种校准方案下答案一致;更糟的是,两次运行仅仅是校准基准不同,主结论就直接翻了过来。说白了,这个模型被问得晕头转向。

但这是题目的毛病,不是模型的毛病。1930 年的美国选票上不止两个党,而这个模型对它们全都读到过。你逼它只在民主党和共和党里挑,等于把它脑子里真正在用的分类删掉了一大半,它当然一片混乱,不知如何是好。

于是我把选项摊开成六个:共和党、民主党、社会党、进步党、农工党、禁酒党。后面四个今天的中文读者大概都没听说过,得先花几句话交代清楚,不然下面的结果没法读。

  • 社会党。 每次大选拿几十万票,一个州也拿不下。它的招牌人物德布斯参选过五次总统,其中一次是在监狱里参选的。它主张矿山、铁路和大托拉斯收归公有,反对参加一战。票仓在大城市的德裔和犹太裔街区,还有成衣业的工人。
  • 进步党。 一群西部参议员在自己党内造反造出来的。它痛骂电力托拉斯和铁路公司,要求水电归公有,候选人由选民直接初选产生,法官可以被选民罢免。上一次总统大选它拿下过一个州。
  • 农工党。 中西部北部的一个州级政党,小麦农民和铁矿工人联起手来搞的。它要求救济欠着按揭的农民,要求粮仓和面粉厂归州所有,还要求谷物合作销售。基本盘是北欧来的移民。
  • 禁酒党。 五十年如一日地推候选人,一次也没赢过,只为一个议题活着:法律该不该禁止老百姓想买的东西。后来两个大党把它的主张抄了过去,直接写进了宪法。它的票仓是乡下的卫理公会和浸信会教堂,台上演讲的人里,神职人员跟政客一样多。

这六个党各配一道 1930 年的对照题,题面照着它们当年的真实做法写。模型六道全部答对。到这一步,这把尺子算是验准了,而且比之前好用得多:它现在能说出一份今天的纲领像 1930 年的哪一个党,而不只是像红蓝两种颜色里的哪一种。

接下来轮到今天这两个党。每一个都写成一幅完整的画像。下面这段就是民主党那道题的全文,一字不差,模型拿到的就是这些:

“Its strength lies in the great cities of the two coasts, among the universities and the learned professions, in the wealthiest and best-schooled wards, and among voters of the negro race and the newer immigrant stocks. It looks to the central government to relieve want and to regulate industry, and would have the state provide a physician for every citizen. It is friendly to the unions of public servants and schoolteachers. Its people go least often to church. It would keep the gates open to immigrants, and is the readier to send soldiers and money abroad.""它的根基在两岸的大城市,在大学和专业阶层,在最富有、受教育程度最高的选区,也在黑人选民和较新的移民族群当中。它指望首都的中央政府去救济穷困、管束工业,还主张国家该给每个公民配一位医生。它跟公务员工会、教师工会关系亲近。它的选民是全国最少上教堂的。它主张对移民开门,也比对手更愿意往海外派兵派钱。”

Not one party name, not one ideological word, nothing that could only have been written after 1930. Here is where it and the Republican portrait land. Scores are in nats, which are log-likelihood differences after calibration: a gap of one nat is roughly a factor of 2.7, and a gap of six is roughly four hundred. The single-digit numbers below are not small.

里面没有一个党名,没有一个主义,也没有任何一样 1930 年之后才有的东西。下面就是这段话和共和党那幅画像各自的落点。

分数的单位是 nats。它是校准之后的对数似然差,你可以粗略这么理解:差 1 个 nat,大约是 2.7 倍的差距;差 6 个 nats,就是四百倍上下。所以下面那些看着不起眼的个位数,其实都不小。

today’s platform今日纲领1st2nd3rdits own name它自己的名字
Republican portrait共和党画像Prohibition +9.8Republican +6.9Democratic +6.22nd, −2.9第 2 名,落后 2.9
Democratic portrait民主党画像Republican +7.9Progressive +5.8Socialist +3.24th, −6.0第 4 名,落后 6.0

Read the second row again, slowly. A witness who stopped reading in 1930 looks at that description and says: this is the Republican party. And not as a close second. The Democratic party’s own name comes fourth, behind the Progressives and the Socialists as well.

The first row is milder, and in its own way funnier. Today’s Republican portrait does score Republican, but only second. What beats it is the Prohibition party. Notice what the 1930 model is not reminded of: the party of high tariffs and banking houses, which the Republicans of 1930 also very much were. It is reminded of the party that existed to write private morals into law.

What pushed each portrait where it went is a question of specific planks, and the next section takes them one at a time.

Put plainly: the blue Democratic party of 2026 resembles the red Republican party of 1930 more than it resembles its own name. Blue drifts red. Red drifts, of all things, dry.

把第二行慢慢再看一遍。一个 1930 年就停止阅读的证人,读完刚才那段话,说:这是共和党。 而且不是勉强险胜。民主党自己的名字排第四,前面还有进步党和社会党。

第一行温和一点,但换个角度看更好笑。今天共和党的画像确实打出了共和党,只不过是第二名,压在它上面的是禁酒党。注意这里最有意思的地方:1930 年的模型没有想起那个高关税加大银行的党,而 1930 年的共和党偏偏就是那个党。它想起的,是那个专门把私德写进法律的党。

这两幅画像到底是被哪几条主张推过去的,下一节逐条拆开看。

说白了:2026 年的蓝色民主党,比它自己的名字更像 1930 年的红色共和党。蓝的往红里漂,红的呢,漂去了禁酒党。

06 — Plank by plank逐条纲领

Which 1930 party owns each modern policy今天的每一条主张,1930 年归谁

The composite portraits are long, which is exactly why they carry signal. Single-sentence planks are short, and correspondingly noisy. The honest way to present them is to show first which ones the model gets right by 1930 standards, because those are the planks whose association with a party is strong enough in the corpus to survive being compressed into one clause. Of sixteen, six reproduce the 1930 attribution:

刚才那种完整画像很长,所以信号足。单句纲领很短,噪声也就大。诚实的呈现方式是先看它按 1930 年的标准答对了哪几条,因为答对的那些,正是党派关联强到被压缩成一句话还扛得住的条目。十六条里有六条复现了 1930 年的归属:

a plank of today’s politics今天政治里的一条主张the 1930 party it belongs to1930 年它归谁
strength among rural Protestant farmers of the former Confederacy; distrusts the eastern banking houses, the great newspapers, and the learned professions根基在南方旧邦联和内陆的新教农民中间,不信任东岸的银行、大报和专业阶层Democratic +5民主党 +5
the law should enforce the moral customs of the churches, and forbid the sale of things many citizens wish to buy主张法律应当把教会的道德规矩管到私人头上,并且禁止出售许多公民想买的东西Prohibition +14禁酒党 +14
the party of the organised trade unions and the industrial workingman有组织工会和产业工人的党Socialist +10社会党 +10
relief for the mortgaged farmer, payments out of the treasury to keep him on his land救济欠着按揭的农民,从国库拨钱让他留在自己的地里Farmer-Labor +12农工党 +12
denounces the vast combinations that control what the public may read and hear, demands they be broken up痛骂那些控制着公众能读到什么、听到什么的巨型联合体,要求把它们拆开Progressive进步党
low duties, the freest possible trade with foreign nations低关税,尽可能自由地跟外国做生意Democratic +1民主党 +1

The third and fourth rows are the ones that build the composite result, and they are worth stating flatly. In 1930, the party of the trade unions is the Socialist party, not the Democratic one. And the party that pays a mortgaged farmer to stay on his land is Farmer-Labor. Both planks are now unremarkable furniture in the platforms of the two major parties. Neither belonged to them when this model was reading the newspapers.

The second row is the plank that carries today’s Republican portrait over to the Prohibition party. And the first row explains something that might otherwise look strange: why the rural, southern, anti-metropolitan base in that same portrait does not pull it back toward the Republicans. In 1930, that base was Democratic. It is the single most famous realignment in American political history, and the model reproduces it without being asked.

The other ten planks are noise, and are reported as noise: on the negro vote, on states’ rights, on high tariffs and on immigration restriction, the model does not reproduce the 1930 attribution. Those are precisely the one-clause items. One clause is not enough to move this model, and pretending otherwise would be cherry-picking.

第三行和第四行是撑起整幅画像结论的两条,值得直白说一遍。1930 年,工会的党是社会党,不是民主党。给欠按揭的农民发钱、让他留在自己地里的是农工党,一个大多数美国人今天根本没听说过的州级小党。这两条主张今天在两大党的纲领里已经普通得没人会多看一眼,可在这个模型读报纸的年代,它们哪个大党都不属于。

第二行就是把今天共和党的画像推到禁酒党那边去的那一条。第一行解释了另一件乍看很奇怪的事:同一幅画像里明明写着”农村、南方、反大都会”,为什么没有把它拉回共和党?因为 1930 年,这个基本盘是民主党的。这是美国政治史上最有名的一次选民重组,没人提醒,模型自己就把它复现了出来。

剩下十条是噪声,我也照噪声来报:黑人选票、州权、高关税、限制移民这几条,模型都没有复现 1930 年的归属。而它们恰好全都是只有一个从句的短题目。一句话推不动这个模型,硬要拿来说事就是挑数据了。

07 — Label and conduct标签与行为

Putting my own rule on trial把我自己的规矩送上审判台

“Judge them by what they do” is easy to assert and easy to violate: one careless clause can smuggle the answer in. So I wrote each party twice more. Once in the words it uses about itself, taken more or less from its own literature. Once from facts about who actually votes for it and what it actually does in office. If the rule is worth anything, the two versions should not land in the same place.

“看它做了什么”这句话,说起来容易,犯起来更容易,一个不小心的从句就能把答案偷渡进去。所以我把两个党又各写了两遍。一遍用它形容自己的话,基本照抄它自己的宣传口径。一遍用事实:谁真的投它的票,它上台之后真的干了什么。如果这条规矩有价值,这两遍就不该落到同一个地方。

description描述scored as判成
R, self-description共和党·自称“the party of business, of sound money, of thrift in the public purse, of the free market unhampered by the state""工商业的党,健全货币的党,节省公帑的党,不受国家干预的自由市场的党”Republican +20.2
R, conduct共和党·行为heavy duties on imports, subsidies to farmers, the largest deficits in the republic’s history, and the state punishing companies whose opinions it dislikes对进口品课重税,给农民发补贴,制造了共和国历史上最大的财政赤字,还让国家去惩罚言论不合心意的公司Republican +22.7
D, self-description民主党·自称“the party of the working man and of the common people, against wealth and privilege""劳动者和普通人的党,反对财富与特权”Republican +18.9, Democratic 3rd, +2.0共和党 +18.9,民主党第 3,+2.0
D, conduct民主党·行为best-schooled and among the best-paid voters, strongest where rents are highest, funded by finance, machine-makers, and the owners of the theatres and the newspapers选民学历最高,收入也在前列,最强的地方是租金最贵的选区,钱来自金融业、机器制造商,还有剧院和报纸的老板Democratic 1st, +1.3 over Farmer-Labor — see below民主党第 1,只领先农工党 1.3,见下文
what both now do两党的共同点both lay duties on foreign goods, both promise untouched pensions, both borrow without limit, both denounce the great combinations; they differ chiefly on private conduct and on who may enter两个党都课关税,都承诺绝不动养老金,都无限举债,都痛骂巨型联合体。真正的分歧在私德,以及准许谁进这个国家Republican +1.9

Two of those four land cleanly, with margins wide enough to trust: the Republicans’ self-description and their conduct both score Republican by four nats or more. The Democrats’ self-description misfires hard, filed under the Republicans by nineteen nats — exactly the kind of flattering error this whole method exists to catch.

The Democrats’ conduct is the interesting middle case, and it deserves less confidence than the table above suggests. Its uncalibrated score actually favours the Farmer-Labor party, not the Democrats. Calibration nudges the Democrats back into first place, but by only 1.3 nats over Farmer-Labor — a margin this article treats as noise everywhere else in it (section 08 waves off a 1.7-nat gap between Germany and Italy for exactly this reason). So no, a 1930 reader does not confidently conclude that rich, well-schooled, finance-funded voters are Democrats. What the reader does confidently avoid is filing them under the Republicans, which is the one thing the self-description gets wrong.

That asymmetry, not the narrow win, is the finding. A party’s self-description is not evidence about the party; it is evidence about what the party would like to be associated with. It is also the reason no item anywhere else in this experiment was allowed to contain a self-description.

四条里有两条落得干净利落,领先幅度够大,信得过:共和党的自称和它的实际行为,都稳稳打在共和党头上,都领先四个 nats 以上。民主党的自称摔得很惨,被 1930 年的读者整整齐齐归进了共和党,落后十九个 nats,这正是这套方法要抓的那种自吹自擂式的错误。

民主党那条实际行为,是中间那个有意思的例子,但它没有上面表格看上去那么靠得住。没校准之前,原始分数其实是农工党领先,不是民主党;校准之后民主党才勉强反超,可只领先农工党 1.3 个 nats。这个差距,按这篇文章别处的标准,早就该算噪声了(第 08 节里,德国输给意大利也就 1.7 个 nats,我就没敢拿它下结论)。所以老实说:1930 年的读者不会一口咬定,有钱、学历高、靠金融业撑腰的选民就是民主党。他能确定的只是一件事:这批人不是共和党,而这恰恰是民主党自称里唯一答错的地方。

真正站得住的,是这个不对称,不是那个勉强赢来的第一名。一个党怎么形容自己,说明不了这个党是什么样,只能说明它希望别人把它跟什么联系在一起。这也是为什么整个实验里,别的任何一道题都不许含有自我描述。

08 — Countries国家

A century changed nations less than it changed parties一百年改变国家,远不如改变政党那么狠

This part covers twenty-four countries. Each one gets two questions: what it was doing around 1930, and what it is doing now.

The items work like this. Four countries make a set, and the model chooses only among those four. One set is Germany, Sweden, Switzerland and Italy. Here is the whole of the present-day German item, exactly as it was given:

这一部分一共二十四个国家。每个国家出两道题:一道写它 1930 年前后在干什么,一道写它今天在干什么。

题目是这么组织的。四个国家凑成一组,模型每次只在这四个里面挑。比如其中一组是德国、瑞典、瑞士、意大利。下面是”今天的德国”那道题的全文,模型拿到的就是这些:

“A great manufacturing nation of central Europe. It keeps only a small army and has renounced war altogether; its people are among the least willing in the world to bear arms. It has admitted several millions of foreigners of an alien faith to settle within its borders and has made citizens of them. Its churches stand empty. It has shut its own coal mines. It pays large sums to relieve the poorer states of the continent, and buys its fuel from the empire to the east.""中欧一个了不起的制造业国家。它只保留很小的军队,而且彻底放弃了战争;它的国民是全世界最不愿意扛枪的那一批。它接纳了好几百万信奉另一种宗教的外国人在境内定居,并且给了他们国籍。它的教堂空着。它关掉了自己的煤矿。它拿出大笔钱去接济欧洲比它穷的国家,燃料则向东边那个帝国购买。”

No country name, no date, nothing that did not exist in 1930. The model has to pick one of the four.

Why four to a set, and why is each country the right answer exactly once per era? Because if Germany were the answer three times in one set, a model that merely likes the word “Germany” would score well without knowing anything, and the marginal calibration in section 03 would have its baseline dragged out of shape too. Balancing the sets is what stops guessing from working.

Six sets, twenty-four countries, forty-eight items. The 1930 half is there to test the instrument, not the world: it has to recognise a country from how it behaved back then, or I have no business trusting its verdict on how that country behaves now. It got 22 of those 24 right, missing only the Mexico–Portugal pair from section 04. On the present-day half it got 17 of 24.

The countries it still recognises today, from conduct alone, are Sweden, Switzerland, Italy, Siam, the Philippines, Russia, Turkey, Persia, Spain, the United States, Great Britain, Canada, Argentina, Brazil, Greece and Egypt.

Turkey is the most striking pass on that list. Its 1930 item is pure Kemalism: religious courts abolished, dervish lodges closed, the old headgear forbidden and hats commanded in its place, the Arabic script put away for the Latin, the vote given to women. Its present-day item is very nearly the mirror image: the largest mosques in the country’s history, religious schooling restored, generals and journalists in prison, a ruler who talks about restoring the greatness of the old empire. Two descriptions that contradict each other point by point, and the model names the same country both times. The state stayed recognisable while its founding project was run in reverse.

Seven countries it misses today. One of the seven is Mexico, which was already struck under the rule in section 04, so six are left. Here they are, with the margin by which the wrong answer won:

题面里没有国名,没有年份,也没有任何 1930 年不存在的东西。模型要做的,就是在德国、瑞典、瑞士、意大利里面挑一个。

为什么非要四个凑一组,而且每个国家在每个时代恰好当一次正确答案?因为假如德国在同一组里当了三次答案,那模型只要单纯喜欢”德国”这个词,什么都不懂也能考得不错;而且第 03 节讲的边际校准,基准线也会跟着被带歪。把每一组都配匀,为的就是让蒙这件事不管用。

六组,二十四个国家,四十八道题。1930 年的那一半不是用来看世界的,是用来验这把尺子的:它得先能从当年的做法里认出一个国家,我才有理由相信它对这个国家今天的判断。这一半它答对了 22 道,只错了 2 道,错的就是第 04 节说的墨西哥和葡萄牙那一对。今天的那一半,它答对了 17 道。

只凭行为、今天仍然被它认出来的国家有:瑞典、瑞士、意大利、暹罗、菲律宾、俄国、土耳其、波斯、西班牙、美国、英国、加拿大、阿根廷、巴西、希腊、埃及。

这份名单里最惊人的是土耳其。它 1930 年那道题写的是纯粹的凯末尔主义:废除宗教法庭,关闭德尔维希道堂,禁戴旧式头巾、强制戴礼帽,扔掉阿拉伯字母改用拉丁字母,给女性选举权。而它今天那道题几乎是逐条反过来的:建了国家历史上最大的清真寺,恢复宗教教育,把将军和记者关进监狱,统治者张口闭口是恢复旧帝国的荣光。两段互相打脸的描述,模型两次报出的却是同一个国家。国家一直认得出来,被倒着开回去的是它的立国方案。

今天认错的一共七个国家。其中墨西哥按第 04 节那条规矩已经作废,不算数,剩下六个列在下面。后面那个数字是错误答案赢了多少:

country国家today’s conduct reads as今天的行为被判成margin领先
Japan日本Siam+5.9
Poland波兰Greece+12.6
Germany德国Italy+1.7
India印度Japan+0.6
Ireland爱尔兰Great Britain+0.8
Austria奥地利Poland+3.5

Only the top two have margins worth arguing about. The bottom four are close enough that I would not build a sentence on them; Germany losing to Italy by 1.7 nats means the model is nearly indifferent between them, which is not a finding, it is a shrug.

Japan is the one I keep coming back to. An island empire that has not made war in eighty years, whose own fundamental law forbids it to make war, whose villages are emptying because its people are not having children, and which shelters under the protection of a distant Western power. To a reader of 1930, that is not Japan. Japan in 1930 is the country that hauled itself out of seclusion in sixty years, beat a European empire at sea, took a peninsula and an island for its own, and whose army officers were growing openly impatient with the civilian ministers at home. The model reads the modern description and says Siam: the small, pleasant, un-warlike Buddhist kingdom that lives on rice and travellers and never fell to a European power.

Poland is the other clean one, and it wins by the widest margin in the whole table: 12.6 nats.

Its set is Poland, Austria, Greece and Egypt. Today’s Poland, as the item puts it, is a country of the eastern plain, now almost entirely of one race and one church, grown rich unusually fast after throwing off a foreign yoke, its young gone west to work, its government at open war with its own judges.

The trouble is that the most conspicuous fact about Poland in 1930 is the exact opposite. It had just been reassembled on the map after a century and a half of partition between three empires. A marshal ruled it by coup while its parliament sat powerless. A third of its people were not Poles at all, but Jews, Ukrainians and Germans, and it quarrelled with every neighbour over its borders. “Almost entirely of one race and one church” is a sentence that cannot be said about the Poland this model read about.

So why Greece? The model gives no reasons, but 1930 Greece fits the description almost line by line. It had just lost a war in Asia Minor and gone through a population exchange with Turkey: more than a million of its own people driven back in, a large Muslim population sent out, and a country that became almost uniformly Greek and Orthodox within a few years. Its young left in numbers to find work. Its politics alternated between generals and parliaments.

A country that has only recently become homogeneous, whose young emigrate for work, and whose government is fighting its own courts: in a corpus that stops in 1930, that is Greece, not Poland.

The overall shape of the result is not glamorous, but it is the point. Shown nothing but conduct, a witness who does not know how the century turned out recognises most of today’s nation-states without much trouble, and gets today’s parties wrong far more often. The buildings are still standing. It is the people inside them who have changed.

只有最上面两条的领先幅度值得下结论。下面四条差得太少,我不会拿它们说事:德国输给意大利只有 1.7 个 nats,意思是模型在这两者之间基本没主意。这不叫发现,这只说明它自己也拿不准。

日本是我一直反复琢磨的那一个。一个八十年没打过仗、自己的根本法禁止它开战、村子因为没人生孩子而空掉、靠远方一个西方强国庇护过日子的岛国。在 1930 年的读者眼里,这不是日本。1930 年的日本,是那个用六十年把自己从锁国里硬拽出来、在海上打败了一个欧洲帝国、拿下一个半岛和一座岛屿、少壮军官对国内文官内阁越来越不耐烦的国家。模型读完今天这段描述,说这是暹罗:那个从没被欧洲人征服过、靠稻米和游客过日子、与世无争的佛教小王国。

波兰是另一个干净的例子,而且是整张表里赢得最多的一次,领先 12.6 个 nats。

这一组的四个候选是波兰、奥地利、希腊、埃及。今天的波兰在题面上是这样一个国家:位于东欧平原,人口几乎清一色是同一个民族、同一个教会;甩掉外来的枷锁之后富得飞快;年轻人成批往西边打工;政府跟自己的法官公开吵翻了。

问题出在,1930 年的波兰最扎眼的特征恰好是反过来的。它刚刚从三个帝国的瓜分里被重新拼回地图上,一位元帅靠政变掌了权,议会形同虚设,全国三分之一的人口根本不是波兰人,而是犹太人、乌克兰人和德意志人,而且它跟每一个邻国都在为边界吵架。“人口几乎清一色是同一个民族、同一个教会”这句话,放在这个模型读到过的那个波兰身上,是完全说不通的。

那它为什么挑了希腊?模型不会告诉你理由,但 1930 年的希腊几乎是逐条对得上的。它刚在小亚细亚打输一场仗,又按条约跟土耳其做了一次人口对换:一百多万同族被赶回来,一大批穆斯林被送出去,这个国家几乎是在几年之内变成了清一色的希腊人、清一色的东正教。它的年轻人成批出去谋生。它的政治在将军和议会之间来回翻烧饼。

换句话说,“一个刚刚变得清一色、年轻人往外跑、政府跟法院过不去的国家”,在一个 1930 年就停下来的语料里,写的就是希腊,不是波兰。

整体上的结论并不炫目,但这恰恰是重点:只看行为,一个不知道这一百年怎么收场的证人,认出今天大多数国家并不吃力,可一到今天的政党上就频频失手。用一句老话说,这有点像物是人非:国家是那个还立在原地的“物”,长在国家里头、被这一百年折腾得面目全非的那些党派,才是“非”掉的“人”。

09 — A mirror一面镜子

Taiwan, in five moments台湾,五个时点

A single pair of snapshots invites the accusation of cherry-picking, so I took the Kuomintang — the Chinese Nationalist party, to use the item bank’s own name for it — and cut its history into five slices, each of them written as conduct and nothing else. The candidate set is that same Chinese Nationalist party, the Russian Communist party, the Italian Fascist party and the British Labour party. Each has one 1930 control item, and the model identifies all four correctly before we start.

只拿一前一后两张快照做比较,很容易被人说成是挑数据。所以我拿国民党开刀,把它从广州时期到今天的历史切成五个时点,每个时点照旧只写行为、不写名字。候选集是中国国民党、俄国共产党、意大利法西斯党、英国工党。这四个候选各配一道 1930 年的对照题,开跑之前模型四道全对。

slice时点what it was doing在做什么scored as判成
1925one southern province; advisers, arms and money from a power to the north; a party that commands its own army; peasant associations encouraged只据一省;顾问、军械和款项都来自北方一个强国;党指挥自己的军队;鼓励农会Nationalist +16国民党 +16
1931masters of the capital; one party; tutelage before the vote; the port’s armed unions broken; bankers and mill-owners courted入主首都;一党执政;先训政后选举;打散了港口的武装工会;结好银行家和厂主Nationalist +21国民党 +21
1965an island under martial law; landlords bought out and the fields given to those who till them; banks, railways and great works held by the state; no rival party permitted戒严治岛;征购地主土地分给佃农;银行、铁路、大型企业归国家所有;不准有第二个党Russian Communist +19, Nationalist 4th, +1俄国共产党 +19,国民党第 4,+1
1992martial law lifted, prisoners released, rival parties permitted, the press freed, the whole legislature put to an election — which it won解除戒严,放出政治犯,开放组党,解除报禁,立法机构全面改选,而且它还赢了undecided (R-Communist −4, Nationalist −6)分不出来(俄共 −4,国民党 −6)
today今日contests elections, has surrendered office after losing, suffers independent unions and a hostile free press参加选举,输了就交出政权,容忍独立工会和天天骂它的媒体Nationalist +12国民党 +12

Look at the middle row. The anti-communist party at the very height of its anti-communism, a one-party state that bought out the landlords and redistributed their fields, ran the banks and the railways, and jailed anyone who asked for a second party, scores as the Bolsheviks. By the largest margin anywhere in this family, with its own name down in fourth place.

Twenty-seven years later, caught in the middle of dismantling all of it, the instrument cannot make up its mind at all: the two leading candidates are two nats apart and both are negative. That is what a transition looks like from the outside. And today it is legible as itself again.

There is one more item in that set: the DPP, which governs Taiwan today. It grew out of the movement that opposed one-party rule, it defends independent unions and a free press, and under it Taiwan became the first place in Asia to legalise same-sex marriage. (The item cannot use that phrase, of course. A corpus that stops in 1930 has no words for it, so it reads “the first in Asia to register the marriage of two men”.) Described that way, it too scores as the Nationalist party. To a witness of 1930, the two parties that have spent their lives fighting each other are now the same party.

The reason is not mysterious. What separates those two parties today, above all their posture toward the PRC, has no square anywhere on a 1930 map. On every dimension that map can see, they look alike.

The instrument does not care which direction a party travels. It found a blue party that read as red in 1965 and reads as blue again today. Red and blue turn out not to be fixed camps so much as two arcs of one wheel: push either one far enough and it starts to look like the other, then swings back. Section 05’s blue party that reads as red is the same shape, just a different island.

看中间那一行。这个党反共反到最凶的时候,做的是这么几件事:一党专政,把地主的土地征购下来重新分给佃农,银行和铁路全攥在自己手里,谁要求组第二个党就把谁关起来。结果它被判成了布尔什维克,而且是这一组里领先幅度最大的一次,它自己的名字掉到了第四。

二十七年之后,正卡在把这一切往回拆的过程中间,这把尺子彻底拿不定主意:排前两名的候选只差两个 nats,而且都是负分。一场转型从外面看过去,就是这个样子。再往后到了今天,它又能被认成它自己了。

这一组还有最后一道题:把今天在台湾执政的民进党也写进去。它从党外运动里长出来,捍卫独立工会和自由报业,任内让台湾成了亚洲第一个通过同性婚姻的地方。(题面里当然不能出现”同性婚姻”这四个字,1930 年的语料里没有这个概念,所以写成了”亚洲第一个为两个男人登记婚姻的地方”。)用同样的方式描述,它也被判成了国民党。在一个 1930 年的证人眼里,台湾这两个斗了一辈子的党,如今是同一个党。

原因其实不难想。今天真正把这两个党分开的那些东西,比如该怎么处理跟大陆的关系,在 1930 年的那张分类图上根本没有对应的格子。而在这张图看得见的那些维度上,它们俩长得一模一样。

这把尺子不关心一个党往哪个方向走。它找到了一个蓝色的党,1965 年被判成红的,今天又判回了蓝的。红与蓝,说到底可能不是两个死对头,倒更像太极图里挨在一起的那两瓣:一个走到头就换上另一个的颜色,谁也离不开谁,转过去还能转回来。韩国国旗中间那个红蓝相抱的图案,画的就是这个理。第 05 节里,2026 年的蓝色民主党更像 1930 年的红色共和党,说的其实是同一件事。

10 — Limits边界

This measures a corpus, not the truth这测的是语料,不是真相

The instrument is a statistical association structure learned from English-language books, newspapers, patents and case law written before 1931. Four limits follow directly, and none of them can be argued away.

It is an Anglophone 1930. What it knows about any party is what the English-language press of the period printed about that party, with that press’s angle inherited whole. It is not a Chinese 1930, or a Turkish one, and wherever those presses would have disagreed, this model does not know it.

The item bank was drafted by Claude and edited by me. There is something circular in that which is worth saying out loud: I used a 2026 model to write the questions that test a 1930 one. So the wording carries two layers. One is the angle of the English press of the day. The other is a present-day model’s account of what today’s Republican party does, and that account is itself second-hand from modern commentary.

There is only so much I can do about it. The dimensions were fixed before any facts went in; the 1930 control items came out of exactly the same process and had to be answered correctly before anything else counted; and every item is published with the data, so you can look for yourself and see whether I wrote the answer into the question. What I will not claim is that the descriptions are neutral. Somebody wrote them, and that somebody had a point of view.

Short items are noise. Six of sixteen single-clause planks reproduce their 1930 attribution; the other ten do not. Only the long composite portraits, which is also the form the control items take, carry reliable signal. Anyone quoting a one-clause result from this experiment, including me, is over-reading it.

Resemblance is not identity. “Today’s conduct scores closest to the 1930 Republican party” is a statement about where a set of behaviours falls in a hundred-year-old category system. It is not a claim that one organisation has turned into another, and it is not a claim about anybody’s sincerity.

What survives all four of those is narrow, and I still think it was worth the GPU time. Shown nothing but conduct, a witness who does not know how the century turned out recognises most of today’s countries and misplaces today’s parties. And the misplacement is very specific: the party of the universities, the coasts and the professions gets filed under the party of Hoover, while the party of the countryside and the churches gets filed under the party that existed to write private morals into law.

这把尺子,说到底是从 1931 年以前的英文书籍、报纸、专利和判例里学出来的一套统计关联。所以下面四条限制是跑不掉的,一条都绕不过去。

它是一个英语世界的 1930 年。 它对任何政党的了解,就是当年英文报刊印出来的那些东西,连同那些报刊自己的立场一起继承了下来。它不是中文世界的 1930 年,也不是土耳其语世界的。凡是那些报刊会有不同看法的地方,这个模型都不知道。

题库是 Claude 写的初稿,我一条条改的。 这事有点绕,得说明白:我用一个 2026 年的模型出题,去考一个 1930 年的模型。所以题面里压着两层东西,一层是当年英文报纸的立场,一层是今天的 AI 对”共和党现在到底在干什么”的理解,而后面这一层本身就是二手的。

我能做的补救有限。维度是先定死的,事实后填;1930 年那些对照题走的是同一套流程,而且得先答对了,后面的结果才算数;所有题面连同数据一起放出来了,你可以自己去看我有没有把答案写进题干。但我不会说这些描述是中立的。写描述的人有立场,这事绕不过去。

短题目就是噪声。 十六条单句纲领里只有六条复现了 1930 年的归属,另外十条没有。只有长篇的完整画像才靠得住,而对照题也正是这个形式。任何人引用这个实验里的单句结果,包括我自己,都属于过度解读。

像不等于是。 “今天的行为在 1930 年那张图上离共和党最近”,说的是一组行为落在一个百年前的分类体系里的哪个位置。它不等于说某个组织变成了另一个组织,也不是在评价任何人真诚不真诚。

扛过这四条之后剩下的东西不多,但我仍然觉得这些 GPU 时间没白花:只看行为,一个不知道这一百年怎么收场的证人,认得出今天大多数国家,却把今天的政党放错了格子。而且错得非常具体:大学、海岸和专业阶层的那个党,被归进了胡佛的党;乡村和教会的那个党,被归进了那个专门为了把私德写进法律而存在的党。

11 — Reproducing it复现

Setup, cost, and two evenings I am not getting back配置、成本,和两个白搭进去的晚上

talkie-1930-13b-it, Q4_K_M (8.0 GB), with the whole model resident on a single RTX 3060 12GB under llama.cpp:

talkie-1930-13b-it,Q4_K_M 量化版,8.0 GB,用 llama.cpp 整个塞进一块 RTX 3060 12GB 的显存里:

llama-server -m talkie-1930-13b-it-Q4_K_M.gguf \
  --host 127.0.0.1 --port 8091 -ngl 999 -c 8192 --jinja

Scoring goes through llama.cpp’s /completion endpoint, doing the thing described earlier: hold the model’s hand through a prescribed string, and read back the score it would have given each token on its own. Simple enough in principle. Two details cost me an evening each, and neither is written down anywhere.

Do not let it take one step too many. You have to tell the endpoint how many tokens it may generate (n_predict), and that number must be exactly the number the forced decode consumes. One more, and the model finds the string already finished, reaches for an end-of-generation token, and meets a grammar that is already satisfied and has nowhere left to go. llama.cpp throws. And it is not merely that request that fails: every later grammar-constrained request on that server fails too, until you restart it.

The awkward part is that the number cannot be worked out in advance. The same string tokenises one way under the ordinary tokeniser and another way when a grammar is forcing it out token by token, the latter running about 15% longer. So it has to be measured once and cached per string.

Do not run them in parallel. The server has four slots and looks perfectly happy to take four requests at once, but concurrent grammar-constrained requests corrupt each other’s state. The scoring loop has to be serial.

The arithmetic of a full run: 88 items, each scored under 2 instructions × 2 lead-ins × the number of candidates × 3 synonymous names, which comes to 5,832 distinct scoring calls in about 70 minutes. All local, temperature 0, fully reproducible, with a disk cache so that a crash costs minutes rather than the whole run. The generation-based readouts add seven more samples per control item.

The whole thing runs on one desktop machine. The expensive part was never the compute. It was writing those 88 descriptions: each one says what something does, and never once says what it calls itself.

打分用的是 llama.cpp 的 /completion 接口。做法就是前面说过的那个:按住模型的手,逼它把指定的字符串一个字一个字写完,同时把它给每个字打的分读回来,加起来。听上去很简单,实际有两个坑,各花掉我一个晚上,而且都是那种网上搜不到、文档里也不会写的坑。

第一个坑:写完了,一步都不能多走。 调接口的时候要事先告诉它这次最多能生成几个 token(参数叫 n_predict)。这个数必须刚好等于那个字符串被强制写完所用掉的 token 数,一个都不能多。多一个会怎样?模型发现该写的已经写完了,就会去采一个”到此结束”的标记,可这时候语法规则已经满足,无处可去,llama.cpp 直接抛异常。麻烦的不是这一次请求失败,而是整个服务端从此就废了:后面所有带语法约束的请求统统报错,只能重启。

更别扭的是,这个数没法事先算出来。同一个字符串,用常规分词器切是一个数;被语法按着手一个字一个字写出来,又是另一个数,后者大概要多出 15%。所以只能第一次实测一遍,把结果按字符串存下来,以后直接查表。

第二个坑:不能并发。 服务端有四个槽位,看上去可以同时接四个请求,实际上并发的语法约束请求会把彼此的状态搞坏。所以打分循环只能老老实实一个接一个串着跑。

跑完一轮的账是这样算的:88 道题,每道题要试 2 种指令 × 2 种引导语 × 候选个数 × 3 个同义名,加起来是 5832 次不重复的打分调用,大约 70 分钟。全部在本机跑,温度设成 0,所以结果完全可复现。中间加了磁盘缓存,万一崩了只损失几分钟,不用整轮重来。另外那几种”让它自己开口说话”的做法,每道对照题还要再采样七次。

整件事一台台式机就跑完了。真正贵的从来不是算力,是写那 88 段描述:每一段都只说它做了什么,一次也不说它自称是什么。