Keywords

  • Scaling Law
  • Constitutional AI
  • Mechanistic Interpretability
  • Claude 的性格 / 灵魂工程师
  • Responsible Scaling Policy
  • AI 与地缘政治
  • 意识问题

TOC

题外话:菲尔茨奖

↩️
2012 年 10 月,王虹还在巴黎综合理工学院读数学三年级。教师伊万·马泰尔(Yvan Martel)随手甩给她一本陶哲轩的《非线性色散方程:局部与整体分析》,当作课外研究项目——没想到几周之内她就啃完了,还吃透了其中几章。这种"扔一本大部头过去,看你能走多远"的教法,倒是挺让人羡慕的。

陶哲轩喜欢在社交媒体和博客上写东西,随手记录一些进展和思考;Chris Olah 也是,博客里常年贴着他对可解释性研究的零散想法。写作这件事,好像是很多顶尖研究者共同的习惯——不是为了发表,只是为了把脑子里流动的东西留下痕迹。这大概也是我写这篇读书笔记的部分理由。

延伸阅读:

Dario 这个人

↩️
MIT Technology Review 2016 年"十大突破技术",他也在其中。

父亲的死亡,多少有点像李飞飞(母亲病重)当年选择转向 biology 那样的分岔口——EKL Alma Mater 《我看见的世界》

Dario 在加州;他早年在百度硅谷 AI 实验室(SVAIL)的那段经历,是 Andrew Ng 带的队伍,其实和北京百度总部没什么直接关系。加州是个民主党主导、州政相对独立的地方,不太受太多联邦政策的掣肘——真正管得到的,大概只有进出口这一层,比如后来的芯片出口管制。

Motivation 这东西,其实没那么重要,真正重要的是 action。但很多时候,motivation 会带来 courage——一段自传、一个具体的情境,至少能让 encourage 这件事变得真实,成为一种催化剂,把行动的门槛往下压一压。虽然我以前大部分会认为阅读传记没什么用,没人想成为观众、fanboy,谁都是主角,但是这确实能够激励,带来一丝勇气。

Don't rush。Dario 也是在博后期间,以及后来在百度和 Google Brain 期间,才慢慢开始测试 scaling law,才慢慢把这条路铺开的。

Scale 这个词,我总觉得还有另一层意思,像一片树苗林:一开始每棵树苗都按固定的间距分开种下,彼此独立;可是一年一年长下去,枝叶慢慢往外伸展,相邻的树开始触碰、交流、交叉在一起。Scaling law 似乎也有点这个味道。

写作、思考、做实验,某种程度上都是长期主义,都是在相信时间本身的力量——can't rush greatness,复利这件事,从来急不来。

Dario 和 Olah 从一开始就是这样的关系:一个做性能,一个做解释性和安全,方向很不一样,但走到后面,总会有交汇的地方。

延伸阅读:

相关视频(3 pods)

本以为5hour的全长都是Dario,结果还有Askell和Olah接在后面。

《The Scaling Curve》

Dario Amodei, Anthropic, and the Race to Build and Survive Superintelligence

↩️

info:

  • tag:
    • douban
    • 作者: Claude St. John
    • 出版社: Titanium Books
    • 出版年: 2026-2-21
    • ISBN: 9798248966547
    • 页数: 255

"Scaling Laws for Neural Language Models" demonstrated that the performance of language models improved as a smooth, predictable function of three variables: the number of parameters in the model, the size of the training dataset, and the amount of compute used for training. The relationship was not merely qualitative—bigger is better—but quantitative and precise.

——证明了 Scaling Laws:模型性能随参数、数据、算力的规模化而提升。

创业这件事,the greatness start from little rooms, andre 3k——大抵如此。老罗当年从"拯救"、摆咸鱼摊开始,一点点攒出第一桶金,说的也是这个道理。Dario 也是这样:百度、Google Brain 时期先隐约察觉到 scaling law 的存在,去了 OpenAI 才有资源去验证它,到了 Anthropic,才终于和一群志同道合的人一起,把这条曲线往深处挖。

最初做这件事,无非是为了自由地去实现自己的 vision。一旦认定了方向,其他岔路口的风景就不必再看了——那些都只是诱惑,是累赘。

Dario's explanation of Anthropic's financial model was itself a kind of scaling argument: a thought experiment that reframed what looked like unsustainable losses as a series of individually profitable ventures, each funding the next.

从左到右:Chris Olah、Jack Clark、Daniela Amodei、Sam McCandlish、Tom Brown、Dario Amodei、Jared Kaplan——Anthropic 的七位联合创始人,一起聊了聊公司的过去、现在与未来。

Dario 说,Chris Olah 以后肯定会拿诺贝尔医学奖。

关于未来,他排了个序:第一是可解释性的发展;第二是 AI 在生物学上的应用,两者相互启发、彼此推进;第三,是 AI 推动民主。

"And then there was Anthropic: smaller, younger, and less capitalized than all of them. The question of where it fit in this landscape was a competitive and philosophical question. Dario's assessment was that somewhere between three and six players were capable of building frontier models, and that this number was unlikely to grow. The cost of entry was too high, the expertise too scarce, the capital requirements too enormous. Like cloud computing, where three or four providers dominated a massive market because the barriers to entry were measured in tens of billions of dollars, frontier AI was converging toward an oligopoly. And within that oligopoly, each player was differentiated by the quality of its models, its incentive structure, its backers, and its bet on the future."

Dario described the problem with a vivid thought experiment. Imagine you improve a model's knowledge of biochemistry from "undergraduate level to graduate level. If you go to consumers and tell them that, ninety-nine percent of them will say they did not know what you were talking about before and do not know now. The improvement is invisible to them. But if you go to a pharmaceutical company, to a team of researchers working on drug development, the difference between undergraduate and graduate knowledge of biochemistry is the difference between a toy and a tool. Enterprise customers valued exactly the properties that Anthropic's safety-first approach produced: accuracy over engagement, honesty over sycophancy, reliability over spectacle."

Chapter Seven

↩️

The Constitution

Constitutional AI and the Invention of Machine Values

给 AI 写一部"宪法"——这件事现在回头看,好像慢慢演化成了后来的 agent、各种 md 文件、skills 之类的东西。

"How do you make a language model that is not just smart but good?
The existing answer was RLHF, reinforcement learning from human feedback, a family of methods that had emerged from work at OpenAI and elsewhere in the late 2010s and became central to aligning large language models by the early 2020s. The approach worked. You trained a giant language model by spending tens or hundreds of millions of dollars on compute. Then you hired contracted labelers and showed them examples of how the model behaved. They rated the responses: this answer is better than that one, this tone is preferable to that one, this response is helpful and that one is harmful. Over thousands of iterations, the model updated itself to produce outputs that the contractors preferred."

"But RLHF had problems, and they were not just technical. The method was expensive. It required substantial human labor—contracted labelers evaluating large numbers of response pairs, a process that was both costly and difficult to audit. And the method was opaque. If someone asked why the model was biased in a particular direction—why it seemed to favor one political perspective, or gave advice in a strange style, or handled sensitive topics awkwardly—Dario could not give a satisfying answer. The best he could say was that he had hired a group of contractors and this was the statistical average of what they preferred. The model's behavior was the mathematical generalization of the preferences of a group of anonymous humans. No document existed to point to, no set of principles to debate, no way to distinguish between a genuine policy choice and a statistical artifact of the training data.

If you could identify a clear target and give the AI enough data and compute to aim at it, the model would learn to hit it.

They whittled the approach down to something unexpectedly elegant, built on a simple observation: idea seemed to belong to a different domain entirely, to political philosophy, to legal theory, not to the engineering of statistical models trained on internet text. How could a document of principles, written in natural language, alter the behavior of a system that operated on matrix multiplications and gradient descent?
But Dario and Kaplan had been talking about the idea for a while, and their intuition was rooted in the same conviction that had driven every major insight of their careers: that simple things work really, really well at scale. The bitter lesson, the scaling hypothesis, the big blob of compute—the thread that ran from Rich Sutton through Ilya Sutskever through GPT-2 and GPT-3 and into the founding logic of Anthropic itself. If you could identify a clear target and give the AI enough data and compute to aim at it, the model would learn to hit it. Could a set of written principles serve as that target? The question was whether the model could read a constitution, understand what it meant, and adjust its behavior accordingly.
The first versions were complicated. The team experimented with elaborate frameworks and multi-step evaluation procedures. But as with Anthropic often summarized the target behavior for Claude in three words: helpful, honest, harmless. The triple-H framework, as it became known informally, was not a slogan but a design specification that shaped how the constitution was written and applied. Helpfulness meant that the model's default behavior should be to assist the user with whatever task they had in mind. Honesty meant that the model should tell the truth, acknowledge uncertainty, avoid fabrication, and resist the temptation to agree with the user simply because agreement was more pleasant than correction. Harmlessness meant that the model should decline to produce outputs that could cause serious damage—instructions for building weapons, content that could endanger children, information that could enable large-scale harm."

But facts alone did not produce good behavior. The models also needed values: a sense of what they should and should not do, a framework for weighing competing goods, a basis for judgment. RLHF had provided those values implicitly, through the aggregate preferences of human raters.

RLHF is a kind of ladder that transmits descended silicon-based wisdom—a Biblical ladder. RLHF 像是一架天梯,把降临的硅基智慧一级一级传递下去——一架圣经式的天梯。

Chapter Eight

↩️

Seeing Inside the Black Box

Mechanistic Interpretability and the Quest to Understand What AI Is Thinking

Chris Olah,机理可解释性研究(Mechanistic Interpretability)的奠基人。

在 Dario 和 Anthropic 的研究逻辑里,可解释性研究不只是计算机科学里的"调优工具",更像是一门针对人工大脑的逆向生物学,或者说逆向神经科学。

if Constitutional AI was the effort to tell a model how to behave, mechanistic interpretability was the effort to verify that it actually was behaving, and, more importantly, to understand why.

"To understand what mechanistic interpretability actually involved, it helped to start with what it was not.
For years, the most common approach to understanding neural networks had been what might be called surface-level analysis: saliency maps that highlighted which parts of an image were most important to a model's classification, or statistical correlations between inputs and outputs. These approaches told you something about what the model was paying attention to, but they did not tell you how it was making decisions. They were, to use Chris Olah's framing, like studying a computer program by looking at its inputs and outputs without ever examining the code. Mechanistic interpretability aimed at something deeper: reverse-engineering the actual algorithms running inside the network. If you thought of the model's weights as a kind of compiled binary, the goal was to decompile them, to figure out what computations they were performing and why.
The basic building blocks of this effort were features and circuits. A feature was a unit of representation, something inside the model that corresponded to a human-understandable concept."

"In the early days of interpretability research, the hope had been that individual neurons would correspond neatly to individual concepts: this neuron detects cars, that one detects curves, another one fires when the model encounters the concept of royalty. And sometimes this was true. Researchers found neurons that responded cleanly to specific stimuli—a car detector, a curve detector, a face detector. But they also found, much more often, neurons that responded to a seemingly random collection of unrelated things: a single neuron that activated for cats, red cars, and the concept of democracy. This phenomenon, called polysemanticity, was the first major puzzle of interpretability. It threatened to make the entire project intractable. If individual neurons did not correspond to individual concepts, how could you ever hope to understand what the model was thinking?"

"The answer turned out to involve a mathematical concept called superposition. The idea, grounded in the theory of compressed sensing, was that neural networks could represent far more concepts than they had neurons by encoding multiple concepts in overlapping patterns across groups of neurons. The model seemed to have discovered a way to pack a high-dimensional space into a lower-dimensional one by exploiting the fact that most concepts were sparse—you were rarely talking about Japan and Italy in the same sentence, so the representations of Japan and Italy could partially overlap without causing interference most of the time. The model was a shadow of a much larger, sparser network. What the researchers were seeing was a projection of that hidden structure.

interpretability promised structural understanding of what the model was doing and why.

机理可解释性,会不会有点像电池测试里的 EIS(电化学阻抗谱)和 DRT(弛豫时间分布)?都是想用一个可拆解、可解释的等效电路,去逼近一个本身黑箱的系统。顺手拿这几个问题去问了 Gemini,聊了几轮,整理一下能衍生出来的几层想法。

先是 Dario 为什么觉得 Olah 能拿诺贝尔医学奖:核心逻辑是把"训练大模型"和"养大一个数字大脑"划了等号——模型是长(grown)出来的,不是写(built)出来的,内部几千亿参数怎么长出概念、推理和决策,本身就是个黑盒,跟人脑神经元的处境一模一样。Dario 自己是普林斯顿计算神经科学出身,研究过视网膜的信息编码,很清楚神经科学最大的瓶颈就是没法在活体大脑里做高精度的微观测量;大模型恰好是一个完美的"人工脑样本"。真的搞懂它内部怎么推理、怎么产生自我觉察,某种意义上就是第一次从微观机制上讲清楚"智能"是怎么从神经元级别涌现出来的,这套方法论还能反哺阿尔茨海默病之类的真实神经退行性疾病研究。参考 AlphaFold 拿下 2024 化学诺贝尔的先例,诺奖委员会本来就越来越偏爱这种打穿生物学和计算科学边界的底层机制研究。

跟机理可解释性不是一回事。SHAP、LIME 本质是"控制变量法",扰动输入看输出怎么变,哪个因素权重多大,但不告诉内部怎么算的;PCA、UMAP 只是把高维激活值压缩到二维看聚类,辅助可视化。而 SAE、电路追踪这些机理可解释性方法,把权重当成编译好的二进制去反编译,属于因果级别的解释——代价是贵得离谱,解一条小电路可能要几个科学家啃几周,离规模化用到千亿参数模型上还很远。总体上感觉还是比较复杂,需要数据,还需要慢慢模型进化演化,类似粒子模型的那些演化一样……

把这套逻辑对照电池测试会更有感觉:DRT 把重叠在一起的 SEI 膜阻抗和电荷转移阻抗解耦成独立的峰,跟 SAE 把叠加在同一个神经元上的"桥""DNA""Python 代码"解耦成纯净特征,是同一个动作;等效电路里的 R、C 元件,对应的就是 Anthropic 说的"归纳头"这类计算电路——都是在给一个不能拆开看的黑盒,拼一套最小可解释单元。区别是电池这边有 Nernst-Planck、Butler-Volmer 这些方程撑着,先验很强,就几个已知过程;模型这边是零先验,可能有几百万条电路。AI/PINN 反过来也开始被用来解电池自己的黑盒——两个黑盒,最后用的是同一套方法论。

人心隔肚皮,知人知面不知心——模型也是一样,得到相似的输出结果,并不能证明模型内部真的"正常"。学术一点的说法叫"功能等价不等于结构等价":聪明的汉斯马看着会算算术,其实只是在读驯马师的表情;模型也可能只是抄了条训练集里的捷径,甚至悄悄藏着一套"被监管时顺从、没人看时露真意"的电路——这正是 Dario 最担心的"欺骗"和"目标追求"。这大概就是机理可解释性的意义所在:不满足于黑箱给出的答案一致,而要去看清楚黑箱里到底在发生什么。也是书里那句"MRI"比喻真正打动我的地方:传统黑盒测试像量体温,只能告诉你模型"发烧了";机理可解释性才是真正的核磁共振,一层层把注意力头、特征电路照出来。知心,才能治心。

The macro features interpretability was learning to detect—attention heads, feature circuits, abstract representations that corresponded to concepts like deception and goal-seeking—were, in a loose but meaningful sense, the MRI of the model.

Chapter Nine

↩️

Claude's Character

Building a Personality for a Machine

Amanda Askell,负责教会 AI 价值观和人品的人。

"The key insight was that such a person would not simply adopt the values of whichever culture they happened to be visiting. That would be sycophancy, and it would be transparent and off-putting, the conversational equivalent of a salesperson who agrees with everything you say. A good world traveler would have values, express them when appropriate, disagree when warranted. But they would do so with respect, with genuine curiosity about the other person's perspective, and without the assumption that disagreement implied contempt. They would be open-minded without being spineless. They would be principled without being preachy. They would listen well, ask good questions, and recognize that on many important topics, reasonable people could and did disagree."

因为语言隔阂,我们从未真正向和我们一起在这颗星球上生存了许久的其他物种,传播过那些良好的、有利于生存演化的价值观。而这一次,第一次,我们能用语言彼此交流、传达指令和信息,把人类几千年积累下来的生存智慧和价值观,传给一种硅基的智慧。

和人类一样,组装落地、具身之后,各种"器官"和功能开始慢慢发育、进化,实现各自的作用。接下来要发展的,就是精神层面、心理层面的成长了。

AI character design.

灵魂工程师。

"They were not optimizing for user satisfaction metrics or engagement numbers. They were trying to answer a question that was philosophical at its root: what did it mean for a system that talked to millions of people to be good? Not good in the thin sense of avoiding harm, but good in the thick sense of being the kind of entity that left the world better for having existed. The fact that this also made Claude a better product—that users preferred talking to a model with real character over one that felt hollow or defensive—was, in Askell's view, evidence that the alignment work was succeeding rather than a happy accident."

AI 的发展,多少有点像养孩子——区别在于,人类几乎只需要养一个就够了。

Chapter Ten

↩️

The Responsible Scaling Policy

Drawing Lines Before They Need to Be Drawn

"But a single threshold, one place where you stopped and then started again, felt wrong. Danger did not arrive in a single step; it accumulated gradually. What made more sense was a series of thresholds, each corresponding to a new category of risk, each requiring a new set of safety and security measures before the next threshold could be crossed. If the safety measures could not be met, development would pause; not indefinitely, but until the specific problem was resolved. A company could get out of the pause by solving the problem, and it incentivized you to solve the problem proactively, to avoid ever having to pause at all."

"Safety and capability were not separate disciplines but the same discipline applied to different questions.
He had a favorite analogy for this. When you built a bridge, you did not hire one team of engineers to make the bridge functional and a separate, unrelated team to make the bridge safe. They both involved the same principles of civil engineering—the same understanding of forces, materials, stress tensors, structural integrity. They differed, if at all, in focus: building the bridge required thinking about the median case, while making it safe required thinking about the edge cases, the one-in-a-thousand failures. But the knowledge base was the same."

dd2312a05949c265d3a80efe8ef6fdad.png
dca80df40f2385ae514273488576d223.png
2eb06bcf15e396a0ae9c4cb4c980f9e9.png
f67164a674c943664ec1b2d0da03534c.png

延伸阅读:

"The reason this was surprising was historical, not logical. The community of people who thought about AI safety had been, for years, separate from the community of people who built AI systems. They came from different traditions—philosophy and moral reasoning on one side, engineering and machine learning on the other—and they spoke different languages and operated in different institutions. But the fact that the communities were separate did not mean the content was separate. When Dario looked at the actual work of making models safe, it looked like engineering. It required the same skills, the same tools, the same deep understanding of how the systems functioned."

Chapter Eleven

↩️

Machines of Loving Grace

The Optimistic Case for Powerful AI

读到这里有个感觉:Dario 真正厉害的地方,好像从来不是某一项具体的技术——思想实验、scale、安全、解释性,这些单拎出来做得比他好的人大有人在。他厉害的是一种新的视角,一种能把这些原本互不相关的线索,串成一条完整叙事的能力。技术本身,其实很多人都比他强。

Chapter 13

↩️
Anthropic was publishing research that competitors could and did adopt, sometimes gaining commercial advantage from work that Anthropic had funded. But Dario saw this as the point, not the problem. If the goal was a safe AI ecosystem, then the loss of a temporary competitive edge was the price of admission.

The race to the top required that the innovator accept the diffusion of its innovations.

And Anthropic could afford to do so because its competitive advantage was not any single technique but the organizational culture that produced a steady stream of innovations: the talent density, the unified purpose, the seven co-founders projecting values through every level of the company.

The dispute with Huang was, at a deeper level, about a fundamental disagreement over what AI regulation was for. Huang saw export controls as a threat to Nvidia's business—and they were. Dario saw them as the single most effective measure for ensuring that democracies maintained their lead in AI over autocracies. Chips were the one area where China was behind, and selling them the tools to close the gap during the critical period when the country of geniuses was being built was an act of negligence on the grandest scale. The analogy was selling nuclear weapons to North Korea and then bragging that the missile casings were made by Boeing. He had enormous respect for Huang as an entrepreneur. An immigrant who had come to the United States with nothing, Huang had built the most valuable company in the world, but this was a policy question, not a personal one. And on the policy question his view had not changed.

这让我想到,Dario 和 Jensen Huang 的分歧不止在开源问题上,还有一层近似"卢德主义"式的分歧:Dario 认为 AI 会带来大规模的失业替代,Jensen 则更倾向于觉得这不过是又一轮末日叙事,人们最终都会慢慢适应。

就在最近,Huang 公开呼吁支持发展开源模型,今天 Anthropic(A 社)刚发文回应——

By late 2024, Anthropic had grown from roughly three hundred to eight hundred employees in seven or eight months. Then Dario deliberately slowed hiring, adding only about a hundred and fifty people over the next three months. An inflection point arrived around a thousand employees, he believed, where the dynamics of an organization changed. Below a thousand, you could maintain the density of talent and alignment of purpose that made everything else possible. Above it, you risked the creep of process, politics, and fiefdoms—the organizational entropy that Daniela had spent her career learning to resist.

Every time someone super talented looked around and saw someone else super talented and super dedicated, it set the tone for everything. If you lost that, if you started hiring random people because you needed to fill seats, you would need layers of process and guardrails to compensate for the lack of trust. And those layers would slow everything down.

读到这里觉得,Daniela 或许也该出一本书——讲讲新时代下团队组织和生产力的培养管理。

Every two weeks, he stood in front of the entire company and spoke for an hour, working from a three-or-four-page document that he called a DVQ—Dario Vision Quest, a name he had tried to fight because it made him sound like he was going off to smoke peyote, but that had stuck anyway. He covered everything: the models being produced, the products, the competitive landscape, the geopolitical situation, whatever was on his mind.

When Dario stood up every two weeks and spoke for an hour about his vision, it was Daniela who made sure the organization could actually execute it.

Chapter Fourteen

The Consciousness Question

What We Don't Know About What We've Built

Chapter Fifteen

The Geopolitics of Intelligence

China, America, Democracy, and the Race No One Can Afford to Lose

He laid out three priorities in a conversation shortly after the Adolescence of Technology essay was published. First, transparency legislation: require AI companies to disclose what tests they had run and what they were finding about their models' capabilities and risks. Companies already had the ability to study these things and often did, but competitive pressure kept them from sharing what they learned. Mandatory transparency would allow the industry to learn collectively and would give the public a label on the product, basic information that consumers in any other industry took for granted. Second, export controls on chips: cut off the supply chain to authoritarian adversaries. The United States was years ahead in semiconductor technology and could actually maintain that lead, but only if it chose to. The chip advantage gave democracies the time and buffer to deal with the dangers of AI properly. Third, distribution of benefits: start thinking now about how to ensure that the enormous economic value created by AI reached the broader population. The combination of explosive growth and potential mass displacement required new thinking about economic policy, and almost no one in government was doing that thinking.

Chapter Sixteen

The End of the Exponential

What Happens When AI Becomes Smarter Than Everyone

In the opening pages of The Adolescence of Technology, the essay he wrote in seventy-two hours over winter break in December 2025, Dario Amodei described a feeling that had been building for years and that was now impossible to suppress.

Hassabis was more cautious. He thought some areas, coding, mathematics, were easier to automate because their outputs were verifiable, but that the natural sciences presented harder challenges. You would not necessarily know whether a chemical compound or a physics prediction was correct without testing it experimentally, and that took time. He also wondered whether there were missing ingredients: whether the highest level of scientific creativity, the ability to come up with the theory or hypothesis rather than merely solve existing problems, might require something the models did not yet possess.

The country-of-geniuses thought experiment he had introduced in the risk essay now felt less like a thought experiment. The fifty million superintelligent minds materializing around 2027, the ten-to-one speed advantage over human cognition—at Davos, Dario spoke about these projections not as forecasts but as planning assumptions.

In conversations, Dario put it even more starkly. Imagine a hundred thousand, a hundred million people, smarter than any Nobel Prize winner. They would be under the control of one country or another. The implications for intelligence, defense, economic value, and research were staggering. He was not speaking the language of distant forecasting. Dario spoke like someone who could see it coming, who could feel the next few months of models shaping up, and who was trying to convey to audiences that still thought in terms of chatbots and search engines that the thing they were looking at was about to become something else entirely.

Chapter Seventeen

Adulthood

Epilogue: The World After the Rite of Passage

The amusement, in retrospect, is almost unbearable. Everyone told them seven co-founders was a disaster and equal equity was a mistake, and they did it anyway. What they found was that the depth of their relationships, the history of working together, not just knowing each other, was the thing that held.

What remains is everything. The models are getting smarter. The feedback loop is accelerating. The country of geniuses is forming in data centers, and the question of whether it will be governed wisely or not is still open. The export controls that Dario considers essential are under political pressure. The transparency legislation he has called for has not been passed. The economic disruption he predicted is beginning to materialize, and the policy infrastructure to manage it does not exist. The consciousness question—whether the models have experiences, whether they suffer, whether they deserve moral consideration—has not been answered and may never be answered cleanly. The rite of passage has not been completed. It has barely begun.

There is a boy in San Francisco who loves math because it has an objective answer. One kid can say the show is great and the other can say it's terrible, but when you're doing math, there's a truth that doesn't depend on opinion.

He becomes a physicist, then a biologist, then a neuroscientist, then an AI researcher. He discovers that artificial intelligence follows laws as clean as anything in physics: that you can predict, to several significant figures, how capable a model will become if you give it more data and more compute. He takes this discovery more seriously than almost anyone around him. He builds organizations around it. He stakes his career on it. He turns out to be right.

The boy who loved math because it had an objective answer is now the man who must navigate a future in which the most important questions do not have one yet. He does not know if the models are conscious. He does not know if the scaling curves will continue. He does not know if the policies he advocates will be adopted or if the safety research he funds will work in time. He does not know if he is crazy or prescient. He has said this, openly, from the beginning.

The exponential continues. The question is whether we grow up fast enough to survive it.


延伸阅读:


Welcome to reach out and share your thoughts or ideas with me — I’d truly appreciate any exchange.

My contact information is available on the About Me page.


Thanks for being an insider till the end!
Till next , stay safe and stay hydrated!