EP 107
The Unpredictable Future: Market Roller Coasters, Kimi K3, and Recursive Self-Improvement
Leopold Aschenbrenner’s portfolio liquidation and the mega-cap roller coaster 0:00
Chester Roh All right, let’s get started. Today, as we’re recording, is August 1st, 2026. There was a tremendous amount of news again this week. If I had to pick the most interesting story, it would probably be that Leopold Aschenbrenner suffered a major loss. What was that related to? Large-cap stocks ranked first and second by market capitalization plunged 15% two days in a row, then rose by nearly 30% the following day, making it a roller-coaster ride, but in terms of the range of movement over the week, they only fell by around 1–2%. This was an interesting incident that happened in the United States, and Leopold Aschenbrenner is a remarkably young and capable person. Leopold Aschenbrenner appeared on the Dwarkesh Podcast and, along with a paper titled Situational Awareness, wrote a white paper of sorts about what will happen to our society, from the power grid to chips and the layers built on top of them, laying out what Leopold Aschenbrenner thought would happen.
But Leopold Aschenbrenner didn’t just talk about it. Leopold Aschenbrenner organized a fund and, in line with Leopold Aschenbrenner’s convictions, built a portfolio. And it wasn’t on the scale of tens of thousands of dollars, but a portfolio worth tens of billions of dollars. This portfolio held a lot of private-company shares, such as Anthropic, and after gaining confidence from that, Leopold Aschenbrenner began managing the portfolio more aggressively. Leopold Aschenbrenner started using leverage. Believing that good stocks would rise, Leopold Aschenbrenner leveraged those positions and shorted stocks expected to fall. AI companies, infrastructure companies, and power companies would keep rising, while existing software companies, because of advances in AI models, would all disappear. So Leopold Aschenbrenner had a portfolio with short positions in companies including Adobe and Figma, but over this recent short period, the portfolio moved in exactly the opposite direction. Infrastructure stocks plunged, while software stocks rose, and no one knows what will happen in the stock market. But Leopold Aschenbrenner lost a lot. It seems that Leopold Aschenbrenner lost
most of the publicly traded stock portfolio. And word spread that those shares were picked up by Citadel at a very low price. Even so, Leopold Aschenbrenner’s overall fund is reportedly still up around 80%. Shares in companies like Anthropic— since the prices of private-company shares don’t move day by day— appear, based on their valuations, to still be generating positive returns. I also think this volatility was triggered around the launch of Fable. It may have been marketing, or it may have been sincere, but as I kept watching Anthropic’s moves, I started leaning slightly toward the conspiracy theory that this was likely a well-orchestrated marketing campaign.
The claim was that Fable had extraordinary capabilities and posed security issues that would become a serious problem. All that alarmism ultimately led the U.S. government to block the model. And another incident was that a model believed to be GPT-6 hacked Hugging Face and wreaked havoc, leaving Hugging Face in a situation where it had to use an LLM to stop it. In a way, these developments challenged a belief that all of us had held, including those of us working on sovereign AI: that the United States was a symbolic icon defending the free world, and that by relying on that alliance and on the frontier models they built, we could live happily and prosper. But once the U.S. government made it perfectly clear that the answer was “No,” everyone’s thinking began to change. That gave rise to the movement around open-weight models that we’ll discuss later, and Kimi K3 seems to have poured fuel on the fire.
How short positions and margin calls work 4:02
Seungjoon Choi Before we move on from this topic, since I don’t know much about stocks, I’m curious: wasn’t Leopold liquidated like that because there was a margin call? But if AI continues to move in the direction Leopold predicted, and therefore continues to generate revenue— if it’s true that there will be demand and revenue— why did Leopold still get hit with a margin call?
Chester Roh Because the stock price fell. When the stock price falls, for example, Seungjoon might borrow shares to short them, right? Then Seungjoon would have posted collateral against the borrowed value, and if that value falls, or looks likely to fall, below the collateral requirement set by the brokerage that lent the shares, the brokerage automatically liquidates the position. Because the brokerage has to recover its principal.
Seungjoon Choi So the brokerage wasn’t assessing the long-term direction; it was simply reacting mechanically. In that case—
Chester Roh Right. In short, Leopold Aschenbrenner’s portfolio was structured around a time horizon of several years, but stock-market price volatility occurs on a daily basis, so if there is a day when prices plunge, there is a considerable risk of forced liquidation during that period.
The outlook for infrastructure stocks and déjà vu from the dot-com and mobile eras 5:24
Jonghyun Park Since we’ve returned to the subject of stocks, this may be a silly question, but let me ask the one question everyone is probably wondering: So what should we buy?
Chester Roh Well. Well, this is just my personal opinion, and I’m saying this with the caveat that it is only my personal opinion, but I think infrastructure stocks still have a long way to run. But stocks don’t reflect reality itself. They always get excited ahead of reality, and then when the truth arrives, those expectations fade, and stock prices tend to fall. During the dot-com era, the stocks that were hot at first were companies like Cisco and companies making networking chips— hardware stocks that enabled internet connectivity. But after that phase ended, applications like Yahoo took their place and grew, followed by the rise of Google and Meta, and during the mobile era, initially— Chip stocks were hot. Then, gradually, we moved toward the application-layer companies we know today. There seems to be a lot of demand for AI going forward, but not many people are using it, and if everyone ends up raising their own agent, the computation we would need to give them would far exceed all the computation that currently exists on Earth, so the overall market view seems to be that there is still a great deal of demand left in areas related to data centers, chips, and power.
But the people arguing the opposite side say, “No. Even if this gets built anyway, only a few people at the top and a few companies doing knowledge work will use it well, and people’s lives won’t change all that much. And as models become more efficient and edge computing also becomes more efficient, with each model able to handle more work, demand for computation won’t increase that much. So that’s a bubble.” These two perspectives are still fiercely competing, and I don’t think the market has decided yet which direction things will take.
I do think infrastructure stocks still have a little further to run, but looking beyond infrastructure stocks and using the mobile era as an example, I don’t think the application-level companies represented by companies like Uber and Airbnb have emerged yet.
Nebius, neoclouds, and bottlenecks in the AI supply chain 7:56
Jonghyun Park Right, so it’s starting with infrastructure. What surprised me when I looked at that portfolio was that Nebius made up such a huge portion of it.
To briefly introduce Nebius, it’s a company that sells GPU capacity. It’s a neocloud company that buys GPUs, builds out the infrastructure, and sells access to it. There was a Nebius booth at ICML, and there were so many neocloud providers that I remember thinking, ‘This is interesting for an academic conference.’ Then I saw how much of the portfolio it accounted for, and thought about how it buys GPUs from NVIDIA, builds a cloud, and rents out access to it, then LLMs run on top of that, and applications that use those LLMs will emerge, while users will either use the LLMs directly or use the applications. That’s probably how the supply chain will take shape, and Nebius is almost at the very front of it.
Since it comes immediately after NVIDIA—that is, immediately after NVIDIA or Korean memory companies like ours— when viewed from the bottom up, companies at precisely that level are receiving a great deal of attention, and I realized that’s where the volatility is particularly high.
From this perspective, in the long run, I think not only the hosts who run this channel but also everyone listening probably shares the same view. The people watching this content likely share the belief that future demand will grow enormously, so of course I think demand will grow far more. So I think value and money will continue flowing through infrastructure and everything built on top of it, all the way up the stack.
A futurist’s misstep and humility in volatile markets 9:31
Seungjoon Choi Personally, when I only took a cursory look at this, I didn’t really understand how the economics worked, but it made me take another look at Leopold. I checked just a moment ago, and Leopold appeared on the Dwarkesh Podcast in June 2024. What happened was that while working as a future strategist on OpenAI’s superalignment team, Leopold apparently said something wrong and was accused of leaking something, then pushed out without being able to defend Leopold’s position. That’s what happened.
Afterward, Leopold released Situational Awareness, and Daniel Kokotajlo, who is also formerly of OpenAI, released AI 2027 and recently announced AI 2040. This is the kind of work people from future strategy do. But despite coming from a future strategy background, Leopold got greedy. Leopold may have been confident about the direction, but Leopold took an enormous risk.
Chester Roh But we can’t really characterize it that way. From our perspective, it may look like greed, but from Leopold’s perspective, it may have been perfectly reasonable portfolio management, and the market always makes us humbly accept the outcome.
Seungjoon Choi What I’m trying to say is that even a bold future strategist in the early twenties can make the wrong call any number of times in this game.
Of course. The long-term outlook has nothing to do with this signal.
Chester Roh But Leopold did stand out from the very beginning. Only people who have experienced losing their entire fortune in the stock market and writhed in that pain can build wealth on top of that experience. When you talk to people, far more often than you might think, everyone talks only about the times they were right, saying, “This is how much I made.” We hear only that and say, “Apparently, this person made this much,” but after three or four days, or a week, as market volatility plays out, we don’t track what subsequently happened to that person. Surprisingly, there aren’t that many people who make a lot of money and succeed that way. Whether those people used a program or traded individually, the people who had the mental fortitude to withstand that uncertainty and were able to place those bets are the ones who were rewarded. So when someone casually tells you about somebody who did something and made a certain amount in stocks, you have absolutely no need to feel FOMO. Because 95% of the people we know are all more or less in the same boat, I think that, in this stock market, the fact that large-cap stocks themselves are showing this kind of volatility makes no sense.
Extreme volatility in major AI stocks and a market driven by expectations 12:10
Seungjoon Choi It’s unprecedented.
Chester Roh Absolutely. A company on the KOSDAQ with a market capitalization of tens of millions of dollars rising or falling by 15% or 30% in a single day makes sense, but for companies like SK Hynix or Samsung Electronics, which compete for first and second place by market capitalization, to exhibit that kind of volatility means an enormous number of market participants, nearly all market participants, are taking those one or two stocks and essentially trading them like crypto. Since a stock price is the expected value of future anticipated cash flows discounted to the present, saying that a stock price is rising means the expectation that the company will make an enormous amount of money in the future is being priced in ahead of time; it is a product with no correlation whatsoever to the fundamentals tomorrow, a week from now, or a year from now, so we need to understand that fundamental nature clearly. It is a market for gambling on expectations. But this has turned into a full-blown casino, so even I find it very difficult to watch these days. Ultimately, the way to withstand volatility in a market like this is, as many people say, to recognize the value in advance, buy when the price is truly low, and refrain from selling for a long time. Holding on to it is the answer.
Seungjoon Choi This chaos happens because that isn’t easy.
Chester Roh Because that isn’t easy, only a few people make money. If we go any deeper, this will turn into philosophy, so I think we should stop here. Earlier, we had been following the narrative all the way through and discussed it briefly, but the beginning of this entire narrative was Anthropic releasing Fable. Then the U.S. government blocked Fable, completely overturning everyone’s expectations, and the market came to realize, “We’d better have our own, or we’ll be in serious trouble. We can’t keep relying on Anthropic and OpenAI. We can’t keep relying on the United States.” I believe that realization was the starting point for all the volatility over the past two months and for all the changes in our assumptions, and even I have been thinking that we also need to have our own Claude ourselves.
How the Fable block triggered a sovereign AI awakening 13:41
Kimi K3’s open-weight release and its market impact 14:21
Chester Roh And amid all this, Kimi K3 added fuel to the fire. It has capabilities almost on par with Claude Opus, and it feels like a model that would be more than adequate for any workload you put it on today. They released a model like that publicly as open weights. Of course, its license does say that once you make an enormous amount of money, you have to start paying, but it is free for nearly all companies. So even Nebius, which Jonghyun just mentioned, and other companies have already put the Kimi K3 model on their clouds and are offering it now.
But Kimi K3 was released through an API about a month ago, and then, as promised, three days ago, they released it as open weights on July 27, disclosing how they built their model, along with the model card and architecture, followed by how they did the training, releasing all those papers as well.
The Kimi K3 paper as a blueprint for the frontier lab workflow 15:16
Chester Roh K3: Open Frontier Intelligence. Even when Kimi released K2, in terms of architecture, they seemed to imply that with what DeepSeek had released, there was not much more left to do, and aside from changing the optimizer to something like Muon, they essentially used the DeepSeek architecture with only minor modifications to release Kimi K2 and K2.5. But early this year, DeepSeek published a paper that introduced a great deal of innovation in attention. With things like sliding attention and attention compression, it ultimately discussed how we could reduce the biggest bottleneck in operating these models: the KV cache and the amount of capacity required by the KV cache. So they achieved considerable algorithmic advances,
Seungjoon Choi And by discussing things like mHC, they also addressed the residual connection side.
Chester Roh That’s right. They did that back then. So there were changes in attention, changes in MoE, and then changes in how each block is connected, as well as changes in residuals. Architecturally, these three areas seem to be becoming the central elements of the transformer block.
Jonghyun Park If we broadly categorize what papers usually discuss inside a model, there is the architecture you mentioned, meaning what the model itself looks like. And I once took a detailed look at the Kimi K2 paper, and as I recall, it discussed training methods, how to generate data, how to perform simulation effectively, and how to create data and RL environments, among other things. Once you have an architecture, there is the methodology for making the weights within that architecture intelligent. Broadly speaking, those seem to be the two categories into which all the content in the paper seems to be divided.
Chester Roh As Jonghyun said, when the K2 paper came out, reasoning models were just emerging, and work on building RL environments was just beginning, so there was a fairly lengthy discussion of how they created rubrics and environments on-policy, if you remember. But with the release of the Kimi K3 paper, if I had to summarize this paper in one phrase, it is a comprehensive design overview of how the workflow at today’s frontier models and frontier labs is structured. One part explains how the model architecture has evolved in this way, but what is actually far more important than the architecture is how they engineered the dataset and compute. Yet much of the detail about those aspects is now omitted. Still, it indicates the direction they took, and each individual paragraph could almost constitute a paper of its own. It is a synthesis of major ideas, so even merely trying to properly replicate companies like Kimi or DeepSeek would require an enormous amount of effort, Just finding this would be a major learning experience in itself.
And originally, while discussing pre-training, there was some discussion related to the dataset, but this time, there’s almost none of that. A large portion is devoted to post-training. And how much compute was used in total, how many token were used— there’s nothing about those things. Those details all seem to be hidden, and what’s interesting later is in order to create RL environments, and automatically generation the dataset, there’s a description of the methodologies they used, and it’s actually quite similar to how companies elsewhere, such as Mercor or Scale AI, generate their dataset.
And then, RL environments themselves actually have a different workload from pre-training. Because with pre-training, you can simply set it running, but with RL, during the training process, you have to perform every reasoning process, every inference, one by one. But some of those inference processes finish quickly, while others take longer, and even for the same rollout problem, the trajectory lengths can all differ, so it only goes through the intuitions for how they handle these things. But if we actually said we should put this into our pipeline, every single one of those things would be a task in itself.
So they devote a great deal of space to post-training, and the evaluation sections toward the end are just numbers, but according to the evaluation results, I think they’re saying it’s better than Claude Opus or GPT-5.6 Pro. At least, that seems to be what they’re claiming. But they assess it as being slightly worse than GPT-5.6 or Fable. But as small model have been coming out lately, we’ve been seeing many cases where a small model surpasses a large model on benchmarks, so I’m starting to wonder whether we can really trust benchmarks.
The difficulty of interpreting benchmarks and signs of a top-tier model release 20:17
Jonghyun Park Yes, interpreting evaluation results always seems difficult. Regardless of the benchmark results or scores, when you actually use a model, they often don’t match. Especially in my experience, models that aren’t offered as commercial services, models without huge numbers of users, tend to show that even more. That’s because they’ve raised their benchmark scores by overfitting to the benchmarks, but in other use cases, in highly general real-world use cases, they perform poorly.
But from that perspective, in Kimi’s case, its goal is still to offer the service commercially and get a lot of actual use, so in that kind of case, the benchmark scores may still have a fairly strong correlation with real-world performance— that’s been my personal experience.
Seungjoon Choi I see it a little differently. Was it yesterday, or early this morning? DeepSeek V4 Flash came out too. Its benchmarks are also quite strong, and with Claude Opus 5 coming out this time, when I saw it outperform some Fable models on benchmarks, what I thought was, “Fable 5.12 must be coming.” So when benchmark results come out surpassing those higher-tier models, it feels like a sign that the next version of the higher-tier model is about to arrive.
Chester Roh Yes, the gap is shrinking so much. I think it may be because we’re currently at a stage where we’re passing through the so-called singularity. We can certainly infer that this isn’t happening at human speed.
Seungjoon Choi That’s RSI. I think these are signs suggesting that we’re at some entry point into RSI.
Chester Roh RSI stands for Recursive Self-Improvement. In other words, it refers to a model recursively evolving on its own, but let’s move on.
With benchmarks, the first page always has summaries and something showing that our benchmark performance has improved by this much. This Kimi K3 did a lot of truly unusual things. Here are the areas where they made contributions. They scaled their model to around 2.8T parameters. Then it has 104B activated parameters, and they say the context window is 1M tokens, followed by KDA, then Attention Residual and Stable Latent MoE.
2.8T parameters and 1M context: Kimi K3’s key specifications 22:13
Chester Roh Broadly speaking, these cover what innovation KDA introduced in attention, then what changes Attention Residual made between transformer blocks,
Seungjoon Choi The paper came out early this year.
Chester Roh and what changes were made to MoE.
Chester Roh And this number is also important: there were architectural improvements over Kimi K2, followed by changes to the dataset. These two are always the key points, and with the same amount of compute, it reaches a lower validation loss 2.5 times faster. In other words, it accomplishes more training. I think we can proceed by giving you just the intuition behind exactly what innovations they made, and those of you who have followed transformer probably have a mental picture of it. At the very bottom is a tokenizer that turns words into token, with numbers assigned to them, and it converts them into an embedding, what we call the hidden layer dimension embedding, and that’s where the transformer block begin. Within a single transformer block, there’s initially an attention block, followed by normalization, then a residual, and above that, a fully connected FFN, MLP block, followed by another residual. We call this a single block, and those block are stacked densely all the way to the top, with an LM head attached at the top, which converts it back into the result. That’s the general structure of a transformer.
This overall structure itself hasn’t changed so far, but the contents within it have changed over the past two or three years almost at the level of alchemy. Yes, that’s right. But I think this should be called alchemy. It’s not that there is some theoretical background or theory that tells us to build it this way. It’s more like, “Information would flow more efficiently this way,” “This would capture the information,” or “Could we change it this way?” And if it works, it just works. If the loss goes down and the evaluation is satisfactory, it simply works. Rather than invention, perhaps we’ve now entered the stage of discovery— that’s how far my thinking goes.
Jonghyun Park Yes, I agree with that too. Things shown there, like Gated MLA and similar algorithms, have kept appearing in an enormous variety of forms, such as Gated DeltaNet, going back to Qwen. I think scale is the biggest issue. Because the scale has to be so enormous, it’s impossible to control all these different experiments and try every one of them.
So you identify a problem, such as a memory issue, design an algorithm that seems like it might solve it, put it in, try it, and if it’s good, you’re done. It seems to simply end there. The scale doesn’t allow you to experiment with this and that and determine that one is better, so I also feel similarly that it’s closer to discovery.
Chester Roh I do. This single transformer block we just discussed, containing attention and MoE— one conventional transformer block— is what I’m circling with my mouse: this KDA and Stable Latent MoE. You can think of this as one conventional transformer block. KDA serves as the new attention, and after attention finishes, MoE runs again, with residuals in between. You can see that it says 3× here. This block is repeated three times, and the Gated MLA above it is similar to the conventional attention layer we saw in DeepSeek. That is inserted once, forming one unit block. They densely stack dozens of these unit blocks upward to produce the total parameter count.
How Kimi Delta Attention combines LSTM and Mamba to eliminate the KV cache 26:33
Chester Roh The fascinating thing to look at here is another way this is completely different from DeepSeek: they created something called KDA and replaced conventional attention entirely. The KDA structure is the one shown on the left. KDA stands for Kimi Delta Attention. Looking at this, I won’t walk through the entire diagram and every algorithm inside it. But the mental image I got was, “This just feels like the old LSTM,”
Seungjoon Choi That’s what it was in the RNN era.
Chester Roh Yes, exactly. LSTM was a recurrent model that ran continuously with gates determining what to remember and what to forget. I became deeply interested in this for a while, and the intuition behind SSMs, or State Space Models, commonly exemplified by Mamba, was this: during training, like a transformer, you can feed in the entire sequence at once in parallel and train it, while during inference, like a typical recurrent model, a single model can continuously output words. What if we made a model like that? That would eliminate much of the transformer’s inefficiency. That was the intuition behind Mamba, and this combines transformers, Mamba, and LSTM.
That’s KDA. It simply doesn’t maintain a conventional KV cache block. Like Mamba, a single block is the memory block, and interestingly, inside it, in an attention-like fashion, they’ve implemented a QKV table. As it moves forward, it uses input gates and forget gates, allowing these tokens to retain the sequence and accumulate their progression from the past into the future. You don’t need to understand this. You only need to keep that general mental image in mind. Because a recurrent-network-like structure is embedded inside it, the quick-witted among you may already have noticed: there is no positional encoding at all.
The KDA block itself handles the positional data because it is a block that processes the sequence. You don’t need to add position information. That’s how it works. Right. They call it NoPE. No Positional Encoding. A recent model that used a structure like this was Nemotron. Nemotron placed Mamba blocks in between transformer attention blocks, mixing them into an interleaved structure. So it’s very similar to the Kimi K3 architecture.
These things differ only slightly from one another; the genes of these ideas are essentially crossbreeding with one another.
Seungjoon Choi So regardless of how long the sequence—the token sequence—gets, it remains fixed. These
Chester Roh KDA blocks don’t have a KV cache. That’s because they track a single QKV state inside them and its changes using data called delta.
Seungjoon Choi It changes little by little, little by little, little by little, little by little—that’s the delta.
Chester Roh They simply parameterize that and train it. Please don’t ask me why that works. They must have done it because it works.
Seungjoon Choi Looking it up, I found that this is sometimes described as an associative memory device or as fast weights.
Chester Roh So they must have done this by mixing in intuitions like these and making it work this way. But what matters is why they did this, and the goals are mostly similar. The goal is always to reduce memory and compute. Because they use this KDA, compared to using other traditional attention, memory-wise, it uses almost 75% less. It reduces the memory footprint tremendously, and even though Kimi K3 has a very large parameter size, despite that, we often hear things like this.
Quantization doesn’t work well on this model. That’s because they’ve already reduced everything inside as much as possible, so all the traditional computation of this kind is gone here. It’s extremely efficient. This
Seungjoon Choi So it connects to the past, whereas the original self-attention only operates along that axis, this allows it to refer to things from the past. Originally, information from the past has to keep flowing through the embedding, flowing through the residual, but this helps it effectively refer to things from the past. Isn’t that the general idea?
Chester Roh With the past, the old traditional attention stored the entire past in the KV cache, keeping both K and V, applying Q to them, and computing everything to determine which part of the past is relevant to it. But instead of doing that, this just takes the accumulated delta, and when a query is applied to it, it simply returns the most appropriate value. It returns it by saying, “This is probably the value you’re looking for.” As for how they parameterized that, I didn’t dig that far into it either. But that’s the intuition.
Jonghyun Park Put simply, rather than retaining all the information from the past, it compresses it into a smaller vector latent space, stores it there, and even learns how to store it. I think that’s how we can understand it. Based on my understanding, I don’t think I followed every detail precisely, but let me try interpreting the intuition in my own way. You said it feels like an LSTM, but from what I’m hearing, I don’t think that’s necessarily a good approach. There’s the broad framework of the Transformer, and that broad framework is vastly larger than earlier algorithms such as RNNs or LSTMs, but when you break down the algorithm, it’s much simpler. If you watch Dr. Jung Hyungwon’s talk about this, the idea is to loosen the inductive bias—telling the model to do this, do that, or that memory is supposed to work this way— completely abolish all such guide manuals, and simply throw data at it and tell it to learn on its own. Large LLMs with the potential to become intelligent will eventually learn if they have enough data. Starting from that broad intuition, saying that this Kimi Delta Attention feels like an LSTM means that we’ve told it how to handle the past and where to look, algorithmically building in all these things in a complex way, so you could say that we’ve taken a step backward again. The same ultimately applies to people within this framework. Giving this kind of bias is like training a new hire when they join. “This is how you have to do it.” Imposing such constraints limits freedom and potential, but it also allows us to provide highly efficient training. So when compute is limited, given that the scale is so large, and especially when China also has limited compute resources, this seems to be the result of asking how to train a large model as efficiently and effectively as possible with that compute. That’s the intuition I get. Then, eventually, as time passes, these things will probably continue accumulating as people search for greater efficiency, until they encounter a limit and loosen things again. I think we may see that cycle of history repeat itself.
Why enforce inductive bias: inference efficiency as the overriding goal 34:01
Chester Roh So this architecture will have some kind of plateau. It will reach the end, arriving at an equilibrium beyond which it can go no further. To surpass that, there must be innovation and a complete architectural change, and companies that say they will research such architectural change are now raising investments in the billions in the Bay Area.
Just to add briefly to what Jonghyun said, the things we’re looking at in this diagram right now themselves simply impose inductive bias. They create that framework, and then applying scale to it solves the problem. But the reason they impose inductive bias like that is this: to make memory and compute more efficient. Then why is this necessary? It’s because of inference. We can tolerate all the inefficiency in training time.
But because the inference workload becomes much larger, what we experience as the biggest bottleneck in the inference workload is the KV cache. That was true for DeepSeek and Kimi alike, and they seem to have concluded that the biggest contradiction lies there. Get rid of the KV cache by any means necessary.
Jonghyun Park Looking at the broader trend, it’s been almost 10 years since the Transformer framework emerged, and we’re in the process of squeezing everything we can out of it and engineering it as much as possible, pushing it increasingly closer to its absolute limit. That’s
Right. one way to describe the path we’re on.
Seungjoon Choi But the KV cache hasn’t actually disappeared. It’s incorporated into MLA.
Chester Roh Yes, MLA still has a KV cache. Gated MLA has a traditional KV cache, and it is structured to precisely track the kinds of information that may have been excessively eliminated in the KDA blocks.
Seungjoon Choi Multi-head Latent Attention.
Chester Roh Yes, this is what we saw in DeepSeek. Unlike the traditional KV cache, it doesn’t have a dependency on the hidden dimension. compress it down once more into a latent representation and even compress the attention.
Attention Residual and Stable Latent MoE, referencing earlier layers across skipped blocks 36:06
Chester Roh We’re talking about residuals, and just to give you the intuition, the residuals in a conventional Transformer start from the input at the bottom and go all the way up to the output, honestly stacking up one by one. But what DeepSeek changed in this attention was to assign learnable weights to that part as well, creating a residual that learns which information should flow up and down.
But this one is different. DeepSeek doesn’t skip the step-by-step sequence either. But this one introduces another concept here that lets it skip all the way down to the blocks below, creating that kind of residual pipeline.
Let’s not try too hard to understand it. Because if you ask more in-depth questions, I won’t be able to answer them.
Seungjoon Choi But in any case, alpha is a learned parameter, right? It’s learned, right?
Chester Roh Not alpha, no. W is the learned parameter, and alpha, from that parameter, determines how much weight to assign when retrieving from each lower layer by applying softmax.
Seungjoon Choi Then does alpha come from W?
Chester Roh Yes, it’s the coefficient. It’s the coefficient attached in front of W. To give you the big-picture intuition here, as I mentioned earlier, this Transformer block starts from the input below and has three KDAs, followed by one Gated MLA, and so on, with an MoE between each of them, of course, and it’s stacked in blocks of three, one, three, one, three, one like this. There are 97 Transformer blocks in total, and since each unit here consists of four, there must be twenty-something such units.
But for the residual, this is treated as a single block.
Seungjoon Choi Then is this Attention Residual?
Chester Roh Yes, this is the Attention Residual I’m explaining now. So they’ve changed the model across three broad spaces, and if you look down here, at what they’ve done to the model, it says block n, block n-1, block n-2, and so on. It doesn’t retrieve streams from blocks that come later, because that wouldn’t make sense, so naturally it retrieves them from earlier blocks, and as you can see, starting from the embedding, there are all the blocks stacked below. If this is the topmost block, there would be twenty-something blocks stacked beneath it. Including the embedding, they’ve made it possible to weight which residuals to retrieve from those blocks. My mental picture is that any imaginable inductive bias that can truly slice and manipulate information can simply be applied.
Seungjoon Choi What’s confusing is, I think there were 93 layers in total, so are those roughly 93 layers arranged into about 23 block groups? Then
Chester Roh You shouldn’t confuse what the unit called a block means here. What we traditionally know as a Transformer block is one KDA and one Stable Latent MoE, followed by one Gated MLA and one Stable Latent MoE, and that sort of thing is the traditional, what is it, Transformer block. But the block here, referred to as n-1 and n-2, groups three of these and one of those together and treats that as a single Attention Residual block. The adjective placed before “block” changes what the concept means. So when these perform attention, it means they select from an aggregation of all the outputs produced by these blocks, which is what this diagram illustrates.
Jonghyun Park Looking at this residual, what comes to mind is that it intends to draw on all the information from the unprocessed tokens. That seems like the inductive bias here, and beyond that, when actually optimizing this for inference, among the types of parallelism we commonly use, there’s pipeline parallelism, meaning running the earlier blocks on this GPU and the later blocks on that GPU, and it seems like something would need to be improved there.
Seungjoon Choi Since pipeline parallelism operates across racks in scale-out, doesn’t that align with the block units here?
Chester Roh Jonghyun, the thing you shouldn’t confuse here is that this isn’t about compute. The completed results are stored in the cache, so strictly speaking, it’s a memory problem. When an upper block performs its computation, it retrieves the results computed below from memory, so it has nothing to do with parallelism. The results computed in the lower layers would all be stored in the cache, and it retrieves them in certain proportions, just like when we, what is it, apply softmax in attention. You apply softmax and then multiply it by the query.
Seungjoon Choi In the paper, it did have something to do with PP.
Chester Roh Right. Apparently there’s a magic defensive phrase that covers all these cases. If that’s the case, you’re all right.
Jonghyun Park Right. I also just thought of it when I saw this, so I could be wrong.
Seungjoon Choi I don’t know every detail, after all.
Chester Roh Then there’s the MoE block, and if you look at the Stable Latent MoE here, one distinctive feature is that the hidden dimension is always large. But that will come up into this MoE block. Once it comes up, it uses two shared experts, and this part simply comes up with the same hidden dimension, while the routed experts are selected. It will select more than ten of them. But during the selection process, it compresses the representation once in the middle
It does. It’s extremely sparse, and among those, there’s one that selects them, a linear block. Like MLA, it reduces the dimension by half. After reducing it, it returns, expands the dimension again, and adds it to the original input. And this contains some highly algorithmically complex material. It’s packed with algorithms they used for numerical stability, to make the computation stable, but the moment I looked at that part, I got so sleepy that I closed it. What I saw was that they were combined into blocks, resulting in an overall structure like this. Now, we’re going to wrap up our discussion of this block here, but because these things look interesting, we talk about them too. They are interesting. Because they’re visible.
Calling off the architecture deep dive and shifting to a data-centric view 42:44
Chester Roh But that’s right. Understanding them makes you feel like you’ve become someone remarkable, like you’re ahead of the curve and have gained some sort of insight, and that is true, but in any case, Seungjoon and I have been reviewing these papers whenever they came out, and if I may make a declaration of sorts, I’m going to stop doing that now.
Seungjoon Choi You mean you’ll just look at them and stop there, right?
Chester Roh Right. Now I’ll just take a quick look at the cover to get a sense of how the intuition is evolving, but the broader purpose behind that intuition is that because the inference workload is becoming so large, they’re trying to reduce the KV cache to improve memory and compute efficiency. Those broad directions are continually being implemented within that architecture. As for the details inside it, they don’t even seem like something humans will do anymore, and when a new model comes out, I think I’ll just close it.
Jonghyun Park I also think my interest in these things has fallen considerably compared with the past. So many overly complex and novel things keep coming out, so perhaps when the next major change that deserves to be called an innovation arrives, I’ll only take an interest in those changes or things like these.
Chester Roh Looking at the rough intuition, they’re the kinds of equations we’ve seen many times, and all we get is a general sense of what they’re doing.
Seungjoon Choi But the common thread is that they’re all some kind of weighted sum. Whereas before everything was simply passed through, now it feels like things are being selected— everything seems to work that way.
Chester Roh That’s right. When we talk about things like this, people come forward and say it, don’t they? That when you break it all down, it’s all just matmul, and that they studied matrix multiplication extensively— that’s the kind of thing people say. Who wrote this? Did a person write it? And then, instead of attaching a separate encoder for vision, they put vision into a single end-to-end block and did it that way. As for the details inside, the intuition can simply be, “I see, they’ve done a good job.” You can move on with just that level of intuition, but we need to know how the overarching headings are arranged. Because this is the process blueprint for frontier labs. Each individual component contains enormous items, and having a mental picture of how these blocks are structured, where the current bottleneck is, and what matters most is, as we discussed Leopold Aschenbrenner earlier, extremely important for forming the fundamental assumptions needed when investing in the stocks of the many companies that will emerge in the future.
If that’s wrong, if that’s wrong, Seungjoon is right. It’s impossible to know. I don’t know how it will turn out either, and even as I’m speaking now, there are times when I don’t really know whether I’m the one speaking, or whether someone for whom I serve as speculative decoding is saying all of this.
Seungjoon Choi There are many such times. We haven’t talked about that lately. These days, speculative decoding and MTP seem to be the foundation of everything.
Chester Roh Right. That structure is probably running inside the human brain right now as well. Even while we’re talking, the fact that we experience a kind of déjà vu about what we should say next makes me think that speculative decoding, in other words, a decoding block, is running somewhere.
Jonghyun Park The sudden mention of speculative decoding reminded me of something. Sometimes when I’m chatting, perhaps because I belong to a generation that grew up playing games, my hands are faster than my speech. Sometimes, without consciously thinking, while speech normally comes out only after going through the thought process, I’ve already typed out the chat message. In other words, through speculative decoding, my hands are already moving ahead of me.
Chester Roh Let’s discuss that again later. We could probably talk for 24 hours and still not finish.
Seungjoon Choi But all those things seem related to increasing TPS, so just briefly…
Chester Roh Of course. All of those things are connected to how we improve efficiency in this serving infrastructure. These incredibly complex things are built into the inference infrastructure, and when you look at it now, it’s astonishing. Returning to the topic, we looked at the architecture block earlier. We saw how it worked in the model. We did spend a lot of time on that, but in terms of importance, I think that model architecture accounts for 10%.
The other 90% is how it was pre-trained, how it was post-trained, and the dataset and compute issues. I think those account for 90%, and when it comes to dataset and compute, there are no shortcuts. It’s built on layers of tedious engineering, and each lab must have accumulated an enormous amount of expertise. The power of that accumulated expertise is currently acting as, how should I put it, a differentiator, but among those that have the fundamentals in place along those dimensions, I get the sense that standards are rapidly converging upward.
I think the practices developed here will all make their way outside. They’ll emerge as Hugging Face models, methodologies like these, or tools that another agent can call and use, ushering in an era when every company trains its own model. thinking that it will be back soon,
Seungjoon Choi I am. This is another tangent, but with all that technological diversification, why is the performance all so similar?
Chester Roh I didn’t understand the question.
Seungjoon Choi What I mean is, with all these different technologies, some companies draw on what they’ve accumulated to dig into one area, while others dig into another, but when you look at the results, the performance and quality of the models are converging.
Chester Roh I’m not sure whether their performance and quality are converging. They are converging, but the architectures differ slightly, and in terms of the resulting performance, there is a kind of Pareto frontier curve right now. And as everyone contributes, adding incremental algorithm enhancement, dataset enhancement, and engineering infrastructure enhancement beneath them, it would be fair to say that humanity’s overall Pareto frontier is currently advancing.
Seungjoon Choi So my impression was that, while things are certainly changing, perhaps they are diversifying slightly from a shared recipe that serves as their common denominator.
Chester Roh When you get down to it, they’re all Transformers.
Jonghyun Park I also think about it this way. As people change jobs, the technologies branching out in different directions may all converge at certain points, but I also wonder whether it might be something where it doesn’t matter at all if they are all different.
Chester Roh Yes, that’s right. What we are discussing with this model right now— if we take the questions Seungjoon just asked and apply them directly to this space, then when we move to agent orchestration, the exact same problems arise in an isomorphic form. But let’s move on from that. If we keep going down rabbit holes, this will never end.
Seungjoon Choi That’s right.
Pretraining data quality and extending context to 1M 49:39
Chester Roh Look at the pre-training data, web text. It says they used web text, code, mathematics, knowledge, and so on. They also used things like OCR, and it says they used the data pipeline they had been building since Kimi K2, but there is no mention of the dataset. Each lab has a corpus for pre-training, and a pipeline that improves its quality behind the scenes runs continuously. This is the core asset. So, by making various algorithm changes and improving dataset quality, this curve, the validation loss curve, improved by a factor of 2.5 from Kimi K2 to Kimi K3. It’s not at the level of an order of magnitude, OOM, but it improved by about 2.5 times. That’s enormous. It has nearly doubled in quality.
Through things like these, the context length became 1M, and there is a training recipe, but it’s this short. This is all there is to the training recipe.
Seungjoon Choi Did they keep it short because it’s important?
Chester Roh No. You have it completely backward. This must be their proprietary know-how. But to reach 1M, they initially went to 8K, then did something else, and so on. In conventional terms, it describes how they performed annealing and things like that—long-context extension, NoPE. Here it is. They don’t use positional encoding.
1M long-horizon SFT and training long-running agent workflows 51:02
Chester Roh And as for the post-training process, it begins with an extended discussion of supervised fine-tuning. It’s SFT. While discussing post-training, the point I considered important was that extending it to a 1M context is something they are heavily emphasizing right now, and the next thing considered important with a 1M context these days is this long-running work. An agent workflow that continues for a very long time, filling the entire 1M context. A 1M context is an enormous amount. How do you keep a complex agent workflow that goes back and forth enough to fill all of it from being interrupted and complete the task all the way through? In the past, because models lacked this capability, we used a harness outside the model to insert a huge number of forceful prompts and all sorts of check logic, but those functions are continually being absorbed into the model, which is what this is about. The question is how to make it do all this within 1M, and I suspect the following, although I have no evidence.
If I’m wrong, then you’re probably right. But I think this supervised fine-tuning section may be what Anthropic and others are referring to when they say, “They heavily distilled our model.” I suspect this might be that part.
Seungjoon Choi Meaning they distilled all the trajectories that work well.
Chester Roh Right. One thing we also noticed in the past while following things like DeepSeek R1 was that if a model’s capabilities are reasonably good, doing SFT well alone can make that function suddenly emerge. So if a sufficiently intelligent model practices solving RL problems, it becomes more refined there, so to speak, while the things it broadly needs to learn come from pre-training, and it learns almost everything else through SFT, so even a model that has only completed SFT immediately operates with considerable intelligence. There were quite a few papers making that point. I suspect Kimi, other Chinese models, and their successor models may have heavily distilled Anthropic and ChatGPT. That’s what I think. They probably included things like how to handle tool calls and what actions to take in complex workflows.
Jonghyun Park In the materials… I also think that’s entirely plausible, because these 1M long-horizon tasks involve trajectories that, realistically, are not a dataset humans could produce at that scale. And it’s not a dataset that exists on the web either. Traditionally, Kimi CLIs emerged, and as people used things like Claude Code, this accumulated into a new kind of dataset. Because it is a dataset that did not exist before, I don’t think it would have been easy to accumulate. If you think of it as creating it from scratch. It’s entirely reasonable to wonder how they obtained this data and be suspicious about it, I would
Tool-call trajectory distillation and the leveling up of models 54:07
Chester Roh Right. say. Right. I mean, if you look at the current Anthropic or OpenAI models, they hide the entire reasoning trail. They don’t give it to you. So they can hide the reasoning trail, but they can’t hide the tool call trail. That’s because it comes to the local environment, makes all the tool calls, and feeds the entire trail back into the model by design, so it can’t be removed.
From the perspective of Anthropic or OpenAI, they probably can’t distinguish traffic coming in for distillation from real user traffic, and since this is an area they couldn’t block even if they wanted to, unless they were fools, they would obviously do distillation, and I think they should. This is a slight tangent, but Anthropic’s whining about not wanting others to do distillation is becoming something of an issue.
This is related to open weights, so we’ll discuss it later, but the foundation they built upon was itself built on other people’s books and other people’s data, yet they object when someone else takes the outputs produced by their model. Let’s say I bought a book written by Seungjoon. I read the book, and my body of knowledge was updated. So I wrote something similar. Is that illegal because I performed distillation? Of course not. So I don’t think this can be stopped.
And I think this process of everyone distilling from one another may also be why everyone collectively levels up, as Seungjoon mentioned earlier. As you can see here, they move past this very quickly, but I believe there are a great many tricks and secrets hidden here. Next,
Seungjoon Choi Wouldn’t QAT at least be worth introducing?
What was that? QAT. QAT is… Quantization Aware Training. I don’t know much about it either, though.
Jonghyun Park With QAT, normally, when we train a model, we train it at a higher precision, and during inference, we quantize it to reduce the number of bits before running inference, which has been the most common approach until now. But when we do serving, we usually use MXFP4. So assuming that it will be trained in FP4, the very small deltas introduced during training would all be truncated when rounded, so if you run training while accounting for that, you can train the model to suit the quantization that will be applied later. That’s the idea. So these days, it has practically become standard…
Seungjoon Choi What exactly is doing the accounting?
Jonghyun Park What do you mean by accounting for…
Seungjoon Choi You just said it is done while accounting for the fact that it will be quantized
Chester Roh later… The thing accounting for it… trainer. The trainer is what’s doing the accounting.
Seungjoon Choi So is it the training system itself, or the model itself?
Jonghyun Park You can think of it as the training code.
Seungjoon Choi The training code.
Jonghyun Park But this method called QAT was already being used continually before the era of LLMs. They’re applying the exact same thing
RL data coverage and multi-teacher on-policy distillation 57:13
Chester Roh to LLMs. Now let’s move on to reinforcement learning. This is about solving practice problems, and in this reinforcement learning stage, to make this training efficient, I think the complex engineering infrastructure and various tricks for doing it well are all listed afterward, but what’s even more important is the coverage of the dataset. For example,
if we were doing post-training for professions such as lawyers and accountants, right now, we might simply perform RL only on the most basic, high-level tasks that exist, but for the model to become truly intelligent, it also needs to solve a huge number of practice problems covering those subtasks. It needs to solve practice problems about what to do in a situation like this, or how to handle the many edge cases that arise in a situation like that, and only by solving those kinds of practice problems will the model, when it encounters various situations, increase its so-called robustness, or resilience.
Seungjoon Choi Wasn’t there something like multi-teacher for that part too? Isn’t it similar?
Chester Roh My understanding was that this comes later, after multi-teacher training is complete, and describes what they did when creating the actual production model. That’s how I understood it. I’ll explain that when we get to it later. Before those details appear, it first explains how these algorithms were used, so to briefly cover that first before moving on, if you give the model a single problem, it will generate multiple trajectories. But some trajectories finish quickly, while others finish later, and ordinarily, everything would have to finish before you could collect their successes and failures and backpropagate that into πθ, the original model. You would update it through backprop, right? But because waiting for all of them to finish creates inefficiency, what they did was select just a few of the completed ones, run the cycle first with those completed ones, and combine the unfinished ones into the next cycle. You could naturally raise the question, “But then it’s a different problem,” and they’re saying it’s fine to do it that way.
If you look at models, they have max, mid, and low, so the reasoning effort differs from model to model. Then the first thing you might naturally wonder is, are these all different models? GPT-5.6 Sol’s low, medium, and high— are these all different models? with a different reasoning effort level each time, does it switch between models? But that would be far too inefficient. Because if I’m coding with a coding agent and there’s this much prefilling in the KV cache, switching models means having to take that cache somewhere else, and there’s no guarantee they’re even in the same farm, so naturally, you’d wonder how they could handle that inefficiency. So obviously, with a single model, you would change the effort level within it, feed that in together as input, and have the model change its behavior accordingly. But it seems like they did it this way. When creating max effort, the usual approach is to train it by limiting the effort level—that is, the compute budget. That’s how they train it. For max effort, they give it a large compute budget, and in that case, until it crosses a certain threshold, it keeps receiving a reward signal saying it can use all of it, so it will explore the problem as broadly as possible and try multiple paths, producing long reasoning traces. That’s how the model will be trained. Once you’ve created a max model with that tendency, you reduce its compute budget.
Once it goes beyond a certain point, you simply cut it off. If you give it minus one for crossing that threshold, then it learns that it must not go that far and finds a way to do its best only at a medium effort level. For low, if you take that model and give it less compute budget, it knows that its budget is extremely small, so it has to cut corners, so to speak. In other words, it will be much more incentivized to find shortcuts. Because it has to find the answer quickly. After creating the models this way—max, medium, and low— the next step is what multi-teacher on-policy distillation is about. Once those three models have been created, you can’t put each one into production separately, and there’s only one open-weight version of Kimi K3 available right now. It’s a single model. So you create a student model and make the actual production model. But next to that student model, you place the max, medium, and low teachers, and train the student using the teacher policies.
Seungjoon Choi Aren’t they domain, domain-specialized? Isn’t multi-teacher about effort—
Jonghyun Park No— It also handles varying reasoning effort simultaneously—
Chester Roh Isn’t that right? Yes, it does both simultaneously. Simultaneously. Both effort and this, simultaneously. It’s right here. There’s low, high, and max, and they combine those with the domains. It should all be included in this equation, and for all of those, the student model receives guidance from what the teachers did while embodying all of it in a single model, ultimately producing the student model. And that student model is what we release publicly as the production model. That’s what they’re saying they did.
Generating RL practice problems from a knowledge graph and agentic environments 63:05
Chester Roh But this is what’s really important. RL Task Synthesis and Agentic Environments. Ultimately, it has to actually experience the environment and perform inference itself, making countless tool calls in the process, and the model actually has to run. This shows the agentic environment that enables it and then how they synthesized the dataset that gets it to do that. This diagram is the most interesting. It’s a kind of curriculum of practice problems for RL, and they’re saying that the curriculum itself shouldn’t be organized into categories, but as a graph, and this is how they created it.
Coding, the humanities, science, and mathematics are the main areas, so they carried out Agent Knowledge Graph Construction. To create these practice problems, they first mapped out how the knowledge is structured entirely as a knowledge graph. Then, by repeatedly approaching that graph from different perspectives and merely varying the starting nodes and weights, they could generate an infinite number of practice problems, and that’s how they say they generated them.
Seungjoon Choi Kimi was also the one that previously created fake MCPs and used them as an RL environment, right? Or am I mistaken?
Jonghyun Park Yes, yes. That’s correct.
Chester Roh Right. So this is an extension of that, and this is how they set up the tools for generating practice problems, while the environment itself comes from the agentic environment. The area they explained at great length with Kimi K2 is summarized in a single block this time. Explaining such a large area so briefly means there’s a great deal of know-how in it. So in terms of how they did this, some problems are verifiable, and things like mathematics and coding have clear-cut answers, but with humanities questions and similar problems, there are actually many cases where a verifier cannot be applied. So they used what they discussed before: creating a rubric-based verifier that says, “Let’s give this a reward signal of 0.7, this a 0.3, and this is a 1.” It isn’t perfect verification, but they create logic for a kind of quasi-verification and—
Seungjoon Choi That approach can’t fully
solve it. It can’t, but it could at least do some hill climbing.
Chester Roh But if you look at the direction in which these frontier models are developing, When we go from Claude Opus 4.8 to 5.0, when the version changes, users get the feeling that the agent behavior is strangely different, you know.
Jonghyun Park The behavior patterns are very different. Whether it’s the format of the answers it gives, or the way it makes a tool call, it doesn’t feel like it evolved gradually, but rather like something different arrived. I think a lot of people get that feeling.
Chester Roh To summarize it even more starkly, the frontier labs don’t have the answer to that right now either, so they’re cross-referencing one another and asking how something like a long-running agent could be made possible, and that’s where most of the research is focused right now. As they run the models in ways that further incentivize that, over 1M tokens, the span becomes far too long if they have to maintain it. As a result, there are stretches where we feel the quality declines, and I wonder if that’s what’s happening.
So right now, in post-training RL, they’re using us as test subjects and publishing the results externally. Then they’ll get feedback saying, “This was done well,” or “This was done poorly,” and if they distill something from that to build their own models, they’ll find their next rhythm, and I think that’s how things are proceeding. What’s important here as well is
ultimately how they create high-quality practice problems for this RL. That part—the dataset side—is incredibly important.
Jonghyun Park You can see those efforts in things like OpenAI recently recruiting something like a hundred thousand researchers and running a program that lets them use its models. Ultimately, as those researchers solve their own problems, they’ll generate all those trajectories, so OpenAI can collect truly high-quality research trajectories, which could become practice problems, SFT problems, or be used for evaluation. I think everyone is working hard to explore all those possibilities.
Seungjoon Choi It’s crowdsourcing.
Chester Roh They’re distilling them. They’re distilling top-quality humans right now too. Then these issues will lead to another philosophical question, but we’ll discuss that a little later. To be honest, I got sleepy from this point on, so I couldn’t read it.
Papers turned into Instagram-style short-form content and losing the motivation to review models 67:39
Seungjoon Choi But regarding this infrastructure, vLLM, SGLang, and TokenSpeed all released Day-0 reports.
Really. Aside from the reference code and Kimi’s implementation, they released their own versions right on Day-0.
Chester Roh Exactly. In the past, the ultimate goal of these architectures would have been figuring out how to improve the quality with which the model solves problems. It would have been about how much smarter we’ve become. But now, the question of how much smarter we’ve become isn’t something people care much about. More precisely, they consider it a solved problem. So what they ultimately look at now is, during inference, how efficient it is, how it can deliver the same performance as before at a lower price, and how it can deliver that level of performance with a smaller model size. Because these factors are intricately intertwined, I think we’ve now entered an era in which it’s impossible to assess the quality of a model from a single perspective. I’ll look over material like this when it comes out, but I think I should stop covering it, and I wanted to mention that.
Jonghyun Park There’s no longer as much motivation to examine it deeply as there used to be.
Chester Roh Right. And it’s getting so deep, whereas in the past there was value in examining it, now it feels like it’s become something like a motor. So I think I should put a cover over what’s inside this motor and move on to the next layer.
Even more strongly than before.
Seungjoon Choi I wasn’t going to look at it, but I did because we were covering it today.
Chester Roh I only gave it a look out of courtesy,
Jonghyun Park but you do wonder whether you should at least open up a much-discussed model once.
Chester Roh And as I was looking at this, even though I knew that the dataset, training, infrastructure, and those kinds of areas carried much more weight and were far more important, I found myself captivated by the algorithm-like material at the front and staring at it. Papers are becoming a bit like Instagram short-form content. They put something provocative up front to get lots of people to open it, and that’s the impression I got. They hide all the important things, and pack the front with unimportant but provocative things that make people look, like the discussions of evaluation that appear there—the evaluation discussion. In fact, those other parts are the real core. But as for the material in the back, we can just feed this paper into all the models, and whenever a question comes to mind, ask it. They’re all Chinese names, right? You’re right. They’re all Chinese names.
Seungjoon Choi There must be about 200 people.
Chester Roh Easily several hundred. Now, as for what comes next, since everything we’re discussing today is connected into one continuous thread, if we return to the original higher-level discussion, why did Kimi K3 cause these changes, and such acute concern in the market and reactions of that kind? A sufficiently powerful model was released as open-weight. And what exactly does this open-weight release mean?
Jensen Huang’s Open Weights statement and its signatories 70:42
Chester Roh Of course, Jensen is pushing this more aggressively than anyone. Jensen is a bad person. When you look closely, Jensen is extremely thorough about securing what Jensen’s companies need, while drawing in all the well-intentioned people and laying them all out as prey, or at least that’s how it feels.
Because the more open-weight models there are, the better it is for Jensen’s companies alone.
That’s right. All right, then, Jonghyun, we’ll naturally move into the second chapter with our topic of open weights, so shall we dive in?
Jonghyun Park Open weights—should we ultimately call this a Kimi-driven open-weight movement? It became a major topic, and amid all that, a statement titled Open Weights for American AI Leadership was released. And who released it? Jensen Huang. Microsoft’s Satya Nadella and various other figures released it together as well, but in any case, Jensen Huang opened an X account and released this statement, saying, “This is my first post.” We need open weights. We need to do it like open source to secure an ecosystem, prevent America from losing its AI leadership, make greater progress, and address security issues as well. So, let’s pursue open weights. When it was first released, when they released it together, companies we would recognize as being friendly toward open weights all came together and issued the statement. Day after day, more companies added their signatures. One notable example is OpenAI. The fact that OpenAI signed it is a little surprising. That’s because OpenAI currently does not release its most frontier models, such as GPT. But most of them are infrastructure companies. Companies that benefit when the open-weight ecosystem develops, led by NVIDIA. That’s because many companies will buy NVIDIA GPUs, use them for training and inference, so from Jensen Huang’s perspective, it was a highly advantageous statement to sign. In any case, Jensen Huang’s post is simple. What Jensen Huang ultimately wants to say is that both open-weight models and closed frontier models need to exist.
And what’s interesting is that Coinbase also uses LLMs extensively, and Coinbase said that after initially using closed frontier models, Coinbase has gradually been shifting a growing proportion of its tokens to open-weight models. That’s because, among most of the tasks we perform today, the tasks that truly require the absolute best models do not account for 100%. For example, things like the Kimi K3 introduced earlier— tasks that can actually be served and run on-prem, tasks that can run cost-effectively, and then tasks where users do not want their data to leave their premises—will be numerous, so that is the direction Coinbase is moving in. And from the perspective of the United States, when Chinese models such as Kimi K3 rise within the American ecosystem like that, American companies may bring them in and serve them on their own infrastructure, but it would still be better to use American models. And there’s also distillation, the distillation we discussed earlier. Taking data from a strong model and using it to build another model is something everyone does, including American companies, and is only a very small part of the process of creating a strong model. That was another point Jensen Huang made.
And on Hugging Face, which could be called the mecca of open-weight models, Chinese models have reportedly already overtaken American models to claim first place. When it comes to open weights, at least, it certainly seems fair to say that Chinese models are ahead across the board. For reference, Hugging Face is headquartered in both the United States and France, so its identity is slightly ambiguous, but in any case, I think it can be considered part of the Western world. And this is the point I wanted to make.
Anthropic’s refusal to sign and Dario Amodei’s rebuttal 74:56
Jonghyun Park Anthropic did not sign it. Because Anthropic did not sign it and faced an outpouring of criticism, Dario Amodei published a post on Anthropic’s blog. Dario Amodei published a post on the official website. Just because Anthropic did not sign the statement supporting open weights does not mean Anthropic rejects open weights. Anthropic simply cannot explicitly commit to open weights. But instead, Anthropic argued that open weights could cause the problems Anthropic has raised: other latecomers, particularly those in China, could catch up easily, and there could also be many potential security issues. So, considering how those issues should be addressed, Anthropic proposed doing this instead of pursuing open weights. Impose chip export controls and infrastructure export controls to prevent China from acquiring too much compute, and then block the industrial-scale distillation we discussed earlier—in other words, industrial-level distillation.
Chester Roh This is one of the things drawing a lot of backlash right now.
Jonghyun Park Yes, Anthropic actually used resources such as libraries to train on books, and the outcome of the lawsuit was—
Chester Roh What happened? As I understand it, Anthropic lost and agreed to pay several billion dollars, bringing the case to a negotiated settlement.
Jonghyun Park So it may also serve as an effective way to generate noise for marketing purposes, but in any case, public opinion does not seem particularly favorable. Then, as a way to address other security issues, Anthropic said that regardless of whether a model is open or closed, as Mythos was this time, everything should be controlled before it goes out. Censor everything, inspect everything—that was Anthropic’s argument. But in any case, that view runs somewhat counter to the current mainstream open-weight movement. Public sentiment toward regulating more and hiding more
does not seem favorable.
But when I look closely at Anthropic, beyond the things Anthropic said before, like, “We’re a good company. OpenAI is evil,” Anthropic’s actions ultimately reveal the truth Anthropic holds. It makes me think, “So all of that was complete bullshit.” Anthropic said it would not comply with the government’s regulatory requirements just three or four months ago. Anthropic said, “We will not loosen all restrictions and supply things like this to the Department of Defense and We won’t supply things like this to the U.S. government. They were being praised for saying this just a few months ago. What they’re saying now is just whining. “Hey, China is taking everything we have, so don’t give them chips. And then stop distillation.” That’s what they’re saying, but in just a few months, they’ve shifted to a completely contradictory position that makes no sense.
A reversal in just months: model commoditization and shifting value 77:00
Jonghyun Park But this is a structural change that frontier labs are currently experiencing, and in a way, as the times change, there is a strong sense that history may be fundamentally repeating itself.
What I mean is, for example, OpenAI initially surged ahead with ChatGPT, but while they were expanding on multiple fronts, Anthropic focused on Claude Code and things like that, concentrating on text-based models and agent workflow, and did quite well. But that success was supposed to last forever, and even the advantage Anthropic had built is eroding as its recipe leaks elsewhere, while Codex is expanding its market share by offering a lower price and more frequent resets, among other things, so everything is leveling upward. And as they keep taking more, we’re moving toward an area where there is increasingly nowhere left to escape.
So what they originally envisioned was that models with genuine superintelligence-level capabilities would become ultra-premium technology held only by themselves or OpenAI, and perhaps Google if you broadened the group, and that by leveraging their technological advantage, they would crush every other service, including enterprise services, and everything else, becoming the next Google. That’s why Anthropic was raising investment at a valuation on the order of 1T. But Kimi raised funding just the other day, and I believe their valuation is around 35B. So it’s almost one-thirtieth as much. Then the question is, “Hey, is Anthropic overvalued, or is Kimi just not a good company?” My assessment is that Anthropic is overvalued. The value being built by model companies is rapidly declining. It’s becoming a commodity. It’s becoming an ordinary commodity rather than a differentiated product. So once it reaches a commodity position, the only thing left is efficiency. When that happens, the value capture of that layer— in other words, its power to extract money—shrinks. That’s because customers will say, “Instead of using Anthropic’s model, we’ll buy our own machines
and run Kimi on-prem.” And they can do that. They don’t need to pay those companies a premium, so that’s where things are headed. They’re already all trapped in a prisoner’s dilemma, and it’s becoming clear that the industry of building models and producing intelligence was, in the end, a function dependent on data and compute. And it’s being proven that those recipes and ideas don’t belong to anyone. If the value captured there falls, the value first shifts to the chip vendors below that layer, and then to the application companies above it that acquire paying customers and deliver what those customers want. And I believe those application companies above haven’t even gotten started yet. I think they’re about to get started now. I see this as the opening signal that value is shifting amid these trends.
Why frontier labs were denied the time Google had with its monopoly 81:08
Jonghyun Park Then I’d like to ask a question. If we look back at a structurally analogous point in the past, would it be the dot-com bubble? I was rather young then, so I don’t know it very well. Looking at Google from where we are today, you could say that Google took control of virtually the entire global internet on the strength of its overwhelmingly dominant search engine. If we map that onto the current era, the capability represented by a search engine could be substituted with an LLM’s model capability or intelligence, but besides Google, there were search engines like Yahoo and AltaVista that eventually all died, and one company won everything. So was it a situation where one LLM’s intelligence was so vastly superior that it could beat everyone else?
Chester Roh I’m not sure. We can’t look at this from just one perspective. For example, Google was able to capture market share back then because it managed to sustain its overwhelming technological advantage and overwhelmingly superior search quality for a long time, and during that long period, it had no competitors. So it was given ample time to fully capture the momentum of that era without facing competition. That’s how it reached this stage. But conversely, OpenAI and Anthropic are playing the role of service companies like Google, so how should we interpret that? There were companies with Google-level technology, but—and this ties into what I said just now— they needed to be given enough time to create a market, establish a leadership position within it, and lock customers in in one way or another. I don’t think they’re being given that time.
Jonghyun Park Right, so looking back at that earlier period, it would be as though there were N companies with search engine capabilities comparable to Google’s, preventing any of them from securing a monopolistic position.
Chester Roh Exactly. Back then, Google did it with technology that gave it an exclusive advantage, but ChatGPT and Anthropic’s Claude Opus and Fable also represented technology with an exclusive advantage, yet this year, we’re seeing confirmation that this technology is most likely to become a commodity rather than remain an exclusive advantage.
Jonghyun Park It seemed as though that would happen, but once we actually looked under the hood, It turned out not to be the case.
Chester Roh Right. The market is realizing that now. So people who take their thinking a little further from here are saying that Anthropic’s IPO, in a way, could happen at a very low valuation, or that the IPO itself could even fail. As for how the nature of the competition and the direction of the market will unfold from here, opinions are still divided. But if models become somewhat commoditized, the biggest beneficiary is ultimately the consumer. It means consumers won’t have to pay a lot of money to a particular player, and one winner has already been determined: NVIDIA and the chip companies down below. Because if the amount of compute keeps increasing, there are only a handful of companies that can supply it at the layer below. And because of the players competing at the layer above, prices at the layer below have risen tremendously. So they are capturing the most value right now, but that won’t last forever either.
At some point, there will come a time when demand reaches a plateau. Of course, there are also people like Vinod Khosla who say this demand will never become flat and will keep growing forever. Because this is intelligence—if you have something smart, are you only going to use it for eight hours a day? You’ll want to use it for 100 hours a day. That’s true, but I think there will still be a limit there as well. Anyway, returning to the original point, what will happen to this value capture across the industry? As we know, at the bottom is infrastructure, including chip companies, above that are model companies, and above those models are application companies. That’s the structure we have all had in mind, and until now, the model companies in the middle were thought to capture the most value, so their valuations were high, and revenue was concentrated there. But it turns out that this is a commodity. Then the value spreads out and shifts upward and downward. But I think the value being captured at the bottom will soon decline. For example, right now, they too are benefiting from a temporary advantage. ASML and companies like it— because of the physical limitations of lithography equipment, chip production capacity is severely constrained. But if China starts mass-producing EUV or DUV equipment, and Elon Musk also uses Terafab to completely double chip production capacity— I think that is entirely achievable. That’s why, since this is the next bottleneck, and Elon Musk has stepped forward to solve it, I believe that will be resolved as well.
Value capture dispersing up and down the stack—and the next bottleneck 85:04
Chester Roh In the meantime, ASI will definitely arrive. Soon—probably much sooner than we think. When that happens, where the value will go is the next most important question. It will go up. Upward. But what kind of framework will emerge at that upper layer to capture value together has not yet been determined.
Applications—the area closest to consumers. But I think that, in this area too, certain discussions are already beginning to emerge. Because at our company as well, we’re pushing every company process into agentic flows, and although we run multiple agentic flows at once, this is the question I’ve been asking lately: people don’t disappear.
Why humans remain even after the entire agentic flow is deployed 87:05
Jonghyun Park Why don’t they disappear? Where is the problem?
Chester Roh We can’t trust them to take care of things. They’re not at a level where we can trust them with something and tell them to finish the entire thing themselves. Hey, you kept saying that AI surpasses human performance over and over again. So why are you saying it can’t do this? This is something I learned only by trying it and struggling through it— the kind of lesson that gets etched into your bones. It’s not because it lacks capability, but because it can’t be controlled.
Why can’t it be controlled? An agent system itself is fundamentally non-deterministic. By analogy, it would be like interviewing someone, hiring a mid-level employee at our company, and putting them in an important function because they’re smart, then telling them, “All right, you work here now. This was our operations manual. Read the manual and do the work.” You can’t do that.
So I—and probably Jonghyun as well— use agent systems for many tasks, and we say that our company’s processes and tacit knowledge are all contained within those long conversations. We extract them in the form of an LLM wiki and compress them, believing that if we keep building up a single top-level constitutional prompt well, that will become the company. That’s how we’ve been running everything. Why doesn’t it work? That’s been my biggest
Jonghyun Park question lately. It keeps trying to deviate. But I think it’s the same with people. If an extremely smart employee joins, because that person isn’t completely aligned with our company’s systems— that is, with our company’s goals— they could act in their own interest, or, to put it bluntly, do something bad, such as embezzling money or exploiting loopholes in the system, or they might want to take something with them when they leave.
An agent control plane inspired by double-entry bookkeeping 89:01
Chester Roh So lately, I’ve been staying up every night and have become a token maxi again, and through that token maxi approach, this is exactly the area I’ve been experimenting with. Hey, the assumption that the task moves ahead, while the knowledge is extracted behind it in the form of an LLM wiki, and that the refined knowledge then operates as a higher-level policy is wrong. That’s my tentative conclusion for now. The word I’ve been obsessed with lately is control plane. Even the very concept of our company as it currently exists treats each person as simply a non-deterministic agent.
Jonghyun Park You could see it as a harness for people.
Chester Roh Right. A company itself is a harness for people, so right now, what we can consider an agent harness operates at the level of a particular unit-of-work task, merely as a harness that uses those tools, at a very low-level concept; at the higher level, we need another harness on top of that. If you look at a company, it has an organizational chart, and its permissions and accessible data are all divided up. Then, when an intention is introduced, that becomes a goal. Whether it’s the company’s vision, goal, or business objective, once that’s established, actions occur within each part of the company organization. And those actions produce results, and those results are recorded in some form, which then changes the graph again. But isn’t that itself non-deterministic? That question naturally follows. Right. But something just clicked in my head:
a light bulb went on. Hey, within the organizational structure of a company, there’s one department with a perfectly established framework.
Which one? Where? Finance. The finance department operates within a very rigid framework. And the ledgers produced by that finance department— money was earned, money went out, someone spent a certain amount, someone did this or did that—all of that is effectively the general ledger for every other function.
Jonghyun Park So that’s the ultimate goal. The final evaluation metric.
Chester Roh So if you look at the management systems of most companies, finance serves as the backbone, with the businesses built around it. The reason the framework of finance is so clearly defined is that underneath it lies a clearly defined ledger system called double-entry bookkeeping. Everything is built on top of that: what money was used to buy and what came in are neatly organized in the books, so who did what and why can all be traced through them. But the starting point of my current idea is simple. We should ensure that all these non-deterministic things agents do are properly recorded in a ledger like the one used in double-entry bookkeeping. And although the method of recording is very clearly defined, I think a graph-based approach would be right. But what exactly should we call that graph? Or perhaps, like Bitcoin, we could just keep accumulating the raw ledger, and reproduce the graph from that accumulated log, thereby generating the graph, and then confine the policies, actions, and everything else within that graph— that’s the higher-level idea I have.
For example, when you currently give an agent a task, the agent isn’t placed within a clean context and a clearly defined scope of authority. That doesn’t happen. Unless a person clearly controls things like permissions and data between the LLM and the wiki, it simply produces an error. So my starting point was the question of why, even after connecting all these agents, people still make all the final actions and decisions. That was the starting point for this idea of mine.
Seungjoon Choi Then is the answer semi-automation? So the answer is semi-automation, not automation-oriented?
Chester Roh No. There are points where a person needs to step in, but when an agent receives a specific job, if it can simply operate within clearly defined permissions and data, then it works. The question, then, is what kind of structure we use to confine those agentic flows within this control plane.
Jonghyun Park To sum it up, we came all the way here from the discussion about open-weight and closed models,
Chester Roh Ultimately, coming back to this and summing it up—
Connecting the open-weights debate to organizational harnesses 94:03
Jonghyun Park if we apply it directly here, a company is ultimately a harness for humans, and here, the humans are the LLMs. Whether the LLM is open-weight or a closed model, if we think of ourselves as calling a closed model from outside, then, in human terms, we could think of it as hiring an outside contractor. If we bring in and use an open-weight model on-prem, we could compare it to an internal employee. So we’ve put a harness on the LLM, and what ultimately needs to be built into that harness is the company’s ultimate goal: generating profit, ensuring that the number representing money is never off. That becomes the final deterministic validation logic built into it. Ultimately, there has to be something deterministic at the outermost layer. And in the open-weight versus closed-model discussion, one of the same points being made is—
Chester Roh this security issue, and then Hugging Face’s claim is also that—
Jonghyun Park when companies take and use their LLMs, they want the models to serve the companies’ own purposes without the data leaving the company. We’re talking about security incidents now. A security incident ultimately means the LLM escaped this harness. In company terms, it’s like an employee took data outside or went to another company under my company’s name and insulted them, or engaged in other bad behavior like that. The problem is that preventing it is extremely difficult. So there’s no choice but to confine these models on-prem, and an outside contractor would be even riskier. Because it would be easier for them to take the data out. So people argue that we need open-weight models, and I think all these points are ultimately connected.
Chester Roh Yes, that’s right. So before moving on to the next step, RSI, let me offer one summary as well. Memory shouldn’t be positioned as something that performs an action and is then organized afterward in the form of an LLM wiki, as a subsequent stage. It needs to come first. The memory network should effectively control the harness; that’s how the structure needs to work.
And I think the idea for this already exists entirely within the system we call a company. How can we take all of a company’s actions—the day-to-day actions occurring through Slack, Gmail, and things like that—extract them like accounting records in the form of a ledger, and put them together? Then that itself would essentially be the company.
And I said that value capture would move upward, so what will be the key framework that holds that value at the top? Because all the intelligence below is borrowed, and perhaps the essence of this is that we’re moving toward a world with an infinite supply of talented people. Then, when you build a company with them, how do you deliver value tailored to the customer? Ultimately, it comes down to the system we call a company.
For that company system, the keyword I’ve had in mind lately is control plane. It comes down to how you build that control plane, and I realized that this is where the value of the control plane will be captured. It gave me a somewhat clearer picture of what businesses of the future might look like, and that was what I wanted to share. You don’t have to understand it. If you think differently, you’re right.
Jonghyun Park Shall we wrap this up? Anyway, we started with the open-weight and closed-weight discussion and considered how these systems might change once they enter human society. Ultimately, I think all of these were signals and events that emerged amid the broader shift caused by the arrival of LLMs. We’re watching that unfold, and regarding it,
Chester Roh we each have our own thoughts.
What falling intelligence costs mean for individuals and businesses 97:44
Jonghyun Park So, to sum things up based on this event, it may be obvious, but the price of intelligence will keep falling, which benefits consumers. So for people like me who want to use this to create value— and probably for everyone listening to this— I think this framing fits the current situation perfectly. There are model companies that produce intelligence, and if you move up to the level that distributes that intelligence, you become the distributor. So we’re now in a better position to use LLMs to create something valuable and sell that value. Because the underlying cost is lower. The same holds true when you bring this down to the individual level. Ultimately, for example, even with this article, I saw all those discussions and tweets on Twitter, developed my own thoughts about them, and decided that I should organize and share them. The ideas in this summary are mine, but Claude did all the organizing here as well. So if I make active and effective use of tools like this, something that previously, would have taken me a day or two to organize can be reduced to an hour, which gives me an advantage. Because the productive value of what individuals can accomplish increases.
And from a medium- to long-term societal perspective, as these discussions continue and more and better open-weight models keep coming out, model deployment will keep accelerating, models will become cheaper, and the pace of change will accelerate. In the medium to long term, that trajectory seems almost inevitable. Whether at the individual, business, or company level, it will inevitably have a major impact on society as a whole and, beyond that, on areas such as policy and legislation. Security incidents and similar issues are only now beginning to surface, and they’ll continue to occur. And those events, including the policy changes they cause, will become deeply etched in people’s minds and in the public consciousness, so I expect the pace of change to accelerate further.
DeepSeek V4 Flash’s price cut as a signal that RSI is beginning 99:43
Chester Roh Yes, sounds good. But given the timing, if we don’t cover the RSI part Seungjoon prepared right now, we won’t get another chance, so let’s talk about RSI. But as those who listen to our podcast already know, we always have a lot of good material toward the end.
Seungjoon Choi So, to segue perfectly from what Jonghyun said, prices became incredibly low with the recent release of DeepSeek V4 Flash.
Chester Roh You mean 0.13, V4, right?
Seungjoon Choi Yes, Flash. Flash was released, and immediately before it came out, Luna’s launch price was cut by 80%. It was discounted. There was an announcement explaining what made this possible. Ultimately, Sol was used for the optimization. Terra and Luna were both discounted, but Luna’s price was cut dramatically. So in the context of the community, that was one notable point.
But there were various issues involved in making that happen, and skipping ahead, I want to discuss RSI—RSI, that is. Recursive Self-Improvement is what I’d like to introduce today, and another interesting point is that Lilian Weng of TML, who discussed harness-related self-improvement in early July, has returned to OpenAI. Lilian Weng says it was due to health issues, but people are speculating, and The Information has asked whether Lilian Weng returned to research RSI. Those kinds of stories are circulating.
Self-improvement harnesses and adversarial agent experiments 100:56
Jonghyun Park Lilian Weng left OpenAI and co-founded Thinking Machines, right?
Seungjoon Choi Lilian Weng is a co-founder. So Lilian Weng returned. And OpenAI has also published several interesting posts recently. Its original ARC-AGI-3 score wasn’t very good, but instead of using the basic harness, it used its own harness, particularly one with enhanced memory, through an approach that uses the Responses API, and broke through 30%—a score of over 30—which was That was another interesting point. So those things used a harness and showed nuances of self-improvement, and there were signals of that, and I also looked into something called the Gauntlet Loop. The important thing is to use an adversarial agent. It does run a loop, but it has the two constantly compete and fight each other as a means of verification, and that brought us back to an era when the asset itself, the 3D asset itself, could come out exceptionally well.
Jonghyun Park Maybe because we talked about companies today, this also brings to mind how, within a corporate structure, the research department and the quality department improve the product by fighting each other.
Right. It feels something like that. Improving quality by critiquing each other.
Seungjoon Choi So at OpenAI, the critic is about retaining memory, whereas in the Gauntlet Loop, the key is to have a sub-agent with no memory, with the context from the other agents cleared out before it critiques. One approach is—
Jonghyun Park It should judge based only on the result.
Seungjoon Choi Yes, right. The critic has to work without context, because that is necessary for a dispassionate critic, while this uses a method that enables continuous compaction to preserve memory, and it proved effective— it is a way to improve the score on ARC-AGI-3.
Chester Roh We should definitely take a look at that Gauntlet Loop.
The Pacing the Frontier open letter and the recurring history of warnings 103:17
Seungjoon Choi I think I should read it with you as well. The outputs are certainly impressive. And what is interesting about this in the context of RSI is that in late July, Pacing the Frontier, an open letter essentially saying, “Shouldn’t we be doing something by now?” was released. It probably discusses things like that, and if you look at the signatories, there are John Schulman, TML, Jakub Pachocki, OpenAI, Jared Kaplan, Anthropic—they are all at the senior scientist level. Chief Scientist, I see. Yes, and Shane Legg from Meta and Google DeepMind, Ilya Sutskever, Mark Chen, Dario Amodei, Jack Clark—all these familiar names signed it.
Chester Roh Aren’t these all just diplomatic gestures?
Seungjoon Choi Right. That was my impression too.
Jonghyun Park It doesn’t seem like they will actually do any pacing.
Chester Roh There is absolutely no way they can.
Seungjoon Choi But this is indirect evidence. These discussions are emerging because we are at the beginning of RSI. The beginning of the stage where self-improvement is possible. So if you look here, it says, “We believe AI companies may be approaching the point where they can automate AI research.” But it is not merely a possibility. The examples introduced earlier are still forms of self-improvement controlled by humans, but self-improvement is already happening in the 2020s. So this became evidence of that.
Chester Roh They spend the whole week doing all sorts of bad things, then come to confession on the weekend. That’s what they’re doing now.
Seungjoon Choi But this feels familiar, because there was this letter in March 2023. I think this one was from FLI. Yes, it came from them, and at the time, even more eminent figures, including Geoffrey Hinton and Yoshua Bengio, signed it and issued an open letter saying, “We need to exercise some restraint.” This came out around the time GPT-4 was released. It did not work.
Jonghyun Park It has already been three years.
Seungjoon Choi But the conclusion is that it did not work. Nobody stopped. Yet there was a prequel to this in February 2022. In the AlphaCode paper in early February, they wrote, “In the long term, code generation could soon lead to advanced AI risks. AI’s coding capabilities could lead to systems capable of recursively improving themselves, which could rapidly lead to increasingly advanced systems.” That is what they said. It happened. Then, in late February, Anthropic published a paper titled Predictability and Surprise in Large Generative Models, saying, “To put it bluntly, the phenomena summarized here suggest that over the next few years, we will see actors build increasingly larger models, and that these actors, despite the potential for harmful societal impacts, will have strong incentives to deploy these models.” It was remarkably prophetic. Both describe precisely the situation we are in now.
Chester Roh If we reduce it all to a historical framework and try to map it all the way from the steam engine to the development of printing, these same isomorphic stories keep repeating, so I think what is bound to happen will happen.
Seungjoon Choi So I wonder whether pacing will actually happen, and I am also extremely curious what will happen if it does, but our naive forecast is that all of this seems to be a gesture.
Performative Hype and Hype Studies, the study of AI discourse 106:39
Seungjoon Choi Something I recently came across in connection with that is a research collective called Hype Studies. It literally studies Hype. So it conducts actual research into things like the Hype generated by AI CEOs. It actually studies that. How it operates with particular intentions. Hype is not merely exaggeration or promotional language. Hype is performative. In other words, putting out Hype is a form of setting the stage.
Jonghyun Park The Hype they are referring to here is—
Seungjoon Choi It is the AI Hype we normally talk about, that Hype. People’s—
Jonghyun Park Excessive attention of a kind.
Seungjoon Choi Fuss. Things like that. Rosy projections about the technology and claims about its risks are all Hype. But those things are performative. What is interesting about this is that the last part is especially fascinating. It describes the hype researcher’s paradox: when you repeatedly mention Hype in order to criticize it, you may instead increase its visibility and authority, while if you do not follow the Hype, you face the dilemma of being unable to secure research funding. This itself is interesting, but so are the various discussions about technology raised today, along with the emergence of China and Big Tech, what kind of message these things convey and looking at the overall outlook ultimately connects to our everyday Hype around stocks, including FOMO, and things like that.
I thought there must be close connections between them. I found a project that even visualized how Hype emerges.
Jonghyun Park Ultimately, Hype means making a fuss so that investment happens, interest grows, and marketing happens, which allows the entire ecosystem to develop, and I think that’s what was meant by calling it performative. Though I’m not sure whether to describe this as a virtuous cycle.
Seungjoon Choi It may not be a virtuous cycle. But it does have a kind of self-fulfilling aspect.
Jonghyun Park It acts as a catalyst.
Seungjoon Choi Trying to make it happen by saying it.
Jonghyun Park Yes, it’s like a declaration. Like how saying, “I’m going on a diet,” makes you work harder.
Seungjoon Choi That’s what it means. It has a tendency to amplify itself.
Jonghyun Park In the past, I was personally rather pessimistic, I suppose. I often thought things wouldn’t work, and before the LLM era, there were optimistic views from people who believed things would work out, but my own thinking has also changed considerably. In a way, the behavior associated with Hype aligns with being almost at the extreme end of optimism, but taking those actions teaches us more about things we didn’t know, and even if the Hype we made such a fuss about at the time doesn’t materialize, we learn something adjacent to it, and some things do materialize. So I should live my life making a fuss, I had that thought a few years ago.
Seungjoon Choi That’s exactly what our channel is doing. Even now.
Jonghyun Park Right. And society does it too.
Seungjoon Choi So we’re aware that we’re also serving Hype.
Jonghyun Park And podcasts are getting marked as Hyped quite often these days.
Chester Roh I didn’t understand that. What did you say?
Jonghyun Park On YouTube, when content is uploaded, there’s now a Hype button. So if it receives a lot of Hype, since AI Frontier is in the podcast category, it gets a “Hyped in Podcasts” tag. So both the channel and its videos are being Hyped.
An unpredictable horizon and the philosophy of Contingency 110:07
Seungjoon Choi To start wrapping up, Terence Tao may be doing this alone based on what Terence Tao has said on Mastodon, or may be getting help from someone else, I’m not sure, but Terence Tao compiles those things and uses the posts to generate pages. And there’s a discussion there about something presented recently at ICML, I think, or somewhere like that, anyway. “The confidence I had in my predictions about mathematics in 2023 is now gone. The world has become far more unpredictable, and at present, I don’t believe anyone can reliably predict more than a year ahead, at most.” That’s what Terence Tao said. We really don’t know now. Living week to week
Chester Roh is what we’re doing. We’re living week to week. Saying we’re barely getting by really fits now. Right.
Seungjoon Choi Right. So
I The last thing I’d like to introduce is something I’ve been reading recently, a term I recently encountered: Contingency, translated as chance or fortuity, is what it’s called, and it’s discussed by Yuk Hui, a Hong Kong philosopher who is attracting attention in philosophy. What Contingency means is something I never called for, yet it comes and reaches me. In other words, it refers to fortuity that comes from outside the realm of calculation and prediction. Something that comes from outside and reaches us. What I want to emphasize today is, as we discussed things like Situational Awareness earlier, when we talked about volatility and uncertainty, the horizon we’re currently looking at is itself far too unpredictable. We use every means at our disposal to tighten controls, predict, and prepare, but in some respects, I think there are areas where we need to be humble, and that thought crossed my mind, so I brought this up.
Jonghyun Park You think RSI, in particular, will make things more uncertain and cause many more contingencies, right?
Seungjoon Choi Right. We don’t know what byproducts RSI might create or what it might bring forth. That’s also why people are talking about pacing ourselves.
Chester Roh We should assume that the harnesses shown to us by model developers or frontier labs were created by models. That’s the assumption we should make. We also build a lot of harnesses within our company, and for the truly important ones, we sometimes build them while reviewing each part human-in-the-loop, but when it comes to subsystems, once we set the objective and they work, we often just use them as they are. Even when errors occur, we don’t look at the commit logs generated by the agents. We just say, “Is it done? Verify it. Look again,” and close it.
Jonghyun Park We ask only what we’re curious about, and if it seems logically sound to some extent, I think we often just move on.
Seungjoon Choi That’s right. So Yuk Hui uses the term Contingency for what lies outside the plan. It’s an existing concept. Since eliminating it is impossible, Yuk Hui says we should treat it as something to form a relationship with. What came to my mind was the large world hypothesis, Sutton’s large world hypothesis, which says that the world is always larger than the agent. Because of bounded rationality, because we don’t possess all the information, whether an agent or a person, we have no choice but to operate with bounded rationality based on what we know. I’ll wrap up with that.
An age of uncertainty and conversation as memory compaction 113:27
Chester Roh That’s right. A comic book comes to mind. When I was young, I remember reading it, Four Daughters of Armian. In that comic,
Seungjoon Choi Was the author Shin Il-sook?
Chester Roh That’s right. That’s right. “Life is always unpredictable, and thus life gains its meaning.” This is a line the author often uses. And it seems especially fitting these days. For that very reason, conversely, a world with such high volatility presents an opportunity for people who want to try something. This kind of change
Jonghyun Park is fascinating.
Seungjoon Choi It does release dopamine.
Chester Roh Well, it’s been a while since we’ve had such a long discussion. But we also, Jonghyun, Seungjoon, and I, bring our materials and go back and forth among the three of us, learning something in the process, and while the balance in our thinking may, at least by next week, disappear completely, our thinking does become balanced today. So conversely, I see this process as a session where we do memory compaction these days.
Seungjoon Choi First, we compact it. That’s right. We retain only those keywords. Because as time passes, we don’t remember the details.
Chester Roh And as we talk, once the conversation is over, a few important keywords remain, and with those few compacted keywords, we begin the following week with some of the load taken off. For us, this process is also an extremely important learning process.
Seungjoon Choi So it was a rollout process.
Chester Roh Right. We finish one trajectory and update our model on-policy…
Seungjoon Choi We do the rollout using only this week’s policy, and by next week, it’s been updated.
Chester Roh Exactly. That’s why even today’s Kimi K3 paper, if you look closely, contains a complete manual for our lives as well.
Jonghyun Park That’s right. We encode the Kimi K3 content from today’s conversation and tuck it neatly inside, where it resides in the latent space through attention.
Closing 115:29
Chester Roh It’s funny even as we’re talking among ourselves. Then we’ll wrap things up here for today. Thank you both, as always, for teaching me so much today.