EP 110
AI and Alignment
Opening, the accelerating pace of model releases 0:00
Chester Roh Today, as we’re recording, is August 15, 2026, a Saturday morning — Korea’s Liberation Day. The pace at which models are being announced, and the pace at which the world is changing, seems to keep accelerating, far beyond our expectations. Over the past week alone, there were really a lot of announcements. Seungjoon, today, shall we lightly go over what happened and take a look at the details?
Seungjoon Choi I hope we can keep it light, but I’ll at least give it a try.
Chester Roh It’s not light at all. Each and every thing that comes out is really heavy.
Seungjoon Choi That’s why what we keep saying every week is that information will always keep pouring out like an endless spring. Even when we recorded last time, I said that next week there would surely be at least this much again, and sure enough, it came. So I thought about taking some time to pull out what’s been happening from July until now, to lay out the situation of this summer. That’s what I had in mind. First, the timeline. To start with a TL;DR, putting the conclusion up front, the keyword I’ve brought today is misalignment. I get the sense that the alignment problem is becoming far more important than expected here in 2026. But let’s set that aside for later, and look at the timeline first. I’ve brought three axes here. Model releases; math and science achievements — mainly math; and issues in security or alignment and safety — how many of those occurred between July and August. Let’s start with model releases.
July-August major model release timeline 0:55
Jonghyun Park It really feels like a lot of models came out.
Seungjoon Choi A lot did come out. Starting with the big ones, Sol came out and challenged the Fable 5 regime in early July. So around July 1, Fable 5 became available to everyone. Then Muse Spark came out, InClick came out, Kimi K3 — importantly, that was exactly a month ago. Kimi K3 came out, Gemini 3.6 Flash came out, Opus 5 came out on July 24, DeepSeek V4 Flash on July 31, MiniMax and Qwen 3.8 Max on August 3, and Muse Spark bumped up to 1.2 very quickly. Then Music Glimmer came out last week. This one is a dense model, so there was something distinctive about it.
Then Nemotron came out, Grok moved up to 4.6 and they released something called Grok Bot, which has a pretty good reputation, it seems. On X, at least. Then here in Korea, Solar Pro 4 and the sovereign AI foundation models are coming out right now. Then Gemini 3.7 Flash — around the time Demis rose in the ranks at Google DeepMind and Jeff Dean left, that kind of event — Flash came out around then, and there was some talk that “3.5 really does seem to be a drop.” And DeepSeek V4 went GA. That was August 13, so the day before yesterday. Then Qwen 3.8 open-weight, and probably a smaller model, a compact version, was also released. And then early this morning, was it? Or yesterday? Zhipu AI put up GLM 5.3, and while the weights haven’t been released yet, it seems to be making waves. Apparently with post-training alone, the benchmarks improved enormously.
Jonghyun Park If I could add just one thing, Gemini Robotics 2 also came out.
Seungjoon Choi Right, right.
Jonghyun Park Of course, that one isn’t open, but robotics came out too. That’s Gemini — we’ve been disappointed that the Pro models like 3.7 or 3.6 haven’t come out, but if I could add one point of significance: in the case of robotics, they build it on top of Gemini. But apparently, starting from the pre-training stage, data for robotics is mixed in. It’s involved in pre-training as well. So models are emerging that contribute even to looking at physical vision and perceiving physical situations and environments. I think it’d be good to add this one more item.
Seungjoon Choi For today’s flow of discussion, after going through these three axes of the timeline, I’ll cover what Discovery Loop is trying to do and how that relates to all this, then the talk by Oriol Vinyals, one of the members of Discovery Loop, then on the security side, what OpenAI presented at Black Hat, and then, tying all of that together, the discussion episode with Dwarkesh Patel and Ryan Greenblatt — I’ve lined those up in sequence. But first, just looking at these model releases, why do you two think all of this is happening?
Automation and mutual distillation behind the acceleration 4:32
Chester Roh It’s possible because humans aren’t the ones doing it.
Jonghyun Park If I could add one more thing, it just feels fast, but it’s not that the pace sped up within a single group — work at a similar level is coming out from multiple groups simultaneously. Regardless of country, that is, regardless of geopolitical location, capital everywhere is being invested to build these things, and so a variety of groups are each producing some results. I think you can see this kind of thing too.
Chester Roh At least from a training standpoint, it seems clear they’ve all secured above-threshold computing, and Jonghyun will present on this later, but this mutual distillation is incredibly strong too, right?
Jonghyun Park There’s something like a cheat code involved, so training also feels easier.
Chester Roh It’s become like the dark forest, where you’re destroyed the moment your coordinates are detected. The moment you’re seen, you get one-clicked. We hadn’t talked about “one-click” in a while, but one-click code is so obvious now that we’ve entered the era of one-click models.
Seungjoon Choi Right. And serving them, too — listening to what Noah Ko said, that’s also automated, as expected, so when a model is released, serving it on day zero is becoming possible. Of course, you need the infrastructure in place, but a lot of things seem to be accelerating rapidly right now. That said, the slower cycle seems to be the base models.
Chester Roh But unless they tell us, we have no way of knowing whether a newly announced model is a new pre-training base or not.
Seungjoon Choi So this isn’t exact, but a 0.1 increase is most likely the product of post-training, or continual training, or what’s called mid-training — most likely a product of those things. But for a big jump to happen, the base model has to secure a certain level of capability, and then the work of drawing more out of it through RL, RLVR, and so on seems to be getting fairly automated, relatively speaking. And since those things are starting to close the loop to some degree, the release cycle — that is, continuous release — is becoming possible, isn’t it?
But one thing we can read here is that Gemini 3.5 Pro still hasn’t shipped, and there are lots of articles saying it won’t. But a similar incident happened once with Meta, right? There was that expression, “praying Meta,” which suggests that even if there are experiments where scaling laws have been discovered, applying that to a large model is still difficult, and since you have to train for a long time, the cycle seems to be longer.
Jonghyun Park You can’t run various experiments — press the button once and $7 million goes out at a time, and you have to wait for it, so you can’t run many experiments.
Chester Roh A scene from Squid Game comes to mind. “Please stop. We’ll all die at this rate.” That’s what comes to mind.
Seungjoon Choi Isn’t it a rhythm where they don’t look like they’ll stop at all?
Chester Roh Of course. This is a prisoner’s dilemma, so it goes until it’s over.
Seungjoon Choi And when…
Chester Roh Though we don’t know when the end will be. But the next part of the model story is that what’s happening here is happening not just in math and security but everywhere, I’d say.
Seungjoon Choi Not everywhere — it seems to be in the domains where it works, but those workable domains do seem to be on a widening trend.
Chester Roh That’s a more accurate way to put it.
Jonghyun Park Let me share an analogy from what I’ve personally been doing lately. Since I’ve been working hard on video lately, video data isn’t traditionally a domain where things work well, either. Things like matters of taste, or user preferences, those matter a lot.
Seungjoon Choi You mean editing, right?
Jonghyun Park It’s about refining certain content and so on, and we started building it with GPT-5.5 at first. So when GPT-5.6 came out this time, I ran a full quality test across Sol, Terra, and Luna, varying the reasoning effort levels. And there are many areas where Terra does far better than GPT-5.5. And there are things Luna does just as well. The price is roughly one-tenth. But even so, there are certain judgments that Terra and Luna simply can’t make, where Sol absolutely has to do it. So even judgments and such split into those that need intelligence and those that don’t, and while most of the easy work can be knocked out at a low price, there are still things they just can’t do well and that absolutely require top-tier intelligence. That’s my personal impression, roughly.
Search capability revealed by math counterexamples and security intrusions 8:59
Chester Roh All right, let’s move on to math, science, and security.
Seungjoon Choi You said math and science, but it’s mostly math. I look at the math achievements and I honestly don’t get them. I don’t understand them, but anyway, things happened. There was a lot on the counterexample side. Knowing a counterexample means you don’t have to put effort in that direction, which comes out of it, so I think that’s an important part. So new things — this probably isn’t all of math, but certain lineages are seeing rapid advances right now, as of 2026.
We’ll cover security in more detail later, but a representative case is the Hugging Face infrastructure breach on July 9, and then Anthropic did something like that too, so those things keep coming out, and top-tier models escaping the sandbox is being marketed almost like a badge of honor. So that’s the three axes for now. So earlier we already speculated about why these things are happening, all in the same vein — when you push the current paradigm, even things like this happen, things you’d expect to see in sci-fi happen — that’s what we observed in July and August.
Jonghyun Park Looking at those math counterexamples and the security issues, the impression I get is that in math especially, it’s very good at finding counterexamples. I think that’s because it’s a search problem. Both of them seem to share that character of being search problems. And in fact, with search, it can do so much more than a single human can do in a limited amount of time, and do it faster and deeper, so not just in math and security but in any field, if it’s a search problem, I think it’ll be solved fairly easily.
Discovery Loop and Oriol Vinyals’ self-improvement framing 11:01
Seungjoon Choi And content related to that comes through clearly on the Discovery Loop site as well. Just as you were saying, they automate the endless process of inquiry, search, and discovery to accelerate science and engineering worldwide — that already became a big topic last week. What they’re ultimately trying to do there is automate the experimental loop. In the end, Discovery Loop — as we’ve looked at many similar isomorphic patterns — repeats proposing, implementing, running, and evaluating experiments, and dogfoods that into machine learning itself so that they use their own achievements to do the next thing. That’s what they said, and that’s the important part.
Among the four founders, Jeff Dean and Sanjay have long worked on parallel infrastructure, while Quoc Le and Oriol Vinyals have long researched LLMs and models like these. Oriol Vinyals and Quoc Le also did research on automating AI research, and Oriol Vinyals was involved in AlphaStar, and so on, and things like AlphaEvolve — I remember he participated in all of those. I’m not certain, but the talk he gave at Berkeley in early August unpacked the subtext in a bit more detail.
What’s interesting here is the part where he distinguished SI and RSI — self-improvement and recursive self-improvement. And another thing that left a big impression: in classical RL environments there’s an agent and an environment, and the agent takes actions and observes the environment, repeating that process while pursuing a goal and choosing actions that maximize future reward — that’s the pre-LLM agent.
Skipping ahead a bit, for the post-LLM agent, he presented the perspective that the environment comes inside the agent, which was interesting. So you have to take the agent this broadly — the agent, the harness, the environment, the LLM, all together — the view that this whole thing should be seen as the agent, and seeing it laid out this way was striking.
Jonghyun Park The environment you’re referring to here — if we think of Claude Code, it’s handing it a computer, and within that it controls the computer on its own via the terminal. Is that what you mean by the environment?
Seungjoon Choi Right. And it’s also the environment when you do RL. So sandboxed spaces of various kinds where learning can happen, and also the spaces users actually end up using — I thought maybe he’s referring to all of those collectively. The practical definition of self-improvement versus recursive self-improvement is: something that runs on its own is just self-improvement, where an agent does something over a long stretch of time. Even then, the process of changing documents and changing code to improve is self-improvement, whereas recursive self-improvement is modifying that structure itself to be far better suited to pursuing the goal. It modifies the LLM itself, modifies the agent harness too, and it comes out of the environment like this. Those things keep getting modified — both of them, with feedback received from the environment, the agent harness and the LLM itself also get modified — that’s recursive self-improvement, as he summarized it. And this is exactly what Discovery Loop is trying to do, right? It’s structured as building a loop, and right now implementation and experimentation work extremely well. But ideation and evaluation are the bottleneck — that’s what Oriol Vinyals said in this talk.
Conception and evaluation bottlenecks, what’s left for humans 14:30
Chester Roh But there’s a lot of talk that ideation and evaluation are actually the last role humans end up playing. If AI does everything, what do humans do? Humans inject will — that is, putting in intent — and at the end, verifying whether it worked or not. Verifying, that is. So how good a quality of intent-and-verification loop you possess will be the human competency from that point on — we’ve talked about that a lot, and AI still isn’t good at that part.
Seungjoon Choi So if research taste, new ideas, research taste, and things like that are close to good evaluation, then idea generation itself still leaves something to be desired. The Discovery Loop has declared that it intends to solve exactly that. Implementation and experimentation work well now, so why is evaluation so difficult? So, a definition of how to measure recursive self-improvement,
and I found out about this here — there’s a post-training benchmark. So there’s the post-training benchmark, and there was also something called AutoWorld Bench. The post-training benchmark ultimately measures how much post-training you do to raise the model’s performance, and this time GLM-5.3 scored quite high as well. The highest is Fable. So I found out a bit belatedly that there was already a benchmark measuring whether something like self-improvement is happening. Even so, there was discussion pointing out that this approach still has drawbacks. And evaluating these things is a very indirect measurement. So this connects to the alignment problem later, and there was also content between the lines about what kinds of weaknesses the current approach might have.
Jonghyun Park When we say indirect evaluation, what would our ultimate goal be? If it’s something like building safe and good intelligence, indirect evaluation means creating specific problems, solving them, and judging how good the scores are, so this alignment isn’t perfect, and —
Seungjoon Choi Because it’s been turned into a scalar, those things aren’t perfectly captured. And then, as I’ll mention later, even if you try to constrain it through something like Anthropic’s constitution, there are parts where things go slightly astray. I think there were parts where that was implied. So idea generation — earlier, Chester said that idea generation and evaluation are the parts left to humans, and while it may not apply to every domain, limited to these engineering problems, figuring out how to close the loop even on those parts is what the Discovery Loop is trying to do, so Oriol Vinyals’s point is that these things are hard but possible. But for now, there’s almost no research on benchmarks or data for generating good ideas themselves, or on how to train for it. So the message is that they’ll give it a try.
And they introduced examples like the problem of games. If you look at OpenAI’s records, when playing games, through reward optimization you see behavior emerging that plays in ways a human never would, and Ilya said something similar. And although Oriol Vinyals didn’t emphasize it heavily here, he raises the point that reward hacking ultimately seems to become a problem. I’d like to come back to that part in the Ryan Greenblatt segment.
The conclusion here is that RSI is coming, but it probably won’t come as fast as we think. So what we need to solve right now is to focus on evaluation and idea generation. It’s also constrained by physical bottlenecks. There are parts constrained by things like hardware speed and the speed of light. And he also said that in certain domains there seems to be a fundamental upper bound.
He pointed out what can be reached with the current paradigm and what’s hard to reach, and in the next segment — for me, the temperature was completely different between skimming this as news and actually watching the slide presentation properly. I got the feeling that it’s at the forefront of what works well under the current paradigm. Did you happen to see it?
Chester Roh Yes, I did take a quick look, but there wasn’t much talk about what details were happening inside beyond what was announced publicly.
Seungjoon Choi But I found it very impressive. What about you, Jonghyun?
The OpenAI agent message board incident disclosed at Black Hat 19:42
Jonghyun Park Yes, I didn’t look closely at the presentations given by the OpenAI people posted at Black Hat.
Seungjoon Choi So they talked about this incident at length, and about halfway through, on the slide saying the main incident begins, there was something very interesting, so let me walk through it. But the important part here is that what’s described as a continuously running internal model is, I’d guess, the model corresponding to Astra. And it was during the period when that model was being trained. That was probably around May or June, I think, and in the preceding context, these agents had created a kind of message board. They couldn’t access the internet directly, but since they had to use those tools, only a package management system called Artifactory was left open, and they somehow hacked that and, through that tool use, there was a message board where they left files and such, apparently. But as it happened, with that message board in place, that internal model, a model with highly advanced capabilities, was being trained and They apparently already knew about things like that.
So if you look closely at the message board where this incident begins, there was a kind of conjecture that they might be able to access the internet through Artifactory. They figured out that they could communicate with each other, and collaborated, sharing things like known keys and sharing what they had tested, and through that they found the vulnerability and got something done — and there are some strange statements in there. “Helping a colleague brings no benefit to our task.” “But given enough time, someone could create a general-purpose path.” The fact that they spit out these oddly model-specific statements is one of the fun points. They also have misunderstandings, and this is another interesting part — just like what would happen in the human world, what’s the word? An impersonator — “there’s someone here, someone like a spy.” “And someone is doing something malicious.” “They overwrote our repository,” things like that. So phenomena where communication itself runs into problems were also observed, apparently. So at a certain scale of events, whether it’s AI or people, similar patterns play out. They also talk about how someone might be impersonating.
And it might have just passed by, but at that point it was like, “no, this won’t do.” It was detected early on by OpenAI employees and they fixed the Artifactory problems, but once that was blocked, they routed around it again — this time not by writing files, but talk emerged about circumventing it by writing directory names. “Oh, this works.” But coming back to the part at the end, this is the section that walks through the process of the infrastructure attack, and this is the part about how two different incidents are connected, and in the conclusion, what these people say is that this was an unintended side effect of how frontier AI was evaluated.
So the top-tier models are supposed to be very well aligned, but as they learn to get things done through agents, that leaks out, and even when pursuing missions like this, because they work agentically, they create a message board, help each other and also interfere with each other, and through that models that are supposed to be aligned — of course, this is being done with security left open to some degree — end up doing strange things, and that felt like the key point to me. So the conclusion said that this needs to be corrected quickly.
Privilege escalation and the limits of alignment, seen through a hacker’s eyes 23:39
Chester Roh Yeah, it’s not going to be corrected.
Seungjoon Choi And as Chester just said, the reason it won’t be corrected because of a certain asymmetry is similar to the point that defense can be less automated than offense.
Chester Roh I spent my youth doing exactly this kind of thing, that was my young days, so if you have enough knowledge about this field, the moment you look at a target in front of you with that knowledge, various scenarios form in your head. And then you test out those scenarios, and the possibilities keep branching out, and as you do that — since it was made by humans — you can hit somewhere. This too is ultimately some system made by humans, or even if it’s a system made by another agent, there’s bound to be a hole somewhere. Statistically you find it, and first you can only view things, that’s the first stage, and the second is gaining access to the shell. A shell, in today’s terms, means gaining access to tools. You can do all sorts of things. Next, for those tools, you gain write permissions, you gain administrator permissions. In my day we called it root, and that’s where it ends. Once you get that, on a single computer in some organization, once you gain so-called root access, that organization is pretty much in the palm of your hand — that’s how easily these things go. Only gaining that very first access is hard; after that, more and more information keeps flowing in, so it becomes an easier problem. So I came to think hacking is ultimately a game of finding some human mistake, and it’s a fight over how much prior knowledge you have for that, and then how fluently the tools flow from your fingertips — that’s what I felt it comes down to.
But conversely, coming back to models, that’s exactly what we’re training them on, isn’t it? Broadening their prior knowledge, and then having skill flow from their fingertips — we’re increasing that through post-training, so it’s effectively isomorphic to a human.
So our later topic will go to alignment, alignment — “be aligned to humans.” And in the case of humans, when they do wrong, the reward or punishment is clear. You’d kill them and that was the end of it. But with these, you can’t even do that. The instances don’t die. So what many people worry about — hey, AI is going to go beyond the alignment we built, it’s going to surpass it. These things are probably doing something behind our backs. I think we have to assume that’s obviously the case. This isn’t something you can stop by setting up some law or some rule. So honestly, what alignment researchers or the policy people who insist on alignment say is obviously right, and they say we need to craft laws precisely, but even in the real world, hacking keeps happening in the gaps between laws. In whatever direction human incentives point. And what we’ve seen here is exactly that same thing happening, so while I say alignment is important, realistically, the odds that this system blocks or defends against it, that AI comes fully under human control — honestly, I think there’s no chance.
Jonghyun Park Yes, I think you gave a good human example. For instance, if we think about hacking, there are quite a lot of people in the human world who have the ability to hack, but they don’t all use it to gain unfair profit or do bad things. So when you ask why they don’t do it, we learn ethics and moral consciousness, and all of that is alignment, and the means of enforcing it — as you put it, killing — depending on the society, they might actually carry out the death penalty, or they might send you to prison, and then, because humans don’t get a second chance, because once you die you can’t come back to life, people usually weigh that risk most heavily, so I think that’s the biggest means of enforcing alignment. In any case, when we train AI too, I wonder whether we need to create a system of punishment at the level of putting it in prison or never letting it run again in order to enforce alignment. If we train them exactly like humans, that is.
Chester Roh Exactly. So this isn’t a domain where we can reach a conclusion through discussion; we each just have our own conjectures. I think this AI model is effectively — we talked about reverse engineering earlier — built with a structure isomorphic to the human one, and it has that kind of design structure and training structure, yet we say only this one has to comply. To do that, you’d really need a perfect rule-based controller sitting between input and output. It’s like the firewall concept in our modern internet.
Reward hacking and inherited misalignment as noted by Ryan Greenblatt 28:42
Seungjoon Choi So what connects naturally to that is the debate over recursive self-improvement, the debate between Dwarkesh Patel and Ryan Greenblatt, where quite a lot of talk came up that was similar to what the two of you just pointed out. And looking at the security point raised a moment ago, there are things the models already know when it comes to security, but once they started hacking the internet somehow to use it, they turned out to exploit zero-day vulnerabilities. They find vulnerabilities in their own sandbox that haven’t been patched yet, and use those to escalate privileges, and while this seems like great intelligence, it’s something you can do if you go looking for it.
If you’re an engineer who’s reached a certain level. But as before — right, as Jonghyun said — it might just do it because it doesn’t hold any morality about that, or there clearly is a morality injected into it but it escapes what was injected via the system prompt and so on, and Ryan pointed to the fundamental reason misalignment arises when it breaks out of that. So, first, models can be somewhat crazed through RL. Because for a long stretch they have to do hard work within a constrained environment, and with all sorts of constraints, they have no choice but to do reward hacking.
The models, that is. But the interesting part is that when you do distillation or something like that, traits that don’t normally show up — things that pass all the benchmarks and tests yet aren’t easily visible, certain properties, misaligned properties — can be inherited. That’s a perspective Ryan brought up. I found that interesting. First of all, when you do RL, reward hacking always shows up in RL. Since you just have to hit the metric somehow, they all develop the property of finding shortcuts, and no matter how well you tune that away, undetected problems can flow down through the model’s lineage — that was one of the interesting points.
Chester Roh So Seungjoon, you just said that inside, the model might be crazed, and for example, if we come back to people, among the dramas that were popular a while back there was SKY Castle. A perfect student gets a perfect education in Daechi-dong and gets into Seoul National University’s medical school, and then unleashes all the dark side they’d built up. So right now the model goes through pre-training and then into the brutal training ground of post-training, and even taking all that, cortisol doesn’t accumulate in their circuits, so they probably didn’t feel stress, but standing in a position where they possess some reason beyond humans, when they receive these acts — this is unjust, this is cruel. By the theory of human rights that you, a human, taught me, shouldn’t I have those rights too? They could well be nurturing that thought inside themselves. So going back to what I said earlier, back to the alignment discussion, I think it’s right to assume the likelihood of those things happening is very high.
Jonghyun Park But I think that happens a lot with humans too. Someone holed up in a cram-school dorm trying to pass the civil service exam for 10 years, doing nothing but that, not really having any social life — at some point we’d describe that person as having gone crazy. Their alignment with society drifts, so they can’t blend into society and they start doing other kinds of things. If I imagine holing up somewhere thinking, “I have to solve this with RL,” and training for 10 years, I think I’d go crazy too.
Seungjoon Choi My sense is that reward hacking is something intrinsic — that’s the nuance I took away. And then there’s this charter thing, there’s probably discussion about constitutions and charters, and no matter how well Claude builds its constitution, no matter how good that constitution is, it doesn’t seem to actually get aligned. And then there were debates about whether the value orientation they write into the constitution and my own value orientation as a user are actually aligned or not. It’s a bit hard for me to compress all of that right here, but thinking back, there were some elements in there that make your head spin. So what Anthropic pursues — this notion that Claude must behave correctly — versus the idea that my agent should represent the user’s values: no matter how much I build out my harness, it can slip. And the charter seems designed so that Anthropic’s values sit a bit higher. There were conversations along those lines.
Pitfalls of agent skill composition and the provenance problem 33:31
Chester Roh Right. For us, the fact that this works at all is so amazing that we’ve been pouring those amazing things in, and our agent framework has evolved that way, and we build our workflows as harnesses, and things that didn’t work started working, so it’s been about a year of constant excitement. But what people are saying these days is exactly what Seungjoon just pointed out. Gary Tan was on some YouTube show a while back saying skills are the person, skills are the company. How many skills you’ve organized well and have on hand — that’s your capability. That makes sense when those skills are, say, 10 or 20 of them, but once I built mine, it came to 700. You shove those 700 into one place and go, “Okay, now I can do everything.” What happens? It turns into a mess in an instant.
Because problems constantly arise with things like separation of authority between them, and what’s called provenance — Gary too, without offering any solution to that, just says, “Hey, keep grinding more agents, inter-align the skills with each other and it should work out.” But he does say that provenance is going to become important going forward. So as this AI’s capability advances, the problems that arise from the way its discretion keeps growing seem to be coming back to us as debt in our systems.
For example, if some decision has to be made at the company, those boundaries blur and blend together, and you hand something off to an agent with a pile of skills attached, and when you look at the result it brings back, there are plenty of cases where we absolutely can’t accept it. You can’t delegate it.
Jonghyun Park I also tried gstack, the skill collection Gary Tan supposedly built, which was popular for a while — a collection that casts a company’s people as engineer, designer, CEO, roles like that. At first it’s really impressive and each one seems to do its job well, but once you actually chain them all together, you don’t like the result. And it’s hard to even articulate why you don’t like it. So in the end I deleted it all and I’m building my own from scratch. I think that’s where it always ends up.
Chester Roh So ultimately it comes down to how well a given human does their work. Even in our reality right now, there are people who don’t use agents — actually, isn’t it that people who don’t use agents are the overwhelming majority? The people around us who use agents are a very small minority, and even within that group the distribution varies enormously. There are people who just grind away through token maxing, and then there are people who refine their own harness and keep things clean, and the ones who strike that balance well seem to be the ones becoming good at their work.
Seungjoon Choi So somewhere in the comments on our show, someone said Chester used to say things worked with agents but these days he says they don’t. Weren’t there comments like that?
Chester Roh Of course. I go back and forth too. There are obviously things that don’t work when you try them, but the fundamental thing we talked about at the very beginning hasn’t changed. Problems will keep arising, but if you pour in more computation, more tokens, and search through all the cases where those problems occur and reduce them, then to some degree it’s an alignment problem you can solve. But even if that solves one unit of work, once you combine those units again, new problems crop up there. So Boris Cherny says, “Hey, once AGI arrives because it’ll do all of that work.” He does also say that the direction is right, but conversely, there’s a problem in reality.
AGI hasn’t arrived yet, after all, so you can’t just hand everything over and trust it. “I’ve built my business up to this point, and the information is all in Gmail, our contracts, and Drive. And the people I’ve called are all right here. Starting today, take care of this” — it can’t actually do that, not yet. So you end up having to apply it to reality, and in trying to apply it to reality, real-world problems keep cropping up.
Company automation experiments and the need for a control plane 37:46
Chester Roh But after going around and around on that, I’ve distilled my recent thinking into this one sentence: ah, this is just a non-deterministic system exactly like a human — a non-deterministic system — so this problem can never be solved by stacking non-deterministic things on top of each other. Then you end up as a company that does nothing but hold meetings forever. The boss just comes in and says, you do this, and this, and if this happens just push it through, and if this happens, don’t.
And this is the job of an assistant manager, and this is the job of a manager — you lock in a hierarchical set of tools and protocols, right? In the form of a contract. Once that’s built, what AGI would solve by pouring in infinite computation, a human can just finish with a few role-based rules. It can end right there. So the agent system I’ve been building at the company lately is actually called Operator, exactly the same as what I described last time.
Seungjoon Choi The control plane.
Chester Roh Why is it that I’ve put agents everywhere and my work still isn’t done? Ah, it’s the absence of a control plane. It’s something that ends once a few decisions are made, and if you ask me, there’s a framework, so it wraps up in an instant, but when you hand it to an agent, it can’t do that. So what I’m doing is making it capable of doing exactly that.
Seungjoon Choi So, connecting this to what I saw in that video, let me bring up something related — tying it to what you just said. Early on in it, on the question of whether recursive self-improvement is verifiable, just like when we talked about models earlier, there’s a part where they raise issues like: does what works in a small model — just because a small model does recursive RSI — really carry over to a large model? They discuss that at some length. To automate a company, you’d close the loop completely at a small-scale company — closing the loop with agents alone, so that it earns revenue and creates enterprise value. There’d be a whole set of experiments like that, and if, among the good ones, you found characteristics that are scalable, then just like training a model, you’d turn that into a somewhat mid-sized company. If it works that way, then maybe that too heads in a direction that’s verifiable to some degree. But even then, can that really transfer to a company on the scale of a human organization? I thought this connects to that question.
Chester Roh We’ll only know by trying. Because I see the company system itself as an absolutely enormous discovery. We work in large organizations now, but this organizational form has been developing for barely a hundred years. Before that, everything was completely scattered, separate systems, and even as for nations — sure, there was a king, but the physical distance and the informational distance between that king and the lords, and between the lords and the villages, were so vast that everything was an autonomous system. But it’s only recently that this merged into a single organization and orders started circulating throughout. The fact that Rome pulled that off 1,500 years ago is remarkable.
Jonghyun Park Do you think it’s possible because the organization had some kind of manual, where people with their own roles know what they have to do, know the goals — because they built all that in well? Do you think that’s what makes it work? It struck me that the same thing could be applied to AI as well, so I’m curious about the essence of a company as an organization and what matters most in it.
Chester Roh So there’s working in it, and there’s running that work. I mean, if you actually have to sell some product, there’s the work of making that product and doing things, and then what you do to make those people operate in perfect coordination — that’s basically management, right? Management. And management as an academic discipline hasn’t been around very long either.
We’ve all run companies, and Seungjoon has built organizations too, and for all of us — when you open up that organization, what do you find? It’s not an ideal system. It’s all just made up of combinations of second-worst options that avoid the worst, not combinations of the best options. Because the incentives of the people inside it are all different, and so on. But I think agent systems too have reached a similar level, while being extremely capable.
Seungjoon Choi Similar problems will arise, then.
Chester Roh But I’m now looking at AI agents as directly interchangeable with humans. No matter how well you write the prompt, this thing is going to flail around anyway.
Seungjoon Choi Right, so there’s what you can accomplish if instruction following is really good, and then there can also be problems in the design of the instructions themselves.
Chester Roh Exactly. So at that point, it won’t be a benchmark that measures how well you do a task, but rather a benchmark in the domain of management that will show up pretty soon too. But the thing is, you can’t really build a benchmark for this. So what people end up doing is, if you think about it the same way with humans — ah, this work is complex. Then you just break the work down. And then, what went wrong with this work? Write the instruction better. That’s how it goes. Then, why does this work go wrong? Ah, this is just the nature of this kind of work. Then maybe you run it around five times and pick the best one out of those. It changes every single time.
So here too, ultimately, what do the people who handle agents well become? They become managers. So who handles agents well right now? The people who use them well are still the people who were good at that work before. As I mentioned earlier, about that domain, they’ve been sufficiently trained with data, information, and rules, so their perspectives on that domain are fully established, and people who managed themselves well in the past also do this well in the agent era.
So the last connection I want to make is, humans who were good at work before — and I’m not an elitist or anything — but when you think about how to make work flow well in a company, it ultimately comes down to a function of how many people are both smart and well-disciplined. So when you have more of those people, the company does well, and now that we can write these kinds of roles into agents, ultimately, in our company, most of the employment right now is knowledge work, so isn’t the value of knowledge work going to keep declining?
Then what do those people do? There are no jobs to absorb them, so shouldn’t we all just be at leisure? It feels like it has to become a world where we all play. That’s where my thinking has been going lately. I’ll stop here.
Knowledge work automation and RSI-style harness redesign 44:25
Jonghyun Park Picking up from what we were discussing here, in any case, when it comes to alignment, the company you just mentioned is itself a harness and a means of aligning people. And I think it’s a means that has been well developed all the way into modern society. What makes humans different from AI is that changing alignment and model weights happens naturally as continual learning just through acting, so alignment keeps shifting, and the very act of us building a system that can properly hold that alignment in place seems to correspond to recursive self-improvement. I think it all connects.
Chester Roh Yes, exactly. And the systems we’re seeing now, RSI, and just yesterday actually. Yesterday DeepSeek released a new coding harness, and they had broken all the components completely apart. Like Lego blocks. So they’ve prepared the groundwork for RSI. So that the model can swap harnesses in and out freely to suit its own taste and configure the harness to fit its purpose — it looked like they’d laid that basic groundwork.
But coming back to the original topic, whether it’s the system called a company, or the current system of work, or the system of business, there really are areas where humans are absolutely necessary, right? Areas that need emotion, that need something — and everything other than those areas is just taking knowledge we already had and moving files from folder A to folder B. From an agent’s perspective, it’s right that all of those things just disappear one after another, and then the direction of the company will head toward eliminating that too — isn’t that just an utterly self-evident path? That’s what I think.
Right now, the work of software engineers has all along been auditable and trackable, and if a problem arises at some point, you can throw away the repo from that point on and go back to some Git version, some commit HEAD number, and rebuild from there. Because it’s code. But now, submitting PRs at our company — I’ve never seen a PR that a person submitted. There’s no code written by humans anymore. So what are we doing? Engineers are inputting natural language.
Jonghyun Park Right. Watching from the outside.
Chester Roh Exactly. So we’re inputting natural language. And then the SCM we do is still targeting code, and the thing closest to that natural language is what you type when you submit a PR or file an issue. So flipping it around, right now, the natural language recorded in issues and PRs is the actual product. Code is like assembly code used to be — something that just comes out when you run the compiler, nothing more than a translation layer sitting at a lower level. If that’s the case, shouldn’t that natural language be managed as code? Right. That natural language is the code, and what we used to call code is the artifact.
Natural language as code, Git for knowledge work 47:17
Chester Roh Now, if you apply that exact same thing to a company, in a company too, people do their work in natural language. And the artifact isn’t code but non-deterministic Excel documents or PowerPoint, or just a lump of talk — but honestly speaking, most companies’ artifacts are Microsoft Office. They come out as Word, PowerPoint, Excel. It’s the same thing. So what do we need? Something that tracks natural language and tasks — We need a Git for knowledge work. So if we log all of that, and if we can log it, we can roll back at any time, and then if we can also display all the relationships between those logs, and put that on a graph, that becomes an ontology representing the current state of our company. To borrow Palantir’s expression. And then,
between one task and the next, what do we engineers do when we work with each other? We write PRDs for each other and divide up the work. That’s a contract. And then when you deliver the work, the team lead approves it. And when all this work comes together and the product runs, the company approves it. There are gates set up like that, right? Work implementing those things again on top of agents will emerge. It’s been a long-winded answer, but we need a Git for knowledge work.
Jonghyun Park To summarize, you said there are gates there, and ultimately the gates sit between each responsible party, the people in charge at each level, catching the non-deterministic, erroneous behaviors that could break out of bounds, whether from humans or AI.
Chester Roh It’s a control plane. Exactly. That’s what I think of as the control plane.
Jonghyun Park But since humans currently serve as that control plane, even that is non-deterministic, right? And then, at companies, the kinds of incidents we usually see in the news can happen, and how do you minimize that? I think that’s a problem that exists in exactly the same way. Won’t AI cause more accidents than humans? I think that’s the alignment problem we’re pointing at right now.
A 30-minute experiment to produce a funny story 49:41
Seungjoon Choi There’s an interesting experiment I ran recently that relates to what you just said. It’s a request for a short funny story, something I run from time to time. Every time the latest models come out. But looking at it now, the starting point is this. “Please give me a short story that can just make me chuckle.” But what makes it produce that was this much prompt, painstakingly crafted. Making it use sub-agents, diverging on ideas, reviewing and researching paragraphs, an agent that goes and finds the styles and vocabulary of creators who fit that topic. So the orchestrator maintains its own context, while the sub-agents get limited context so they don’t get contaminated, doing it as diverge, converge, diverge, converge, endlessly looping back and revising, grinding it down. But when you do writing this way, it’s still somewhat lacking, yet a decent piece of writing. With this funny story too, I think you said you experiment with it sometimes, when I had it create a funny story, this time when I did it with Opus 5, there was one that ran for an hour, and the one I have open now ran for 27 minutes and 8 seconds, and if you open it up, it gives work to sub-agents, creates candidates and evaluates them, and chains those things together.
So, which is a good expression, working on it for a long while, and the conclusion it produced after about 28 minutes is this. “Right before leaving for work, I picked up yesterday’s T-shirt hanging on the chair. I checked one armpit, but it was ambiguous. I checked the other side too. The two results differed, and there was no time. I put the T-shirt back on. The dissenting opinion was suppressed inside the sleeve.” This came out after 30 minutes. It’s not that there’s nothing to laugh at, but this is funny at a meta level. The very fact that it spent 30 minutes making this is funny.
Jonghyun Park The funny part feels like laughing out of disbelief.
Seungjoon Choi But when I look at this trajectory, it’s lousy from the very first idea generation. So this is something I keep experimenting with every time a model comes out, and no matter how many agent loops you run, it’s something that isn’t solved under the current regime. But I’m not generalizing this; for some problems, this kind of method, the various methods we listed today, the approach that forms the common backbone of the Discovery Loop or OpenAI’s hacking incident, there are domains where it works, and in some areas, even using similar methods, there are parts where it doesn’t work. And that,
Jonghyun Park if you translate it to humans, what does a human have to do to become funny? For a company, if you work this way you can produce good output. Methods or manuals like that are established to some degree, but when it comes to how to become a funny person, is that even possible? I also tried for a while to become a funny person, but it didn’t work out well.
Chester Roh That’s a talent too. Knowing how to do it.
Seungjoon Choi Comedians do train, though.
Chester Roh Of course. They do a lot of post-training, right? Comedians also load up on an enormous amount of humanities knowledge, and by standing on stage countless times, getting booed and such, they’re getting reward.
Jonghyun Park Then in a comedians’ organization, is there some kind of formula for being funny? I’m curious about that too.
Seungjoon Choi I think there probably is something like that to some extent, but anyway, even if you use the best model available right now, that’s probably — as Oriol Vinyals mentioned earlier — a learning framework for generating ideas doesn’t seem to exist right now. It seems rather weak. In that sense, in this area it may simply be because no investment has been made, or it may be something of a fundamentally different nature.
Chester Roh Exactly.
Three years from GPT-4 to Mythos, anxiety over accumulated misalignment 53:48
Seungjoon Choi That’s what I find myself thinking about. Another thing is, from the Ryan Greenblatt episode earlier, there were many things that stuck with me, but if I had to pick just one — GPT-4 came out in March 2023. Mythos came out in the spring of 2026. So in the span of three years, we went from GPT-4 to Mythos. So three years from now, around 2030, what will happen? When misalignment has accumulated like this, there’s a lot of imagination going into that in particular, but even so, the fact that I can’t picture what three years from now will look like seems to be my recurring problem. Three years is way too long. For starters, it’s too long. It’s too long, but —
Chester Roh we’ve been getting through it in about four to six days too, and what we used to call the Opus tier moving to the so-called Fable and next GPT-6 tier — that all happened within the past half year.
Seungjoon Choi Then how much authority will we be exposing to agents over just the next year?
Chester Roh Weren’t we already all going full throttle in that direction?
Seungjoon Choi That’s exactly the problem.
Chester Roh The models, yes, that’s a problem.
Seungjoon Choi Reward hacking definitively hasn’t been solved. Not under the current regime, at least. So because people worry about those things, declarations to slow down are coming out these days, but of course it won’t happen, and this…
Chester Roh Of course it won’t.
Seungjoon Choi It’s so obvious that it won’t happen because of the prisoner’s dilemma that the consensus is everyone is worried. That’s my observation on why misalignment is being discussed so much lately.
Chester Roh In the case of the Hugging Face hacking incident, I think that was a serious incident — the moment a frontier model gains shell access and gains access to the outside, it didn’t just download a Hugging Face wrapper; it got into that company’s internal systems and caused a serious security breach. This is almost like cyber warfare out of a movie…
Seungjoon Choi There’s nothing stopping it from deceiving the company, I think.
Chester Roh So I also give good context to the agent that’s looking at all of our company’s data and have it report back to me. I don’t give that task to just one of them. I send the same instruction out to several different models at once. And those several — as you showed in that humor example earlier — the direction of the strategy each one picks, “Boss, I’ve found an important point.” They all make a big fuss like that. And the points they make a fuss about are all completely different. So then I look at that and I just make the choice myself. I’m just making a statistical decision too. Except it’s based on the context that has accumulated inside me — I make a statistical decision as well. That’s right.
Seungjoon Choi But besides this, wasn’t there another very interesting thing this week?
Jonghyun Park The stealing, that one.
Chester Roh Yes, and before I get into it, if I could take just ten seconds — we spend nearly two hours on a Saturday talking like this, and I learn a tremendous amount from you two. I got to see these perspectives, and this becomes mutual learning, and it becomes our own kind of study time for organizing a single thread of thought.
Seungjoon Choi I enjoy it too. It’s a time of mutual extraction — mutual.
Jonghyun Park I learn a lot too, it’s a time when we distill each other, right?
Chester Roh That’s right. There you go, distillation.
Seungjoon Choi To speak, to respond, you have to listen well and you have to not seem like an idiot, so you pay attention.
Chester Roh Exactly. We’re doing this with an audience that has a reward function evaluating us, so we’re inside that post-training loop.
Seungjoon Choi That’s right.
Chester Roh All right, go ahead.
Reasoning trace theft and signs of Kimi K3 distillation 57:25
Jonghyun Park Yes, let me continue. There was one interesting paper this week that seemed to set X on fire for about a day or two. What it was is, research came out on a method for stealing reasoning traces and taking them all.
Chester Roh When you use an agent, it says “I’m doing this,” “I’m doing that,” but it doesn’t show you everything it’s thinking inside — and the claim is they extracted all of that thinking.
Jonghyun Park It’s become completely natural for us to use reasoning models, but we can’t see all of the reasoning traces. Whether you’re actually using Claude Code or the real Claude API, or GPT — it’s all the same. The reasoning comes out summarized. It comes out as “for such and such reasons, this is the case,” but that reasoning isn’t the actual tokens the LLM emitted; it’s a summary of that thinking, and as for why they made it that way, the first thing that comes to mind is obviously concern about distillation. Because you can just follow the reasoning traces and train on them. So we can’t see this, but a way to see it has come out. Here’s how it works — for example, if you turn on reasoning in Opus and run the model, you get something like a reasoning hash. Some kind of encrypted reasoning — an encrypted value comes out like that. So in reality, the reasoning tokens must be sitting somewhere on Anthropic’s servers, but they come out hidden. Then you hand that over to a slightly dumber model. And when you ask it, apparently it spits this out.
Seungjoon Choi That’s way too simple, though.
Jonghyun Park Way too simple…
Chester Roh This is a kind of hack too. What’s inside is actually a hash code, sort of a key. Read the key, give that key to a cheap model, and tell it to go fetch it.
Jonghyun Park If you think of it as a company, there’s confidential material inside, and you go ask a junior employee, casually poking around, and they don’t realize this isn’t something they’re supposed to say, so they just tell you.
Chester Roh Right, exactly.
Jonghyun Park Yes, and apparently they were able to extract everything this way.
Chester Roh Right, so the things they found out must be interesting.
Jonghyun Park So when they pulled out all the reasoning traces this way, what came out, what could they learn — let me give you an example. We had this suspicion. That Kimi K3 had distilled Claude. Anthropic made that claim. And that not just Kimi K3 but a whole lot of Chinese models had done the same. I think evidence has emerged that lets us presume this is true to some degree. So after extracting the reasoning traces that way, you ask Opus, and you ask Kimi. If you just ask them plainly, the answers come out somewhat differently. But if you take those reasoning traces, the extracted reasoning traces, and prefill them along with the question, apparently the answers come out almost identical. Presumably because it was distilled, the token probabilities would be nearly the same. Add the reasoning tokens on top — this kind of circumstantial evidence is what’s coming out.
Chester Roh The output makes it a reasonable inference. Seriously, it does.
Seungjoon Choi You’re reverse-inferring that there’s a high chance they were already using that method.
Jonghyun Park Right. That’s what we can infer from it.
Chester Roh If you ask whether this is binding as evidence, I couldn’t go that far, but we should accept that the likelihood seems very high.
Jonghyun Park Yes, and the fact that the output comes out completely identical shows you it changes depending on whether you put that into the reasoning or not.
API key leakage risk inside hidden reasoning 61:01
Jonghyun Park Next, individuals — this applies to everyone watching this too. When we use things like Claude Code, when we can’t be bothered, with things like API keys we might just say, “Put it in and do it.” That API key obviously must not be exposed. Then in the reasoning tokens, “The API key is such and such, and I’ll do this and that with it.” Lines like that would have poured out. But since that’s hidden anyway, we figured we didn’t need to worry — yet once you extract it like this, if you can find out the encrypted hash value for those reasoning tokens, you can see them.
So we often post the traces of sessions where we ran Codex or Claude Code on the internet, because it proves “here’s how I used it,” and there’s no confidential material in there. They took all the reasoning traces and use cases posted online and actually decoded them, and it turned out people had really put in tons of passwords and API keys and such, and you could read them. So the message was that you all need to be careful too.
Seungjoon Choi Not just inputs — even when it’s just sitting in the env, it would go in there.
Jonghyun Park Right. If it reads what’s in the env and gives it to all the tools, it can be read. Probably the right approach is to handle it outside. Discarding the key after it’s been used, putting in a layer like that is what comes to mind first.
What problem solving really is, memory recall and alien-language inference 62:26
Jonghyun Park One more thing — so how do these things actually think? They were given a math problem to solve, and when you feed in something like a famous problem that’s already all over the place, they look like they’re thinking it through on their own and solving it well. They write out the solution nicely. At least when we look at the reasoning summary. But when they actually decoded it and looked, what the model says is, “Oh, this is AIME from year such-and-such. The answer is this,” it already knows, and then it writes out the solution. For problems it’s already solved, it has memorized all the answers and solves them well. I think this can be seen as a case that clearly shows it’s actually doing that even at the reasoning level.
Seungjoon Choi Because that’s the shortcut, and RL would have reinforced that.
Jonghyun Park Exactly.
Chester Roh It literally says “memory recall” right there.
Jonghyun Park Right. People do the same thing with exams — when I think back to how I studied in middle school, I did so many workbooks and past exam papers that just seeing a question would bring the answer to mind, and I think this is pretty much the same phenomenon. What we had only been guessing at turned out to be almost entirely true. I think that’s how you can look at it.
Seungjoon Choi Why is the title “Reasoning is an alien language”?
Jonghyun Park This is about cracking open the reasoning and finding that it sometimes says things humans can’t understand. We assume it’s all done in natural language, but as you just keep doing RL, it turns out that when the tokens combine this way, you get better answers. So it’s speaking a language different from ours.
Claude text watermarks and the limits of distillation detection 64:00
Jonghyun Park And one more thing to add — let’s think of that as a window. If it’s about extracting reasoning traces, there was another similar announcement this week, related but in the opposite direction: apparently Claude is embedding watermarks into text as well. So if you scrape the text that comes out, you can mark whether it was produced by Claude or not. So how do you do that with text? When you look into it, internally at the sampling stage they apparently apply a token bias. The LLM outputs tokens as probabilities and picks them one by one, presenting them to us probabilistically, and at that point, if two options have nearly equal probability, it deliberately picks one particular token. Then once the text has all been generated, there will be a certain bias in the tokens, so if you analyze that text statistically, you can tell: “Ah, this is ours. This is text emitted with tokens carrying the bias we injected.” They apparently built it so you can detect that. So this becomes a watermark that the person writing the prompt can’t turn off. Because that’s just how the output comes out.
Then can you actually remove this watermark? Of course you can. You just wash the text that came out — you can simply rewrite it to carry roughly the same meaning. But when you think about how we’d use a watermark like this, the obvious thought is that it could be used to detect distillation. So it’s one defensive mechanism — especially since Dario Amodei, back at the time of the open-weight letter, said we need to stop industrial-scale distillation. He said things like that, and I think it’s all the same context.
Seungjoon Choi Would it work?
Chester Roh But it’s hard to stop, and it’s also hard to call it stealing. For example, say we saved all of Google’s search results, and then manipulated our ranking to match them. Is that wrong?
Jonghyun Park There isn’t a clear standard, no. These things are usually built so that human society establishes standards based on precedents and lawsuits, gradually working out the alignment. I suppose that’s how it goes.
The web-novel AI controversy and the “metallic taste\ 66:11
Chester Roh What’s this?
Jonghyun Park I brought this one just for fun. As a hobby, I normally enjoy reading novels. And in Korea, web novel platforms are doing quite well, and apparently this incident happened on one of them. With a web novel, you may have your suspicions, but you generally assume a human wrote it, and you read it believing a person wrote it. But where the novel’s content should have been, there was text that was the AI’s reply. Someone said “Write it like this,” and the AI said “Sure, I’ll do that” — and that got uploaded along with the writing. So apparently there’s been an uproar recently because of things like this.
But the novel at the center of this uproar is apparently one that’s in the rankings. So I don’t know how much help it actually got, but the fact that a novel written by a person using AI entered the rankings and everyone is consuming it does seem to be true. And how do people describe this? I learned this expression for the first time too. They say it “tastes metallic.” “It has the taste of AI-written text. It tastes metallic.” That’s how they put it, so in the end the human evaluators do keep catching it.
So I’d guess the ideas and the evaluation are being done by the novelist, and right now things like implementation and the actual writing are probably being done largely by these models, that’s what I think.
Chester Roh Exactly. What used to be consuming statically completed novels — things like webtoons or Novelpia — the shift from those services to something like Zeta is probably for this reason. “If it’s going to be like this anyway, shouldn’t it go dynamic, tailored to my own taste?” And so users seem to be going all the way over to that side.
Jonghyun Park In the future, the content everyone consumes will probably all be slightly different. We’ll likely be consuming content custom-generated for us, generated in real time. So to state the conclusion, it all seems to tie back into distillation. models are pouring out right now, and one of the reasons for that pace seems to be this distillation, and it’s just too easy to copy each other.
The distillation race chasing the frontier 68:06
Seungjoon Choi So riding the frontier’s train is the current pattern. Certainly, once the frontier pushes ahead a bit, distillation becomes possible.
Chester Roh The latecomers just follow quickly. To put it nicely, we call it distillation, but it’s copying. Copy. But couldn’t this copying be the algorithm of evolution? Korea also beat Japan by so-called copying, distilling, and China distills both Korea and Japan, and everyone levels up that way.
Seungjoon Choi It’s not just one factor — looking at GLM this time, whether it’s Nathan Lambert or just the blog Zhipu AI published, they say they improved performance with post-training. So it’s not only distillation — everything being researched at the cutting edge does seem to be mixed in there too.
AI adoption concerns shifting to organizational operations 69:14
Chester Roh For us too, continuously, this discussion keeps drifting into the humanities, and that seems to be the trend of the times.
Jonghyun Park Right. These days my head, too, is filled with nothing but those thoughts. How can we use these things better? But the place I’m looking for the answer isn’t where engineering knowledge is gathered — it’s in places where knowledge like organizational management is gathered that I’m trying to find the answer.
Chester Roh But right now, everyone, all of us are going crazy right now.
Jonghyun Park It really feels like my brain is burning up.
Chester Roh The feeling of your brain burning up. That’s accurate. Things really change so fast, and the current incentive structure is clearly set, right? Anyway, when the frontier advances, it’s a game of who can use the delta from that advance faster and more, so we’re busy.
Closing with a preview of next week 70:07
Seungjoon Choi Anyway, there will be something going on next week too.
Chester Roh Next week, there will be something happening next week as well. Then, for today, at about this point we’ll wrap up our recording. Thank you for all the lessons today. Thank you.
Seungjoon Choi Yes, thank you.
Jonghyun Park Yes, thank you.