逐字稿图文版:How always-on agents work
Lee Robinson · Stanford CS146S · 2026 年 10 月 8 日 · 约 45 分钟 · 原帖 x.com/leerob · ← 返回中文总结
说明:英文原文为机器转写(本地 Whisper small 模型)后人工校对拼接的逐字稿,基本保持原话,仅去掉个别乱码;方括号 [ ] 内为整理者补全或推测的词。中文为整理者翻译(译),力求忠实,非官方译文。人名/产品名可能被误识别(如 Grok Bot 被识别为 “Grockba/Grapbot”),译文中已注明。时间戳点击可得到段落链接。
开场:你的 agent 和“一直在线”的 agent0:00

So as mentioned, I'm Le. We're going to talk a little bit today about how this new category of. Proactive, persistent agents work and kind of give you some of the understanding of the building blocks for how these are made.
译正如介绍的,我是 Lee。今天我们要聊一聊这类新的、主动的、持久运行的 agent 是怎么运作的,带大家理解构建它们的基础组件。

Just as a, you know, fun statement up here. Some of the things we'll talk about are in progress. Yeah. Space X. And in week one, you all built this very kind of simple
译先放一个有点好玩的声明:我们会聊到的一些东西还在开发中。是的,SpaceX。第一周你们都做过这个非常简单的

kind of simple harness that helps you understand the core idea, which is that you have a model, it can call tools in a loop, and that's actually very powerful. It unlocks a lot of new capabilities for the models, but it does come with some drawbacks. Notably, you close your laptop and the process stops. It's no longer running. And I want to talk about how you can take this simple idea, that is very powerful for how to get model models and extend it
译非常简单的 harness,它帮你理解核心思想:你有一个模型,它可以在循环里调用工具,这其实非常强大。它为模型解锁了很多新能力,但也有缺点。最明显的是:你合上笔记本,进程就停了,不再运行。我想讲讲怎么把这个对模型来说非常强大的简单想法扩展开来,

to this new type of agent, which is an always-on agent. So rather than just running in your terminal, this can run on a server. It can run in the cloud. You can close your computer, it can restart, it can survive crashes for example. This agent has its own computer. So not just your laptop but a Linux machine in the cloud where it has its own desktop, files, browser, and more. It can save files on the machine. It can understand more about your [progress] — and you need to know and understand more about your progress, and how you work with it. Unlike a harness in the terminal that is responding to every message that you send, this type of always on agent is more like a colleague where it doesn't always immediately send you a message in the background a little bit and only come back to you when there's an important update. And it can be triggered more than just you typing. So it's not just you prompting it in the terminal. Maybe there was a Slack message or an email or some other type of event that wakes up this agent to go do work. And ideally the hope with the type of always-on agents is that they can do work over very long periods of time. Not just hours or days, but maybe even weeks or months. And that just requires a little bit of a different architecture
译变成一种新型的 agent,也就是一直在线(always-on)的 agent。它不只是在你的终端里跑,而是可以跑在服务器上、跑在云端。你可以合上电脑,它能重启,比如能从崩溃中恢复。这个 agent 有自己的电脑——不是你的笔记本,而是云端的一台 Linux 机器,有自己的桌面、文件、浏览器等等。它可以在机器上保存文件,能更多地了解你的……——你需要更多地了解你的进展、以及你如何与它协作。和终端里那种你发一条它回一条的 harness 不同,这种一直在线的 agent 更像一位同事:它不会总是立刻给你回消息,而是在后台做一会儿,只有在有重要进展时才回来找你。而且触发它的不只是你打字,不只是你在终端里给它发提示。可能是一条 Slack 消息、一封邮件或其他某种事件唤醒这个 agent 去干活。理想情况下,我们希望这类一直在线的 agent 能在很长时间里持续工作——不只是几小时或几天,甚至是几周或几个月。这就需要一种稍有不同的架构,

which we'll get into. So we're going to go decently in depth on how we built this architecture, how we built the harness, kind of what's got us here today and where we're going in the future, and we'll get into some of the specifics of how we've engineered this system to be proactive, to be persistent, and to follow instructions very well. And of course time at the end for questions if you all want to ask anything about this or anything outside of what we've covered today. first how we got here.
译这些我们之后会讲到。我们会比较深入地讲我们是怎么搭建这个架构的、怎么做这个 harness 的、是什么把我们带到了今天、未来往哪走,还会讲一些具体的工程细节:我们如何让这个系统做到主动、持久、并且很好地遵循指令。当然最后也留了时间提问,关于今天的内容或其他话题都可以。首先,我们是怎么走到这一步的。
一、我们是怎么走到这一步的2:39

There's kind of been four eras of developers working with AI so far, I would say. The first one was copy-pasting code out of chatGPT and putting it into your editor of choice, which feels like decades ago at this point, but it was really only a few years ago. And then, once the model started to get better at being able to have kind of that simplified the hardest we talked about in the last lecture, they were able to edit files, run shell commands, kind of have this very simple interface for how you worked with them in the terminal. And this really took off because all of a sudden, the model had this new capability where anyhow do more proactive. Work for you and be more productive. Then this third era was, taking the agents and where they were kind of simple to giving them these dedicated apps. And at this point people were starting to run many agents in parallel, where they have their own computer. They have their own memory. They learn new skills, on the job, and they keep working after you go to bed or go do other work. A good example of this is we've kind of moved away from this world of prompting an agent every single day, although we still do that sometimes. But you can set up agents where it will just run in the background. So maybe in the before times it was you would prompt an agent, you would wait, prompt an agent, you would wait. And if it asked you to run a shell command, you'd have to approve or disallow that command.
译我认为到目前为止,开发者与 AI 协作大概经历了四个时代。第一个是从 ChatGPT 里复制粘贴代码,放进你喜欢的编辑器——现在感觉像是几十年前的事,但其实只是几年前。后来模型变得更强,能用上我们上节课讲的那种简化的 harness,它们可以编辑文件、运行 shell 命令,在终端里用一个非常简单的界面和它们协作。这一下子火了,因为模型突然有了新能力,可以更主动地为你干活、让你更高效。第三个时代,是把原本很简单的 agent 放进专门的应用里。这时人们开始并行运行很多 agent,它们有自己的电脑、自己的记忆。它们在工作中学习新技能,在你睡觉或去忙别的事情时继续工作。一个好例子是,我们已经告别了每天都要给 agent 发提示的那个世界——虽然有时我们还会这么做。但现在你可以设置 agent 让它在后台运行。以前可能是你给 agent 发提示、等待,再发提示、再等待;如果它要求运行一条 shell 命令,你得批准或拒绝。

Now, instead of kind of queuing commands, you can steer the agent. And you might say, hey, go check this thread, and then you can have, oh, by the way, you know, actually go do that on Thursday. And prior generations of models would get kind of confused when you would immediately interrupt it and give it some other command. But now they kind of understand just like how you might ramble on talking to a friend and they can kind of parse that these two things are related. The models and the harnesses can now handle that. And if you say, hey, go send this email, go send this Slack message, there are additional safety checks in place, which a lot of companies are calling auto review, which is essentially a separate AI model that's checking all of the commands that get run. So funnily enough, that's actually more secure than the fatigue of having to watch hundreds of shell commands and approve each one. We've reached the point where if you did a blind test with what the AI model is checking on the commands and what a human is checking on the commands, most of the time the model is actually doing a better job.
译现在,你不再是把命令排队,而是可以“掌舵”agent。你可能会说:“嘿,去看看这个帖子”,然后又补一句:“哦对了,其实那件事改到周四做吧。”上一代模型在你立刻打断它、给它另一条命令时会有点糊涂。但现在它们能理解,就像你跟朋友絮絮叨叨说话一样,它们能分辨出这两件事是相关的。模型和 harness 现在都能处理这种情况。如果你说“去发这封邮件、去发这条 Slack 消息”,还会有额外的安全检查,很多公司把它叫做 auto review(自动审查),本质上是另一个 AI 模型在检查所有要执行的命令。有意思的是,这其实比让人疲惫地盯着几百条 shell 命令、一条条批准更安全。我们已经到了这样一个阶段:如果做盲测,比较 AI 模型检查命令和人检查命令,大多数时候模型其实做得更好。

Kind of of paradoxically, but as the models have gotten better, the interface has gotten quite a bit more simple actually. And this type of always-on agent, we're calling a bot, but as you'll notice, there's actually quite a few of these new types of products popping up right now and gaining a lot of traction. And I-I think the main reason is because we all know how to text. You know, everyone knows how to text. Everyone is pretty proficient with using their phones. And we've kind of distilled down this way of working with AI that feels like you're just texting a friend, but it can actually go and do real work for you. It's not just answering questions in a Q&A style setting anymore. actually going and doing longer tasks like you would hand off to a co-worker for example. So there's been a few changes specifically in the models that have enabled this to happen.
译有点矛盾的是,随着模型变强,界面反而变得简单多了。这种一直在线的 agent,我们叫它 bot,你会发现现在冒出了不少这类新产品,而且很受欢迎。我认为主要原因是我们都会发消息。每个人都会发短信,每个人都很会用手机。我们把与 AI 协作的方式提炼成了一种感觉:就像在给朋友发消息,但它真的能去替你干实事。它不再只是问答式地回答问题,而是真的去做更长的任务,就像你交给同事去做一样。所以模型本身有几项变化让这成为可能。
二、模型到底变了什么6:21

I'll go over six but there's you know quite a bit more. The first one is that the models just over the past year or two years have really gotten better at being able to follow instructions over a longer period of time where they don't get confused or make mistakes and now models can pretty confidently run for hours or days with the right compression and and compaction. And a little bit. Two, models now can essentially hand off to helpers. And those helpers are called subagents. And the reason reason why a subagent is helpful is the model has kind of, it's working memory of the process that it's on. And that's going to fill up eventually. So it's nice if it can hand off to these helpers that have a fresh set of working memory. that effectively can give it an unlimited context, basically, where it can go and do work that might be expensive in terms of the tokens it's using or the tools that it's calling, and then come back to the main agent and report its updates, and I'll have some diagrams of that later.
译我讲六项,但其实还有很多。第一,过去一两年里,模型在长时间遵循指令方面进步很大,不会糊涂或犯错;现在配合恰当的压缩(compaction),模型可以相当可靠地连续运行几个小时甚至几天。第二,模型现在基本上可以把任务交给帮手,这些帮手叫子 agent(subagent)。子 agent 有用的原因是:模型对当前进程有一份“工作记忆”,它最终会被填满。所以如果能交给拥有全新工作记忆的帮手,就很好——这实际上相当于给了它无限的上下文:帮手可以去做那些在 token 或工具调用上很昂贵的工作,然后回到主 agent 汇报进展。后面我会放一些示意图。
Models now not only are just better at calling tools, I mean just a few years ago, they basically couldn't call tools very well and they would hallucinate all the time. It's funny, you don the people talk about hallucinations anymore because it's kind of a solved problem, mostly. But not only can they call tools, very powerful tools, the same way as if you onboarded somebody onto your team so obviously running shell commands and files but also you give them a full computer a full Linux computer you allow them to make demo videos and such they can click around on a screen just like human would and they're getting, increasingly better at doing that otherwise-known as computer use which basically means that any SAS product or any website that didn't expose an an API or an MCP server can be used headlessly, basically. So autonomously through computer use by these models and these bots. You can store all of the information in files. And the reason you can do that is because it's actually very easy for the models then to recall and search this information, which we'll talk about.
译现在的模型不只是更会调用工具——几年前它们基本调不好工具,总是产生幻觉。有意思的是,现在大家都不怎么谈幻觉了,因为这基本上算是解决了。它们不仅能调用工具,而且是非常强大的工具,就像你给团队新人做入职一样:当然有运行 shell 命令和文件,你还可以给它们一整台电脑、一台完整的 Linux 电脑,让它们做演示视频之类的。它们可以像人一样在屏幕上点来点去,而且越来越擅长,这就是所谓的 computer use(电脑操作)。这基本意味着任何没有开放 API 或 MCP 服务器的 SaaS 产品或网站,都可以被这些模型和 bot 以无界面的方式自主使用。你还可以把所有信息存在文件里,原因是模型回忆和搜索这些信息其实非常容易,我们后面会讲。
And then finally, you can just show the bot how to do some work. And it will turn that into skills or memories or routines or things that it can reuse later.
译最后,你可以直接演示给 bot 看怎么做某项工作,它会把它变成技能、记忆或 routine,之后可以复用。

So on the file thing, for example, or you might imagine that when you're working with a bot, under the hood, it might look something like this. You build up this profile or this understanding of the user, and that's going to be included in every single prompt. Maybe you like lowercase text-like messaging, and every single message should have that. Also, as you're working with these bots over time, you're building app a working log, just like if you were working with a colleague and they were writing down notes every day of the work they've done. That's done. That's all being put into file that the model can then decide to search or use later. If you say, hey nouve what did I do last week? It just can go and look that up. And the models are very good at using Linux commands, running shell commands to go grep and search and look through these files. Similarly with skills, you are essentially taking your knowledge of how to do a specific task and encoding it into a markdown file.
译拿文件这件事来说,你可以想象,当你和 bot 协作时,它底层可能长这样:你会逐渐建立起一份用户档案或对用户的理解,它会被包含在每一条提示词里。比如你喜欢小写的、像发短信一样的消息风格,那每一条消息都应该这样。另外,随着你长期和这些 bot 协作,你会积累一份工作日志,就像同事每天把做过的工作记成笔记一样。这些都被放进文件里,模型之后可以决定去搜索或使用。如果你问:“嘿,我上周做了什么?”它就能去查。而且模型非常擅长使用 Linux 命令、运行 shell 命令去 grep、搜索、翻这些文件。技能也类似,本质上是把你关于如何完成某项具体任务的知识编码进一个 Markdown 文件里。
specific process you need to do, the things that you believe make that successful or unsuccessful. That's not just your skills. There's a whole ecosystem of companies and developers writing these skills for how to send a great email or how to build a great API or many different things. For example, in Grockba, there is a whole system of plugins. So basically any tool that you want to use, there's a plugin you can install, which uses skills or servers. Then also routines, which if you're familiar with, Cron jobs. running some prompt basically on some kind of schedule which is pretty useful as well. So basically everything is computer. Part three. So let's go a little bit further into the kind of the mind of this always on agent or going
译你需要遵循的具体流程,你认为什么会让它成功或失败。而且不只是你自己的技能,有一整个由公司和开发者组成的生态在写这类技能,比如怎么写一封好邮件、怎么做一个好 API,等等。比如在 Grok Bot(原音识别为“Grockba”)里,有一整套插件系统,基本上你想用的任何工具都有插件可装,插件使用的是技能或服务器。还有 routine,如果你熟悉 cron 任务的话——就是按某种计划运行某个提示词,这也很有用。所以基本上一切都是电脑。第三部分。我们再深入一点,看看这个一直在线的 agent 的“大脑”,或者说
三、always-on agent 的内部结构10:11

behind the scenes. What does it mean for the agent to be always on? You can kind of read this in chart from left to right but it starts out with the server being a sleep, and also this computer that the agent can use being asleep. And so whether you send in a message or some kind of trigger decides to make the agent start up, which we'll talk about in a second, that's basically loading up and booting up all of the state for this agent and it's running on a server. If it decides that it needs to use a computer, maybe it doesn't use... Maybe it doesn't use the computer's browser or go click around on the computer, it can start up the computer as well. And then once you for done using it, it also has the ability then to basically put those to sleep. Why is this important? It would be expensive and also, there's not enough computers to just have those computers running all of the time for everyone as well as that's a lot of server usage too. And so this state can then be persisted to a database, the conversation logs, the memories, the routines, all of that.
译幕后发生了什么。agent 一直在线到底意味着什么?这张图可以从左往右读:一开始服务器在休眠,agent 能用的那台电脑也在休眠。无论是你发来一条消息,还是某个触发器决定让 agent 启动(我们马上会讲),都会加载、启动这个 agent 的全部状态,它运行在服务器上。如果它决定需要用电脑——也许它不需要用电脑上的浏览器或去点来点去——它也可以把电脑启动起来。用完之后,它也能把它们重新休眠。为什么这很重要?因为成本很高,而且也没有足够多的电脑可以给每个人一直开着,那样服务器用量也很大。所以这些状态可以持久化到数据库里:对话日志、记忆、routine,所有这些。

How does the agent get woken up then? Well, obviously you can prompt it directly, But I think what's increasinglyly interesting is having other things running in the background that can then decide when you want to go reach out to your bots. Maybe that is somebody pinging you on Slack and your bots can live in Slack. Maybe that's a phone call or a meeting that you put your bot in. It's increasingly common now at SpaceX where we see bots joining meetings for employees where they could make it to the meeting but they want their bot to go and just take notes and come back to them, which is kind of funny. Maybe you had some routine or some background work that was taking hours or days to do, and then you want to come back and give you a summary. That all happens. Then we have basically a router that dedupes that, and figures out how to put it into the conversation. And then it has this heuristic of, do I need to bug the user about this or not? Because not everything is actually important. Just like if you have a coworker working on something, it would be pretty annoying if every single time they made progress on something, they came and bugged you, like, hey, check it out, I made progress on this thing. You really only need a status update. Maybe, I don't know, every day, depending on the task.
译那 agent 是怎么被唤醒的?当然你可以直接给它发提示。但我觉得越来越有意思的是,有其他东西在后台运行,它们可以决定什么时候去联系你的 bot。可能是有人在 Slack 上 @ 你,而你的 bot 可以住在 Slack 里;可能是一通电话,或者你让 bot 参加的会议。现在在 SpaceX 越来越常见的是 bot 替员工参加会议——员工本来能去开会,但他们想让 bot 去做笔记再回来汇报,挺好笑的。也可能是你有某个 routine 或后台工作,要花几小时或几天,做完后你希望它回来给你一份总结。这些都会发生。然后我们基本上有一个路由器来去重,并决定怎么把它放进对话里。接着它有一个启发式判断:这件事需要打扰用户吗?因为并非所有事都真的重要。就像同事在做某件事,如果他每取得一点进展都跑来烦你,“嘿,看看,我这件事有进展了”,那会很烦。你真正需要的只是一个进度更新,也许每天一次,取决于任务。

So I talked about how the agent has this computer, this Linux box, this is running a firecracker virtual machine and the reason this is really helpful is because it's very easy then to essentially snapshot the state of the machine and and shut it down and start it back up. And on this machine you of, of course, have files, but you can kind of connect into the machine and have a live view of the machine and actually control it yourself if The machine can of course run commands, which is great if you're going to run some shell commands. Maybe you don't want it on your personal laptop. You want it in this other environment that is secured and locked down. And then another interesting thing is of course it has Chrome and it can use a browser, but But you can also make it essentially have virtual desktops. So if you're running five bots, 10, 10 bots, and it needs their own browser, you can essentially let them have different screens. And all of this, again, is kind of a shared infrastructure between, these bots where they can go call on these computers.
译我讲过 agent 有这台电脑、这台 Linux 机器,它运行在 Firecracker 虚拟机里。这很有帮助,因为这样很容易给机器状态做快照、关机、再启动起来。这台机器上当然有文件,你还可以连进去看机器的实时画面,甚至自己操控它。机器当然可以运行命令,如果你要跑一些 shell 命令,这很好——也许你不想在自己的笔记本上跑,而是想在另一个安全、锁定的环境里跑。另一个有意思的点是,它当然有 Chrome、可以用浏览器,但你还可以让它拥有虚拟桌面。所以如果你跑五个、十个 bot,各自需要自己的浏览器,你可以让它们拥有不同的屏幕。所有这些也是 bot 之间共享的基础设施,它们可以调用这些电脑。
So we talked about the triggers that will wake up the agent. That's kind of on the left, so different apps or Slack messages or emails. Then you. You have this middle piece, which is the server, and then on the computer that we just talked about where you can run files and commands and other things. So let's talk a little bit more about the server. I already mentioned commands will come in, they get put into this queue, and then it calls the purple box, which is a agent loop. And this is what we covered last week. This is a model calling tools in a loop and then when you shut it down you have the state to the database or storage. So it's interesting, I think, to look at this and realize there's this one small piece that can then have all of these abstractions built on top to create a very powerful product when you give it all of these helpful tools and you connect it to all of these different services. The reason the agent runs on the server is a pretty common practice in computing where you want [to]
译我们讲了会唤醒 agent 的触发器,就是左边这些:不同的应用、Slack 消息、邮件。然后中间这部分是服务器,再往右是我们刚讲的电脑,你可以在上面操作文件、运行命令等等。我们再多讲讲服务器。前面说过命令进来后会被放进队列,然后调用紫色的框,也就是 agent 循环。这就是我们上周讲的:模型在循环里调用工具;当你关掉它时,状态会存到数据库或存储里。我觉得有意思的是,看着这张图你会意识到:就是这么一小块东西,可以在上面搭起所有这些抽象,当你给它各种有用的工具、把它连到各种服务上时,就能做出一个非常强大的产品。agent 跑在服务器上,是计算领域一种很常见的做法:你希望
have a thin client and a thick server and computing is kind of oscillated between these things over the years. But for example, a few good reasons for this is let's say you need to push an update to the agent code. It would kind of suck if you had to push an app store update just to get that code out there versus just pushing a bunch of server updates very quickly. Also sometimes you need to do heavy computation work like transcription or image generation or something. That's probably going to be faster on a server than on device, depending on what the devices are. And there's very established ways that we know how to scale servers, especially with variable traffic and variable growth. But the most critical thing here is that if you put the agent loop on the virtual machine, then you can't shut down the virtual machine independently from the
译采用瘦客户端、厚服务器——计算这些年一直在这两者之间来回摆动。举几个好理由:比如你需要推送 agent 代码的更新,如果每次都得发一个 App Store 更新才能把代码推出去就太糟了,相比之下推一批服务器更新快得多。另外有时你需要做繁重的计算,比如转录或图像生成,这在服务器上通常比在设备上快,取决于是什么设备。而且我们已经有非常成熟的方法来扩展服务器,尤其是面对波动的流量和增长。但这里最关键的一点是:如果你把 agent 循环放在虚拟机上,那你就没法独立于 agent 服务器去关闭虚拟机。

agent server. Yeah, you want to save everything on the server so you can always kind of rebuild that virtual machine. For example, let's say you need to do some critical zero day Linux update on the VM. Well, well, it would be great to be able to rebuild that and not blow away the entire agent loop and have downtime in the application. When you hit send, this is maybe a little in the weeds, so I'll just kind of gloss through it, but basically, if you want to dig more into learning more about distributed systems and thinking about strong versus eventual consistency, which are just fun topics to get into. You send a message and you kind of want the client to immediately, identify that and acknowledge it. Then when you go into the turn, you're obviously putting these tools in a loop with the model, saving that to the database. And then the clients are actually putting subscriptions with all of the servers.
译对,你希望把一切都存在服务器上,这样你总能重建那台虚拟机。比如你需要在虚拟机上做某个紧急的 Linux 零日漏洞更新,那么能重建它、同时不把整个 agent 循环搞崩、不让应用停机,就很好。你点击发送时——这部分可能有点太细了,我就简单带过——基本上如果你想深入学习分布式系统,思考强一致性和最终一致性,这些都是很有意思的话题。你发出一条消息,你希望客户端立刻识别并确认它。然后进入这一轮(turn)时,你当然是让模型在循环里调用工具,并把结果存进数据库。而客户端其实会对所有服务器建立订阅。
So they know something changed and they're listening for these signals. If they get a signal, then they can go and essentially fetch the latest information from the server, and there's some fancy deduping you can do to make sure, if there were some kind of retries, that you would not show multiple messages in the list. And so the underlying infrastructure primitive to think about here is a workflow versus a job queue. And the main difference here is with a job queue,
译这样它们就知道有东西变了,它们在监听这些信号。一旦收到信号,就可以去服务器拉取最新信息。这里还可以做一些巧妙的去重,确保如果发生了重试,列表里不会显示多条重复消息。所以这里要考虑的底层基础设施原语,是工作流(workflow)与任务队列(job queue)。主要区别在于,用任务队列时,

Let's say you're kind of going along on steps, you ask your agent to go and do some kind of complicated task. And it's making its way from step one to step two to step three. But then for some reason there was an infrastructure issue, there was some problem with a normal queue if that crashes and you go to restart it. Well now you have to replay steps one one, two and three which is obviously that could be pretty problematic. So what you really want here is you want a durable workflow which means that it's safe for you to actually retry that. And when you retry after a failure it picks right up on step three, and this is kind of a solved problem in some ways in distributed systems and infrastructure, so we use an open source tool called Temporal, which is interesting to look into if you're curious, so that we don't have to rebuild a lot of this very complicated stuff from scratch. Of course you still have to run the infrastructure for it, but [those are] some of the pieces to put together there.
译假设你在一步步推进,你让 agent 去做某个复杂任务,它从第一步做到第二步、第三步。但由于某种原因出了基础设施问题,如果用的是普通队列,它崩了,你去重启,那就得重放第一、二、三步,这显然可能很成问题。所以这里你真正需要的是持久化工作流(durable workflow),这意味着重试是安全的。失败后重试时,它会直接从第三步接着做。这在分布式系统和基础设施里某种程度上是已解决的问题,所以我们用了一个叫 Temporal 的开源工具,感兴趣的话值得研究一下,这样我们就不用从零重建很多非常复杂的东西。当然你还是得为它运行基础设施,但这些就是要拼起来的一些组件。

So let's say a new message comes in. There's different priorities based on whether I'm messaging my agents of messaging my bots versus if it's in a group chat or a Slack thread or some background routine, for example. So if I say, oh, actually change the date for this event, we actually want that agent loop to stop and my message is gonna have priority over everything else that's happening which kind of makes sense. For everybody else, you're kind of putting things into the queue in the back of the line. And if the agent gets, you know, the agent gets, you know, it's very quickly quickly, it's able to kind of batch those together, so it's just one turn, you can kind of figure out whatever heuristic you want there. And then it kind of goes through and does each turn at a time. Okay, so let's talk about the harness itself. And this is going to kind of take the basic harness that we talked about last week and add some new tools, some new concepts to it. [One]
译假设来了一条新消息。根据是我在给我的 agent/我的 bot 发消息,还是在群聊、Slack 帖子里,或者是某个后台 routine,优先级是不同的。所以如果我说“哦,其实把这个活动的日期改一下”,我们其实希望 agent 循环停下来,我的消息优先于正在发生的其他一切,这很合理。对其他人的消息,基本上就是放进队列、排到后面。如果 agent 很快收到好几条,它可以把它们合并成一批,作为一轮来处理——具体用什么启发式规则可以自己定。然后它一轮一轮地处理。好,我们来讲 harness 本身。这部分会在上周讲的基础 harness 上,加上一些新工具、新概念。一个
四、Harness:模型外面的那层“外壳”18:21

difference with the Grapbot harness and probably some of the other tools like this on the market is the way we've done the thin client fixed server architecture is there's only one tool that communicates between the client and the server and it's just send user and the great thing about this is that it's actually very architecturally simple where the client is just texting the server in some ways and you're putting a lot of these heavier tools on the server and there's many of those tools. There's, running shell commands, reading files, but also calling on a cloud agent or recording your screen or doing demos for example. And this, I think has worked pretty well because now, regardless of where you've spent up a new client, whether it's a mobile app, a server, a desktop app, an email slack. It all just has this one single tool that communicates back and forth with the server.
译Grok Bot(原音识别为“Grapbot”)的 harness 和市面上一些类似工具的不同之处,在于我们做瘦客户端、厚服务器架构的方式:客户端和服务器之间只有一个通信工具,就是 send user(发消息给用户)。这样做的好处是架构上非常简单——某种意义上客户端只是在给服务器发消息,而你把很多更重的工具放在服务器上,那里有很多工具:运行 shell 命令、读文件,还有调用云端 agent、录屏或做演示等等。我认为这个方案效果不错,因为现在无论你新起一个什么客户端——移动应用、服务器、桌面应用、邮件、Slack——都只有这一个工具和服务器来回通信。

On the client though, even though it just has this When it receives the tool, it can decide how it wants to display that in an interactive way. That makes the most sense for the user. If it gets information back from the server, that the user is trying to log in to something. It can show these unique sign-in forms and connect and connect to tools like OnePassword. If you're trying to do something that is non, you can't reverse it sending an e-mail or a Slack message, it might come back and ask you, you know, edit and approve a draft before you send it. And then sometimes it just, you know, you can just leave an emoji reaction if you don't actually need to send a message.
译不过在客户端这边,虽然它只有这一个工具,但收到工具调用时,它可以决定如何以交互的方式展示,怎么对用户最合理就怎么来。如果它从服务器拿到的信息是用户正在尝试登录某个东西,它可以展示专门的登录表单,并连接 1Password 这样的工具。如果你要做的是不可逆的操作,比如发邮件或 Slack 消息,它可能会回来让你先编辑、批准草稿再发送。还有些时候,如果其实不需要发消息,它就只留一个表情回应。

So going back to what I was saying with dynamic tool discovery. So for any model, there's a context,, which is kind of it's working memory. And that's some limit in tokens. And if you just jam every single tool possible, into that context window, you're not going to have a lot of space. And the way the tools work is you have the name of the tool, you have the schema, you have all these different details. And if you start installing your Google Drive connector your and Gmail connector, your YouTube connector, like all these different services that use MCP servers, which all have tools, you're going to start looking at something like the the top, where a lot of the blocks are already filled. And the problem with that is it doesn't give the model as much space to actually do the work of the conversation. So one kind of context engineering or harness trick that a lot of the popular harnesses now use is called dynamic context or some other folks call it the tool search tool, which is a fun name.
译回到我刚说的动态工具发现。对任何模型来说,都有一个上下文(窗口),它就像模型的工作记忆,有一个 token 上限。如果你把所有可能的工具都塞进上下文窗口,就没剩多少空间了。工具的构成是:工具名、schema,以及各种细节。如果你开始安装 Google Drive 连接器、Gmail 连接器、YouTube 连接器,这些都是用 MCP 服务器的服务,各自都带一堆工具,你就会看到像上面那张图一样,很多格子已经被填满。问题在于,这样模型就没有足够空间去真正完成对话中的工作。所以现在很多流行的 harness 都用的一种上下文工程或 harness 技巧,叫动态上下文,也有人叫它“工具搜索工具”(tool search tool),名字挺好玩。
But basically you're only putting the names of the tools into every message that gets sent to the model. And then the model can decide, okay, I actually wanna go use this tool, I can go read the information from a file. It's kind of files in a file system all the way down for all of these things. And that just saves on your context usage, which then also helps save on cost, save on usage, just better for everybody.
译但基本上,你只在发给模型的每条消息里放工具的名字。然后模型可以决定:好,我其实想用这个工具,我可以去从一个文件里读取它的信息。所有这些东西归根结底都是文件系统里的文件。这样就节省了上下文用量,也节省成本和用量,对大家都更好。

In general, the principle here really is to try the cheapest option or the most reliable option first and kind of work your way up the ladder here. So the easiest one is just what the agent or what the bot already knows based on the conversation history. If it doesn't, if it's not in a immediate context. Maybe there's a plugin, an MCP server, some kind of API that it can go and pull your banking information from PLAD, for example. Maybe that doesn't work and it needs to go do some kind of web search to see what was the score, score of the Cubs game last night. Maybe if it doesn't have that, it actually needs to go to the computer and open up the browser and do some searches on there there. And finally, like maybe it needs to go and actually use the entire desktop, so run scripts or run commands or build things on that Linux computer. And only if all of those fail, then go bug the user and ask them to go give input or approve something. And this really helps it feel like what it would be like to work with a colleague is like they're not going to bug you for all these different things. Ideally they're going to try a lot of these options first.
译总的来说,这里的原则是先尝试最便宜或最可靠的选项,然后沿着阶梯一级级往上。最简单的就是 agent/bot 根据对话历史已经知道的东西。如果不在当前上下文里,也许有插件、MCP 服务器或某个 API 可以用,比如从 Plaid(原音识别为“PLAD”)拉取你的银行信息。如果不行,也许需要做网页搜索,看看昨晚小熊队(Cubs)比赛的比分。如果还没有,它可能需要去电脑上打开浏览器搜索。最后,也许它需要真正使用整个桌面,在那台 Linux 电脑上跑脚本、运行命令或构建东西。只有在这些都失败时,才去打扰用户,请他提供输入或批准某件事。这真的让它感觉像和一位同事共事——他们不会为了各种事情来烦你,理想情况下会先尝试很多选项。
On that computer, so the Linux machine that the agent is able to use, there's kind of three different ways that it can retrieve information from it. Of course it can just take a screenshot of the screen, but there's more efficient ways of doing this. It could also look at the HTML of the page, but you can also look at the accessibility tree, which is a nice way of like cutting down a lot of the token usage and basically the buttons or links or other elements that you want the browser to click on. Of course, if none of that works, you can just click on pixels. So some websites, for example, have modals or dialogues that pop up that you need to click. And so your model, your agent needs to have that functionality as well. as well. Or again, kind of worst case on that ladder is you go back to the user and like,, oh, there's a two factor auth code or some kind of capture or something like that.
译在那台电脑上,也就是 agent 能用的那台 Linux 机器上,它大致有三种方式获取信息。当然它可以直接截屏,但还有更高效的方法。它也可以看页面的 HTML,还可以看无障碍树(accessibility tree),这是一种很好的方式,可以省下大量 token,基本上就是你想让浏览器点击的按钮、链接或其他元素。当然,如果这些都不行,它可以直接点击像素。比如有些网站会弹出需要点击的模态框或对话框,所以你的模型、你的 agent 也需要具备这种能力。再说一次,阶梯上最坏的情况是回去找用户说:“哦,这里有个双因素验证码,或者某种验证码(CAPTCHA)之类的。”

So the idea is that kind of putting the pieces together. You have this main agent and it's kind of like your orchestrator. And the main agent, you want to keep the conversation as sparse as possible because it now has the ability to go call on all of these different helpers. And the helpers are just sub-agents. And so this is helpful because if you hand it off to a helper to go run a bunch of shell commands or do a lot of expensive tool calls, all of that work that ultimately is ending up in some answer to a question or a task, none of that is bloating the conversation in the context of the main agent. It's only coming back with the result. And this is how you get to a product that feels like, like it has infinite context. Obviously, their infinite context is not a real thing. Every model has a context limit. But there are these tricks and workarounds such that, especially when you train models to get better at this, that it doesn't feel like you're running up against the limits of the context window.
译所以整体思路就是把这些拼起来。你有一个主 agent,它有点像你的调度者(orchestrator)。对主 agent,你希望对话尽可能精简,因为它现在可以调用所有这些不同的帮手,帮手就是子 agent。这很有用:如果你把跑一堆 shell 命令、做大量昂贵工具调用的工作交给帮手,那么所有这些最终得出某个问题答案或完成某个任务的过程,都不会撑大主 agent 的对话和上下文,回来的只有结果。这就是你如何做出一个感觉像拥有无限上下文的产品。当然,无限上下文并不真实存在,每个模型都有上下文上限。但有这些技巧和变通办法,尤其是当你训练模型更擅长这些时,就不会感觉自己撞上了上下文窗口的上限。
And you don't feel a degradation of the model's quality as you get further along that as well. So we have these different helpers. Each with different skills. Each one of these is its own durable workflow. So going back to the cues versus workflows thing. That means if, if it fails, the video helper crashes, something, you know, uses up too much memory, or something,, that I can restart and keep going without losing progress along the way. And then the main agent, it is kind of like a coworker, where it is managing a to-do list. You have a few things you can do, you can go check on it, you can send it a message, you can stop the helpers for example. And so a lot of this is modeled after the, So, just a few little interesting tidbits from building this out on the team. I think we'll talk a little bit more about prompt caching here, but ideally you don't add or remove new tools during turns.
译随着进行得越久,你也不会感到模型质量下降。所以我们有这些不同的帮手,各有不同的技能,每一个都是自己的持久化工作流。回到队列与工作流的对比:这意味着如果它失败了——视频帮手崩了,或者某个东西用了太多内存之类的——它可以重启、继续,而不会丢失进度。主 agent 就有点像一位同事,它在管理一份待办清单。你可以做几件事:去查看进展、给它发消息、比如停止某些帮手。很多这些都是模仿……[此处原音不完整]接下来是我们团队在搭建过程中的一些有趣小经验。我想之后会再多讲一点 prompt caching,但理想情况下,你不要在一轮进行中增加或删除工具。
And just like I mentioned with the slide with all of the boxes, if you have extremely long tool descriptions or schemas that's gonna take a lot of space in the context window, which is bad for pricing, it's bad for usage. you want to minimize that. So there's plenty of optimizations you can make here on the engineering side to just trim that down. Another interesting one is like you might hear people talk about how, oh, like we can have a harness that just has one tool. And that one tool is, like, running shell commands. And this is true and you can model a lot of things this way. But if you see the model doing, you know, 90% of its time doing shell commands for this one thing, it actually can be more efficient to make a dedicated tool for some of that stuff. And it's a bit more legible for people actually reviewing how the product is working. And the last one's kind of funny too. Obviously, if you're writing errors for humans, it's helpful to make it very descriptive. Like, okay, this thing failed. Like, click on this link or here's what you need to do next. Turns out, that's also very helpful for agents.
译就像我在那张满是格子的幻灯片上提到的,如果工具描述或 schema 特别长,会占用上下文窗口大量空间,对价格不好、对用量也不好,你要尽量减少它。工程上有很多优化可以做来精简它。另一个有意思的点是,你可能听人说:“哦,我们可以做一个只有一个工具的 harness,这个工具就是运行 shell 命令。”这是真的,很多事情都可以这样建模。但如果你看到模型 90% 的时间都在为某件事跑 shell 命令,那么为其中一些事专门做一个工具其实会更高效,而且对真正审查产品如何运行的人来说也更易读。最后一点也挺好笑:显然,如果你给人写报错信息,写得非常具体会很有帮助,比如“这个失败了,点这个链接,或者下一步该这样做”。结果发现,这对 agent 也非常有帮助。
So there is some DX in designing good error messages for agents too. And when you're trusting your bots to go and do work over a long period of time, it is very important that you trust them to be secure with your information, to be secure with how it operates. So even working on its own separate machine, not polluting anything on your machine, still you would like it if a model reviews every shell command before it goes. So auto review. It would be great if when these untrusted inputs, on just making sure that that is properly protected from prompt injections. For these kind of non-reversible things like sending emails, you want to really ask the user first.
译所以为 agent 设计好的报错信息,也有一些开发者体验(DX)上的讲究。当你信任 bot 长时间去干活时,非常重要的是你能信任它们安全地处理你的信息、安全地运行。所以即便它在自己独立的机器上运行、不会污染你机器上的任何东西,你仍然会希望有一个模型在每条 shell 命令执行之前审查它,也就是 auto review。对于不可信的输入,最好确保它们得到妥善防护,避免提示词注入。对于发邮件这类不可逆操作,你真的要先问用户。
And then also, you can build in some workflows into the product to keep passwords and other things outside of the model conversations.
译另外,你还可以在产品里内置一些流程,让密码等信息不进入模型的对话。
五、上下文工程27:26

OK. Context engineering is also known as harness engineering. I feel like it's kind of the same thing. They've evolved over time. Basically, what are some tips and tricks we can do to minimize the amount of context that's being used in these harnesses, which is good for efficiency, for reducing costs, and other reasons? I talked about prompt caching. This is a good example of why it's really important. So if you think about when you're working with a coding agent or any agent, most of the tokens used are actually ideally cached tokens, because you're sending it a conversation, and then you're doing a change, you're adding one more message, and then you're basically resending the entire conversation over again., so if you look at the pricing for models, there's huge discounts, of course, on if a token is cached, and you wanna keep them in the cached window as long as possible. and so a way to do this is you try to keep the tool definitions, the a system prompts-prompts, all this information as static as possible between different calls so that you don't break the cash and then have to pay for the uncached tokens every single time.
译好。上下文工程也被称作 harness 工程,我觉得基本是一回事,它们随时间演变而来。基本上就是:我们有哪些技巧可以减少这些 harness 使用的上下文量,这有利于效率、降低成本,以及其他原因。我讲过 prompt caching,这是一个很好的例子,说明它为什么重要。想想你在用编程 agent 或任何 agent 时,理想情况下大多数 token 其实都是缓存过的 token,因为你发出一段对话,然后做一个改动、加一条新消息,然后基本上把整段对话再重新发一遍。所以如果你看模型的定价,缓存命中的 token 当然有很大折扣,你会希望它们尽可能久地留在缓存窗口里。做法之一是,尽量让工具定义、系统提示词等所有信息在不同调用之间保持静态,这样就不会打破缓存,不必每次都为未缓存的 token 付费。
So if you have that fixed part atop, then when you get down to the bottom and you only have the newest message that comes in, that's the only delta that you're paying the uncached tokens for. A whole art and science to this as well. Basically because that main agent in the most ideal sense, this is the one you're having this very, very, very long conversation with that's compacting or compressing, so, summarizing the context, many, many times over, without you even noticing sometimes, which we'll show an example of. So you really wanna keep that context small. Some ways of doing that is anytime there's a very long output rather than stuffing that in the context, you can just put it in a file, and the agents are very good at reading files. So this actually works pretty well, which is just is just like reading a skill or reading a memory for example. If you give the model a screenshot or some other very expensive work you just delegate that out to a helper again, so it goes to the sub agent and and it doesn't affect the main agent's, context and then also in an ideal world you want to kind of be always gardening and trimming and removing unnecessary stuff from this conversation from this contact window as you're going both for cost but for also recall and just making sure the model's performing well.
译所以如果顶部是那段固定的部分,等到底部只有最新进来的那条消息时,那就是你唯一需要为未缓存 token 付费的增量。这里面也有一整套艺术和科学。基本上,因为在最理想的情况下,主 agent 就是你一直与之进行非常非常长对话的那个,它会压缩、总结上下文很多很多次,有时你甚至都察觉不到——我们会展示一个例子。所以你真的要让那个上下文保持小。一些做法是:每当有非常长的输出,与其把它塞进上下文,不如直接放进文件,agent 很擅长读文件。所以这其实效果很好,这就像读取一个技能或一条记忆。如果你给模型一张截图或其他非常昂贵的工作,就再把它委派给一个帮手,交给子 agent,这样就不影响主 agent 的上下文。另外,理想情况下你要一边进行一边不断“修剪园子”,从对话、从上下文窗口里修剪、移除不必要的东西,既为了成本,也为了回忆能力,确保模型表现良好。

So if there's a cache miss, it's kind of like a bug in the system, really. And if you go all the way back to January of this year when OpenClaw was really getting popular and people were using it inside of subscription plans for model providers, this was a very new usage pattern that model providers hadn't really seen. the harness hadn't really yet been optimized for this type of workflow. So there was a lot of cache misses and something like this where agents are running all the time. That's obviously problematic. So lots of fixes had to go in to make that robust to keeping the prompt cache as much as possible. So a fun trick here is just trying to keep everything in the prompt in order such that even if you deploy a new change in the system prompt or add a new tool.
译所以如果出现缓存未命中,那其实有点像系统里的 bug。回到今年一月,OpenClaw 正火的时候,人们在模型服务商的订阅套餐里使用它,这是模型服务商从未见过的全新使用模式,harness 也还没为这种工作流做优化。所以出现了大量缓存未命中,而在 agent 一直运行的场景下,这显然是个问题。于是必须做很多修复,让它稳健地尽可能保持 prompt 缓存。这里有个好玩的技巧,就是尽量让提示词里的一切保持顺序,这样即便你在系统提示词里部署了新改动,或者加了一个新工具,
You can almost hash each section and always keep a fixed order. Like let's say you're doing an A-B test and you want to put a new piece of the prompt in there. You really don't want that to screw up the prompt order for everybody else. Otherwise you just busted the cache and then that's problematic. Another fun one is when you open the chat and you start typing, the server can actually go and prepare the prompt early to basically warm the cache almost like if you're on a website and you go to hover on a link it can so there's like tons and tons of little engineering optimizations that compound how far these models can go in terms of token usage and how smart they can be when you use them for very long periods of time
译你几乎可以对每个段落做哈希,并始终保持固定顺序。比如你在做 A/B 测试,想往提示词里加一段新内容,你绝对不希望这把其他所有人的提示词顺序搞乱,否则你就把缓存打破了,那就成问题了。另一个好玩的是:当你打开聊天窗口、开始打字时,服务器其实可以提前去准备提示词,基本上就是预热缓存,有点像你在网站上把鼠标悬停在链接上时,它可以(预加载)……所以有非常多这样的小工程优化,它们叠加起来,决定了这些模型在 token 用量上能走多远,以及长时间使用时能有多聪明。

and the key thing is that when you're working over a very long period of time you essentially have to get very good at compaction or summarization because let's say the model has a 200,000 token context window, you're going and going and going, eventually it's gonna hit a point where it needs to summarize, and all compression is lossy. So you have to figure out the best algorithm to do the compression compression and not lose important details. And there's, I'm sure, entire PhDs just for this problem. Like it's very, very tough, tough to do, right? But just to show an example, like let's say you're going back and forth with your bot, you have over 100,000 tokens, And then you go and do something else. Well, there is this time period where model providers have like a cash window when the cash is still warm. Maybe it's 10 minutes, maybe it's 60 minutes. One trick that we do, which I think is really helpful, is, is still within the warm cash period, you should probably do the compaction right now.
译关键在于,当你长时间工作时,你本质上必须非常擅长压缩或总结。因为假设模型有 20 万 token 的上下文窗口,你一直用、一直用,最终会到达一个需要总结的点,而所有压缩都是有损的。所以你得找到最好的算法来做压缩,同时不丢掉重要细节。我敢说光是这个问题就有人在读整个博士。这非常非常难做,对吧?但举个例子:假设你和 bot 来回聊,已经超过 10 万 token,然后你去做别的事了。模型服务商有一个缓存窗口期,缓存还是热的,可能 10 分钟,也可能 60 分钟。我们用的一个技巧,我觉得非常有用:趁缓存还热,你大概应该现在就做压缩。
Because you have this long conversation, it's going to be much cheaper. So just go ahead, do the compaction. And now when the user comes back in 30 minutes or whatever, they're starting from this summary. And the summary now is going to be much cheaper for them to continue adding messages on. So lots of work to do there to make the algorithm very high quality for how you do summarization. When you talk about compaction and what information you choose to include in the agent's memory, I kind of think about it in a few different layers. The first one is facts about the user that are very important to be included every single time. Going back to earlier when I said, oh, the user always wants to have lowercase text. That probably needs to be in every single message, otherwise it would be really weird if one message now has uppercase letters, if you're going full Sam Altman style. Then you probably also have some facts that you want to include in there.
译因为你有这么长的对话,(现在压缩)会便宜得多。所以就直接去压缩。等用户 30 分钟后回来,他们就从这份摘要开始,在摘要上继续添加消息会便宜得多。所以要让总结算法达到很高质量,还有很多工作要做。说到压缩,以及你选择把哪些信息放进 agent 的记忆,我会从几个层次来考虑。第一层是关于用户的、非常重要、每次都必须包含的事实。回到前面我说的“用户总是想要小写文本”,这大概需要出现在每一条消息里,否则如果某条消息突然出现了大写字母就会很奇怪——如果你走的是 Sam Altman 那种全小写风格的话。然后你大概还有一些想放进去的事实。
And there's many different ways you can make this algorithm to determine what are the most important facts to keep. So ideally there's this kind of garbage collection of this fact. And if you think about it, this is kind of our brains work. It's like this fact right now is important to keep. It's relevant for this conversation today. You're probably not thinking about a soccer match right now. And your brain is very good at keeping those ones in context right now and remembering them for later. We try to model some of that here where the harness has, of course, these core facts. It also has these dated logs of kind of your past conversation history, essentially. And then also these kind of scratch pad notes. So there we have just like little things that you've been working on right now that are designed to be ephemeral and they will go away, but they might be relevant just for the next 30 minutes or an hour. And everything else can be found through files. This is the magic of files of the models being good at running commands is like, they can always just go look stuff up and they're actually very good at doing that.
译有很多种方法可以设计这个算法,来决定保留哪些最重要的事实。理想情况下,对这些事实有一种“垃圾回收”。想想看,我们的大脑也是这样工作的:这条事实现在很重要,需要保留,它和今天的对话相关;你现在大概不会在想一场足球比赛。你的大脑很擅长把这些放在当前语境里,并为以后记住它们。我们在这里尝试模仿一部分:harness 当然有这些核心事实,还有按日期记录的日志,本质上是你过去的对话历史,还有一些草稿本笔记——就是你当前在做的小事,它们被设计成短暂的、会消失,但可能在接下来 30 分钟或一小时内有用。其他一切都可以通过文件找到。这就是文件的魔力、也是模型擅长运行命令的魔力:它们总能去查,而且它们真的很擅长这件事。
六、未来方向34:29

There is this flywheel of training the model to get better at the product. And this is a high level of refresher on kind of the stages of training a large language model. First you have pre-training where the model learns a general understanding of the world from large amounts of internet text or other private data. Then you have supervised fine tuning, where you're teaching the model how to behave in a certain way, maybe as a chat assistant or as an agent. And then finally you have reinforcement learning, where it's kind of like teaching the model how to play games and get better at those games. And in the context of these always on agents, it's very helpful for example, to train the model to see examples of the harness and to know how to work inside of the harness in SFT. so it can understand the tools that you have available. It's helpful to train the model to be better at instruction following during reinforcement learning.
译有这样一个飞轮:训练模型,让它在产品上变得更好。这是对训练大语言模型各阶段的一个高层回顾。首先是预训练,模型从大量互联网文本或其他私有数据中学到对世界的一般理解。然后是监督微调(SFT),教模型以某种方式行事,比如作为聊天助手或 agent。最后是强化学习(RL),有点像教模型玩游戏并越玩越好。在一直在线 agent 的场景下,比如在 SFT 阶段让模型看到 harness 的示例、知道如何在 harness 里工作,这样它就能理解你有哪些可用工具,这很有帮助。在强化学习阶段训练模型更好地遵循指令,也很有帮助。
How to understand the right skill to call when there's hundreds of skills, for example. How to click on the right spots when you're given in a browser and you need to kind of click around. And not only the training itself, but also how you have served the model, I'm sure. The inference that you do. The main agent you're talking with, you know, ideally you want to have very low latency. You want to come back to you right away, give you really fast replies, and you can tune the harness and the inference to do that. But then when you go hand it off to a helper and you're asking it to go do some work for a very, very long period of time, the difference between 10 minutes and 12 minutes is kind of negligible at that point. So that allows you to kind of control how you optimize things. The way that this kind of flywheel works is that let's say your agent gets something wrong. Ideally then that's in some kind of evaluation. The model accidentally called this tool and it wasn't supposed to.
译比如在有几百个技能时,如何判断该调用哪个技能;当你在浏览器里需要点来点去时,如何点到正确的位置。不仅是训练本身,还有你如何部署(serve)模型,也就是你做的推理。和你对话的主 agent,理想情况下延迟要非常低,要马上回复你、回得很快,你可以调 harness 和推理来做到这一点。但当你把任务交给帮手、让它去做很长很长时间的工作时,10 分钟和 12 分钟的差别到那时就可以忽略不计了。这让你可以控制如何做优化。这个飞轮的运作方式是:假设你的 agent 做错了某件事,理想情况下它会进入某种评测(eval)——模型不小心调用了这个工具,而它本不该调用。
Then on the back end you can go and fix the harness, you can fix the bug, you can fix fix the inference issue, you can check and make sure it's actually fixed in the evals and then shift the change. The teams working on these products, they're kind of doing this every single day. They're finding bugs, they're fixing it, they're measuring it with evals. That's kind of this interloop that you're always getting better at. But then there's this outer, which is like every few months or every month or however long, you're training new models. And ideally the new models are learning from all of these bugs and failures and they're kind of climbing on the evals to get better at the things that you care about. the new versions into the product and make sure that they're improving at all these different things. And the fun thing about this is when the new models come out they have new capability jumps. It allows you to use the product in new ways. There hasn't really been a lot of training data, for example, on using bots in a group chat. This is kind of an emergent thing that's being figured out, and it's important to get it right.
译然后在后端你可以去修 harness、修 bug、修推理问题,在评测里检查确认确实修好了,然后发布改动。做这些产品的团队基本每天都在做这件事:找 bug、修 bug、用评测衡量。这就是你不断进步的内循环。还有一个外循环:每隔几个月、或每个月、或多久都行,你会训练新模型。理想情况下,新模型会从所有这些 bug 和失败中学习,在评测上不断爬升,在你在乎的事情上变得更好。(然后把)新版本(放)进产品,确保它们在这些方面都在进步。有趣的是,新模型出来时会有新的能力跃升,让你能以新的方式使用产品。比如关于在群聊里使用 bot,其实还没有太多训练数据,这是一个正在被摸索的涌现现象,把它做对很重要。
Obviously, if you've seen things like Multbook or the Huggy Face incident, there are ways that that goes poorly. So it's important to be able to train the models how to behave well in a multi-agent system where different bots are talking to each other, for example. And so ideally, you're always kind of running these inner loops and outer loops as a product team building these type of harnesses at a company working on AI together.
译显然,如果你看过 Moltbook(原音识别为“Multbook”)或 Hugging Face(原音识别为“Huggy Face”)事件之类的事情,就知道这件事可能会搞砸。所以能训练模型在多 agent 系统里——比如不同 bot 互相对话时——表现良好,非常重要。因此理想情况下,作为一家公司里一起做 AI、构建这类 harness 的产品团队,你要一直运行这些内循环和外循环。

One of the, the principles that we had at Cursor and now at, SpaceX is to delete the product. And what we mean by this is, you kind of have to internalize and assume that the models today are the worst they'll ever be. Which means that in six months, you might need to completely rebuild the UI. Like most of the stuff that I've I've talked about today, it's a pretty big update from, if I would have given this talk six months ago or definitely a year ago, because we've learned a a lot of new practical ways of building harnesses and doing context engineering. So a good strategy here is you really want to try to build for the next generation of the model. Like, how do you build as little as possible and get by while making just a really, really simple interface to work with these models and allow them to think an act for you. anyhow some things i think will happen in the next six to twelve months we'll see i think these are easily safe that's but uh... you know right now let's say you have your agents of your bots working for maybe a few hours maybe a day depending on how tough it is i think it's very possible that pretty soon uh... your agents will be working for weeks or months uh... it does bring up some kind of interesting questions in terms of how we train and evaluate those models I think the models are gonna get better and better about how they recall on past information in conversations.
译我们在 Cursor、现在在 SpaceX 的一条原则是:把产品删掉(delete the product)。意思是你必须内化并假设:今天的模型是它们有史以来最差的。这意味着六个月后你可能需要彻底重建 UI。我今天讲的大部分内容,如果是六个月前、更别说一年前来讲,都会有很大不同,因为我们学到了很多构建 harness 和做上下文工程的新实用方法。所以一个好策略是:你真的要为下一代模型去构建。比如,怎样尽可能少地构建、还能撑得住,同时做出一个非常非常简单的界面来和这些模型协作,让它们为你思考和行动。总之,我认为未来 6 到 12 个月会发生一些事,我们拭目以待,我觉得这些算是比较稳妥的判断:现在你的 agent/bot 可能工作几个小时,也许一天,取决于任务多难;我认为很可能很快你的 agent 会工作几周甚至几个月。这确实带来一些有意思的问题:我们如何训练和评测这些模型。我认为模型在回忆过往对话信息方面会越来越好。
Plenty of different approaches in research here for the type of algorithms used or the representation of the data under the hood. Maybe it's a graph, maybe it's files maybe something else. I think ideally if you have this digital colleague you give it literally all the tools that you have if you were onboarding a real human. And so I think organizations They're still warming up to that idea, of course. So there's more to do there. I think the interfaces will maybe get even easier. I think some people will prefer to just talk to their kind of bots or agents in a speech-to-speech way. Maybe they'll have their like Johnny I've device that is hopefully really cool and they can like talk to that all day. I think the models and the products will get really good at learning skills on the job and then encoding that into files and instructions, skills, taking that raw intelligence and making [it useful]. [Now some open] problems that I think could just spark ideas for you all to research or to look into. The first one is going back to this idea of agents
译研究里有很多不同方法,涉及所用算法的类型,或底层数据的表示方式——也许是图,也许是文件,也许是别的。我认为理想情况下,如果你有这样一位数字同事,你应该给它你所有的工具,就像你给真人入职那样。当然,各个组织对这个想法还在慢慢接受,所以这方面还有更多要做。我认为界面也许会变得更简单。我觉得有些人会更喜欢用语音对语音的方式和他们的 bot 或 agent 对话,也许他们会有 Jony Ive(原音识别为“Johnny I've”)设计的那种设备,希望它非常酷,他们可以整天和它说话。我认为模型和产品会变得非常擅长在工作中学习技能,然后把它们编码进文件、指令、技能中,把原始智能变得有用。接下来是一些开放问题,我觉得也许能给你们带来启发,去研究或深入了解。第一个,回到“运行一个月的 agent”这个想法:

that run for a month, what do you do when you need to evaluate an agent that runs for a month but the models ship every month? It's like you can't fully evaluate that the model is sufficiently behaving as expected so you either have to figure out how to do a better eval or you know not ship models as fast and I think that's gonna be a problem that the industry has to figure out in the next six the six months. Memory, obviously we talked about the different algorithms. I think there's so much to do on the security side for how to make a secure harness, how to make the infrastructure around these agents very secure, and so they can't be taken over by bad actors. I think we didn't really talk about it too much but the heuristics around when the agent is silent or when it kind of bugs you, I don't think that we have it perfect yet and I think trying to thread the needle between being proactive and suggesting like oh hey like I saw that you missed class yesterday, like I had a recording there and I took the notes and like here's the thing, like maybe that's helpful or maybe that's kind of annoying, you know, and like trying to figure out the right balance there. I also think that it's a pretty safe bet to assume that at this time next year there will be significantly, significantly more tokens flowing in the world, which means that optimizing the inference, optimizing the the models, optimizing the harnesses, trying to drive down the tokens to be more efficient.
译当你需要评测一个运行一个月的 agent,而模型每个月都在发布时,你怎么办?你没法完全评测模型是否表现得足够符合预期,所以要么你得想出更好的评测方式,要么就别那么快发布模型。我认为这是业界在接下来六个月里必须解决的问题。记忆,我们显然讲了不同的算法。我认为在安全方面还有大量工作要做:如何做出安全的 harness,如何让这些 agent 周围的基础设施非常安全,不被坏人接管。我们没怎么谈到的一点,是关于 agent 什么时候保持安静、什么时候来打扰你的启发式规则,我认为我们还没做到完美。要在“主动”和“打扰”之间找到平衡——比如它说“嘿,我看到你昨天缺了课,我在那里录了音、做了笔记,这是要点”——这也许有用,也许有点烦,你知道的,要找到合适的平衡。我也认为可以相当稳妥地假设:明年这个时候,世界上流动的 token 会多得多得多,这意味着优化推理、优化模型、优化 harness、努力降低 token 以提高效率,
Will also the token usage to be more efficient, I think will also be very important.
译让 token 用量更高效,我认为也会非常重要。

So there's kind of these new building blocks to think about. I believe it was Alex from Open Router that I saw on X. He had a really good take on this. There's people saying like, oh, every app today looks the same. You've got the sidebar. You've got the agent. You've got the chat box. Like everybody's building the same thing. And he was like, yeah, that's like in 2005, you said everybody's building the same thing. They've got a server, they've got a database, they've got some like crud interface to like update and delete things. And the reality is this is just kind of the new normal now where every new product is gonna have to have some kind of a genetic side to it, either integrating. of it will become a, you know, a headless piece of data for another to use. And so a lot of these products today, durable workflows, sandboxes, computers in the cloud, memory skills, connecting to other services, doing voice, payments, identity, evals, observability.
译所以有这些新的构建模块值得思考。我记得是 OpenRouter 的 Alex,我在 X 上看到他对此有个很好的观点。有人说:“哦,现在每个应用都长一个样:有侧边栏、有 agent、有聊天框,大家都在做同样的东西。”他说:“是啊,这就像 2005 年你说大家都在做同样的东西:都有服务器、都有数据库、都有某种 CRUD 界面来更新和删除东西。”现实是,这就是新常态:每个新产品都必须有某种 agent 化(agentic,原音识别为“a genetic”)的一面,要么是集成(agent),要么它会变成供其他(agent)使用的无界面数据。所以今天的很多产品——持久化工作流、沙箱、云端电脑、记忆、技能、连接其他服务、语音、支付、身份、评测、可观测性——
Each one of those boxes is like multiple billions of dollars of VC capital, probably trillions. There's 20 startups for each one of those. And there's so many different things to explore and all of those because they kind of are the new primitives for the next generation of companies, the next generation of products. and they want to integrate and be part of harnesses. So there's definitely a lot to do there. A few things to take away and then we'll have time for questions. Obviously you all are engineers, you're thinking about how to write software. It's more important than ever to think about the broader system and the architecture of the code that you're working on because that, the decisions you make there will really compound. So I'm still a pro looking at code. I still look at the code that I'm writing, some people are not looking at the code. I still think it's helpful, especially to understand the architecture that you're shipping. It's crazy how much that's changed in just a couple years. Secondly, it sounds like y'all were kind of messing around with building owned versions or understanding of how the harness worked in last week.
译其中每个格子背后都有几十亿美元的风险投资,可能是上万亿。每个方向都有 20 家创业公司。在所有这些方向上都有太多东西可以探索,因为它们某种程度上是下一代公司、下一代产品的新原语(primitives),而它们都想集成进、成为 harness 的一部分。所以那里肯定有很多事可做。几点收获,然后留时间提问。显然你们都是工程师,在思考如何写软件。比以往任何时候都更重要的是,去思考更宏观的系统和你所写代码的架构,因为你在那里做的决定会不断复利。所以我仍然支持看代码,我仍然会看自己写的代码,有些人已经不看代码了。我仍然认为这有帮助,尤其是为了理解你要发布的架构。短短几年里变化之大真是疯狂。第二,听起来你们上周在折腾构建自己的版本,或者去理解 harness 是怎么运作的。
I encourage you to think about some of the things that I did today and try to apply that and think about how you might build a version for yourself. I'm pretty sure there's like open Grockbot on GitHub somewhere where somebody's like, you know, slot forked our code and made a version. You could probably set up the GUI for that with a sandbox somewhere, a VM, and all sorts of different things. That could be interesting, you want to learn about how these pieces work. Finally, I would not shy away from just learning a little about how LLM's work. We gave a little bit of the TLDR here of how they're trained. So much more you can go into here without necessarily going deep in ML or deep into the science part of it. But I think it's helpful to understand, just like when you're building a web app, it's helpful to know how databases work. Because increasingly, there are so many decisions you make at the product and the harness level that are influenced by the models themselves.
译我鼓励你们想想我今天讲的一些东西,试着应用它们,想想你可能如何给自己做一个版本。我很确定 GitHub 上某处有个“open Grok Bot”之类的项目,有人 fork 了我们的代码做了一个版本。你大概可以在某个沙箱、虚拟机里给它搭起图形界面,等等。如果你想了解这些组件如何运作,这可能会很有意思。最后,我不会回避去学一点 LLM 是怎么运作的。我们这里给了一个关于它们如何训练的简短版(TL;DR),你还可以深入很多,而不一定要深入机器学习或科学部分。但我认为理解它是有帮助的,就像你做 Web 应用时了解数据库如何运作很有帮助一样。因为在产品和 harness 层面,你做的越来越多的决定都受到模型本身的影响。
And having an intuition of how that works I think will make you a very well-well position for whatever role you're operating in.
译对它如何运作有直觉,我认为会让你在无论什么岗位上都处于非常有利的位置。
