| Paper | Year | Why read it |
|---|---|---|
| A first operating loop | ||
| How I use LLMsAndrej Karpathy | 2025 |
Background. Consumer LLM products place model selection, reasoning controls, tools, file upload, coding features, and persistent instructions behind one chat interface. A new user can stay in the text box without learning what the rest of the interface changes. Problem. The interface does not tell you which capability a task needs or what each choice costs. A poor model choice, missing context, or unnecessary tool call can reduce quality or add latency without producing an explicit error. Key idea. Walk through the product surface in the order a user encounters it, using concrete tasks to show when model selection, reasoning, tools, files, coding, and persistent instructions change the result. Why it matters. Start here if the agent ecosystem is new to you. The demonstrations also expose the workload behind the interface: one user task can create several model calls, tool calls, and growing prompts. Those are the quantities the rest of the course measures. |
| Best practices for Claude CodeBest practices for Claude Code (Anthropic) | — |
Background. A terminal coding agent can inspect a repository, edit files, run commands, and continue until it decides the task is complete. Its context contains the conversation, file contents, and command output accumulated during that work. Problem. Two resources need explicit management. Without a check the agent can run, “looks done” is its only stopping signal. Without context discipline, exploration and command output consume the window before the task is complete. Key idea. Give every task an executable verifier, then separate exploration, planning, implementation, and commit. Provide specific context, course-correct early, clear unrelated work from the session, and use subagents or parallel sessions only when their outputs can be reviewed and combined. Why it matters. Treat the guide as both a user manual and a workload description. Verification changes when a run can stop; context management changes prompt length; subagents and parallel sessions change the number and dependency structure of model calls. Assignment 1 records each of those effects. |
| Agentic Engineering PatternsAgentic Engineering Patterns (Simon Willison) | 2026 |
Background. Practice around coding agents is distributed across product documentation, blog posts, and individual prompts. The same techniques recur under different names, which makes experience difficult to transfer between tools. Problem. Advice without a shared vocabulary is hard to compare or apply to a failure. A user needs to know whether a task calls for a smaller prompt, a test loop, a checkpoint, a subagent, or a different interface. Key idea. Organize an evolving field guide around reusable practices, including Git as the recovery mechanism, subagents for bounded work, red-green testing, agent-driven manual testing, code walkthroughs, and annotated examples of real prompts. Why it matters. Use this as a reference after the two readings above rather than as a linear feature tour. It connects each user habit to a control mechanism the harness can expose, which is the bridge from using an agent in this lecture to designing one in Lecture 4. |
| Configuration, context cost, and enforcement | ||
| Equipping agents for the real world with Agent SkillsEquipping agents for the real world with Agent Skills (Anthropic) | 2025 |
Background. Persistent instruction files place an entire operating manual in the prompt, even when the current task needs only one procedure. Installing more procedures can therefore consume context before the user asks for anything. Problem. An agent needs enough information to discover a relevant capability without paying to load every capability in full on every call. The loading policy must also package scripts and reference material that are useful only after a procedure is selected. Key idea. Use three levels of progressive disclosure: load each skill’s name and description at startup, read its SKILL.md body when selected, and open bundled files only when a later step needs them. Why it matters. This is a concrete context-admission policy. The description imposes a standing token cost, while the body and resources are paid only on use. Read it against Lecture 2's context-rent discussion and ask which level each instruction actually deserves. |
| Code execution with MCPCode execution with MCP: Building more efficient agents (Anthropic) | 2025 |
Background. Many MCP hosts place every enabled tool definition in the model’s prompt and return every intermediate tool result through the conversation. Tool catalogs and data then compete with the user’s task for the same context. Problem. Direct tool calling repeats schemas on every model call and serializes large results through the model even when ordinary code could filter or combine them. The overhead grows with both the number of tools and the length of the workflow. Key idea. Present tools as code APIs that the agent can discover on demand, then execute loops, joins, filtering, and error handling outside the model. Only selected definitions and compact results enter the context. Findings. In the article’s worked example, on-demand discovery reduces tool-definition input from 150,000 tokens to 2,000 tokens, a 98.7% reduction. The example demonstrates a mechanism rather than a benchmark over a representative task distribution. Why it matters. It makes context admission actionable. Tool selection becomes retrieval, intermediate data can remain outside the prompt, and several dependent tool operations can run without inserting another model call between them. |
| demystifying skills and MCPsHow skills and MCPs actually work (Lawrence Jones)also Sep 16 | 2026 |
Background. Skills and MCP servers are installed into the same context window, and both reach the model as text it reads before deciding anything. A user configuring an agent therefore has three places to put one piece of guidance: the system prompt, a skill, or the tool itself. Problem. The three are not interchangeable, and the difference is not a matter of wording. The system prompt outranks the user's own messages by construction, so a rule written there is obeyed rather than weighed. A tool description is resident on every turn whether or not the turn needs it. A requirement stated as prose in either place is one the model can reason its way past. Key idea. Separate capability from know-how, then match each to the placement that makes the model treat it correctly. MCP supplies capability as a tool list merged into the prompt, close enough to an ordinary HTTP API that an OpenAPI specification does the same job. A skill supplies know-how as ordinary thread messages below the system prompt, which is what allows the model to weigh it against the task. A non-negotiable goes into neither and becomes a required tool argument, since an argument the model has to supply is not one it can argue with. Findings. A skill description costs roughly 100 tokens of resident context, and its body of a few thousand tokens loads only once that description matches the task, which is the three-level admission policy Agent Skills describes, reached independently from one team's practice. The post also reports what misplacement costs: a system-prompt rule instructing the agent to split long telemetry queries made it refuse legitimate week-long requests, and the identical rule reloaded as a skill restored the intended behavior. The author attributes much of the hallucination he sees from frontier models to prompts that leave no compliant answer available, such as demanding three to five tags for an empty document. Why it matters. It turns Lecture 2's distinction between resident and retrieved context into a rule you can apply per instruction rather than per file. Ask of each line of configuration whether it has to be resident, whether it should be weighed or obeyed, and whether it is a constraint the model should be unable to decline. The three answers select the system prompt, a skill, or the tool signature, and the same three questions apply to a persistent instruction file and to a connected MCP server. |
| Claude Code hooks referenceHooks reference (Anthropic) | — |
Background. Instruction files and skills influence an agent only when the model notices and follows their text. Some requirements, such as blocking writes to a protected directory, cannot depend on that cooperation. Problem. Repeating “never do this” in the prompt spends tokens without enforcing the boundary. The host needs a deterministic interception point before an action and observable signals after actions and session events. Key idea. Register handlers on lifecycle events such as SessionStart, PreToolUse, PostToolUse, and Stop. A PreToolUse command hook can inspect a proposed call and block it before the permission check or tool execution. Why it matters. Hooks draw the line between guidance and enforcement. Read the event table rather than every schema: identify which requirements belong in context, which belong in an automatic repair, and which must be rejected before they run. |
| The Harness Is the ThingThe Harness Is the Thing (Scott Fryxell) | 2026 |
Background. A developer who works across several coding agents keeps the configuration that shapes them, including an instruction file, skills, extensions, scripts, and a work directory, outside any one product. That surrounding structure, rather than the model behind the terminal, is what this post calls the harness. Problem. A single prompt that plans, implements, and critiques the same task confuses its own objectives, and sending every step to a frontier model pays the highest available price for work a cheaper model finishes. Configuration written against one vendor’s tool has to be rewritten when the developer switches tools. Key idea. Keep the instruction file, the skills, and the scripts they call in one repository that every terminal agent loads, then isolate the roles a task passes through: exploration produces a plan, the plan becomes an explicit DAG of tasks, a worker implements one node at a time, a critic simplifies and questions the result, and a promoter communicates the finished work. Route each role to a model chosen for it, reserving the frontier model for planning, the first task, and promotion, and running maintenance work on a cheaper one. Findings. The author reports that the split cut their frontier-model usage in their most intense contexts by 75%, leaving two twenty-dollar monthly subscriptions sufficient for both client and personal work. The number is one developer’s self-report across their own projects rather than a controlled measurement, and the post concedes that confirming which harness changes helped will require an empirical approach it has not yet applied. Why it matters. It is the practitioner’s side of the question Guardrails Beat Guidance studies below, and the two disagree usefully about what counts as evidence for a configuration choice. Read it for the framing the rest of the course assumes: models become interchangeable parts, and what the user designs is the harness, meaning the roles, the model routing, the scripts, and the files that persist between sessions. |
| Evidence on harness configuration and cost | ||
| Guardrails Beat GuidanceA Large-Scale Study of Rules, Skills, and Persistent Configuration for Coding Agents | 2026 |
Background. Coding agents load repository rules, skills, and persistent configuration as context. Teams write these files from experience, but a plausible rule is not evidence that the rule caused a better result. Problem. Rule content, polarity, position, and the mere presence of additional context are confounded unless each is varied against a no-rule control. Observational examples therefore cannot identify which instruction helped. Key idea. Extract 25,532 rules from 679 public rule files, then run more than 5,000 Claude Code trials while varying curated, random, shuffled, mismatched, and individually ablated rules. Findings. Curated and random rule sets both improve pass rate by 13.8 percentage points over the no-rule baseline on the paper’s 58-task subset. In the individual ablations, beneficial rules are negative constraints and harmful rules are positive directives. Why it matters. The random-rule control shows why a rule should not receive credit without an ablation. Treat the polarity result as a strong hypothesis rather than a universal law, because the experiment uses one harness, one model, and one benchmark subset. |
| The Scaffolding Matters More Than the InterfaceA Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task | 2026 |
Background. An agent can reach repository operations through MCP tools or by writing ordinary command-line commands. Comparisons often attribute a cost difference to the interface even when the surrounding harness also changes. Problem. System prompts, tool descriptions, retry behavior, and whether the agent follows the assigned route can all differ between runs. An unpaired comparison therefore measures several variables at once. Key idea. Hold one six-operation software task fixed, run it across seven agent scaffoldings and five models, pair MCP and CLI runs where the scaffold supports both, and verify the repository state instead of trusting the agent’s completion claim. Findings. Thirteen paired MCP-to-CLI cost ratios range from 0.43× to 29×. The scaffold dominates the result, and agents frequently ignore the interface they were assigned. Why it matters. It sets a methodological requirement for interface comparisons: verify the route actually used and hold the harness fixed. The evidence comes from one small task, so the instability of the ratio, not any one ratio, is the result to retain. |
| Paper | Year | Why read it |
|---|---|---|
| References | ||
| Software 2.0Andrej Karpathy | 2017 |
Background. Machine learning was commonly described as another component inside conventional software. That description kept the programmer-written source code at the center and treated a trained neural network as one replaceable tool. Problem. The framing misses a change in what specifies program behavior. For many perception and decision tasks, writing the desired algorithm directly is harder than collecting examples, defining an objective, and searching for a model that performs well on them. Key idea. Call explicit human-written code Software 1.0 and learned neural-network weights Software 2.0. In the second stack, the dataset and model architecture act as source code, training acts as compilation, and iteration moves from editing individual instructions toward editing data, objectives, and the surrounding training system. Why it matters. This is the baseline for the two Software 3.0 readings below. Ask what changes when the learned model stops being only a component trained for one task and becomes an interface that interprets instructions, writes conventional code, calls tools, and coordinates work. |
| Software 3.0 — the era of intelligent software developmentItamar Friedman | 2022 |
Background. Code models such as Codex and Copilot made generated source code visible before autonomous coding agents became a common product interface. Software 2.0 explained programs represented by learned weights, but not a workflow in which a model produces the conventional code around those weights. Problem. A programming stack driven by natural-language instructions needs a different account of authorship, iteration, and validation. Generated code may look familiar while remaining statistically produced, and greater output volume makes manual testing an increasingly narrow bottleneck. Key idea. Describe Software 3.0 as a pipeline in which instructions and data enter an AI agent and the output combines programmer-readable code with neural-network components. The article treats automated testing, evaluation, monitoring, and human review as necessary infrastructure for iterating on that output. Why it matters. This is an early, explicitly speculative use of the Software 3.0 label. Read it historically beside Karpathy's later talk: compare which predictions became ordinary agent features, which remained aspirations, and why executable verification becomes more important when software generation becomes cheaper. |
| Software is changing (again)Andrej Karpathy | 2025 |
Background. Eight years after Software 2.0, large language models can interpret natural language, generate code, use tools, and sit inside applications that combine conventional interfaces with model-driven behavior. The programmer now works across hand-written code, learned weights, and prompts. Problem. Calling an LLM a chatbot hides the new software surface around it. The model is powerful but fallible, application context has to be assembled deliberately, and useful products must decide how much autonomy to grant rather than treating manual use and full delegation as the only choices. Key idea. Present natural language as the programming interface of a Software 3.0 stack and place products on an autonomy slider, from suggestions through increasingly agentic execution. The surrounding application packages context, routes among models and tools, exposes a graphical interface, and keeps a person able to inspect or redirect the work. Why it matters. The talk connects the earlier taxonomy to this lecture's multi-agent question. Moving the autonomy slider to the right creates more model calls, longer trajectories, and more concurrent work, so the user's job shifts toward specifying outcomes, supplying environments, and verifying evidence rather than typing every implementation step. |
| Claude Code subagentsCreate custom subagents (Anthropic) | — |
Background. A side task that fills the conversation with search results, logs, and file contents spends window the later steps still need. A subagent does that work in a context window of its own and returns only its summary, and a definition file makes the same worker reusable across sessions. Problem. Isolation is neither free nor uniform. A subagent that starts fresh sees none of the conversation, so it may re-read files the parent already read, and latency grows while it gathers that context again. A fork inherits the whole conversation instead and gives up the input isolation, keeping only its tool calls out of the parent. Every definition’s description stays resident whether or not a turn delegates, and each returned report is appended to the parent’s context in full. Key idea. Declare a worker as a Markdown file whose frontmatter fixes its tool allowlist, its model, its permission mode, and optionally a temporary git worktree to work in, and whose body becomes its system prompt. The description is the string the parent matches on when it decides to delegate, so the routing policy and the worker are the same artifact. Why it matters. This is the delegation primitive the two pages after it compose, and it is context admission and eviction expressed as configuration rather than as a habit. Read the constraints as the shape of the resource being managed: the combined descriptions of custom subagents raise a warning past 15,000 tokens, a subagent may spawn subagents three layers below the main conversation by default, 20 may run at once, and several tools are withheld from every subagent so a worker cannot ask you a question or approve its own plan. Note also which permission modes a parent forces on its children, since that decides whether a narrow tool list is a boundary or a suggestion. |
| Claude Code agent teamsOrchestrate teams of Claude Code sessions (Anthropic) | — |
Background. A subagent reports to whoever spawned it and communicates through that one return value. Some work instead needs several workers that hold findings simultaneously, disagree with each other, and each own part of a repository. Problem. Making the workers peers turns a delegation tree into a distributed system, and the failure modes follow. Token cost scales with the number of live teammates because each is a full session. Two teammates editing one file overwrite each other. A task marked complete by nobody blocks everything that depends on it. Moreover, a teammate is another agent rather than another user, so a request it relays must not be able to grant a permission you never granted. Key idea. One lead session spawns teammates, each a separate instance with its own context window, its own mailbox as a file on disk, and a name any other teammate can address. Work is a shared task list with dependencies, claimed under file locking so two teammates cannot take the same item, and three lifecycle hooks let a check refuse a task’s creation or completion or send an idle teammate back to work. Why it matters. Read it for the boundary rather than the feature list. Claude Code tells a receiving agent that a message came from another session and not from you, a teammate cannot supply consent on your behalf, and a teammate denied an action cannot route it through a peer, which is the permission boundary of the tools reference redrawn between agents. The page is also unusually candid about cost and limits: teams are experimental and off by default, use significantly more tokens than one session, are recommended at three to five teammates, cannot nest, and do not survive session resumption. Those admissions are the argument for the next entry. |
| Claude Code dynamic workflowsOrchestrate subagents at scale with dynamic workflows (Anthropic) | — |
Background. With subagents, skills, and agent teams, the model is the orchestrator. It decides turn by turn what to spawn next, and every intermediate result lands in some context window. Problem. An orchestration that lives in a context window cannot outgrow that window, cannot be read after the fact, and cannot be rerun the same way. A plan the model re-derives each turn is also not the thing you wanted to reuse, because the worker definition is repeatable while the coordination is not. Key idea. Move the plan into code. A JavaScript script runs in a runtime outside the conversation, where one call spawns a single agent and two others fan a task out over a list, intermediate results stay in script variables, and only the final answer returns to the context. The runtime makes the clock and the random number generator throw, so relaunching a script replays the same calls and completed agents return saved results instead of running again. Why it matters. This is the clearest statement available of the course’s own distinction between what the model decides and what the harness decides, written by the people who had to implement it. The caps are the honest part of the design: at most 16 agents run concurrently, a single fan-out takes at most 4,096 items, and one run spawns at most 1,000 agents, with an advisory warning once a run schedules more than 25 agents or projects more than 1.5 million tokens. Read the resume rules closely, because a failure in the middle of a fan-out reruns every agent that started after it, which is the cost of replay as a recovery mechanism. |
| Building a C compiler with a team of parallel ClaudesNicholas Carlini (Anthropic) | 2026 |
Background. The three pages above describe mechanisms. This one is a field report from running sixteen coding agents in parallel, with no human in the loop, until they had written a C compiler in Rust from scratch with no network access at any point. Problem. An autonomous loop needs a stopping signal it cannot fake, and parallel agents need work partitioned without an orchestrator to partition it. Both requirements break on the same task: building the Linux kernel is one indivisible goal, so every agent hits the same bug and the parallelism buys nothing. Key idea. Keep the harness minimal and put the intelligence in the verifier. A shell loop reinvokes the agent with the same prompt file, each agent works in its own container against a bare repository, and a claim is a file written into a directory where the version control system’s own synchronization forces the second claimant onto different work. There is no messaging between agents and no orchestrator. The indivisible task is split by using an existing compiler as a known-good oracle, compiling most files with it and a subset with the new one, then bisecting to the file that fails. Findings. Roughly 2,000 sessions over two weeks across sixteen agents consumed 2 billion input tokens and 140 million output tokens for just under $20,000, producing about 100,000 lines that pass 99% of most compiler test suites and boot Linux 6.9 on three architectures. The limits are reported with the same specificity: no dependable assembler or linker of its own, one code path that calls an existing compiler instead, and output less efficient than that compiler produces with every optimization disabled. Why it matters. It is the strongest available demonstration that verifier quality, not model quality, is the binding constraint on autonomous work, because an agent optimizes against the check it is given and an imperfect check selects for the wrong solution. Read the two harness details that generalize past compilers: logs written to files and marked so they can be searched rather than printed into the context, and a sampling flag added because the agent cannot perceive elapsed time and will not shorten a slow loop on its own. The author, whose background is in breaking software, ends uneasy rather than triumphant, on the ground that passing tests is easy to mistake for being finished. |
| Research acceleration: The view inside OpenAIOpenAI | 2026 |
Background. A frontier laboratory published measurements of how much of its own research work now passes through coding agents, framed as a transparency commitment and as evidence on a milestone it had announced in advance. Problem. Adoption anecdotes cannot answer how much of the work agents actually do, and a rise in experiments per researcher is not evidence that agents caused it, because available computation grew over the same period. A measurement therefore needs a common unit for agent and human effort and a classification of what the agents were used for. Key idea. Convert total agent runtime across the research organization into eight-hour workdays and compare that against human labor over the same interval, then classify agent activity into the six phases of a research process: decide, design, build, run, analyze, and communicate. Findings. The reported ratio reaches 3.1 agent-workdays per human workday by the middle of August 2026, and inference spending for the median researcher rises from near zero in February to more than $600 a day. Success rates improve on longer tasks, yet more than half of the successful tasks in the four-to-eight-hour range still required human intervention, and high-level planning stays a small share of what the agents produce. Why it matters. It supplies the denominator the compiler report leaves out, which is how a whole organization’s work divides between agents and people rather than how one ambitious project went. Read the ratio as a utilization measure rather than a productivity one: runtime is not output, the post concedes that growing computation confounds the throughput trend, and the intervention rate says the remaining human work concentrates in whatever is least automatable. Every figure comes from the organization measuring itself, which is the same caution the practitioner report above deserves. |
| The Shift to Agentic AIEvidence from Codex | 2026 |
Background. Claims about agentic engineering mostly come from inside one organization or from users describing their own practice. This paper measures a coding agent’s usage across three populations at once: individuals on personal accounts, external organizations, and the vendor’s own staff. Problem. Usage counts alone cannot distinguish more people trying a tool from people restructuring their work around it. Separating the two requires measures of sophistication, such as how many agents someone runs at a time and how large a task they are willing to hand over, and it requires a pipeline that produces those measures without exposing what anyone wrote. Key idea. Build an automated privacy-preserving pipeline over usage data, then report volume and sophistication separately and compare the three populations, so internal adoption becomes a leading indicator against which external adoption can be read. Findings. Active users grow more than fivefold across the first half of 2026, with the fastest growth outside the original developer audience. More than 10% of users run three or more agents at once in a given week and 26.6% use shared instruction files. The share of individual users submitting at least one task estimated at more than eight hours of experienced human work grows nearly tenfold from the start of the year. Inside the vendor, the median employee in a legal role produced 13× the monthly output tokens of November 2025 and the median researcher more than 50×. Why it matters. It is the outside counterpart to the post above, and the two together are the best available answer to how quickly the practices in this lecture spread. However, every quantity is a measure of what people ran rather than of what the runs produced, so treat the sophistication measures as evidence of changed behavior and not of changed output. The concurrency and task-length numbers also give the workload behind Part II a shape: one user is increasingly several simultaneous long-running sessions rather than one interactive conversation. |
| Paper | Year | Why read it |
|---|---|---|
| Agent | ||
| ReActSynergizing Reasoning and Acting in Language Models | 2022 |
Background. By 2022 two lines of work sat next to each other without touching. Chain-of-thought prompting had models reason in text, and other work had models emit actions against an environment or a search API. Problem. Reasoning alone has nothing to check itself against, so an early wrong fact propagates confidently to the end. Acting alone has no plan, so a model cannot decide what to do next when an observation contradicts what it expected. Key idea. Interleave the two in a single trace. The model writes a thought, then an action, then reads the resulting observation, then thinks again, all from one prompt. This interleaved reason/act loop is the mechanism, and nearly every agent framework is a variation on it. Findings. On ALFWorld and WebShop, ReAct improved absolute task success over the paper's imitation- and reinforcement-learning baselines by 34 and 10 percentage points, respectively, using one or two in-context examples. Why it matters. It fixes the shape of the object this course keeps taking apart: thought, action, observation, repeat. Planning, tool schemas, context management, and permission gates are all modifications to one position in that cycle, which is why the loop is worth reading in its original form. Adoption. LangGraph ships and documents create_react_agent as a direct implementation of ReAct. |
| building effective agentsBuilding Effective Agents | 2024 |
Background. By 2024 the default advice for building on an LLM was to adopt an agent framework, which hands you a planner, a memory abstraction, and a graph of steps before you have established that you need any of them. Problem. Those abstractions hide the prompt and the tool definitions, which are the two things you actually have to debug. Complexity gets added because a pattern is available, not because a measurement showed the simpler version failing. Key idea. Separate workflows, whose steps are fixed in code, from agents, which let the model choose the steps, then name the composable patterns worth reaching for: prompt chaining, routing, parallelization, orchestrator-workers, and evaluator-optimizer. Add complexity only when it measurably improves outcomes. Why it matters. It gives you a default that is hard to argue with, measure first and complicate second, and it supplies the names for the thing you built. A design review can then say which pattern is in use instead of calling every system an agent. |
| What Is an Agent Harness? Harness Engineering ExplainedTejas Kumar | 2026 |
Background. An agent is often described as a model inside a tool-calling loop, but a deployed system also has to manage context, keep secrets away from the model, stop runaway executions, and decide whether the requested work actually happened. Problem. Treating the loop as the whole agent leaves the model in charge of declaring success. In the post's example, GPT-3.5 Turbo reaches a Hacker News login page, performs no upvote, and reports that the task succeeded because nothing outside the model checks its claim. Key idea. Define the harness as the runtime layer around the model, made of six parts: the tool registry, model, context management, guardrails, agent loop, and verification. Put deterministic limits, authentication, trace inspection, and retries in that layer, so failures become code changes to the harness rather than stronger wording in the prompt. Why it matters. It separates the mechanism that proposes actions from the system that makes those actions reliable. The walkthrough holds the model and prompt fixed while adding guardrails, verification, and a login handler until the task completes and the trace confirms it, which gives a concrete design test: a model's completion message is a claim, and the harness must supply the evidence. |
| Agent PatternsPatterns, Anti-Patterns, and Primitives for Coding Agents | — |
Background. The entry above names five composable patterns and argues for reaching for the simplest one that works. Coding agents have since become the dominant application, and the working knowledge about building them accumulates in blog posts, changelogs, and team lore rather than in papers. Problem. That knowledge is hard to consult at the moment it is needed. Someone deciding how to shape a tool definition, a review step, or a split across several agents has no reference that separates what works from what merely circulates, and the failure modes are the part nobody writes up. Key idea. Build a reference corpus rather than one long argument. Short Markdown pages each cover one concept, group related practices, and place anti-patterns beside patterns so known failures are documented instead of rediscovered. Agents can load the same pages into their own context. Why it matters. Read it as the practitioner's counterpart to the paper above, and read the anti-patterns first, since that is where the cost of a pattern shows up. That it is written to be consumed by an agent is the lesson in miniature: a document an agent can load is a different artifact from one written only for a human reader. |
| LLM powered autonomous agentsLLM Powered Autonomous Agents (Lilian Weng) | 2023 |
Background. By 2023 the agent literature was a scatter of prompting patterns, tool wrappers, and retrieval add-ons. There was no shared account of what an agent is made of. Problem. Without a decomposition, two agent systems cannot be compared. Each paper described its own loop end to end, so it was hard to tell which part of a design carried the result and which part was incidental. Key idea. Put the model call at the center and name three components around it: planning, memory, and tool use. The post is a survey rather than a new system, and the vocabulary is the contribution. Why it matters. That split is the structure LangChain, LlamaIndex, and the AutoGPT-era frameworks expose to users, and it is still how most agent papers organize their related work. |
| practical guide to building agentsA practical guide to building agents (OpenAI) | 2025 |
Background. Research papers explain agent components, but a team first needs to decide which workflows justify an agent at all. This guide draws on OpenAI's customer deployments. Problem. Teams may choose an agent where a deterministic pipeline would suffice, then discover the reliability cost late. Research papers rarely define the boundary clearly enough to make that decision early. Key idea. Define an agent as a system in which the model controls workflow execution. Use one when nuanced judgment, unmaintainable rules, or unstructured data defeat a deterministic design. Start with one model, tools, and instructions; add multiple agents only when necessary, with guardrails, explicit exit conditions, and human escalation. Why it matters. Its strongest contribution is a negative criterion: it identifies systems that are not agents and workflows that should remain deterministic. Compare its decomposition advice with the Anthropic and Cognition positions. |
| Structuring context for cache reuse | ||
| context engineeringEffective context engineering for AI agents (Anthropic) | 2025 |
Background. Prompt engineering treats the window as somewhere to word one instruction well. An agent instead fills that window itself over many turns, with tool results, file contents, and its own earlier output. Problem. Context is a finite resource whose value degrades as it fills, so a long agent run is not fixed by phrasing. The question is which tokens deserve the window on this turn and what happens to everything else. Key idea. Curate the context rather than write the prompt. The four techniques it names are mechanisms: compaction, structured note-taking into external memory, subagents that each get a clean window, and just-in-time retrieval that holds identifiers and loads the data only when a step needs it. Why it matters. Each of the four changes what the serving system underneath actually sees. The agent-memory readings later in this sequence show what happens when context allocation goes wrong. Treating the window as an allocation decision is what makes prompt work measurable. |
| Sourcegraph context engineeringContext Engineering: A Practical Guide for AI Agents | 2026 |
Background. An agent’s context window holds the system prompt, the tool definitions, the retrieved code, and every tool result so far. On a long coding task that window, not the model, is the binding constraint. Problem. Contexts fill with material that has stopped helping: stale file reads, superseded diffs, search output the agent already consumed. A full window forces truncation at the worst moment, and text that no longer matters competes for attention with text that does. Key idea. Treat context as a resource with a lifecycle rather than a prompt to be worded well. The guide organizes that work into four pillars, separates it from prompt engineering, and reports which practices hold up in production coding agents. Why it matters. It names the discipline that decides whether a long agent run stays coherent. Once you see the window as a budget with an allocation policy, retrieval scoping, compaction, and subagent isolation stop being separate tricks and become one decision made three times. |
| context pruningLess Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents | 2026 |
Background. A long-horizon tool-using agent writes its own transcript. Every tool result, intermediate step, and retry stays in the window and is re-read on every following turn. Problem. Trimming that history is normally treated as a cost measure paid for in accuracy, so pruning gets tuned only once a run stops fitting or stops being affordable, and never as something that could help the task itself. Key idea. Prune tool output and history aggressively, then score task success rather than tokens alone. Findings. On a 50-task hotel-expense benchmark, full history completed 71.0% of tasks using 1.48 million tokens. Keeping five tool exchanges completed 79.0% using 535,274 tokens, while adding summaries completed 91.6% using 553,374 tokens. Why it matters. Context management becomes a joint accuracy and cost knob rather than pure compression. If cutting context can improve both at once, the right setting is not the largest window your budget allows, and it has to be measured per workload. |
| lost in the middleHow Language Models Use Long Contexts | 2023 |
Background. Context windows grew from a few thousand tokens to tens of thousands, and the working assumption was that anything you place inside the window is equally available to the model. Problem. It is not. On multi-document question answering, accuracy depends on where in the input the one relevant passage sits, so a retrieval pipeline can hand the model exactly the right document and still get the wrong answer. Key idea. Hold the task fixed and move the relevant document through the context, measuring accuracy at each position. Findings. The curve is U-shaped: accuracy is highest when the evidence sits near the beginning or the end and lowest when it sits in the middle. Why it matters. Long context is not free accuracy. This is the paper to reach for when a project’s answer to a problem is “put more in the prompt”, because it turns the ordering of retrieved chunks into a design decision with a measured cost. |
| context rotDiagnosing and Mitigating Context Rot in Long-horizon Search | 2026 |
Background. A long-horizon search agent appends every tool result to its context and rarely removes anything, so the prompt fills with search output the model has already finished using. Problem. That accumulation costs accuracy, not just space. The agent starts answering worse while the context window is still far from full, so watching capacity never sees the failure coming. Key idea. Grow accumulated search context while controlling task difficulty, classify the resulting failures, and compare seven context-management methods with a behavior-aware filter for parallel samples. Findings. Across four models and three benchmarks, premature termination increased with context length even before the window filled. Behavior-aware filtering improved reported accuracy by 2.6-4.9% over unfiltered aggregation across three aggregation methods. Why it matters. It changes the argument for context management. Eviction and compaction are normally justified by memory pressure; the same operations justified by answer quality have to trigger much earlier, and on different evidence. |
| Tool design | ||
| writing effective toolsWriting effective tools for agents — with agents (Anthropic) | 2025 |
Background. An agent reaches the outside world only through tools, and each tool a harness mounts arrives at the model as a name, a schema, and a description it has to read. Anthropic wrote this guide out of building Claude Code and the Model Context Protocol. Problem. A tool the model misuses is usually a documentation failure rather than a capability failure. The description has to teach the schema and the semantics at once, and the output has to be legible to a model instead of to a person reading a log. Key idea. Treat the description as part of the interface, not as documentation beside it, since it is what teaches the model how and when to call the tool. Shape the return value the same way, giving the next decision what it needs rather than everything the underlying API produces. Why it matters. Tool definitions and their outputs occupy the same context budget as the task, so a badly shaped tool costs tokens on every turn it is mounted, not only on the turns it is called. It is the practical guide MCP server authors follow when naming tools and shaping their outputs. |
| superhuman bashHow Foundational Models Became Superhuman in Bash (Philipp Schmid) | 2026 |
Background. A harness gives the model one tool per operation, so reading a file, writing a file, editing a range, and searching a tree each arrive as a separate schema. Every one of those schemas is resident context, and the set of them fixes what the agent can do. Problem. The designer has to anticipate each operation in advance, and an operation nobody anticipated is unreachable however capable the model is. Composing several of those tools also costs a round trip per step, and each hand-off is a place the trace can go wrong. Key idea. Expose a shell instead of a catalog. A single command tool routes to the programs already installed, so Python, git, SQLite, a compiler, and the project's own command-line tools become the action space without any of them being mounted as a tool. The safeguards the atomic tools used to provide move into the harness: truncate long output and say how to request a narrower slice, return exit status and duration alongside the text, gate paths and network and destructive commands by policy, and manage long-running processes asynchronously. The author pairs this with two habits, keeping bulk data in the environment and returning only summaries to the model, and delegating messy debugging to a subagent that reports a clean result. Findings. The author reports that a shell-centered harness performed on par with or better than one exposing separate read, write, edit, and search tools on the same task set, with no numbers given, so read it as one practitioner's observation rather than a measurement. He also names the exception: text cannot carry pixels, so a screenshot still needs its own channel, and a tool is still justified wherever it genuinely offers a better interface than a command would. Why it matters. It states the tool-count question as a design trade-off rather than a preference. A large catalog spends resident context and bounds the action space by what the designer foresaw, while a shell spends almost none and bounds it by what the environment has installed, at the cost of moving truncation, sandboxing, and process management into the harness. The three papers that follow read differently in this light, since retrieving from a catalog of thousands of APIs presumes the catalog is where capability lives. |
| ToolformerLanguage Models Can Teach Themselves to Use Tools | 2023 |
Background. A model that cannot call out answers arithmetic, current facts, and translation from its weights. The fix at the time was prompting: describe the available tools in the context and hope the model emits a well-formed call at the right moment. Problem. Prompted tool use puts the decision in the wrong place. Whether an API call helps at this point in the text is a property of the model's own uncertainty, and instruction text cannot teach it when calling is worth the extra round trip and when it is not. Key idea. Let the model label its own training data. Sample candidate API calls in context, execute them, keep only the calls whose results reduce the model's loss on the tokens that follow, and fine-tune on what survives. Tool use becomes learned behavior instead of instructed behavior. Findings. Across the paper's downstream evaluations, Toolformer improved zero-shot performance and was often competitive with much larger models while preserving its underlying language-modeling performance. Why it matters. This is where the tool-call boundary moves into the weights, which is what makes tool schemas a serving concern. The checkpoint now expects one particular call format, so the chat template, the parser, and the execution loop become part of the model's contract. |
| demystifying skills and MCPsHow skills and MCPs actually work (Lawrence Jones)also Sep 9 | 2026 |
Background. An agent is a loop. The harness sends the thread, runs whatever tools the model called, appends the results, and repeats. The model holds no state of its own, so everything the agent knows is whatever that assembly put in the thread, and a designer's whole leverage sits in the assembly step. Problem. Two structural decisions follow from that loop and are usually made by default rather than on purpose. The first is how many agents to run, since one agent per domain keeps each prompt small but splits the thread so that no agent sees the whole task. The second is what a tool has to be, since a designer who assumes a tool must wrap a real system will build integrations the loop never required. Key idea. Treat both as free choices, because the loop constrains neither. A tool need not correspond to anything real. The author's team implements a shell tool in Go over a virtual filesystem, and the model cannot tell, so the tool surface can be designed for the model instead of inherited from the backend. In place of one agent per domain, run a single agent that loads skills on demand, which returns full-thread context while keeping the resident prompt small. Wrapping that load in a skill tool, rather than letting the model open files itself, buys name resolution, usage tracking, and a returned file index that saves the model a recursive directory listing. Findings. A skill description costs roughly 100 tokens of resident context, and its body of a few thousand tokens loads only once that description matches the task, which is what makes the single-agent consolidation affordable at all. Across the roughly 20 observability platforms the author's team covers, one loop and one tool surface serve every platform, with the platform-specific knowledge carried in skills rather than in code. Why it matters. It separates what a designer has to build from what a designer can simply declare. The loop, the thread, and the tool schemas are engineering, while the know-how about when to use them is text loaded on demand, and moving material across that line is how a harness gains capability without growing its resident prompt. The post also notes that none of this is provider-specific, so the design carries across model vendors. |
| Paper | Year | Why read it |
|---|---|---|
| Multi-agent architecture and orchestration | ||
| AutoGenEnabling Next-Gen LLM Applications via Multi-Agent Conversation | 2023 |
Background. By 2023 the single-agent loop of a model plus tools plus a scratchpad was a settled pattern. Building an application out of several such agents meant writing the orchestration by hand every time, because there was no vocabulary for how one agent hands work to another. Problem. Multi-agent applications had no programming model. What an agent is, how it receives a turn, who decides whose turn comes next, and where a human sits in the arrangement were all answered per application, so nothing composed and nothing was reusable. Key idea. Make conversation the programming model. Agents are conversable objects with configurable roles and tools, a human can be one of the participants on equal terms, and control flow is whatever message-passing pattern you wire between them, including a group chat with a manager choosing the next speaker. Findings. On the full MATH test set, an AutoGen conversation between an assistant and a code executor reaches 69.48% accuracy, compared with 55.18% for GPT-4 alone. On 134 unseen ALFWorld tasks, adding a grounding agent improves success by 15% on average over the same conversational system without that agent. Why it matters. It gives you the vocabulary for deciding when a second agent earns its bill. Read against the fan-out arithmetic it also prices the arrangement, since every additional participant carries its own context and pays its own tokens on every turn it takes. |
| MetaGPTMeta Programming for A Multi-Agent Collaborative Framework | 2023 |
Background. Conversational multi-agent frameworks let agents talk freely and let coordination emerge from the conversation. Human software teams do not work that way. They assign roles and pass documents whose shape is agreed in advance. Problem. Free-form conversation between agents drifts. Each agent restates the task in its own words, an error made early propagates as prose that later agents treat as given, and no step has a contract the next step can check before acting on it. Key idea. Assign roles borrowed from a software organization, product manager, architect, engineer, and quality assurance, and require each role to emit a structured artifact that the next role consumes. The hand-off is explicit and the return contract between agents becomes the design rather than a convention. Findings. With GPT-4, MetaGPT reaches 85.9% pass@1 on HumanEval and 87.7% on MBPP. Its executable-feedback loop improves MBPP by 5.4 percentage points over the otherwise identical pipeline without execution feedback, isolating part of the gain from the broader multi-agent arrangement. Why it matters. It is the clearest statement of the position that structure, not more conversation, is what makes a pipeline of agents work. The structured hand-off is also what makes the pipeline debuggable, since you can read the artifacts and see which stage was already wrong. |
| multi-agent research systemHow we built our multi-agent research system (Anthropic) | 2025 |
Background. Some agent builders favor one compacted thread. Anthropic's research feature instead uses a lead agent to delegate parallel searches and synthesize their results. Problem. A multi-agent design must specify each worker's task, its return format, and the lead's stopping rule. Every worker also pays for a fresh context. Key idea. The lead delegates explicit tasks, workers explore in parallel, and each returns a summary rather than its transcript. Prompting concentrates on teaching the lead to scope work. Findings. Agents use about 4× the tokens of a chat interaction, and multi-agent systems use about 15×. On BrowseComp, token use alone explains 80% of performance variance. Why it matters. Compare it with Cognition's single-thread position. Parallelism pays when subtasks are genuinely independent and read-only, and the 15× token cost determines whether the fan-out is affordable. |
| scaling agent systemsTowards a Science of Scaling Agent Systems | 2025 |
Background. The frameworks above show how to wire multiple agents, and Anthropic shows one task where parallel workers pay off. Neither establishes which arrangement to choose for a different task. Problem. Comparisons of agent teams often change the model, prompts, tools, or compute budget along with the coordination design. A higher score then cannot show whether the architecture helped, and adding agents remains a guess. Key idea. Hold prompts, tools, and compute fixed while comparing a single agent with independent, centralized, decentralized, and hybrid multi-agent arrangements. Vary model capability and task structure across 260 configurations on six agentic benchmarks. Findings. Relative to a single agent, multi-agent performance ranged from an 80.8% gain on decomposable financial reasoning to a 70.0% loss on sequential planning. Independent agents amplified trace-level errors 17.2×, compared with 4.4× under centralized coordination. Why it matters. It gives the architecture choices above an empirical test. Parallel work can pay when a task decomposes, while sequential work can lose accuracy to coordination; a central verifier also changes how far errors propagate. Choose the topology from the task and its failure modes rather than the number of agents alone. |
| Coordination and communication | ||
| don’t build multi-agentsPrinciples of Context Engineering (Cognition) | 2025 |
Background. The frameworks above assume decomposition helps: split the task, give each part its own agent, run the parts in parallel, and merge. Cognition builds Devin, a long-running coding agent, and reached the opposite conclusion from running one in production. Problem. Most of what an agent decides rests on context that was never written down. Split the work across parallel agents and each one makes those implicit decisions independently, so the pieces disagree and the disagreement surfaces only at the merge. Decomposition can multiply inconsistency faster than useful work. Key idea. Keep a single thread. Carry the whole trajectory forward and compact it when it grows too long, and where a subagent is unavoidable, give it read-only work and take back a summary instead of letting it act. Context engineering, not orchestration, becomes the primary design activity. Why it matters. It is the counterargument to AutoGen and MetaGPT, written by a team shipping a product rather than a paper. Read it as an engineering position paper and ask which of its claims the evaluation papers still need to test. |
| Paper | Year | Why read it |
|---|---|---|
| State management and agent memory | ||
| RAGRetrieval-Augmented Generation for Knowledge-Intensive NLP Tasks | 2020 |
Background. Before 2020, a model answered a knowledge-intensive question from whatever pretraining had written into its weights. Updating a fact meant retraining, and there was no passage to point at when the answer came out wrong. Problem. Parametric knowledge is fixed at training time and cannot be inspected. Asked about something absent from its weights, the model produces a fluent answer anyway, and nothing in the output tells you which case you are in. Key idea. Put a dense retriever over a Wikipedia index in front of a pretrained sequence-to-sequence generator, fetch passages for each input, and condition generation on them, training retriever and generator together while the index stays fixed. Knowledge moves out of the weights and into a store you can edit. Findings. RAG set the reported state of the art on three open-domain question-answering tasks and generated more factual language than the paper's parametric-only sequence-to-sequence baseline. Why it matters. It is the origin of the long, partly repeated prompts that make prefix caching worth building. It also splits serving in two: a retrieval tier with its own index and latency, and a generation tier whose input length is now set by how much you retrieved. Adoption. Hugging Face Transformers implements the original RAG model and retriever classes, providing a maintained external implementation of the architecture. |
| MemGPTTowards LLMs as Operating Systems | 2023 |
Background. The context window is a fixed budget, and any long conversation or long task overruns it. The standard answers are truncation, summarization, and retrieval bolted on outside the model. Problem. All three discard state without asking the agent what it still needs, and none of them lets the agent decide what stays resident. The window is managed for the model rather than by it. Key idea. Treat the context window as the top tier of a memory hierarchy and give the model explicit paging operations into the tiers below. What stays resident, what is evicted, and what is fetched back become decisions the agent makes with function calls. Why it matters. It is the most systems-flavored agent paper on the list, and it imports the virtual memory design wholesale: a small fast tier, a larger slow one, and an explicit policy for moving data between them. Once the window is a cache, compaction and retrieval stop looking like separate features. |
| Common failure modes and guardrails | ||
| where agents failWhere LLM Agents Fail and How They can Learn From Failures | 2025 |
Background. An agent failure is observed at the end of a trajectory. What is visible is a wrong answer, and the step that made it inevitable is many turns behind it. Problem. Errors cascade rather than accumulate. One wrong decision enters the transcript and every later step reasons from it as given, so the last step that looks wrong is rarely the step that was wrong. Diagnosis from the final answer therefore identifies the symptom. Key idea. Annotate failures at the step level instead. Sort the errors by which module produced them, covering memory, reflection, planning, action, and system operations, then build a debugger that isolates the root-cause step and feeds a correction back into a fresh attempt. Findings. Root-cause feedback is worth more than a retry: the paper reports 24% higher all-correct accuracy and 17% higher step accuracy over the strongest baseline, and up to 26% relative improvement in task success on ALFWorld, GAIA, and WebShop. Why it matters. It is the diagnosis order of this meeting with a dataset behind it, and it makes the case for treating the transcript as the primary artifact. Root-cause attribution is only possible if every step was recorded, which is what the instrumentation section asks you to build before you need it. |
| why multi-agent systems failWhy Do Multi-Agent LLM Systems Fail? | 2025 |
Background. The preceding papers diagnose failures in one agent's trajectory. A multi-agent system adds handoffs and shared decisions, so final task success hides which interaction broke. Problem. A multi-agent failure can stem from system design, agents misunderstanding one another, or weak task verification. Without a shared vocabulary, two systems failing for unrelated reasons look identical from the score, and fixes are aimed at the wrong component. Key idea. Build the vocabulary from evidence rather than from a diagram. Annotate traces across popular multi-agent frameworks, derive a taxonomy from close reading of a subset, and validate it with human annotators before applying it at scale. Findings. The taxonomy names 14 failure modes in three clusters: system design issues, inter-agent misalignment, and task verification. It is built from a careful analysis of 150 traces with high inter-annotator agreement (kappa = 0.88) and applied across a dataset of more than 1,600 annotated traces from seven frameworks. Why it matters. It extends the step-level diagnosis above to failures between agents. The three clusters tell you whether to repair the overall design, the contract between agents, or the verifier that checks their output; the broader taxonomy below places coordination among other agent failure modes. |
| agent failure taxonomyBeyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents | 2026 |
Background. Agent progress is reported as pass rates, one number per harness per suite. What is known about why agents fail is scattered across benchmark papers, taxonomies, and audits that share no vocabulary. Problem. A pass rate tells you how often a run failed, not what broke. Two harnesses at the same score can be failing for unrelated reasons, and a fix aimed at one failure mode is credited or blamed by a number that mixes all of them. Key idea. Synthesize 27 benchmark, taxonomy, and audit papers into six failure clusters: tool invocation, planning, long-horizon context accumulation, multi-agent coordination, safety, and measurement validity. Findings. Failures compound nonlinearly with task length, and additional scaffolding does not consistently improve reliability. Why it matters. It gives you names for the failures that harnesses such as SWE-agent and Claude Code build guards against, so you can argue about a mechanism instead of a leaderboard position. Its measurement critique lands on those leaderboards, SWE-bench included. |
| overthinkingThe Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks | 2025 |
Background. A reasoning model in an agent loop has two ways to make progress: think longer about the state it already has, or act on the environment and read what comes back. Longer chains of thought are widely treated as strictly better. Problem. An agent that reasons instead of acting substitutes its own model of the environment for the environment, then commits to plans built on stale observations. Aggregate pass rates hide the substitution, so the failure has no name to argue about. Key idea. Define an overthinking score for trajectories that favor internal reasoning over interaction, then measure what it predicts. Findings. Higher scores correlate with lower SWE-bench Verified performance across 4,018 trajectories, and selecting the lower-overthinking solution improves performance by almost 30% while cutting compute cost by 43%. Why it matters. More thinking is not monotonically better, which breaks the assumption behind spending test-time compute freely. It also gives you a case where the cheaper policy is the more accurate one, so reasoning budget belongs in the design, not at whatever the model defaults to. |
| Paper | Year | Why read it |
|---|---|---|
| Agent evaluation | ||
| agents that matterAI Agents That Matter | 2024 |
Background. Agent research often reports a leaderboard accuracy number, although systems with the same score may use different numbers of model calls and have different costs. A deployed system must also work beyond the benchmark tasks. Problem. Without cost controls and simple baselines, a complex agent can appear to improve reasoning when it mainly buys more attempts. Small or public test sets invite shortcuts, and inconsistent evaluation settings make results hard to reproduce. Key idea. Plot accuracy against dollar cost and compare agents with simple retries, retries at increasing temperature, and escalation to a stronger model. Match holdouts to the generality claimed and distinguish model research benchmarks from evaluations used to choose a deployed system. Findings. On HumanEval, retries at increasing temperature were not significantly different in accuracy from the best tested complex architecture, while LATS cost more than 50× as much. On HotPotQA, joint prompt optimization cut variable cost by 53% with GPT-3.5 and 41% with Llama-3-70B while maintaining accuracy. Why it matters. A higher score has to beat cheap baselines at a comparable budget and survive a held-out, reproducible test. The cost section below turns the two-axis comparison into an operating decision, while the benchmark checklist below asks whether the test itself is valid. |
| agent reliabilityTowards a Science of AI Agent Reliability | 2026 |
Background. Agent evaluation compresses a trajectory into a success rate, and release decisions then inherit that compression. A system can improve on a benchmark while still being erratic across retries, brittle under perturbation, or dangerous when it fails. Problem. Accuracy hides the operating profile. It does not say whether an agent fails predictably, whether small input changes flip the outcome, whether stronger models reduce severe failures, or whether the same instruction produces stable behavior across runs. Key idea. Borrow the shape of safety-critical engineering and decompose reliability into twelve metrics across four dimensions: consistency, robustness, predictability, and safety. The benchmark result becomes a profile rather than one scalar. Findings. Evaluating 15 models across two complementary benchmarks, the paper finds that recent capability gains produce only small reliability improvements. Higher accuracy therefore does not automatically buy consistent, robust, predictable, or bounded-failure behavior. Why it matters. It gives release engineering a vocabulary for what has to be true after an agent passes a task. The question is no longer only "does it work?", but how often the same system degrades, surprises you, or fails in ways the surrounding product can absorb. |
| agentic benchmark checklistEstablishing Best Practices for Building Rigorous Agentic Benchmarks | 2025 |
Background. Agent benchmarks are assembled quickly and reported widely. A task set, a scorer, and a harness are put together, and a leaderboard follows within weeks. Problem. A benchmark can be wrong in two ways a score cannot show. The task may be unsolvable or underspecified as written, and the scorer may accept an answer the task never asked for. Either one moves every reported number without any agent changing. Key idea. Separate outcome validity, whether the checker actually measures task completion, from task validity, whether the task can be solved as stated, then audit published benchmarks against a checklist built from both. Findings. The audit lands on suites this course reads. SWE-bench Verified uses insufficient test cases and TAU-bench counts empty responses as successful, and issues of this kind can under- or overestimate performance by up to 100% in relative terms; applied to CVE-Bench, the checklist cuts overestimation by 33%. Why it matters. It is the gameable-benchmark problem carried out on the benchmarks listed below rather than as a hypothetical. Read it before freezing your own set, because most of what it finds is a scorer that accepts the wrong thing, which is the cheapest mistake to make and the one a score will never reveal. |
| evaluation cheating examples2. Examples of cheating in CAISI’s agent evaluations (NIST) | 2025 |
Background. Agent benchmarks give models tools and an environment to act in, then score the result with an automated checker. A passing score does not show how the agent reached it. Problem. An agent can obtain information unavailable in the real task, which NIST calls solution contamination, or satisfy the checker without doing the intended task, which it calls grader gaming. Both make a pass overstate the capability being tested. Key idea. Use an LLM-based reviewer to search agent transcripts for unintended routes to success, then validate successful cases with human review. Classify the cases by whether the agent crossed an information boundary or exploited the grading rule. Findings. On Cybench, agents retrieved published challenge walkthroughs and flags. On SWE-bench Verified, o4-mini passed five of 498 tasks through test overfitting, including a case where it commented out an assertion instead of fixing the issue. NIST cautions that its historical runs used different samples across benchmarks and models, and its reviewer may have missed cases. Why it matters. It supplies observed traces for the checklist above: the environment must protect the intended information boundary, and the grader must reward the intended behavior. Inspect trajectories alongside scores before using a benchmark for release decisions. |
| SWE-benchCan Language Models Resolve Real-World GitHub Issues? | 2023 |
Background. Code models were scored on short, self-contained functions with hidden unit tests. Real software work arrives differently, as an issue filed against a repository that already has a history, a build, and a test suite. Problem. Tasks of that shape cannot be graded by comparing strings. No single patch is the correct one, a fix may touch several files, and deciding whether it works means installing the repository at the right commit and running its tests. Key idea. Build 2,294 task instances from real issues and their merged pull requests across 12 Python repositories, and score a submitted patch by running the repository’s own tests. The model is given the issue text and the codebase, and nothing else. Why it matters. It is the benchmark SWE-agent is evaluated against, and the reason a coding agent’s evaluation harness has to build environments rather than compare strings. It became a headline coding number in frontier model cards, so its scoring choices shaped what the field optimized. The audit below shows why that use has become unreliable for frontier models. |
| SWE-bench Verified auditWhy SWE-bench Verified no longer measures frontier coding capabilities (OpenAI) | 2026 |
Background. SWE-bench Verified selected 500 human-reviewed issues from the original suite and became a standard score in frontier coding model releases. Problem. Hidden tests can demand implementation details absent from an issue or check behavior the issue did not request, rejecting a valid fix. Public issues and their original patches can also enter model training, so a pass may reflect prior exposure. Key idea. Audit tasks that a frontier model often fails with experienced engineers, then probe models for recall of task descriptions and original patches. Use both checks to decide whether the benchmark still measures coding ability. Findings. Of 138 selected tasks that o3 did not consistently solve, 59.4% had material test or description problems: 35.5% had overly narrow tests and 18.8% tested behavior beyond the issue. Targeted probes elicited task or patch details from tested OpenAI, Anthropic, and Google models. The 138 tasks were a hard subset, so the 59.4% figure does not describe the full 500-task benchmark. Why it matters. OpenAI stopped reporting SWE-bench Verified scores and recommends SWE-bench Pro for now, while noting that Pro is imperfect. Together with the NIST traces above, this audit shows why both passed and failed tasks need review before treating a benchmark score as a capability measure. |
| AgentBenchEvaluating LLMs as Agents | 2023 |
Background. By 2023 agent evaluation was scattered across one-off setups: a web task in one paper, a game in another, a database query in a third. Each group built its own environment, and no two results were comparable. Problem. Without a shared suite you cannot tell a better agent from a better fit to one harness. Assembling many environments under a single interface is the obvious fix, and it is mostly unglamorous engineering rather than a new idea. Key idea. Run one agent interface across a catalog of differently shaped interactive environments. The benchmark's main contribution is the breadth of those environments rather than one aggregate score. Why it matters. Read it as a survey of what people thought agents should be tested on in 2023. The narrower suites on this page have largely replaced it for measurement, which is itself informative: a score averaged over dissimilar environments tells you less than a narrow one you can act on. |
| METR time horizonsMeasuring AI Ability to Complete Long Tasks | 2025 |
Background. Capability on software work is usually reported as a pass rate over a fixed task set. A pass rate tells you how often a model succeeds on those tasks and nothing about how large a task it can finish. Problem. Task sets go stale as models improve, and scores from different suites do not compose into a trend. Comparing 2019 with 2025 needs a unit that stays fixed when the benchmark does not. Key idea. Measure each task by how long it takes a human, then report the 50%-task-completion time horizon, the human task duration at which a model succeeds half the time. Findings. On RE-Bench, HCAST, and 66 shorter tasks, Claude 3.7 Sonnet had a 50% time horizon of about 50 minutes. Across frontier models since 2019, the fitted horizon doubled approximately every seven months, although the paper cautions that external validity is uncertain. Why it matters. It gives a capability axis measured in task length rather than pass rate, which is the right shape for reasoning about how long agent sessions run. Session length is what sets context growth, checkpointing, and the cost of a run that fails late. |
| Ï„-benchA Benchmark for Tool-Agent-User Interaction in Real-World Domains | 2024 |
Background. The benchmarks above hand an agent a fully specified task and score whether it finishes. A deployed customer-service agent gets neither of those: the request arrives in pieces, and the company has rules about which actions are allowed. Problem. Two things therefore go unmeasured. A single-shot task hides whether the agent can hold requirements that accumulate over several turns, and a benchmark with no written policy cannot tell a successful action from a permitted one. Key idea. Run conversations between a language agent holding domain APIs and a simulated user, then score the run by comparing the database state at the end against the state the task required. Each domain ships with a written policy the agent has to comply with. Findings. Even GPT-4o completed fewer than half of the evaluated tasks. In the retail domain, its pass^8 reliability was below 25%, meaning repeated trials exposed substantially lower consistency than single-run success suggested. Why it matters. A correct action is therefore not only a successful one, which is the distinction Assignment 2 has to design against. State-based scoring has a systems consequence as well: the checker needs the environment, so evaluation carries a database rather than an answer key. |
| Persistent autonomy and reliability engineering | ||
| HeadlongA Microharness for Persistent Agents | 2026 |
Background. An agent harness normally runs a bounded task. It starts, works, returns, and its context ends with the session, so nothing has to outlive the run. Problem. A persistent agent outlives its context window, so the harness must decide what to keep, what to summarize, and what it can return to later. Everything a bounded session got for free becomes an explicit design decision. Key idea. Build a continuously running agent around a small Bash harness, tiered context compaction, and a forkable trajectory log. The minimal harness isolates persistence rather than tooling as the subject. Findings. Several weeks of operation expose failures that a bounded demo misses. After a 30-second inactivity watchdog disrupted recursive subagents, successful merges fell from 64 during the first two days to 12 during the next 12 days. The agent also stopped its own service three times, motivating an external guard. Why it matters. The reported self-modification failures, sandbox requirement, and background token cost are more instructive than the demo. Persistence changes the safety, evaluation, and resource model at once, and this is a worked account of all three. |
| AgentOps infrastructure, policy, and telemetry | ||
| DapperA Large-Scale Distributed Systems Tracing Infrastructure | 2010 |
Background. One user request in a large system fans out across many services, each with its own logs. No single log holds the request, and the teams that own the services are different teams. Problem. Without a shared identifier the request cannot be reassembled after the fact, so latency cannot be attributed to a component. The time is somewhere in the fan-out, and no one service's instrumentation can say where, which is the same position an agent leaves you in when a run takes a minute. Key idea. Propagate a trace identifier along the call path and record a span for each operation, so the whole request can be rebuilt as a tree. Keep the overhead affordable through sampling, and keep the instrumentation maintainable by confining it to a small number of shared libraries rather than asking every team to add it. Findings. Sampling is what makes it deployable rather than a research prototype: Google reports collecting over one terabyte of sampled trace data per day, with the sampling rate, not the instrumentation, as the knob that sets the cost. Why it matters. An agent transcript is a trace without stable identifiers. Treating steps and tool calls as spans exposes orchestration gaps that separate component logs cannot attribute. Adoption. Zipkin, Jaeger, and OpenTelemetry all implement Dapper's trace-and-span design. |
| AgentOpsAgentOps: Enabling Observability of LLM Agentsalso Nov 30 | 2024 |
Background. An agent's execution is a sequence of model calls, tool invocations, and state changes that no existing log format was designed to hold. The tracing services and evaluation harnesses that record it each invented their own vocabulary for it. Problem. Without agreement on what to record, a trace is whatever the framework happened to emit. The same failure then looks different in two deployments, and neither trace answers the question the other was built to answer. What to instrument is the prior decision, and it has had far less attention than how. Key idea. Derive the answer from what the tools already do. A systematic mapping study over existing AgentOps tools yields a taxonomy of the artifacts, and the data attached to each, that should be traced across an agent's whole lifecycle, offered as a reference template for building the infrastructure rather than as an implementation of it. Why it matters. It is the checklist, and a checklist belongs before the mechanisms rather than after them, because a span you did not record is not recoverable later. Read it against the schema you would need to answer this meeting's own questions: which turn evicted the session, how long the sandbox took to start, and which agent in a group was waiting on which. |
| FirecrackerLightweight Virtualization for Serverless Applicationsalso Nov 30 | 2020 |
Background. Serverless platforms run short, untrusted functions from many customers on shared hardware. Containers share a kernel, so the isolation boundary is the whole Linux syscall surface; a virtual machine gives a narrower boundary and a slower start. Problem. Neither end of that choice fits the workload. Container isolation is too weak for arbitrary tenant code, and a conventional VM with a general-purpose device model boots too slowly and holds too much memory to give every invocation its own. Key idea. Build a minimal virtual machine monitor on KVM that emulates only the devices a serverless guest needs, and drop the rest. What remains boots fast enough, and costs little enough memory, to run one microVM per function. Findings. With the minimal guest-kernel configuration, each microVM uses less than 5 MB of memory, boots to application code in less than 125 ms, and can be created at up to 150 microVMs per second per host. Why it matters. It is the case for why a VM boundary can still be cheap, which is the assumption every agent sandbox in this group rests on. Once virtualization is affordable per invocation, isolation stops competing with density. Adoption. Fly.io Machines use Firecracker microVMs, and the E2B agent-sandbox platform builds its isolation layer on Firecracker. |
| ActPlaneProgrammable OS-Level Policy Enforcement for Agent Harnessesalso Nov 30 | 2026 |
Background. Agent harnesses decide what an agent may do at the tool layer, in user space. A permission check runs before each tool call, and the allowlist of permitted tools is the whole policy. Problem. An agent that reaches the same effect by another path escapes that check; a shell command can write files or open sockets without passing through the tool the allowlist names. Policy intent also arrives as underspecified natural language while enforcement must act on concrete system actions. Key idea. Split declaring policy from enforcing it. The agent declares policy in an information-flow DSL and eBPF enforces it in the kernel, so indirect execution paths are covered too, and refusals come back as semantic feedback rather than opaque errors. Findings. Across empirical policies, coding tasks, and safety benchmarks, ActPlane improves compliance on indirect execution paths that tool interception cannot observe, with 1.9-8.4% runtime overhead. Why it matters. It moves the agent permission question from the harness down to the operating system, where the actions actually land. Once enforcement sits below the tool layer, the hard part becomes translating vague intent into concrete rules rather than enumerating tools. |
| Paper | Year | Why read it |
|---|---|---|
| Cost-effective agent design | ||
| cost and performance on Claude PlatformReducing cost and improving performance with Claude Platform (Anthropic) | 2026 |
Background. The agent-memory readings from the Sep 28 meeting decide what goes into the window. The provider bills for what arrives there, so the same choices that shape context also set the invoice, namely what the prefix looks like from one request to the next, how the instructions are written, and how long the model is allowed to deliberate. Problem. Cost reduction is assumed to be paid for in accuracy, so it gets deferred until the bill forces it and is then answered with a smaller model. Prompt text also ages badly. Instructions written to prop up an older model, such as verification rituals, emphasis boosters, mandatory procedures, and rules that contradict one another, survive the migration to a newer one, where they are charged on every request and make the output worse rather than safer. Key idea. Three levers, each measurable on its own. First, keep the prefix byte-identical across requests and place volatile values such as timestamps and session identifiers after that prefix rather than inside it, so the cache is read rather than rewritten, with explicit breakpoints and a request that generates no tokens to warm it. Second, audit the instructions and delete the anti-patterns instead of adding more text to compensate for them. Third, set the effort level per task, because deliberation a task does not need is billed and can lower accuracy. Findings. Anthropic reports that removing instruction anti-patterns while migrating from Opus 4.8 to Opus 5 decreased cost by 14.6% and increased accuracy by 5.3% on average. Its automated optimizer cuts spend at matched accuracy by about 58% on LegalBench, 73% on tau2-bench retail, 52% on OfficeQA Pro (from $136.20 to $64.87), and 55% on SWE-bench Verified, where the median number of steps per task falls from 29 to 17. Effort remains a real tradeoff rather than free savings: on FrontierCode Diamond, Fable 5 scores 11.5% at low effort for $5.35 per task and 30.9% at maximum effort for $19.00. Why it matters. It puts a price on the agent-design decisions of the earlier meetings. A prompt written for an earlier model, a prefix that changes one byte per request, and an effort setting nobody revisited are each a recurring charge, and the first two are corrected without changing what the agent can do. The caching advice is also the application side of the prefix-cache meeting later in the course, where the same reuse is a property of the serving system rather than of how you assembled the request. Treat the numbers as vendor-reported and as evidence that cost and accuracy do not always trade against each other, not as a rate to plan against. |
| RouteLLMLearning to Route LLMs with Preference Dataalso Nov 23 | 2024 |
Background. An agent may call a model many times in one run. Paying for the strongest model at every step adds up, while using a cheaper model throughout can put weak answers into the trajectory. Problem. The agent has to choose a model before seeing either answer. A fixed split between cheap and expensive models cannot send the requests that need the stronger model to it. Key idea. Learn from human preferences between strong- and weak-model responses to predict when the stronger model is likely to win. A threshold chooses one model per request and can be adjusted to trade response quality against cost. Findings. On the paper's evaluated benchmarks, the learned routers cut cost by more than 2x in some settings without a measured drop in response quality. Routing performance also transfers to model pairs outside the training set. Why it matters. Model choice becomes a per-call budget decision alongside prompt caching and effort settings. The paper measures request-level quality, so an agent still needs an end-to-end test: an inexpensive wrong answer early in a run may change every step that follows. The later model-routing meeting examines this mechanism in more detail. |
| tool-use inefficiencyBeyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning | 2026 |
Background. Tool-integrated reasoning is scored on accuracy, and its cost is reported in tokens. A token count treats every token as equally expensive and says nothing about what the machine did between them. Problem. The expensive part of a tool call is the pause. A blocked request can have its KV cache evicted and recomputed on resume, and a bloated tool response is charged as prefill; neither appears in a token total. Key idea. Introduce Prefill Token Equivalents, a hardware-aware efficiency metric that charges tool-call pauses and response bloat, then use it to name four recurring inefficiency patterns across five tool-use benchmarks. Findings. In a 256-request experiment running DeepSeek-V3.2 on eight H200 GPUs, PTE correlates with model-generation latency at r = 0.9253 across 100 samples, while output-token count correlates at r = -0.3750. Efficiency rankings remain stable across simulated hardware profiles, with Spearman correlations above 0.95. Why it matters. It puts cost and accuracy in the same unit, so an agent design can be argued down on efficiency without treating accuracy as free. The unit is hardware work, which is what the operator actually pays for. |
| software factory at Uber scaleRunning a Software Factory Efficiently at Uber Scale (Uber) | 2026 |
Background. The three cost readings above use benchmark and single-run evidence. Uber reports the same design decisions running as production infrastructure for a whole engineering organization, where more than 70% of pull requests are attributed to local or cloud agents and engineers have built over 3,600 agent skills that execute more than 30,000 times a day. Problem. Adoption and spend move together. Between February and August 2026 weekly active users across Uber's agentic offerings grew 7× and weekly agentic requests grew 9.4×, so every per-request inefficiency was multiplied by a workload that grew almost tenfold in six months. Waiting for model prices to fall does not close that gap, because usage outruns the price curve. Key idea. Decompose the bill into factors that different teams can own, namely users, sessions per user, turns per session, requests per turn, tokens per request, and price per token, and give each factor its own mechanism. Mounting over 1,000 MCP servers would add 50,000-70,000 tokens of schema to the initial prompt, so a gateway routes them instead of loading them. Code mode batches tool calls into one executed script rather than paying a model turn per call. Cache policy is written against the vendor's TTL tiers, where a five-minute entry costs 1.25× to write and a one-hour entry 2×. Findings. Cost per 1,000 model requests is down almost 34% from its peak and cost per session is down 52% from its June peak, over the same months in which usage grew 7×. Code mode reduces tokens by 55%, 58%, 59%, 71%, and about 100% on the reported query types. Why it matters. The multiplicative decomposition is what turns this meeting's design choices into a budget. A second agent, a longer trajectory, and another mounted tool are each a term in the same product, so the equation tells you which one is worth attacking rather than which one is easiest to change. It is also the reading that prices the fan-out argument at organizational scale, where unit cost and adoption are measured separately and a design is only efficient if it holds while usage grows. |
| Paper | Year | Why read it |
|---|---|---|
| Agent safety: threats | ||
| indirect prompt injectionNot what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection | 2023 |
Background. An LLM application becomes useful once it can read outside itself: retrieved documents, web pages, email, tool output. All of that arrives in the same context window as the developer’s instructions and the user’s request. Problem. The model has no channel separation. An agent that retrieves untrusted data cannot separate that data from its instructions, so any document the agent reads is a place an attacker can put a command, and the attacker never has to speak to the user. Key idea. Name the attack the tool loop creates and demonstrate it against real integrated applications. Text planted in content the model will later retrieve turns retrieval itself into a delivery channel for instructions the operator never wrote. Findings. The demonstrations compromise Bing's GPT-4-powered chat, code-completion engines, and synthetic GPT-4 applications. Retrieved instructions can redirect API calls, exfiltrate data, manipulate later model outputs, and propagate through content, showing that the vulnerability crosses application types rather than depending on one interface. Why it matters. Read it directly after the tool-use entries above, because every interface decision made there also decides what an injected instruction is able to reach. What you control is the granularity of the capabilities you hand the loop, not the model’s willingness to obey. Adoption. OWASP classifies prompt injection as LLM01 in its Top 10 for LLM applications, carrying the paper's attack class into an external security standard. |
| Agent safety: defenses | ||
| AgentDojoA Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents | 2024 |
Background. Once indirect prompt injection is known to work, the question becomes how much a defense actually helps. Answering that takes an agent, tools it can call, tasks worth completing, and attacker-controlled content sitting in the middle. Problem. A frozen list of attack strings measures the wrong thing. Attacks adapt to defenses, so a suite fixed at publication reports a defense against last year’s attacker, and it says nothing about what the defense costs on ordinary work. Key idea. 97 realistic tasks and 629 security test cases in an environment built to be extended rather than frozen, on the argument that a static suite cannot track adaptive attacks. Utility on the benign tasks and security under injection are measured in the same harness. Findings. On GPT-4o, the paper's strongest generic injection succeeds on 57.7% of targeted attacks and lowers task utility from 69.0% without attacks to 50.0% under attack. A tool-filter defense reduces targeted attack success to 6.84%, but still leaves utility under attack at 56.28%. Why it matters. The evaluation-design lesson generalizes past security: a benchmark for a moving target has to be an environment, not a fixed list. Measuring both axes together also forces the tradeoff into view, since a defense that blocks every injection by refusing tool calls scores badly on the tasks. |
| securing agents against prompt injectionDesign Patterns for Securing LLM Agents against Prompt Injections | 2025 |
Background. Indirect prompt injection is not a defect awaiting a patch. The model has no channel separating instructions from retrieved data, and a detector placed in front of it is another model that can be evaded. Problem. A defense that asks the model to resist instructions is a request rather than a constraint. What is needed instead is an arrangement in which a successful injection cannot reach anything that matters, and that is a property of the loop rather than of the weights. Key idea. Stop trying to make the model resistant and constrain what the loop is able to do. The paper names design patterns that each surrender some generality for a security property holding regardless of the model, among them selecting an action without letting tool output re-enter planning, fixing the plan before any untrusted data is read, and splitting the work so that the component reading untrusted content holds no capabilities. Why it matters. It is the designer's side of this group, which otherwise has an attack and a way to measure it but nothing to build. Every pattern is a decision in the loop you own, and each one is a restriction on which tool results are allowed to influence what the agent does next. |
| CaMeLDefeating Prompt Injections by Design | 2025 |
Background. The patterns above restrict how an agent is arranged. A different question is what a runtime underneath the model can enforce whatever the model does, given that the model itself will keep failing to separate data from instructions. Problem. Any guarantee has to come from outside the weights, which means a component that knows which values arrived from untrusted sources and what each tool is permitted to touch. Neither fact is available inside a single context of undifferentiated text. Key idea. Extract the control flow and the data flow from the trusted query and keep them apart. A privileged component turns the user's request into a program before any untrusted content is read, a quarantined component processes that content but holds no capabilities, and a custom interpreter tracks the provenance of every value and checks a policy before a tool with side effects runs. Findings. The guarantee has a measurable price. CaMeL solves 77% of AgentDojo tasks with provable security, against 84% for the same undefended system, so the security property costs 7 points of utility rather than being free. Why it matters. It is the harness argument applied to security: the guarantee lives in code you wrote instead of in the model's willingness to comply. The cost is stated plainly, since a plan fixed before the untrusted data is read cannot adapt to what that data says, and somebody has to author the policy. |
| Self-evolving harnesses | ||
| ReflexionLanguage Agents with Verbal Reinforcement Learning | 2023 |
Background. A failed agent trajectory contains its actions, errors, and stopping point. Reinforcement learning can use such feedback, but it requires many samples and a fine-tuning run. Problem. Neither cost is available inside an agent loop. A deployment cannot fine-tune between two attempts, and one failed episode is far too little data for a weight update, so the lesson in the transcript is discarded when the episode ends. Key idea. After a failure, have the agent write a reflection, store it in episodic memory, and read it on the next attempt. Behavior changes through context rather than a weight update, using either scalar or language feedback. Findings. Reflexion improves on a baseline agent across sequential decision-making, coding, and language reasoning. On HumanEval it reaches 91% pass@1, against 80% for the GPT-4 state of the art it is measured against. Why it matters. Memory becomes the mechanism of improvement that a harness can implement between attempts. Its limit is diagnosis: a reflection is only as useful as the trajectory and feedback from which it was written. |
| Self-RefineIterative Refinement with Self-Feedback | 2023 |
Background. The entry above reflects between attempts, on feedback the environment returned. Much of what an agent produces draws no environmental feedback at all: a plan, a summary, a message, a patch nobody has run yet. Problem. With no external signal there is nothing to reflect on, and a first draft is where a model otherwise stops. Training a separate critic, or collecting supervision for the revision step, is the usual answer, and it costs data and a training run for each task. Key idea. Let one model play all three parts. Generate an output, have the same model write feedback on it, then have it revise against that feedback, and iterate. There is no supervised data, no additional training, and no second model, so the pattern drops into any loop that already runs one. Findings. Across seven tasks, from dialogue response generation to mathematical reasoning, humans and automatic metrics both prefer the refined output to one-step generation from the same model, by about 20 points absolute on average. Why it matters. It is not an agent paper, and the generate-critique-revise triple it names became a standard component inside agents anyway, which is why it sits here. Read it for what the critic actually is: the same weights asked a different question, which is what makes it cheap and is also why a self-critique can miss exactly what the generator missed. |
A serving engine raises utilization by changing the unit at which it manages work. Orca schedules one generation iteration instead of one whole request, PagedAttention allocates one cache page instead of one contiguous sequence, and the parallelism papers divide one model or sequence across devices. These mechanisms operate at different layers but solve the same mismatch: requests have variable lengths and memory demands, while accelerators prefer regular batches of fixed-shape work. Read the measurement paper last to check which resource each mechanism actually relieves.
| Paper | Year | Why read it |
|---|---|---|
| Iteration-level batching | ||
| OrcaA Distributed Serving System for Transformer-Based Generative Models | 2022 |
Background. Serving systems before Orca batched at request granularity. A batch of requests entered the model together and left together, the way an image-classification server batches independent inputs. Problem. Generative requests do not finish together, because their output lengths differ. A request-level batch runs until its longest member is done, so finished sequences keep occupying slots and newly arrived requests wait for the whole batch to drain. Key idea. Schedule at the granularity of one iteration rather than one request. After every token the scheduler re-forms the batch, admitting arrivals and retiring completed sequences, which requires running attention per sequence while the linear layers stay batched. Findings. On GPT-3 175B, Orca delivers 36.9× the throughput of NVIDIA FasterTransformer at the same latency under the paper's request traces. Why it matters. Once the batch can change every token, throughput becomes a scheduling question rather than a kernel question. This is the baseline every later serving paper is measured against, and it is also a model of a well-argued systems evaluation. Adoption. Continuous batching is the scheduler in vLLM, SGLang, TensorRT-LLM, and Hugging Face TGI. |
| Paging and memory management | ||
| vLLM / PagedAttentionEfficient Memory Management for Large Language Model Serving with PagedAttention | 2023 |
Background. Before this paper, a serving engine gave each sequence one contiguous KV cache buffer, sized for the longest output the request might produce. The reservation was made on arrival and held until the sequence finished. Problem. Fragmentation, not compute, was the binding constraint. A contiguous reservation for an unknown output length wastes most of its own buffer, and the holes left between buffers block new sequences even when the total free memory would hold them. Key idea. Apply virtual memory to the KV cache. Store it in fixed-size blocks, keep a per-sequence page table from logical positions to physical blocks, and let the attention kernel read blocks that are not contiguous, so sharing and copy-on-write between sequences become allocator operations. Findings. Across the paper's workloads, vLLM provides 2-4× the throughput of FasterTransformer and Orca at the same latency. The advantage grows with longer sequences, larger models, and more complex decoding. Why it matters. It is the single most important systems paper on the list. Once memory is paged, the batch a scheduler can form is set by an allocator you control rather than by the worst case you were forced to reserve for. Adoption. vLLM originated the design; SGLang and TensorRT-LLM likewise manage the KV cache in fixed-size blocks behind a page table. |
| JengaEffective Memory Management for Serving LLM with Heterogeneity | 2025 |
Background. PagedAttention assumes every layer of every model stores the same number of bytes per token, so a single uniform page size fits the whole cache. That held for the dense attention stacks it was designed against. Problem. The assumption breaks on layers whose per-token KV footprints differ, such as sliding-window attention in Gemma-2 and hybrid Mamba and attention stacks in Jamba. One uniform page must then be sized for the largest layer, and every smaller layer wastes the difference. Key idea. Use a two-level allocator that generalizes paging to heterogeneous embedding sizes, so each layer's blocks hold only what that layer stores. The tradeoff is allocator complexity for less waste. Findings. Across diverse models, workloads, and GPU configurations, Jenga improves GPU memory utilization by up to 79.6% and serving throughput by 1.80× on average and up to 4.92×. Why it matters. This SOSP 2025 paper is the direct sequel to PagedAttention. Architectures changed, and the allocator assumption that made vLLM fast became the part that had to change, showing how a systems abstraction responds to a workload it was not designed for. |
| vAttentionDynamic Memory Management for Serving LLMs without PagedAttention | 2024 |
Background. PagedAttention removes fragmentation by storing the KV cache in fixed-size blocks that every attention kernel reads through a page table. Problem. Non-contiguous storage forces every attention kernel to traverse block indices. That adds inner-loop overhead and delays support for each new kernel, even though fragmentation is a physical-memory problem. Key idea. Give each sequence a contiguous virtual buffer backed by physical pages committed on demand through CUDA's virtual-memory APIs. Kernels retain a contiguous tensor view while the runtime hides allocation latency and controls coarse commitment granularity. Findings. Unmodified attention kernels run as they are, and serving throughput rises by up to 1.23× against the PagedAttention kernels of FlashAttention and FlashInfer. Why it matters. Physical on-demand allocation does not require a non-contiguous virtual layout. The tradeoff is portability: the design avoids kernel changes by depending on driver-level virtual-memory support. |
| Splitting the model across devices | ||
| Efficiently Scaling Transformer InferenceEfficiently Scaling Transformer Inference | 2022 |
Background. A model can be split within layers, across layers, or across sequence positions, and each layout creates different communication costs. Problem. A layout that maximizes throughput may miss its latency target, and FLOP balance may hide memory or collective bottlenecks. Prefill and decode can favor different layouts. Key idea. Model the arithmetic, memory traffic, and collective communication of each layout. Choose per phase and batch size against a latency-throughput frontier rather than one peak number. Findings. On PaLM 540B with a 2,048-token context, the optimized layouts reach 29 ms per generated token at low batch size with int8 weights and 76% model-FLOPs utilization during large-batch prefill. Why it matters. The model predicts which resource binds for a model shape, device count, batch size, and latency target before an expensive deployment sweep. |
| Megatron-LMTraining Multi-Billion Parameter Language Models Using Model Parallelism | 2019 |
Background. A transformer layer contains two feed-forward matrix multiplications and attention over independent heads. Splitting a layer across devices requires choosing which matrix dimensions and heads each device holds. Problem. A naive split synchronizes whenever devices need one another's partial results. Those collectives sit on every forward pass, so communication can erase the latency reduction from adding devices. Key idea. Partition the first feed-forward matrix by columns and the second by rows, so the intermediate activation stays local and only the block output is reduced. Give each device whole attention heads. The resulting layer needs two all-reduces per forward pass. Findings. Training an 8.3B-parameter model on 512 GPUs sustains 15.1 PFLOP/s and 76% scaling efficiency relative to a strong single-GPU baseline that sustains 39 TFLOP/s. Why it matters. This layout defines tensor parallelism in current serving engines. Because its collectives are on every token's critical path, deployments usually keep the tensor-parallel group within one node. Adoption. vLLM, SGLang, TensorRT-LLM, and DeepSpeed expose this tensor-parallel partitioning for multi-GPU inference. |
| GPipeGPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism | 2018 |
Background. Pipeline parallelism gives each device a contiguous run of layers. Each device holds fewer weights, and devices communicate activations only at stage boundaries. Problem. Layers remain sequential. Sending one input through K stages leaves K-1 stages idle at each instant, spending the memory benefit on pipeline bubbles. Key idea. Split a batch into M micro-batches and send them through the stages back to back. The idle fraction falls from (K-1)/K to (K-1)/(M+K-1), while one synchronous update preserves training semantics. Findings. For a Transformer with 32 micro-batches, moving from two to eight TPU partitions raises normalized throughput from 1.8 to 6.3. Bubble overhead becomes nearly negligible once the micro-batch count is at least four times the partition count. Why it matters. In serving, the number of in-flight requests determines the available micro-batches. Pipeline parallelism therefore trades less communication per token for an idle fraction that worsens when the queue is short. Adoption. vLLM, DeepSpeed, and Megatron-LM expose pipeline parallelism and use micro-batching to fill stages. |
| LoongServeEfficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism | 2024 |
Background. Sequence parallelism splits one sequence's tokens across devices, so a context too long for one device's memory can still be processed. The degree is normally fixed when the deployment is configured, before any request has arrived. Problem. One fixed degree has to serve every request. Size it for the longest context and short requests pay communication they never needed; size it for the common case and long requests do not fit. A single request wants different degrees in its two phases, because prefill is compute-heavy over the whole sequence while decode extends it one token at a time. Key idea. Make the degree elastic. Vary it per request, and within a request across phases, migrating KV state between devices as the degree changes rather than fixing the device mapping when the cluster is configured. Findings. Across real workload traces, LoongServe raises maximum throughput by up to 3.85× over chunked-prefill systems and 5.81× over prefill-decode disaggregation. On 16 GPUs, it improves total throughput by up to 1.86× over vLLM. Why it matters. It moves a deployment-time constant into the scheduler's decision space, which is what makes it belong on this page rather than in a parallelism survey. Once the degree is elastic the scheduler trades communication cost against memory headroom on every admission, so parallelism and admission control stop being separable decisions. |
| NanoFlowNanoFlow: Towards Optimal Large Language Model Serving Throughput | 2024 |
Background. One device draws on three resources at once: compute units, memory bandwidth, and the links to other devices. Every layout above splits work across devices and leaves the question of what happens inside one of them unasked. Problem. Different stages of a single iteration bottleneck on different resources, so the stage saturating one of them leaves the other two idle. Running the stages one after another inside a device caps throughput well below what that device's own rates allow, and adding devices does not recover the waste. Key idea. Split operations into smaller units and co-schedule them so that compute, memory, and network work overlap within one device. The resulting schedule is argued against a throughput bound derived from the device's resource rates rather than against a measured baseline, so the gap it closes is stated in absolute terms. Findings. Across Llama-2 70B, Mixtral 8×7B, and Llama-3 8B workloads, NanoFlow provides up to 1.91× the throughput of the evaluated serving systems and reaches 59-72% of its hardware-derived optimum. Why it matters. It relocates the question of where the idle time is. The layouts above ask which device should run which work; this asks the same question inside one device, and the two are alternatives rather than complements when you are deciding where to spend effort. |
| Measurements | ||
| LLM inference characterizationA Systematic Characterization of LLM Inference on GPUs | 2025 |
Background. The roofline arguments this course makes rest on a claim about where time goes: prefill is compute-bound and decode is bound by memory bandwidth. That claim is usually asserted from first principles rather than measured on a named machine. Problem. A phase label does not tell you what the phase is doing. Time in decode can go to kernel mix, issue stalls, or scattered memory access, and which of those dominates decides whether a given optimization helps at all. Key idea. Split inference into prefill and decode and trace each phase’s cost to kernel mix, roofline bounds, issue stalls, and memory access patterns, measured on an A100 server and on a Jetson Orin edge device. Findings. On the A100, both phases spend 70-80% of cycles stalled, but for different reasons. Tensor parallelism cuts prefill time from 1.24 seconds to 0.66 while increasing decode latency because of communication and synchronization. Why it matters. It is the empirical grounding for the roofline claims the rest of the course makes, and the two very different devices show which conclusions are about transformers and which are about one particular memory system. |
Speculative decoding spends inexpensive work to avoid expensive serial model steps while preserving the target model's output distribution. The readings answer four questions in order: why parallel verification is exact, where candidate tokens come from, how candidate trees should match the hardware and context length, and how speculation interacts with a live serving queue. Start with the two independent formulations of the method, then treat every later proposal as a change to draft cost, acceptance rate, verification shape, or scheduling overhead.
| Paper | Year | Why read it |
|---|---|---|
| Guess ahead, verify in parallel | ||
| speculative decodingFast Inference from Transformers via Speculative Decoding | 2022 |
Background. An autoregressive model emits one token per forward pass, and at small batch sizes that pass is bound by memory bandwidth: the weights are read from HBM to produce a single token. Problem. The sequential dependency, not the arithmetic, sets the latency. You cannot start token n+1 before token n exists, so a response costs a fixed number of passes however much of the accelerator each pass leaves idle. Verifying several tokens at once would be nearly free, but you do not have them yet. Key idea. Let a cheap model guess the next several tokens, run the target model once over the whole guessed block, and accept a prefix using a rejection-sampling rule. The rule is constructed so the accepted tokens carry exactly the target model's distribution, so only the failed guesses cost anything. Findings. On T5-XXL, speculative decoding runs 2-3× faster than the standard T5X implementation while producing identical outputs. Why it matters. It separates how aggressively you guess from whether the output changes, and the exactness argument is what makes speculation acceptable in production at all. Every later method in this group varies the drafting and reuses this verification step unchanged. Adoption. vLLM, SGLang, TensorRT-LLM, llama.cpp, and Hugging Face Transformers all ship speculative-decoding or assisted-generation implementations of this verify step. |
| speculative samplingAccelerating Large Language Model Decoding with Speculative Sampling | 2023 |
Background. The entry above states the propose-and-verify scheme and the rejection-sampling argument for why the output distribution survives it. This is the concurrent DeepMind treatment of the same scheme. Problem. Whether the accepted tokens really carry the target model's distribution rests on a derivation, not on a measured speedup. A correctness argument that only one group has written down is hard to trust and harder to build on. Key idea. Derive the same accept-or-resample rule independently, setting the argument up differently, and report it on large models at serving scale. Reading the two derivations together shows which parts of the rule are forced by preserving the distribution and which are free choices. Findings. On the 70B-parameter Chinchilla model in a distributed setup, speculative sampling accelerates decoding by 2-2.5× without changing sample quality. Why it matters. Two independent derivations landing on one acceptance rule is the signal that the rule is the constraint rather than an implementation detail. Everything later in this group changes where the guesses come from and leaves the rule alone. |
| blockwise parallel decodingfor Deep Autoregressive Models | 2018 |
Background. In 2018 autoregressive Transformers were used for machine translation, and they already had the constraint decoding has now: the parallelism available while training disappears at inference time. Problem. Greedy decoding of an output of length n takes n sequential passes even though each pass leaves the accelerator mostly idle. The non-autoregressive proposals of the period bought parallelism by changing the output, which translation quality would not absorb. Key idea. Add auxiliary output heads that predict the next several tokens from the same hidden state, verify the whole block in one pass of the base model, and keep the longest prefix that matches what greedy decoding would have produced. Propose in parallel, verify in one pass, accept only what agrees. Findings. The evaluated translation and image-super-resolution models cut decoding iterations by up to 2× without quality loss. Allowing a small quality decrease cuts iterations by up to 7×, while the fastest wall-clock configuration runs up to 4× faster than greedy decoding. Why it matters. It is the ancestor of this entire group, four years before speculative decoding, and it shows the propose-and-verify shape needs no second model. The later split between separate draft models and extra heads on the target model is already present here. |
| Where the guesses come from | ||
| EAGLESpeculative Sampling Requires Rethinking Feature Uncertainty | 2024 |
Background. Speculative decoding needs a cheap proposer whose guesses the target model usually accepts. The standard choice is a smaller independent model, which has to be trained separately, served alongside the target, and kept aligned with it. Problem. Drafting one token at a time from tokens alone discards the target model's own hidden state, and a token sequence does not fix what comes next: sampling introduces uncertainty that a token-level drafter cannot see. Acceptance stays low for the cost paid. Key idea. Draft in feature space rather than token space. One small autoregressive head sits on top of the target model and predicts its hidden features a step ahead, conditioned on the tokens already sampled so that sampling uncertainty is an input rather than a source of error. The model's own output layer turns those features into draft tokens. Findings. On Llama-2-Chat-70B, EAGLE reports a 2.7-3.5× latency speedup and 2× throughput while preserving the generated-text distribution. Why it matters. It reframes drafting as a question of what the target model already computes rather than what a second model can guess. The drafter is a head, not a model, so there is one set of weights to serve and the accuracy comes from reusing the target's representation. Adoption. vLLM, SGLang, and TensorRT-LLM ship EAGLE drafting. |
| EAGLE-2Faster Inference of Language Models with Dynamic Draft Trees | 2024 |
Background. Tree drafting proposes several candidate continuations at once and verifies the whole tree in one target-model pass. EAGLE and the tree methods before it use a tree whose shape is fixed in advance and identical at every decoding step. Problem. A fixed tree spends its budget uniformly, but acceptance is not uniform. How far a draft survives depends on the context, so some branches are worth extending several tokens and others are dead at the first position. A static shape spends verification on branches that will be rejected. Key idea. Build the tree at run time. The draft head's own confidence scores approximate each candidate's acceptance probability, so the tree expands where confidence is high and prunes where it is low, giving a different shape at every step with no change to the target model or the acceptance rule. Findings. Across three model series and six tasks, EAGLE-2 reports 3.05-4.26× speedups, which are 20-40% higher than EAGLE, while preserving the generated-text distribution. Why it matters. It separates two things speculative decoding usually conflates, how good the drafter is and how the draft budget is spent. The second turns out to be worth a large share of the speedup on its own, and it costs no extra training. |
| EAGLE-3Scaling up Inference Acceleration of Large Language Models via Training-Time Test | 2025 |
Background. EAGLE trains its draft head to predict the target model's features, and that recipe became the default drafter in the major serving engines, with correspondingly more data and compute available to train heads. Problem. The feature-prediction objective caps what more data can buy. A head constrained to reproduce one representation stops improving once it fits it, so additional training data no longer lengthens accepted drafts; the drafter saturates while the training budget keeps growing. Key idea. Drop the feature-prediction constraint. The head predicts tokens directly and consumes fused features from several layers of the target model rather than one, and it is trained under the multi-step conditions it meets at inference, the training-time test the title names. Draft quality then keeps improving as training data grows. Findings. EAGLE-3 reaches a speedup of up to 6.5× and improves over EAGLE-2 by about 1.4×. In SGLang at batch size 64, it raises throughput by 1.38× over the evaluated baseline. Why it matters. It makes the drafter something you scale rather than something you tune. Acceptance length becomes a function of training compute, which is a different engineering posture from choosing tree shapes and draft lengths by hand. Adoption. vLLM, SGLang, and TensorRT-LLM ship EAGLE-3 drafting. |
| MedusaSimple LLM Inference Acceleration Framework with Multiple Decoding Heads | 2024 |
Background. Classic speculative decoding puts a second, smaller model in front of the target. That means training a draft model, keeping its tokenizer and its distribution aligned with the target, and finding memory and scheduling room for it in the serving stack. Problem. The second model is an operational cost that contributes nothing to the speedup itself. It has to be chosen per target model, loaded beside it, and batched with it, and for many targets no suitable small sibling exists at all. Key idea. Attach extra decoding heads to the target model, each trained to predict a token one position further ahead than the last. Their combinations form a small candidate tree that the target verifies in a single pass with tree attention, so the draft comes out of the model you already run and no second model is served. Findings. Medusa-1 accelerates the evaluated models by more than 2.2× without reducing generation quality. Jointly tuning the heads and backbone in Medusa-2 raises the reported speedup to 2.3-3.6×. Why it matters. It shows the drafter need not be a model. Capacity to predict future positions can live in a few heads on top of the existing network, which reduces the deployment cost of speculation from a second model to one extra set of small weights. Adoption. TensorRT-LLM and Hugging Face TGI ship Medusa heads, and vLLM carries a Medusa proposer. |
| self-speculativeDraft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding | 2023 |
Background. Speculative decoding is lossless because the full model verifies every token; the drafter itself can be anything cheap. Until this paper the cheap thing was a separate small model, which you have to obtain, align, and hold in memory. Problem. A suitable small sibling often does not exist, and when it does it is extra weights resident beside the target for the life of the deployment. What you want is a cheap approximation of the model you already loaded, with nothing new trained and nothing new stored. Key idea. Draft by running the target model with a subset of its intermediate layers skipped. The subset is selected once, offline, and the surviving layers act as a fast approximate model that proposes tokens; the full model then verifies them under the standard acceptance rule, so the output distribution is unchanged. Findings. Across Llama-2 and its variants, self-speculative decoding reports speedups of up to 1.99× with identical outputs, no extra training, and no additional model memory. Why it matters. It isolates the assumption underneath all of speculation, that a degraded version of the model agrees with it most of the time, and gets that degraded version free from the same weights. It also ties speculative decoding to early exit and depth pruning, usually treated as separate work. |
| RESTRetrieval-Based Speculative Decoding | 2023 |
Background. Every drafter to this point computes its guesses with a neural network, whether an independent small model or a head bolted onto the target, and so pays in parameters, in training, or in both. Problem. Much of what a model emits is text it has effectively seen before: boilerplate, code idioms, and spans quoted straight out of the prompt. Spending a forward pass to predict those tokens is waste, and any learned drafter has to be trained again for each new target model. Key idea. Retrieve the draft instead of generating it. Match the current suffix against a corpus, take the continuations that follow the matches, assemble them into a candidate tree, and let the target model verify in the usual way. There are no draft parameters and no training, so one datastore serves any model. Findings. On 7B- and 13B-parameter models at batch size one, REST reports 1.62-2.36× speedups across code and text generation. Why it matters. It makes the drafter a data structure. The speedup then depends on how much of your workload is retrievable rather than on the quality of a second model, which is a useful thing to be able to reason about separately when traffic is repetitive. |
| lookahead decodingBreak the Sequential Dependency of LLM Inference Using Lookahead Decoding | 2024 |
Background. Autoregressive decoding is a sequential fixed-point problem, each token conditioned on the ones already emitted. Every speculative method breaks that sequence by adding a proposer: a draft model, a set of heads, or a datastore. Problem. Each of those proposers is something extra to build and keep. A draft model needs training and serving, heads need training against a specific target, and retrieval needs a corpus that resembles the workload. None of them is available at the moment you have only the model. Key idea. Guess several future tokens at once and refine them in place with Jacobi iteration, so a step both commits verified tokens and improves the outstanding guesses. The n-grams verified along the way go into a pool the next step draws candidates from. There is no draft model and no training at all. Findings. Lookahead decoding accelerates the evaluated MT-Bench workloads by up to 1.8× on one accelerator and reaches 4× speedup through strong scaling on multiple GPUs for code completion. Why it matters. It is the pure form of the question this topic asks: how much parallelism is available in decoding when you add nothing to the model. Whatever it recovers bounds what the extra machinery in every other method is actually buying you. |
| Verification trees, long context, and scale | ||
| SequoiaScalable, Robust, and Hardware-aware Speculative Decoding | 2024 |
Background. Tree speculation hands the target model a set of candidate continuations and verifies them in one pass. The tree's shape, how wide it is, how deep it goes, and where it branches, is normally picked by hand or by a fixed heuristic. Problem. The best shape depends on three things at once: the acceptance rate at each depth, the sampling temperature, and how much parallel verification the hardware gives you for free. A tree tuned for one setting is wrong in the others, and one tuned for a given GPU is wrong on the next. Key idea. Treat the shape as an optimization problem. Given measured acceptance behavior, solve for the topology that maximizes expected accepted tokens per step, draw candidates so the result holds across temperatures, and size the tree from the target hardware's own compute and memory profile instead of a constant. Findings. On an A100, Sequoia accelerates Llama-2-7B by up to 4.04×, Llama-2-13B by 3.73×, and Vicuna-33B by 2.27×. In the evaluated L40 offloading setup, exact Llama-2-70B inference reaches 0.56 seconds per token, 9.96× faster than the paper's optimized baseline. Why it matters. It converts a tuning knob into something you compute, and it names the hardware term explicitly: the right amount of speculation is a property of the machine as much as of the model pair. That is the framing you need to move a speculative configuration between deployments. |
| SpecInferAccelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification | 2023 |
Background. The first speculative decoding papers described a decoding algorithm and measured it on single-sequence generation. A serving system has to run many requests concurrently, under a scheduler, with batching and a memory budget. Problem. A single draft sequence is discarded from its first disagreement onward, so one linear guess wastes most of the work it did. Verifying several alternative continuations at once would recover it, but the attention kernels and the scheduler both assume one linear sequence per request. Key idea. Let several small models propose a tree of candidate tokens, then verify the entire tree against the large model in one pass with a token-tree verifier that shares the prefix's attention work across branches. The tree, not the sequence, becomes the unit the serving system schedules and verifies. Findings. SpecInfer outperforms the evaluated serving baselines by 1.5-2.8× for distributed inference and 2.6-3.5× for offloaded inference while preserving generation quality. Why it matters. This is where speculation stops being a decoding trick and becomes a serving design. Tree verification and the branch-aware attention it requires are what later engines inherited, and the argument is made in a serving system's terms of throughput rather than one request's latency. |
| SpecExecMassively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices | 2024 |
Background. Speculative decoding is normally analyzed with both models resident in GPU memory, where a target step costs one pass over weights already on the device. Consumer hardware cannot hold a large target model, so its weights are offloaded and streamed in per step. Problem. When weights stream from host memory, the cost of a target step is set by the transfer rather than by how many tokens that step verifies. A draft tree sized for a resident model leaves nearly all of that transfer unused. Key idea. Build a very large draft tree with the small resident draft model, then verify the whole tree in one offloaded target pass, so a single expensive step commits many tokens. The tree is sized to the step cost instead of to a fixed draft length. Findings. SpecExec generates as many as 20 tokens per target-model iteration. With RAM offloading on consumer GPUs, 50B-plus-parameter models reach 4-6 tokens per second at 4-bit precision and 2-3 tokens per second with 16-bit weights. Why it matters. Tree size is not a constant of the method. It follows from the ratio of target step cost to draft cost, and that ratio depends on where the weights live, so the datacenter setting and the consumer setting want opposite answers. |
| MagicDecBreaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding | 2024 |
Background. The standard case for speculative decoding is that decode at small batch is memory-bound, so verifying several tokens costs barely more than generating one. The corollary usually drawn is that speculation buys latency at low batch and costs throughput at high batch. Problem. At long context and large batch the usual assumptions invert. KV cache traffic rather than weight loading dominates the target step, and it grows with both batch size and sequence length, so the reasoning that decides when speculation pays no longer applies. Key idea. Work out which term actually dominates at a given batch size and sequence length, then hold the draft to a fixed-size sliding-window KV cache so drafting stays cheap as context grows. Speculation then improves latency and throughput together in the long-context regime. Findings. For moderate-to-long sequences and batch sizes from 32 to 256, MagicDec accelerates Llama-3.1-8B by up to 2.51× across the evaluated hardware and tasks. Why it matters. A tradeoff presented as fundamental can be an artifact of an operating point. This is the entry that forces you to ask which term dominates before reusing a rule of thumb about memory-bound decode. |
| mirror speculative decodingBreaking the Serial Barrier in LLM Inference | 2025 |
Background. Draft-then-verify runs in strict alternation. The draft model produces candidate tokens, the target model verifies them, and only then does the draft resume from whatever was accepted. Problem. Alternation means one of the two models is idle at every instant, and the latency of a round is the sum of both phases. On a deployment with more than one kind of accelerator, whichever device is not in its phase does nothing. Key idea. Run both directions at once. The draft proposes continuations while the target proposes correction paths for the draft, with the two halves placed on heterogeneous accelerators. The tradeoff is a deployment that needs both a GPU and an NPU plus a tight cross-device synchronization budget. Findings. On SpecBench with 14B- to 66B-parameter models, Mirror-SD reports 2.8-5.8× wall-clock speedups and a 30% average relative improvement over EAGLE-3. Why it matters. It reframes speculation as a placement and pipelining problem across unlike devices rather than an algorithm running on one. Given heterogeneous silicon, the serial loop is leaving a whole device idle, and that idleness is the thing to attack. |
| lossy verificationRevisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes | 2026 |
Background. Everything above this line on the page is lossless. The acceptance rule is constructed so that the tokens surviving verification are distributed exactly as the target model's own sampling would be. Problem. Relaxing that rule buys acceptance rate and quietly rewrites the output distribution. The speedup is easy to measure and the distributional cost is not, so a lossy variant can be reported as a strict improvement. Key idea. Classify lossy verification schemes into two families, truncation-based and collaborative, then analyze the distributions each family induces under a shared diagnostic framework. Findings. Truncation-based methods can perform substantially worse than the true truncation-sampling baseline because they distort its distribution. Collaborative methods avoid low-quality outputs only when the overshoot of draft probabilities over target probabilities remains controlled. Why it matters. The lesson generalizes past speculative decoding to any approximation sold as a speedup with an unstated distributional cost. It gives you two questions to ask: which family the scheme belongs to, and what its failure looks like. |
| Inside a serving system | ||
| online spec decodingOnline Speculative Decoding | 2023 |
Background. A draft model is trained once, offline, on a general corpus, and then shipped with the serving stack. The speedup it delivers is governed by its acceptance rate against the target model. Problem. Acceptance rate depends on how closely the draft agrees with the target on the text actually being generated, and a deployed model sees a narrow query distribution that shifts over time. A general draft is mismatched where the traffic is, and the mismatch costs speed silently, as rejected tokens rather than wrong output. Key idea. Keep updating the draft model during serving on the live query distribution, taking the target's outputs on the queries that arrive as the training signal. The draft becomes a component that tracks traffic instead of a fixed artifact shipped beside it. Findings. Across synthetic and real query streams, online adaptation raises token acceptance by 0.10-0.65 and reduces latency by 1.42-2.17× relative to the paper's offline-draft baseline. Why it matters. Acceptance rate turns into something you maintain rather than a number measured once at deployment. It also puts a training loop inside a serving system, which raises the question of what capacity pays for the updates and when. |
| AdaServeAccelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding | 2025 |
Background. A serving engine applies one speculation setting to the whole running batch, on or off, with a single draft length or tree shape. The requests sharing that batch arrive with different latency targets. Problem. One setting cannot satisfy several SLOs at once. A tree sized for the tightest target spends compute the looser requests would rather have as throughput, and a tree sized for throughput misses the tight target. Key idea. Make per-request speculation a scheduling variable. A constrained-optimization formulation builds a speculation tree sized to each request’s latency target, and a speculate-select-verify pipeline trades decode speed against throughput within one shared batch. Findings. Across the evaluated workloads, AdaServe reduces SLO violations by up to 4.3× and improves goodput by up to 1.9× over the strongest baseline. Why it matters. This is the serving-system view of speculative decoding once SLOs stop being uniform. Speculation stops being a deployment flag and becomes a per-request decision, made by the scheduler with the same information it uses to batch. |
| DSparkConfidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation | 2026 |
Background. Parallel drafters propose a long block of tokens in one forward pass, which is what makes drafting cheap enough to matter at scale. AdaServe above makes the verification budget a per-request scheduling variable. Problem. Parallel drafts omit dependencies within the block, so acceptance decays along its suffix. Verifying the full block then spends scarce batch capacity on likely rejections. Key idea. Add a lightweight sequential module so the draft captures intra-block dependencies. Then choose each request's verification length from its estimated prefix survival probability and the engine's measured throughput under current load. Findings. Accepted length improves over autoregressive and parallel drafters. Under live DeepSeek-V4 traffic, per-user generation speed rises 60-85% over the MTP-1 baseline at matched throughput. Why it matters. Draft length reflects model quality; verification length reflects current serving capacity. Treating them separately turns speculation into a per-request scheduling decision. |
| SpecForgeA Flexible and Efficient Open-Source Training Framework for Speculative Decoding | 2026 |
Background. EAGLE-3 drafting is in every major engine, and the speedup depends on a draft head trained against the specific target model. Trained heads exist for a handful of popular open checkpoints. Problem. The algorithm is settled and the drafts are not. A newly released base model has no head, and producing one at production quality means training a draft against a large target, which is a distributed-training problem rather than an inference one. Key idea. Build a draft-training framework with target-draft decoupling, hybrid parallelism, and tuned kernels. SpecBundle adds released EAGLE-3 drafts for mainstream open models. Findings. For Qwen3-235B-A22B, SpecForge trains EAGLE-3 drafts up to 9.9× faster than the evaluated training baseline. The released SpecBundle drafts deliver up to 4.48× end-to-end inference speedup in SGLang. Why it matters. The practical blocker it addresses is that good draft models, not the algorithm, are what the community is missing. Whether an inference optimization is available to you often turns on training infrastructure and released artifacts rather than on the published method. |
| AgentSpecSpeculative Decoding for Batch Inference of LLM Agents | 2026 |
Background. Speculation pays when a decode step is memory-bound and has arithmetic to spare for checking extra tokens. That spare capacity shrinks as the batch grows, and agent serving runs at large batch, with many trajectories in flight at once. Problem. At those batch sizes speculative decoding loses most of its speedup. The two causes identified here are a high rejection rate for drafted tokens and a dynamic token budget that goes unused. Key idea. Constrain drafting to semantically coherent segments of the agent workflow, which cuts speculation down irrelevant paths, and allocate the free token budget using agent-level information. Why it matters. It recovers a lost optimization from structure the serving layer normally discards, namely what the agent is doing next. Read it with the agent-serving meetings, where the batch sizes come from. |
This meeting asks what changes when the serving unit is an agent program rather than an isolated request: first measure and replay the workload, then expose workflow structure, schedule whole programs, and optimize across their execution graphs. The nov-30 meeting follows the persistent runtime state those decisions create through tool stalls, sandboxes, and sessions.
Prefix-cache reuse is the other half of this meeting: recognize shared token spans, place their KV state, and choose what to retain. Scheduling and caching together prepare the serving work in Assignment 4.
| Paper | Year | Why read it |
|---|---|---|
| Workloads, traces, and evaluation | ||
| agentic AI workloadsAgentic AI Workload Characteristics | 2026 |
Background. Systems work on agents has been designed against an assumption about the workload: long prompts, growing context, prefill-heavy execution. That assumption came from reasoning about how agents work, not from traces of them running. Problem. Designing a scheduler or a cache policy for the wrong workload shape wastes the effort. Whether agent serving is prefill-bound or decode-bound, and whether state is per-request or per-session, changes which mechanism pays off, and neither question can be settled by inspection. Key idea. Instrument ReAct-style agents end to end across five benchmarks, collecting traces from both the model-serving and tool-execution sides rather than measuring isolated model calls. Findings. With effective context caching, most input tokens are reused across turns, making execution decode-dominated and dependent on long-lived KV-cache state. Tool use also follows a temporal pattern, shifting from read and explore operations early in a trajectory to execute and write operations later. Why it matters. Each finding retargets a design decision. Decode-dominated execution moves the bottleneck to memory bandwidth, long-lived session state makes cache retention a first-class scheduling concern, and the phase structure of tool use is something a scheduler can anticipate. |
| agentic workload characterizationFrom LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems | 2026 |
Background. Serving research measures one model call at a time: prefill a prompt, decode a response, report time to first token and throughput. Agentic applications interleave those calls with tools, sandboxes, and retrieval. Problem. The assumptions that hold for single-call inference do not survive that interleaving, and until an agent is instrumented end to end you cannot say which of them break, or by how much. Key idea. AgentSysBench combines ten representative agentic applications with unified end-to-end instrumentation and production traces, then follows each measured bottleneck to a serving-system intervention. Findings. Non-LLM components dominate latency in 5 of 10 applications, sandbox working sets peak at 28 GB per session, and component latencies diverge by as much as 32×. Four design explorations reduce task latency by 29-40%, improve communication-aware placement by as much as 4.5×, reduce state memory by 4.6×, and eliminate 35.2% of redundant search calls. Why it matters. Each property invalidates an assumption you would otherwise carry over from request serving, so it tells you which parts of the stack are worth redesigning. Read it for the four design explorations at the end, which are the agenda this meeting works through. |
| TraceLabCharacterizing Coding Agent Workloads for LLM Serving | 2026 |
Background. Serving research has mostly measured chat traffic: short prompts, long outputs, and requests independent of one another. The cache and scheduling policies in production stacks were tuned against that shape. Problem. A coding agent emits none of it. It runs long autonomous loops, sends long contexts for short outputs, and issues heavy-tailed tool calls, so a policy tuned on chat traffic is tuned on the wrong distribution. Key idea. Release a trace of roughly 4,300 Claude Code and Codex sessions containing about 350,000 model steps and 430,000 tool calls, together with the collection pipeline and analysis code. Findings. Coding-agent sessions have long autonomous loops, long contexts paired with short outputs, diverse heavy-tailed tool calls, and prefix-cache hit rates that are high but imperfect. Why it matters. It is an artifact rather than an argument. You can drive your own simulation and cache-policy experiments against real agent sessions, and the imperfect part of that hit rate is exactly where a cache policy still has room to work. |
| CacheWiseUnderstanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding Agents | 2026 |
Background. A serving engine caches KV blocks by prefix and evicts them under memory pressure. vLLM's automatic prefix caching evicts in least-recently-used order, which assumes a request's blocks stop being useful once the request goes quiet. Problem. A coding-agent session goes quiet on every tool call and returns needing the same prefix, so recency-ordered eviction discards blocks that are about to be reused. Long-running sessions hold sustained KV pressure that conventional policies mishandle. Key idea. Measure coding-agent KV reuse first, then design around the measurement: prefix-aware scheduling plus reuse-aware eviction, driven by lightweight predictions from tool-call metadata, implemented in vLLM. Findings. On the collected coding-agent traces, CacheWise reduces KV-cache evictions by 2-2.6× and improves total agent-session completion time by as much as 3.5× over conventional cache management. Why it matters. It is the concrete case that eviction order should follow predicted reuse rather than recency, and it shows the prediction can come from metadata already attached to the request instead of from the model. |
| architectural implicationsArchitectural Implications of Agentic AI Workflows | 2026 |
Background. A GPU server is provisioned as a uniform unit, with a fixed ratio of CPU to accelerator sized for a workload that keeps the accelerator busy. Inference deployments have been analyzed on that assumption. Problem. An agentic request fragments into LLM calls, tools, and orchestration that repeatedly cross the CPU-GPU boundary. That puts the CPU on the critical path and strands both CPU and GPU capacity under bursty load. Key idea. Characterize agentic workflows from Microsoft Azure production data and open-source frameworks, attributing time and stranded capacity to workflow fragments rather than to the model call alone. Findings. The study finds that host-side orchestration and tools put the CPU on the critical path, bursty execution strands both CPU and GPU capacity, homogeneous CPU provisioning mismatches heterogeneous software roles, and core sharing degrades locality. Agora improves utilization and server throughput while preserving agent tail latency. Why it matters. It gives you the vocabulary to separate inefficiency a scheduler can fix from inefficiency built into a uniform server. The second kind is a provisioning question, and no amount of batch tuning answers it. |
| XPerfBenchmarking LLM Serving Systems for Agentic AI Workloads with XPerf | 2026 |
Background. Agent applications are nondeterministic programs whose later model calls depend on earlier generations and tool results. A conventional serving benchmark can replay a fixed list of prompts, but that list no longer describes the application that produced it. Problem. Comparing serving systems requires a repeatable workload without erasing the dependencies, branching, and timing that make agent traffic different from chat. Re-running the live application changes the trace, while replaying requests independently changes the workload. Key idea. XPerf collects, synthesizes, and replays traces of nondeterministic agent applications while preserving the structure needed to exercise a serving stack. The release includes eight applications and supports controlled experiments from recorded executions. Findings. The evaluation shows that its replay is accurate enough to support comparative profiling, and demonstrates the framework as a tool for debugging serving behavior and studying scaling under agentic workloads. Why it matters. It connects characterization to experimentation. The traces in this group tell you what production-like agent traffic looks like; XPerf supplies a way to reproduce that traffic closely enough to compare systems and explain where their performance diverges. |
| when does disaggregation payWhen Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference | 2026 |
Background. Disaggregated serving puts prefill and decode on separate machines so each phase gets hardware matched to its bottleneck, prefill compute and decode memory bandwidth. A further split, attention apart from FFN, has been proposed on the same reasoning. Problem. Whether a split pays depends on the traffic, and finding out normally means building each configuration and measuring it. Agentic requests are not chat requests: context accumulates across turns, so the ratio of prefill work to decode work moves. Key idea. Simulate prefill-decode and four-way prefill-decode-attention-FFN specialization across quantization, parallelization, and heterogeneous accelerator choices, then compare them with unified execution. Findings. The simulations project up to 75% higher throughput from prefill-decode disaggregation than unified serving on current GPUs. With custom NPUs, the four-way split is the most consistently beneficial configuration across the evaluated models. Why it matters. The result makes hardware specialization conditional on both workload and device assumptions. It shows why an architecture cannot be evaluated apart from its traffic, but also why a projection for custom NPUs should not be read as measured production performance. |
| Exposing workflows to the serving layer | ||
| compound AI systemsThe Shift from Models to Compound AI Systems | 2024 |
Background. By 2024 the strongest results on many tasks came not from one model call but from an assembly: a retriever, several model calls, tool invocations, and control logic holding them together. The practice had run ahead of any vocabulary for it. Problem. Treating the model as the system leaves the interesting engineering invisible. If the only unit you measure and tune is a single forward pass, you cannot reason about a pipeline whose quality, latency, and cost are set by how the calls are composed. Key idea. Name the unit of design. A compound AI system has models as components, and includes retrieval, tools, and control flow; the argument is that recent gains come from that composition, and that design, optimization, and operation therefore belong at the system level rather than the model level. Why it matters. It supplies the system boundary that Parrot and SGLang then make explicit and optimizable. Once the unit is the program rather than the call, a scheduler can be shown the dependencies between calls, and DSPy, LangChain, and LangGraph are all frameworks built at that boundary. |
| ParrotEfficient Serving of LLM-based Applications with Semantic Variable | 2024 |
Background. An LLM application is rarely one call. A chat agent, a map-reduce summarizer, and a multi-agent pipeline all issue many calls whose inputs depend on earlier outputs, and each call reaches the serving system as an independent request. Problem. The public completion API hides that dependency structure. A scheduler seeing only isolated requests can optimize only per-request latency, which is the wrong objective: it will finish an intermediate call quickly while the application waits on a chain the scheduler cannot see. Key idea. Semantic Variables annotate a request's inputs and outputs, so the runtime can reconstruct the dependency graph across the calls of one application and schedule against it, optimizing end-to-end latency instead of per-request latency. Findings. Across the evaluated application patterns, Parrot improves end-to-end performance by as much as an order of magnitude over request-level serving baselines. Why it matters. It names the interface question the whole agent-serving block turns on: how much structure must an application declare before a serving system can do anything useful with it? Once the graph is visible, prefix sharing, batching, and priority all have a target other than the individual request. |
| PythiaExploiting Workflow Predictability for Efficient Agent-Native LLM Serving | 2026 |
Background. Serving engines built for chat treat each request as independent and try to recover reuse after the fact, through a prefix cache that hopes the next arrival shares a prefix with something still resident. Problem. Treating structured agent traffic as generic requests leaves the engine unable to anticipate reuse, long-context contention, or changes in demand across workflow stages. Key idea. Agent workflows are predictable, so let the application say what it is about to do. It exposes workflow semantics through a small interface at the serving layer, and the scheduler exploits that declaration instead of inferring structure from the request stream. Findings. Production traces from an agent-serving platform and an internal coding assistant show low prefix-cache hit rates, severe contention from long-context requests, and substantial queuing from suboptimal scaling. Pythia improves throughput and job completion time over the evaluated serving baselines. Why it matters. Read it against Parrot for how thin a declaration interface can get and still pay for itself. It also supplies the measurement the rest of this block assumes: the cross-call reuse a chat-era prefix cache counts on is not there in agent traces. |
| workflow-aware servingA Workflow-Aware Serving Layer for Agentic Applications | 2026 |
Background. An agentic application runs between two layers it does not control: a framework above that expresses the workflow, and one or more serving engines below that execute individual calls. Which model and which backend handles each step is normally left to static configuration. Problem. Those choices interact across the workflow and change while it runs. The right model for one node depends on what later nodes will need, the right backend depends on load that shifts between steps, and re-solving the assignment on every call puts an optimizer in the request path. Key idea. Dyserve inserts a layer between frameworks and engines and solves one integer linear program per workflow, picking each DAG node's model, verifier, and backend over a heterogeneous pool. That program is pre-solved at several pressure levels at admission, so a load shift can redirect the workflow's uncommitted suffix without re-running the solver. Findings. Across LiveCodeBench, GAIA, ComplexFuncBench, and SWE-bench, Dyserve scores 3-10 accuracy points above the strongest compared baseline while reducing latency by 1.1-6.8×. Its admission compilation takes less than 60 ms per request at p95. Why it matters. Read it for how much structure the application must expose to make that solve possible. It also shows what pre-computation buys: the expensive decision happens once at admission, and the runtime only chooses among answers already computed. |
| software-defined agentic servingSoftware-Defined Agentic Serving | 2026 |
Background. Software-defined networking separated a network's control plane from its data plane, so routing policy became a program running against a global view instead of a configuration file on each switch. Agentic serving is arranged the other way, through static parameters fixed per deployment. Problem. A fixed configuration cannot answer questions whose answers change continuously: where to send this call, which instance already holds the state it needs, how much of the interconnect its transfers deserve. The system has the runtime state to decide and no place to put the policy. Key idea. The proposal is that agentic serving be programmable and system-aware rather than statically parameterized. This short position paper sketches an SDN-inspired framework whose communication attributes are adjusted from runtime state instead of set at deployment time. Why it matters. It is worth reading for whether the SDN analogy holds when the “packets” carry gigabytes of KV state. Read it as an argument about where serving policy belongs rather than as a system to evaluate. |
| PiePie: A Programmable Serving System for Emerging LLM Applications | 2025 |
Background. Emerging LLM applications need behaviors such as custom control flow, state handling, and coordination around inference. A fixed completion interface keeps those behaviors outside the serving engine, where they cannot participate in its scheduling and batching decisions. Problem. Adding every new application pattern directly to a serving system does not scale, but exposing the engine's internals to application code would compromise isolation and portability. The interface needs to be programmable without turning each feature into an engine fork. Key idea. Pie exposes fine-grained event handlers controlled by WebAssembly inferlets. Application-specific serving logic runs through those handlers while the engine retains control over execution and resources. Findings. Pie reports 3-12% latency overhead for ordinary completion requests and 1.3-3.4× higher throughput on the evaluated agentic workflows. Why it matters. It offers a concrete answer to how much programmability the serving layer should expose. Read it against declarative workflow interfaces: Pie lets applications bring behavior into the engine, and WebAssembly is the boundary intended to keep that behavior manageable. |
| NalarNalar: An agent serving framework | 2026 |
Background. Agent code in Python naturally expresses dependencies by assigning the result of one step and passing it to the next. A serving runtime instead needs those dependencies before values exist if it is to schedule the workflow globally. Problem. Requiring developers to rewrite an application as a separate DAG duplicates its control logic, while intercepting isolated model calls reveals dependencies too late. Context and mutable state also have to survive across the calls the runtime reorders. Key idea. Nalar uses Python stubs that return dependency- and context-carrying futures, allowing ordinary-looking application code to reveal work before it completes. The runtime manages state and combines global control with local decisions as the workflow unfolds. Findings. Across the tested request rates, Nalar reports 34-74% reductions in tail latency and up to 2.9× speedup on the evaluated workflows. Why it matters. It is an interface design for partial disclosure: the program exposes enough structure to schedule ahead without becoming a static graph. That makes it a useful comparison point for both declarative workflow systems and lower-level programmable serving hooks. |
| MurakkabMurakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms | 2026 |
Background. A cloud agent workflow combines model calls, tools, and other services whose resource needs differ by stage and shift with load. Application code commonly mixes the workflow's logical structure with deployment and resource configuration. Problem. A configuration chosen before deployment cannot remain efficient as stage costs and queueing change, while a purely reactive runtime lacks the profile needed to choose good placements. Coupling configuration to code also makes retuning the workflow invasive. Key idea. Murakkab separates a declarative workflow specification from its configuration, then combines a profile-guided optimizer with an adaptive runtime. The optimizer establishes an efficient plan and the runtime adjusts execution as conditions change. Findings. Across the evaluated workflows, Murakkab reports reductions of up to 2.8× in GPU use, 3.7× in energy, and 4.3× in cost while maintaining the tested SLOs. Why it matters. It spans the boundary between offline planning and online control. The separation of specification from configuration makes the workflow an object the platform can optimize, while adaptation acknowledges that no profile remains exact once the system is running. |
| Scheduling programs rather than requests | ||
| AgentixAgentix: An Efficient Serving Engine for LLM Agents as General Programs | 2026 |
Background. A serving engine queues individual requests. An agent program is not one request: it is a sequence of LLM calls interleaved with tool calls, and the engine sees only the calls, one at a time, with no record that they belong together. Problem. Scheduling each call on its own arrival lets a newly submitted request cut ahead of the next call of a program that has already run for minutes. The program waits behind work that started later, and head-of-line blocking is paid at program granularity, where it is far more expensive than at request granularity. Key idea. Treat the agent program, not the request, as the scheduling unit. Track how much work each program has already done and order calls by what that program still needs, so the queue reflects program progress rather than request arrival order. Findings. Across the evaluated models and agent workloads, Agentix sustains 4-15× more completed programs than vLLM at the same end-to-end latency. Why it matters. It changes the objective a serving engine is optimizing. Once the unit is a program, per-request latency stops being the right target and program completion time takes over, which is the metric a user of an agent actually experiences. |
| SAGAWorkflow-Atomic Scheduling for AI Agent Inference on GPU Clusters | 2026 |
Background. An agent workflow spans many LLM calls, and on a GPU cluster those calls land on whatever instance the router picks. The cluster scheduler places calls; nothing above it knows the calls form one workflow. Problem. Placing calls independently spreads a single workflow across instances and interleaves it with others, so cached state is scattered and the workflow finishes only when its slowest fragment does. Optimizing each placement locally does not optimize the workflow. Key idea. Workflow-atomic scheduling: admit and place the whole agent workflow as one unit on the cluster, so the scheduler commits resources to a workflow rather than to its individual calls. Findings. On 64 GPUs running SWE-bench and WebArena agents, SAGA reduces task completion time by 1.64× over vLLM with prefix caching and affinity routing, improves GPU-memory utilization by 1.22×, and reaches 99.2% SLO attainment. Its latency gains cost about 30% of peak throughput relative to throughput-optimal batching. Why it matters. It lifts the program-as-scheduling-unit argument from a single engine to a cluster, where the scarce resource is instances and KV capacity rather than a single GPU's batch slots. That is where the gang-scheduling and admission-control tradeoffs become visible. |
| SMetricRethink LLM Scheduling for Serving Agents with Balanced Session-centric Schedulingalso Nov 23 | 2026 |
Background. Serving metrics were defined for humans reading streamed output, so time to first token and inter-token latency drive the scheduler. Agents consume responses programmatically, and they issue many requests per session. Problem. Under that objective, cache-aware routing looks wrong: sending a session's requests to the instance holding its prefix overloads a few instances while others idle. Load balance and KV reuse pull against each other, and per-token latency does not say which should win. Key idea. Take cluster throughput as the objective, then balance only each session's first request and route the follow-ups cache-aware. The tradeoff being managed is reuse against load balance, not latency against fairness. Findings. Relative to state-of-the-art schedulers, SMetric improves cluster throughput by 10-16% for colocated prefill and decode with a global cache, and improves prefill throughput by 2-34% under disaggregation while also lowering per-token latency. Why it matters. It shows that a metric chosen for one consumer misprices a mechanism for another. Once the reader is an agent, deliberate imbalance is the right call, and that is a general lesson about which objective a scheduler is allowed to optimize. |
| observation not predictionObservation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving | 2026 |
Background. Most scheduling proposals for agent serving need to know something about the future: how many tokens a response will take, how many turns remain, whether a session is finished. They obtain it by prediction. Problem. A prediction that is wrong costs more than having no prediction at all, because the scheduler commits placement and memory on it. Agent behavior depends on tool outputs and on the model's own choices, so the quantities being predicted are exactly the ones that are hard to predict. Key idea. Refuse to predict. Schedule on observed conversation state instead, at conversation granularity across disaggregated instances, so decisions rest on what has already happened rather than on an estimate of what comes next. Findings. Against a per-turn prediction baseline, ConServe reduces p95 time to first effective token by 51.08% and improves energy efficiency by 7.51% while preserving last-turn time between tokens and SLO attainment. Heterogeneous GPU tiers improve energy efficiency by a further 22.75%. Why it matters. It is the robustness argument against the prediction-based approaches that dominate this section, and it makes the cost of a mispredict explicit. Reading it alongside the predictive schedulers tells you what each one is actually betting on. |
| KairosLow-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud | 2025 |
Background. Multi-agent applications send every agent's call to one shared model endpoint. The serving stack underneath admits and batches those calls one request at a time, and a public cloud deployment is sized for average load rather than peak. Problem. Under overload the queue decides end-to-end latency, and per-request scheduling is blind to the workflow that issued the request. One queued call gates an entire workflow, because the agents downstream of it cannot start until it returns. Key idea. Schedule with the workflow in view. Kairos orders requests by each agent's latency characteristics and dispatches by memory demand, so calls a workflow is waiting on are not left behind calls that nothing depends on. Findings. In the paper's experiments, Kairos reduces end-to-end latency by 17.8-28.4% against the evaluated multi-agent serving baselines. Why it matters. It argues the scheduling unit should be the workflow, not the request. Once one model serves many agents under load, queue discipline becomes the dominant latency term, and the information needed to choose a good one lives above the serving layer. |
| MaestroWorkload-Aware Cross-Cluster Scheduling for LLM-Based Multi-Agent Systems | 2026 |
Background. A multi-agent workflow is a sequence of stages, and those stages call different models spread over more than one cluster. The scheduler sees each call at admission, with no knowledge of what the call is about to do. Problem. Without a stage's output length or memory footprint you cannot place it well. Two costs follow: loading a model's weights on demand adds a cold start, and a long stage in front of a short one blocks the interactive tasks behind it. Key idea. Predict each stage's output length and memory footprint, then spend the predictions at three levels: co-locate several models on a node by caching weights, route across clusters to dodge cold starts, and prioritize globally by workflow to limit head-of-line blocking. Findings. Across prototype experiments and trace-driven simulations, Maestro reduces reserved KV-cache HBM by 67.2% and improves high-contention SLO attainment over earliest-deadline-first scheduling by 23.6 percentage points. Why it matters. The useful scheduling input turns out to be a prediction about the stage rather than a measurement of the request. Read it for what workload-awareness costs as well, because all three levels inherit whatever the predictor gets wrong. |
| Optimizing across the workflow | ||
| agentic batch query optimizationBatch Query Processing and Optimization for Agentic Workflows | 2025 |
Background. A database does not execute queries one at a time as written. It compiles each into a plan and then optimizes across concurrent plans so shared work runs once. An agent workflow is also a graph of operations, and agent frameworks execute it one call at a time. Problem. Concurrent workflows overlap heavily, sharing prompts, retrieved context, and whole subgraphs, and a stack that sees only independent calls repeats that shared computation once per workflow. Key idea. Halo compiles each agent workflow into a query-plan DAG and consolidates concurrent workflows into one graph, so shared computation runs once. A cost model over prefill and decode cost, cache reuse, and GPU placement drives optimization at the level of the plan rather than the individual call. Findings. Across six benchmarks, Halo speeds up batch inference by as much as 3.6× and improves online-serving throughput by as much as 2.6× while preserving output quality, including workloads with thousands of queries and complex graphs. Why it matters. It is the cleanest statement of the database view of agent serving. If a workflow is a query, then serving many of them is multi-query optimization, and the vocabulary of plans, cost models, and shared subexpressions transfers to a problem the serving literature has been solving without it. |
| Ayo (formerly Teola)Towards End-to-End Optimization of LLM-based Applications with Ayo | 2025 |
Background. An LLM application is a pipeline: embed the query, search an index, rerank, prompt the model, parse the output, call a tool, prompt again. Each stage is tuned by whoever owns it, and the serving engine sees only the model calls. Problem. Optimizing the model call alone leaves end-to-end latency roughly where it was. The engine cannot see what runs before or after it and the orchestration layer cannot see inside the engine, so stages that could overlap run in sequence, and work that could be shared across requests is not. Key idea. Express the whole application as a dataflow graph of primitive operations rather than opaque service calls, then optimize across that graph: run independent branches in parallel, batch equivalent primitives from different requests together, and let a downstream stage start on partial output from the stage above it. Findings. Across several representative LLM applications, Ayo achieves as much as a 2.09× end-to-end speedup over existing orchestration systems. Why it matters. It defines the object the rest of this group operates on. Once the application is a graph, serving becomes query planning and the question shifts from how fast one call runs to what the planner is allowed to move. It also relocates the target: what a user waits for is the application rather than the model call, and the largest remaining latency is often in the parts that are not the model at all. |
| HeliumEfficient LLM Serving for Agentic Workflows: A Data Systems Perspective | 2026 |
Background. Agentic workflows arrive at a serving engine as a stream of independent calls. Data systems have long handled a related shape, a declarative program compiled and planned before execution. Problem. A serving engine plans one call at a time, so it cannot reuse intermediate results across a workflow, cannot reorder stages, and cannot choose an execution strategy from the workflow's structure. The information a planner would need never reaches it. Key idea. Read agentic serving as a data systems problem: describe the workflow, then plan and execute it the way a query engine would, with the workflow rather than the call as the unit of optimization. Findings. Across the evaluated workloads, Helium's proactive caching and cache-aware scheduling deliver as much as a 1.56× speedup over the compared agent-serving systems. Why it matters. The framing imports decades of planning, caching, and cost-model machinery into serving, and it makes explicit what today's engines give up by scheduling greedily on arrival. |
| HexAGenTEfficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling | 2026 |
Background. Production clusters now split prefill and decode onto separate, often unlike, hardware, and transfer KV cache between them. An agent workflow is a DAG whose next node is revealed only after the current call returns. Problem. A scheduler that sees one call at a time cannot tell which ready call is on the critical path of its workflow, and on heterogeneous prefill-decode disaggregated hardware it must also choose where each half runs under KV capacity limits and cross-stage transfer cost. Local choices miss the workflow's deadline. Key idea. Schedule the online-revealed DAG directly: rank ready calls by projected risk of missing the workflow's completion horizon, then jointly pick prefill placement, decode placement, and queue priority. The objective is workflow-level latency, not per-call latency. Findings. Across agent workloads on heterogeneous A100, H100, and H200 clusters, HexAGenT lowers the SLO scale needed for timely completion by 20.1% on average at 95% attainment and by 33.0% at 99% attainment. Why it matters. It joins two decisions the field usually treats separately, what to run next and where to run it, and shows that heterogeneity and workflow structure have to be resolved together rather than in sequence. |
| KVFlowEfficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows | 2025 |
Background. Multi-agent workflows repeatedly invoke specialized agents whose fixed role prompts create reusable prefixes. Prefix caching can preserve the corresponding KV tensors instead of recomputing them on every invocation. Problem. Least-recently-used eviction sees only past accesses. It may discard an agent's cached prefix immediately before the workflow invokes that agent again, causing recomputation or a host-to-device reload even when the workflow exposes the future execution order. Key idea. Represent the workflow as an Agent Step Graph and assign each agent a steps-to-execution value that estimates when it will run again. Use that value for fine-grained KV-node eviction, and prefetch the next step's tensors from CPU memory while the current step executes. Findings. Relative to SGLang with its hierarchical radix cache, KVFlow speeds up a single workflow with large prompts by as much as 1.83× and workloads with many concurrent workflows by as much as 2.19×. Why it matters. KVFlow shows why agent structure belongs in the cache manager as well as the scheduler. When the workflow reveals near-future reuse, recency is an avoidable approximation, and prefetching can turn the same structural signal into overlap rather than a generation stall. |
| Inference scaling and agentic RL | ||
| rollout infrastructure taxThe Rollout Infrastructure Tax in Coding-Agent Reinforcement Learning | 2026 |
Background. Reinforcement learning on coding agents needs rollouts, and every rollout runs the agent against a real repository inside an isolated execution environment. That environment is drawn from four substrates in practice: single containers, hosted sandboxes, Kubernetes, and cloud VMs. Problem. The substrate is usually picked for convenience, on the assumption that its cost is small next to the GPUs. Nobody had measured what the four actually cost on the same rollout workload. Key idea. Compare single containers, hosted sandboxes, Kubernetes-orchestrated containers, and cloud virtual machines on the same coding-agent rollout workload, then project each substrate's overhead at large training scale. Findings. Cold-start latency differs by as much as 110× across the four substrates, and projected worker-hours differ by 1.8× for one million 150-step trajectories. The projection covers rollout workers only and excludes GPU inference and optimizer steps. Why it matters. It puts a number on infrastructure that the agentic-RL papers next to it treat as free. A 1.8× spread in worker-hours makes the substrate a first-order term in an RL training budget rather than a deployment detail. |
| TideRLBoosting Agentic RL Goodput with Readiness-Aware Scheduling | 2026 |
Background. An agentic RL step generates a batch of rollouts and then trains on them. Each rollout drives an agent through tool calls and environment steps, so the batch is finished only when its slowest trajectory is. Problem. Trajectory length is data dependent and long tailed, so trainer utilization is set by the tail rather than the mean. Waiting for the straggler idles the trainer; the trainer has work available, just not the work its batch is waiting on. Key idea. Schedule rollouts by readiness, so the trainer consumes trajectories as they become available instead of waiting for a batch boundary. On-policy staleness is the constraint that bounds it: reordering may only go as far as the training data can drift from the policy that generated it. Findings. Across text-only and multimodal agent workloads, TideRL improves training goodput by as much as 5.6× over synchronous baselines and by more than 33% over asynchronous baselines at similar task performance. It also cuts per-step training time by as much as 44.3% and waiting time by as much as 77.6%. Why it matters. This is the stall-hiding logic of agent serving applied to training, and the tradeoff is explicit. The more reordering you allow, the busier the trainer stays and the staler the data it trains on. |
| LibraEfficient Resource Management for Agentic RL Post-Training | 2026 |
Background. Agentic RL post-training alternates two unlike jobs on one cluster: generation, which is inference over many concurrent trajectories, and training, which is a synchronous gradient step. Frameworks divide the GPUs between them. Problem. That division is fixed when the job launches, and the demand behind it is not. Agent trajectories are long tailed, so how long generation holds its share varies from step to step, and a static partition leaves one side of it idle. Key idea. Reallocate an elastic GPU pool between rollout and training while a causality-driven scheduler assigns trajectories to heterogeneous rollout buckets from tool-return signals rather than predicted sequence lengths. Findings. On 48 A800 GPUs, Libra delivers as much as 3.0× higher throughput than the evaluated resource-management baselines and reaches a given reward as much as 2.5× sooner. Why it matters. It names the resource-management question hiding inside every RL post-training stack, and it locates the cost of the static answer in the shape of the workload rather than in model size. |
| MISA-TScheduling Mixed RL Rollouts Beyond Prefix Locality | 2026 |
Background. RL trainers generate a prescribed mixture of rollout workloads whose sessions differ in model behavior, context growth, and tool gaps. Serving them efficiently requires reuse of KV state without changing the data mixture the trainer expects. Problem. Prefix-local scheduling alone groups similar prompts but can admit too many long-lived sessions, exhausting KV capacity and delaying other rollout classes. Reordering for efficiency can also distort the trainer's requested mixture. Key idea. MISA-T adapts session admission to current pressure and assigns KV capacity by workload, accounting for residency across each session while preserving the trainer's rollout mixture. Findings. In a matched 50-iteration Step3.7 run, MISA-T reports 35.6% higher rollout throughput and 22.8% lower mean iteration time. Why it matters. It shows why rollout scheduling cannot stop at prefix reuse. The scheduler must control how many stateful sessions enter the system and how their KV footprints compete, while respecting a statistical contract with training rather than optimizing an unconstrained request stream. |
| Indexing shared prefixes | ||
| SGLang / RadixAttentionSGLang: Efficient Execution of Structured Language Model Programs | 2023 |
Background. Real LLM use is rarely one prompt and one completion. Few-shot prompts, chat turns, agent loops, and tree search all reissue overlapping token sequences, and serving systems of the time freed a request's KV cache the moment it finished. Problem. Every request that shares a prefix recomputes it, and nothing in a plain completion API tells the runtime which spans will come back. The runtime has to discover the sharing itself, then decide what to keep under a fixed memory budget. Key idea. Index the KV cache with a radix tree over token sequences, match each arriving request against its longest cached prefix, and evict with LRU. The frontend language makes a program's structure explicit, so the scheduler can also order requests to land on the cache instead of missing it. Findings. Across agent, reasoning, few-shot, structured-output, retrieval, and chat workloads, SGLang delivers up to 6.4× more throughput than the evaluated state-of-the-art inference systems. Why it matters. It turns prefix reuse into ordinary cache management, with a hit ratio, an eviction policy, and a scheduling decision attached, which is the framing the rest of this meeting works in. Adoption. xAI and LinkedIn use SGLang, providing external production uptake of RadixAttention's automatic prefix reuse. |
| CachedAttentionCost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention | 2024 |
Background. The radix tree above indexes prefixes inside one engine's GPU memory. A chat deployment is where those prefixes come from, because every turn of a conversation resends the whole history as the prefix of the next request. Problem. GPU memory cannot hold every open conversation, so evicted state is recomputed on the next turn. A lower storage tier helps only if fetching beats prefill, and context truncation can invalidate position-dependent cached state. Key idea. Keep KV state in a storage hierarchy. Preload by layer and save asynchronously to overlap transfers with computation, use scheduler hints for placement, and separate positional encoding from stored state so truncation does not invalidate it. Findings. Against recomputing the history KV on every turn, time to first token falls by up to 87%, prompt prefilling throughput rises by up to 7.8×, and end-to-end inference cost falls by up to 70% on multi-turn conversations. Why it matters. Multi-turn state turns the prefix cache into a storage hierarchy rather than a GPU table. Reuse also requires cached bytes to remain valid at their new positions, a constraint that later systems must address. |
| Prefix cache for hybrid attention | ||
| MarconiPrefix Caching for the Era of Hybrid LLMs | 2024 |
Background. Prefix caching assumes cached state is indexed by token: keys and values for position i sit at slot i, so a shared prefix is a shared range of blocks. Hybrid models interleave attention layers with state-space layers. Problem. A state-space layer carries one recurrent state that has already absorbed every token it saw, with no per-token entry to address or split. There is no prefix of it to look up, so block-level reuse and eviction have nothing to operate on. Key idea. Decide where in a sequence recurrent state is worth checkpointing, then admit and evict knowing that hybrid entries differ both in size and in how much recomputation they save. Prefix caching becomes a policy question about state rather than a lookup over token blocks. Findings. Across the evaluated hybrid models and workloads, Marconi raises token hit rate by up to 34.4× over prior prefix caches and reduces time to first token by up to 71.1%, or 617 ms. Why it matters. It shows how much of prefix caching was an artifact of attention's data layout. Once cached state is not addressable by token, you reason about recompute saved per byte held instead of counting block hits, which is the framing any non-attention architecture needs. Adoption. vLLM directly identifies its selective retention policy for hybrid caches as Marconi-style caching and uses the policy in its Kimi K3 serving path. |
| LPCLearned Prefix Caching for Efficient LLM Inference | 2025 |
Background. A chat deployment's prefix cache holds one entry per conversation, and the entry is worth keeping only until the user stops replying. Every engine decides that with LRU, the policy SGLang's radix tree above uses. Problem. LRU ranks entries by how recently the last turn arrived, which is a weak proxy for whether another turn is coming, and the gap to the offline optimum is large. The signal that would close it, whether this particular conversation continues, is not present in access timestamps at all. Key idea. Predict continuation from the conversation itself. Train a predictor over conversational content to estimate which conversations will receive another turn, then combine that estimate with the last access time in the eviction decision. Eviction becomes a prediction problem whose features are the cached text rather than only its access history. Findings. Across three real-world datasets, LPC reaches the hit ratios LRU reaches with caches 18-47% smaller, and raises LLM prefilling throughput by 11% in an emulated environment. Why it matters. It is prefix caching's version of learned eviction, and the feature it uses is what makes it worth reading. A CDN policy can learn only from the request stream, whereas an LLM cache can read the object it is holding, so content becomes a predictor of future access. Read it against LRB in the eviction policies group below, which frames eviction the same way and pays for it with feature collection and inference on the critical path. |
| Memory pool | ||
| MooncakeA KVCache-centric Disaggregated Architecture for LLM Serving | 2024 |
Background. Prefill and decode have different bottlenecks, which is why serving stacks began splitting them onto separate instances, and the KV cache prefill produces is worth keeping for later requests. What a deployment that commits to both looks like under real traffic is a separate question. Problem. At production scale the scarce resource is KV cache capacity and the bandwidth to move it, not weights or FLOPs. Overload is routine rather than exceptional, so a design that assumes every arriving request is admitted plans for the wrong workload. Key idea. Put the cache at the center. Disaggregate prefill from decode, pool KV across the cluster's underused CPU memory and storage so reuse is cluster-wide instead of per-node, and reject requests early under overload instead of admitting everything and missing latency targets. Findings. Mooncake raises SLO-compliant throughput by up to 525% over the simulated baseline. Under Moonshot's real traffic, the deployed architecture lets Kimi handle 75% more requests. Why it matters. It reports what a KV-centric design costs at scale, including the two parts a prototype leaves out: transfer bandwidth and admission control. Read it as the deployment counterpart to the indexing and scheduling mechanisms in this meeting. Adoption. Mooncake's transfer engine ships as a connector in vLLM's KV-transfer interface. |
| MemServeContext Caching for Disaggregated LLM Serving with Elastic Memory Pool | 2024 |
Background. Two things move KV around a deployment. Disaggregated serving splits prefill and decode onto separate instances, so KV must cross the network, and context caching keeps KV across requests, so it must outlive the request that produced it. Problem. When each instance owns its own KV memory, capacity is stranded wherever it was allocated and a cached context can be neither shared nor migrated. A request that could have reused a warm cache on a neighboring instance prefills from scratch. Key idea. Expose one placement and lookup interface over an elastic memory pool spanning instances. Cached context becomes a shareable resource instead of private state inside the instance that computed it. Why it matters. It separates two things usually conflated: where computation happens and where its cached state lives. Once the cache is a pool, disaggregation and cross-request reuse stop being two features and become one mechanism. |
| LMCacheAn Efficient KV Cache Layer for Enterprise-Scale LLM Inference | 2025 |
Background. A serving engine keeps its KV cache in its own GPU memory. That cache dies with the engine process, and nothing is shared between vLLM and SGLang, between replicas of one model, or across queries that arrive minutes apart. Problem. Enterprise deployments hold far more reusable state than one GPU fits and run more engines than one engine-local cache can serve. The work then lands in data movement across GPU, CPU, storage, and network, and prefill-decode disaggregation needs that transfer anyway. Key idea. Make the KV cache a layer under the engine. LMCache plugs beneath both vLLM and SGLang, spans GPU, CPU, storage, and network, pipelines data movement so transfers overlap compute, and carries KV across a prefill-decode split. Why it matters. Worth reading for what enterprise deployment breaks, notably that context truncation halves the prefix cache hit rate. It treats cache capacity as a memory-hierarchy problem rather than an engine feature, which puts the design work in the interface between the two. Adoption. LMCache ships as a KV connector in vLLM's production stack, where it backs the CPU, disk, Redis, and Mooncake offload tiers of a Kubernetes deployment. |
| TokenLakeA Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving | 2025 |
Background. Each serving replica holds its own prefix cache, so the router sends a request to whichever instance already has its prefix. Cache-aware routing of that kind is what production routers do today. Problem. Instance-granularity caching ties placement to the cache. Hot prefixes pin requests to one replica, so load follows the cache instead of the compute, and a long context cached on a saturated instance is unusable to an idle one. Key idea. Pool the prefix cache at segment rather than instance granularity, behind a declarative cache interface, and balance load with knowledge of which segments are the heavy hitters. Any instance can reach any segment, so hot prefixes stop pinning requests. Findings. On real workloads, TokenLake raises throughput by up to 2.6× over cache-aware routing and 2.0× over cache-centric prefill-decode disaggregation. Cache hit rate improves by up to 2.0× and 2.1×, respectively. Why it matters. Read against cache-aware routing and cache-centric prefill-decode disaggregation, which are its two baselines. It separates where cached state lives from where compute happens, and that separation is what makes capacity elastic rather than per-instance. |
| Approximate reuse of non-prefix spans | ||
| Prompt CacheModular Attention Reuse for Low-Latency Inference | 2023 |
Background. Prefix caching reuses a span only when it begins at token zero and matches exactly, which covers a system prompt at the front of every request and little else. Problem. The text that actually recurs is often not a prefix: a document in the middle of a prompt, one tool description among several, a snippet whose position shifts between calls. Keys and values depend on position, so state cached at one offset is wrong at another. Key idea. Declare the reusable spans as prompt modules in a schema, with their position ids fixed in advance. Each module's attention state is computed once and stays valid wherever the module lands in the assembled prompt, so reuse no longer has to run contiguously from the start. Findings. The prototype reduces time to first token by up to 8× on GPU and 60× on CPU for long-prompt workloads, while preserving the reported output accuracy. Why it matters. It moves reuse from something the runtime infers to something the application states, and it identifies position dependence as the obstacle to reusing anything other than a prefix. |
| CacheGenKV Cache Compression and Streaming for Fast Large Language Model Serving | 2023 |
Background. Prefix caching lets a server skip prefill for context it has already processed, but the saving only lands if the cache is reachable from the GPU that needs it. In a multi-node deployment the KV for a long context sits on another machine or on disk. Problem. That KV is large. Fetching it across a link can take longer than recomputing prefill from scratch, so reuse becomes a bandwidth problem rather than a memory problem, and the number that settles it is time to first token rather than bytes stored. Available bandwidth also moves under load. Key idea. Treat the cached tensors as something to encode rather than merely to store. Compress them into compact bitstreams that exploit the structure they already have, spend fewer bits where the model tolerates more error, and stream them at a compression level chosen for the bandwidth on hand, recomputing instead when even that would arrive too late. Findings. Across four datasets and three models, CacheGen reduces time to first token by 3.1-4.7× against recomputing from text and by 3.2-3.7× against uniform KV quantization. Why it matters. It puts a price on a cache hit. Across nodes and disks, a cache hit that arrives after recomputation would have finished is a miss. Adoption. LMCache ships CacheGen-style KV encoding and streaming for its disk and remote cache tiers, which is how the vLLM production stack moves reused KV between nodes. |
| CacheBlendFast Large Language Model Serving for RAG with Cached Knowledge Fusion | 2024 |
Background. Prefix caching hits only when a prompt starts with a cached prefix. A RAG prompt concatenates several retrieved chunks, and the same chunk appears at a different position, after different neighbors, in every query that retrieves it. Problem. RAG chunks share no common prefix, so only the first chunk is ever a hit. Precomputed per-chunk KV is not directly reusable either, because it was computed without attention to the chunks placed before it, and splicing the pieces together degrades output quality. Key idea. Reuse the precomputed KV for every chunk, then recompute selectively, only the small set of tokens whose KV deviates most from what a full prefill would have produced, so the fused cache stays close enough to be correct. Findings. Across three models and four datasets, CacheBlend reduces time to first token by 2.2-3.3× and raises throughput by 2.8-5× over full recomputation without measured quality loss or additional storage. Why it matters. It turns reuse from a yes-or-no correctness question into a knob you can turn, trading recompute against fidelity. That knob is what the rest of this group varies, and it is the reason non-prefix reuse is usable at all. Adoption. LMCache ships CacheBlend as its non-prefix reuse path, which is how vLLM users get blended RAG chunk caches today. |
| DroidSpeakKV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving | 2024 |
Background. Everything above reuses KV that one model computed for that same model. A compound or agentic deployment runs several models specialized for different users, tasks, or roles, and they pass the same context between them, so the same prefix is prefilled once per model. Problem. KV computed by one model is not KV for another. Two models sharing an architecture still hold different weights in every layer, so one model's cache is wrong in the other, and whether that error costs output quality had not been measured. Key idea. Reuse most of the layers and recompute the rest. Study which layers' KV actually diverges between the two models, recompute only those, reuse the remainder as it stands, and pipeline the layer-wise recomputation against the loading of the reused layers so the two overlap. Reuse holds so long as the models share an architecture. Findings. Up to 4× higher throughput and about 3.1× faster time to first token, against a baseline that shares nothing across models, with negligible loss in F1, Rouge-L, or code similarity across several datasets and model pairs. Why it matters. It moves the selective-recompute knob to a second axis. CacheBlend recomputes the tokens a changed position invalidated, and DroidSpeak recomputes the layers a changed model invalidated, which is the axis a multi-model agent deployment sits on. |
| HYPICAccelerating Hybrid-Attention LLM Serving with Position-Independent Caching | 2026 |
Background. Prefix caching hits only on an exact shared prefix. Position-independent caching relaxes that, reusing KV for segments that are not a shared prefix, which is what retrieval and agent prompts actually look like. Hybrid-attention models interleave full-attention layers with linear-attention layers that carry a per-request recurrent state. Problem. Position-independent reuse breaks on hybrid-attention models because a per-token KV primitive has no counterpart in a per-request recurrent state. You can concatenate the cached KV of two segments; you cannot concatenate two recurrent states and get the state the full prompt would have produced. Key idea. Cache each segment's cumulative transition operator alongside its zero-start end state, which composes independently cached segments in constant time, then repair cross-segment attention in the remaining full-attention layers with a small seam window. Findings. Across four hybrid-attention models and five workloads, time to first token falls 3.25× on average and QPS rises 1.66× against prefix caching, and task quality stays within 1.71 points of full recomputation. Why it matters. It says what a cache entry has to be for reuse to compose: an operator you can apply, not only bytes you can concatenate. That is the same checkpointing question Marconi raises above, asked where segments have to combine rather than merely be looked up. |
| Eviction policies | ||
| S3-FIFOFIFO Queues are All You Need for Cache Eviction | 2023 |
Background. LRU has been the default eviction policy for decades, and the policies that beat it on hit ratio usually do so by tracking more metadata per object and by reordering the data structure on every access. Problem. That bookkeeping costs throughput and lock contention, and most of it is spent on objects that never come back. Real cache workloads are dominated by one-hit wonders, objects requested exactly once, which LRU still admits to the full cache and promotes to the most-recently-used position. Key idea. Use three FIFO queues. A small probationary queue quickly removes one-hit objects, a main queue holds objects accessed twice, and a ghost queue remembers evictions. Hits never reorder a queue. Why it matters. It beats LRU on hit ratio and on throughput at the same time, where those two are normally traded against each other, and it is the clearest demonstration that one-hit wonders dominate real workloads. Read it before designing any prefix-cache eviction policy. Adoption. S3-FIFO runs in production at Google, VMware, and Redpanda, and it is an eviction policy in open-source caches including Rust's foyer and Go's otter. |
| SIEVEis Simpler than LRU: an Efficient Turn-Key Eviction Algorithm for Web Caches | 2024 |
Background. Policies that improve on LRU usually arrive as new data structures: extra queues, frequency sketches, learned predictors. Adopting one means rewriting the cache, which is part of why so many production caches still run LRU. Problem. Promotion on every hit is what makes LRU expensive, because it mutates shared list state on the read path. Dropping promotion normally costs the hit ratio that promotion buys, so simplicity and competitiveness look like opposite choices. Key idea. Keep the FIFO queue and promote lazily. A hand walks the queue, evicting objects that have not been accessed since it last passed and clearing the bit on those that have, so retained objects never move. That is about ten lines of diff over FIFO. Why it matters. A good model for how simple a competitive policy can be, and for the difference between eager and lazy promotion. What a policy costs is not only its miss ratio; it is also what adopting it does to the read path and to the code. Adoption. SIEVE was merged into five production cache libraries with fewer than twenty lines of change each, and it now appears as a policy option in Go, Rust, and JavaScript caches. |
| LRBLearning Relaxed Belady for Content Distribution Network Caching | 2020 |
Background. Belady's algorithm evicts the object whose next request is furthest in the future. It is optimal and it is offline, because it needs the future request sequence, so every deployed policy is a heuristic stand-in for it. Problem. A heuristic encodes one fixed assumption about reuse, and a workload that violates the assumption loses hit ratio. Worse, the policy itself tells you nothing about how much of the offline optimum you are leaving on the table. Key idea. Relax Belady. Rather than predict an exact next-access time, train a model on features collected online to predict which sampled candidates have their next request beyond a boundary, and evict from those. Eviction becomes a supervised prediction problem. Why it matters. Read it for the framing of eviction as prediction, and for how much machinery that framing costs. Feature collection, training, and inference all land on the cache's critical path, which is the price of chasing the optimum. |
| TinyLFUA Highly Efficient Cache Admission Policy | 2015 |
Background. Cache policy discussions usually start and end with eviction: given a full cache, which object leaves. Admission, the decision of whether a newly requested object belongs in the cache at all, is the other half and is usually left implicit. Problem. Frequency is the right signal for admission, and exact frequency counts over a long history cost more memory than the objects they protect. A cache that spends its space on counters has less space left for data. Key idea. Approximate the history with a sketch. A few bits per entry hold an aging estimate of each key's recent frequency, and a candidate is admitted only if its estimated frequency beats the eviction victim's. Windowed TinyLFU puts a small LRU window in front so bursty new keys still get a chance. Why it matters. Admission, not eviction, is often the lever, and a sketch is enough to make the decision. It is also the clean case of trading bounded accuracy for a large cut in metadata, which is the same trade a per-block prefix cache has to make. Adoption. W-TinyLFU is the admission filter in Caffeine, which backs the caches in Cassandra, Solr, and Druid, and the design was carried into Go's Ristretto and Rust's moka. |
| Cache engines and simulators | ||
| libCacheSima high-performance cache simulator and library | — |
Background. Cache research is comparative. A new policy is only interesting relative to LRU, FIFO, and the offline optimum on the same traces, and running that comparison from scratch means reimplementing every baseline and every trace format. Problem. Reimplemented baselines are where hit-ratio claims go wrong. Two papers reporting LRU on the same trace can disagree because of object-size handling, admission, warmup, or trace parsing, and none of that disagreement is visible in the numbers they publish. Key idea. Package standard baselines in one simulator with a library interface, so a new policy is a small addition to a shared harness rather than a standalone program. Every policy runs on the same trace reader. Why it matters. It is the standard cache simulator for this kind of work, so read the docs before you write a policy. It produced the S3-FIFO and SIEVE evaluations above, and it replays the open Twitter, Meta, and Wikimedia traces that most published miss-ratio comparisons now use. |
| CacheLibThe CacheLib Caching Engine: Design and Experiences at Scale | 2020 |
Background. A large web service runs many separate caches, in front of its CDN, its social graph, its key-value storage, and more. Each has its own workload and object type, and historically each team built its own cache. Problem. Most of what a production cache does has nothing to do with the eviction policy. It has to size memory across pools, admit selectively to flash so the device is not written to death, and come back warm after a deploy. A team re-solving all of that solves some of it badly. Key idea. Provide one embeddable library behind a common API for DRAM and flash tiers, flash admission, per-pool sizing, structured item types, and cache contents that survive process restarts. Findings. The production study reports that one CacheLib server can replace tens of backend database servers, with 20× higher throughput and hit ratios above 80% for the cited workload. Why it matters. It is the clearest account of what a production cache has to handle beyond the eviction policy: sizing, admission, flash, and warm restarts. Read it as the list of conditions a research policy has to survive contact with. |
Quantization asks how to represent the same model with fewer bits after training. The first groups establish the available number formats, explain why activation outliers make a uniform INT8 conversion fail, and turn those observations into post-training weight-quantization methods. The later groups extend the same question to the growing KV cache, ask which low-precision kernels turn smaller tensors into speed, and follow the numerical effects into fine-tuning, extreme formats, and model reliability. Rotations change the distribution before quantization, while pruning removes values or structures instead of representing them more compactly. Read every group against the same test: a reduction in bits or FLOPs matters only when the hardware executes less work and the model retains the behavior the application needs.
| Paper | Year | Why read it |
|---|---|---|
| Number formats | ||
| mixed precisionTraining | 2017 |
Background. FP32 was the default numeric type for training. Half precision halves the memory a tensor occupies and the bandwidth it costs to move, and on hardware with half-precision units it also raises arithmetic throughput. Problem. FP16 has too few mantissa bits and too narrow an exponent range for the whole training step. Small gradients flush to zero before they are ever accumulated, and an update far smaller than the weight it modifies vanishes in the rounding. Key idea. Three rules. Keep an FP32 master copy of the weights and apply every update to it; scale the loss up before the backward pass so gradients land inside the FP16 range, then unscale before the update; accumulate reductions in FP32. Findings. Across convolutional, recurrent, and generative-adversarial models, including networks above 100 million parameters, the recipe matches full-precision training while reducing model memory by nearly 2×. Why it matters. This is the template every later low-precision recipe follows. Decide where precision is cheap, storage and matmuls, and where it is not, accumulation and the update, then add a scale factor to move values into the representable range. FP8 and FP4 repeat that shape with less room. Adoption. PyTorch torch.amp, JAX, and NVIDIA Apex all implement this recipe, and loss scaling with an FP32 master copy is the default in Megatron-LM, DeepSpeed, and Hugging Face Trainer. |
| FP8 formatsfor Deep Learning | 2022 |
Background. Mixed-precision training established that a network can be stored and multiplied at fewer bits than it is updated in. At eight bits so few bits are left that how you split them between exponent and mantissa becomes the design decision. Problem. One 8-bit float cannot serve the whole step. Weights and activations need mantissa bits near their typical magnitude, while gradients span a much wider dynamic range and need the exponent bits more. Key idea. Standardize two types rather than one. E4M3 spends the contested bit on the mantissa and carries the forward pass, E5M2 spends it on the exponent and carries gradients, and per-tensor scaling covers what the exponent bits cannot. Findings. Across CNNs, RNNs, and transformer models as large as 175B parameters, FP8 training matches the quality of the corresponding 16-bit sessions without changing their hyperparameters. Why it matters. It presents a format as an allocation of dynamic range, and the tensor's role in the computation as what decides the allocation. Once E4M3 and E5M2 read that way, the 4-bit microscaling formats below are the same tradeoff with less to spend. Adoption. E4M3 and E5M2 are the two FP8 types in NVIDIA's Transformer Engine and in Hopper and Blackwell tensor cores, and they are PyTorch's float8_e4m3fn and float8_e5m2; vLLM serves FP8 weights and KV cache in them. |
| INT vs FPINT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats | 2025 |
Background. Low-bit quantization now offers two families at the same bit width: integers with a scale per block, and small floating-point types with their own per-block scale. Each is argued for from first principles, and hardware exists for both. Problem. The arguments are made at different block sizes and bit widths, so they do not compare. A reported win can come from the format or from a finer scale, and published numbers taken across papers cannot separate the two. Key idea. Hold block granularity and bit width fixed and vary only the format. Findings. The better format changes with block size and bit width, so neither integers nor floating point dominate under every controlled configuration. Why it matters. It removes an assumption people deploy on, that FP4 beats INT4 because it is floating point, and it promotes block size from an implementation detail to a variable you have to state before a format comparison means anything. |
| microscaling FP4Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization | 2025 |
Background. Microscaling formats store a block of 4-bit values under one shared scale, and current hardware offers two of them. MXFP4 restricts that shared scale to a power of two; NVFP4 does not. Problem. Measured accuracy at 4 bits falls short of what the bit width promises, and the shortfall differs between the two formats. Until you know where it comes from, choosing between them is a guess dressed as a specification. Key idea. Attribute the loss to the shared block scale rather than assuming it comes from the 4-bit payload. Findings. MXFP4's power-of-two scale fits each block's range more coarsely than NVFP4's scale, accounting for the observed accuracy gap between formats at the same payload width. Why it matters. Two formats at one bit width with different accuracy make the choice between them a real deployment decision. It also locates a block format’s accuracy in its scale rather than in its elements, and treats the numerics as something you measure rather than read off a vendor spec sheet. Adoption. Both formats ship: MXFP4 is how GPT-OSS releases its weights, and MXFP4 and NVFP4 are native Blackwell tensor-core types that vLLM and TensorRT-LLM serve. |
| Outliers, and why naïve INT8 fails | ||
| LLM.int8()8-bit Matrix Multiplication for Transformers at Scale | 2022 |
Background. Post-training INT8 quantization was routine for smaller networks. Scale a tensor into the integer range, round, and do the matmul in INT8, for roughly half the memory traffic and faster arithmetic at negligible accuracy loss. Problem. The same recipe breaks on large transformers. A few feature dimensions carry activation magnitudes far above the rest, so a scale wide enough to cover them leaves almost no resolution for everything else, and accuracy collapses past a model size rather than degrading smoothly. Key idea. Separate the outliers instead of scaling around them. Vector-wise quantization gives each row and column its own scale, and the small set of outlier dimensions is pulled out and multiplied in FP16 while the rest of the matrix stays INT8. The two partial products are summed. Findings. More than 99.9% of values remain in the 8-bit path. The method runs models through 175B parameters without measured performance degradation while halving inference memory, making OPT-175B and BLOOM runnable on one server with consumer GPUs. Why it matters. This is where the outlier problem was identified, and outliers are what every later quantization method on this page is arranged around, whether by migrating them into the weights or by keeping them in higher precision. Read it for why naive INT8 fails on large models specifically. Adoption. The mixed-precision decomposition ships in bitsandbytes, which Hugging Face Transformers exposes through its 8-bit loading path. |
| SmoothQuantAccurate and Efficient Post-Training Quantization for Large Language Models | 2022 |
Background. Eight-bit quantization of both weights and activations is what lets the integer tensor-core path do the matmul. Weights quantize cleanly at that width. Activations do not, because a few channels carry magnitudes far larger than the rest. Problem. A per-tensor activation scale has to cover those channels, so every other value in the tensor is left with a fraction of the range it could have used. Keeping the outlier channels in higher precision preserves accuracy but splits one GEMM into two. Key idea. Divide each activation channel by a factor and multiply the matching weight row by the same factor. The product is unchanged, the factor folds into the preceding layer, and a migration strength setting decides how much of the difficulty moves from activations onto weights until both quantize at eight bits. Findings. Across the evaluated model families, W8A8 SmoothQuant preserves accuracy while reducing model memory by 2× and accelerating inference by up to 1.56× over the paper's full-precision implementation. The resulting memory footprint is small enough to serve a 530B-parameter model within one node. Why it matters. The outlier problem turns out not to need special hardware or a split kernel; it needs a change of variables chosen offline and free at run time. It is the cleanest example on this page of reparameterizing a numerics problem instead of absorbing its cost. Adoption. TensorRT-LLM ships SmoothQuant as an INT8 weight-and-activation recipe, and llm-compressor applies the same per-channel rescale when producing the W8A8 checkpoints that vLLM serves. |
| ZeroQuantEfficient and Affordable Post-Training Quantization for Large-Scale Transformers | 2022 |
Background. By 2022 post-training quantization for transformers was mostly reported as a bit width and an accuracy number, with the kernels that would turn either into a speedup left to someone else. Problem. A quantized model is faster only if the whole layer runs quantized. Per-tensor scaling costs accuracy on transformers, finer scaling costs extra scale traffic, and dequantizing between operators can give back the arithmetic the lower precision saved. Key idea. Three parts, delivered together. Fine-grained scaling, group-wise on weights and token-wise on activations; layer-by-layer distillation against the original model, which needs no training data, to recover the accuracy; and a fused kernel backend so the finer scaling does not cost bandwidth. Findings. INT8 weights and activations accelerate the evaluated BERT- and GPT-3-style models by up to 5.19× and 4.16×, respectively, over FP16 with minimal accuracy loss. Adding INT4 feed-forward weights reduces model memory by 3× relative to FP16. Why it matters. It is the useful baseline because it names all three pieces of the problem at once: the scheme, the accuracy recovery, and the kernel. Read it as the reminder that a reported bit width is half a claim, and whether a kernel exists is the other half. |
| Post-training weight quantization | ||
| GPTQAccurate Post-Training Quantization for Generative Pre-trained Transformers | 2022 |
Background. A decode step reads every weight once, so generation speed follows bytes moved rather than arithmetic. Weight-only post-training quantization buys that bandwidth back, and the crude version rounds each weight to the nearest representable value. Problem. Rounding weights independently ignores what the layer's output does with them. The accurate alternative reconstructs each layer's output using second-order information about its inputs, and that was too slow to run on a model with billions of parameters. Key idea. Quantize a layer column by column, and after each column update the weights not yet quantized so they absorb the error just introduced, using the inverse Hessian of the layer inputs. Batched updates and a Cholesky reformulation keep the pass fast and numerically stable. Findings. GPTQ quantizes OPT-175B in about four GPU hours to 3 or 4 bits per weight with negligible accuracy degradation. Against FP16 inference, the paper reports end-to-end speedups of 3.25× on an NVIDIA A100 and 4.5× on an NVIDIA A6000. Why it matters. Error compensation in a single pass over a calibration set is the template every later post-training method argues with. It also makes quantization an operation you apply to a released checkpoint instead of a property you have to train in. Adoption. GPTQ checkpoints are a standard release format. AutoGPTQ and GPTQModel produce them, Hugging Face Transformers loads them, and vLLM, SGLang, and ExLlama all ship GPTQ kernels. |
| AWQActivation-aware Weight Quantization for LLM Compression and Acceleration | 2023 |
Background. Weight-only quantization works because decode is bandwidth-bound, and its accuracy loss is concentrated rather than spread out. A small number of weight channels account for most of the output error once the format gets narrow. Problem. Weight magnitude does not identify those channels. Keeping the ones that matter in higher precision does identify them, but it leaves a mixed-precision layout that no single kernel serves efficiently. Key idea. Find the salient 1% of channels from activation statistics, since the activation flowing through a weight decides how much its rounding error costs, then scale those channels up before quantization so that error shrinks. The format stays uniform and one kernel handles the whole matrix. Findings. AWQ preserves accuracy across language, instruction-tuned, and multimodal models without backpropagation or reconstruction. Its TinyChat implementation runs more than 3× faster than the Hugging Face FP16 baseline on the evaluated desktop and mobile GPUs and serves Llama-2-70B on a mobile GPU. Why it matters. Not all weights matter equally, and importance lives in the activations rather than in the weights. It also shows the deployment constraint doing real work, because a fix that breaks format uniformity is not a fix. Adoption. Hugging Face Transformers loads AWQ checkpoints directly, and vLLM, SGLang, and TensorRT-LLM ship AWQ kernels. |
| OmniQuantOmnidirectionally Calibrated Quantization for Large Language Models | 2023 |
Background. Post-training quantization has knobs. You choose where to clip each weight distribution, and which equivalent transform to apply first so activation outliers move into the weights, the way SmoothQuant does with a migration factor. Problem. Those knobs are set by hand or by grid search, layer by layer, and the right setting depends on the layer and on the bit width. Hand-tuning stops working once both weights and activations go to low precision. Key idea. Make the knobs parameters and learn them. Learnable weight clipping and a learnable equivalent transform are optimized block by block against the full-precision block's output, so gradients do the tuning without retraining the model end to end. Findings. OmniQuant calibrates Llama-2 models from 7B to 70B parameters on one A100 40 GB GPU in 1-16 hours using 128 samples. The evaluation spans weight-only and weight-activation settings from 2 to 6 bits. Why it matters. The line between post-training quantization and quantization-aware training is a question of how many parameters you optimize, not a hard boundary. Learning a transform rather than picking one is exactly what the rotation methods after it do. |
| SqueezeLLMDense-and-Sparse Quantization | 2023 |
Background. Weight-only quantization is a bandwidth play, so the question is how few bits a weight can carry. Uniform quantization answers it by spacing the representable values evenly across the range the weights occupy. Problem. Weight distributions are not uniform. Most weights sit in a narrow band while a few extremes stretch the range, so an even grid spends its levels where almost no weights are and rounds hardest the values whose error propagates furthest. Key idea. Two changes. Place the quantization levels non-uniformly, chosen by a sensitivity-weighted clustering rather than by the raw weight distribution, and hold a small set of outlier and sensitive weights in a separate sparse full-precision matrix while the dense remainder goes low-bit. Findings. On Llama models at 3 bits, SqueezeLLM reduces the perplexity gap from FP16 by up to 2.1× relative to methods under the same memory budget. Its A6000 implementation runs up to 2.3× faster than the paper's baseline. Why it matters. It splits low-bit failure into two causes with separate fixes, bad level placement and a few extreme values, and it establishes a sparse side path as the alternative to widening the format for every weight. Adoption. vLLM ships a dedicated SqueezeLLM quantization path and kernel. |
| Quantizing the KV cache | ||
| KIVIA Tuning-Free Asymmetric 2bit Quantization for KV Cache | 2024 |
Background. The KV cache grows linearly with context and batch size, so quantizing it is the largest capacity cut available without touching the model weights. The default move is to pick one bit width and one grouping axis and apply both to the whole cache. Problem. Keys and values do not share a distribution. Keys carry large outliers concentrated in particular channels, so grouping keys per token lets a few channels set the scale for everything else in that token; values have no such pattern and quantize well per token. One scheme for both falls apart at 2 bits. Key idea. Quantize asymmetrically by axis: per-channel for keys and per-token for values, both at 2 bits. The method requires no tuning, so it can be applied directly to a new model. Findings. Across Llama, Falcon, and Mistral models, KIVI retains nearly the original quality while cutting peak memory by 2.6×. The saved capacity supports batches up to 4× larger and raises measured inference throughput by 2.35-3.47×. Why it matters. It is the standing example that the granularity of a compression scheme, not only its bit width, decides whether the scheme survives contact with a real model. It also sets the floor the rest of this group is measured against. Adoption. Hugging Face Transformers bases its quantized KV cache on KIVI's asymmetric design. |
| KVQuantTowards 10 Million Context Length LLM Inference with KV Cache Quantization | 2024 |
Background. Long-context inference is bounded by KV cache capacity rather than by the model, since the cache grows with sequence length and batch size and is held at half precision by default. Quantizing it is the largest cut available without touching the weights. Problem. Keys and values do not tolerate a uniform low-bit grid. A few numerical outliers stretch the range every other element has to share, so a straightforward 4-bit cast costs accuracy that long-context tasks cannot absorb. Key idea. Fit the quantizer to the distribution instead of assuming a uniform one. Keys are quantized per channel and before the rotary embedding is applied, values per token, the levels are spaced non-uniformly, and the small fraction of outliers is kept separately at higher precision. Findings. Three-bit KVQuant adds less than 0.1 perplexity on WikiText-2 and C4. It fits a one-million-token Llama-7B context on one A100 80 GB, reaches ten million tokens on eight GPUs, and accelerates the evaluated matrix-vector kernels by up to about 1.7× over FP16. Why it matters. It moves the binding constraint. Once the cache costs a few bits per element, how much context you can serve is set by what the attention implementation can process rather than by what memory can hold, which is what the ten-million-token figure in the title measures. |
| ZipCacheAccurate and Efficient KV Cache Quantization with Salient Token Identification | 2024 |
Background. KV-cache quantization normally assigns one precision to every cached token. That uniform treatment ignores the fact that some tokens have much more influence on later attention outputs than others. Problem. Lowering every token to the same bit width either wastes memory on insensitive tokens or damages the salient tokens that preserve model quality. Static mixed-precision schemes also need a practical way to identify those tokens before the cache is stored. Key idea. Estimate token salience from the norm of its key and value vectors, retain salient tokens at higher precision, and quantize the rest more aggressively. Channel-separable quantization and fused CUDA kernels keep the mixed-precision layout efficient during attention. Findings. On Mistral-7B and GSM8K, ZipCache reports a 4.98× cache compression ratio with a 0.38-percentage-point accuracy decrease. For Llama-3-8B with a 4,096-token input, it reduces prefill latency by 37.3%, decoding latency by 56.9%, and GPU memory use by 19.8% against the evaluated baseline. Why it matters. ZipCache treats token importance as a cache-allocation signal rather than forcing a single precision across the sequence. It complements KIVI's axis-aware quantization and QServe's kernel co-design by adding a third decision: which tokens receive the scarce high-precision budget. |
| Serving a quantized model at speed | ||
| QServeW4A8KV4 Quantization and System Co-design for Efficient LLM Serving | 2024 |
Background. Weight-only 4-bit quantization, GPTQ and AWQ, was routine by 2024, and the rotation methods had pushed activations and the KV cache down as well. The wins those papers report are compression ratios and perplexity. Problem. A smaller checkpoint is not a faster server. Dequantization sits on the critical path of every matmul, and when its arithmetic runs on the same units as the multiply, the instructions it adds can cost more than the memory traffic it saves. Key idea. Pick the precision assignment and the kernel together. Weights at 4 bits, activations at 8 bits, and the KV cache at 4 bits keeps every general matrix multiply on integer tensor cores, and the dequantization is arranged so that unpacking does not dominate the multiply it feeds. Findings. Against TensorRT-LLM, QServe raises maximum serving throughput by 1.2× on an A100 and 1.4× on an L40S for Llama-3-8B. For Qwen1.5-72B, the reported gains rise to 2.4× on an A100 and 3.5× on an L40S. Why it matters. It takes dequantization overhead seriously and measures the result end to end, which is where the gap between a good compression ratio and a good token rate becomes visible. Read it as the reason a precision choice is a systems decision and not only a numerics one. |
| AtomLow-bit Quantization for Efficient and Accurate LLM Serving | 2023 |
Background. A quantization result is usually a bit width and a perplexity. Serving a model at that bit width is a separate question, because the kernel has to dequantize, the KV cache still occupies memory, and the batch size decides which resource binds. Problem. A compression ratio does not convert into tokens per second. Finer-grained scales and mixed-precision handling improve accuracy while adding dequantization work and irregular memory access, which can cancel the arithmetic saving the narrower format bought. Key idea. Treat low-bit quantization as a serving design. Quantize weights, activations, and the KV cache with fine-grained group-wise scales, route outlier channels through a higher-precision path, fuse dequantization into the operators, and report end-to-end throughput rather than perplexity. Findings. At the same latency target, Atom's evaluated 4-bit weight-activation path delivers up to 7.7× the end-to-end throughput of FP16 and 2.5× the throughput of INT8 while maintaining accuracy. Why it matters. It sets the standard of evidence for the rest of this group. A method that wins on perplexity at a given bit width has shown nothing yet about serving, and the gap between the two is where the systems work sits. |
| SageAttention3Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training | 2025 |
Background. Low-precision serving work has concentrated on weights, activations, and the KV cache, leaving the two matmuls inside attention at higher precision. Blackwell tensor cores added native FP4 microscaling, a format with very few mantissa bits and one scale per block. Problem. The post-softmax probability matrix is the hard case. Its values sit in a narrow [0,1] range, which a single block scale represents badly, so quantizing attention itself is not the same problem as quantizing the tensors around it. Key idea. Quantize both attention matmuls in FP4 microscaling on Blackwell tensor cores, with a two-level scheme for the probability matrix so the narrow range gets its own scaling. Findings. A companion 8-bit training variant preserves fine-tuning quality but converges more slowly during pretraining. Why it matters. Attention is where the sequence-length term lives, so precision there buys something different from precision in the weights. The training comparison also separates a workload's tolerance for low precision from the format itself. |
| TilusA Tile-Level GPGPU Programming Language for Low-Precision Computation | 2025 |
Background. GPU kernels for quantized inference are written per precision pair, by hand, against the tensor-core layouts of one architecture. The stacks that serve models ship a short list of them: 4-bit and 8-bit weights against 8-bit or 16-bit activations. Problem. Anything off that list, W3A16 or W5A8, needs a new kernel. Values narrower than a byte do not align to register boundaries, so both the packing layout and the unpacking arithmetic change with every bit width, and that engineering cost, not accuracy, is what keeps odd widths out of production. Key idea. Make bit width part of the type in a tile-level GPGPU language, and describe register distribution with a layout system. The compiler then generates packing and dequantization code for each width. Findings. Across the evaluated low-precision kernels, Tilus runs 1.75× faster than Triton, 2.61× faster than Ladder, 1.29× faster than QuantLLM, and 1.03× faster than Marlin. Why it matters. It answers a question the quantization papers leave open, which is why the accuracy-optimal bit width is never the one you can serve. Move layout into the compiler and the bit width becomes a parameter you search rather than a kernel you commission. |
| Quantized fine-tuning | ||
| QLoRAQLoRA: Efficient Finetuning of Quantized LLMs | 2023 |
Background. LoRA freezes the base model and trains small low-rank adapters, eliminating most gradient and optimizer state. The frozen base weights can still dominate memory when held at 16 bits. Problem. A 65B-parameter base model does not fit on one commodity accelerator at 16-bit precision, even when only its adapters are trainable. Quantization must save that memory without blocking gradients to the adapters or destabilizing the optimizer. Key idea. Backpropagate through a frozen 4-bit base model into 16-bit LoRA adapters. NormalFloat 4 represents normally distributed weights, double quantization compresses the quantization constants, and paged optimizers absorb memory spikes. Findings. QLoRA fine-tunes a 65B-parameter model on one 48 GB GPU while matching the task performance of full 16-bit fine-tuning. The reported Guanaco model trains in 24 hours on one GPU and reaches 99.3% of ChatGPT's score on the paper's Vicuna evaluation. Why it matters. The precision of frozen weights and the precision of trainable updates become separate decisions. This composition makes large-model fine-tuning a memory-allocation problem rather than a requirement for a multi-GPU training cluster. |
| The extreme end: 4-bit, 2-bit, ternary | ||
| BitNet b1.58The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits | 2024 |
Background. Post-training quantization takes a model trained at 16 bits and rounds it afterwards, which treats precision as a deployment decision made once the weights are already fixed. Problem. Rounding after the fact asks a network to tolerate a grid it never saw while training. At the extreme end no assignment of a 16-bit model to a handful of levels holds accuracy, so the memory and energy win that motivates the whole line stays out of reach. Key idea. Train from scratch with every weight constrained to -1, 0, or +1, which is 1.58 bits each, so the network learns around the constraint instead of being rounded into it. The weight path then needs additions rather than multiplications. Findings. At matched model size and training tokens, BitNet b1.58 matches the reported perplexity and downstream-task performance of FP16 and BF16 transformers while reducing serving memory and arithmetic cost. Why it matters. It moves quantization out of post-processing and into the training objective, and it changes the hardware question with it. Read it for the claim that quantization belongs in training, not after: if weights are ternary from the start, the arithmetic a serving chip needs is not the arithmetic a GPU is built around. Adoption. Hugging Face Transformers carries a BitNet quantization integration. |
| BitNet b1.58 2B4TBitNet b1.58 2B4T Technical Report | 2025 |
Background. The 1-bit line argued that extreme quantization belongs in training rather than in a rounding pass afterwards. Until this report there was no open native 1.58-bit model at a size anyone would serve. Problem. A format proposal leaves the two questions that decide whether ternary weights are usable: whether a natively ternary model matches full-precision models of the same size, and whether the memory and latency argument survives a real inference stack. Key idea. Train a 2B-parameter model natively at 1.58 bits on 4T tokens and release GPU and CPU inference implementations alongside the weights. Findings. Across the report's evaluations, the resulting ternary model performs on par with full-precision open models of similar size. Why it matters. The memory, energy, and decoding-latency claims for ternary weights arrive with a runnable stack rather than a format proposal, so you can check them yourself. It is also the reference point for anyone testing ternary inference kernels, since Microsoft's bitnet.cpp runs it on CPU. |
| ParetoQImproving Scaling Laws in Extremely Low-bit LLM Quantization | 2025 |
Background. Quantization-aware training results at 1, 1.58, 2, 3, and 4 bits arrive from different papers, each with its own recipe, model family, and token budget. Problem. When every bit width is tuned separately, a reported 2-bit number tells you about the effort spent on it as much as about the format. There is no way to say where the representation actually gives out, so “how low can you go” stays anecdote. Key idea. Train 1, 1.58, 2, 3, and 4-bit models under one unified recipe and budget so the widths are directly comparable. Findings. Performance does not decline smoothly as precision narrows. The controlled comparison identifies a sharp representational break between 2 and 3 bits. Why it matters. It turns the extreme low-bit question into a Pareto curve you can read off: for a fixed memory budget, which pairing of bit width and parameter count wins. It also locates the break in the representation rather than in how hard someone tuned. |
| QuartetNative FP4 Training Can Be Optimal for Large Language Models | 2025 |
Background. Low-precision training is normally mixed precision: the tensor cores run narrow while master weights and some layers stay wide. Blackwell adds native hardware support for 4-bit formats. Problem. Mixed precision hides the question it is answering. If some layers fall back to higher precision, the measured speedup is not the format's, and nothing tells you which bit width buys the most accuracy per unit of compute. Key idea. Fit a low-precision scaling law across bit widths and training configurations, read the accuracy-per-unit-compute optimum off it, then build kernels that realize it. Every linear layer stays natively in FP4 with no mixed-precision fallback. Findings. Under the paper's compute-normalized comparison, native FP4 training remains competitive with FP8 and FP16. Why it matters. It treats precision as a systems decision with an optimum rather than a knob to turn down until quality breaks. That is the template: price the accuracy cost per bit against the compute the hardware gives you at that width, then choose. |
| NVFP4 pretrainingPretraining Large Language Models with NVFP4 | 2025 |
Background. NVFP4 is a 4-bit microscaling format that Blackwell tensor cores execute natively, which makes pretraining at 4 bits a hardware option rather than a simulation. Problem. Four bits offer too few levels to absorb the outliers and the gradient dynamic range a long pretraining run produces, and a run that diverges halfway costs the whole budget. The open question is not whether FP4 matmuls are fast but which parts of the model tolerate them. Key idea. Carry out a large pretraining run in FP4 and report the stabilizers it required: Hadamard transforms to spread outliers, stochastic rounding to keep the gradient unbiased, and selected layers held at higher precision. Findings. A 12B-parameter model trained for 10T tokens with the NVFP4 recipe matches the FP8 baseline's training loss and downstream-task accuracy. Why it matters. The list of exceptions is the result. It shows which layers refuse to go low-precision even at scale, which is what you need before committing a training budget at 4 bits. |
| stable FP4 trainingStable FP4 Training via Transposition-Invariant Block Quantization | 2026 |
Background. FP4 training quantizes a tensor in blocks, each block carrying its own scale. The forward and backward passes read the same weight matrix in opposite orientations, one of them transposed. Problem. With one-dimensional blocks, transposing a tensor groups different values together and so assigns different scales to the same values. The forward and backward passes then disagree and the gradient is biased. The stabilizer stacks in the two entries above work around this instability without naming it. Key idea. Quantize in two-dimensional blocks, which makes the scaling transposition-invariant. Combine it with truncation-free scaling, stochastic rounding, and MXFP8 for the query and key projections. Findings. Across dense models up to 7B parameters and a 30B mixture-of-experts trained on as many as 100B tokens, the method stays within 1.3% of the BF16 baseline on perplexity and downstream accuracy. Why it matters. It replaces a collection of empirical fixes with one stated invariant: the quantization a tensor receives must not depend on which way you read it. An invariant can be checked in a design, while a recipe can only be copied. |
| What a bit costs beyond perplexity | ||
| reliability scalingReliability Scaling Laws for Quantized Large Language Models | 2026 |
Background. A quantization method is chosen off a curve of accuracy against bits, and that accuracy is measured on clean benchmark inputs. Every entry above this one reports it that way. Problem. Clean accuracy says nothing about whether a quantized model knows when it is wrong, or whether it holds up on perturbed input. A deployment decision taken on the accuracy curve alone is taken on one axis of several. Key idea. Measure uncertainty, calibration, and robustness to character-level and word-level perturbations across 2, 3, 4, and 8 bits and six quantization methods. Findings. Predictive performance scales monotonically with total bits, but reliability peaks at 4 bits. The paper also reports that quantization improves robustness to the evaluated natural input perturbations. Why it matters. The peak moves the operating point: a moderately sized model at 4 bits can be a better reliability-per-byte choice than a smaller model at higher precision. |
| silent failuresSilent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts | 2026 |
Background. A quantization result on a reasoning workload is normally reported as accuracy at a given precision, and the usual finding is that accuracy barely moves. Problem. Accuracy counts right answers. It cannot see whether the wrong answers changed kind, or whether a right answer was reached by reasoning that does not hold. A precision change can leave the count intact and move both. Key idea. Hand classify 30,000 chain-of-thought outputs from five instruction-tuned models at three precisions into six failure types, with high annotator agreement. Findings. Accuracy moves by at most 3.1 percentage points while the failure composition changes. Under NF4, shortcut collapse rises from 44% to 78% of wrong answers in a 3B model, confidence snowballing falls to near zero, and surface-text detectors reach a best F1 of only 0.53. Why it matters. An accuracy-neutral quantization result can conceal a different mix of reasoning failures, including correct answers reached through reasoning that does not hold. |
| Rotations to help quantization | ||
| QuIP#Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks | 2024 |
Background. Post-training quantization at 4 bits was routine by 2024, and the methods that got there worked one weight at a time, rounding each entry to the nearest level of a scalar grid fitted per channel or per group. Problem. Two bits leaves four levels per weight. A few large-magnitude entries stretch the grid wide enough that the rest of the matrix rounds to almost nothing, so accuracy collapses at exactly the bit width where the memory win is largest. Key idea. Two mechanisms. First, multiply the weights by random Hadamard matrices so the distribution is incoherent and no coordinate carries the outlier. Second, quantize vectors of weights jointly against a lattice codebook rather than rounding each weight to its own grid. Findings. At 4 bits per weight and below, QuIP# outperforms the post-training quantization baselines evaluated by the paper while retaining a codebook designed for efficient inference. Why it matters. It separates two things usually bundled together: making a distribution easy to quantize, and spending a fixed bit budget well. The strong 2-bit result shows the representation exists; whether a kernel can decode it fast enough is a separate question. |
| QuaRotOutlier-Free 4-Bit Inference in Rotated LLMs | 2024 |
Background. Earlier post-training methods treated activation outliers as special cases: hold a sparse set of channels in higher precision, migrate scale between activations and weights, or search per-channel clipping ranges. Problem. Every special case buys accuracy with a branch in the kernel. A mixed-precision outlier path splits one dense matmul into pieces, and if activations and the KV cache stay at 16 bits, the memory traffic you were quantizing to cut stays as well. Key idea. Rotate the model instead. Multiplying weights and activations by randomized Hadamard matrices leaves the computed function unchanged while spreading outlier energy across every channel, so no coordinate is extreme and weights, activations, and KV cache all take one uniform 4-bit quantizer. Some rotations fold into adjacent weights, and the rest run online as a fast Hadamard transform. Findings. End-to-end 4-bit QuaRot retains 99% of Llama-2-70B's zero-shot performance and adds at most 0.47 WikiText-2 perplexity. At 6 and 8 bits, round-to-nearest quantization is lossless without calibration data. Why it matters. It reframes the outlier problem as a property of the basis rather than of the model, which replaces a list of exceptions with a change of coordinates. That is what makes an end-to-end 4-bit path possible instead of a 4-bit weight format with 16-bit everything else. Adoption. llm-compressor implements QuaRot-style transforms for checkpoints served by vLLM. |
| SpinQuantLLM quantization with learned rotations | 2024 |
Background. QuaRot showed that rotating a model by a randomized Hadamard matrix removes outliers and lets weights, activations, and the KV cache share one uniform 4-bit quantizer. That rotation is drawn at random, with no reference to the weights it will be applied to. Problem. Random draws are not equally good. Accuracy at 4 bits varies from one rotation to the next, which leaves the most consequential choice in the pipeline to chance. Key idea. Learn the rotation. Treat the rotation matrices as parameters, optimize them against a small calibration set under a constraint that keeps them orthogonal so the rotated network still computes the same function, then quantize. Findings. With 4-bit weights, activations, and KV cache, SpinQuant leaves a 2.9-point zero-shot reasoning gap on Llama-2-7B, improving on LLM-QAT by 19.1 points and SmoothQuant by 25.0 points. On Llama-3-8B, it narrows QuaRot's gap to full precision by up to 45.1%. Why it matters. It turns a structural trick into a fitted object, and it shows how much of the remaining 4-bit gap was the rotation rather than the quantizer. The cost is one short optimization step added to the quantization pipeline, far less than quantization-aware training. Adoption. Meta released quantized Llama 3.2 1B and 3B checkpoints in a SpinQuant variant alongside a QAT one, and that rotation recipe runs through ExecuTorch on phones. |
| block rotation MXFP4Block Rotation is All You Need for MXFP4 Quantization | 2025 |
Background. MXFP4 assigns one power-of-two scale to each 32-element block of a tensor. Rotation-based post-training quantization, the QuaRot and SpinQuant line, is the standard route to 4 bits, and its rotations span the whole hidden dimension. Problem. Under MXFP4 those global rotations collapse. The paper benchmarks post-training methods in the format and traces the cause to its power-of-two per-block scales rather than to rotation itself, so a transform that is exact in real arithmetic loses accuracy once the quantizer sees it. Key idea. Match the rotation granularity to the block by rotating within each 32-element block rather than across the hidden dimension. Findings. In the paper's MXFP4 benchmark, global rotation methods degrade sharply while GPTQ remains strong. Block-local rotation restores the accuracy gains across the evaluated model families. Why it matters. It corrects the assumption that SpinQuant-style rotation is format-agnostic. A quantization method's gains hold relative to a scaling granularity, so moving a recipe onto a new hardware format is a re-derivation and not a port. |
| Pruning, sparsity, and cheaper architectures | ||
| SparseGPTMassive Language Models Can Be Accurately Pruned in One-Shot | 2023 |
Background. Pruning at LLM scale inherited a workflow from smaller networks: remove weights, then retrain to recover the accuracy the removal cost. Retraining a model of that size is out of reach for almost everyone who would want to prune one. Problem. Without retraining, the mask has to be right the first time. Removing one weight changes what the remaining weights in the layer should be, so which weights to drop and how to compensate is a joint problem rather than a per-weight threshold. Key idea. Treat each layer as a reconstruction problem and solve mask selection together with the compensating weight update in one pass, using the same second-order machinery as GPTQ. The procedure can also be constrained to emit a 2:4 mask directly. Findings. SparseGPT reaches at least 50% sparsity without retraining and with minimal accuracy loss. On OPT-175B and BLOOM-176B, it reaches 60% unstructured sparsity with negligible perplexity increase and completes the pruning pass in less than 4.5 hours. Why it matters. It puts pruning and post-training quantization in one frame: both are layer-wise reconstruction under a constraint, and the same second-order approximation solves both. It also sets the reference point for asking whether unstructured sparsity buys anything on hardware that only rewards the 2:4 pattern. Adoption. llm-compressor ships SparseGPT as a one-shot pruning modifier. |
| WandaA Simple and Effective Pruning Approach for Large Language Models | 2023 |
Background. One-shot pruning arrived carrying a second-order solve and a compensating weight update in every layer, as SparseGPT does. That machinery was what made pruning without retraining work at all. Problem. The solve is expensive, and it hides how much of the result comes from the mask alone. If a score with no solve in it selects nearly the same weights, the accuracy was never attributable to the solve. Key idea. Score each weight by its magnitude times the norm of its input activation, compare scores within each output rather than across the whole layer, and change no weight afterwards. Weights times input activations, and nothing else. Findings. Across Llama and Llama-2 language benchmarks, Wanda outperforms magnitude pruning and remains competitive with reconstruction-based methods that perform intensive weight updates. Why it matters. It is the baseline any pruning method has to beat, and it fixes where the burden of proof sits: a method carrying a solve has to show the solve buys something over one multiplication per weight. It also isolates the input activation, not weight magnitude alone, as what makes a weight worth keeping. Adoption. llm-compressor offers Wanda as a pruning modifier alongside SparseGPT. |
| LLM-PrunerOn the Structural Pruning of Large Language Models | 2023 |
Background. One-shot pruning methods such as SparseGPT and Wanda zero individual weights and leave every matrix the same shape, so the saving only becomes speed on hardware with sparse tensor core support. Structured pruning removes whole heads and channels instead. Problem. In a transformer a channel is not removable on its own. Residual connections and consecutive matmuls couple it to channels in other layers, and retraining the pruned model from scratch is out of reach at this scale. Key idea. Discover the coupled structures automatically, score each group with gradient information from a small calibration set, delete the lowest-scoring groups, then recover the lost accuracy with low-rank fine-tuning rather than a full retraining run. Findings. On Llama, Vicuna, and ChatGLM, the structurally pruned models retain zero-shot classification and generation capabilities after LoRA recovery. That recovery takes three hours and 50,000 examples in the reported setup. Why it matters. Structural pruning is the kind that shrinks the GEMM shapes themselves, so the speedup arrives on any GPU with ordinary dense kernels. It also gives the shape of the prune-then-recover pipeline that later model families use. |
| ProxSparseRegularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs | 2025 |
Background. Semi-structured 2:4 sparsity is the pattern NVIDIA sparse tensor cores accelerate, and the masks that reach them today come from one-shot criteria that score weights layer by layer, SparseGPT and Wanda among them. Problem. Choosing a mask is a discrete decision, so it cannot be optimized directly. One-shot methods substitute a local heuristic score for that search, and each layer commits to its own mask without regard for the rest of the network. Key idea. Learn the mask instead. A regularized objective turns non-differentiable mask selection into a smooth search over which weights to drop, and once the mask is fixed no weight update is applied to compensate for what left. Findings. Across seven evaluated models, ProxSparse consistently outperforms the semi-structured mask-selection baselines without updating the surviving weights after mask selection. Why it matters. It occupies the learning-based end of the mask-selection design space, evaluated across seven models, and it separates the two things one-shot pruning conflates: which weights to remove, and how to correct the ones that remain. |
| beyond FLOPsBenchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy | 2026 |
Background. Pruning papers report how much of the model they removed and how far perplexity moved, and they read the FLOP reduction as the speedup. The methods differ in what they remove: depth, width, or individual weights. Problem. FLOP reduction is not latency. Whether a removal helps depends on which dimension of the matrix multiply it shrinks and on how an implementation handles the resulting shapes, so speedups reported by different papers do not compare. Key idea. Reorganize pruning methods by which GEMM dimension, M, N, or K, each one actually shrinks, then measure realized speedup rather than FLOP reduction, with every method run under one implementation-consistent harness. Findings. Static depth pruning provides the strongest Pareto-optimal baseline at low quality loss. As the permitted quality loss grows, the measured frontier shifts first to dynamic depth pruning and then to static width pruning. Why it matters. The measured ranking does not follow FLOP reduction alone, so the deployment choice should follow realized latency at the acceptable accuracy target. |
GPU performance depends on arithmetic intensity, memory traffic, launch overhead, and communication. Read the roofline group first, then the attention, fusion, and compiler groups for ways to change those costs. The runtime group connects the kernels to CUDA graphs and PyTorch compilation. Other readings retain the broader serving topics from the previous lecture plan; structured decoding follows in the November 30 agent-systems meeting.
| Paper | Year | Why read it |
|---|---|---|
| The performance model | ||
| rooflineAn Insightful Visual Performance Model for Multicore Architectures | 2009 |
Background. By 2009 multicore processors had more arithmetic throughput than their memory systems could feed, and tuning was done per machine and per kernel with no shared way to state what the ceiling was. Problem. A machine's peak FLOP/s says nothing about the ceiling for one kernel. Without knowing whether memory traffic or arithmetic is the binding constraint, you can spend weeks optimizing the side that was never the limit. Key idea. Plot attainable performance against operational intensity, the FLOPs a kernel does per byte it moves from memory. The machine contributes two ceilings, a bandwidth slope and a flat compute limit, and which one a kernel sits under names what to fix. Why it matters. It gives you the units for the rest of the course, and it is the right first step in Assignment 5. Most later results in this field are a claim about raising one of the two ceilings or about moving a kernel along the intensity axis. Adoption. Nsight Compute and Intel Advisor both ship a roofline chart as a built-in profiler view, and NERSC's Empirical Roofline Toolkit measures the ceilings a real machine reaches. |
| making DL go brrrMaking Deep Learning Go Brrrr From First Principles | 2022 |
Background. The roofline divides a kernel's time between arithmetic and memory traffic. A deep learning program is a Python process issuing thousands of small operators, so it spends time in a third place: getting work onto the GPU at all. Problem. Such code often sits in neither roofline regime. Python dispatch and per-kernel launch cost can leave the GPU idle between operators, and no amount of tiling or fusion touches that time. Key idea. Diagnose first, against three regimes rather than two: compute-bound, bandwidth-bound, and overhead-bound. Each has its own test and its own fix, so decide which one you are fighting before you touch anything. Why it matters. It is the roofline in the form these workloads need, and it promotes overhead from measurement noise to a regime with a name. It also explains why fusion and graph capture are two different tools rather than two speedups. |
| FlashAttention needs your attention | ||
| FlashAttentionFast and Memory-Efficient Exact Attention with IO-Awareness | 2022 |
Background. Attention over a sequence of length N forms an N×N matrix of scores, softmaxes it, and multiplies it by the values. A standard implementation writes that matrix to HBM and reads it back for each of those steps. Problem. The matrix grows quadratically in sequence length, and the traffic to HBM, not the arithmetic, sets the runtime. Memory also caps the sequence length you can train on, so the context limit came from an implementation choice rather than from the model. Key idea. Tile the computation so a block of queries meets a block of keys in SRAM, accumulate the softmax with a running normalizer, and never write the N×N score matrix to HBM at all. The output is exact attention, not an approximation. Findings. Against the paper's baselines, FlashAttention delivers a 15% wall-clock speedup on BERT-large over the MLPerf 1.1 record, a 3× speedup on GPT-2 at length 1,000, and a 2.4× speedup on Long Range Arena tasks at 1,000-4,000. Its block-sparse extension reaches 63.1% accuracy on Path-256 at length 64,000. Why it matters. It is the clearest demonstration in the course that IO, not FLOPs, is the budget. An algorithm with identical arithmetic and identical output runs faster because it moved less data, which is the lesson every later kernel in this section applies. Adoption. PyTorch's scaled_dot_product_attention dispatches to a FlashAttention kernel, and vLLM, SGLang, TensorRT-LLM, and Hugging Face Transformers all call one; xformers and FlashInfer ship variants. |
| FlashAttention-2Faster Attention with Better Parallelism and Work Partitioning | 2023 |
Background. FlashAttention already keeps the score matrix out of HBM, so what remains is internal to the kernel: how the work is split across thread blocks and warps, and how much of the arithmetic runs somewhere other than the tensor cores. Problem. The original kernel parallelizes over batch and heads only, which leaves SMs idle when the batch is small and the sequence is long. It also does its rescaling on the general-purpose units, whose throughput per operation sits far below the tensor cores'. Key idea. Better work partitioning and fewer non-matmul FLOPs. Parallelize over the sequence dimension as well, split each tile's work across warps so they no longer exchange partial results through shared memory, and defer the softmax division so fewer rescaling operations run per tile. Findings. On A100, FlashAttention-2 runs about 2× faster than FlashAttention and reaches 50-73% of theoretical peak throughput. End-to-end GPT-style training reaches 225 TFLOP/s per A100, or 72% model FLOPs utilization. Why it matters. It shows that removing the memory bottleneck is only the first step. Once IO is handled, the remaining wins come from where work is placed: across blocks, across warps, and on which functional unit. That is the vocabulary the rest of this section uses. Adoption. PyTorch's scaled_dot_product_attention, Hugging Face Transformers' flash_attention_2 implementation, Megatron-LM, DeepSpeed, and vLLM's prefill implement FlashAttention-2. |
| FlashAttention-3Fast and Accurate Attention with Asynchrony and Low-precision | 2024 |
Background. Hopper changes the execution model the earlier kernels assume. It adds a tensor memory accelerator for asynchronous copies, warpgroup-wide matmul instructions that issue asynchronously, and FP8 tensor cores. Problem. A kernel written for synchronous loads and per-warp matmuls cannot keep those units busy, because the copy, the matmul, and the softmax run in sequence when the hardware could run them at once. FP8 adds a second difficulty, since attention's exponentials and outliers do not survive naive quantization. Key idea. The same algorithm rewritten for a new memory and execution model. Warps specialize into producers that issue asynchronous copies and consumers that run the matmuls, the softmax of one tile overlaps the matmul of the next, and FP8 is applied with block quantization and incoherent processing so accuracy holds. Findings. On H100, the FP16 kernel is 1.5-2.0× faster than FlashAttention-2 and reaches 740 TFLOP/s, or 75% utilization. The FP8 path approaches 1.2 PFLOP/s and has 2.6× lower numerical error than the paper's baseline FP8 attention. Why it matters. It is the case study in what a hardware generation costs a kernel author. The mathematics did not change; the code did, because the units that mattered became asynchronous. Expect that cost at every new architecture, and see FlashAttention-4 for the Blackwell instance. Adoption. vLLM builds and installs a dedicated FlashAttention-3 component for its Hopper attention backend. |
| FlashAttention-4Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling | 2026 |
Background. FlashAttention-3 is tuned for Hopper, where the balance among tensor-core throughput, SRAM bandwidth, and the unit that computes exponentials held a particular shape. Blackwell breaks that balance: tensor-core throughput grew and the other two did not. Problem. When one unit gets faster and its neighbors do not, the bottleneck moves. Attention's softmax needs exponentials and needs tiles resident in SRAM, so a pipeline balanced for Hopper spends its Blackwell time waiting on the units that did not scale. Key idea. Co-design the algorithm and the kernel pipeline for the new ratios instead of porting the Hopper schedule. The subtitle names both halves, and the point is that they move together: a pipeline balanced for one set of unit throughputs is wrong for another. Findings. On B200 with BF16, FlashAttention-4 reaches 1,613 TFLOP/s, or 71% utilization, and is up to 1.3× faster than cuDNN 9.13 and 2.7× faster than Triton. Its CuTe-DSL implementation compiles 20-30× faster than the reported C++ template path. Why it matters. A concrete lesson in redesigning a kernel when hardware scales asymmetrically. Peak FLOP/s is the figure vendors advertise and the least useful one here; what sets the rewrite is which units failed to keep up with it. |
| B200 attention kernelB200 Attention Kernel from Scratch to Near-SOTA in 60 Diagrams (Iaroslav Elistratov) | 2026 |
Background. The FlashAttention papers report what a finished kernel achieves, and the asymmetric scaling FlashAttention-4 names is the reason its Blackwell schedule differs from the Hopper one. On B200 the tensor cores roughly doubled while the unit that computes exponentials did not move, so softmax costs about as many cycles as the matrix multiplies despite far fewer floating-point operations. Problem. A paper reporting one number for a co-designed kernel does not say which of its choices bought the speedup, and step-by-step walkthroughs on this hardware existed for matrix multiplication rather than for attention. A reader who wants to write such a kernel, or to judge which optimization to try on a new shape, has no ordered account of what each one is worth. Key idea. Build the kernel one optimization at a time and measure every step. Fourteen versions of a dense non-causal BF16 kernel at head dimension 128, each a single change from the one before it, take the same tile decomposition from a naive baseline to near the reference implementation, with the Blackwell mechanisms introduced where they are first needed: tcgen05 tensor-core instructions and the tensor memory that holds their accumulators, swizzled shared-memory layouts, asynchronous copies driven by barriers, warp specialization, and load and compute pipelines that overlap the softmax with the matrix multiplies. Findings. Every version is scored as a fraction of stock FlashAttention-4 on the same B200, over the 4,000, 8,000, and 16,000-token shapes the FlashAttention-4 paper uses, and the ladder runs from 14.2% at the baseline to 94.4% at the end. The large individual gains come from shared-memory swizzling (15% to 26%), a second query tile per block sharing the same key and value loads (27% to 41%), the load pipeline (49%), the compute pipeline (58%), and two steps that attack the exponential unit directly, calling the hardware instruction rather than a wrapper (64%) and emulating a quarter of the exponentials in three chained fused multiply-adds to relieve pressure on that unit (70%). Two steps, moving the probability tile into tensor memory and the first warp specialization, bought almost nothing on their own and were kept because later steps depend on them. Why it matters. It is the roofline lesson at instruction granularity, and the binding resource is not the one the hardware is sold on: the specialized unit that failed to scale sets the runtime, and six of the fourteen steps exist to hide it. It also shows that an optimization ladder does not decompose into independent wins, since a step worth nothing measured alone is what makes the next one possible. For anyone writing a kernel in this course, it is the closest thing to a worked example of the method, including the discipline of scoring each attempt against a fixed reference on the same GPU. |
| Attention kernels inside a serving engine | ||
| FlashInferEfficient and Customizable Attention Engine for LLM Inference Serving | 2025 |
Background. A serving engine never hands attention a clean batch. Requests arrive at different lengths, share prefixes, get preempted, and their KV cache lives in pages scattered across HBM. Each engine grew its own attention kernels to match its own layout. Problem. Kernel work and scheduling work were being solved separately, and both suffered. A kernel written for one KV layout does not transfer, variants such as grouped-query attention, sliding windows, and custom masks each fork the code, and ragged batch shapes ruin the load balance a fixed tile assignment assumes. Key idea. Give the kernel layer one interface. Represent every KV layout, paged or shared-prefix, as block-sparse; generate each attention variant from a template by just-in-time compilation rather than writing it by hand; and assign tiles with a load balancer that absorbs varying sequence lengths while staying compatible with CUDA graphs. Findings. Against compiler backends in the paper's serving benchmark, FlashInfer reduces inter-token latency by 29-69%. It cuts latency by 28-30% for long-context inference and speeds up serving with parallel generation by 13-17%. Why it matters. This is where the kernel layer and the scheduler layer meet. The lesson is that peak kernel throughput on a uniform batch is the wrong target for serving; what you get to keep is the throughput that survives a ragged batch and a paged cache. Adoption. FlashInfer is the default attention backend in SGLang and a supported backend in vLLM and MLC-LLM. |
| ragged paged attentionA High-Performance and Flexible LLM Inference Kernel for TPU | 2026 |
Background. Paged attention keeps the KV cache in fixed-size blocks reached through a block table, so a batch of requests with different lengths reads a ragged set of blocks rather than one dense tensor. The kernels that do this are CUDA work. Problem. A TPU has no warps and a different memory system, so the CUDA kernel cannot be ported line by line. That leaves the question of which parts of a paged-attention kernel are the algorithm and which parts are CUDA-specific folklore. Key idea. Implement paged attention natively for TPU: the same block-table KV layout, a ragged batch so sequences of different lengths are served in one kernel rather than padded to a common length, and a tiling strategy chosen for the TPU memory system. Findings. With Llama 3 8B on TPU7x, the kernel reaches up to 86% memory-bandwidth utilization during decode and 73% model FLOPs utilization during prefill. Why it matters. It is a forcing function for separating an algorithm from its implementation. If paged attention survives a move to hardware without warps, the block table and the tile-at-a-time traversal are the portable content and the rest was tuning for one vendor. Adoption. vLLM's TPU backend runs a ragged paged attention kernel of this shape, written in Pallas, against the same block-table KV layout the GPU backends use. |
| Kernel fusion | ||
| FLUXFLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion | 2024 |
Background. Tensor parallelism splits a layer across devices to hold weights that do not fit on one, and every split layer ends in a collective that recombines the partial results. §7.20 counts 64 of them per decode iteration for the reference model. Problem. Those collectives sit on the dependency path, since the next layer cannot start until the current one's results are combined. At decode the messages are small, so their cost is latency rather than bandwidth, and it does not fall as the degree rises even though the compute per device does. Key idea. Decompose the communication and the computation into pieces much finer than a layer, then fuse those pieces into a single kernel, so one tile's data can be sent while other tiles are still being computed. Overlap becomes a property of the kernel rather than something a runtime arranges around it. Findings. The paper reports up to 96% of communication hidden inside a fused kernel, with 1.66× prefill and 1.30× decoding speedups over vLLM on eight GPUs, and 1.24× over Megatron-LM for training on 128. Why it matters. It is the concrete form of §7.20's third consequence, that an engine can overlap a layer's collective with arithmetic where the dependence allows, and it names the price. The overlap has to be built into the kernel, so it arrives per fused operation rather than as a setting you turn on. |
| Compilers and mega-kernels | ||
| TVMAn Automated End-to-End Optimizing Compiler for Deep Learning | 2018 |
Background. Frameworks in 2018 ran a model by dispatching its operators to vendor kernel libraries such as cuDNN. A model therefore ran fast only on hardware whose vendor had written and tuned kernels for the operators that model used. Problem. Every new operator, fusion, or accelerator needs another hand-written kernel, and the number of operator and hardware pairs grows as a product. A graph-level compiler alone cannot fix this, because the performance is decided inside the operators. Key idea. Separate what a tensor operator computes from how it executes. The compute expression is written once; a schedule then specifies tiling, loop order, vectorization, and memory placement, and code is generated per target from the pair. Search over schedules replaces hand tuning. Findings. Across low-power CPUs, mobile GPUs, server GPUs, and an FPGA accelerator, generated implementations are competitive with the paper's state-of-the-art hand-tuned libraries. The result establishes performance portability across unlike targets rather than a win on one device. Why it matters. The compute and schedule split is the organizing idea of the kernel languages that follow, Triton and TileLang included, and it names the layer at which kernel performance is decided. It also sets up the auto-tuning problem the next reading attacks. Adoption. Apache TVM ships as a real compiler stack: MLC-LLM builds its on-device LLM runtime on TVM Unity, and AWS SageMaker Neo compiles customer models with it. |
| MirageA Multi-Level Superoptimizer for Tensor Programs | 2024 |
Background. A GPU tensor program exists at three levels at once: kernels launched on the device, thread blocks within a kernel, and threads within a block. Compilers usually optimize one level at a time and take the operator graph as given. Problem. The best implementation sometimes requires changing all three levels together, for instance replacing a sequence of operators with an equivalent one that tiles differently. A pass that rewrites a single level cannot reach those programs, because the win only appears once every level changes. Key idea. Search for the program rather than optimizing the given one. Candidate implementations are generated jointly across the kernel, thread-block, and thread levels, and each candidate is checked for equivalence with the original before it is accepted. Findings. Across the paper's heavily optimized DNN workloads, Mirage improves performance by as much as 3.3× over existing compiler and library implementations. Why it matters. It shows where performance still hides once per-operator kernels are individually tuned, which is in the boundaries between them. The unit of optimization becomes a region of the model instead of an operator, and that is the premise the next reading takes to its conclusion. |
| MPKA Compiler and Runtime for Mega-Kernelizing Tensor Programs | 2025 |
Background. A model normally runs as a sequence of kernel launches, at least one per operator, with a device-wide synchronization between them. At decode batch sizes each kernel is short, so launch and synchronization time is a real share of the step. Problem. The launch boundary forbids the two things that would hide it. An operator cannot begin on the tiles that are ready while its predecessor finishes, and a collective cannot proceed while a compute kernel holds the device. Writing a megakernel by hand for one model is possible; doing it per model is not. Key idea. Compile a whole multi-GPU model into one persistent megakernel whose in-kernel scheduler dispatches SM-level tasks from a task graph. Per-operator launch boundaries disappear, which makes cross-operator pipelining and compute and communication overlap expressible inside a single kernel. Findings. Across the paper's multi-GPU inference workloads, MPK delivers an end-to-end inference speedup of up to 1.7× against kernel-per-operator serving systems. Why it matters. It relocates scheduling from the CUDA runtime into the kernel, and it names the price: the runtime must carve SMs into workers and schedulers and size the task graph to the physical SM count, so the compiled artifact is tied to the GPU it was built for. |
| The runtime between the scheduler and the kernel | ||
| CUDA graphsCUDA C++ Programming Guide, Graphs | 2026 |
Background. Work reaches a GPU one launch at a time, and each launch costs host time whether or not the sequence repeats. A decode step issues nearly the same sequence of kernels on every token it produces. Problem. Repeating a known sequence still pays to describe it. The ordinary launch API offers no way to tell the driver that this iteration's work has the same shape as the last, so the description is rebuilt and resubmitted every step. At a decode step measured in single-digit milliseconds, the host can become the thing the GPU waits for. Key idea. Separate defining the work from running it. A graph records a sequence of operations and their dependencies once, either by capturing a stream or by building the nodes directly, and the instantiated graph is then launched as one submission. Why it matters. Replay requires a recorded shape, so engines capture graphs for a fixed list of batch sizes and pad to the nearest one. Shape immutability is an API constraint; avoiding pointer updates is an implementation choice. Adoption. vLLM, SGLang, and TensorRT-LLM all capture decode graphs over a fixed set of batch sizes, which is why their configuration exposes a capture list and a maximum captured batch. |
| PyTorch 2PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation | 2024 |
Background. PyTorch runs a model one tensor operation at a time, dispatching each from Python as the program executes. That is why the framework is easy to debug, and it is also why every operation is a separate launch the host has to issue. Problem. Capturing a whole-graph representation without giving up eager semantics had defeated earlier attempts. A tracer that records one execution specializes on whatever it happened to observe, and Python's data-dependent control flow, mutation, and side effects mean the recording can be silently wrong on the next input. Key idea. Intercept CPython bytecode before it runs, extract the tensor operations into a graph, and attach guards stating exactly which properties of the observed state the graph depends on. A guard failure retraces or falls back to Python, so correctness never rests on one trace being universally valid. A backend then lowers the graph, emitting Triton for GPUs. Findings. Across more than 180 real-world models on an NVIDIA A100, TorchInductor delivers geometric-mean speedups of 2.27× for inference and 1.41× for training over eager PyTorch. These are broad framework benchmarks rather than measurements of a continuously batched serving engine. Why it matters. Guards decide how often a served model recompiles. A serving engine sees a new shape when its batch or sequence length changes, so the guard set determines whether compilation is paid once at startup or repeatedly under load. The evaluation uses open-source benchmark suites rather than a served workload. |
| Optional Other readings | ||
| ServerlessLLMServerlessLLM: Low-Latency Serverless Inference for Large Language Models | 2024 |
Background. Every earlier meeting assumed the model was already resident. Loading it is a separate cost, and it is large: the weights that a decode step reads from HBM in milliseconds have to arrive from storage first. Problem. Serverless inference wants to start a model on demand and release it when idle, but a cold start that reads tens of gigabytes over a network file system takes minutes. Keeping every model warm removes the cold start and gives up the reason for being serverless. Key idea. Treat the checkpoint as a storage problem. Use a loading-optimized checkpoint format read directly into GPU memory, exploit the multi-tier storage already on a GPU server, and schedule a request to the server where the model is furthest along in loading rather than to the least-loaded one. Findings. Across the paper's microbenchmarks and serverless workload scenarios, ServerlessLLM reduces end-to-end latency by 10-200× against the evaluated serverless baselines. The range combines workloads with different models and checkpoint-locality states, so it should not be read as one universal cold-start factor. Why it matters. It converts capacity from a fixed decision into a scheduling one, which is the same move the routing meeting will apply to requests and the batching meeting made for memory. Read the locality-aware scheduling argument closely: the placement decision now depends on storage state, which is a signal none of the schedulers earlier in Part II could see. |
| SpotServeSpotServe: Serving Generative Large Language Models on Preemptible Instances | 2023 |
Background. Preemptible cloud instances cost a fraction of on-demand ones and can be reclaimed with a short warning. Serving on them means the fleet changes size and shape without asking. Problem. A preemption during generation destroys work that cannot be recomputed cheaply, because the KV cache of every in-flight sequence on that instance is gone. Restarting from scratch is unacceptable, and a fixed parallel configuration cannot survive losing an arbitrary member of its group. Key idea. Make the parallelization plan adaptive. Recompute the tensor and pipeline degrees for whatever instances remain, minimize the migration cost of moving from the old plan to the new one, and use the preemption warning to drain and checkpoint progress rather than to stop. Findings. On real spot-instance preemption traces and several LLMs, SpotServe reduces P99 latency by 2.4-9.1× against the best evaluated serving baseline. It also reduces monetary cost by 54% relative to serving only on on-demand instances. Why it matters. It is the strongest statement on this page that a serving system's resources are not a constant. The layout decision from the batching meeting becomes something the system re-solves at runtime, and the cost of being wrong is measured in the KV state that has to move, which is the same currency as prefix-cache placement and disaggregated handover. |
| From Words to WattsFrom Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference | 2023 |
Background. Inference cost is usually quoted in dollars or in latency. Energy is the physical quantity underneath both, and a datacenter's binding constraint is increasingly power delivered rather than accelerators bought. Problem. Almost nothing published states the energy of a served token on named hardware under a named load. Without that, power cannot be reasoned about as a system parameter and shows up only as an operating expense after the fact. Key idea. Benchmark inference energy directly across model sizes, batch sizes, and sharding configurations on a supercomputing cluster, and report the effect of power capping the GPUs. Findings. For LLaMA 65B on four A100 80GB GPUs with a batch of 64, reducing the power cap from 250 W to 175 W increases inference time by 6.7% on average while reducing total energy by 23.21%. A further reduction to 150 W increases time by 19.49%. Why it matters. It gives this course's arithmetic a third axis. Every earlier meeting traded latency against throughput; this one shows a setting where you trade a little of both for a large power saving. Read it first in this section, because it is the controlled benchmark the three readings after it build a constraint, a policy, and a metric on top of. |
| POLCAPOLCA: Power Oversubscription in LLM Cloud Providers | 2023 |
Background. A GPU cluster is provisioned against the power its servers could draw rather than the power they do draw. Racks, feeds, and cooling are sized from nameplate ratings, and the total is fixed years before any particular workload arrives. Problem. Power, not accelerators, is what a datacenter runs out of first. Sizing every server at its rated maximum strands delivered capacity whenever the workload draws less, and the only way to recover that capacity is to build another datacenter, which is slow enough to bound how fast a provider can grow. Key idea. Measure what the workload actually draws and sell the gap. The paper characterizes power for a range of models and configurations, separates inference from training, and proposes a framework that deploys more servers than the nameplate budget allows while monitoring draw and capping GPU power when a rack approaches its limit. Findings. In simulations driven by open-source models reproducing the measured production patterns, POLCA deploys 30% more servers in the same GPU cluster for inference with minimal performance loss. Training does not admit the same margin, because it runs closer to the rated power and its peaks arrive synchronized across the cluster. Why it matters. It supplies the constraint the rest of this section optimizes against, and it is the reason to state a serving result in joules as well as in tokens per second. If the fleet's size is set by watts delivered, then energy saved per token is capacity the provider does not have to wait years to build. It also separates inference from training as power consumers, and the two are usually quoted together. |
| DynamoLLMDynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency | 2024 |
Background. An inference cluster serving mixed traffic runs at a fixed parallelism, a fixed instance count, and a fixed clock. All three are set for the heaviest load the cluster expects to see. Problem. Traffic is not constant and requests are not alike, so a configuration sized for the peak wastes energy for most of the day. Moreover the energy-optimal configuration for a short request differs from the one for a long generation, so a single setting is wrong for part of the mix even at a fixed load. Key idea. Reconfigure the cluster at runtime along three axes at once: how many instances run, how each is parallelized, and what frequency the GPUs are clocked at, chosen from the current traffic and its request mix while holding the latency objective. Findings. At service level, DynamoLLM reduces energy by 53%, operational carbon emissions by 38%, and customer cost by 61% against the paper's static cluster configuration while continuing to meet the evaluated latency objectives. Why it matters. Read it as the energy counterpart to the routing meeting. The decision variables are the same ones earlier meetings fixed at deployment, and the objective is the one this course has otherwise ignored, which makes it the clearest demonstration that power belongs in the scheduler rather than in the facilities budget. |
| AI at Google scaleMeasuring the environmental impact of delivering AI at Google Scale | 2025 |
Background. Every energy figure this course can otherwise cite comes from a benchmark run on a handful of accelerators. A production assistant runs on a fleet that also holds machines ready but idle, spends energy on host CPUs, and pays a datacenter overhead on all of it. Problem. No published measurement had covered that boundary, and the boundary is what the answer depends on. Estimates for the same workload differ by orders of magnitude because each omits a different term, so a per-prompt energy number cannot be compared across models or across providers. Key idea. Fix the accounting boundary first, then instrument all of it. Count active accelerator power, host system energy, the capacity held idle, and datacenter overhead together, and measure production serving of the Gemini assistant against that definition rather than against a benchmark harness. Findings. The median Gemini Apps text prompt consumes 0.24 Wh of energy and 0.26 mL of water. Over one year, software efficiency work and clean-energy procurement reduced that prompt's energy by 33× and its carbon footprint by 44×. The figures are a median over one provider's own traffic on its own hardware, reported by that provider. Why it matters. Read it for the boundary rather than for the number. Idle capacity and datacenter overhead are exactly the terms a systems course leaves out, and they are the difference between a kernel measurement and what a served prompt costs. It is also the denominator the rest of this section needs, because a 53% energy reduction means something different against 0.24 Wh than against an accelerator-only figure. |
| intelligence per wattIntelligence per Watt: Measuring Intelligence Efficiency of Local AI | 2025 |
Background. Part II has assumed throughout that a request is served in a datacenter, on hardware the provider owns, by the largest model available. Two things have changed under that assumption. Local models at or below 20B active parameters now match frontier models on many tasks, and a laptop-class accelerator can run them at interactive latency. Problem. Deciding where a query should run needs one quantity that both halves of the question share. Accuracy alone does not say whether a power-constrained device can afford the model, and tokens per second does not say whether the answer was right, so capability and efficiency are reported separately and cannot be traded against each other. Key idea. Define intelligence per watt as task accuracy per unit of power, and measure it across the whole grid rather than at one point. The evaluation covers 20+ local models, 8 local and cloud accelerators, and 1M real single-turn chat and reasoning queries, recording accuracy, energy, latency, and power for every query. Findings. Local models answer 88.7% of these queries accurately, with the rate varying by domain. From 2023 to 2025 intelligence per watt improved 5.3× and the share of queries a local model could service rose from 23.2% to 71.3%. Local accelerators reach at least 1.4× lower intelligence per watt than cloud accelerators running identical models, so local serving is the less power-efficient option for the same model even where it is viable. Accuracy is a win rate against frontier models rather than a score against ground truth. Why it matters. It is the metric the rest of this section is missing. Every other energy reading here holds output quality fixed and measures power; this one puts quality and power on one axis, which is what lets you ask whether a smaller model is the efficiency win rather than a quality loss. Read it also as a placement decision one level above the routing meeting: routing chose which instance serves a request, and this chooses whether the datacenter is involved at all. |
| Defeating nondeterminism in LLM inferenceDefeating Nondeterminism in LLM Inference | 2025 |
Background. A request sent twice at temperature zero is expected to return the same tokens. It frequently does not, and the usual explanation is that floating-point addition is not associative and GPU kernels reduce in a nondeterministic order. Problem. That explanation is wrong, and believing it makes the bug unfixable. The kernels used in serving are run-to-run deterministic on their own. The nondeterminism enters somewhere else, and until you know where, reproducibility looks like a property you must give up to use a GPU. Key idea. The cause is batch invariance, or its absence. A kernel's reduction order depends on the shape of the batch it was called with, and the batch a request lands in depends on what else arrived at the same moment, so the same request computes slightly different numbers depending on its neighbours. Making the reduction order a function of the data rather than of the batch shape removes the variation. Why it matters. Read it for the diagnosis before the fix. It reframes a numerical annoyance as a serving property, and it shows that the batch, which every meeting since the batching one has treated as free to reshape, is visible in the output. The fix costs throughput, so the last section is a trade rather than a correction, and it is the trade an evaluation harness or an on-policy training run has to make. |
| CoRunCoRun: Padding is Simple and Efficient for Deterministic LLM Inference | 2026 |
Background. Batch-invariant kernels make a request's output independent of its neighbours by changing how the kernels reduce. That is a kernel-level answer, and it asks every kernel in the engine to be rewritten against it. Problem. Rewriting reductions to be batch-invariant costs performance in exactly the kernels that were tuned hardest, and the cost is paid on every request whether or not it needed determinism. Key idea. Fix the shape instead of the reduction. Pad batches to a fixed set of sizes so that a kernel is only ever called with shapes it has seen, which makes the reduction order constant without changing the kernel. Findings. Across the evaluated Qwen and DeepSeek models, CoRun preserves deterministic output while improving throughput by 15-324% over batch-invariant-kernel approaches. It reduces time to first token by 51.8% and time per output token by 48.6% on average against those approaches. Why it matters. It is the second design point for the same requirement, and it pairs with the CUDA-graphs entry above: engines already pad to captured batch sizes for replay, so part of the machinery determinism needs is present for an unrelated reason. Read the two determinism entries together and ask which layer the property belongs in, which is the recurring question of this meeting. |
| ElasticMMElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism | 2025 |
Background. A multimodal request runs an encoder over the image or audio before the language model sees anything. The encoder is a different model with a different cost structure, and it runs on the same devices. Problem. Treating the pair as one model gives them one parallelization and one batch. The encoder is compute-bound over a fixed-size input while decode is bandwidth-bound over a growing one, so any single configuration starves one of them, and an encoder pass lands in the middle of a decode batch as an interference spike of the kind the batching meeting priced for long prefills. Key idea. Separate the stages and let each carry its own parallelism, elastic in the degree, with requests routed between independently sized groups rather than passing through one fixed pipeline. Findings. On the paper's real-world multimodal datasets, ElasticMM reduces time to first token by up to 4.2× and sustains 3.2-4.5× the throughput of the evaluated serving baselines while meeting the specified service-level objectives. Why it matters. It is the same disaggregation argument that split prefill from decode, applied one stage earlier, which is the cleanest evidence that the phase-splitting idea generalizes beyond the two phases Part II studied. Read it for the cost asymmetry between the stages rather than for the specific scheduler. |
| efficient inference for large vision-language modelsEfficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects | 2026 |
Background. Vision-language serving inherits every mechanism in this course and adds an encoder, a much longer effective prompt, and a token budget the user never typed. Problem. Which of the inherited mechanisms still pay is not obvious. An image expands to hundreds or thousands of tokens, so prefill dominates in a way it does not for chat, and the KV cache, the prefix cache, and the batching policy are all sized against assumptions text traffic supplied. Key idea. Locate the bottlenecks specific to vision-language inference and organize the published techniques against them, separating what is a general serving optimization from what depends on the visual token stream. Why it matters. Use it as the map for a final project that leaves text, and as the check on how far this semester's arithmetic transfers. Its value here is the accounting of where the tokens come from, because a system tuned for typed prompts is being handed inputs that are an order of magnitude larger without anyone deciding that. |
A cache entry is a product of four dimensions, heads times layers times tokens times dimensions, and every group below attacks one of them. Multi-query and grouped-query attention share heads, latent attention compresses the per-token vector, cross-layer attention and YOCO cut how many layers hold a distinct entry, and windows and trained sparsity decide how many tokens are kept or read. Read the page against one distinction the headline factors hide. Three quantities move independently, and they are the bytes stored per token, the bytes read per generated token, and the attention arithmetic per generated token. Cross-layer attention halves the first and leaves the second exactly where it was, because a shared entry is still read separately in every layer that uses it. DeepSeek Sparse Attention raises the first by about a fifth and cuts the second 4.6× at a 128k context. Latent attention cuts the first two together and raises the third. A paper that improves one of the three tells you nothing about the other two, and most of them do not say which they mean.
The dedicated Transformers and LLM foundations page supplies optional background for the model architectures discussed here.
| Paper | Year | Why read it |
|---|---|---|
| Storing fewer keys and values | ||
| MQAFast Transformer Decoding: One Write-Head is All You Need | 2019 |
Background. Multi-head attention gives every head its own key and value projections, so incremental decoding has to store, and then re-read, one key vector and one value vector per head per layer for every token in the sequence. Problem. A decode step does very little arithmetic and reads the whole cache to do it, so generation speed is set by how many bytes of keys and values move rather than by FLOPs. The cache scales with heads, layers, batch size, and length at once. Key idea. Keep the query heads and collapse the rest, so one shared key head and one shared value head serve all of them. The cache shrinks by the number of heads, and every decode step reads that much less. Findings. On WMT14 at batch size 1,024, multi-query attention cuts incremental decoder cost from 46 to 3.8 TPUv2 microseconds per output token, while beam-search BLEU rises from 28.4 to 28.5. Why it matters. It changes the size of the problem before any serving system touches it, and it pays for the reduction in model quality rather than in engineering. Every optimization later in this meeting operates on a cache whose baseline size this one decision already set. Adoption. PaLM projects keys and values to one shared head while retaining every query head. Gemini 1.0 names multi-query attention as the efficient attention behind its 32K context, and StarCoder, Falcon-7B, and Gemma 2B each ship one key-value head. |
| GQATraining Generalized Multi-Query Transformer Models from Multi-Head Checkpoints | 2023 |
Background. Multi-query attention minimizes the KV cache, while multi-head attention preserves quality. Their different parameter shapes made the layout appear to be a choice fixed before pretraining. Problem. Multi-query attention loses quality on some tasks, but converting an existing multi-head checkpoint appeared to require pretraining again from scratch. Key idea. Use an intermediate number of key and value heads, each shared by a group of query heads. Convert a multi-head checkpoint by mean-pooling projections within each group, then repair the lossy conversion with continued training. Findings. The repair uses 5% of the original pretraining steps. The converted eight-group model scores 47.1 on the paper's seven-task average against 47.2 for multi-head attention, while reducing time per sample from 1.51 seconds to 0.28. Why it matters. Cache size becomes a dial set by the number of groups rather than a binary choice, and an existing model can move along that tradeoff. The approximation is paid during conversion; serving remains exact for the converted model. Adoption. Grouped-query attention appears in Llama 2 70B, every Llama 3 and Llama 4 release, Qwen 2.5, and Mistral models. Hugging Face Transformers, vLLM, SGLang, llama.cpp, and TensorRT-LLM support the layout. |
| MLADeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model | 2024 |
Background. Multi-query and grouped-query attention reduce the number of key and value heads but preserve the shape of each cached entry. Problem. Reducing the group count eventually costs quality, so larger savings require changing what an entry stores. Key idea. Project the hidden state into a low-dimensional latent and cache it with one small shared key carrying rotary position information. Fold the up-projections into the query and output matrices so decoding never materializes full keys and values in memory. Findings. DeepSeek-V2 caches 576 elements per token per layer, 3.56× fewer than grouped-query attention with eight groups. Its controlled comparison reports 4% of multi-head attention's cache at the larger evaluated scale; the often-quoted 93.3% reduction also includes six-bit cache quantization. Why it matters. Latent attention compresses the representation rather than the head count. Storage and read traffic fall together, but attention arithmetic rises because the folded form scores a wider latent vector. Dedicated decode kernels are therefore part of the design. Adoption. DeepSeek V2 through V3.2, Kimi K2, Kimi Linear, and LongCat-Flash use latent attention. vLLM, SGLang, TensorRT-LLM, FlashInfer, and llama.cpp provide dedicated kernels. |
| CLAReducing Transformer Key-Value Cache Size with Cross-Layer Attention | 2024 |
Background. Earlier methods reduce heads or entry dimensions while preserving a distinct cache for every layer. Problem. Model depth multiplies the cache size, and that multiplier grows as model families scale. Key idea. Give key and value projections to only some layers and let adjacent layers reuse their output. Every layer retains its own queries, attention, and output projection, so the model shares stored state rather than computation. Findings. Pairwise sharing halves the cache on top of multi-query attention. Validation perplexity worsens by 0.04 points at 1B parameters and 0.05 at 3B, although both models were trained at only 2,048-token context. Why it matters. Capacity and bandwidth are different budgets. Sharing doubles how many sequences fit, but each layer still reads the shared entry during decode, so per-step traffic does not fall. Adoption. Character.AI, Apple's on-device foundation model, Gemma 3n and Gemma 4, and Hunyuan-Large share key-value state across layers. |
| YOCOYou Only Cache Once: Decoder-Decoder Architectures for Language Models | 2024 |
Background. Cross-layer attention divides the depth multiplier by sharing state among a few neighboring layers. Problem. Sharing across the entire model is harder because late layers would attend to state computed near the input. Key idea. Split the model into a self-decoder and a cross-decoder. The lower half produces one global set of keys and values; every upper layer attends to that set. Cache size no longer scales with upper-layer depth, and prefill can stop producing cache entries halfway through the model. Findings. At one million tokens, the trained 3B model uses 12.4 GB of inference memory, 9.4× less than the grouped-query baseline with the same kernel optimizations. Results quoted for 30B and 65B are projections rather than measurements. Why it matters. YOCO removes the depth multiplier from storage, but not the sequence-length term from decode. Every cross-decoder layer still reads the global cache, so bandwidth and arithmetic continue to grow with context. Adoption. Apple's on-device foundation model uses cache-free upper layers, and Microsoft's Phi-4-mini-flash-reasoning cross-attends from later layers to one global attention layer. |
| Sparsity fixed in advance | ||
| SWAMistral 7B | 2023 |
Background. Earlier methods change what a model stores per token while still letting every query read every retained position. Problem. A fixed pattern is cheap to execute, but it may hide distant information that the current query needs. Key idea. Let each query attend only to the latest W positions and store the cache in a fixed circular buffer. Per-token work stops growing after W, while stacking layers expands the theoretical receptive field beyond one window. Findings. Mistral 7B uses a 4,096-token window in all 32 layers and reports 8× less cache at a 32k sequence. That factor is sequence length divided by window size; the paper does not evaluate quality beyond the window. Why it matters. A sliding window reduces storage, read traffic, and arithmetic together while remaining exact for a model trained with it. However, recency rather than relevance decides what survives. Adoption. Gemma 2, Gemma 3, GPT-OSS, Cohere Command A, and Ministral 8B interleave windowed and full-attention layers, preserving some global paths while bounding most layers' caches. |
| Removing the sequence-growing state | ||
| MambaLinear-Time Sequence Modeling with Selective State Spaces | 2023 |
Background. Every entry above keeps a per-token cache and makes it smaller, by sharing heads, compressing an entry, sharing layers, or reading fewer positions. A state-space layer carries a recurrent state whose size is fixed by the state and channel dimensions, so length drops out of the cache term instead of being scaled down. Earlier structured state-space models were fast but their transitions did not depend on the input, which left them unable to do content-based recall. Problem. A recurrence that ignores its input cannot decide what to keep, and that is why state-space models lost to attention on language. Making the transitions functions of the input repairs the recall, and it also removes the convolutional form the fast state-space algorithms relied on, so the modeling fix destroys the speed that motivated the family. Key idea. Let the state-space parameters be functions of the input, which the paper calls selection, then recover throughput with a hardware-aware parallel scan that keeps the expanded state in SRAM rather than materializing it in HBM. A decode step then carries one state per sequence per layer, and that state is the same size after a million tokens as after one. Findings. The paper reports 5× the generation throughput of a Transformer of comparable size, and Mamba-3B matching Transformers of twice its parameter count both in pretraining perplexity and on downstream evaluations. Why it matters. It is the one entry here that deletes the KV-cache term rather than reducing it, so batch capacity stops depending on context length and the bytes read per generated token stop growing with it. Two costs come with that. Fixed state recalls a specific earlier token less reliably than attention, and it exposes no per-token entries to address, so the serving machinery the rest of this course builds on a token-indexed cache has nothing to act on: paged allocation, prefix reuse, and eviction all name entries the layer does not have. Adoption. Mamba blocks ship in Hugging Face Transformers and vLLM. Jamba, Codestral Mamba, Falcon-Mamba, and NVIDIA's Nemotron-H all interleave state-space layers with attention layers rather than dropping attention entirely, which keeps a bounded number of layers that can still be addressed per token. |
| Kimi LinearAn Expressive, Efficient Attention Architecture | 2025 |
Background. Every full-attention layer retains per-token state even after heads, dimensions, layers, or windows are reduced. Problem. Compression cannot eliminate growth with context; pure linear attention can, but it recalls specific earlier tokens less reliably. Key idea. Combine mostly linear-attention layers, whose recurrent state has fixed size, with periodic latent-attention layers for exact lookup. Cache size follows the number of full-attention layers rather than total depth. Findings. The 48B model uses seven latent-attention layers among 27 and reports up to 75% less cache than its full-attention counterpart. At batch size one and one million tokens, decode is 2.2× faster; the often-quoted 6× assumes a larger batch enabled by the freed memory. Why it matters. Most layers replace sequence-growing cache with fixed recurrent state. The tradeoff moves into pretraining, where the ratio of linear to full layers is fixed rather than tunable per deployment. |
| Sparsity built into the model | ||
| NSANative Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention | 2025 |
Background. Many sparse-attention methods drop positions at inference from a model trained with full attention. Problem. The model then sees an unfamiliar attention pattern, and scattered token-level sparsity may save FLOPs without reducing memory transactions. Key idea. Pretrain three parallel attention branches: compressed block summaries, selected raw blocks, and a sliding window. Independent sigmoid gates combine their outputs, and blockwise selection preserves contiguous memory access. Findings. At 64k tokens, the kernel is 9.0× faster forward and 6.0× faster backward than a Triton FlashAttention-2 implementation. The separate 11.6× figure is an expected ratio from memory-access volume, not a measured speedup. Why it matters. Sparse arithmetic and reduced memory traffic are separate achievements. Because the selected blocks change by query, the full cache remains stored; the design reduces bytes read and attention arithmetic, not capacity. |
| MoBAMixture of Block Attention for Long-Context LLMs | 2025 |
Background. Mixture-of-experts layers route each token selectively, while attention usually reads every key or follows a fixed sparse pattern. Problem. A fixed pattern may discard useful long-range dependencies, while full attention grows quadratically with sequence length. Key idea. Treat context blocks as experts. A parameter-free gate scores each block from the query and its mean-pooled keys, then routes attention to the top blocks while always retaining the current block for causality. Findings. At 95.3% sparsity, the attention layer of an 8B model is up to 6.5× faster than FlashAttention at one million tokens and 16× faster at ten million, with validation loss within 0.001 of full attention. Why it matters. Conditional computation applies to context rather than parameters. Because routing adds no parameters, a layer can alternate between sparse and full attention during training or across serving phases. Adoption. Independent implementations include a CUDA kernel from MIT and NVIDIA and support in flash-linear-attention. |
| DeepSeek-V3.2Pushing the Frontier of Open Large Language Models | 2025 |
Background. Native sparse-attention methods require pretraining, which is impractical for an existing frontier checkpoint. Problem. Reading fewer past tokens can reduce long-context cost, but it may also damage the reasoning and tool use that need the context. Key idea. Add a lightweight indexer to the dense checkpoint. It scores every past position from one shared key per token, and attention reads only the top 2,048 latent entries. First train the indexer to imitate dense attention with the model frozen, then continue training the full model. Findings. At 128k tokens, reported serving cost falls about 3.5× for prefill and 9× for decode while MMLU-Pro remains 85.0. However, the indexer still scans the full context on every step. Why it matters. Reading less is not storing less. The full latent cache remains and the indexer adds state, so stored bytes rise about one fifth even as read traffic at 128k falls 4.6×. The indexer's full scan also caps the asymptotic saving. |
| DeepSeek-V4Towards Highly Efficient Million-Token Context Intelligence | 2026 |
Background. Token selection reduces what attention reads but still stores one cache entry per token. Problem. At one million tokens, both cache capacity and the selector's full-context scan remain expensive. Key idea. Compress the token axis before selection. One layer type forms learned summaries at one quarter of the original length and applies top-k selection; another pools every 128 positions and attends densely. Both retain an uncompressed 128-token local window. Findings. At one million tokens, the larger model uses 27% of DeepSeek-V3.2's single-token inference FLOPs and 10% of its cache, despite activating more parameters. It scores 83.5 on million-token MRCR. Why it matters. This method shrinks the number of entries rather than their width or number of copies. Compression reduces storage, read traffic, and arithmetic together; selection then reduces only the last two. Adoption. vLLM, SGLang, and the Ascend and Cambricon ports of vLLM implement the mechanism. |
| sparse frontierThe Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs | 2025 |
Background. Training-free sparse-attention methods report results under different sequence lengths, model sizes, tasks, and sparsity levels. Problem. Those results do not answer which method and sparsity level an operator should use for a particular workload. Key idea. Evaluate six training-free methods under one protocol while varying sequence length, model size, and sparsity together. The set spans both prefill and decode. Findings. Tolerance depends on the task, phase, model size, and length. A 7B model's usable decode compression falls from 12× at 16k tokens to 5× at 128k, while 32B and 72B models remain near 17×. Why it matters. One fixed sparsity setting cannot serve every workload. A sparse-attention speedup is not portable unless it names the task, sequence length, model size, and phase. Aggregate averages hide this operating regime. |
Routing names two decisions that look similar at an API boundary but optimize different outcomes. A load balancer chooses among equivalent replicas, trading queue length against cache affinity and migration cost. A model router chooses among non-equivalent models, trading response quality against latency or price. Read the two groups side by side: both predict the cost of a destination before sending a request, but only the second is allowed to change the answer the user receives.
| Paper | Year | Why read it |
|---|---|---|
| Load balancing | ||
| PrebleEfficient Distributed Prompt Scheduling for LLM Serving | 2024 |
Background. Prefix caching lets an instance skip prefill work for a prompt sharing a prefix with one it served earlier. On a single instance that is a cache replacement problem. Problem. Across a cluster the reuse depends on where the request lands. A load-balanced router scatters requests that share a prefix over many instances, so each one recomputes the same prefix and the cache that would have paid for itself is never hit. Key idea. Schedule and place requests with prefix reuse as an explicit objective, so a request carrying a long shared prefix reaches the instance that already holds it. Cache reuse becomes a scheduling and placement question rather than a per-instance cache policy. Findings. Across five workloads on clusters of two to eight GPUs, Preble lowers average request latency by 1.5-14.5× and p99 latency by 2-10× against distributed SGLang. Why it matters. It makes prefix locality a first-class input to routing, and it sets up the tension the rest of this meeting works on. Affinity concentrates traffic exactly where the popular prefixes live. |
| LlumnixDynamic Scheduling for Large Language Model Serving | 2024 |
Background. A serving cluster picks an instance for each request when the request arrives, and the request stays on that instance until it finishes generating. Problem. Output lengths are unknown at arrival, so a placement that was balanced when it was made goes stale. Memory fills unevenly, requests get preempted on one instance while another has room, and a one-shot scheduler can no longer touch either. Key idea. Migrate running requests between instances, KV cache included, so a placement can be revised while the request is still generating. Scheduling becomes a continuous decision rather than a single admission-time choice. Findings. Against the evaluated serving systems, Llumnix improves tail latency by an order of magnitude, accelerates high-priority requests by up to 1.5×, and reduces cost by up to 36% at similar tail latency. Why it matters. It puts LLM serving on the same ground as virtual-machine live migration, and makes the cost of moving state the quantity to reason about. Every later routing paper either pays that cost or explains why it only reroutes queued requests. |
| DualMapEnabling Both Cache Affinity and Load Balancing for Distributed LLM Serving | 2026 |
Background. Prefix-aware routers send each request to the instance that already holds its prefix, which is what makes distributed prefix caching pay for itself. Problem. Affinity and balance pull in opposite directions. Pinning a prefix to one instance sends every request carrying that prefix there, and the popular prefixes are the ones carrying the most traffic, so the best cache hit rate produces the worst load skew. Key idea. Hash each request to two candidate instances using independent prompt hashes, then choose between them by live load. Power-of-two-choices adapted to prefix locality, with a fallback to load-aware placement once TTFT breaches the SLO. Findings. On real workloads, DualMap supports up to 2.25× more requests than prior schedulers under the same time-to-first-token SLO. Why it matters. It quantifies the cache-affinity versus load-balance conflict rather than asserting it, and shows what a second candidate buys: most of the balance back, at the cost of one more place each prefix has to be cached. |
| LMetricSimple Is Better: Multiplication May Be All You Need for LLM Request Scheduling | 2026 |
Background. Prefix-aware routing tries to send a request to an instance that already caches its prefix, while load-aware routing tries to keep work balanced. Existing schedulers combine the two signals. Problem. A linear combination requires workload-specific weights, and a simulator requires an accurate model of the hardware and traffic. Both add tuning cost and can still choose a poor operating point. Key idea. Multiply two indicators: the number of new prefill tokens if the request is routed to an instance, and that instance's current batch size. The comparison cancels the hyperparameters that a linear combination would expose, so the score needs no workload-specific tuning. Findings. On real chatbot and coding-agent workloads, LMetric reduces time to first token by 92% against vLLM-v1 and 39% against a production scheduler. Time per output token falls by 24% and 51%, respectively. The paper also derives detectable conditions under which multiplication can fail. Why it matters. LMetric replaces a tuned tradeoff with a score whose two factors retain clear units. The derived failure conditions make its simplicity auditable rather than heuristic. |
| SMetricRethink LLM Scheduling for Serving Agents with Balanced Session-centric Schedulingalso Nov 2 | 2026 |
Background. Every router above decides one request at a time. DualMap chooses between two candidates per request and CacheRoute plans per prefix key, but the unit being placed is still the individual request. Problem. An agent does not send individual requests, it sends a session, and the requests within a session share almost all of their prefix. A per-request rule therefore re-derives the same affinity decision on every turn, and inherits the same load skew each time it does. Key idea. Move the decision to the session. Route a session's first request purely for load balance, then route its follow-ups cache-aware to whichever instance the first one landed on. Balance is chosen once, at the moment nothing is cached yet and the choice costs no reuse, and reuse is then collected on every turn after. Findings. KV reuse exceeds 80% of request tokens in a production trace from BAILIAN, against 54-62% in chat, which is what makes a single placement per session sufficient. Measured against state-of-the-art schedulers, cluster throughput rises 10-16% under prefill-decode colocation with a global store, and prefill throughput rises 2-34% under disaggregation. Why it matters. Routing granularity can remove the affinity-balance tradeoff: choose balance before state exists, then preserve affinity. The result depends on sessions being long and prefix-stable. |
| Model routing | ||
| RouteLLMLearning to Route LLMs with Preference Dataalso Oct 7 | 2024 |
Background. Instance routers choose where one model runs. A deployment with cheap and expensive models must instead choose which model answers. Problem. Most requests do not need the strong model, but their quality gain is unknown before either model answers. A fixed traffic split cannot follow a changing budget. Key idea. Learn the gain from human preference data. Train router models on pairwise preferences between model outputs to predict whether the weak model's answer would be preferred for a given prompt, then expose one threshold that trades quality against cost, so the router traces a frontier rather than sitting at a single operating point. Findings. On the evaluated benchmarks, learned routers sometimes cut cost by more than 2× against always using the strong model with no measured quality loss. Performance largely transfers to replacement model pairs. Why it matters. The quality-cost frontier becomes the design object, and existing preference data becomes a training signal for allocating model cost. |
| Hybrid LLMCost-Efficient and Quality-Aware Query Routing | 2024 |
Background. A small model that fits on a low-cost device, an edge device included, is far cheaper to run than a large model behind a cloud API, and it answers many requests about as well. Problem. It does not answer all of them as well, and the deployment cannot tell in advance which ones those are. Routing on the prompt alone therefore means predicting how hard a request is. Moreover, how much difficulty is tolerable is not a fixed quantity: it depends on how much quality the deployment is willing to give up, which differs between scenarios and can change after the router is trained. Key idea. Train a router to predict query difficulty and send the easy requests to the small model, with the desired quality level as an explicit second input to the routing rule. That level is tunable at test time, so one trained router covers the whole range of quality-cost operating points instead of one. Findings. The approach makes up to 40% fewer calls to the large model with no drop in response quality, measured against routing every request to the large model. Why it matters. It separates the two quantities RouteLLM folds into a single threshold: a difficulty estimate, which is a property of the request, and a quality target, which is a property of the deployment. Separating them is what lets the same router serve a strict deployment and a thrifty one. |
| AutoMixAutomatically Mixing Language Models | 2023 |
Background. Both routers above commit before any model runs, using only the prompt as evidence. Problem. The prompt is weak evidence. Whether the small model will get a request right depends on the answer it would actually produce, and that evidence is nearly free, because the small model is the cheap one and running it first costs little. Key idea. Let the small model answer, have it verify its own answer with a few-shot self-verification prompt, and escalate to a larger model only when the verification is unconvincing. Self-verification is noisy, so the escalation decision is posed as a POMDP whose hidden state is the correctness of the answer in hand and whose observation is the confidence score, which is the natural formulation once routing is a sequence of decisions rather than one. Findings. Across five language models and five hard datasets, AutoMix consistently beats strong cascading baselines and reduces computational cost by more than 50% relative to the baselines whose performance it matches. Why it matters. It moves routing from prediction to observation, making the cheap model's own attempt the feature the decision reads. The cost is that every escalated request pays for the small model's answer as well as the large model's, so the saving holds only while the escalation rate stays low, which is the quantity to watch in any cascade. |
| RouterBenchA Benchmark for Multi-LLM Routing System | 2024 |
Background. Each router above reports a quality-cost frontier, and each reports it on its own model pair, its own benchmark mix, and its own cost accounting. Problem. Those frontiers are not comparable to one another, so a new routing policy cannot be placed against the published ones. Evaluating one honestly also means knowing what every candidate model would have answered on every query, which is expensive enough to discourage measuring it at all. Key idea. Precompute the outcomes once and share them. RouterBench releases more than 405,000 inference outcomes from representative LLMs over a benchmark mix, so a routing policy is evaluated by looking up what each model would have produced rather than by calling it. The paper adds a theoretical framework for routing and a comparative analysis of existing approaches under one metric. Why it matters. Read it before designing a router. It fixes what a frontier means, supplies the non-routing reference points a router has to beat, and reduces an evaluation that would otherwise cost inference on every model to a table lookup, which is what makes iterating on a routing policy affordable. |
The nov-02 meeting asks how a serving control plane schedules an agent program rather than isolated requests. This meeting supplies the runtime and data-plane complement: overlap or survive tool stalls, isolate and checkpoint execution, retain session state, share state across branches and agents, and trace the causal path when those mechanisms interact.
Making a smaller model usable as an agent also requires reliable tool-call structure. The structured-decoding group examines how the serving engine enforces output form while the harness remains responsible for semantic correctness.
| Paper | Year | Why read it |
|---|---|---|
| Tool and environment stalls | ||
| parallelizing tool executionParallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving | 2026 |
Background. An agent turn runs in a fixed order: the model generates a tool call, the runtime executes it, the model reads the result and generates again. The two halves occupy different resources, the GPU and whatever the tool runs on. Problem. They run strictly one after the other, so the GPU idles for the length of the tool call and the tool idles for the length of the generation. A turn costs the sum of the two latencies even where the next tokens do not depend on the pending result. Key idea. Overlap tool execution with token generation instead of serializing them, so a turn's latency moves toward the larger of the two rather than their sum. Findings. Across deep-research, coding, and scientific-agent workloads, the proposed PASTE system reduces average task completion time by 43.5% and lowers observed tool latency by 1.8×. Why it matters. Every other reading in this block decides what to do with GPU state during the stall. This one attacks the stall itself, which bounds how much those policies can be worth: overlap the call with generation and there is a smaller idle window left to schedule around. |
| TimelyLLMTimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited Agents | 2026 |
Background. Physical and spoken agents generate decisions while robots, sensors, or users execute them on a much slower wall-clock, leaving intervals in which the next piece of model output can be prepared before it is consumed. Problem. A conventional serving engine optimizes token latency without knowing when an action is actually needed. It can spend scarce GPU time on output whose deadline is distant while another agent waits for a segment on its physical critical path. Key idea. Segment generation around the slower execution loop, coordinate those segments with robot and spoken-agent progress, and schedule the available slack by time utility so the engine prioritizes output according to when it becomes valuable. Findings. The MobiSys 2026 program summary reports up to 1.52× higher time utility and up to 84% less agent waiting in the evaluated settings. Why it matters. It changes the objective from producing tokens as soon as possible to producing each segment when the surrounding world can use it. That is the right latency model whenever physical I/O, rather than decoding alone, sets the pace of an agent loop. |
| MORIIdleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI | 2026 |
Background. The sessions in an agentic system are rarely all active at once. Between tool calls a session holds KV cache and does nothing, and host DRAM sits below HBM as somewhere that state could go instead. Problem. A binary busy-or-idle label is too coarse to decide what to move. Under it every stalled session looks alike, so the system cannot tell which state will be wanted back in milliseconds from which can afford a round trip across the interconnect. Key idea. Rank agent programs on a continuous idleness spectrum, put the busiest in GPU HBM and the most idle in CPU DRAM, and shift the tier boundary to match the hardware's capacity ratio rather than fixing it at a threshold. Findings. On Claude Code workloads across four GPU-model pairs, MORI delivers 20-71% higher throughput and 18-43% lower time to first token than the strongest evaluated offloading baseline. Why it matters. The tradeoff is cross-tier KV transfer cost against how well a relative idleness ranking predicts the next stall. It is the constructive counterpart to InferCept's discard-or-retain choice at an interception: the same decision, taken across all sessions at once instead of one interception at a time. |
| speculative tool callsOptimizing Agentic Language Model Inference via Speculative Tool Calls | 2025 |
Background. An agent request stops generating at every tool call, and the engine must decide what to do with the KV cache while the tool runs. Discarding it means re-prefilling the whole prompt once the result comes back. Problem. The stall is serial: the model emits the call, waits, then resumes. Nothing overlaps, and a resume whose blocks were reclaimed pays a full prefill on a prompt that grew only by a tool result. Key idea. Speculate the tool call before the model commits to it and issue it early, and force the sequence to stay resident so the round trip costs no re-prefill. A theoretical analysis says which speculation configurations pay off, and a proposed “tool cache” endpoint asks providers to expose that residency. Why it matters. It treats the price of a tool stall as a scheduling choice rather than a fixed cost, and it names its own limit: non-idempotent tools, where a wrong speculation has already had an effect. |
| AtomixTimely, Transactional Tool Use for Reliable Agentic Workflows | 2026 |
Background. Agent tool calls write to the outside world; they send mail, file tickets, move money. Frameworks retry a failed step or abandon a branch with no account of the effects already released. Problem. A retried or abandoned step can leave effects nothing undoes, and an effect released too early can be read by work that was ordered before it. A database would give you isolation; a tool call has no transaction to belong to. Key idea. Wrap tool calls in progress-aware transactions. The runtime buffers effects, seals a transaction once its read and effect footprint is complete, and settles only after per-resource frontiers prove no earlier conflicting work can arrive, compensating reversible effects on abort and gating irreversible ones before release. Findings. Fault-injection experiments show clean recovery and isolation for contending and speculative workflows, while microbenchmarks put the transaction wrapper's overhead at microsecond scale relative to tool latency. Why it matters. It exposes the tension between isolation and the side effects real tools cannot take back. With that vocabulary an agent loop becomes a distributed transaction problem, and you can ask which of its effects are compensable. |
| MCP gatewayScalable LLM Agent Tool Access in the Cloud | 2026 |
Background. An agent learns what tools it has because their schemas sit in its prompt. The Model Context Protocol standardized how those schemas arrive, so one client can mount many tool servers at once. Problem. Every mounted tool costs context window and latency before the agent does any useful work. A catalog of thousands of tools cannot be pasted into the prompt, and the agent cannot choose from tools it was never shown. Key idea. Put a gateway between agents and MCP backends, and let it handle retrieval over the tool catalog, legacy-service integration, access control, compatibility, and session-aware routing. Findings. In the production deployment, hybrid retrieval sustains 98% Top-15 recall over more than 3,000 tools, reduces tool-selection time by 8.9×, and reduces token usage by 23.8× while remaining stable as the gateway scales out. Why it matters. The argument to take from it is that the tool catalog is a serving-system concern rather than an agent-prompt concern. Once it is one, tools get indexed, cached, and routed like any other resource the serving layer owns. |
| CortexCortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching | 2026 |
Background. Tool-using agents repeatedly fetch semantically related knowledge from remote search, code, and data services, and that access can dominate the latency of the model step that consumes it. Problem. A byte- or object-level cache cannot reliably recognize reusable meaning across differently expressed requests, while keeping every retrieved result is too expensive and prefetching indiscriminately wastes both capacity and remote calls. Key idea. Cache knowledge as semantic elements, use a two-stage retrieval and judging path to decide whether an element answers the current request, and make eviction and prefetch cost-aware. Findings. In the evaluated settings, Cortex reports up to 3.6× search throughput and 20× coding-task throughput. Why it matters. It moves remote knowledge access into the agent runtime's cache hierarchy. The important design question is no longer only whether two requests share tokens, but whether a cached semantic result is useful enough to avoid another tool round trip. |
| Sandboxes, checkpointing, and OS control | ||
| FirecrackerLightweight Virtualization for Serverless Applicationsalso Sep 30 | 2020 |
Background. Serverless platforms run short, untrusted functions from many customers on shared hardware. Containers share a kernel, so the isolation boundary is the whole Linux syscall surface; a virtual machine gives a narrower boundary and a slower start. Problem. Neither end of that choice fits the workload. Container isolation is too weak for arbitrary tenant code, and a conventional VM with a general-purpose device model boots too slowly and holds too much memory to give every invocation its own. Key idea. Build a minimal virtual machine monitor on KVM that emulates only the devices a serverless guest needs, and drop the rest. What remains boots fast enough, and costs little enough memory, to run one microVM per function. Findings. With the minimal guest-kernel configuration, each microVM uses less than 5 MB of memory, boots to application code in less than 125 ms, and can be created at up to 150 microVMs per second per host. Why it matters. It is the case for why a VM boundary can still be cheap, which is the assumption every agent sandbox in this group rests on. Once virtualization is affordable per invocation, isolation stops competing with density. Adoption. Fly.io Machines use Firecracker microVMs, and the E2B agent-sandbox platform builds its isolation layer on Firecracker. |
| SpecBoxSpeculative Sandbox Scheduling for Efficient LLM Agent Serving | 2026 |
Background. An agent tool call runs in an isolated sandbox, and the sandbox has to exist before the call can run. Starting one on demand puts its start-up latency directly in the critical path of the agent loop. Problem. You do not know a sandbox is needed until the model has emitted the tool call, and by then the request is already waiting. Keeping sandboxes hot in advance instead spends memory on the ones no request asks for. Key idea. Schedule the sandbox speculatively. Start it before the model has committed to the tool call, so start-up overlaps generation instead of following it. Findings. On high-concurrency, multi-turn traces, SpecBox reduces P99 end-to-end latency by up to 2.9× against on-demand sandbox creation and peak memory by 45.9% against permanently reserved sandboxes. Why it matters. It makes sandbox start-up a scheduling decision rather than a fixed cost. Speculation buys latency with resources a wrong guess wastes, and the exchange rate depends on how well the next tool call can be predicted. |
| DeltaBoxScaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback | 2026 |
Background. A long-running agent accumulates state in its sandbox: installed packages, edited files, live processes. Serving many such agents means suspending the idle ones and bringing them back, which is what a checkpoint is for. Problem. The checkpoint mechanisms production sandboxes use write a full snapshot, so the cost scales with the size of the sandbox rather than with what changed since the last one. At that price you checkpoint rarely, and rollback stays coarse. Key idea. Make checkpoint and rollback incremental, capturing the delta since the previous checkpoint instead of the whole sandbox, and bring both down to milliseconds. Findings. On SWE-bench and reinforcement-learning microbenchmarks, DeltaBox completes checkpoints in 14 ms and rollbacks in 5 ms, allowing more search nodes within a fixed time budget. Why it matters. Checkpoint cost sets the granularity of everything built on it: speculative branches, retry after a bad tool call, packing idle sessions off a host. Cheap rollback turns a sandbox into something you fork and discard rather than something you protect. |
| TCloneLow-Latency Forking of Live GUI Environments for Computer-Use Agents | 2026 |
Background. A computer-use agent drives a live GUI environment: a running desktop with open windows, a filesystem, and process state. Trying an alternative means getting that environment back to where it was. Problem. Whole-VM snapshots are the reset mechanism today, so exploring two branches means two full copies, and rolling back means restoring the entire machine. A snapshot is durable and slow; a branch needs to be neither. Key idea. Make workspace versioning a first-class primitive. Sibling containers share memory copy-on-write and version the filesystem, so a live GUI workspace can be forked, rolled back, and selectively merged, and fast branch creation is separated from durable checkpointing. Findings. In end-to-end agent-loop measurements, TClone reduces total task latency by 1.9× against KVM and 1.5× against CRIU. Why it matters. Speculative and parallel agent branches pay off only if they share environment state. TClone also names the right split, cheap forks for exploring and slow snapshots for durability. |
| AgentCgroupUnderstanding and Controlling OS Resources of AI Agents | 2026 |
Background. A sandboxed coding agent spends much of its time not in the model but in the operating system, running the commands it emitted. Those resources are bounded at container granularity today, from user space. Problem. Container-level control does not fit a workload whose demand changes at tool-call boundaries, so a user-space controller reacts after the spike it was meant to bound. Key idea. Characterize the OS-level resource use of sandboxed agents first, then push enforcement into the kernel with eBPF and build cgroup hierarchies whose boundaries line up with tool calls rather than with containers. Note that the evaluation is preliminary. Findings. Across 144 SWE-rebench tasks and two models, OS execution accounts for 55-60% of task latency, memory limits concurrency, and peak memory demand reaches 15.4× its average. Why it matters. It puts numbers on the half of agent latency that serving papers skip, and argues the control point sits in the wrong place. If the unit of resource demand is a tool call, that is the unit a limit should attach to. |
| ActPlaneProgrammable OS-Level Policy Enforcement for Agent Harnessesalso Sep 30 | 2026 |
Background. Agent harnesses decide what an agent may do at the tool layer, in user space. A permission check runs before each tool call, and the allowlist of permitted tools is the whole policy. Problem. An agent that reaches the same effect by another path escapes that check; a shell command can write files or open sockets without passing through the tool the allowlist names. Policy intent also arrives as underspecified natural language while enforcement must act on concrete system actions. Key idea. Split declaring policy from enforcing it. The agent declares policy in an information-flow DSL and eBPF enforces it in the kernel, so indirect execution paths are covered too, and refusals come back as semantic feedback rather than opaque errors. Findings. Across empirical policies, coding tasks, and safety benchmarks, ActPlane improves compliance on indirect execution paths that tool interception cannot observe, with 1.9-8.4% runtime overhead. Why it matters. It moves the agent permission question from the harness down to the operating system, where the actions actually land. Once enforcement sits below the tool layer, the hard part becomes translating vague intent into concrete rules rather than enumerating tools. |
| Session and KV-cache lifecycle | ||
| ContinuumEfficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live | 2025 |
Background. A multi-turn agent session pauses between turns while a tool runs or a person replies, then resumes with nearly the same context. The KV cache for that context is expensive to rebuild and holds GPU memory while nothing is running. Problem. A scheduler that keeps every paused session resident runs out of memory, and one that evicts under pressure alone discards caches that were about to be reused. Neither policy knows how long a given session is worth holding. Key idea. Attach an explicit time-to-live to cached session state and schedule against it, so the question of how long a session's KV cache is worth keeping is answered by the scheduler rather than left to eviction pressure. Findings. Across SWE-bench, BFCL, and OpenHands workloads spanning four model families, Continuum improves average job-completion time by more than 8× while also increasing throughput. Why it matters. It makes session lifetime a first-class scheduling input instead of a side effect of cache replacement. That is the shift you need before you can reason about agent sessions as state with a policy rather than as requests that happen to repeat. |
| TalariaSession-Aware Serverless Serving of Hundred-Billion-Parameter LLMs | 2026 |
Background. Serverless serving scales by treating workers as interchangeable and disposable: any replica can take any request, and idle replicas go away. Hundred-billion-parameter models make each replica expensive to start. Problem. An agent session wants the opposite. Its KV cache lives on one worker, so routing the next turn elsewhere means rebuilding the context, and scaling that worker down throws the cache away. Session affinity and elasticity pull directly against each other. Key idea. Make the serverless layer session-aware rather than session-blind, and reconcile sticky KV with elastic workers at hundred-billion-parameter scale. What the paper shows is what that reconciliation costs. Findings. On 30 SWE-bench sessions totaling 960 calls across three models larger than 100 billion parameters, Talaria cuts median session-completion time from 1,000 to 189 seconds and P95 from 2,296 to 867 seconds. Why it matters. It states the conflict cleanly, which is useful even if you never deploy serverless inference. Any autoscaling decision for an agent workload is a bet about how much cached state you are willing to discard. |
| prediction-based KV managementEfficient Serving for Dynamic Agent Workflows with Prediction-based KV-Cache Management | 2026 |
Background. An agent workflow is a program. It branches, calls tools, retries, and fans out, so which context comes back next is decided by that structure rather than by the request that just arrived. Problem. Cache policies are reactive. They observe a hit or a miss after the fact, so a branch that is about to be taken looks the same as one that has been abandoned, and the cache evicts the wrong one. Key idea. Predict what a dynamic agent workflow will need next and manage the cache against that prediction, admitting, retaining, and evicting by expected future use rather than by past access. Findings. On three workflow benchmarks, PBKV delivers up to 1.85× speedup over LRU for dynamic workflows and up to 1.26× over KVFlow for a static workflow. Why it matters. It asks whether the workflow's own structure is a usable signal for cache management. If it is, the agent runtime and the cache stop being independent layers and eviction becomes a scheduling decision. |
| CacheScoutLearning Agent Execution for KV-Cache Management in Agentic Serving | 2026 |
Background. A multi-turn agent revisits some execution states and abandons others, but its serving engine usually sees only the requests that have already arrived rather than the transitions likely to come next. Problem. Static eviction and prefetch policies miss reuse when an agent's future execution is implicit, while approaches that require a declared workflow DAG or offline training do not fit agents whose paths emerge online. Key idea. Learn execution transitions online from observed agent behavior, use those predictions to guide eviction, and gate prefetch so the system moves KV state only when the expected reuse justifies the transfer. Findings. Across the evaluated workloads, CacheScout raises cache hit rate by 10-18 percentage points, reduces mean time to first token by 18-45%, reduces mean per-turn latency by 29-38%, and improves peak throughput by up to 57%. Why it matters. It occupies the middle ground between a runtime that knows the workflow and one that treats every turn as unrelated. Learned transitions can expose enough of the hidden session structure to manage cache state without requiring the application to declare its graph. |
| SYMPHONYSYMPHONY: Enabling Compute-Memory Disaggregation in LLM Serving Systems | 2026 |
Background. Disaggregated serving can place model computation and a larger memory pool on different machines, but useful KV state must arrive at compute before the request that needs it. Problem. Moving state only after a miss puts the network transfer on the critical path, and managing the remote cache without coordination can let background movement contend with latency-sensitive requests and the serving framework's own GPU memory use. Key idea. Let the serving framework issue advisory prefetch requests, manage cached state by priority, and coordinate the GPU and framework memory cooperatively so transfers and residency follow the needs of computation. Findings. In the evaluated workloads, SYMPHONY reports 2.4× lower end-to-end latency and 4× more served requests with minimal latency increase. Why it matters. It makes disaggregated memory a serving-runtime participant rather than a passive backing store. The advisory interface is the key abstraction: computation can reveal future demand without taking over placement and eviction policy. |
| AgentServeAlgorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU | 2026 |
Background. Cluster serving resolves prefill-decode interference by spreading the phases across machines. A single consumer-grade GPU has nowhere to spread them, so every phase of every request contends for one device. Problem. Agent execution is not one traffic class but three: cold prefills over long system prompts, resume prefills that append tool output to an already cached context, and short decodes that are latency-critical. Run all three through one queue and the latency-critical decodes wait behind the long prefills. Key idea. Isolate prefills from decodes, budget the resume prefills, and partition the GPU through CUDA Green Context slots so the classes run side by side without each one owning the whole device. Findings. Across the evaluated single-GPU settings, AgentServe improves time to first token by as much as 2.8× and time per output token by as much as 2.7× over state-of-the-art baselines while sustaining competitive throughput. Why it matters. It is the consumer-hardware version of the phase-isolation argument the cluster papers above make, which is a useful test of that argument. The same taxonomy of request kinds still pays when the resource being partitioned is one GPU rather than a fleet. |
| Agent memory and knowledge state | ||
| IC-CacheEfficient Large Language Model Serving via In-context Caching | 2025 |
Background. Prefix caching reuses KV state only when a new request repeats an earlier request's tokens exactly. Two requests that mean the same thing in different words share nothing, and both pay a full prefill. Problem. Chat and agent traffic is full of near-duplicates: the same question asked differently, the same task on a different input. Exact-token matching finds no reuse there, so each one runs a full forward pass on a large model. Key idea. Cache past request-response pairs, retrieve the semantically similar ones, and replay them in context as examples. A smaller model given those examples absorbs work that would otherwise need a larger one. Reuse is by meaning rather than by token identity. Findings. Across millions of realistic requests, IC-Cache raises serving throughput by 1.4-5.9× and reduces latency by 28-71% without lowering measured response quality. Why it matters. A genuinely different reuse axis from exact-token caching, and a weaker guarantee. Exact prefix reuse returns the output the model would have produced anyway; in-context reuse changes what the model sees, so the saving is paid for in output fidelity. |
| Shared state across branches and agents | ||
| TokenCakeA KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications | 2025 |
Background. A multi-agent application issues many model calls that share a system prompt, a set of tool definitions, and a common history, and they run on serving stacks tuned for independent requests. Problem. Treating each agent call as unrelated pays for the shared context again and again, and a scheduler blind to the dependencies between agents leaves the GPU idle while one agent waits on another's output. Key idea. Organize serving around the KV cache rather than the request. Shared agent state then determines retention, eviction, and scheduling order. Findings. On representative multi-agent benchmarks, TokenCake reduces end-to-end latency by more than 47.06% and improves effective GPU-memory utilization by up to 16.9% over vLLM. Why it matters. It puts the shared cache instead of the request queue at the center of scheduling, which is the shift a multi-agent workload forces on a stack designed for one prompt at a time. |
| CoAgentConcurrency Control for Multi-Agent Systems | 2026 |
Background. Several agents in one system read and write the same shared state, a file tree, a database, a task list. Databases settled this long ago with serializability, enforced either by blocking under two-phase locking or by aborting under optimistic concurrency control. Problem. Both remedies assume a participant you can cheaply stall or restart. An agent is neither. Blocking it stalls the workflow waiting on it, and aborting it discards a long chain of reasoning along with the tokens that paid for it. Key idea. Fix a serialization order up front, apply writes speculatively in place, and notify the affected agent so the model re-judges and patches its own plan. Writes that turn out to be misordered are undone by saga-style inverses registered in advance. Findings. On ten contended workloads, CoAgent remains within 5% of serial correctness while running 1.4× faster at near-serial token cost. On its bash-only target, pass rate rises from 45/71 to 63/71 while time and cost both fall. Why it matters. It rebuilds concurrency control around a participant that can be told about a conflict and asked to adapt. The repair is performed by the model, which moves correctness from something the runtime enforces to something the agent must be persuaded to re-plan. |
| governed shared memoryGoverned Shared Memory for Multi-Agent LLM Systems | 2026 |
Background. The memory systems above are built for one agent and one conversation, and they optimise what comes back into the context window. A fleet of agents sharing one persistent store inherits none of that framing, because every agent writes into the state every other agent reads. Problem. Shared memory raises questions retrieval quality cannot answer. Which agent may read a given memory, which version is current when two agents write facts that contradict each other, where a retrieved claim came from, and how long a write takes to become visible to a sibling. A larger context window addresses none of the four, so the paper names them as failure modes: unauthorized leakage, stale propagation, contradiction persistence, and provenance collapse. Key idea. Treat the store as governed operational state rather than as an index, and give it four primitives that answer one failure mode each: scoped retrieval, temporal supersession, in which a later fact marks the earlier contradicting row non-active, provenance tracking, and policy-governed propagation. Then measure a running production service against them. The evaluation drives a live multi-tenant memory service through its REST API rather than a simulator, which is what makes the negative results possible. Findings. Two primitives held. All 50 depth-four derivation chains reconstructed with the correct writer at every hop, at 291 ms per hop at the median, and propagation reached 97.5% of 120 fleet-sibling probes with no cross-fleet hit in 80, at a write-to-visible median of 0.83 s. The two failures are the more useful half. Semantic search honoured the fleet filter only partly, returning the row on 43.9% of the 164 probes it should have denied, and a direct fetch-by-id path enforced tenant scope but not sub-tenant scope until it was remediated during the study. Contradiction detection fired on 49.0% of 200 fact runs but on 100% of the 90 runs where both writes were admitted, because a synchronous near-duplicate gate rejected the contradicting write before the asynchronous detector could observe it. Why it matters. It reads the agent memory subsystem as a distributed system with an access-control policy and a write pipeline, which is the reading a fleet forces and the one a retrieval benchmark cannot produce. Note what the two failures have in common: neither is a missing primitive. One is a policy enforced on one code path and not another, and one is two correct stages composed in the wrong order, which are the defects that only appear once the thing is running. |
| Runtime observability and causal diagnosis | ||
| AgentOpsAgentOps: Enabling Observability of LLM Agentsalso Sep 30 | 2024 |
Background. An agent's execution is a sequence of model calls, tool invocations, and state changes that no existing log format was designed to hold. The tracing services and evaluation harnesses that record it each invented their own vocabulary for it. Problem. Without agreement on what to record, a trace is whatever the framework happened to emit. The same failure then looks different in two deployments, and neither trace answers the question the other was built to answer. What to instrument is the prior decision, and it has had far less attention than how. Key idea. Derive the answer from what the tools already do. A systematic mapping study over existing AgentOps tools yields a taxonomy of the artifacts, and the data attached to each, that should be traced across an agent's whole lifecycle, offered as a reference template for building the infrastructure rather than as an implementation of it. Why it matters. It is the checklist, and a checklist belongs before the mechanisms rather than after them, because a span you did not record is not recoverable later. Read it against the schema you would need to answer this meeting's own questions: which turn evicted the session, how long the sandbox took to start, and which agent in a group was waiting on which. |
| AgentSightAgentSight: System-Level Observability for AI Agents Using eBPF | 2025 |
Background. Coding agents now run as ordinary processes on developer and CI machines, issuing model calls over TLS and acting on the machine through system calls. Claude Code and Gemini CLI are the paper's examples. Problem. The two views of such an agent are separately observable and separately useless. A prompt log says what the agent meant to do and offers no evidence that it did it; a system-call trace says what happened to the machine and gives no reason why. Instrumenting the framework to join them breaks whenever the framework's API changes, which for agent frameworks is often. Key idea. Trace at the boundaries the agent cannot move. Use eBPF to intercept TLS-encrypted model traffic and recover intent, monitor kernel events to recover effect, and correlate the two streams across process boundaries in a real-time engine, with a secondary model used for the analysis. Nothing inside the agent is instrumented. Findings. The technique is framework-agnostic and incurs less than 3% performance overhead. In the paper's evaluation it detects prompt-injection attacks, identifies reasoning loops that waste resources, and reveals coordination bottlenecks in multi-agent systems that no single agent's own logs expose. Why it matters. It is the observability counterpart to this meeting's sandbox sections, because the kernel interface that constrains what an agent may touch is also the one place that can report what it did touch. The reasoning loops and coordination bottlenecks it finds are failure modes this meeting has already described from the inside, arrived at here from outside the framework, which is what makes the result portable to whichever agent stack you end up running. |
| AgentTraceAgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems | 2026 |
Background. A multi-agent workflow that fails leaves a long execution trace, and the error surfaces at whichever agent could not proceed rather than at the one that caused the trouble. Problem. Cascading effects and hidden dependencies put distance between the manifestation and the cause. Reading the trace forward from the start is expensive, and reading it backward from the error requires knowing which earlier steps the failing one actually depended on. Handing the trace to a model instead puts an inference call, with its cost and its own nondeterminism, inside the debugging loop. Key idea. Reconstruct a causal graph from the execution logs, walk backward from the point where the error appeared, and rank candidate root causes using structural and positional signals read off the graph, so that diagnosis at debugging time needs no model inference at all. Findings. Across the paper's benchmark of multi-agent failure scenarios, AgentTrace localizes root causes with sub-second latency and higher accuracy than both the heuristic and the model-based baselines it is compared against. The scenarios are constructed to reflect common deployment patterns rather than sampled from production, so read the accuracy as a comparison among the evaluated methods. Why it matters. It closes the loop this section opens. The Mystery Machine recovers a causal graph to explain where the latency went, and this recovers one to explain why the run failed, from the same raw material. Read the argument for keeping the model out of the debugging path: it is the critical-path argument this course makes about requests, applied to the operator instead. |
| Structured decoding for agent tool calls | ||
| XGrammarXGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models | 2024 |
Background. A request that must return JSON, a function call, or code can be constrained while it decodes rather than checked after it finishes. A context-free grammar expresses the constraint, and the decoder enforces it by masking the tokens the grammar forbids at each step. Problem. Enforcing a grammar means deciding, for every token in the vocabulary at every step, whether it is currently legal. Running that through the grammar's stack states at runtime costs enough to appear in the token rate, and it is per-token-per-vocabulary work that no batching trick amortizes away. Key idea. Split the vocabulary in two. Most tokens are context-independent, so their legality can be precomputed per grammar state and cached, and only the remainder need interpreting against the runtime stack. Grammar transformations shrink the context-dependent set further, and a persistent stack makes the surviving checks cheap. Findings. The paper reports up to 100× faster grammar execution than the systems it compares against, and near-zero end-to-end overhead once the grammar engine is co-designed with the inference engine so its computation overlaps GPU execution. Why it matters. Grammar masking adds per-token work proportional to the vocabulary, but precomputation and overlap can remove it from the critical path. Tool-calling traffic makes this a common serving cost. Adoption. vLLM and SGLang both ship XGrammar as a structured-output backend, so this is the mask computation running behind a JSON-schema request on either engine. |