CS2680 Modern AI Systems: Agents and System Optimizations
Readings

Sep 28 Agents design III: agent memory and reliability

Paper Year Why read it
State management and agent memory
RAGRetrieval-Augmented Generation for Knowledge-Intensive NLP Tasks 2020

Background. Before 2020, a model answered a knowledge-intensive question from whatever pretraining had written into its weights. Updating a fact meant retraining, and there was no passage to point at when the answer came out wrong.

Problem. Parametric knowledge is fixed at training time and cannot be inspected. Asked about something absent from its weights, the model produces a fluent answer anyway, and nothing in the output tells you which case you are in.

Key idea. Put a dense retriever over a Wikipedia index in front of a pretrained sequence-to-sequence generator, fetch passages for each input, and condition generation on them, training retriever and generator together while the index stays fixed. Knowledge moves out of the weights and into a store you can edit.

Findings. RAG set the reported state of the art on three open-domain question-answering tasks and generated more factual language than the paper's parametric-only sequence-to-sequence baseline.

Why it matters. It is the origin of the long, partly repeated prompts that make prefix caching worth building. It also splits serving in two: a retrieval tier with its own index and latency, and a generation tier whose input length is now set by how much you retrieved.

Adoption. Hugging Face Transformers implements the original RAG model and retriever classes, providing a maintained external implementation of the architecture.

MemGPTTowards LLMs as Operating Systems 2023

Background. The context window is a fixed budget, and any long conversation or long task overruns it. The standard answers are truncation, summarization, and retrieval bolted on outside the model.

Problem. All three discard state without asking the agent what it still needs, and none of them lets the agent decide what stays resident. The window is managed for the model rather than by it.

Key idea. Treat the context window as the top tier of a memory hierarchy and give the model explicit paging operations into the tiers below. What stays resident, what is evicted, and what is fetched back become decisions the agent makes with function calls.

Why it matters. It is the most systems-flavored agent paper on the list, and it imports the virtual memory design wholesale: a small fast tier, a larger slow one, and an explicit policy for moving data between them. Once the window is a cache, compaction and retrieval stop looking like separate features.

Common failure modes and guardrails
where agents failWhere LLM Agents Fail and How They can Learn From Failures 2025

Background. An agent failure is observed at the end of a trajectory. What is visible is a wrong answer, and the step that made it inevitable is many turns behind it.

Problem. Errors cascade rather than accumulate. One wrong decision enters the transcript and every later step reasons from it as given, so the last step that looks wrong is rarely the step that was wrong. Diagnosis from the final answer therefore identifies the symptom.

Key idea. Annotate failures at the step level instead. Sort the errors by which module produced them, covering memory, reflection, planning, action, and system operations, then build a debugger that isolates the root-cause step and feeds a correction back into a fresh attempt.

Findings. Root-cause feedback is worth more than a retry: the paper reports 24% higher all-correct accuracy and 17% higher step accuracy over the strongest baseline, and up to 26% relative improvement in task success on ALFWorld, GAIA, and WebShop.

Why it matters. It is the diagnosis order of this meeting with a dataset behind it, and it makes the case for treating the transcript as the primary artifact. Root-cause attribution is only possible if every step was recorded, which is what the instrumentation section asks you to build before you need it.

why multi-agent systems failWhy Do Multi-Agent LLM Systems Fail? 2025

Background. The preceding papers diagnose failures in one agent's trajectory. A multi-agent system adds handoffs and shared decisions, so final task success hides which interaction broke.

Problem. A multi-agent failure can stem from system design, agents misunderstanding one another, or weak task verification. Without a shared vocabulary, two systems failing for unrelated reasons look identical from the score, and fixes are aimed at the wrong component.

Key idea. Build the vocabulary from evidence rather than from a diagram. Annotate traces across popular multi-agent frameworks, derive a taxonomy from close reading of a subset, and validate it with human annotators before applying it at scale.

Findings. The taxonomy names 14 failure modes in three clusters: system design issues, inter-agent misalignment, and task verification. It is built from a careful analysis of 150 traces with high inter-annotator agreement (kappa = 0.88) and applied across a dataset of more than 1,600 annotated traces from seven frameworks.

Why it matters. It extends the step-level diagnosis above to failures between agents. The three clusters tell you whether to repair the overall design, the contract between agents, or the verifier that checks their output; the broader taxonomy below places coordination among other agent failure modes.

agent failure taxonomyBeyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents 2026

Background. Agent progress is reported as pass rates, one number per harness per suite. What is known about why agents fail is scattered across benchmark papers, taxonomies, and audits that share no vocabulary.

Problem. A pass rate tells you how often a run failed, not what broke. Two harnesses at the same score can be failing for unrelated reasons, and a fix aimed at one failure mode is credited or blamed by a number that mixes all of them.

Key idea. Synthesize 27 benchmark, taxonomy, and audit papers into six failure clusters: tool invocation, planning, long-horizon context accumulation, multi-agent coordination, safety, and measurement validity.

Findings. Failures compound nonlinearly with task length, and additional scaffolding does not consistently improve reliability.

Why it matters. It gives you names for the failures that harnesses such as SWE-agent and Claude Code build guards against, so you can argue about a mechanism instead of a leaderboard position. Its measurement critique lands on those leaderboards, SWE-bench included.

overthinkingThe Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks 2025

Background. A reasoning model in an agent loop has two ways to make progress: think longer about the state it already has, or act on the environment and read what comes back. Longer chains of thought are widely treated as strictly better.

Problem. An agent that reasons instead of acting substitutes its own model of the environment for the environment, then commits to plans built on stale observations. Aggregate pass rates hide the substitution, so the failure has no name to argue about.

Key idea. Define an overthinking score for trajectories that favor internal reasoning over interaction, then measure what it predicts.

Findings. Higher scores correlate with lower SWE-bench Verified performance across 4,018 trajectories, and selecting the lower-overthinking solution improves performance by almost 30% while cutting compute cost by 43%.

Why it matters. More thinking is not monotonically better, which breaks the assumption behind spending test-time compute freely. It also gives you a case where the cheaper policy is the more accurate one, so reasoning budget belongs in the design, not at whatever the model defaults to.