How Coding Agents Manage Context

Context and Memory Management are difference:

Context management = per-turn. “What fits in this prompt right now?” Deals with the model’s finite window this request.

Memory management = cross-turn/cross-session. “What should the agent remember for later?” Deals with persistence beyond the current request.

Context management keeps the prompt usable now; memory management keeps knowledge alive later; summarization is the bridge where they overlap. why we say summarization is the bridge? When context overflows, you summarize old turns into a compact message. That summary is simultaneously: A memory technique (preserves knowledge of past work for future turns); A context technique (frees window space now).

There are 7 real techniques in production use to manage CONTEXT:

First, OpenCode’s isOverflow (compaction.ts:30-39) is the simplest form. it counts input + cache_read + output against the usable budget. The cache_read matters because cached tokens still occupy context window space even though they’re cheap.

# token_budget.py
"""Accurate token budgeting. OpenCode uses chars/4 for pruning estimates
but real budgeting needs the actual tokenizer."""
from dataclasses import dataclass
from typing import Any
import tiktoken
@dataclass
class ModelLimit:
context: int # total context window
input: int # max input (sometimes < context)
output: int # max output tokens
@dataclass
class TokenUsage:
input: int = 0
output: int = 0
cache_read: int = 0
cache_write: int = 0
class TokenBudgeter:
"""
Tracks token usage and decides when context is full.
Port of SessionCompaction.isOverflow + accurate counting.
"""
def __init__(self, model: ModelLimit, output_token_max: int = 32_000):
self.model = model
self.output_token_max = output_token_max
# model-specific encoder; cl100k is a reasonable fallback
try:
self.encoder = tiktoken.encoding_for_model(model.api_id)
except KeyError:
self.encoder = tiktoken.get_encoding("cl100k_base")
def count(self, text: str) -> int:
return len(self.encoder.encode(text))
def count_messages(self, messages: list[dict]) -> int:
"""Count tokens for a full message list (approximation)."""
total = 0
for msg in messages:
total += 4 # role + structural overhead per message
content = msg.get("content", "")
if isinstance(content, str):
total += self.count(content)
elif isinstance(content, list):
for part in content:
if isinstance(part, dict) and "text" in part:
total += self.count(part["text"])
return total
def is_overflow(self, last_usage: TokenUsage) -> bool:
"""Direct port of OpenCode's isOverflow."""
if self.model.context == 0:
return False
count = last_usage.input + last_usage.cache_read + last_usage.output
output = min(self.model.output, self.output_token_max) or self.output_token_max
usable = self.model.input or (self.model.context - output)
return count > usable
def remaining(self, current_usage: TokenUsage) -> int:
"""How many input tokens are left before overflow."""
output = min(self.model.output, self.output_token_max) or self.output_token_max
usable = self.model.input or (self.model.context - output)
used = current_usage.input + current_usage.cache_read
return max(0, usable - used)

Second technique: Tool output pruning. OpenCode protects the most recent 40k tokens of tool output and the skill tool specifically (compaction.ts:41-44). Note the broader industry adds: Cursor and Continue do semantic pruning — they track which file reads are still “alive” (the file was edited after the read) and drop reads of files that haven’t been touched. OpenCode’s approach is simpler: pure recency + token count.

Technique 3: Conversation summarization (compaction) , this is realized by spawning a subagent with below prompts. Concretely, Claude Code: same approach, but also saves the summary to a file so it survives session restart. Aider: uses a “repo map” (tree-sitter AST summary) instead of LLM summarization — deterministic but less nuanced. LangChain ConversationSummaryBufferMemory: incremental summarization — summarizes each chunk as it ages, rather than one big summary at overflow.

COMPACTION_SYSTEM = """\
You are a helpful AI assistant tasked with summarizing conversations.
When asked to summarize, provide a detailed but concise summary.
Focus on:
- What was done
- What is currently being worked on
- Which files are being modified
- What needs to be done next
- Key user requests, constraints, or preferences that should persist
- Important technical decisions and why they were made
"""
COMPACTION_USER_PROMPT = """\
Provide a detailed prompt for continuing our conversation above. \
Focus on information that would be helpful for continuing the conversation, \
including what we did, what we're doing, which files we're working on, \
and what we're going to do next considering new session will not have \
access to our conversation."""

Technique 4: Prompt caching (the highest-leverage technique)

This is where OpenCode is genuinely well-engineered. transform.ts:164-198 shows it marks the first 2 system messages and last 2 conversation messages as cacheable. This means the stable system prompt is cached (huge savings) and the recent turn boundary is cached (so the next turn reuses most of the prefix). This matters more than summarization: caching doesn’t lose information. A 200k-token cached prefix costs ~10% of the uncached price and has zero latency on cache hit. Summarization loses detail. The industry trend is: cache aggressively, summarize only when you must.

But How prompt caching actually works (Anthropic)?

Anthropic’s cache is a prefix cache. When you mark message N with cache_control, everything from the start of the prompt up to and including message N gets cached. On the next request, if that prefix is identical, it’s served from cache at ~10% cost and near-zero latency. The constraint: Anthropic allows a maximum of 4 cache breakpoints. Note the cache concept in LLM is different than used to be, When an LLM processes a prompt, it computes KV tensors (key-value attention states) for every token — that’s the expensive part, not the text itself. Prompt caching means the provider stores those computed KV tensors and reuses them on the next request if the prefix matches.

# prompt_caching.py
"""
Prompt caching: mark stable parts of the prompt so the provider
serves them from cache instead of reprocessing.
Anthropic: cache_control breakpoints (up to 4)
OpenAI: automatic prefix caching (no markers needed, just keep prefix stable)
Google: explicit context caching API
OpenCode marks system[0:2] + final[0:2] messages as ephemeral cache points.
"""
from typing import Any
def apply_anthropic_caching(messages: list[dict]) -> list[dict]:
"""
Port of OpenCode's applyCaching for Anthropic.
Strategy: cache the system prompt + the last 2 messages.
Anthropic allows up to 4 cache breakpoints.
"""
system_msgs = [m for m in messages if m["role"] == "system"][:2]
non_system = [m for m in messages if m["role"] != "system"]
final_msgs = non_system[-2:] if len(non_system) >= 2 else non_system
cache_targets = system_msgs + final_msgs
for msg in cache_targets:
if isinstance(msg.get("content"), list):
# Anthropic expects cache_control on the last content block
if msg["content"]:
last_block = msg["content"][-1]
if isinstance(last_block, dict):
last_block["cache_control"] = {"type": "ephemeral"}
else:
# String content: wrap in a content block to add cache_control
msg["content"] = [
{"type": "text", "text": msg["content"]},
{"type": "text", "text": "", "cache_control": {"type": "ephemeral"}},
]
return messages
def apply_openai_caching(messages: list[dict], cache_key: str | None = None) -> dict:
"""
OpenAI does automatic prefix caching — no content markers needed.
Just keep the message prefix stable and pass a prompt_cache_key
for better hit rates on recurring prefixes.
"""
extra_body = {}
if cache_key:
extra_body["prompt_cache_key"] = cache_key
return extra_body
def apply_caching(messages: list[dict], provider: str) -> tuple[list[dict], dict]:
"""
Unified entry point. Returns (modified_messages, extra_kwargs).
OpenCode's transform.ts dispatches by providerID the same way.
"""
if provider == "anthropic":
return apply_anthropic_caching(messages), {}
elif provider == "openai":
return messages, apply_openai_caching(messages, cache_key="session")
elif provider == "google":
# Google's context caching is a separate API call to create a cache,
# then you reference it by ID. More complex; omitted here.
return messages, {}
else:
# OpenRouter / Bedrock pass through to the underlying provider
return apply_anthropic_caching(messages), {}

Technique 5: Sliding window truncation (what OpenCode avoids)

OpenCode deliberately does NOT do this — it uses compaction instead. But it’s the most common approach in general agent frameworks.

Technique 6: RAG context injection (file-level).

OpenCode does this implicitly via its read tool — the agent reads only what it needs. But Cursor and Continue do it at the context-building layer: before the agent even acts, they retrieve relevant code chunks.

Technique 7: Prompt compression (LLMLingua)

Not used by OpenCode. Research-grade but increasingly practical.

Here’s what a production context manager does each turn, combining all techniques:

# context_manager.py
"""
The full per-turn context management pipeline.
This is the Python equivalent of what OpenCode's prompt.ts loop does,
extended with techniques OpenCode doesn't use.
"""
from dataclasses import dataclass
from token_budget import TokenBudgeter, TokenUsage
from tool_pruning import prune_tool_outputs
from summarization import compact_conversation, filter_to_post_compaction, build_post_compaction_history
from prompt_caching import apply_caching
from rag_context import CodeChunkIndex
from prompt_compression import compress_messages
@dataclass
class ContextConfig:
enable_pruning: bool = True
enable_compaction: bool = True
enable_caching: bool = True
enable_rag: bool = False # OpenCode: False (uses tools instead)
enable_compression: bool = False # OpenCode: False
compaction_model: str = "gpt-4o-mini"
class ContextManager:
"""
Called each turn before sending to the model.
Returns the final message list that fits the context window.
"""
def __init__(
self,
budgeter: TokenBudgeter,
config: ContextConfig,
code_index: CodeChunkIndex | None = None,
):
self.budgeter = budgeter
self.config = config
self.code_index = code_index
self._compaction_count = 0
async def build_context(
self,
messages: list[dict],
system_prompt: str,
user_query: str,
last_usage: TokenUsage,
provider: str,
) -> tuple[list[dict], dict]:
"""
The main pipeline. Returns (messages_to_send, extra_kwargs).
"""
extra_kwargs: dict = {}
# 1. Check if we need compaction (token budget exceeded)
if self.config.enable_compaction and self.budgeter.is_overflow(last_usage):
messages = await self._do_compaction(messages)
# Reset usage tracking after compaction
last_usage = TokenUsage()
# 2. Filter to post-compaction history
messages = filter_to_post_compaction(messages)
# 3. Prune old tool outputs
if self.config.enable_pruning:
prune_tool_outputs(messages)
# 4. RAG: inject relevant code chunks (if enabled)
if self.config.enable_rag and self.code_index:
rag_block = self.code_index.build_context_block(user_query)
if rag_block:
system_prompt = system_prompt + "\n\n" + rag_block
# 5. Assemble final message list
final = [{"role": "system", "content": system_prompt}, *messages]
# 6. Compress older messages (if enabled)
if self.config.enable_compression:
final = await compress_messages(final)
# 7. Apply prompt caching
if self.config.enable_caching:
final, cache_kwargs = apply_caching(final, provider)
extra_kwargs.update(cache_kwargs)
# 8. Final budget check
total = self.budgeter.count_messages(final)
if total > self.budgeter.model.context:
# Emergency: hard truncate oldest non-system messages
final = self._hard_truncate(final)
return final, extra_kwargs
async def _do_compaction(self, messages: list[dict]) -> list[dict]:
"""Summarize and replace old history."""
result = await compact_conversation(
messages, model=self.config.compaction_model
)
# Keep the last few messages as-is, replace the rest with summary
recent = messages[-4:] # keep last 2 turns
self._compaction_count += 1
return build_post_compaction_history(result.summary, recent)
def _hard_truncate(self, messages: list[dict]) -> list[dict]:
"""Emergency truncation: keep system + last N messages."""
system = [m for m in messages if m["role"] == "system"]
rest = [m for m in messages if m["role"] != "system"]
# Keep removing from the front until we fit
while rest and self.budgeter.count_messages([*system, *rest]) > self.budgeter.model.context:
rest.pop(0)
return [*system, *rest]

Leave a Reply