Context cost & limits
Window size, token spend, compression, and what to do when the budget runs out.
1 answer · 12 articles · 6 definitions
Definitions in this subject
MCP server context window cost: what to cut first
Disconnect servers before you optimize them. Tool definitions are paid on every turn whether the agent calls them or not, so removing servers nobody uses is usually the largest single saving. After that, gate the remaining tools per session, defer schema loading until a tool is selected, and move occasional capabilities out of the tool list entirely.
ArticleContext window blindness: agents can't see their limits
Context window blindness: four frontier models misjudged their own context size by 43 to 84%. Why compaction drops the wrong things, and what fixes it.
ArticleContext engineering: what replaces prompt engineering
Prompt engineering has a new successor: context engineering. Learn why Karpathy and Tobi Lütke made the switch, and what it means for production AI systems.
ArticleContext pruning helps agents, until it doesn't
Context pruning helps AI agents in one regime and hurts in another. A 2026 study of models from 4B to 284B maps when to prune stale context and when not to.
ArticleContext bloat: why long-running agents break
Context bloat is when accumulated tool-call output crowds out an agent's task. Tool calls, not window size, break long-running agents. Here is the fix.
ArticleHow agents manage their own context window
Agent context management is shifting from fixed harness rules to learned, runtime decisions an agent makes about its own window. What the 2026 research shows.
ArticleAI notetakers ship the wrong artifact
AI notetakers ship transcripts, but downstream work needs decisions, drafts, or handoffs. The artifact gap is a context engineering problem, not transcription.
ArticleLong context tripled hallucinations in 35 open models
A 172-billion-token study across 35 open models found hallucination rates triple from 32K to 128K context, and exceed 10% at 200K for every model tested.
ArticleContext Rot: Why AI Performance Degrades With More Information
Research shows LLMs drop from 95% to 60% accuracy as context grows stale. Here's how context rot degrades AI performance and why bigger windows won't help.
ArticleContext budgets: how to allocate tokens for AI agents
A practical guide to context budgets for AI agents. How to allocate tokens across system prompts, tools, retrieval, history, and a buffer in production.
ArticleWhy your AI costs are a context problem
Token prices fell 280x but enterprise AI spend rose 320%. Poor context architecture drives 60-70% of total AI costs. Here is where the money actually goes.
ArticleContext compression: why less context means better AI
Context compression reduces AI agent memory usage by 26-54% while preserving task performance. Here's how it works and why bigger context windows aren't the answer.
Give your agents one place to read and write.
Put in your docs, data and notes, then connect Claude, Codex or anything else with one link.
Create a container