Learn

Context cost & limits

Window size, token spend, compression, and what to do when the budget runs out.

1 answer · 12 articles · 6 definitions

Answer

MCP server context window cost: what to cut first

Disconnect servers before you optimize them. Tool definitions are paid on every turn whether the agent calls them or not, so removing servers nobody uses is usually the largest single saving. After that, gate the remaining tools per session, defer schema loading until a tool is selected, and move occasional capabilities out of the tool list entirely.

Article

Context window blindness: agents can't see their limits

Context window blindness: four frontier models misjudged their own context size by 43 to 84%. Why compaction drops the wrong things, and what fixes it.

Article

Context engineering: what replaces prompt engineering

Prompt engineering has a new successor: context engineering. Learn why Karpathy and Tobi Lütke made the switch, and what it means for production AI systems.

Article

Context pruning helps agents, until it doesn't

Context pruning helps AI agents in one regime and hurts in another. A 2026 study of models from 4B to 284B maps when to prune stale context and when not to.

Article

Context bloat: why long-running agents break

Context bloat is when accumulated tool-call output crowds out an agent's task. Tool calls, not window size, break long-running agents. Here is the fix.

Article

How agents manage their own context window

Agent context management is shifting from fixed harness rules to learned, runtime decisions an agent makes about its own window. What the 2026 research shows.

Article

AI notetakers ship the wrong artifact

AI notetakers ship transcripts, but downstream work needs decisions, drafts, or handoffs. The artifact gap is a context engineering problem, not transcription.

Article

Long context tripled hallucinations in 35 open models

A 172-billion-token study across 35 open models found hallucination rates triple from 32K to 128K context, and exceed 10% at 200K for every model tested.

Article

Context Rot: Why AI Performance Degrades With More Information

Research shows LLMs drop from 95% to 60% accuracy as context grows stale. Here's how context rot degrades AI performance and why bigger windows won't help.

Article

Context budgets: how to allocate tokens for AI agents

A practical guide to context budgets for AI agents. How to allocate tokens across system prompts, tools, retrieval, history, and a buffer in production.

Article

Why your AI costs are a context problem

Token prices fell 280x but enterprise AI spend rose 320%. Poor context architecture drives 60-70% of total AI costs. Here is where the money actually goes.

Article

Context compression: why less context means better AI

Context compression reduces AI agent memory usage by 26-54% while preserving task performance. Here's how it works and why bigger context windows aren't the answer.

Give your agents one place to read and write.

Put in your docs, data and notes, then connect Claude, Codex or anything else with one link.

Create a container