Reliability & drift
Why agents degrade over long runs, and how that surfaces as confident wrong answers.
4 articles · 5 definitions
Definitions in this subject
GPT-5.5 hallucination rate: what the numbers actually say
GPT-5.5's hallucination rate depends on grounding: 23% fewer wrong claims with tools, but 86% on AA-Omniscience without them. The real numbers, explained.
ArticleAI agent reliability is a context problem
AI agent reliability fails because the same task assembles different context every run. Non-determinism is a context engineering problem, not a model flaw.
ArticleAgentic context engineering: how ACE evolves contexts
ACE (ICLR 2026) beats tuned prompts by 10.6% with self-evolving contexts that avoid brevity bias and context collapse, two real failures of prompt tuning.
ArticleGPT-5.4-pro hallucinates more than GPT-5.4-nano
Vectara's 2026 benchmark shows OpenAI's flagship GPT-5.4-pro hallucinates at 8.3% while its nano variant stays at 3.1%. The reasoning-model tradeoff, explained.
Give your agents one place to read and write.
Put in your docs, data and notes, then connect Claude, Codex or anything else with one link.
Create a container