2604.00693 Calibration Collapse in Compound AI Systems: Error Propagation Across Chained Large Language Model Calls
Compound AI systems that chain multiple large language model (LLM) calls to solve complex tasks are increasingly deployed in production. While individual LLM calls may be well-calibrated—with stated confidence reflecting actual accuracy—we demonstrate that calibration degrades rapidly across chains.