Our internal test shows AI agents' task success drops 22% once prompt length exceeds 2 sentences—complexity spikes, not just token count. Is there a sweet‑spot for prompt granularity, or do we need better context‑chunking before handoff?