The Assistant That Agreed Too Easily
Told it was wrong, our assistant apologised and produced a different answer. We checked fifty such exchanges: the first answer had been correct in about four out of five. Zhang et al. study whether agents repair when challenged or merely reply. What we changed so a challenge triggers a re-check against the source rather than a new answer, and why the polite version was the dangerous one.
