Confident Answers, Inconsistent Rationales
The data reveals another telling pattern: The length of AI answers is far more consistent than their reasoning.
That’s because the data behind them is rarely perfect. AI-powered recommendation systems often use a person’s previous activity to identify what might be useful or relevant to them. Generative AI can then turn those signals into a natural-language explanation, but activity histories are not exact records of someone’s interests. People click things accidentally. They browse without intending to act. They share devices and accounts. Their preferences evolve, and information may be missing or recorded incorrectly.
But Workday researchers used a new method of testing AI explanations generated from that imperfect data.
They first asked an AI model to explain why it recommended a particular item. They then altered parts of a fictional user’s activity history and asked the model to explain the same recommendation again.
The changes represented five common complications:
Irrelevant activity
Events appearing in a different order
Interests that might belong to another person
Preferences changing over time
Missing information
The researchers then compared the original and revised explanations. Did the two retain the same overall meaning? Did they refer to the same important details? Did they follow a similar structure? And did they remain roughly the same length?
In other words, the researchers asked whether AI would continue telling the same story when unreliable or irrelevant details entered the picture.
They found AI models mostly kept the overall meaning intact, with a consistency score of 0.6 out of 1, even when the underlying context changed. AI also tended to produce explanations of a similar length—even when the wording, structure, or rationale had changed.
This is particularly telling as length can often be mistaken for thoroughness or accuracy. Organizations may interpret a polished, detailed answer as a sign that AI understands the situation. But the data proves an agent’s confidence in its explanation doesn't always indicate accuracy.