Are AI Explanations Dependable? New Workday Study Puts This Question to The Test

AI can now explain the rationale behind its recommendations in natural language. A new Workday study gives organizations a way to test whether those explanations are dependable: do they shift appropriately when the evidence they’re based on changes?

As enterprise AI continues to evolve and act on new data, how can organizations ensure AI-generated suggestions are dependable? A new Workday study answers that question. 

In the study, Workday AI Research—a new team dedicated to advancing trustworthy, useful, and efficient enterprise AI—examined so-called “AI explanations,” the rationale AI provides for its suggestions or answers. To test whether AI could distinguish meaningful changes from irrelevant noise, the researchers altered the underlying user activity data behind an AI recommendation—such as by adding stray clicks or making trivial changes to browsing history. Across the four models tested, explanation consistency averaged just 0.51 on a zero-to-one scale, meaning AI explanations were substantially rewritten even when the change to the underlying data was trivial or irrelevant.

This insight gives organizations reason to pressure-test AI explanations to ensure they are not just plausible, but truly dependable.

Across the four models tested, AI explanations were substantially rewritten even when the change to the underlying data was trivial or irrelevant.

Confident Answers, Inconsistent Rationales

The data reveals another telling pattern: The length of AI answers is far more consistent than their reasoning. 

That’s because the data behind them is rarely perfect. AI-powered recommendation systems often use a person’s previous activity to identify what might be useful or relevant to them. Generative AI can then turn those signals into a natural-language explanation, but activity histories are not exact records of someone’s interests. People click things accidentally. They browse without intending to act. They share devices and accounts. Their preferences evolve, and information may be missing or recorded incorrectly.

But Workday researchers used a new method of testing AI explanations generated from that imperfect data.

They first asked an AI model to explain why it recommended a particular item. They then altered parts of a fictional user’s activity history and asked the model to explain the same recommendation again.

The changes represented five common complications: 

  1. Irrelevant activity

  2. Events appearing in a different order

  3. Interests that might belong to another person

  4. Preferences changing over time

  5. Missing information

The researchers then compared the original and revised explanations. Did the two retain the same overall meaning? Did they refer to the same important details? Did they follow a similar structure? And did they remain roughly the same length?

In other words, the researchers asked whether AI would continue telling the same story when unreliable or irrelevant details entered the picture.

They found AI models mostly kept the overall meaning intact, with a consistency score of 0.6 out of 1, even when the underlying context changed. AI also tended to produce explanations of a similar length—even when the wording, structure, or rationale had changed. 

This is particularly telling as length can often be mistaken for thoroughness or accuracy. Organizations may interpret a polished, detailed answer as a sign that AI understands the situation. But the data proves an agent’s confidence in its explanation doesn't always indicate accuracy.

When the underlying context changed, AI kept its answer length consistent while its reasoning shifted. Scores are on a 0–1 scale; higher means more consistent.

The length of AI answers is far more consistent than their reasoning.

When Should an Explanation Change?

Consistency is not always the right goal.

If someone develops a new interest, acquires a new skill, changes roles, or provides meaningful new information, AI should take that change into account. A system that repeats the same explanation regardless of new evidence isn’t dependable.

The harder challenge is distinguishing meaningful change from irrelevant noise. In the study, the severity of a change barely mattered: severe alterations disrupted explanations only about 1.7% more than mild ones—a pattern that held across every model tested. 

The researchers tested four language models of different sizes. All four reacted to the mere presence of a change, not its size. The largest model was approximately 8% more stable than the smaller models. But greater scale did not eliminate the problem; even the largest model remained only moderately consistent. 

An accidental click shouldn’t redefine someone’s interests. At the same time, a sustained change in behavior might. Similarly, activity from another person using the same account shouldn’t be treated in the same way as evidence that someone’s preferences have genuinely evolved.

Workday’s study provides a structured way to measure how sensitive an AI explanation is to different kinds of change. That measurement can help researchers and organizations identify where models are stable, where they are overly reactive, and where further refinement is needed.

The severity of a change barely mattered: severe alterations disrupted explanations only about 1.7% more than mild ones.

3 Takeaways for Enterprise AI Leaders

Workplace data is incomplete and constantly evolving, but enterprise AI must be stable and trustworthy to provide real value. Agents using inconsistent and changing context must recognize and adapt to meaningful developments without being unduly influenced by isolated or irrelevant details.

For leaders deploying and governing enterprise AI, the research points to three key takeaways:

  1. Organizations should test systems with imperfect information, not just carefully prepared examples. A false positive poses just as much risk—if not more—than a clearly wrong explanation. 

  2. The explanation should be evaluated separately from the recommendation it accompanies. This is because a recommendation can remain unchanged while its stated rationale shifts. 

  3. Companies shouldn’t assume that using a larger model will automatically make an explanation more dependable.

The research also reinforces the importance of giving employees tools and training to question or correct the information behind an AI-generated suggestion. An explanation can provide a useful starting point for that interaction, but it shouldn’t create a false impression that the system possesses context it does not have.

Companies shouldn’t assume that using a larger model will automatically make an explanation more dependable.

From Plausible to Dependable

People increasingly experience an AI recommendation and its explanation as a single interaction. The explanation may influence whether they understand, accept, question, or act on what the system suggests.

As AI interfaces become more conversational, language will play a larger role in how people judge automated systems. A polished, articulate rationale can make uncertainty read like confidence, and sensitivity to noise come across as personalization.

Dependable AI explanations should continue reflecting the evidence that matters when irrelevant details change. They should also respond appropriately when circumstances genuinely evolve, communicate their limitations, and leave room for people to provide context or make corrections.

Workday’s study provides a framework for how enterprises can test their own AI-generated explanations. The ability to generate an explanation is no longer the difficult part. The harder task is ensuring that the explanation reflects the relevant evidence—and changes for the right reasons.

The next time AI explains why something is right for you, the most important question isn’t whether the answer sounds plausible. It’s whether the system would tell the same story if one irrelevant detail changed.

A note on the research: Workday sponsored the research, conducted by Guilin Zhang, Kai Zhao, Jeffrey Friedman, and Xu Chu as part of Workday's new AI Research program. The study introduces RobustExplain and was evaluated on a controlled, synthetic shopping dataset; the enterprise and workplace examples in this piece are illustrative and were not directly tested.

 

Looking for more insights on how to make AI trustworthy, useful, and efficient for the enterprise? Visit workday.com/ai-research for a complete list of Workday AI research papers.

More Reading