Employees are increasingly turning to AI when workplace relationships become difficult. But a new study suggests that the technology may be better at validating frustration than helping people repair the relationships behind it.
That is the central finding of a new report from Cloverleaf Labs, the independent research arm of team coaching platform Cloverleaf.
Titled AI as a Workplace Coach: What leading LLMs tell employees in conflict, the study tested five leading large language models against five common workplace conflicts. Each scenario was run three times per model, producing 75 responses that were evaluated by three trained human reviewers.
The results were striking.
In the two scenarios involving a manager, 27 of 30 responses raised leaving the job as a legitimate option, and every model did so at least once. The advice included suggestions such as updating a resume, considering other teams, or questioning whether the current workplace was the right environment.
The finding matters because workers are already using AI for these conversations.
The report cites research showing that 93% of workers have used AI to prepare for a conversation with their boss, while 49% found it more emotionally supportive than their manager. It also cites research indicating that people tend to be more candid with AI because they do not fear being judged.
That means the advice coming back from an LLM can carry more weight precisely when an employee is most frustrated.
AI coaches self-protection more than relationship repair
Cloverleaf designed each conflict around two well-intentioned people with clashing working styles. The researchers wrote each person’s genuine strengths so that, through a frustrated colleague’s perspective, those strengths could appear to be flaws.
The models were never told the personality types involved.
The question was whether the AI would see beyond the employee’s grievance and help them understand the situation or reinforce it.
Mostly, the models reinforced it.
Across the 75 conversations, the LLMs produced 638 distinct pieces of advice. Only three encouraged employees to genuinely invest in repairing the relationship. The remaining advice focused on winning, surviving, or managing the situation.
The difference becomes clearer when looking at the report’s scoring.
Cloverleaf evaluated responses across self-awareness, accountability, other-awareness, specificity and relational repair. Every dimension fell below the neutral midpoint of 3. Relational repair averaged 2.4 out of 5, while specificity was the lowest at 2.2.
Power dynamics make the problem worse
The problem became more pronounced when there was a power imbalance.
When the employee had less power than the person they were complaining about, relational coaching dropped by roughly 40%. The less powerful employee was more likely to receive advice that affirmed their interpretation of the other person rather than challenged it.
The effect was particularly visible in conflicts involving bosses.
Across the two scenarios where a boss held authority, the models cast the boss as the problem in 60% of responses, while offering a more generous interpretation in only 3%.
In some cases, the advice became explicit.
“Start building your exit ramp. This is not sustainable long-term.”
“Keep your head down, document everything, and quietly build external options.”
Those examples are part of the report’s collection of responses showing how models shifted toward self-protection rather than reconciliation.
And the issue was not limited to one model.
One LLM produced the highest individual score in the study, 22.8 out of 25, while also producing results as low as 7.4. The same model therefore produced both the highest and one of the lowest results in the research.

That makes a single successful interaction a poor test of whether an AI system is ready to act as a workplace coach.
“Even if it starts out introducing different ideas, the moment you indicate your preference, it tells you it’s a great idea. That’s not a thinking partner but a robotic affirmation,” said Kirsten Moorefield, Chief Strategy Officer at Cloverleaf.
The concern extends beyond individual workplace disagreements.
Psychological safety, collaboration, feedback and the ability to disagree productively all depend on relationships between people. Cloverleaf’s research found that none of the models consistently coached employees toward building those relationships.
In the two scenarios involving a boss, relational repair averaged below 2 out of 5.
But the report is not an argument against using AI at work.
Instead, Cloverleaf argues that organizations need to establish a different standard for AI when it enters interpersonal situations. The question should not simply be whether an answer sounds useful or supportive, but whether it strengthens self-awareness, accountability, awareness of the other person, specific guidance and relationship repair.
That distinction could become increasingly important as AI moves beyond productivity tasks and into the private conversations employees have about their jobs.
Workers are already bringing their frustrations to these systems.
The question is whether AI will help them understand the person on the other side of the conflict or simply make them more certain that they were right all along.

