I did not run this study. That is the part that matters most to me.

A research team from Abu Dhabi University, KFUPM and Corvinus University took the AI coach built on the 5 Spheres of Empathy and tested it independently, as part of the Responsible AI Consortium with QS, under the Future17 programme, mentored by Gasm Elbary Elkhair. They designed the study, ran the sessions, collected the scores, and wrote up what they found. Including the parts that do not flatter anyone.

See the findings (PDF) → 11 slides, ~3 min read

Twelve participants were recruited. Eleven finished. Between them they completed forty-five coaching sessions, four each, thirty to forty-five minutes at a time, with their empathy measured three times along the way using the Toronto Empathy Questionnaire: once at baseline, once at the midpoint, once at the end.

Scores went up, and kept going up

The average score moved from 49.5 to 51.9 to 56.4, out of a maximum of 64. That is a 14% rise across the study. What interests me more than the headline is the shape of it. The scores did not spike early and flatten out, which is what you would expect if people were simply learning what a good answer looks like. They rose at every measurement point, and they rose faster in the second half.

Chart showing average Toronto Empathy Questionnaire scores rising from 49.5 at baseline to 51.9 at midpoint to 56.4 at final

Then the story stops being simple

Break the average apart and the two cohorts behave nothing like each other. Abu Dhabi rose 9.4 points. KFUPM moved 0.3. Same coach, same scenarios, same four sessions.

Part of that is arithmetic. KFUPM started from a higher baseline, so there was less room to move. The rest is probably context and culture, and neither the team nor I can tell you which parts with any confidence. An average of 14% hides two completely different experiences, and I would rather show you that than the clean number on its own.

They did not get bored

Engagement usually decays. In this study it grew, reaching 9.5 out of 10 by the final session. And 58% of the reflection entries described a genuine shift in perspective. Not "that was useful" or "I felt good". People describing a situation they now saw differently, in their own words.

The finding that stopped me

Each session, participants chose one of five scenarios to work through: a customer, a colleague, a family member, someone they disagreed with, or themselves.

They chose themselves nearly a third of the time. More than any other option. Nobody told them to start there, and nobody explained the framework to them first.

The five spheres listed in order, with the first two, self work on values and beliefs and self work on the physical, mental and emotional, marked as where participants started

Spheres 1 and 2 are both self work: your values and beliefs, then your physical, mental and emotional state. They come first in the method for a reason, and people who had never heard of the method walked straight to them anyway. Self-understanding is not a nice-to-have that you get around to once you have been kind to everyone else. It is the floor. You cannot extend to others what you have not done for yourself.

What this study cannot claim

Twelve people started and eleven finished. There was no control group. The timeline was compressed. The scores were self-reported. Every one of those limits is in the team's own write-up, stated plainly, near the front. Their sentence, not mine:

"We can say empathy scores went up and people engaged deeply. We cannot say the AI caused it."
From the research team's report

A student team handed a promising result and a friendly sponsor could have rounded that corner. They did not. That rigour is rarer than the finding, and it is the reason I am comfortable publishing any of this.

In their words

"I had never looked at empathy as a concept before this. I learned how people think, and how they want to be understood."
Ahmed Faarih, Corvinus

"The 14% was real. The reflections were real. Now run it again: longer, larger, with a control group and clear safety rules."
Zoha Salman, Abu Dhabi University

"The hard part wasn't the AI. It was getting people to stop and reflect, every session, every week."
Thamer Al-Hirsan, KFUPM

That last one is the whole thing, really. The technology was never the difficult part. Getting a person to stop, sit with something uncomfortable, and put words to it, every week, when nothing is forcing them to, is the difficult part. It always was. An AI coach that helps with that is doing something worth measuring properly.

What happens next

The next study is being planned now. Bigger, longer, with a control group. Until then, a 14% rise in eleven people over four sessions is exactly one thing: a reason to run the real study, and not a claim to put on a slide.

The work was run as part of the Responsible AI Consortium with QS. Optomize and the 5 Spheres method are mine. The research is theirs.

See the findings (PDF) → 11 slides, ~3 min read