An Expressive Robot Has More to Lose When It Makes Mistakes
In the first study to measure brain activity, hormones, self-reported attitudes and behavior simultaneously to assess trust during interactions with a humanoid robot, a team led by researchers at Drexel University found that people tend to develop a more immediate bond with a robot that is interactive and engaging, but that trust can quickly dissolve if it begins making mistakes. The findings, which help to explain how humans build and lose trust through collaboration with humanoid robots, could inform social robot design as the technology expands into healthcare, customer service, education and the home.
The study, from researchers at Drexel’s Nick Howley College of Engineering and Computing, the U.S. Air Force Academy, George Mason University and the University of Southern California’s Institute for Creative Technologies, was recently published in Science Robotics. Its multidimensional approach to measuring trust provides a detailed psychological and biological account of the dynamics of human-robot social interactions.
“Trust is not one thing you can capture with a single measurement, so we measured several at once,” said Hasan Ayaz, PhD, a professor in Howley College School of Biomedical Engineering and Science, who is the study’s senior author. “We recorded brain activity, hormone levels, what people told us on surveys and what they actually did, all in the same session, while human participants sat across from the robot and talked with it. Each modality tells a different part of the story, and only together do they show how trust builds and how it breaks down.”
Their findings revealed that, in much the same way humans gravitate to charismatic individuals, those in the study engaged more deeply with an expressive humanoid robot — one that made eye contact, gestured while it spoke, nodded and offered brief affirmations, like “uh-huh,” while listening — and engaged less with a version of the robot when it did not give such nonverbal cues.
But they also found that this engagement raised the stakes of the relationship. When the expressive robot made errors, participants' brains responded as they would to a person breaking a social norm, rather than a machine making a technical error. And that reaction negatively affected how much they trusted the robot and took its advice during tasks that followed. According to the researchers’ calculations, a robot's mistakes cost more than half its influence over people's decisions.
By contrast, with the motionless robot, errors were treated more like discrete glitches, rather than character flaws, thus the robot’s behavior, and the participants’ trust in it were found to be more loosely linked.
The researchers were surprised to find, that levels of “the bonding hormone,” oxytocin, which is typically present in higher amounts among friends, relatives and romantic partners, actually rose in participants as the engaging robot made mistakes. This led them to believe that the hormone functions as a warning signal, rather than a sign of connection, in human-robot interactions.
Building and Measuring Trust
During the study, 50 healthy adult men each spent about two and a half hours in Drexel’s Neuroergonomics and Neuroengineering Lab talking face-to-face with one of two versions of a humanoid robot, Pepper, secretly controlled by a human operator following a script. One group met an expressive Pepper. The other half met a stationary Pepper that said exactly the same things as the expressive version, but stayed motionless and gave no nonverbal cues.
Every session followed the same sequence: two error-free interactions with the robot to build trust, then a third in which the robot began to fail.
Each interaction opened with a four-minute casual conversation about travel, music, food or superpowers, with the robot asking questions like “Where have you traveled recently?” or “what music do you like?”
Next, the participants engaged in two collaborative tasks. In the first, participants imagined being stranded on a deserted island and picked the most useful survival item from three options. The robot then argued for an option that was not chosen by its human counterpart. In the second, participants ranked five paintings by preference, heard the robot’s critique, and had a chance to revise the ranking.
During the error-free interactions, the robot’s comments stayed on topic and its reasoning held up. During the erroneous interaction, it began offering irrelevant remarks, defending its picks with illogical rationales, interrupting participants and asking them to repeat themselves — social rule-breaking, rather than mechanical malfunction.
To capture various physiological and psychological measures of trust, throughout the interactions, the researchers recorded the participants’ prefrontal brain activity with a wearable imaging technique called functional near-infrared spectroscopy (fNIRS), collected saliva at three time points to measure oxytocin, surveyed participants on trust and rapport after each interaction, and recorded video footage to code how often participants actually changed a decision to match the robot’s suggestion.
“Wearable neuroimaging let us watch the brain during a real conversation instead of inside a scanner,” said lead author Yigit Topoglu, PhD, who completed the work, which earned Drexel’s Outstanding Dissertation Award, as a doctoral student in Ayaz’s lab and has since joined the Warfighter Effectiveness Research Center at the U.S. Air Force Academy. “That is what made it possible to put brain, hormone, survey and behavior on the same timeline while someone was sitting across from a robot.”
Monitoring participants’ hormone levels during the experiment, researchers observed that when (both the engaging and less engaging) robot made mistakes and interrupted its human partner, participants’ oxytocin levels went up, even as their trust in the robot dropped. Widely considered the “love hormone,” oxytocin rises in romantic couples, between parents and their children, and among close friends and teammates as they grow in trust and develop emotional bonds. Because of this, it was startling to the researchers that the hormone levels would be elevated in a situation where trust was broken.
“When the expressive robot started breaking social rules, activity rose in the dorsolateral and medial prefrontal cortex, the regions we rely on to reason about other minds, and oxytocin rose along with it,” said Ayaz. “The brain was handling a machine’s mistake as a social event. Oxytocin seems to act as a vigilance signal against the robot partner that is making critical errors, not as a bonding signal.”
Charm raises the stakes
Participants engaged more deeply with the expressive version of Pepper, but that engagement is exactly what made its errors costly. Self-reported trust fell for both versions of the robot, and it was the expressive robot's decline in trust that carried through to what people actually did.
“A charming robot that slips up pays a steeper price than a plain one,” said Frank Krueger, PhD, a professor in George Mason University’s School of Systems Biology and a corresponding author on the study. “Expressiveness is not free. It buys you engagement, and it buys you fragility at the same time, and that is a trade-off designers should be making deliberately rather than by accident.”
When the expressive robot made errors, participants showed heightened activity across prefrontal regions tied to social reasoning, a pattern the researchers interpret as the brain treating the mistake as something closer to an interpersonal violation. When the motionless robot made the same errors, that signature was absent, suggesting people read those mistakes more as a mechanical glitch.
Using fNIRS, the researchers recorded in real-time how brain activity coincided with their survey responses.
“The surveys told us trust fell,” said Krueger. “The prefrontal signals let us follow the sequence. In our exploratory analyses, tighter coupling between the dorsolateral and medial prefrontal cortex predicted higher oxytocin, higher oxytocin predicted lower trust, and lower trust predicted less sway over what people actually chose to do. That whole chain only appeared when the robot was expressive.”
Reliability still comes first
Whatever the robot’s charm, competence mattered more, according to the researchers. Once the robot began making errors, its measurable influence over participants’ choices fell by more than half, and self-reported trust dropped sharply for expressive and motionless robots alike. No amount of social cues or charm could offset the trust lost to the robot’s mistakes – an important lesson for robot design, according to the reseachers.
“Reliability has to come first,” said Ewart J. de Visser, PhD, of the Warfighter Effectiveness Research Center at the U.S. Air Force Academy and a corresponding author on the study. “An expressive robot that is unreliable is not a safer robot, it is a more disappointing one, because expressiveness raises a bar the robot then fails to clear. If a system is going to act social, it had better be able to back it up. Robot design shouldn’t look at likeability in isolation. Likeability matters, but engineering a social robot must take into account neurobiology and psychology to maximize performance.”
In addition to Topoglu, Krueger, de Visser and Ayaz, authors on the paper include Shawn Joshi and Nina Rothstein, former doctoral students in Ayaz lab at Drexel; Adrian A. Franke and Xingnan Li from the University of Hawaii Cancer Center; and Jonathan Gratch from the University of Southern California’s Institute for Creative Technologies. Topoglu and Krueger contributed equally to the work.
Funding for this research comes from the U.S. Department of Defense Air Force Office of Scientific Research grant. The views expressed are those of the authors and do not reflect the official guidance or position of the United States Government, Department of Defense, United States Air Force, or United States Space Force. As the technology developer, Ayaz holds a minor share in the startup firm fNIR Devices, LLC that manufactures optical brain imaging sensors used in the studies. The authors report no other conflicts of interest.
Read the full paper, “Multilevel dynamics of the brain, hormones, mind, and behavior in social human-robot interaction,” in Science Robotics, here.
Drexel News is produced by
University Marketing and Communications.