Master's Thesis

I spent two years on a question most teams wave past: how do you actually measure whether someone is relying on a machine the right amount?

Answering it took a 189-person study, a simulation game I built in Unity, and a new metric for reading reliance as a pattern rather than a number. Three documents came out of it. Start wherever you like, or read on for how the study worked.

The documents

The thesis · ProQuest, 2025

Quantifying Calibration: Bridging Trust and Reliance in Automation Across Cultural Values and Dispositional Factors. The full document, from the gap in the literature through the conceptual model, the study, the path analysis, and the temporal work behind Use-in-Range.

Thesis title page Thesis page 25 Thesis page 80 Thesis page 100

Preview only. Selected pages shown.

The workshop paper · CHI ’25 AutomationXP, CEUR Proceedings

A short first-author paper laying out the conceptual model before the data came in, and the study design meant to test it.

CHI workshop paper, first page Workshop paper page 2 Workshop paper page 5 Workshop paper page 8

Preview only. Selected pages shown.

The defense · August 2025

The visual companion to the oral defense. The same arc in condensed form, ending on what calibrated use would mean for systems already in the field.

Defense deck, title slide Defense slide 2 Defense slide 36 Defense slide 70

Preview only. Selected slides shown.

How the study worked

Collaborative automation lets the person decide how much to hand over. Partially automated driving, AI clinical decision support, a robot that takes half of your task. Decades of research looked at what the system does and what the situation is. Far less looked at who the operator is, and almost none of it had a rigorous way to say whether the reliance that followed was appropriate.

So I built one. Calibratio is a Unity simulation game where you sort items alongside an adaptable robot teammate named Otto, whose capability shifts underneath you. Leading a team of seven, I ran 189 people through it. Surveys captured cultural values, propensity to trust, faith in technology, and baseline trust before anyone touched the game. The game itself captured continuous reliance behavior, plus repeated in-task trust and self-rated performance.

Robust structural equation modeling in R tied it together. Baseline trust was shaped by collectivism, uncertainty avoidance, and faith in general technology. In-task trust tracked how capable Otto actually was, and moved inversely with how well people thought they were doing on their own. Reliance was predicted by power distance, by Otto’s capability, and by self-rated performance.

Out of that came Use-in-Range, a metric I adapted from time-in-range in diabetes care. It reads reliance as a pattern over time rather than a single number, and classifies it as calibrated use, misuse, or disuse. If your lab is working on how people trust and rely on automation, I’d like to hear what you’re measuring.

Degree
M.S. Human Factors Engineering, Tufts University
Advisor
Dave B. Miller, Ph.D.
Committee
Daniel J. Hannon, Ph.D., Holly Taylor, Ph.D.
Participants
N = 189
Analysis
Robust structural equation modeling in R (lavaan)
Built with
Unity (C#), Qualtrics, Prolific
Funding
Wittich Grant, Trefethen Fellowship
Defended
August 2025, passed

Get in touch →

← Back to the shelf