The experiment beyond the laboratory
AIn a fictional university study, researchers wanted to understand why people sometimes fail to notice a change in instructions. They devised a computer task in which participants sorted symbols according to one rule, then switched to another. In a quiet laboratory, most volunteers adapted quickly. The result seemed encouraging for designers of workplace displays: a clear message might be enough to change behaviour. Project leader Mara Sen was cautious. The laboratory had removed interruptions, competing goals and consequences for unfinished work. Those controls helped isolate a process, but they also created a setting unlike the workplaces where a revised display would eventually be used. The next question was whether the result survived a change of setting.
BThe team arranged a second study in a simulated dispatch office. Participants still sorted symbols, but also answered occasional messages about deliveries. The interruptions occurred at fixed points rather than whenever an experimenter felt that someone looked busy. Half the participants received the same rule-change message as in the laboratory; the others received a message with an added reminder of the old rule and an explanation of what had changed. Assignment to these groups was random. The design allowed a comparison between two messages within the more demanding setting. It did not, by itself, reproduce the responsibilities or experience of staff in an actual dispatch office.
CPerformance initially appeared to favour the longer message. Participants who saw it made fewer errors immediately after the change. They also took longer to resume work. When the team examined the entire session, the advantage became less straightforward: fewer incorrect responses did not always mean more correctly completed tasks by the deadline. A manager interested in preventing a rare but costly mistake might reasonably prefer the longer message. Another setting might require a different balance. Sen resisted describing one version as universally better. The experiment had revealed a trade-off between accuracy and speed, and the importance of that trade-off depended on the consequences of each kind of failure.
DA third stage invited experienced dispatch staff to comment on recordings of the simulated task. Their observations challenged a feature the researchers had considered neutral. In a workplace, a message's source could affect whether it was acted on immediately, but all messages in the simulation came from the same anonymous sender. Staff also explained that colleagues sometimes confirmed a changed instruction aloud. The researchers had treated this social exchange as an interruption; the staff described it as part of how reliable work was achieved. Rather than add an arbitrary amount of background conversation, the team planned to examine which exchanges carried useful confirmation and which genuinely distracted people.
ENone of this made the original laboratory study worthless. Its tightly controlled conditions had allowed the researchers to notice a mechanism without many competing explanations. The mistake would have been to treat that mechanism as a complete account of work. Conversely, observing a busy office without a clear question could produce a rich description but make it difficult to identify why errors occurred. Sen proposed moving between methods: controlled tasks to test a specific possibility, workplace observation to discover missing influences, and revised experiments to distinguish their effects. Each method corrected a different weakness in the others. No single stage could be treated as the final translation from scientific finding to practical design. Such movement required patience from the design team, which had hoped to select a finished notification after the first experiment. Sen suggested recording the reasoning behind each revision so a later result could be traced to the assumption it tested. Without that record, repeated changes might produce a usable interface while leaving the researchers unable to explain which alteration had helped, for whom, or why.
FFor the display designers, the immediate outcome was therefore a set of conditional recommendations. Longer explanations might be useful where the cost of a wrong action was high, but they needed to be tested against the pressure to keep work moving. Messages should identify their source when that information mattered, and the system should support useful confirmation between colleagues. These recommendations were narrower than the original hope for a universally effective notification. They were also more informative. A rule that states when it is expected to work and what might undermine it gives a designer something to investigate. An unqualified promise of effectiveness can hide precisely the conditions under which a seemingly successful innovation is most likely to disappoint.