06Chapter 6 of 13

Measuring the Invisible: Detecting Cognitive Distance in Real Systems

From Cognitive Load to Cognitive Distance · Izaias Cavalcanti · about 12 minutes to read

Cognitive distance is easy to recognize once someone points it out. People hesitate, second-guess their actions, and complete tasks without being able to explain how the system works. Measuring it is harder (Hornbæk 2006; Paas et al. 2003). Traditional usability metrics focus on observable outcomes such as task completion, time on task and error frequency. They capture performance but not how people conceptualize the system. Cognitive distance sits in the gap between visible behavior and internal understanding, and detecting that gap requires different research strategies.

6.1 Why Traditional Usability Metrics Fall Short

Task success hides conceptual confusion

In most usability tests, participants complete set tasks while researchers observe, and if a participant reaches the correct outcome the attempt is recorded as a success. Success does not imply comprehension. A participant may succeed by exploring until they stumble on the right option, or by following visual cues without knowing why those cues lead to the result. When the interface later changes slightly, the same person may struggle again. The test recorded success, but the conceptual model never stabilized. The pattern is common in complex enterprise systems, where people memorize procedures rather than understand workflows.

Efficiency metrics capture speed, not alignment

Faster completion suggests efficient interaction, but speed alone says nothing about how people interpret the system. Someone may be quick because they have memorized the steps, an interaction closer to muscle memory than to understanding. When the system changes or the scenario shifts slightly, the learned sequence fails and they must reconstruct their strategy. Apparent efficiency can hide considerable cognitive distance.

Error rates reveal symptoms, not causes

Errors are rightly read as signals that an interface needs work, but error counts do not say why the mistakes happened. A person may have misread a term, misunderstood the structure of the workflow, or held a wrong model of how the system processes information. All of these produce the same observable error. The cause lies in the conceptual gap between the person’s reasoning and the system’s organization, and without examining that gap directly, teams fix surface issues and leave the real problem in place.

6.2 Observing Mental Models

Think-aloud protocols

One of the most useful techniques is to ask participants to say what they are thinking while they use a system (Ericsson and Simon 1993). In a think-aloud session, people describe what they expect to happen before acting, explain why they choose particular options, and voice their assumptions about the system’s behavior. These verbalizations expose the mental model guiding the interaction, and moments of uncertainty or surprise show researchers exactly where cognitive distance appears. Think-aloud studies are long established in usability research and cognitive psychology, and they supply qualitative evidence that quantitative metrics cannot.

Concept mapping

Another approach asks participants to describe a system’s structure after using it, for example by drawing how they believe it is organized. These diagrams often differ strikingly from the system’s actual architecture. People group features by the goals they associate with them and describe relationships that do not exist in the design. Comparing the drawings with the real structure makes cognitive distance directly visible: large discrepancies indicate that the system’s organization is hard to infer.

Interviews that probe understanding

Post-task interviews offer a third window. After a scenario, participants can be asked questions such as: How does this system decide what to show you? Why do you think this step comes before the next one? What would happen if you did these steps in a different order? Such questions push people to articulate their models. When they struggle to explain the system’s behavior, or give conflicting explanations, researchers have evidence that alignment is weak, evidence that rarely surfaces in traditional metrics.

6.3 Behavioral Signals of Cognitive Distance

Hesitation during navigation

One of the earliest observable signs is hesitation. People move the cursor across several options without choosing one, pause to reread labels, or scan the screen repeatedly. They are trying to map their intention onto the structures the interface offers, and when that mapping is unclear, hesitation grows. Watching for these moments shows where alignment breaks down.

Repeated backtracking

Frequent backtracking is another signal. People move forward through several steps, then return to earlier screens to reconsider their choices. They are testing hypotheses about how the system works: trying a path, observing the outcome, revising their understanding. Some exploration is a natural part of learning, but excessive backtracking means the system’s structure is hard to infer and people must keep revising their model as they go.

Reliance on external memory

When distance becomes large, people compensate with external aids. They write down instructions, keep reference notes and build personal checklists for routine tasks. This is common in enterprise environments, where employees maintain private documentation for actions they perform every week. Those notes become an external mental model. The system never became conceptually transparent, so its users built their own interpretation outside it.

6.4 Quantifying Conceptual Alignment

Mental model similarity analysis

Researchers have developed ways to compare users’ mental models with system structures. One approach asks participants to sort interface elements into categories that make sense to them and compares their groupings with the interface’s actual organization. High similarity suggests people perceive the system as it was designed; low similarity suggests substantial distance. The method turns conceptual alignment into data that can be compared across designs.

Cognitive walkthroughs

The cognitive walkthrough offers a structured way to evaluate how well an interface supports learning and first use, and teams can apply it during design without recruiting participants (Polson et al. 1992). Evaluators step through each action from the perspective of a new user and ask: Will the user know what to do at this point? Will they understand why the action is needed? Will they recognize that the outcome matches their goal? Examining each step this way reveals the moments where alignment is likely to fail.

Longitudinal observation

Some forms of distance only appear over time. A system may seem easy in initial testing and reveal conceptual weaknesses as people meet new scenarios. Longitudinal studies follow people over extended periods to see whether they develop stable mental models or keep relying on procedural shortcuts. If people repeatedly struggle with new tasks in a system they already use, cognitive distance is likely present. Long-term observation is an important complement to short usability tests.

6.5 The Cognitive Distance Index

The techniques above reveal cognitive distance, but they do not put a number on it. Teams that want to compare two designs, track a system across releases, or report to people who were not in the research sessions need a measure. The Cognitive Distance Index (CDI) is one.

Cognitive distance cannot be observed directly. It has to be inferred from indicators that together reflect the gap between system logic and user understanding, and no single indicator is enough. Self-reported understanding is the most direct, but people do not always know what they do not know, and a successful task can inflate confidence when comprehension is shallow. Prediction and comprehension questions are less vulnerable to that bias but only test what they ask about. Open explanations capture the actual content of understanding but must be scored by hand. The CDI combines four dimensions, shown in Table 6.1, so that no single measure carries the construct alone.

Table 6.1. The four dimensions of the Cognitive Distance Index.
Dimension Question it answers Typical instrument
Perceived understanding Can the user explain, in plain words, what the system just did? Agreement items after the task, plus an open explanation (“In your own words, what just happened?”) scored against a rubric.
Predictive accuracy Can the user anticipate the outcome before it is shown? Multiple-choice prediction questions asked after the user confirms but before the result appears.
Interaction confidence Does the user feel certain about what happened and what to do next? Confidence ratings on each prediction, plus an “Is the task done?” item (certain / think so / not sure).
Conceptual transparency Can the user identify why the outcome occurred, not only what it was? Comprehension questions that test the rule behind the outcome, plus the causal content of the open explanation.

Two details of the instrument matter more than they first appear. The prediction questions are asked before the result is shown. A person who predicts correctly beforehand and explains correctly afterwards holds a causal model of the system; a person who answers correctly only after seeing the result may simply be repeating what the screen said. And the comprehension questions test causal rather than factual knowledge. Knowing that an appointment is on Thursday at 14:30 is recall. Knowing why the system booked a general consultation instead of the specialist the user asked for requires understanding the rule that governs its behavior.

Computing the score. Each dimension is converted to a distance score from 0 to 100, where 0 means complete alignment and 100 means maximal misalignment. Predictive accuracy, for example, becomes 100 minus the percentage of correct predictions. The CDI for a participant or group is the mean of the available dimension scores, and it is not reported on fewer than two dimensions, because partial composites are hard to compare. Table 6.2 gives the interpretation bands.

Table 6.2. CDI interpretation bands.
CDI Band What it looks like Response
0–25 Strong alignment Users explain outcomes correctly and predict future behavior reliably. Monitor.
–45 Early warning Scattered confusion; some users cannot account for outcomes. Review terminology and feedback wording; run targeted tests.
–65 Needs attention Misunderstanding recurs across users. Redesign result screens, explanatory copy and causal labels.
–100 Serious misalignment Users routinely succeed operationally while failing interpretively. Restructure how outcomes are communicated and how information is organized.

The bands are a working proposal, not a validated standard. They are anchored to the logical extremes of the construct and to pilot observations, and the equal weighting of the four dimensions has not yet been tested against alternatives. The CDI is also baseline-relative: it is most informative when used to compare versions of the same system, or the same system across user groups, rather than to rank unrelated products. Teams should treat a CDI score the way a clinician treats a single blood test: as a reason to look closer, not as a diagnosis.

The four dimensions can also diverge, and when they do the pattern is diagnostic. The most concerning combination is what this book calls the calibration gap: moderate or high confidence alongside prediction accuracy near chance. People who know they do not understand will ask for help. People who are wrong but confident cannot correct themselves, because nothing tells them anything needs correcting (Kruger and Dunning 1999). In a license renewal or a medical booking, that is the person who leaves believing the task is finished when it is not.

6.6 Testing the Idea: The Everyday Services Study

A construct earns its place by making predictions that could turn out to be wrong. The Everyday Services Study is designed to test the central claim of this book under controlled conditions (Cavalcanti 2026).

The study isolates one variable: the language a system uses to report what it has done. Participants complete three everyday tasks on working prototypes: a bank transfer, a medical appointment booking and a driver’s license renewal. Half see Version A, written in plain language. Half see Version B, written in the institutional register common in deployed public-service, banking and healthcare systems (Li and Liu 2025). Everything else is identical: task structure, number of clicks, form fields and actual outcomes. Only the language changes, and it changes most on the result screens, where the system has its last chance to pass on what it knows (Table 6.3).

Table 6.3. Result screens in the Everyday Services Study. Task structure and outcome are identical; only the language differs.
Service Version A: plain language Version B: institutional language
Banking Done. Mary receives $200.00 today by 18:00. The money has left your account. You do not need to do anything else. Operation processed successfully. Authentication: 8F2-9921-B. Ref: TX-29381.
Medical Booked: General consultation, Thursday 14:30, Central Clinic, Dr. Lima. Bring your ID. At this visit, ask the doctor for the skin specialist referral. Booking registered. Protocol 5512A. Type: Triage, General Practice. Status: Pending allocation.
License Payment received. You are done. We review your renewal within 5 days, then mail your new license, which arrives in about 10 days. Request filed under protocol 2026-88412. Status: Stage 2 of 6, Administrative instance.

The Version B screens are not caricatures. Each is a pattern found in real systems, and each leaves unanswered questions that matter to the person reading it. The license screen does not say whether the fee was paid, whether the renewal was approved, or whether anything else is required. Version A answers all three.

The study runs online, unsupervised, in English, Spanish and Brazilian Portuguese, because cognitive distance matters most precisely when people use systems alone. Every participant completes all three services in one of six counterbalanced orders. Measurement is built into the flow using the CDI instruments described above: prediction questions after the user confirms but before the result appears, comprehension questions immediately after, and a short questionnaire ending with an open explanation that is scored without knowing which version the participant saw.

The study makes five directional predictions. Compared with Version B, participants who see Version A should report higher understanding, give better open explanations, predict outcomes more accurately and feel more confident, and their CDI should be lower in all three services. Version B participants are expected to predict at or near the 25 percent chance level. Just as important are three checks that must not differ between versions: completion rates, self-reported ease and effort, and task duration. If Version B were simply harder, the study would be measuring load. The design only tests distance if effort stays equal while understanding diverges.

What the first pilot showed. An earlier iteration of the study ran as a small qualitative pilot. Its results are not evidence in a statistical sense, but they illustrate the pattern this book describes. Participants using the institutional version completed tasks at normal speed and reported no unusual difficulty. Their open explanations told a different story. Many described what had happened in the system’s own words without interpreting them: “it said the protocol was filed,” “it showed stage 2 of 6.” These are explanations that quote the system back to itself. They are what high cognitive distance sounds like.

The pilot also changed the study. The institutional screens were revised so that their ambiguity reflects real systems rather than exaggerated ones, task instructions were standardized so that any advantage for Version A comes from the result screens alone, and the medical comprehension questions were rewritten after almost everyone answered them correctly in both versions.

Limits. The study uses prototypes rather than live systems, measures distance at one moment rather than over months of use, and relies on hand-scored explanations that will need a second independent scorer. Whichever way its results fall, it will answer a narrow question well rather than a broad question loosely. That is the right order in which to build evidence for a new idea.

6.7 Diagnosing Cognitive Distance in Practice

Design teams do not always have the resources for a controlled study, but many signs of cognitive distance can be found with simple practices: watching for hesitation during usability tests, holding short interviews about how people believe the system works, and asking participants to explain a workflow in their own words. Adding even two or three prediction questions to an existing test, asked before the result appears, brings a lightweight version of the CDI into everyday practice. These small steps reveal conceptual gaps that traditional metrics overlook, and once teams recognize the patterns they can start addressing the causes. Table 6.4 is a practical checklist for spotting conceptual friction early.

Table 6.4. Diagnostic signals of conceptual friction and how to detect them.
Signal What it looks like Likely cause How to measure
Hesitation Long pauses before obvious actions Unclear mapping between labels and intent Time to first action; think-aloud transcripts
Backtracking Frequent returns to earlier steps Navigation does not match the task model Path analysis; loop detection
Procedural dependence Users repeat sequences but cannot explain them Completion without understanding Explanation request after the task
Support reliance Help desk calls for routine tasks The system cannot explain itself Support ticket categories and trends
Terminology mismatch Users rephrase interface language to make sense of it Internal vocabulary in user-facing text Interview coding; language comparison
Calibration gap Confident users who predict wrongly Outcomes reported as states, not consequences Prediction accuracy against confidence