A note on language
The document below is in English. The results spreadsheet is still in Spanish, like the other spreadsheets here: its sheet names, its columns and its dropdown values are Spanish, and this page keeps them as they are, with the English beside them, so that what you read matches what you open.
When to use it
During the session, with the prototype or the product in front of you. It is what the moderator and the observer keep open while someone tries to get something done.
It is not the conversation script. That is the interview guide, which already covers the guided walkthrough. This is the layer the guide does not have: what was expected of each task, whether it happened, what it cost, and what was observed.
It is not the plan either. Who gets recruited, with what method and for what decision is agreed beforehand, in the research plan.
It works for moderated testing, in person or remote. Unmoderated testing has tasks and metrics but no observation, so half the sheet stays empty and you are better off going straight to the spreadsheet.
Before you use it
The success criterion is written before you test. That is the rule everything else rests on. With no prior criterion, "the task failed" is an opinion to be argued about in the results meeting; with one, it is a fact you can check. And a criterion written after seeing the data is always met.
Separate the scenario-level goal from the site-level one. "Finds the form in under 2 minutes" is measured on one task; "90 % complete the purchase unaided" is measured on the product. Mixing them produces goals nobody can check.
"With help" is not success. It is its own category. Without it, a task the moderator rescued with a hint gets recorded as completed, and the number is inflated exactly where the problem was.
Time is only recorded if the task was completed. Averaging the time of failed tasks means nothing: someone who gives up after 20 seconds is "faster" than someone who succeeds in three minutes.
The moderator does not take notes. Or takes very few. If you are writing, you are not watching. This sheet is meant to be filled in by the observer, with the moderator marking only what cannot wait.
Everything in brackets gets replaced.
The document
One sheet per task, repeated as many times as the session has tasks. It is printed or filled in on screen, one copy per participant.
Project: [client · product or flow] Participant: [ID, never the name] · Profile: [per the screener] Moderator: [name] · Observer: [name] Date: [DD/MM/YYYY] · Planned length: [60] min What gets put in front of them: [low-fidelity prototype / high-fidelity prototype / product in production] Format: [in person · remote] · [moderated · unmoderated] · Tool: [which one]
The last three lines are not decided here: they come from the research plan. They are copied onto the sheet because whoever is observing needs them in view, and because they are what explains, on re-reading, why a result says what it says.
Measurable goals for the test
Filled in before the first session. They are the same ones that go in the Metas sheet of the spreadsheet, which is in Spanish: Tiempo, Precisión, Éxito, Satisfacción.
| Category | Level | Task | Priority | Measurable requirement |
|---|---|---|---|---|
| Time | Scenario | [task] | [High/Medium/Low] | [90 % find X in under 3 minutes] |
| Accuracy | Scenario | [task] | [ ] | [90 % reach X with fewer than 2 wrong clicks] |
| Success | Site | [main task] | [ ] | [90 % complete the task unaided] |
| Satisfaction | Site | — | [ ] | [80 % score 4 or more on a 1 to 5 scale] |
Before you start
- Consent signed and received
- Recording authorised and running
- Prototype open and tested on the session's device
- Plan B ready if the test environment goes down
What you say at the opening: the script is in the interview guide. The essentials: there are no right answers, we are evaluating the product and not the person, and they can stop whenever they want.
Task [n] · [task name]
Scenario read out loud
[A situation, not an instruction. "Imagine you want X and you land on the site…", never "click the blue button".]
| Goal | [what the person should manage to do] |
| Success criterion | [what has to happen for it to count as completed, written beforehand] |
| Starting point | [screen or state they start from] |
| Data they need | [what they need to hand: email, number, file] |
| Expected path | [the steps the team assumes they will follow] |
| Assumption being tested | [the team belief this task puts to the test] |
Record
| Completed | Yes · With help · No |
| Time (s) | [only if completed] |
| Errors | [n of wrong actions before getting back on track] |
| Where they stopped | [exact screen or step] |
| What they said there | [verbatim quote, not paraphrased] |
| What their body did | [hesitation, sighs, going back, re-reading, hunting for the mouse] |
Observer's notes
[Whatever does not fit in the boxes above.]
After the tasks
The questions go open first and scaled afterwards, so the number does not anchor what they tell you.
- How did the experience feel overall?
- Which task was the hardest? What made it hard?
- Was there a moment when you did not know what to do? Tell me about that moment.
- If you could change one single thing, what would it be?
- Is there anything you expected to find and was not there?
- On a scale of 1 to 5, where 5 is best, how would you rate the experience? · [ ]
- Why that number and not one higher?
Question 7 is what rescues the scale. A 4 with no explanation says nothing; a 4 with "because the email step made me uneasy" is a finding.
Closing the session
Obstacles the person ran into: [ ]
Questions they asked during the test: [what they asked is what the product did not answer for them]
Questions they asked at the end: [ ]
What the observer takes away, in one sentence: [ ]
Note on method and limitations
With five or six participants this describes what happened in those sessions, not proportions of a population: "4 out of 5" is not "80 % of users".
Of the three indicators, effectiveness is the most solid. Time depends a lot on the device and on the session's context, and satisfaction declared at the end of a moderated session tends to run higher than the real thing, because the person taking part is talking to a person.
Findings that come out of here get prioritised separately, in the severity matrix. This sheet records; it does not decide what gets fixed first.
It expires. Review on [date]: a protocol written for a prototype that has already changed measures something else.
The results spreadsheet
The sheets above are filled in by hand, one per participant. This spreadsheet is what turns them into numbers: it downloads separately, has five tabs and calculates on its own.
What the spreadsheet contains
Metas (goals) — the four categories of measurable goal, with their level and their threshold. The requirement is written in prose so it can be read, and also as a number —maximum time and minimum success— because a spreadsheet cannot compare against a sentence.
| Categoría | In English | Level | What it measures |
|---|---|---|---|
| Tiempo | Time | Scenario | How much it takes to get there |
| Precisión | Accuracy | Scenario | How many detours there were along the way |
| Éxito | Success | Site | Whether the main task gets completed |
| Satisfacción | Satisfaction | Site | How the person felt at the end |
Participantes (participants) — one row per person: ID, profile and their satisfaction score from 1 to 5, validated so a 7 cannot get in.
Sesiones (sessions) — one row per participant and task. It is what the hand-filled protocol brings: whether they completed it, how long it took, how many errors, and what happened. The Cumple la meta (meets the goal) column checks the time against that task's threshold, and says Sin meta (no goal) when no threshold was written, instead of inventing a verdict.
Resumen (summary) — recalculates on its own and excludes the example rows:
| Indicador | In English | What it is | Where it comes from |
|---|---|---|---|
| Eficacia | Effectiveness | % of tasks completed successfully | Sesiones · column C |
| Eficiencia | Efficiency | Average time and errors per task | Sesiones · columns D and E |
| Satisfacción | Satisfaction | Average of the 1 to 5 scale | Participantes · column C |
Plus the per-task breakdown, which is the range the chart is built from.
How to use it
- Write the goals before the first session.
- One row per participant, with their satisfaction at the close.
- One row per participant and task, transcribing the sheets.
- Delete the rows marked
EJbefore sharing. The summary already leaves them out of the calculation.
Check before you use it
- Are the goals written before the first session, or were they written while looking at the results?
- Does every task have a success criterion that does not depend on the watcher's judgement?
- Do the scenarios describe a situation, or do they tell the person where to click?
- Is "with help" kept separate from "yes"?
- Are the moderator and the observer different people?
- Does the participant ID replace the name everywhere on the sheet?
- Did you write down verbatim quotes, or summaries of what you hoped to hear?
- Did you delete the
EJrows from the spreadsheet before sharing it?
Grounding
The measurable-goals structure — the time, accuracy, success and satisfaction categories, and the distinction between scenario level and site level — is adapted from the Measurable Usability Goals template by Usability.gov (U.S. Department of Health and Human Services), a U.S. government work and therefore in the public domain. The examples and the rest of the document are my own.
The per-task sheet and the closing questions come from a protocol I use on retail projects; the five session stages, from a usability-testing workshop I put together for a product team in 2022.
What surrounds this template: the research plan agrees the study, the screener recruits, the informed consent authorises the recording, the interview guide is the session script, and the severity matrix prioritises whatever comes out of this.