How to Build a Usability-Test Severity Codebook from Session Notes

Usability codebooks fail when they invent a P4 quote and an 80 percent success rate. This is not an interview theme codebook. Severity comes from the scale you pasted. Quotes come from the notes.
The matching generator is the Usability Test Severity Codebook from Session Notes (No Invented Quotes) prompt. Browse related cards in the PromptDig library (Browse more prompts). When a filled run survives, share the version you actually use (Share a prompt).
Ledger participants and tasks that actually appear
Build a usability-test severity codebook from pasted session notes. No invented quotes, task times, or participant IDs. Start by filling Inputs, not by asking the model to remember last week's run. If a field is blank, write NONE or NOT IN INPUTS and leave it blank through Generate. The card is built so the model cannot honestly invent a number, owner, URL, or command that you did not paste.
Paste these fields before you hit run:
Study name and prototype version: [Study]
Tasks I actually ran: [Tasks]
Severity scale I will use (paste definitions): [Scale]
Pasted session notes (who, task, observation): [Notes]
Participant labels as in the notes: [Who]
Metrics I may use only if in Notes (time, success): [Metrics]
Out of scope methods: [Out]
Words I must not use: [Banned]
Max findings: [Count or 8]
Audience for the codebook: [Audience]
That inventory is the honesty ledger. Anything that does not appear there is forbidden in the draft. If you catch yourself adding a nice-to-have after the run, you are no longer using the card. You are ghostwriting. Put the extra fact in Inputs and run again.
Cap findings at what the notes support
Generate is numbered on purpose. Do not skip a step because the first paragraph looked done. The early steps exist to stop later prose from smuggling claims.
Walk the Generate list in order:
- Honesty ledger: study, tasks, scale definitions, participant labels, every quote-ready clause in Notes. Forbidden: quotes not in Notes, times not in Notes, SUS not in Notes.
- Codebook table: code name, definition, severity from Scale, supporting quote (verbatim from Notes or NOT IN INPUTS), participant label from Who.
- Cap at Count findings. Merge duplicates. Do not invent a P0 to make the study look dramatic.
- Task map: which Task each finding came from. If Notes omit the task, write task unknown.
- Metrics: only Metrics present in Notes. If no time was recorded, write time not recorded.
- Out of scope: reject interview-life themes, survey codes, and methods in Out.
- Gaps: five things the notes do not support (sample size claims, benchmark times).
- Compliance pass: quote Banned words, invented quotes, fake SUS. Cut them.
If a step asks for a version lock, quote the version from Inputs in the output. If a step asks for a refuse list, keep the refuse list in the published artifact, not in a sidebar you delete. Reviewers should see what the model was not allowed to do.
Do not compute SUS unless the instrument is in the notes
Most failures are the same shape: a missing field gets a confident fill. A conversion rate appears. A Gradle task appears. A flash point appears. A caption appears on a job that asked for slide text only. Your review is to search the draft for numbers, names, and commands, then grep Inputs. No match means cut.
Honor the constraints as hard stops, not vibes:
- Usability-test codebook, not a qualitative interview theme codebook and not a survey codebook.
- Never invent a quote, participant, or task time.
- Severity labels only from Scale.
- Do not compute SUS or NPS unless Notes include the instrument.
- No emojis.
When the card says not legal advice, not certification, not an exam dump, or not a caption engine, that sentence belongs at the top of the output. Deleting it to look more finished is how you inherit risk.
Keep interview and survey codes out of scope
Finish with the compliance pass the prompt already asks for. Quote the banned-word hits. Cut them. Print character counts when the job has a cap. Print word counts when the job has a budget. List gaps as gaps. Five missing facts are more useful than one smooth paragraph.
Tags on the card (usability test severity codebook, session notes synthesis, no invented quotes) are a reminder of the job shape, not an invitation to wander into a neighboring cluster. If you need a different surface, open a different PromptDig card rather than stretching this one.
Fill the card, then run
Replace every bracket. Run on ChatGPT, Claude, or Gemini. Read the ledger first, then the artifact. If the model invents a commit, KPI, DOI, PEL, bid, or logo, discard the run. Tighten Inputs. Run again. Share the filled card that survived, not the first draft that sounded done.