How Wizard of Oz Testing Actually Works When No One Is Watching

You have a product idea that involves some kind of intelligence — a chatbot, a recommendation engine, a voice assistant — and you need to know whether users will interact with it before you spend months building it. The Wizard of Oz approach means you sit behind a curtain and pretend the system is automated while real people use it. The activity sheet is what keeps you from losing your mind during that process. A Wizard of Oz Activity Sheet is simply a structured log that records every interaction during a manual simulation test. It captures what the user said or did, what response you delivered, how long it took, how the user reacted, and any issues that came up. That's it. It's not a research paper. It's a working document.

Where to Get Of Oz Activity Sheets and What They Should Actually Look Like

There isn't one canonical source for these sheets because they're fundamentally custom-built for each test. You'll find templates scattered across UX research blogs, Notion template galleries, and GitHub repos. What matters more than where you get one is how you set it up. Here's the structure I use and have used for probably a dozen tests now. Google Sheets works fine. I create columns for timestamp, session ID, user prompt, scripted response, actual response delivered, latency in seconds, user reaction (coded as positive, neutral, confused, frustrated), and an issues column. One tab per session. Color coding the response column — green for scripted, yellow for improvised, red for broken — saves time during review. I've seen people spend twenty minutes per session trying to figure out what happened because they didn't color code. The exact Of Oz Activity Sheets you end up using will depend on your modality. A voice interaction test needs different columns than a chat interface test. Voice tests require you to track silence duration, interruption timing, and whether the user repeated themselves. Chat tests need turn count and whether the user pivoted topics entirely.

The Part Nobody Talks About: Timing Intervention

Most people set up their activity sheet and immediately start logging responses. That's backwards. The first column you should define is the timing column — specifically, when you triggered the response, not when you logged it. This distinction matters more than it seems. During a voice UI test I ran for a smart home scheduling prototype, we were simulating an AI that could understand natural language requests like "remind me to call mom when I get in the car." The initial sheet tracked response text and user reaction but didn't capture the gap between the user finishing their sentence and the response playing. Users would sometimes start speaking again while the "system" was thinking. That second utterance was never recorded, and it turned out to be significant — roughly a third of users added clarifying information during that pause, and we had no record of it. The fix was adding a start-timestamp column that I filled the moment the user stopped talking, before I even thought about what response to deliver. It takes an extra three seconds per entry and completely changes how readable the data is afterward.

Get the Full Details

Wizard of Oz Printable Activity Pack, Reading Comprehension, Coloring (PDF Sheets) - Etsy
Wizard of Oz Printable Activity Pack, Reading Comprehension, Coloring (PDF Sheets) - Etsy

Counter-Intuitive Things That Actually Matter

Don't separate the user's input from your intervention in different rows. Keep them on the same row. When you're reviewing thirty sessions later and trying to figure out why a particular feature request came up three times, you need to see the full interaction in one glance. Splitting them forces you to match session IDs and line up rows mentally, which is slow and error-prone. Include a notes column even when you think everything fits in the structured fields. The things you can't quantify — the user's tone, the environment they were in, the fact that they seemed distracted because of something outside the test — these end up being the most useful data points during debrief. I once spent an hour trying to explain why users consistently performed poorly on a particular task, only to realize from my notes that every affected user had been testing on a noisy subway. The sheet saved me from drawing a wrong conclusion. Don't pre-write all your responses before the session starts. It feels like good preparation but it limits you. When a user says something unexpected — and they always will — you need to be able to respond naturally while still capturing what you actually said. My approach is to have response guidelines ready, not scripted responses. I know the general direction I want to go, and I write down the exact words after the fact in the activity sheet. This keeps the interaction flowing and preserves accuracy.

When This Method Breaks Down Completely

Wizard of Oz testing with activity sheets has real limitations that most guides don't emphasize enough. The primary bottleneck is the human responder. If your test requires reasoning, personalization, or dynamic content generation, you need someone who can produce that reliably in real time under observation. That's a narrow skill set and it doesn't scale. You'll burn out a good responder after two or three hours of active testing. The method also fails when you need to simulate systems with genuine state management. If your product maintains user history, preferences, or context across interactions, the manual approach introduces friction that distorts results. A user who has to wait twenty seconds for you to look up their previous conversation in a separate document is not a user you're accurately testing. There's also the observer effect to consider. When users know someone is behind the curtain, they behave differently. Some become more patient. Others become more critical. I've seen the same prototype get markedly different feedback depending on whether the responder was visibly present in the room or operating from another location. The activity sheet can note these conditions, but it can't control for them.

If your primary goal is validating whether users want a feature at all, a clickable prototype in Figma or a concierge test where you manually fulfill requests behind the scenes may be faster and just as informative. Activity sheets add real value when you need to understand interaction patterns, timing expectations, and conversation flow — not when you just need to know if the idea resonates. The analysis phase is another place where people underestimate the effort. A single hour of testing with a well-maintained activity sheet typically requires another hour to an hour and a half of clean-up and pattern extraction. The data is only as good as your ability to read it later, and messy sheets from rushed sessions are almost impossible to work with. Budget that time explicitly.

Wizard Of Oz Coloring Sheets
Wizard Of Oz Coloring Sheets