What Wizard Of Oz Dorothy Actually Is
The Wizard Of Oz Dorothy framework is a human-in-the-loop simulation tool for building and testing conversational AI. You set it up so users think they are talking to an AI, but a human operator (or a scripted fallback) handles the responses behind a thin interface. The "Dorothy" part is the dashboard and orchestration layer that routes messages, manages operator seats, and logs conversations. I used it for a healthcare triage chatbot prototype back in 2023. The idea was simple: validate whether our intent classification logic held up under real patient language before investing in actual model training. We ran it for three weeks with two part-time operators. The framework itself is lightweight enough to deploy on a modest VPS, and it pairs well with something like a FastAPI backend and a PostgreSQL database for message history.
Wizard Of Oz Dorothy Setup And Deployment
The installation is straightforward if you already have Python 3.10+ and Docker on your machine. Clone the repository, copy the env template, and adjust the values. The critical settings are OPERATOR_POOL_SIZE, which controls how many simultaneous human responders can be active, and MESSAGE_TIMEOUT_SECONDS, which determines how long a message waits before falling back to a default response if no operator picks it up. Get those two wrong and your UX will look broken within hours. I learned that one the hard way. We set the timeout to 8 seconds because we thought users wouldn't notice a short wait. They noticed immediately. The first-day drop-off rate on the test page was 34 percent. People assumed the system was broken when messages just timed out. I bumped the timeout to 30 seconds, added a loading spinner that actually indicated "operator connecting," and cut the drop-off to under 9 percent. A spinner alone doesn't fix bad timing, but it gives people a visual reason to wait instead of clicking away.
How It Works Under The Hood
When a user sends a message, the system checks whether an operator is available in the pool. If one is, the message gets queued to their dashboard and the user sees a typing indicator. If no operator is free, it either queues the message or returns a fallback, depending on your FALLBACK_MODE setting. The fallback can be a canned message, a redirect to FAQ content, or simply nothing until an operator becomes available. The logging is where this framework actually earns its keep. Every exchange is stored with timestamps, operator ID, response latency, and the full message thread. You export this data later and use it to train your actual model, which is the whole point. Without clean structured logs, the human sim phase is just expensive guessing. One thing most people miss is that the operator dashboard needs to support response templates and quick-reply buttons. If your operators are manually typing every response during a long test, you will burn through your budget and your patience quickly. I built a small set of template variables into our operator view—things like [patient_name], [symptom_category], [urgency_level]—and it cut average response time from about 45 seconds down to roughly 12 seconds per message. That improvement alone made the difference between running the test for three weeks or having to shut it down after five days.
Get the Full Details

Common Pitfalls When Using This Approach
The biggest mistake I see is assuming the Wizard Of Oz phase will produce training data that is representative of real usage. It won't, not unless you design the test carefully. Human operators tend to be more patient, more detailed, and more empathetic than any model you will ever ship. Your training data will be quietly skewed toward longer, warmer responses, and when you try to fine-tune a model on it, the output will feel artificially formal and verbose compared to how actual users talk. To counter this, I started injecting random noise into operator instructions—telling them to give shorter answers sometimes, to use casual language, to occasionally ask clarifying questions instead of answering directly. It sounded silly but it produced training data that was noticeably closer to realistic user-operator exchanges. After fine-tuning, the model was about 20 percent more natural according to our blind evaluation panel, compared to models trained on unmodified operator logs. Another issue is session state management. The framework handles basic session tracking, but if your use case involves multi-turn context that depends on prior messages (which almost everything does), you need to make sure your operators have full conversation history visible before they respond. We had a bug where operators in a concurrent pool could only see the current message, not the previous turn. Users would ask follow-up questions that made no sense without context, operators would answer as if they understood, and the conversation would spiral into nonsense within two or three turns. The fix was enabling the SHOW_CONVERSATION_HISTORY flag and making sure each operator's view included the last ten messages by default.
When Wizard Of Oz Dorothy Won't Help You
This approach is not useful if you already have enough labeled data to train a baseline model. Running a human-in-the-loop simulation for three weeks costs money, coordination, and operational overhead. If you can get comparable signal from existing public datasets or synthetic data augmentation, skip the whole thing. The framework is expensive to run at scale. Each active operator represents a real cost, and you need at least two concurrent operators to avoid unacceptable response latency during anything beyond a tiny test. It also breaks down for highly technical domains where operators cannot reasonably answer questions without external tools or lookup systems. We tried extending our healthcare test to include medication interaction queries and realized within a day that no operator could answer those accurately without a drug database integration. Adding that integration would have been faster than training operators to handle the queries. For domains that require real-time lookup, consider building a RAG pipeline directly instead of simulating the interaction.
Where To Get Wizard Of Oz Dorothy
The project is open source and available on GitHub. You will find the README with installation instructions, a sample configuration file, and documentation on the operator dashboard. The repository is updated periodically, but it is not under heavy development, so check the commit history before committing to it. If your use case requires features that are not in the current release, you will likely need to fork and extend the framework yourself rather than waiting for a patch. I have been running a modified version of it for about two years now, and it still does the job for what it was designed to do. It is not a production-grade deployment tool. It is a prototyping and validation tool, and that distinction matters more than most people realize before they start using it.
