The Reality of Using Math Test Candy
Most people installing this package just want it to run and stop thinking about it. The documentation is thin on the actual mechanics. You download it, you point it at your problem set, and you expect a score. It usually works for standard arithmetic. It breaks in obvious ways when you try to use it for anything involving algebra or geometry. I spent three weeks trying to get it to handle multi-step equations correctly. The output was fine for addition and subtraction, but as soon as variables entered the mix, the parser started dropping coefficients. I ended up writing a pre-processing script to normalize the input into a format the core engine could actually read. That added about ten minutes to my workflow, but it was better than manually grading everything by hand.
Math Test Candy Workflow
The interface looks like a game, which is the main selling point for students. You generate a test, it serves problems one at a time, and it tracks response time alongside correctness. Response time matters more than you might think. The built-in analytics flag answers that take longer than expected even if they are correct. It assumes hesitation means guessing or confusion. You can export results to CSV. The columns include question ID, answer value, user answer, time taken, and a confidence score if you enable the hint system. Do not skip enabling the hint system. The raw confidence score is useless without it because the engine needs to know whether the student asked for help before marking a final answer. I once had a student who finished a five-problem set in forty seconds. Every answer was right. The system flagged it as suspicious activity. It turns out he had memorized the order of operations shortcuts instead of solving the problems. The software could not distinguish between speed due to fluency and speed due to memorization. I had to switch to timed versions with randomized problem sets to stop that loophole.
Where This Actually Fails
The licensing model is the first friction point. Each installation requires a unique key tied to a hardware ID. If you move the program to a different machine, it stops working. You have to contact support to unlock it. Support tickets get answered in two to four business days. If you are running a classroom with ten laptops, plan for at least one broken install per semester. Another issue is the export format. It writes UTF-8 with a BOM. Most spreadsheets handle that fine, but if you pipe the output into a Python script or a database, you will need to strip the BOM first. I spent an hour debugging a script that kept failing to parse the header row. The first three bytes were invisible characters that broke every string comparison. Offline mode works, but only for the current test session. If the internet drops mid-test, the answers queue locally and sync when the connection returns. The queue does not persist across restarts. If the power goes out, that session is gone. There is no auto-save checkpoint feature. I recommend running tests on stable networks or allowing extra time for students who might experience drops.
Get the Full Details

Practical Debugging Tips
Enable verbose logging in the settings panel. The default log level hides the parsing errors that cause silent failures. You can find the log file in your user directory under AppData or .local/share depending on your OS. Search for the string "parser warning" to see where the input got mangled. If you see a mismatch between the number of problems generated and the number recorded in the results file, check your problem source. Some formats allow optional hints that the engine might ignore, causing a desync. Re-export the problem set with hints disabled and see if the numbers align. For advanced classes, disable the timer. The default timer setting penalizes students who think slowly but correctly. It skews the analytics and makes the confidence score unreliable. You can turn off timing in the global preferences. The trade-off is that you lose the response time column in your exports, but you gain accuracy in the correctness metrics.
I prefer using Math Test Candy for quick formative assessments rather than high-stakes exams. It is fast to set up and gives immediate feedback. It is not suitable for standardized testing because the randomization and security features are minimal. Students can share screenshots of the interface and look up the answer keys posted by others online. The engine does not prevent collaboration. If you need secure exams, pair it with a remote proctoring tool or a lockdown browser. That adds complexity and defeats the simplicity of the system. For regular practice and low-pressure quizzes, it works well once you adjust the settings to your actual use case. Most failures come from using the defaults without reading the manual. The manual is short, but it covers the pitfalls. Read it before you distribute the test.