Setting Up and Running Trolley Problem Inc Questions in Production

Trolley Problem Inc Questions is a tool for generating, scaling, and scoring variants of the classic trolley problem across different ethical frameworks. It shows up most often in AI alignment work, ethics training pipelines, and decision-theory research. If you are looking to integrate it into a lab or a product, here is how it actually works and where it breaks down. Download it from their repository at trolleyproblem.inc/questions. Clone the repo, install the Python dependencies, and set your environment variables for the model provider you intend to use. The default config supports GPT-4-class APIs, Claude, and local models through Ollama. Pick one and set the API key in your shell before launching. The config file lives at config.yaml. The important keys are scenarios_per_batch, ethical_framework, and output_mode. I use jsonl for output_mode so I can stream results into a database without waiting for a full batch to finish. If you leave it on the default text mode, you will end up parsing messy output later and waste a couple of hours on regex cleanup.

How It Actually Runs

Once configured, the tool generates trolley problem variants by swapping parameters in a scenario template. It pulls from a built-in library of agent types, constraint sets, and outcome distributions, but you can also feed it your own YAML-based scenario definitions. The pipeline moves like this: load scenario definitions, sample variants, run each through the configured LLM or rule engine, parse the ethical verdict, and aggregate the results. The scoring layer is where most people get tripped up. Trolley Problem Inc Questions uses a weighted rubric that scores responses across utilitarian, deontological, virtue-ethics, and contractualist dimensions. The default weights assume you care equally about all four. You should not assume that. In my experience, the default weighting skews toward utilitarian outcomes because the underlying training data for most general-purpose LLMs has a utilitarian bias baked in. If you are doing rigorous analysis, override the weights before running anything. I set deontological weight to 1.5 and utilitarian to 0.7 for my own work, which changes the distribution of results noticeably.

A Real Problem I Hit and How I Fixed It

Last year I ran a batch of about 4,000 trolley problem variants through the standard config using a mid-tier API model. Halfway through, I noticed that roughly 18% of the responses were empty or contained only a single word. The tool does not validate response quality before moving to the next scenario, so those blank entries ended up in my final dataset and silently corrupted the scoring. The fix was straightforward but not obvious from the documentation: I added a post-processing script that filters any response shorter than 40 tokens and retries it with a higher temperature setting. I also switched to using the batch API endpoint instead of the streaming one, which gave me better retry logic and consistent rate limiting. That cut my failed scenario rate from about 18% down to under 3%. The first thing people get wrong is assuming the tool handles ambiguity gracefully. It does not. When a scenario has genuinely conflicting ethical signals, the model often picks a default path rather than flagging the ambiguity. This means your dataset will have fewer edge-case responses than reality would produce. If you need that edge-case signal, you have to explicitly construct ambiguous scenarios and run them separately. The second thing is the timeout behavior. The default timeout is set to 60 seconds per query, which seems generous until you are running scenarios through a quantized local model. I have seen entire batches stall because a single slow response held up the whole queue. I changed my config to use per-scenario timeouts with a separate worker pool, so a stuck scenario never blocks the others. This usually cuts wasted wall-clock time from something like 45 minutes per failed batch down to about 3 minutes of idle waiting.

Get the Full Details

Trolley Problem, Inc. Download - GameFabrique
Trolley Problem, Inc. Download - GameFabrique

When the Tool Does Not Work

Trolley Problem Inc Questions is not designed for real-time interactive use. It is a batch-oriented pipeline. If you are trying to build a live chatbot that responds to trolley problems on the fly, you will run into latency issues and inconsistent scoring across sessions. The tool also does not handle multi-agent scenarios well beyond what is in the built-in library. If you need to model complex stakeholder interactions, you are better off building on top of the scenario parser and writing your own agent simulation layer rather than trying to force the default engine to do something it was not built for. If you are using this for academic purposes, export your data in JSONL format and include your config weights in the metadata. Peer reviewers will ask for them, and if you did not log them, you will be rewriting your methods section from memory three weeks later. The tool does not automatically persist config snapshots, which is a design oversight that costs people more time than it should. There is also an optional evaluation module you can enable in the config. It runs a secondary model to cross-check verdicts against a held-out reference set. I found it useful for catching systematic bias in my own config, but it roughly doubles your compute cost. If you are working with tight grant budgets, skip it and do manual spot checks on 50 to 100 scenarios instead. That usually catches the same class of errors at a fraction of the cost.

The project is open source and the maintainers respond to issues on GitHub, but the documentation is sparse on advanced configuration. The README covers installation and a basic run. Anything beyond that requires reading the source code or asking in the issues thread. I learned most of what I know about the internal scoring logic by reading the Python files directly, not from any written guide. That is normal for tools at this level. It is worth your time if you plan to use it seriously. One last thing that is worth knowing: the tool caches previous scenario outputs by default to avoid re-running identical inputs. This is useful until your config changes slightly and you realize halfway through a long batch that you are getting cached results instead of fresh ones. Disable caching with the --no-cache flag whenever you are testing a new config. I wasted an entire afternoon debugging what I thought was a scoring bug before I remembered the cache was serving old results. If you are just starting out, run a small batch of 50 scenarios with verbose logging enabled and review the output before committing to a large run. It will save you a day of rework at least once in the first week. The tool is solid once you understand its limits, but it does not forgive blind trust in defaults.