Reading Comprehension 4: A Practical Guide
If you are evaluating large language models or building a reading comprehension pipeline, you have probably run into Reading Comprehension 4 at some point. It is not a widely publicized benchmark, but it shows up in engineering workflows when teams need a straightforward pass/fail check on whether a model can extract answers from structured text. The name comes from an internal naming convention used by a handful of data science teams who wanted a versioned test suite without tying it to a commercial platform. People started calling it "RC4" in Slack channels and issue trackers, and the name stuck in certain corners of the NLP community. At its core, Reading Comprehension 4 is a benchmark-style evaluation harness. It takes a passage, a set of questions, and an answer key, then measures how well a model matches ground-truth answers. It is deliberately simple compared to something like SQuAD or RACE. There are no adversarial examples baked in by default. There are no multi-hop reasoning chains unless you build them yourself. What it does well is give you a baseline number quickly, so you can tell whether a model update made things better or worse without running a full evaluation suite. I have used it to sanity-check fine-tuning runs on domain-specific text. When my team started converting our internal documentation parser to use an open-weight model, I needed a quick way to measure whether the new weights actually understood the source material. RC4 gave me a result in about twenty minutes. A full MMLU or HellaSwag evaluation would have taken overnight on the same hardware.
How It Works in Practice
The evaluation format follows a narrow structure. You prepare a JSON or JSONL file where each entry contains a passage, a question, a list of possible answers (or a free-text ground truth), and sometimes metadata like passage source or difficulty tag. You then pipe your model outputs through a scoring script. The scoring script handles exact match, fuzzy match with string normalization, and optionally partial credit for answers that contain the ground truth as a substring. Here is a minimal setup if you are working in Python: Create an input file called rc4_questions.jsonl with lines like this:
{"passage": "...", "question": "...", "ground_truth": "...", "options": [...]} Then run your model against each question, collect the raw output strings, normalize them by lowercasing and stripping punctuation, and compare against the ground truth. The official scoring script that most people use is a short utility published alongside the benchmark corpus. It uses standard string operations and takes roughly one second to score a thousand items on a modern laptop.
Get the Full Details

Where RC4 Falls Apart
I should be honest about the limitations. This benchmark does not test reasoning. If your model can paraphrase and retrieve, it will score well. If your task requires multi-step logic, temporal understanding, or negation handling, RC4 will give you a misleadingly high number. I learned this the hard way when I evaluated a model that scored above 88 percent on RC4 but failed completely on a real customer support ticket where the answer required chaining three separate sentences from the passage. The dataset also skews heavily toward factual, extractive QA. Narrative comprehension, inference, and tone understanding are not represented. If your application involves any of those areas, you need a different benchmark in addition to RC4. Another issue is the answer format. The benchmark assumes answers appear as exact spans or short phrases. When models generate verbose responses, the normalization step can either over-penalize or over-credit depending on how you configure the scoring threshold. I ended up writing a custom scorer that splits on sentence boundaries and checks each clause separately instead of relying on the default script.
A Real Problem I Hit and How I Fixed It
During a recent project, I ran into a problem where the RC4 scoring script was returning zero for answers that were clearly correct. The issue turned out to be date formatting. The ground truth stored dates as "March 15, 2024" while the model output "15 March 2024". The string comparison script treated them as non-matching. I spent about an hour debugging before realizing the mismatch was purely cosmetic. The workaround was straightforward. I added a date normalization pass using the dateparser library before running the comparison. It converted every date string into ISO format regardless of the original format. After that change, my scores jumped by about four percentage points and looked more realistic. I also added a similar normalization pass for numerical values since the benchmark sometimes stored numbers as written-out words in the ground truth while models output digits.
How to Run a Clean Evaluation
If you want to use Reading Comprehension 4 properly, follow these steps. First, download the latest dataset split from the public repository. The official source is on Hugging Face under the name rc4_dataset. Second, set up a virtual environment with Python 3.10 or later and install the scoring dependencies. Third, format your model outputs into the required JSONL structure. Fourth, run the scoring script with the normalized flag enabled if your use case involves dates or numbers. The whole process takes roughly fifteen to thirty minutes depending on dataset size and your scoring configuration. I usually batch questions in groups of one hundred to keep memory usage low when working with larger language models. This also makes it easier to spot individual failures when something goes wrong.
When to Use RC4 and When to Move On
Use this benchmark when you need a fast, cheap check on whether a model can handle basic extraction from text. It works well for documentation parsers, FAQ systems, and simple knowledge retrieval pipelines. Do not use it as the sole metric for any production system that involves complex reasoning, time-sensitive answers, or ambiguous phrasing. If your task involves those harder categories, I recommend pairing RC4 with something like DROP or BoolQ. Together they cover extraction, numerical reasoning, and logical entailment. That combination gives you a more complete picture without requiring a full custom evaluation build.
Common Mistakes to Avoid
One mistake I see repeatedly is using RC4 results as a proxy for real-world accuracy without checking the error distribution. The benchmark has a narrow topic range, mostly centered on technical and factual passages. A model can score high on RC4 and still perform poorly on conversational or opinion-based text. Another mistake is skipping the normalization step. Raw string comparisons are too strict for most LLM outputs. You will get artificially low scores that do not reflect actual capability. Always run the normalization pass, even if it means writing a small custom script. The third mistake is treating a single RC4 score as definitive. Run multiple splits if you have the data. The benchmark provides train, dev, and test partitions, and performance can vary noticeably between them. I typically report the average across all three to get a stable estimate.
Getting the Data and Scripts
You can find the dataset and reference implementation at the standard public locations. Search for "rc4_dataset" on Hugging Face to pull the JSONL files. The scoring script is included in the repo under utils/. Clone the repository, install requirements, and run the example command from the README. It will guide you through the basic flow in about ten minutes. If you want the raw benchmark files directly, the download link is hosted on the project's GitHub releases page. The current stable release includes the full dataset, a sample submission file, and documentation covering edge cases. I keep a local copy of the latest release because the dataset does not change frequently, and having a fixed version prevents subtle scoring differences when you compare results across weeks or months.

Final Thoughts on Using RC4
Reading Comprehension 4 is a useful tool when you understand what it is and what it is not. It gives you a fast baseline for extraction-style QA. It will not save you from deeper comprehension problems. The best approach is to use it alongside other metrics, normalize your outputs carefully, and check the error cases before trusting the final number. That is how I evaluate now, and it has kept me from making expensive mistakes based on incomplete benchmarks.