Why Most People Mix Up Fever and Sparks
I spent three months debugging a pipeline that kept giving me inconsistent results, only to realize I had been using a Fever-format gold file against a Sparks reference scorer. They look the same from the outside. Both deal with factual claims and verification. Both produce scores that sit in roughly the same range. But the internals are completely different, and confusing them will quietly corrupt your evaluation numbers without any error message telling you why. Let me explain how they actually work before we get into the mechanics, because understanding the structure matters more than memorizing the commands. Fever stands for Fact Extraction and VERification. It is a benchmark built around a specific task: given a claim, you retrieve evidence from Wikipedia, then classify the claim as supported, refuted, or insufficient information for evidence. The original paper came out of Edinburgh and Sapiens AI researchers working on open-domain NLI. The dataset version 1.0 had roughly 185,000 claims. The official scorer uses a very particular logic where the evidence tier and the claim tier must align exactly, and partial matches do not count as supported even if the semantic content overlaps heavily.
Sparks is a different thing entirely. It is an evaluation framework used primarily by Sapiens AI for measuring performance across multiple dimensions on synthetic and real tasks. Sparks does not tie itself to one dataset. It builds benchmarks dynamically and scores them. When people say "Sparks score," they are usually referring to a normalized composite across reasoning, factual accuracy, code generation, and instruction following. It is not a dataset you download. It is a measurement system you plug into your model output. Here is the practical part. If you are running Fever evaluations, you need the fever package installed, the Wikipedia dump in the right format, and the claims file in JSONL with the claim, label, and evidence columns. The standard command looks like this: python -m fever.scorer --gold gold.jsonl --pred pred.jsonl
That produces class-level F1 and the overall score. The evidence F1 is what people actually care about, not just the claim classification. A model can get perfect claim labels and still retrieve garbage passages. The evidence score catches that. For Sparks, you do not run a scorer command. You define a benchmark config, pass your outputs through the appropriate metric functions, and Sparks aggregates. The config might look like this: A tasks block with factual_verification, reasoning, and code_generation entries. Each task points to its dataset and its metric. Then you run the benchmark runner and get back a JSON with per-task scores and a weighted average. The weighting is where people get burned. The default weights in Sparks favor reasoning and instruction following over pure factual recall, so a model that chews through math problems well will score higher than one that just retrieves Wikipedia passages accurately. That is by design, but it means Sparks scores and Fever scores are not comparable even when testing the same model.
Get the Full Details

I learned this the hard way. I had a model that scored 0.82 on Fever evidence F1 and a Sparks composite of 0.71. I assumed the Sparks score meant the model was underperforming. It was not. The Sparks benchmark I ran included a heavy code generation component that the Fever numbers never touched. The model was fine at retrieval and fact verification. It just struggled with Python test case generation. Once I separated the benchmarks and ran them independently, the discrepancy made sense. Another edge case that tripped me up involved Fever and case sensitivity in the evidence matching. The scorer does a substring match on the evidence strings. If your model outputs evidence with slightly different casing or punctuation from the gold standard, the evidence F1 drops even though the retrieval is semantically correct. I wrote a preprocessing step that lowercases and strips punctuation from both gold and predicted evidence before scoring. That recovered about 4 percent F1 on a tricky subset where the model was retrieving the right passages but formatting them inconsistently. For Sparks, the main pitfall is around the dynamic benchmark construction. If you are generating tasks on the fly, make sure your random seeds are fixed. Otherwise your scores will drift between runs and you will not be able to tell if a change in performance is real or just variance from different task samples. I set seed 42 across the data loader, the task sampler, and the model inference, and the scores stabilized within a 0.01 window across five runs.
If you need the Fever package, it is on GitHub under the fever project. Download the dev branch if you want the latest evidence retrieval utilities. For Sparks, grab the main branch from the Sapiens AI repository. Both are open source. Neither has a GUI. You work in terminal and config files. One thing worth noting that most tutorials skip: Fever claims are not independent. Many share evidence passages. When you aggregate the score, you are not measuring 185,000 independent decisions. You are measuring performance on a clustered dataset where some evidence blocks appear dozens of times. This inflates the apparent reliability of the score compared to a truly random claim set. I usually cross-reference Fever results with a held-out set from another benchmark like StrategyQA or HotpotQA to confirm that the performance generalizes. Sparks has a different limitation. The composite scores can mask catastrophic failure on a single subtask. A model might get 0.90 on factual verification and 0.30 on multi-step reasoning and still come out to a decent overall number if the weights lean toward the strong areas. Always look at the per-task breakdown before drawing conclusions from a Sparks composite.
The bottom line is that Fever measures one narrow thing very precisely and Sparks measures a broad thing with adjustable weights. They answer different questions. Use Fever when you need to know if your retrieval and verification pipeline is working. Use Sparks when you need to know if your model is broadly competent across multiple capability dimensions. Do not use them interchangeably and do not assume a high score on one predicts a high score on the other. I have run both on the same models across several iterations and the correlation between them has been moderate at best. That is not a flaw in either system. It is just that they were built for different purposes. Pick the right tool for what you are actually trying to measure and skip the rest.