Bot 2 Scoring Manual: What It Actually Does

The Bot 2 Scoring Manual is a set of rules for evaluating automated agents across a range of benchmark tasks. It assigns weighted scores based on correctness, latency, resource usage, and format compliance. Most people treat it like a black box. It isn't. It runs on a clear rubric, but the rubric has enough edge cases that you will waste a weekend if you don't read the fine print first. I went through this twice across two different projects. The first time I skimmed the manual and submitted scores that were later rejected for "misalignment with weighting rules." I had missed the section on partial credit for structural validity. That cost me three days of appeals and re-runs.

Bot 2 Scoring Manual — Download and Usage Guide

You can find the latest version on the official repository at https://bot2scoring.org/manual. Grab the PDF and the companion Python reference implementation. They don't always agree perfectly, but they point you in the same direction. I keep both open side by side when I'm debugging score mismatches. The manual describes four primary scoring dimensions:

  • Task correctness — final answer matches the ground truth or satisfies the acceptance criteria.
  • Format compliance — structure, tags, ordering, and required fields match the spec.
  • Latency — response time relative to the baseline threshold for each task type.
  • Resource efficiency — token usage, compute steps, and memory footprint within allowed ranges.

Each dimension has its own weight, which changes depending on the task category. A code generation task weights correctness and format heavily, while a reasoning task weights latency less and allocates more to intermediate step validity. Start by isolating one task batch. The manual expects you to run the bot against a closed set of prompts with precomputed ground truth. You feed the prompts through the reference evaluator, capture outputs, then run the scoring script. Don't skip the validation pass. The manual rejects any submission that hasn't been pre-validated against the schema checker. Here's the workflow I use now:

Get the Full Details

Scoring and Interpreting the BOT-2 - YouTube
Scoring and Interpreting the BOT-2 - YouTube
  1. Export the prompt batch from the benchmark tool.
  2. Run the bot and capture raw outputs with timestamps and resource logs.
  3. Validate the output format against the schema provided in Appendix C.
  4. Run the reference scoring script with the matching task category flags.
  5. Review the score breakdown and flag any dimension that looks off.
  6. Compare against the manual's acceptance thresholds before submitting.

This takes about 15 to 20 minutes for a standard batch of 50 prompts. The bottleneck is usually step three, where format validation catches something like a missing closing tag or an out-of-order field. Those errors don't crash the scorer, but they reduce the format compliance score to zero, which tanks the overall result. Pitfall one: assuming the latency baseline is universal. It isn't. Each task type has its own baseline, and the manual publishes them in a separate table. I once used the reasoning baseline for a code generation task and got penalized heavily because the bot appeared artificially slow. The fix was just to switch the flag to the correct baseline category before running the scorer. Pitfall two: ignoring partial credit rules. The manual gives partial points for structurally valid outputs even when the answer is wrong. This matters more than people expect. A bot that consistently formats responses correctly but occasionally misses the target still scores decently because the manual rewards reliability in structure.

Here's a specific edge case I hit. I was scoring a bot that returned correct answers but wrapped them in slightly non-standard JSON keys. The output passed content validation but failed strict schema validation. The manual's rule for this situation says you get a reduced format score, not zero, as long as the content is recoverable. I had to write a small normalization script that mapped the bot's non-standard keys to the expected ones before running the scorer. That changed my score from 0.41 to 0.68. It sounds small, but it crosses the threshold for passing in many task categories. The normalization approach works like this:

  • Read the output schema from Appendix C.
  • Build a key mapping table from observed bot outputs.
  • Apply the mapping before validation.
  • Log every transformation so you can explain it in your submission.

Don't skip the logging. Reviewers check the transformation log to verify you didn't alter ground truth or fabricate compliance. One thing that trips people up is the interaction between resource efficiency and latency. The manual doesn't treat them independently when computing the combined dimension score. If a bot is fast but uses excessive tokens, the resource penalty can outweigh the latency gain. I saw a bot that looked great on speed but scored poorly overall because it burned tokens on verbose chain-of-thought steps. The fix was to trim the output and keep the reasoning compact. Another nuance is how the manual handles repeated identical outputs. If a bot returns the same response for multiple distinct prompts, the scorer flags it as low diversity. The penalty is small per instance, but it adds up. I learned this the hard way when a caching layer in my test runner accidentally reused outputs across prompts. The fix was to disable caching during scoring runs and verify prompt uniqueness beforehand.

BOT-2 Gross Motor Scoring Sheet | PDF
BOT-2 Gross Motor Scoring Sheet | PDF

When the Bot 2 Scoring Manual Falls Short

The manual is reliable for structured benchmark tasks, but it struggles with open-ended creative generation. The format compliance rules are too rigid for free-form outputs, and the resource penalties don't account well for inherently variable workloads. If you're evaluating a bot that does creative writing, code design, or conversational reasoning without strict structure, the scores can feel misleading. In those cases, I recommend pairing the manual with a secondary qualitative review. Run the automated scores first, then have a human grader assess a sample of outputs against criteria like coherence, relevance, and usefulness. The manual alone won't give you a complete picture for non-deterministic tasks. There's also the issue of baseline updates. The latency and resource baselines shift occasionally as new hardware and models arrive. If you're comparing scores across versions of the manual, make sure you're using the same baseline edition. Mixing editions introduces invisible drift that looks like performance regression but is just a numbers game.

If you're looking for something less rigid, the Community Evaluation Framework is worth exploring. It leans more toward human scoring and contextual metrics, which works better for open-ended bots. It's slower and less repeatable, but it captures nuance the manual misses. Use both when you can.

Practical Tips for Running the Score

Run the scoring script in a clean environment. I pin dependencies to the versions listed in the manual's requirements file. Mixing versions causes silent mismatches in how timestamps and resource metrics are parsed. Save the raw outputs before scoring. The manual allows resubmission within a short window, but only if you can prove the original outputs are intact. I keep an archive of every batch with checksums. It sounds excessive until you need to prove that a score drop wasn't caused by output corruption. Review the dimension breakdown, not just the final number. A high overall score can hide a zero in one dimension. That zero matters more in audits and appeals than people realize. The manual reviewers look at the breakdown first, and a suspicious pattern there is enough to trigger a manual re-review of your entire submission.

Bot 2 Manual Supplemet | PDF
Bot 2 Manual Supplemet | PDF

Keep notes on every parameter you change. The manual requires a submission log, and incomplete logs get rejected. I use a simple text file with timestamps, parameter values, and notes on any manual overrides. It takes two minutes to maintain and saves hours when someone asks why your score changed after an update.

Final Thoughts on Using the Manual

The Bot 2 Scoring Manual is thorough but unforgiving. It rewards careful reading and punishes assumptions. If you treat it like a rigid checker, you'll get reasonable scores for standard tasks. If you ignore the edge cases, you'll lose time fixing avoidable issues. My recommendation is to run a small pilot batch before committing to a full evaluation. The pilot reveals which task categories your bot struggles with and highlights any scoring quirks you'll need to work around. It also surfaces formatting issues early, so you don't waste a full batch run on a fixable problem. Download the manual, follow the workflow, and keep good records. The process isn't glamorous, but it produces scores that hold up under scrutiny. That's usually the goal anyway.