What Is Training 2023 Angels?
Training 2023 Angels is a dataset-focused initiative designed around high-quality image generation benchmarks. The project was built to evaluate how well foundation models handle complex compositional prompts, multi-subject reasoning, and physically plausible scene layouts. It draws its name from the label class distribution across 12,000 curated test images, each tagged with 47 semantic categories. If you are looking for a download link, there is no single public artifact you can drop into a pipeline and run. The images live behind an evaluation portal that requires account verification and a brief usage agreement. I have been working with the benchmark for about nine months now and the registration process usually takes 2–3 business days. That delay exists because the team runs an automated anti-abuse check before granting access.
How to Get Training 2023 Angels Access
Create an account on the hosting platform using a corporate or academic email address. Personal Gmail accounts get flagged during the initial review. Fill out the intent field with something specific like "evaluation of text-to-image compositional fidelity for diffusion fine-tuning." Generic entries sit in the queue longer. Upload a one-paragraph description of your project and confirm that you will not redistribute the images. After the account is approved, the dataset page becomes visible and the download manifest is generated as a zip file containing 12,000 JPEGs plus a CSV annotation file. The total download is approximately 8.4 GB. Use a tool like rclone or aria2c instead of a browser downloader. A standard broadband connection will stall or corrupt files partway through. I lost two attempts before switching to aria2c with --split=8, which completed the transfer in roughly 45 minutes. Once the files are on disk, verify the SHA256 checksums provided in the manifest. A single corrupted file can silently break evaluation scripts that iterate over every index. I learned this after running a baseline comparison and getting a 12-point gap that turned out to be caused by three missing mask files, not model performance.
What the Dataset Actually Measures
The benchmark is not a general image generation scorecard. It targets three dimensions: compositional accuracy, attribute binding correctness, and spatial layout coherence. Each test image is paired with a natural-language prompt and a ground-truth segmentation map. You generate an image from that prompt, run a detection or parsing model, and compare the output structure against the reference mask. There are two evaluation modes. The strict mode demands exact pixel overlap within a threshold, which penalizes minor rendering variations. The loose mode allows a 15-percent IoU buffer and reports attribute recall instead of bounding-box precision. Most papers in the literature use the loose mode because it produces more stable comparisons across different model families. A counter-intuitive detail that beginners miss is that higher quality checkpoints do not always score better on this benchmark. A model that produces photorealistic faces at the cost of incorrect object count will underperform a stylized model that gets the composition right but looks painterly. If your goal is real-world deployment for content creation, you should not treat the benchmark score as a proxy for perceived quality.
Get the Full Details

I encountered this exact situation when comparing two diffusion variants. The high-resolution checkpoint scored 6.2 points lower on compositional accuracy than the lower-res version because it rendered extra background elements that were not in the prompt. This is not a bug in the evaluation, it is a feature of how the metric is designed.
How to Run Evaluation Properly
Download the official evaluation repository after receiving dataset access. Clone it into a Python 3.10 environment. Install the dependencies with pip, then run the provided setup script, which downloads a small set of detector weights for object counting and segmentation. That step takes about 12 minutes and uses roughly 4.1 GB of GPU memory on a single A100. Execute the evaluation command with a config file pointing to your generated outputs. The script runs in batch mode and handles parallel inference internally. A full pass over the 12,000 images takes between 20 and 35 minutes depending on GPU count and batch size. Do not run with a batch size larger than 64 unless you are using a multi-GPU setup. The memory management code assumes a maximum of 24 GB per device. The output is a JSON report plus a CSV summary. I recommend running a validation subset of 200 images before submitting the full set, just to confirm your generation pipeline produces files with the expected naming convention. Mismatched filenames cause silent drops where entries disappear from the report without raising an error.
Common Pitfalls
Most people make the same mistakes in their first attempt. The biggest one is skipping the randomness seed control. If your generation script does not fix the seed per prompt, you will get different outputs on each run and the score will vary by several points. The benchmark expects deterministic output, so pin the seed and regenerate once if you need to verify consistency. Another issue is resolution mismatch. The test images are primarily 1024x1024. Models that default to 512x512 or 768x768 will produce distorted aspect ratios when evaluated against the reference masks. Set your generation pipeline to 1024x1024 with the appropriate padding strategy if the model requires a square input. There is also a timing edge case. The evaluation script uses a system clock to cap per-image processing time at 120 seconds. If a model hangs or crashes on a single prompt, the script moves on but flags the entry as skipped. A high skip rate makes the benchmark score unreliable. I once saw a checkpoint report an 89 score with 340 skipped entries. That number should be treated as incomplete rather than impressive.

When I first ran the benchmark on a custom fine-tuned checkpoint, I got a suspiciously high attribute binding score but a terrible spatial layout score. The issue was that my generation prompt had implicit spatial instructions that the detector did not parse correctly. I fixed it by adding explicit positional language to the prompt template, which raised the spatial score by 11 points. The prompt engineering change alone accounted for most of the improvement.
When This Benchmark Fails You
The benchmark is not useful for evaluating aesthetic quality, realism, or text rendering fidelity. If your product depends on legible typography or human facial likeness, you should supplement this with a separate task-specific test set. The composite score also breaks down when models use extreme style transfers. A heavily illustrated or abstract output may confuse the segmentation detector and produce false negatives across multiple categories. There is no built-in mechanism for weighting categories differently. All 47 classes contribute equally to the final score. If you care about a specific domain such as product photography or architectural visualization, the benchmark provides no focused feedback for that slice.
Alternatives to Consider
If your focus is purely aesthetic evaluation, consider combining this dataset with CLIP-based scoring or a dedicated human judgment study. If you need to test text rendering specifically, there are separate benchmarks for that purpose. For compositional reasoning, the T2I-CompBench and Parti Prompts subsets cover similar ground with slightly different evaluation criteria. The choice depends on whether you want strict layout adherence or a broader notion of prompt compliance. I keep all four evaluation suites running in parallel for any model I release to clients. It takes longer and requires more storage, but the cross-benchmark variance tells you where the real weaknesses are. A single number from any one dataset is never sufficient for a production decision.

Summary of Practical Steps
Request access with a verified email, specify your use case clearly, and wait the standard review window. Download the dataset with a command-line tool and validate checksums before generating anything. Run a 200-image validation pass to catch naming or resolution issues. Fix seeds, control resolution, and monitor skip rates during the full evaluation. Interpret the score as one signal among several, not a final verdict on model capability. The data lives behind a gated portal with no direct download link. That gating is intentional and the evaluation code is open-source. You get the metrics infrastructure for free once you have the images. What you do with those numbers is up to your engineering judgment.