Running stimulus preference assessments without losing your mind
I've been doing these assessments for long enough that I've stopped being surprised by most of the edge cases, but the paperwork side still trips people up more than the actual data collection. Most folks I see struggle not because the concept is hard, but because their recording system forces them to do mental math mid-trial instead of just writing numbers down. A stimulus preference assessment data sheet is fundamentally a structured log that tracks which items or activities an individual approaches, maintains contact with, avoids, or ignores across a series of structured presentations. In practice, that means you're recording approach latency, duration of engagement, and choice selections while minimizing cognitive load so you don't miss data because you were calculating percentages in your head. The typical formats you'll encounter include single-item presentations where you track time spent interacting with one item, paired-choice where two items are presented simultaneously and you record the first selection, and multiple-stimulus without replacement where you lay out all available items and remove each one after it's chosen. Each format feeds into different analysis paths.
Stimulus Preference Assessment Data Sheet
Here's what a functional one actually looks like. Your columns should be: Date, ID number, Trial Number, Item A, Item B, First Choice, Latency to First Contact (seconds), Duration of Engagement (seconds), Repeat Selection (yes/no), and Notes. That's it. Some protocols add an "Avoidance Recorded" column, but honestly, avoidance is usually captured in the notes field unless you're running a very rigid protocol that requires it. One thing I wish more people understood: the data sheet design should match your assessment format exactly. If you're doing a multiple-stimulus without replacement, your sheet needs columns for each item position and a sequence tracker. If you're doing paired-choice, you only need two item columns. I once watched someone try to use a single-item data sheet for a paired-choice session and end up with completely unusable data because they were writing down "Item C" in the wrong column three times before realizing what happened. That took forty-five minutes of cleanup work that could have been prevented with five minutes of upfront planning.
Setting it up and running the trial sequence
Before you bring the individual in, you need your stimulus set ready and your data sheet printed or loaded on whatever device you're using. A typical stimulus set ranges from five to ten items depending on the format and the individual's attention span. Write down your stimulus list separately first, before you generate the randomization sequence, so you have a paper trail if anything gets disputed later. When you start the trial, begin timing latency the moment the items are presented. Duration ends when contact ceases for more than five seconds or when the trial is terminated by the assessor for another reason. Record repeat selections separately because the fact that someone chose an item twice in a row versus once matters for your ranking calculation. I recommend writing raw seconds directly rather than converting to minutes or percentages on the sheet. You do the conversions during data analysis, not during data collection. The most common error I see is not standardizing the presentation procedure. If you place items on the left versus the right side inconsistently, or if you vary the distance between items, you're introducing side preference confounds that make the data unreliable. Keep the physical layout identical across trials. Use a template or marked positions on the table surface if you need to.
Get the Full Details

Edge cases that aren't mentioned in the manuals
Here's something I ran into repeatedly and had to figure out through experience rather than reading about it: when an individual picks up an item, immediately drops it, picks up a different item, then picks the first one back up again. On a standard data sheet, this looks like two choices between the same two items, which artificially inflates the apparent preference. The workaround I use is adding a code column where I mark "partial engagement" for any item that receives less than three seconds of sustained contact. Items below that threshold get flagged and aren't counted toward the primary ranking. This changes your numbers significantly more often than you'd expect. Another subtle issue is what I call the novelty bias in the first two to three trials of a new session. New items trigger approach behavior that isn't really preference, it's just novelty. I counter this by running a brief habituation period where items are presented without scoring for the first couple of rounds, then switching to scored trials. It adds about four minutes to the session but prevents the top-ranked items from being wrong items that were just novel.
Analysis basics and where the method falls apart
Once your trials are complete, ranking is straightforward. For paired-choice, count how many times each item was selected first and divide by total trials. For multiple-stimulus without replacement, count the number of times each item was chosen in any position across all trials. Higher frequency equals higher preference ranking. But here's the uncomfortable part that nobody likes to talk about: preference does not equal reinforcement. I've had multiple cases where the highest-ranked item during a stimulus preference assessment was completely non-functional as a reinforcer during actual teaching trials. The individual would approach it enthusiastically during the assessment and then immediately lose interest or use it in a stereotypic way that had zero functional value for increasing target behavior. This is why assessment results should always be followed by a reinforcer potential check, not assumed to be the final answer. SPA data is also unreliable for individuals with severe motor impairments that prevent clear approach or selection responses, or for those with limited discrimination abilities where the presentation format itself becomes the barrier. In those cases, consider a question-based preference assessment for verbal individuals or consult with a colleague about modified procedures. There's no shame in acknowledging the tool's limits and moving to an alternative.
What I actually keep on my desk
The data sheet I use is a simple table with rows for each trial and the columns I mentioned earlier. I print double-sided to save paper. For randomized ordering, I use a basic spreadsheet macro that shuffles the item list each session so position bias doesn't accumulate over weeks of testing. The macro outputs a printable sequence card I reference during the session without having to think about it. If you're looking for a template to adapt, most state autism spectrum disorder resource centers and university behavior analysis programs post free versions online. Look for ones that match your assessment format specifically. A single-item template used for paired-choice is just going to create more work for you than it saves. Make sure the columns align with how you plan to analyze the data before you commit to printing fifty sheets and realizing halfway through that you forgot to include a latency column. Record everything. I mean everything, including the weather in the room if it's noticeably affecting engagement, equipment malfunctions, or if the individual seemed unusually fatigued that day. Future you or anyone reviewing the data later will thank you for the context rather than having to guess why a particular session's results looked nothing like the previous twenty.
