What actually happens when you try to teach someone new

You sit down with a client or student. You have a pile of potential reinforcers — snacks, toys, breaks, activities. You pick what you think they will like most, hand them one, and immediately start teaching. If it works, great. If it doesn't, you move to the next one, and the next, and by session four you have spent ninety minutes figuring out what might work while also trying to teach them something. This is how reinforcer assessments go wrong, and it is also why people confuse the two main types of assessment you should be running. The core difference between a preference assessment and a reinforcer assessment is that one measures what someone seems to want, and the other measures what actually changes their behavior. Preference tells you interest. Reinforcement tells you power. They overlap sometimes, but they are not the same thing, and treating them as interchangeable is one of the most common mistakes I see in practice.

Preference Assessment Vs Reinforcer Assessment

A preference assessment is a structured way to figure out which items or activities an individual shows interest in. The most straightforward version is a forced-choice trial. You put two items side by side. The person picks one. You record which one they picked. You rotate through all the items. Usually you do multiple rounds and tally the counts. Items picked most often are labeled high preference. Items picked least often are low preference or non-preferred. There are several formal protocols for this. Massey’s paired-stimulus preference assessment is the most commonly used. It uses a grid where you present every item against every other item, usually in pairs, and record selections. That can get long quickly. A single stimulus presentation method, sometimes called the free operant or one-item-at-a-time approach, just presents one item and records approach or engagement time. Brief sensory reinforcement assessment is faster but more limited in scope. Each has tradeoffs. A reinforcer assessment is different. Here you do not just watch someone pick something. You set up a contingency. The person has to work for the item. You give them the item contingent on a specific response. Then you measure whether the likelihood of that response increases over time. That increase in response rate is what makes the item a reinforcer. If the response rate does not increase, the item is not a reinforcer for that person, regardless of how many times they chose it during a preference assessment.

I ran into this exact gap a few years ago with a client who scored very high on a paired-stimulus preference assessment for fidget spinners. He picked them consistently. We started using them as rewards for completed trials. His response rate dropped below baseline. Nothing about that made sense on paper. The problem was that the fidget spinner was a sensory reinforcer that competed with the task demand rather than supporting it. Once he had it, he did not need to do anything else to access it, so engagement in the target behavior declined. We switched to an edible and worked the contingency properly. Response rate went up. That experience taught me to always verify preference data with a reinforcement check.

Get the Full Details

Preference Assessment and Reinforcer Surveys Pack for Special Education and ABA
Preference Assessment and Reinforcer Surveys Pack for Special Education and ABA

How to run a preference assessment properly

Start by building a stimulus list. The list should come from caregiver interview, direct observation, and a review of what the person already accesses independently. Do not skip the interview step. I have seen assessors skip it and waste forty-five minutes testing items the person clearly does not interact with. It feels efficient until you realize you missed the top three reinforcers because nobody thought to ask. For a paired-stimulus assessment, arrange your items on a tray or table. Present two items at a time. Let the person choose. Record the choice. Rotate positions left and right to avoid side bias. Do not provide the item unless it will also be used in a trial structure later. The goal here is selection ranking, not delivery. Run through the full grid. With ten items, that is roughly forty-five pairs. In a clinical setting with a participant who engages quickly, that takes about twenty to thirty minutes. With someone who needs prompts or has shorter attention spans, it can take forty-five to sixty minutes. Plan accordingly. If you are working with an adolescent or adult who reads or communicates well verbally, a survivor or single-item method may give you equivalent information in half the time.

At the end, rank the items. High preference items are those selected most frequently. These become your candidates for reinforcement. But again, that is only the first step.

How to run a reinforcer assessment

This is where you test whether those ranked items actually function as reinforcers. The standard approach is a multiple-opportunity reinforcement assessment. You set up three conditions: no reinforcement, small reinforcement, and large reinforcement. In each condition, the person performs a defined number of trials or responses. The reinforcement condition delivers the item contingent on completion. You record the response rate in each condition. If the response rate is significantly higher when the item is available as a reinforcer compared to when it is not, the item is validated as a reinforcer. If the rate is the same across conditions, the item is not a reinforcer for that individual in that context, regardless of preference ranking. The setup usually looks like this. You pick the top-ranked items from your preference assessment. Test them one at a time. For each item, you do about five to ten trials per condition. The no-reinforcement condition means the person works through the trials without getting the item. The small-reinforcement condition gives a short delivery. The large-reinforcement condition gives a longer or more frequent delivery. You compare the rates. The mathematical analysis is typically a simple visual inspection of response counts across conditions, though some programs require a minimum percentage difference, often around twenty percent, to classify an item as a reinforcer.

Preference and Reinforcer Assessment - YouTube
Preference and Reinforcer Assessment - YouTube

In practice, a full reinforcer assessment with five candidate items takes about forty-five to seventy-five minutes. Not every item survives. It is normal for three or four of your top-ranked preference items to fail the reinforcer test. That is not failure on your part. That is the system working as intended.

Practical things that go wrong

One issue I encounter constantly is satiation during the assessment itself. If you present high-value food items back to back without any break, the person will eat through motivation before you finish the trials. I started spacing food items apart and inserting brief non-food trials between them. This kept response rates more stable across the assessment. It also made the data more useful because you were measuring true preference rather than consumption speed. Another problem is that preference assessments assume the person can communicate a choice. If the individual points poorly or avoids eye contact, selection data becomes unreliable. I had a client who consistently selected the item on the right side of the tray across every trial. After switching to a random position schedule and adding a third option, the pattern broke. The real issue was positional responding, not preference. This is why you should always check for side bias when reviewing preference assessment data. There is also the issue of items that function as punishers rather than reinforcers in certain contexts. I once tested a loud musical toy that scored medium on preference. During the reinforcer assessment, the noise increased vocal stereotypy in the participant to the point where task engagement nearly stopped. The toy was a sensory distractor, not a behavioral reinforcer. It felt counterintuitive because the person clearly enjoyed the item, but enjoyment does not equal reinforcement for a target behavior.

Common pitfalls and what to do instead

The biggest pitfall is assuming that preference data alone is enough to select reinforcers. It is not. A preference assessment gives you a ranked list. A reinforcer assessment validates which items on that list actually increase behavior. Running only the preference assessment and skipping the validation step is like ordering food based on the menu description and never tasting it. You might get something you like, but you might also get something you do not, and you just wasted time and opportunity cost. Another pitfall is using only one type of assessment. Some clinicians rely exclusively on caregiver reports. Some rely exclusively on paired-stimulus data. Neither approach alone is sufficient for building a robust reinforcement program. A mixed method that combines interview data, preference assessment, and reinforcer assessment produces the most reliable results. It takes longer upfront, but it reduces the back-and-forth later when items fail to work in actual teaching sessions. A smaller but frequent mistake is testing only food items. Non-food reinforcers are often overlooked. Activities, tactile items, social interactions, and auditory stimuli can be equally effective, especially for individuals who are not food-motivated or who have restricted eating patterns. I started building separate preference lists for tangible items and activities. This broadened the reinforcer menu significantly and made it easier to maintain motivation across longer teaching sessions.

Preference Assessment and Reinforcer Surveys Pack for Special Education and ABA | Life skills ...
Preference Assessment and Reinforcer Surveys Pack for Special Education and ABA | Life skills ...

When these methods do not work

Preference and reinforcer assessments are not universal solutions. They rely on the individual having some level of discriminatory responding and the ability to select or approach items. If a person cannot point, look, or physically reach toward stimuli, the data you collect will be noisy or meaningless. In those cases, you shift toward indirect methods or naturalistic observation. Track what the person engages with independently over several days. Use that data to guide item selection, then test those items in a contingency format with whatever form of responding the person can manage. There is also the issue of generalized reinforcers. Some items, like tokens or preferred activities, function across many contexts and do not require a full reinforcer assessment every time you introduce them. That does not mean you skip validation entirely, but it does mean you can be less rigorous on items that have already demonstrated reinforcing properties in similar settings. The key is tracking whether the item continues to function as expected over time, because reinforcement value can change. Finally, these assessments assume a relatively controlled environment. If you are running them in a chaotic classroom with constant interruptions, the data will be unreliable. I learned this the hard way when a district pushed for preference assessments to be done on the general education floor. The response rates were inconsistent and the rankings made no sense. We moved the assessments to a quiet room and repeated them. The rankings changed completely. The environment matters more than people usually account for.

What I would do differently if starting over

I would validate every top-ranked item before building a teaching schedule around it. I wasted about three months early in my career using items that looked good on paper and failed in practice. I also would have built a larger stimulus repertoire from the beginning instead of relying on a small set of known reinforcers. Broadening the pool early made subsequent teaching sessions more flexible and reduced the impact of satiation when it inevitably showed up. The process is not glamorous. It involves sitting through repetitive trials, recording data by hand or in a spreadsheet, and watching your best guesses get corrected by actual behavioral outcomes. That correction process is the point. The assessments exist to replace assumptions with evidence. Preference Assessment Vs Reinforcer Assessment is not a debate about which is better. It is a reminder that measuring interest and measuring behavioral impact are two separate steps, and skipping the second step is where most programs stall.