Setting Up a Course Psychology Experiment That Doesn't Fall Apart

Most college psychology courses require students to run an experiment. The problem is the gap between what the textbook says should happen and what actually happens when you are standing in front of a room full of people who just want to get out of there. I have watched dozens of student projects derail because someone tried to run a two-hour reaction time study without recruiting enough participants first, or because the ethics board sent the proposal back with fourteen pages of requested revisions. The process is manageable if you treat it like a project plan instead of an afterthought. If you are looking for something with real traction but minimal equipment, here are a few designs that consistently work well at the undergraduate level. The Stroop task with a social media twist — Instead of the standard color-word conflict, give participants a list of words where some are socially charged (words related to social media notifications, likes, follower counts) and others are neutral. Measure reaction time and error rate. It is cheap, it runs in fifteen minutes per person, and it gives you clean data if you use a tool like jsPsych or a well-configured Qualtrics survey with exact millisecond timing. The limitation is that online platforms can introduce timing variance, so if your university has access to any lab computer setup, use it.

A simple bystander intervention study using a written scenario — Describe a situation where someone needs help and vary the number of witnesses mentioned in the vignette. Ask participants how likely they would be to intervene on a Likert scale. This avoids the awkwardness of staging a real emergency and still demonstrates the core principle. I once had a student who tried to do this as a live roleplay with other students playing bystanders, and the results were completely unusable because the actors broke character and laughed. Stick to the written format unless you have a trained confederate and IRB approval for deception, which most undergrad projects will not get. Memory recall with emotional versus neutral word pairs — Show participants a list of words split into emotional and neutral categories, then give them a distractor task for two minutes before asking them to write down everything they remember. The emotional words should show a recall advantage. The tricky part is controlling for word frequency and concreteness, which most students skip. Without that control, your confound is not emotion, it is how commonly the words appear in everyday language. Use a corpus tool to select matched word sets before you ever run a single participant. A brief implicit association style task using image preferences — Pair positive and negative images with concepts like nature versus technology and measure response time to categorize them correctly. It is essentially a simplified IAT, easier to justify ethically, and still teaches the mechanics of implicit measurement. You will need about forty to sixty trials per block to get reliable latency data, and you should run a pilot with at least five people first to check that your stimulus images are actually loading on time across different browsers.

The Setup Steps That Actually Matter

Before you touch a single participant, you need your protocol locked down on paper. Write out the independent variable, the dependent variable, the exact procedure, the inclusion and exclusion criteria, and the debrief script. If your university uses an online IRB module, most proposals for student experiments are exempt or expedited, but they still need to see that you have thought through confidentiality and the right to withdraw. I once submitted a proposal that assumed participants would self-select for a "stress and memory" study without clarifying that I would not be exposing anyone to actual stressors. The reviewer flagged it as potentially misleading. I rewrote the consent form to explicitly state that the task was cognitive and non-stressful, and it went through on the next submission. Participant recruitment is where most projects stall. Posters in the psych building still work, but the conversion rate is low. The faster route is signing up through your department's participation pool if your school has one, or offering course extra credit through the learning management system. Budget about ten to fifteen minutes of your time per recruited participant for check-in and consent. If you are running a twenty-minute experiment, plan for participants to show up at 1.3 times your target sample size, because no-show rates in student studies usually land between twenty and thirty percent.

Common Mistakes That Ruin Your Data Before It Exists

The biggest issue I see is underpowered samples. A lot of students aim for twenty people across two conditions, which gives you almost no statistical power to detect anything. Plan for at least thirty participants per condition if your design is between-subjects, or twenty-five per condition if it is within-subjects, assuming a medium effect size. If that is not feasible due to class size, switch to a within-subjects design where every participant serves as their own control. It halves your participant requirement and increases sensitivity at the same time. Another mistake is poor instruction scripting. If you tell participant A "rate how much you agree" and participant B "tell us your opinion," the wording difference introduces a subtle priming effect. Write out every instruction verbatim and read it to your first five participants to make sure it sounds the same each time. If you notice yourself paraphrasing mid-study, stop and fix the script. Data cleaning is rarely discussed but it is essential. Set up your data file before the first participant arrives. Define your outlier rule upfront — I use the standard approach of removing any reaction time below 200 milliseconds or above three standard deviations from the participant mean. Document this decision in your methods section. Changing the outlier rule after seeing the results is a red flag for reviewers, and it can make your analysis look like p-hacking even if your intentions were clean.

Running the Session

Keep the environment consistent. Same room, same desk setup, same lighting. If you are running online, acknowledge the variability and build it into your limitations section. A quiet room matters more than anything else for reaction time tasks. Background noise in a hallway will add random variance that looks like signal later when you run your analysis. I learned this the hard way during a spring semester when I ran half my Stroop participants in a study room and half in a corner of the library lounge. The within-condition variance was so high that the effect disappeared entirely, even though the between-condition difference looked the same. Moving the rest to the study room brought the effect back, but by then I had already lost a day of scheduling. Debriefing is not optional and it should not be an afterthought. Tell participants the true purpose, explain why any deception was necessary, and answer their questions. If you used a cover story, make sure they leave understanding what actually happened. This is both an ethical requirement and a way to prevent demand characteristics from contaminating future participants who might overhear what the study was really about.

Analysis and Reporting

Run your statistics in whatever software your program uses. R, SPSS, JASP, or even Excel will work depending on the complexity of your design. Report your effect sizes alongside p-values. A significant result with a tiny effect size tells a different story than one with a large effect size, and reviewers notice when only the p-value is reported. If your results are non-significant, say so and discuss possible reasons — small sample, weak manipulation, high variance — rather than burying the outcome. The final report should follow standard empirical paper structure, but do not pad the literature review with generic summaries. Focus on the theories directly relevant to your manipulation and cite primary sources where possible rather than secondhand summaries in textbooks. One page of focused, well-chosen references is worth more than three pages of overview fluff.

Tools Worth Using

For online studies, jsPsych is free and gives you precise timing control if you host it on your own server or through a university platform. Qualtrics is easier to set up if your school already has a license, but its timing precision is limited for reaction time work. Use it for survey-based designs and stick to jsPsych or OpenSesame for behavioral tasks. OpenSesame is free, runs locally, and exports directly to Qualtrics if you need to switch halfway through. None of these require programming experience to use for basic experiments, though knowing how to write a simple loop in JavaScript saves you hours when you are trying to randomize trial order. If you want a ready-made resource directory for experiment templates and stimuli libraries, search for the Open Science Framework and the Psychopy template repository. Both have collections of student-friendly protocols you can adapt rather than building from scratch.

When an Idea Just Will Not Work

Sometimes you pick a design and realize halfway through piloting that it is fundamentally flawed. A participant expectation manipulation that requires deception, for example, may need IRB review that takes longer than your semester allows. A longitudinal design tracking changes over weeks is unrealistic if your timeline is eight weeks. In those cases, pivot early. Swapping to a cross-sectional or single-session design is almost always better than submitting incomplete data. I had a student who spent three weeks trying to recruit a longitudinal sample for a sleep and decision-making study, only to lose half the participants between sessions. She switched to a single-session design measuring sleep quality via self-report the night before and decision-making tasks the next morning, and her final analysis was actually stronger because attrition bias was no longer a problem. Student experiments are not supposed to be groundbreaking. They are supposed to teach you how research actually works — the paperwork, the unexpected failures, the data cleaning, the honest reporting. If you plan ahead, keep the design tight, and accept that something will go wrong, the project will finish on time and the results will be usable whether they support your hypothesis or not.