Field Notes on Watching Animals Do What They Do
You spend a lot of time sitting still, waiting for something interesting to happen, and then what happens is usually pretty boring. That is the reality of ethology. The popular image involves researchers in khaki chasing leopards through the Serengeti, but most animal behavior work is slow, repetitive, and occasionally wet or insect-ridden in ways you did not expect. I want to talk about how the actual process works when you are trying to collect clean behavioral data from wild or semi-wild animals, because the gap between what textbooks say and what happens in the field is where most people get tripped up.
Getting Started with Science Studying Animal Behavior
The first thing you need is a behavioral coding scheme, also called an ethogram. This is a catalog of every distinct behavior you plan to record, with clear operational definitions. If your ethogram says "grooming" includes any instance where the animal touches its own fur with a paw, mouth, or beak, you need to decide upfront whether mutual grooming between two individuals counts as a separate category or falls under the same definition. Getting this right before you collect a single data point saves weeks of reanalysis later. Here is the practical part. You pick your focal animal, start a timer, and record everything it does at set intervals. The two main approaches are focal sampling, where you follow one individual for a continuous period and note every behavior with timestamps, and scan sampling, where you sweep across a group at regular intervals and record what each individual is doing at that exact moment. Focal sampling gives you richer sequences. Scan sampling is faster and better for larger groups, but you lose the detail about transition patterns between behaviors. I spent three weeks trying to code burrowing owl hunting sequences using instantaneous scan sampling every thirty seconds. The problem was that these owls make most of their striking movements in under five seconds. My data missed roughly sixty percent of the actual hunting attempts because the behavior started and ended between my sampling points. I switched to focal sampling with event recording and a handheld audio timer, which cut my observer error significantly and actually gave me usable strike-to-success ratios. The tradeoff was that I could only follow one owl at a time instead of scanning the whole family group, so I lost demographic breadth. That kind of decision is constant in this work.
The Tools People Actually Use
Modern researchers mostly use specialized software instead of paper sheets. The most common options are BORIS, which is free and open source and handles both event and interval sampling with complex code hierarchies, and Animal Behavior Pro, which runs on iPads and is more streamlined for field use. There is also The Observer XT from Noldus, which is the professional-grade option but costs thousands of dollars per license and is overkill for most student or independent projects. If you are starting out, BORIS is the sensible choice. It takes about an hour to learn the interface if you already understand sampling methods. The free download is at bois.grenoble.fr. You build your ethogram inside the program, define your sampling method, connect it to a laptop or tablet, and go. For pure field work without technology, a simple clicker counter and a waterproof notebook still works fine. One researcher I know uses a motorcycle rev counter strapped to her wrist to tally displacement behaviors in captive primates. It is crude but effective, and she reports that the tactile feedback reduces counting errors compared to pencil and paper. That is not an endorsement of motorcycle parts as scientific equipment, but it is a reminder that improvised solutions often outperform expensive ones in messy field conditions.
Get the Full Details

Common Problems and How to Handle Them
Observer drift is the most common issue. This happens when your definition of a behavior subtly changes over time because you stop calibrating against your original criteria. After two weeks of scoring, you might find that what you labeled as "agonistic display" in week one now includes postures you would have coded as "neutral standing" in week three. The fix is regular reliability checks. Have a second observer score the same video simultaneously and calculate Cohen's kappa or percent agreement. If kappa drops below 0.80, you go back and recalibrate your definitions. Another problem that nobody warns you about is the reactance effect, where animals change their behavior because they notice you. This is especially bad with habituated species in long-term study sites. I worked with a population of free-ranging vervet monkeys that had been photographed and fed by researchers for twelve years. By year four of my project, their baseline activity budgets were completely different from what the published literature reported for the same population in the 1990s. Their vigilance rates were half because they had learned that humans meant food, not danger. I had to exclude the first six months of data and compare my results against contemporaneous studies rather than historical baselines. It is hard to admit that your study site is compromised, but publishing anyway with unvalidated comparisons is worse. Sample size is another area where people make mistakes. A common rule of thumb is that you need at least five independent focal follows per individual and ten individuals per treatment group for basic statistical power. But independent follows matter more than total observation hours. Ten hours on the same individual counted as ten data points will inflate your sample size artificially because the observations are autocorrelated. Use mixed-effects models with individual as a random effect if you have repeated measures on the same subjects. Otherwise your p-values are meaningless.
What This Method Does Not Do Well
Behavioral sampling has real limitations. It cannot capture internal states. You can code "resting" or "vocalizing," but you cannot tell from external observation alone whether the animal is stressed, comfortable, hungry, or satisfied without parallel physiological data. If you need mechanistic insight, you have to combine behavior coding with hormone assays, heart rate monitors, or neural recording, which adds cost and complexity. Another limitation is that sampling methods impose artificial structure on continuous behavior. Real animal behavior flows without clean boundaries. Deciding that a three-second stare is "investigation" and a two-second stare is not is arbitrary, even if your ethogram requires it. This is why event-based recording is generally preferred over time-based recording when your behaviors have natural start and stop points. The biggest practical bottleneck is data processing time. Coding ten hours of field video through BORIS typically takes six to eight hours of work if the animals are active and the behaviors are complex. Automated tracking tools like DeepLabCut can reduce this substantially, but they require camera setups, calibration, and computational resources that most behavioral ecologists do not have access to. The tradeoff is real: manual coding is slow but flexible, automated tracking is fast but fragile when animals overlap or leave the frame.
A Note on Ethics and Permitting
You cannot study wild animal behavior without institutional approval. Every country has regulations, and most universities require an IACUC or equivalent ethics review before fieldwork begins. This is not optional paperwork. I know a researcher who published behavioral data from wild raccoons without permitting and had the paper retracted two years later when the journal verified the permit status. The data went unpublished and the funding was clawed back. The process usually takes four to six weeks, so factor that into your timeline before you book travel or buy equipment. Non-invasive methods are preferred whenever possible. Remote cameras, audio recorders, and GPS collars reduce stress on the animals and often produce more reliable data because the observer is not present. The cost is higher initial setup and the loss of real-time behavioral context that a live observer provides. Most important, keep your records organized from day one. Name your files with dates and locations. Back up your ethograms and raw data to at least two separate drives. A corrupted hard drive after three field seasons is a specific kind of professional grief that is difficult to recover from.
