Understanding Physiological Computing and Signal Processing

I spent about two years ago trying to build a real-time stress detection system using heart rate variability and skin conductance data. The hardware was cheap — an Arduino, a pulse sensor, and a couple of Adafruit breakouts. The software side turned out to be way messier than I expected. Not because the math is hard, but because human bodies are noisy, inconsistent, and actively resist being measured cleanly. If you're looking into physiological computing or wondering what is a physiological signal processing pipeline, here's what actually works and what doesn't. A physiological pipeline takes raw biometric readings from the body and turns them into interpretable features. The steps are consistent across most projects: sensor acquisition, signal conditioning, feature extraction, and classification or visualization. The tricky part is between acquisition and conditioning. Raw signals from cheap sensors come with motion artifacts, baseline drift, and electromagnetic interference. You don't fix those problems with more code. You fix them with better hardware placement, appropriate sampling rates, and sometimes just accepting that your data will have gaps. For heart rate variability, you need a clean PPG or ECG signal sampled at 100 Hz or higher. Lower rates introduce aliasing problems that basic filters can't recover from. For skin conductance (GSR), 4 Hz is plenty — the signal changes slowly. Using the same sampling rate for everything is a common mistake. It wastes processing power and complicates your filtering strategy.

Building the Pipeline: From Raw Signals to Features

Here's a practical walkthrough. I'll use Python because it's the standard. The main libraries you'll need are scipy for filtering, numpy for array operations, and biosppy or neurokit2 for signal processing shortcuts. Start with raw data import. If you're working with CSV files exported from a sensor, the first thing to check is whether timestamps are consistent. Gaps in timestamps from Bluetooth sensors are normal — they buffer and resend. Linear interpolation fills small gaps. Anything over two seconds between samples usually means the sensor lost contact, and interpolation introduces artifacts you can't trust. Filtering comes next. For ECG signals, you typically apply a bandpass filter between 0.5 and 40 Hz to remove baseline wander and high-frequency noise. A notched filter at 50 or 60 Hz handles mains interference. scipy.signal.butter combined with scipy.signal.filtfilt does this cleanly. The filtfilt function applies the filter forward and backward, eliminating phase distortion. Phase distortion matters if you're correlating signal features across modalities like heart rate and skin conductance.

I ran into a specific problem once where my GSR signal showed artificial periodicity that matched exactly the sampling jitter of my USB-to-serial converter. The solution wasn't a better filter. It was switching to a different USB cable. The old cable had a ground loop that induced a 120 Hz ripple at half the sampling rate, which aliased down into the signal band. Cheap problem, expensive lesson.

Get the Full Details

What is Physiology? - GeeksforGeeks
What is Physiology? - GeeksforGeeks

Feature Extraction That Actually Works

For HRV, the standard features are SDNN (standard deviation of R-R intervals), RMSSD (root mean square of successive differences), and pNN50 (percentage of successive intervals differing by more than 50 ms). RMSSD is the most reliable for short recordings — you can get a stable reading from 30 seconds of clean data. SDNN needs at least five minutes. If someone tells you to compute SDNN from a 30-second window, they don't know what they're talking about. For GSR, you're looking at skin conductance level (the baseline) and skin conductance responses (the phasic peaks). Detecting peaks requires a derivative threshold. I use a simple approach: compute the first derivative, find points where it exceeds three times the rolling standard deviation, then verify each candidate peak is a local maximum within a 0.5-second window. This catches about 85% of genuine responses in clean data. The remaining 15% are usually respiratory sinus arrhythmia artifacts or motion spikes. One counter-intuitive insight: higher GSR isn't always more arousal. If your subject's baseline conductance drifts upward over a session due to sweating, every subsequent response will look smaller relative to the new baseline. Normalize by subtracting a rolling minimum or using the tonic component estimated through a low-pass filter at 0.05 Hz. This is standard in psychophysiology but gets skipped in most hobbyist implementations.

Classification and Real-World Deployment

If you're building something that classifies states — stress, relaxation, cognitive load — you need labeled training data. This is the part most people underestimate. Collecting 500 data points with manual labels takes longer than training any model on them. I spent three weeks building a decent labeled dataset for a stress detection prototype. The model itself took two days to train. For classification, start simple. A random forest with 100 trees on hand-crafted features usually matches or beats deep learning approaches on small physiological datasets. The reason is sample efficiency. Neural networks need hundreds or thousands of subjects to generalize. Random forests work with dozens. Use scikit-learn. The pipeline is straightforward: FeatureUnion combining HRV features, GSR features, and temporal statistics, then a RandomForestClassifier with grid search on max_depth and n_estimators. Cross-validation strategy matters enormously here. Don't shuffle your data randomly across folds. Shuffle by subject. If subject A's data appears in both training and test folds, your accuracy numbers will be inflated by 15-30% because the model learns individual baselines rather than generalizable patterns. Leave-one-subject-out cross-validation is the gold standard, even though it's computationally expensive.

Pitfalls and Limitations

Physiological signals are highly individual. A resting heart rate of 60 means something different for a trained athlete than for a sedentary person. Your classification thresholds need personal calibration, or they need to be normalized per subject before aggregation. Group-level models often perform worse than per-subject models because they average out the very signal patterns you're trying to detect. Another limitation: consumer-grade sensors are not medical devices. The pulse sensors available for microcontrollers have accuracy specifications in the range of ±5 BPM under ideal conditions. Motion degrades that significantly. If your application involves walking or gesturing, you need motion compensation or a different sensor modality. Accelerometer data fused with PPG can help, but the fusion logic adds complexity that's often not worth it for non-clinical applications. Long-term deployment introduces another issue: skin contact quality degrades over hours. Electrodes dry out. Optical sensors shift position. I've seen systems that performed well in a 20-minute lab session degrade to chance-level accuracy over a four-hour continuous recording. If you're building something meant to run for extended periods, plan for periodic recalibration or use algorithms that adapt their baselines online rather than relying on a fixed preprocessing pipeline.

A.1.1. What is Physiology? - BasicPhysiology.org
A.1.1. What is Physiology? - BasicPhysiology.org

Tools and Resources

For open-source tooling, neurokit2 is the most comprehensive Python library. It handles preprocessing, feature extraction, and visualization for ECG, EDA, EMG, respiratory, and EEG signals. The API is clean and the documentation includes examples. biosppy is another option, slightly more minimal but with solid signal processing routines. Both integrate well with pandas and matplotlib. If you need hardware recommendations, the Empatica E4 wristband is the industry standard for research-grade GSR and accelerometer data, though it's expensive. For prototyping on a budget, the Arduino Nano with an Adafruit MAX30102 (pulse oximeter/heart rate sensor) and a GSR breakout covers the basics. The coding sample rates are adequate for HRV and GSR analysis when you're starting out. The FieldTrip toolbox in MATLAB is worth mentioning if you're working with EEG data. It's MATLAB-licensed, but the signal processing pipelines are among the best documented in the field. If you're doing anything beyond basic processing, the difference between ad-hoc filtering and FieldTrip's pipeline shows up clearly in your spectral estimates.

A Note on Ethics and Data Handling

Physiological data is personal health information. Even anonymized, it can be used to infer mental states, health conditions, and behavioral patterns. If you're collecting this data from other people, get proper consent. IRB approval is required for academic work. For personal projects, at minimum make sure your subjects understand what you're measuring and how you're storing the data. A blood pressure reading doesn't seem sensitive until you realize it can reveal stress patterns, medication effects, and cardiovascular risk factors. The field moves fast. New sensing modalities appear regularly — thermal imaging for stress detection, voice analysis for cognitive load, pupilometry for attention tracking. Each adds complexity to your pipeline. Start narrow. Master one signal type before expanding. The people I've seen succeed in this space are the ones who built deep expertise on ECG first, then added GSR, then layer by layer expanded their capabilities. The ones who tried to build multi-modal systems from day one usually ended up with five mediocre data streams instead of one well-understood one.