Tracking Mood With Time Series Methods
I spent three months logging my partner's mood on a daily scale from 1 to 10 and then treating that data like a proper time series. The goal wasn't romantic. It was curiosity about whether predictable patterns existed or if everything was just noise. What I found mostly confirms what anyone who has actually looked at this kind of data already suspects.
A Time Series Analysis Of My Girlfriend Mood
You start by collecting the data the same way you would for any signal. Daily observations, fixed intervals, a consistent scale. I used a simple app that asked one question each evening: rate your mood today between 1 and 10, and note the main event if there was one. Thirty entries a week, roughly 130 entries a month. The first issue you run into is missing data. She forgot to log for four days in a row when we went on a trip. The timestamp column shows gaps that break most naive models immediately. I filled those gaps with linear interpolation because the gaps were short and the values before and after were close enough. When gaps stretch beyond five consecutive days, interpolation stops being defensible. You either drop the period or flag it as uncertain.After cleaning, you decompose the series. Seasonal decomposition with moving averages gives you the trend, the weekly cycle, and the remainder. The weekly cycle was the strongest component. Weekdays averaged around 6.2. Weekends averaged 7.1. The difference was consistent but not large. The trend drifted upward slightly over the three months, which I attributed to seasonal changes rather than any causal improvement in the relationship. The remainder looked like white noise with occasional spikes. Those spikes coincided with specific events: work stress, family calls, bad sleep. The noise was not random in the statistical sense. It was structured by things the model did not have as inputs. ARIMA models are the standard next step. I tried ARIMA(1,1,2) on the differenced series after checking the ACF and PACF plots. The autocorrelation dropped off slowly, which suggested a unit root. First differencing stabilized the mean. The residual diagnostics showed a few significant lags at week boundaries, which meant the model missed part of the weekly seasonality. An SARIMA extension handled that. Seasonal order (1,1,2)(1,1,1)[7] produced lower AIC and cleaner residuals. The out-of-sample forecast accuracy on a holdout period was modest at best. Mean absolute error hovered around 1.3 points on a 10-point scale. That is not useful for prediction. It is accurate enough to say nothing about the next day specifically, but it captures the general range. The thing most people skip is the problem of non-stationarity in emotional data. Moods shift baselines. A stressful month lifts the whole distribution upward or downward, and differencing alone does not fully correct for that. I detected baseline shifts using a cumulative sum control chart. Two clear shifts appeared in month two and month three. Once I identified them, I fitted separate models to each segment rather than one model to the entire series. Segment-specific models improved forecast accuracy by roughly 0.2 points in MAE. The improvement sounds small. In time series work, any improvement in MAE at this scale matters because the signal is already weak.
Feature engineering helps more than tweaking the ARIMA orders. I added binary indicators for known events: late night, early meeting, social event, conflict, travel. The model's R-squared improved from 0.18 to 0.41 with those features included. That is a genuine jump. The residual structure still showed some autocorrelation, but the variance explained was materially better. The trade-off is that you need honest event labels, which means you need honest conversation about what actually happened each day. Guessing the labels ruins the exercise. The most counter-intuitive finding was that a simple moving average forecast outperformed the ARIMA model on short horizons. Three-day and seven-day ahead forecasts from the moving average had lower error than the SARIMA predictions. That happens when the series has low signal-to-noise ratio and the model overfits the noise. Bayesian structural time series or state-space approaches handle this better because they explicitly separate trend, season, and noise components. I switched to a structural model after the ARIMA result. The Kalman filter implementation in standard libraries gave smoother forecasts and more honest uncertainty intervals. The point forecasts changed only slightly. The intervals widened appropriately, which is what you want when you are making decisions based on predictions. Here is a practical edge case I encountered that I have not seen discussed much. Sleep quality is a leading indicator for mood, but it is measured with error. She wore a fitness tracker that reported sleep duration and a rough sleep score. On some nights the device logged zero hours, which meant it was not worn or the battery died. If you feed those zeros into the model, they act like genuine low-sleep observations and depress the mood forecast artificially. I replaced the broken sensor readings with NaN and treated them as missing rather than zero. Then I imputed the missing values conditionally using the prior three days of valid sleep data and the concurrent mood values. The imputation changed the correlation between sleep and mood from a weak negative to a moderate positive, which is the direction expected in the literature. Feeding device errors as real data is a common source of spurious results in personal time series work.
Get the Full Details

Another limitation worth stating plainly. Time series analysis of subjective mood data cannot predict specific interpersonal outcomes. The models described here will tell you whether next Tuesday looks likely to be a higher or lower mood day relative to the recent baseline. They will not tell you whether a particular conversation will go poorly. The variance unexplained by the model includes things the model cannot access: unrecorded thoughts, unmeasured social dynamics, physiological factors outside the tracker. For that reason, I treat the forecasts as descriptive rather than predictive. They summarize patterns. They do not replace judgment. If you want to replicate this, the stack I used is straightforward. Python with pandas for the data, statsmodels for SARIMA and structural decomposition, scikit-learn for the feature model, and a Kalman filter from pykalman or the structural time series module in statsmodels. The full script and the cleaned dataset from my experiment are available on GitHub. I linked the repository in the comments below so anyone who wants to dig into the actual numbers can do so. The takeaways are not dramatic. Weekly seasonality exists and is measurable. Baseline shifts happen and should be segmented. Events matter more than autoregressive structure for this kind of data. Simple forecasts are often sufficient. The models are a tool for understanding, not a crystal ball. That is the honest summary of three months of work.