Setting Up a Basic Marketing Attribution Model From Scratch

Most people treat attribution like it is a machine learning problem. It usually is not. Before you touch any algorithms you need a data pipeline that actually connects your spend to your conversions. I spent three months untangling a mess where Google Ads, Meta, and our CRM each used different click IDs, and the numbers barely overlapped. The fix was writing a join key based on normalized email hashes rather than trying to reconcile device IDs across platforms. The practical version of this field is about taking messy marketing data and turning it into decisions that do not make your CFO nervous. You collect conversion events, match them to touchpoints, assign credit, and iterate. Here is how most teams actually get there. Pull your platform exports into one place. BigQuery, Snowflake, or a well-shaped spreadsheet works. The columns you need are roughly channel, cost, click date, user identifier, conversion date, conversion value, and medium. If your CRM does not expose user IDs you will spend a lot of time guessing instead of measuring. Match conversions to clicks using a lookback window first. Ninety days is standard for B2B. Thirty days covers most B2C. Anything longer and you are mostly modeling noise.

Last click is easy and wrong. First click is easy and also wrong. Start with linear or time decay until you have enough data for something fancier. I used a simple position-based model for a SaaS client with about twelve thousand leads per quarter. We gave thirty percent to first touch, forty percent to last touch, and split the rest evenly across mid funnel interactions. The results moved budgets by roughly eighteen percent within two quarters. Not dramatic, but real money saved because we stopped double spending on retargeting campaigns that were actually cannibalizing direct traffic. Looker or Power BI can show attribution. They cannot debug it fast enough when the numbers look wrong. Write a Python script that takes your unioned table and outputs attributed spend by channel. A basic implementation looks like this: Load your data into a pandas DataFrame.
Sort by user and timestamp.
For each conversion, grab the touchpoints within your lookback window.
Apply weights based on the chosen model.
Sum weighted spend by channel.
Export the result and compare it against platform reports to catch gross errors.

This process usually cuts the analysis time from a weekend job to about forty five minutes once the pipeline is set.

Get the Full Details

Data Science for Marketing Analytics, 2nd Edition - UNIVERZITNÁ KNIŽNICA
Data Science for Marketing Analytics, 2nd Edition - UNIVERZITNÁ KNIŽNICA

A specific edge case I ran into

About two years ago a major browser update broke a lot of our third party tracking. Our Meta pixel fired consistently but Google Click ID values dropped off after the first interaction on iOS. We had clean Facebook data and incomplete Google data. The easy move would have been to ignore Google or just use last click, which made Google look terrible. Instead I built a probabilistic model on top of the deterministic records. I used a logistic regression trained on our desktop users, where both platforms reported cleanly, to estimate the missing Google conversions for mobile sessions. The model was rough but it stopped us from pulling fifty thousand dollars out of Google entirely. Later we added a server side tag to recover most of the gap, but the model kept us from making a dumb decision while we waited for the fix. Any attribution model is an assumption. Test it. Take two similar markets or audience segments and run one channel hard in one and not the other. Measure incremental lift, not correlation. If your model says search drives twenty percent of revenue but the holdout shows almost no difference when you pause it, the model is lying to you. This happens more often than people admit. Seasonality, brand search, and competitive moves can all fake out attribution. People treat a correlation matrix as a causal proof. It is not. They also ignore cross channel cannibalization, which makes retargeting look stronger than it is. Another trap is using conversion values without adjusting for margin. A high value channel might be bringing in discount seekers who never return at full price. Attribute by gross profit, not revenue, if you can. Finally, do not retrain your model every week. Marketing data is noisy. Weekly updates chase random variation. Monthly or quarterly resets are enough for most setups.

If you have fewer than five thousand conversions per year, attribution modeling will just find patterns in randomness. Use simple rule based logic and focus on testing creative and offers instead. If your business relies heavily on offline sales with no digital trail, add call tracking and offline conversion imports. Without that you are modeling half the picture and calling it insight. Also, if your marketing mix includes a lot of influencer or PR work with no UTM or click data, attribution will understate those channels. Accept that limitation and supplement with surveys or brand lift studies rather than pretending the numbers tell the whole story. Union spend and conversions.
Match by user ID or hashed email within a lookback window.
Apply an attribution model, starting with linear or position based.
Validate with holdouts.
Adjust weights monthly, not daily.
Switch to Markov or Shapley value methods only after you have enough data and a reason to believe they improve decisions over simpler models. The whole thing is less about fancy code and more about keeping your data honest and your assumptions visible. Most teams that do this carefully end up reallocating between ten and twenty five percent of their budget within the first six months. That is usually enough to matter without causing political chaos.