Getting Part Worths Out of Conjoint Without Losing Your Mind
Part worths are the utility values assigned to each level of every attribute in a conjoint study. That's the textbook answer. The practical answer is that they tell you how much relative preference someone has for, say, a $200 price point versus a $300 price point, or brand A versus brand B. You add them up across attributes and you get a composite utility score for each profile. It's arithmetic, not magic. Here's what actually happens when you run a standard hierarchical Bayes or OLS conjoint. Respondents pick between profiles. You feed those choices into a model. The model spits out numeric utilities for each attribute level. Those numbers are your part worths. They're usually normalized to a 0-to-100 scale or mean-centered around zero so they're comparable across attributes. The normalization step is where people screw up most often. If you don't normalize, a four-level price attribute with utilities ranging from -500 to +500 will completely dominate a three-level color attribute ranging from -2 to +2. Your simulated market share predictions will be garbage. Always normalize within a reasonable range. Most packages do this automatically, but check the output. Don't trust defaults blindly.
I ran into a problem a while back on a pharmaceutical conjoint where the part worths for dosage frequency came out completely flat — like, essentially zero variance across all levels. The model had converged, the log-likelihood looked fine, and the fit statistics were acceptable. But the part worths told me respondents couldn't distinguish between once-daily, twice-daily, and three-times-daily dosing. At first I thought it was a sampling issue. It wasn't. The survey was embedded in a longer health attitude instrument, and by the time respondents got to the conjoint section, they were fatigued and just clicking through. The workaround was straightforward: I went back and flagged all the respondents whose choice task response times fell below the 25th percentile. Removing those 18 percent of sessions cleaned up the dosage part worths immediately. The other attributes looked normal the whole time. It was purely a fatigue artifact.
The Mechanics Nobody Warns You About
Part worths are interval-scaled, not ratio-scaled. That means you can compare differences between levels, but you cannot say one utility is "twice as preferred" as another. People do this constantly in client presentations and I've sat through enough of those meetings to know it's worth flagging early. If attribute A has a range of 60 points and attribute B has a range of 15 points, attribute A is roughly four times as important in driving choice. That's a ratio of ranges, not a ratio of individual part worths. Another thing that trips people up: part worths are relative to the reference level in your coding scheme. If you use dummy coding with the lowest price as the base, all other price part worths are measured against that. Switch to effects coding and the interpretation shifts. With effects coding, the part worth tells you how much better or worse a level is compared to the average of all levels, not compared to a single baseline. This matters when you're aggregating across segments or doing meta-analysis across studies. Pick a coding scheme and stick with it unless you have a reason to change. Document which one you used. The computational side is also worth addressing head-on. If you're doing OLS conjoint with 10 attributes and 4 levels each, you're estimating roughly 30 parameters per respondent. That's manageable. Jump to 15 attributes with 5 levels and you're at 60+ parameters. OLS starts to wobble. Hierarchical Bayes handles this better because it borrows strength across respondents through the population-level prior. But even HB has limits. I've seen it produce wildly unstable part worths on attributes with very low choice frequency in the data — maybe three out of 200 respondents ever chose a particular level. The posterior distribution just hasn't been constrained enough. The fix is either to collapse those levels or to constrain the prior. Most software lets you adjust prior variance. Shrinking it toward zero dampens the noise without eliminating the attribute entirely.
Get the Full Details

What Part Worths Can't Tell You
They can't account for interaction effects unless you explicitly model them. Standard main-effects-only conjoint assumes attributes are independent. In reality, people might want a premium brand specifically because it comes in a certain price range. The combo creates utility that isn't captured by adding two separate part worths. If you suspect interactions, you need to include them in the design. That means more profiles, longer surveys, and a bigger sample size. The part worths you get from a main-effects model in the presence of unmodeled interactions are biased estimates. Not catastrophically wrong, but systematically off. I'd rather under-predict an interaction than overstate confidence in main effects alone. They also don't translate directly into dollar certainty. A part worth difference of 20 utility points between two price levels doesn't mean the lower price generates exactly 20 percent more sales. It means something close, depending on the simulation method you use. I prefer Monte Carlo simulation over deterministic simulation because it accounts for parameter uncertainty. Deterministic simulations give you a single point estimate that looks precise but hides the variance. Monte Carlo runs thousands of draws from the posterior and gives you a distribution of market shares. It takes longer — usually 5 to 10 minutes on a decent machine versus 30 seconds for deterministic — but the output is genuinely more useful for decision-making. The biggest limitation is that part worths assume stable preferences. In practice, respondents adapt their decision strategies across the choice tasks. Someone might start by screening on price and then only consider brand as a tiebreaker. The model treats this as random error rather than a strategic shift. Best-worst scaling or choice-based conjoint with task-level covariates can partially address this, but you're still working with an approximation of human behavior. No model fixes that completely.