Getting the Scores Out of a PLS Model Without Losing Your Mind
People keep asking about how to properly score new data with a Partial Least Squares model and whether there is an actual authoritative reference for the workflow. The term Pls Scoring Manual comes up when you search for a clean step-by-step description, but honestly, most of what you find online is either a software help file buried in a vendor PDF or a blog post that skips the part where things go wrong. I have spent too many years watching people apply a PLS model built in Wavelength to real production data and get garbage predictions because they missed one preprocessing step or misunderstood how the scores are calculated. At its core, scoring with a PLS model means taking a new observation, running it through the same preprocessing pipeline used during model training, projecting it onto the latent variables, and then transforming those latent scores back through the regression weights to get your predicted Y values. That sentence sounds simple. It is not always simple in practice. A complete manual would walk through:
- How to carry forward the mean, standard deviation, and multiplicative scatter correction parameters from the calibration set
- How to handle missing values without collapsing the projection
- The exact matrix multiplication order for T and Q so you do not accidentally transpose a weight matrix
- How to calculate prediction intervals using the residual variation, not just the RMSEP
- How to detect outliers in the new sample space using both the T² and Q residual frameworks
- What to do when a new sample falls outside the applicability domain of the original calibration If you are looking for that kind of document, check the original PLS methodology papers by Wold, and then look at the implementation notes from the software you are actually using, whether that is SIMCA, Unscrambler, PLS_Toolbox, or something like scikit-learn. The theory is consistent across platforms. The implementation details are where people trip up.
How to Score New Samples Correctly
Here is the actual procedure, not the abbreviated version you see in most tutorials. Step one is securing the preprocessing parameters. When you build your PLS calibration model, you will typically mean-center and scale the X matrix, possibly apply SNV or MSC if you are working with spectral data, and maybe do a wavelength selection. You must save every single parameter from that process. The mean vector, the scale vector, the MSC coefficients, everything. When you score a new sample later, you apply those exact parameters. You do not recompute them from the new sample. You do not recompute them from a new batch of data. The preprocessing parameters belong to the calibration set, not the sample being scored. Step two is the projection. You take your preprocessed new sample vector and multiply it by the PLS weight matrix W to get the score vector T. If you are using the NIPALS algorithm, which most implementations do, W is stored explicitly and you can compute T directly. If you are using a kernel-based variant or an orthogonal PLS model, the transformation may involve additional matrices. Check your software documentation for the exact form. The common mistake here is using W_star or the final beta weights instead of the original W matrix. They are related but not identical, and swapping them will give you slightly wrong scores that look plausible until you check the residuals.
Get the Full Details

Step three is the regression prediction. You multiply your score vector T by the Y-loading matrix P (or use the direct regression coefficient vector B, which is equivalent if computed correctly) to get the predicted Y values. If you are doing PLS2 with multiple Y variables, you need to make sure you are applying the correct P matrix. PLS1 and PLS2 handle the Y-space differently, and some software packages wrap them in slightly different ways. Step four is the validation. You should never score a new sample and just accept the predicted value. You need to check the distance metrics. The T² statistic tells you whether the sample is inside the modeled variation in the latent space. The Q residual, also called the squared prediction error or SPE, tells you whether the sample contains variation that the PLS model did not capture. A sample can have a normal T² but a huge Q residual, which usually means it is a novel material or has an interference that was not present in the calibration set. Conversely, a normal Q but extreme T² means the sample is far outside the concentration or property range of your calibration. Either way, flag it and do not treat the prediction as reliable.
The Problem I Ran Into and How I Fixed It
Last year I was working with a client who had a PLS model for near-infrared spectroscopy that had been validated and deployed two years earlier. The model was built with about 400 calibration samples spanning the expected concentration ranges. Everything looked solid on paper. RMSEP was low, cross-validation was stable, the loadings made chemical sense. Then they started getting new batches of raw material that occasionally produced predictions that were internally consistent but clearly wrong when they compared them to the reference method. The T² and Q residuals both looked fine. The samples were not outliers by the standard metrics. The model was just predicting the wrong values. Turns out the issue was a subtle baseline drift in the spectrometer. The original calibration set had been collected on Instrument A, and the new batches were being scanned on Instrument B, a newer model from the same manufacturer. The spectra looked similar enough that the PLS model accepted them without triggering any outlier detection. But the baseline offset was systematically shifting the predictions by a small amount, maybe two to three percent of the full range. Over the concentration span they were working with, that difference was significant.
The fix was not to rebuild the model. The fix was to apply a transfer function. We used a piecewise direct standardization approach to map the spectra from Instrument B onto the spectral space of Instrument A before scoring. After that, the predictions aligned with the reference method. The original model, the preprocessing, the weight matrices, everything was still valid. It was just the input domain that had drifted. This is the kind of thing that is never going to be in a textbook. The Pls Scoring Manual you find online will tell you to check T² and Q residuals. It will not tell you that your instrument might be subtly different from the one you calibrated on and that your scores are technically valid but chemically wrong because of it.

Common Pitfalls and What Most People Miss
One thing beginners almost always get wrong is the order of operations between scaling and centering. If your PLS model was mean-centered and scaled, your new sample must be centered and scaled using the calibration means and standard deviations before the projection step. But if you are working with a model that also includes a standardization step like SNV or MSC, that step needs to happen before the mean-centering and scaling, not after. The order matters because these operations are not commutative. Get it wrong and your scores will be systematically biased. Another thing that catches people is the number of latent variables. The optimal number of factors for calibration is determined by cross-validation, but scoring itself does not depend on that choice in the way people think. Once the model is built, you project onto however many latent variables the model retained. The question is whether you are projecting onto too few or too many. Too few and you lose predictive information. Too many and you start projecting noise. If you are scoring new samples regularly, you should periodically re-evaluate whether the original number of latent variables is still appropriate, especially if your sample population has changed. A third issue is the applicability domain. PLS models are interpolation engines. They are not good at extrapolation. If your new samples have any measured property that falls outside the range of the calibration set, the predictions will be unreliable even if the T² and Q residuals look acceptable. This is because the regression coefficients in the Y-space are derived from the calibration range, and pushing beyond that range means you are asking the model to extrapolate linear relationships that were never validated there.
Software and Resources
Depending on what you are working with, the actual manual you need is usually tied to your software package. If you are using MATLAB with PLS_Toolbox, the documentation is thorough and covers the exact matrix operations. If you are using Python with scikit-learn or sklearn-ppls, the API is more minimal and you will need to handle some of the intermediate steps yourself, which actually gives you more visibility into what is happening. If you are using an instrument-specific package like those from Bruker, Agilent, or Shimadzu, the scoring procedure is often automated behind a button, which is convenient but dangerous if you do not understand what is happening under the hood. For a comprehensive theoretical reference, the original work by S. Andersson is solid, and the practical implementations covered in various chemometrics textbooks will give you the matrix algebra details that vendor manuals often skip. There is no single universal Pls Scoring Manual because the field is fragmented across chemistry, biology, process engineering, and marketing analytics, and each community has adapted PLS scoring to its own needs. The underlying mathematics is the same, but the preprocessing, validation, and outlier detection conventions vary enough that a one-size-fits-all document would not be useful.
When PLS Scoring Will Not Save You
There are scenarios where PLS is the wrong tool regardless of how well you follow the scoring procedure. If your data has strong nonlinear relationships, PLS will underperform compared to a kernel-based method or a neural network, no matter how carefully you preprocess or validate. If your X matrix has severe collinearity that is not captured by the latent variable structure, adding more factors will not help and may hurt. If your Y variables are highly correlated in a way that introduces numerical instability during the regression step, you may need to switch to PLS2 or a different multivariate approach entirely. PLS also struggles with very high-dimensional data where the number of variables far exceeds the number of samples unless you incorporate variable selection or regularization. Standard PLS handles this better than PCR, but it is not a magic bullet. If you are working with datasets that have tens of thousands of features, you should consider whether a sparse PLS variant or a completely different modeling strategy would be more appropriate before investing time in fine-tuning a standard PLS pipeline. The bottom line is that scoring a PLS model is straightforward if you respect the preprocessing chain, validate the predictions against the calibration domain, and check both the latent space distance and the residual space distance for every new sample. Most of the problems people encounter come from cutting corners on one of those three items, not from any flaw in the method itself.
