Skip to content

Core Model Algorithms in Detail โ€‹

The core capabilities of StarWay Data Insight are built on five classic multivariate statistical models: PCA (Principal Component Analysis), PLS (Partial Least Squares Regression), PLS-DA (Partial Least Squares Discriminant Analysis), OPLS (Orthogonal Partial Least Squares Regression), and OPLS-DA (Orthogonal Partial Least Squares Discriminant Analysis).

This chapter takes a deep look at the principles, applicable scenarios, mathematical essence, and concrete application in the platform of these five algorithms. Understanding these models will help you interpret analysis results better and make more precise data decisions.


๐Ÿ“Š Model Family Overview โ€‹

Before going into the details, let's first see the relationship among the five with one diagram:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                     Multivariate Data Analysis Model Family                      โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                                                  โ”‚
โ”‚   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”               โ”‚
โ”‚   โ”‚ Unsupervised โ”‚            โ”‚            Supervised            โ”‚               โ”‚
โ”‚   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
โ”‚          โ”‚                                     โ”‚                                 โ”‚
โ”‚          โ–ผ                                     โ–ผ                                 โ”‚
โ”‚   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”               โ”‚
โ”‚   โ”‚     PCA      โ”‚            โ”‚      Is Y a continuous value?    โ”‚               โ”‚
โ”‚   โ”‚ Explore the  โ”‚            โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
โ”‚   โ”‚ structure &  โ”‚                             โ”‚                                 โ”‚
โ”‚   โ”‚ find patternsโ”‚                  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                      โ”‚
โ”‚   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                  โ–ผ                     โ–ผ                      โ”‚
โ”‚                                    Yes                   No                      โ”‚
โ”‚                               (regression)        (classification)               โ”‚
โ”‚                                     โ”‚                     โ”‚                      โ”‚
โ”‚                                     โ–ผ                     โ–ผ                      โ”‚
โ”‚                            โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚                            โ”‚       PLS        โ”‚  โ”‚      PLS-DA      โ”‚            โ”‚
โ”‚                            โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ”‚                                     โ”‚                     โ”‚                      โ”‚
โ”‚                                     โ–ผ                     โ–ผ                      โ”‚
โ”‚                            โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚
โ”‚                            โ”‚       OPLS       โ”‚  โ”‚     OPLS-DA      โ”‚            โ”‚
โ”‚                            โ”‚ Strips orthogonalโ”‚  โ”‚ Strips orthogonalโ”‚            โ”‚
โ”‚                            โ”‚      noise       โ”‚  โ”‚      noise       โ”‚            โ”‚
โ”‚                            โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜            โ”‚
โ”‚                                                                                  โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

One-sentence summary:

  • PCA: "What does the data look like?" โ†’ Explore the intrinsic structure
  • PLS: "How does X affect Y?" โ†’ Establish a predictive relationship
  • PLS-DA: "Which class does it belong to?" โ†’ Perform classification and discrimination
  • OPLS: "How does X affect Y, and how do we separate the noise out?" โ†’ Regression that is easier to interpret
  • OPLS-DA: "Which class does it belong to, and how do we separate out the within-group noise?" โ†’ Classification that is easier to interpret

๐Ÿ’ก The relationship between PLS and OPLS: OPLS is not a "more accurate" model; it splits the variation of X into a part "related to Y" and a part "unrelated to Y", making the results easier to explain. In predictive ability the two are usually close, and the main reason to choose OPLS is readability and the clean score plot brought by a single predictive component.


๐Ÿ”ฌ PCA (Principal Component Analysis) โ€‹

What is PCA? โ€‹

PCA (Principal Component Analysis) is an unsupervised dimensionality reduction technique. Its core idea is: use fewer new variables (principal components) to retain as much of the information in the original data as possible.

Imagine you have a set of three-dimensional data points. PCA finds the best two-dimensional plane such that, after the points are projected onto it, they are "spread out" the most โ€” this way the information loss is minimal.

PCA principle illustration

Reading the figure above:

  • The blue points represent the original high-dimensional (3D) data
  • The red plane is the best projection plane found by PCA (spanned by PC1 and PC2)
  • The green dashed lines show the process of projecting the data points onto the plane
  • The red crosses after projection retain the information of the direction of maximum variance in the data

Core principles โ€‹

1. Variance is information โ€‹

PCA holds that: the direction in which the data varies the most contains the most information.

  • If a variable is about the same across all samples (small variance), it carries little information
  • If a variable differs greatly (large variance), it carries important information

2. Construction of the principal components โ€‹

Through a linear transformation, PCA converts the original correlated variables into mutually uncorrelated new variables (principal components):

Where:

  • are called principal components
  • is the loading, representing the contribution of the original variable to the new component
  • The principal components are mutually uncorrelated (orthogonal)

3. Properties of the principal components โ€‹

  • The first principal component (PC1): explains the direction of maximum variance in the data
  • The second principal component (PC2): explains the direction of maximum remaining variance, subject to being orthogonal to PC1
  • And so on...

Mathematical essence (simplified) โ€‹

The mathematical essence of PCA is eigenvalue decomposition of the covariance matrix:

  1. Center the data: subtract the mean of each variable
  2. Compute the covariance matrix:
  3. Eigenvalue decomposition:
    • (eigenvalue): represents the amount of variance explained by that principal component
    • (eigenvector): represents the direction of the principal component (that is, the loadings)
  4. Select principal components: sort the eigenvalues from largest to smallest and take the first

Application in the platform โ€‹

Applicable scenarios โ€‹

  • Data exploration: get a preliminary understanding of the overall structure and distribution of the data
  • Anomaly detection: find abnormal samples through the Tยฒ and SPE statistics
  • Dimensionality reduction visualization: project high-dimensional data into 2D/3D space for observation
  • Denoising: remove noisy components and keep the main signal

PCA score plot example โ€‹

The figure below shows a typical score plot from PCA analysis. Each point represents one sample, so the distribution pattern of the samples and any abnormal points can be seen directly:

PCA score plot example

How to interpret it:

  • Each gray point represents one sample, positioned by its scores on PC1 and PC2
  • The elliptical region represents the confidence interval (usually 95%); points outside it may be abnormal samples
  • Points clustered together represent similar samples, while scattered points indicate greater differences

Automatic trigger conditions in the platform โ€‹

When you have configured only X variables (no Y variable), the platform automatically uses the PCA model:

X columns only โ†’ PCA is selected automatically โ†’ explore the data structure

Key output metrics โ€‹

MetricMeaningInterpretation
RยฒXCumulative explained rate of XHow much information in X the model explains; the closer to 1 the better
Cumulative contribution rateThe explained proportion of the first k principal componentsUsually 80%~95% is enough
LoadingThe relationship between variables and principal componentsSee which variables dominate that component
ScoreThe coordinates of samples in the new coordinate systemUsed to draw scatter plots and observe the sample distribution

Choosing the number of principal components โ€‹

The platform automatically selects the optimal number of principal components through cross-validation, but you can adjust it manually with C+1/C-1:

  • Too few: underfitting, with serious information loss
  • Too many: overfitting, introducing noise
  • Rule of thumb: consider stopping when the contribution of an additional component to RยฒX is < 5%

Pros and cons of PCA โ€‹

โœ… Advantages:

  • Unsupervised, no labeled data needed
  • Computationally efficient, with interpretable results
  • Effective dimensionality reduction, removing correlation between variables
  • Good visualization results

โš ๏ธ Limitations:

  • Only looks at the variance structure of X and does not consider Y
  • Sensitive to outliers
  • Principal components are linear combinations, and may lack business meaning
  • Assumes the principal components are orthogonal, which real data may not satisfy

๐Ÿ”— PLS (Partial Least Squares Regression) โ€‹

What is PLS? โ€‹

PLS (Partial Least Squares Regression) is a supervised regression method. Unlike PCA, PLS considers both X (features) and Y (target) when modeling, looking for a latent variable space that best explains the relationship between the two.

Simply put: PCA asks "how does X vary", while PLS asks "how does X affect Y".

PLS principle illustration

Reading the figure above:

  • The blue box on the left is the X variables (X1-X5), and the green box on the right is the Y variable (target)
  • The yellow box in the middle is the extracted latent variables (LV1, LV2), which explain both X and Y
  • The goal of PLS is to maximize the covariance between the X latent variables and the Y latent variables
  • A predictive relationship X โ†’ Y is established through the latent variables

Core principles โ€‹

1. Decompose X and Y simultaneously โ€‹

PLS decomposes X and Y at the same time, but requires their latent variables to be maximally correlated:

Where:

  • : the score matrix of X (similar to the scores of PCA)
  • : the loading matrix of X
  • : the score matrix of Y
  • : the loading matrix of Y
  • : the residual matrices

2. Maximize the covariance โ€‹

The core optimization objective of PLS is: find the latent variables of X and Y such that their covariance is maximized.

This means the components extracted by PLS must both represent the variation of X and be closely related to Y.

3. Extract components iteratively โ€‹

PLS extracts latent variables one by one through an iterative algorithm (such as NIPALS):

  1. Find the direction of maximum covariance between X and Y as the first pair of latent variables
  2. Subtract the already explained part from X and Y (deflation)
  3. Repeat until enough components have been extracted

Mathematical essence (simplified) โ€‹

The mathematical core of PLS is covariance maximization:

  1. Initialization: start from some column of Y or a random vector
  2. Iterative optimization:
    • (find the weights of X from the scores of Y)
    • (compute the scores of X)
    • (find the weights of Y from the scores of X)
    • (compute the scores of Y)
  3. After convergence: compute the loadings
  4. Deflation: ,

Application in the platform โ€‹

Applicable scenarios โ€‹

  • Regression prediction: build a predictive model X โ†’ Y
  • Variable screening: use VIP to find the X variables with the greatest influence on Y
  • Multi-response problems: Y can be several columns (multiple response variables)
  • Collinearity handling: remains stable when X variables are highly correlated

Automatic trigger conditions in the platform โ€‹

When you have configured both X variables and a Y variable, and Y is a continuous value:

X + Y (continuous values) โ†’ PLS is selected automatically โ†’ build a regression model

๐Ÿ’ก You can also manually specify the model type in the data configuration, switching freely among PCA / PLS / PLS-DA / OPLS / OPLS-DA.

Key output metrics โ€‹

MetricMeaningInterpretation
RยฒXCumulative explained rate of XThe proportion of X variance captured by the model
RยฒYCumulative explained rate of YThe proportion of Y variation explained by the model; the higher the better
QยฒYCross-validated predictive ability for YThe most critical one! Reflects generalization ability; > 0.5 is acceptable, > 0.9 is excellent
RMSERoot mean square errorThe average deviation between predicted and true values; the smaller the better
VIPVariable importance in projection> 1 indicates an important variable, < 0.5 can be ignored

Choosing the number of latent variables โ€‹

The platform automatically selects the optimal number of latent variables through cross-validation, on the following principles:

  • The number of components is optimal when QยฒY reaches its peak
  • If RยฒY is very high but QยฒY is low โ†’ overfitting, and the number of components should be reduced
  • If both are low โ†’ underfitting, and you may need more components or a check of the data

VIP analysis โ€‹

VIP (Variable Importance in Projection) is an important output of PLS; it tells you which X variables matter most for predicting Y.

The figure below shows a typical VIP analysis result: the taller the bar, the greater the influence of that variable on Y:

VIP variable importance plot

How to interpret it:

  • The red reference line (VIP = 1) is the threshold for an important variable
  • Variables above the red line (such as x3 and x5 in the example) contribute significantly to Y
  • Variables below the red line can be considered for removal to simplify the model

The VIP formula:

Interpretation criteria:

  • VIP > 1: an important variable with a significant contribution to Y
  • 0.5 < VIP < 1: moderately important
  • VIP < 0.5: negligible, consider removing it

Pros and cons of PLS โ€‹

โœ… Advantages:

  • Handles X and Y at the same time, with strong predictive power
  • Effectively solves the multicollinearity problem
  • Supports multiple response variables (multi-response)
  • Provides VIP for variable screening
  • Still works when the number of samples is smaller than the number of variables

โš ๏ธ Limitations:

  • Requires labeled data (Y)
  • Model interpretation is more complex than PCA
  • Limited ability to model nonlinear relationships
  • Sensitive to outliers

๐ŸŽฏ PLS-DA (Partial Least Squares Discriminant Analysis) โ€‹

What is PLS-DA? โ€‹

PLS-DA (Partial Least Squares Discriminant Analysis) is an extension of PLS, dedicated to classification problems. When Y is a class label (such as "qualified/unqualified" or "Class A/Class B/Class C"), PLS-DA is your choice.

Simply put: PLS predicts values, PLS-DA predicts classes.

PLS-DA classification principle illustration

Reading the figure above:

  • The blue circles represent Class A and the orange squares represent Class B
  • The samples of the two classes form clearly separated clusters in the latent variable space (LV1-LV2)
  • The green dashed line is the decision boundary used to separate the two classes
  • The shaded ellipses represent the confidence regions of each class; the less they overlap, the better the classification

Core principles โ€‹

1. Turn the classification problem into a regression problem โ€‹

The clever trick of PLS-DA is: convert the class labels into dummy variables, then use PLS for regression.

For example, a three-class problem (A, B, C) is converted into:

SampleOriginal labelY1(A)Y2(B)Y3(C)
1A100
2B010
3C001

Then standard PLS is applied to this multi-response Y matrix.

2. Discriminant rule โ€‹

At prediction time, PLS-DA outputs a "score" for each class, and the sample is assigned to the class with the highest score:

3. Visualization advantages โ€‹

The score plot of PLS-DA is naturally suited to showing classification performance:

  • Samples of different classes should form separated clusters in the plot
  • The first latent variable is usually most related to the between-group difference
  • The second latent variable shows the within-group variation

Mathematical essence โ€‹

The mathematics of PLS-DA is almost the same as PLS; the difference lies in how the Y matrix is constructed:

  1. Encoding: convert the class labels into an indicator matrix
  2. PLS regression: run standard PLS on X and the encoded Y
  3. Discrimination: at prediction time, choose the class with the largest response value

Class encoding methods:

  • Binary classification: Y = 0/1 or -1/+1
  • Multiclass classification: one-hot encoding (one column per class)

Application in the platform โ€‹

Applicable scenarios โ€‹

  • Binary classification: qualified/unqualified, positive/negative, normal/abnormal
  • Multiclass classification: raw material grading, product typing, variety identification
  • Feature selection: find the key variables that distinguish different classes
  • Biomarker discovery: medical and omics data analysis

Automatic trigger conditions in the platform โ€‹

When you have configured X variables and a Y variable, and Y is a class label (text or discrete values):

X + Y (class labels) โ†’ PLS-DA is selected automatically โ†’ build a classification model

Key output metrics โ€‹

MetricMeaningInterpretation
RยฒXCumulative explained rate of XThe X variance captured by the model
AccuracyClassification accuracyThe proportion of correct predictions; watch out for the trap of class imbalance
F1 ScoreHarmonic mean of precision and recallMore reliable than Accuracy when classes are imbalanced
AUCArea under the ROC curveDiscrimination ability: 0.5 is random, 1.0 is perfect, > 0.8 is good
Confusion matrixPredicted vs. actual classification tableSee at a glance which class is easily confused
VIPVariable importanceFind the key variables that distinguish the classes

Classification performance evaluation โ€‹

Reading the confusion matrix:

The figure below shows the confusion matrix of a PLS-DA model, directly displaying the model's prediction performance on each class:

PLS-DA confusion matrix

How to interpret it:

  • The values on the diagonal (dark color) are the number of correctly classified samples
  • The off-diagonal values are the number of misclassified samples
  • Ideally, all samples should be concentrated on the diagonal

The confusion matrix in table form:

                  Predicted
           Positive      Negative
        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  True  โ”‚     TP      โ”‚     FN      โ”‚
Positiveโ”‚ (True Pos.) โ”‚ (False Neg.)โ”‚
        โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
  True  โ”‚     FP      โ”‚     TN      โ”‚
Negativeโ”‚ (False Pos.)โ”‚ (True Neg.) โ”‚
        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Derived metrics:

  • Precision: โ€” of the ones predicted positive, how many are real
  • Recall: โ€” of the truly positive ones, how many were found
  • F1 Score: โ€” a combined metric

The trap of class imbalance:

If 95% of the samples are Class A and 5% are Class B:

  • Even if the model predicts Class A for everything, Accuracy is still 95%
  • But this is completely useless for Class B!
  • Solution: look at the F1 Score and AUC, or adjust the class weights

ROC and AUC โ€‹

ROC curve: true positive rate vs. false positive rate at different thresholds

ROC curve example

How to interpret it:

  • The closer the curve is to the top-left corner, the better the model
  • The diagonal = random guessing

AUC evaluation criteria:

  • AUC = 0.5: random level (no discrimination ability)
  • 0.7 โ‰ค AUC < 0.8: acceptable
  • 0.8 โ‰ค AUC < 0.9: good
  • AUC โ‰ฅ 0.9: excellent

Pros and cons of PLS-DA โ€‹

โœ… Advantages:

  • Suitable for high-dimensional small-sample data (number of variables > number of samples)
  • Handles multicollinearity
  • Provides visualization (the score plot shows class separation)
  • Gives VIP for screening discriminant variables
  • More robust than LDA (linear discriminant analysis)

โš ๏ธ Limitations:

  • Assumes the boundary between classes is linear
  • Sensitive to class imbalance
  • Risk of overfitting (when there are too many components)
  • Requires strict evaluation by cross-validation

๐Ÿงญ OPLS (Orthogonal Partial Least Squares Regression) โ€‹

What is OPLS? โ€‹

OPLS (Orthogonal Projections to Latent Structures) was proposed by Trygg and Wold in 2002 and is an interpretation-improved version of PLS. Its core idea is:

Split the variation of X into two blocks โ€” "predictive variation related to Y" and "orthogonal variation unrelated to Y" โ€” and keep only the former for prediction.

In reality, industrial data often contains a lot of systematic fluctuation unrelated to quality: ambient temperature drift, equipment aging, batch-to-batch baseline drift, instrument state changes... These fluctuations make the components of PLS "go off track", and the score plot looks like a tangled mess.

OPLS first strips this noise out of X through orthogonal signal correction (OSC), and then runs PLS on the clean data.

Core principles โ€‹

1. Split X into two parts โ€‹

OPLS decomposes the variation of X into two parts:

  • (predictive variation): related to Y and carrying the true causal relationship, with only one predictive component
  • (orthogonal variation): unrelated to Y; it is systematic noise and is filtered out

2. Filter first, then regress โ€‹

The implementation path in the platform is:

Original X  โ†’  OPLS orthogonal correction  โ†’  denoised Z  โ†’  run PLS on Z  โ†’  predict Y
  1. Orthogonal correction: iteratively find the direction orthogonal to Y (), compute the orthogonal scores , and subtract them from X
  2. Obtain the net signal: after repeating N times, what remains is , containing only predictive information
  3. Build the regression: run standard PLS on to obtain the predictive model

The orthogonal weights are obtained iteratively from the following expression ( being the PLS weights):

3. โš ๏ธ What "number of components" means in the platform (important) โ€‹

This is the point where it is easiest to trip up:

ModelMeaning of "number of components (C+1 / C-1)"
PCANumber of principal components
PLS / PLS-DANumber of latent variables
OPLS / OPLS-DANumber of "orthogonal components" stripped off

๐Ÿ’ก Why? OPLS always has exactly 1 predictive component (which is also why it can draw a clean, single-direction score plot). When you click C+1 on OPLS, what increases is not "predictive power" but "one more layer of noise stripped off".

Therefore: C+1 on OPLS should not be understood as "the model became more complex", but as "the noise filtering became more thorough".

4. Equivalent coefficients (used for SHAP and the regression coefficient plot) โ€‹

Because the prediction of OPLS happens in the corrected Z space, restoring it to coefficients on the original variables requires a projection transformation:

Internally, the platform uses exactly these equivalent coefficients to draw the regression coefficient plot and to compute SHAP values, ensuring that the explanation corresponds to the original process variables rather than abstract latent variables.

Application in the platform โ€‹

Applicable scenarios โ€‹

  • Industrial data with strong noise: obvious baseline drift, environmental drift, batch systematic bias
  • Spectroscopic data: near-infrared, Raman, chromatography, and other data with scattering/baseline interference
  • Situations requiring clear explanation: when you need to explain to process or management people "which variable is actually at work"
  • Marker screening: combined with S-Plot to quickly pinpoint candidate markers

Key output metrics โ€‹

MetricMeaningInterpretation
RยฒXCumulative explained rate of XIncludes both the predictive and the orthogonal part
RยฒYCumulative explained rate of YThe higher the better
QยฒYCross-validated predictive abilityThe core metric of generalization ability; > 0.5 is acceptable
VIPVariable importance> 1 indicates an important variable
S-PlotCovariance vs. correlation coefficientVariables in the top-right / bottom-left corners are candidate markers

Dedicated chart: S-Plot โ€‹

OPLS has a "killer" chart that PLS does not have โ€” the S-Plot: the horizontal axis is covariance (size of contribution) and the vertical axis is the correlation coefficient (reliability). Variables falling in the top-right corner (positive correlation) or the bottom-left corner (negative correlation) both contribute a lot and are stable, making them ideal marker candidates.

See: S-Plot Marker Plot

Automatic trigger conditions in the platform โ€‹

OPLS is not triggered automatically; it must be selected manually in the data configuration:

X + Y (continuous values) + manually select OPLS โ†’ orthogonally corrected regression model

Pros and cons of OPLS โ€‹

โœ… Advantages:

  • Strong interpretability: a single predictive component makes the score plot and loading plot clearly one-directional instead of "a tangled mess"
  • Strips systematic noise: particularly effective against interference such as baseline drift and environmental drift
  • Retains all the advantages of PLS: collinearity, small samples, and multiple responses are all no problem
  • Comes with S-Plot: marker screening becomes more intuitive

โš ๏ธ Limitations:

  • Predictive ability is usually no better than PLS; the gain lies mainly in interpretability
  • The number of orthogonal components must be chosen carefully: stripping too much may throw away useful information along with the noise
  • The "business meaning" of orthogonal components is often unclear; do not over-interpret them
  • It is still a linear method

๐Ÿงญ OPLS-DA (Orthogonal Partial Least Squares Discriminant Analysis) โ€‹

What is OPLS-DA? โ€‹

OPLS-DA is simply the classification version of OPLS: first encode the class labels as dummy variables, then build the model with OPLS.

It is almost standard practice in fields such as metabolomics, food authenticity identification, and raw material grading, for a very direct reason โ€” in these scenarios, the within-group noise is often larger than the between-group difference. OPLS-DA can suppress the within-group noise and make the question "can the two classes really be separated" obvious at a glance.

Core differences from PLS-DA โ€‹

DimensionPLS-DAOPLS-DA
Predictive componentsMultiple latent variablesOnly 1 predictive component
Score plotRequires picking LV1/LV2 and may be skewed by noiseA single predictive axis, making between-group separation more intuitive
Noise handlingCannot be separated explicitlyExplicitly stripped into orthogonal components
S-PlotUsableMore suitable (the recommended usage)
Applicable dataGeneral classification dataClassification data with large within-group variation

Discriminant rule โ€‹

The same as PLS-DA: take the class with the largest response value:

Key output metrics โ€‹

MetricMeaningInterpretation
RยฒXCumulative explained rate of XIncludes both the predictive and the orthogonal part
AccuracyClassification accuracyWatch out for the class imbalance trap
F1 ScoreHarmonic mean of precision and recallMore reliable when classes are imbalanced
AUCArea under the ROC curve> 0.8 good, โ‰ฅ 0.9 excellent
Confusion matrixPredicted vs. actualSee which class is easily confused
S-PlotCovariance vs. correlationPinpoint the marker that distinguishes the two classes

Automatic trigger conditions in the platform โ€‹

X + Y (class labels) + manually select OPLS-DA โ†’ orthogonally corrected classification model

Pros and cons of OPLS-DA โ€‹

โœ… Advantages:

  • Between-group separation visualization is significantly better than PLS-DA
  • Few parameters (only the number of orthogonal components needs tuning)
  • Suitable for high-dimensional data with large within-group variation

โš ๏ธ Limitations:

  • Higher risk of overfitting: with too many orthogonal components it easily "learns the noise as a pattern", so Qยฒ / AUC must be checked repeatedly
  • Still sensitive to class imbalance
  • Still a linear boundary

๐Ÿ”„ Comparison and Selection of the Five โ€‹

Quick selection guide โ€‹

                        Start
                          โ”‚
                          โ–ผ
                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                 โ”‚  Is there a Y    โ”‚
                 โ”‚    variable?     โ”‚
                 โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                          โ”‚
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ–ผ                        โ–ผ
             No                       Yes
              โ”‚                        โ”‚
              โ–ผ                        โ–ผ
      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
      โ”‚      PCA       โ”‚    โ”‚   What type of Y?    โ”‚
      โ”‚    Explore     โ”‚    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
      โ”‚    structure   โ”‚               โ”‚
      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
                                       โ”‚
                     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                     โ–ผ                                   โ–ผ
                Continuous                         Class labels
                     โ”‚                                   โ”‚
                     โ–ผ                                   โ–ผ
         โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
         โ”‚   Obvious systematic   โ”‚          โ”‚   Obvious systematic   โ”‚
         โ”‚   noise / baseline     โ”‚          โ”‚   noise / baseline     โ”‚
         โ”‚   drift?               โ”‚          โ”‚   drift?               โ”‚
         โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
            โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
            โ–ผ                 โ–ผ                 โ–ผ                 โ–ผ
           No                Yes               No                Yes
            โ”‚                 โ”‚                 โ”‚                 โ”‚
            โ–ผ                 โ–ผ                 โ–ผ                 โ–ผ
    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
    โ”‚      PLS      โ”‚ โ”‚     OPLS      โ”‚ โ”‚    PLS-DA     โ”‚ โ”‚    OPLS-DA    โ”‚
    โ”‚  Regression   โ”‚ โ”‚  Explanation  โ”‚ โ”‚ Classificationโ”‚ โ”‚  Separation   โ”‚
    โ”‚  prediction   โ”‚ โ”‚  comes first  โ”‚ โ”‚ discriminant  โ”‚ โ”‚  comes first  โ”‚
    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Detailed comparison table โ€‹

FeaturePCAPLSPLS-DAOPLSOPLS-DA
Learning typeUnsupervisedSupervisedSupervisedSupervisedSupervised
Y variableNot neededContinuous valuesClass labelsContinuous valuesClass labels
Main purposeDimensionality reduction, explorationRegression predictionClassification and discriminationRegression + better interpretabilityClassification + better separation
Optimization objectiveMaximum X varianceMaximum X-Y covarianceMaximum class separationKeep predictive variation, strip orthogonalKeep predictive variation, strip orthogonal
Number of predictive componentsk principal componentsk latent variablesk latent variablesFixed at 1Fixed at 1
Meaning of C+1Add a principal componentAdd a latent variableAdd a latent variableAdd an orthogonal stripping layerAdd an orthogonal stripping layer
Key metricsRยฒXRยฒY, QยฒY, RMSEAccuracy, F1, AUCRยฒY, QยฒY, VIPAccuracy, F1, AUC
Signature chartsScore plot, loading plotPrediction plot, VIP plotROC, confusion matrixS-Plot, equivalent coefficientsS-Plot, confusion matrix
Sample requirementsNo restrictionMore samples than variables is betterBalanced samples per class is betterSame as PLSSame as PLS-DA
Main advantageNo labels neededStrong predictive powerThe standard approach to classificationResults easy to interpretClear separation between groups

Combined usage strategies โ€‹

In real projects, these models are often used in combination:

Scenario 1: explore first, then model

1. PCA explores the data structure โ†’ find anomalies, understand the distribution
2. Clean the data โ†’ remove abnormal samples
3. PLS / PLS-DA (or OPLS / OPLS-DA) modeling โ†’ build a prediction / classification model

Scenario 2: model diagnosis

1. PLS trains the model
2. Read the score plot with a PCA mindset โ†’ check the sample distribution
3. Combine with Tยฒ/SPE โ†’ identify abnormal samples

Scenario 3: variable screening

1. PLS / OPLS computes VIP
2. Remove variables with VIP < 0.5
3. Check the S-Plot with OPLS โ†’ pinpoint markers that both contribute a lot and are stable
4. Rebuild the model โ†’ simplify the model and improve generalization

Scenario 4: model upgrade (PLS โ†’ OPLS)

1. Build a PLS model and find that the score plot is "a tangled mess" and that Qยฒ is acceptable but hard to explain
2. Switch to OPLS โ†’ orthogonal components strip the systematic noise
3. Compare QยฒY: if it is basically unchanged while the plots are clearly cleaner โ†’ adopt OPLS
4. Then look at the S-Plot โ†’ find candidate markers and go back to business validation

๐Ÿ› ๏ธ Modeling Practice in the Platform โ€‹

Modeling workflow โ€‹

Parameter tuning tips โ€‹

Choosing the number of components / latent variables โ€‹

The platform provides the C+1/C-1 buttons for manual adjustment:

SymptomCauseSolution
High Rยฒ, low QยฒOverfittingReduce the number of components
Both Rยฒ and Qยฒ lowUnderfittingIncrease the number of components
Qยฒ decreases as components are addedNoise is being introducedChoose the number of components at the Qยฒ peak

โš ๏ธ Note for OPLS / OPLS-DA: here the "number of components" is the number of orthogonal components. The criterion is still to look at QยฒY (or AUC) โ€” if adding too many orthogonal components causes QยฒY to drop, then you have started to lose useful information.

Cross-validation settings โ€‹

  • K-Fold: used when the sample size is large (default 7 folds, adjustable in the system settings)
  • Leave-one-out (LOO): used when the sample size is small
  • Random seed: fix the seed to guarantee reproducible results

See: Model Training Settings

Model diagnosis checklist โ€‹

After training a model, check it against the following lists:

General checks (all models):

  • [ ] Whether RยฒX is reasonable (> 0.5 is usually acceptable)
  • [ ] Whether there are obvious outliers in the score plot
  • [ ] Whether Tยฒ/SPE exceed their limits

Checks specific to PLS / OPLS:

  • [ ] QยฒY > 0.5 (the minimum threshold)
  • [ ] Gap between RยฒY and QยฒY < 0.2 (overfitting prevention)
  • [ ] Whether the high-VIP variables make business sense
  • [ ] Whether the prediction scatter plot is distributed along the diagonal

Checks specific to PLS-DA / OPLS-DA:

  • [ ] Accuracy > 0.8 (depending on the difficulty of the task)
  • [ ] A reasonable F1 Score (a must-check when classes are imbalanced)
  • [ ] AUC > 0.8
  • [ ] Whether some class performs particularly badly in the confusion matrix
  • [ ] Whether the classes are clearly separated in the score plot

Extra checks for the OPLS family:

  • [ ] After switching to OPLS, QยฒY (or AUC) has not dropped noticeably โ€” otherwise useful information was stripped away
  • [ ] Whether the high-contribution variables at both ends of the S-Plot are consistent with business understanding
  • [ ] Whether there are too many orthogonal components (usually 1~3 is enough)

๐Ÿ“š Further Reading โ€‹

If you want to understand these algorithms more deeply, the following resources are recommended:

Classic literature:

  • Wold, S. et al. (2001). PLS-regression: a basic tool of chemometrics
  • Trygg, J. & Wold, S. (2002). Orthogonal projections to latent structures (O-PLS)
  • Bylesjรถ, M. et al. (2006). OPLS discriminant analysis: combining the strengths of PLS-DA and SIMCA classification

Related charts in the platform:


๐Ÿ’ก Summary โ€‹

ModelUnderstanding it in one sentenceWhen to use it
PCAWhat does the data look like?Structure exploration, dimensionality reduction, anomaly detection
PLSHow does X predict Y?Regression problems, building a prediction equation
PLS-DAWhich class does it belong to?Classification problems, discriminant analysis
OPLSHow does X predict Y, and where is the noise?Regression with systematic noise that needs clear explanation
OPLS-DAWhich class does it belong to, and what is the within-group noise?Classification with large within-group variation, marker screening

Master these five models and you have mastered the core of StarWay Data Insight. Remember: models are tools; business understanding is the soul. Good analysis = the right model + clean data + deep domain knowledge.

May your data exploration journey go smoothly! ๐Ÿš€

Let data speak, make decisions simpler.