BWIQ is B&W Tek's chemometric software
designed for both on and off -line quantitative and qualitative spectroscopy measurements
Today we are going to demonstrate how to build a classification model with Raman spectral data
In this video, we will import the spectra into the BWIQ software
define the class for each sample
classify the spectra using a Principal Component Analysis-Mahalanobis Distance algorithm
and finally use the model to classify new samples
In classification models, spectra of samples corresponding to specific groups are collected
and then modeled according to their similarity to form a class
For this example, we will create a model using the Raman spectra of amino acids l-alanine
l-aspartic acid and l-cysteine hydrochloride
that were collected on B&W Tek's NanoRam handheld Raman system
20 spectra were collected for each amino acid
To import the spectra into the BWIQ software, click File - Import
and select the file format you wish to use
Locate the data set and import the files into BWIQ
The data files now appear in the Spectra window along with the corresponding spectra
When developing classification models, we need to define the class for each spectrum in the software
Click the blue addition button in the upper left corner of the Spectra window
a third column will appear
You can rename the column as "Class"
by right-clicking the header and selecting "Rename" from the dropdown menu
Here, we will enter an integer value for each distinct class
so that all aspartic acid spectra are designated class 1
all alanine spectra are class 2
and all l-cysteine hydrochloride spectra are class 3
There are several ways to designate spectra files as calibration or validation samples in BWIQ
Spectra files can be manually designated as calibration, validation, or ignored files
by clicking the drop-down button from the "Usage" column
Files designated "Ignored" will not be included in the final analysis
There are also several sampling algorithms available in the BWIQ software
which can be selected under "Sampling"
The parameters for the algorithms selected will appear in the algorithm properties panel
In the algorithm properties panel, you can change the ratio of calibration files to validation files
for instance, designating 60% of the files as calibration files, and 40% as validation files
To apply the sampling algorithm to the data
click the blue "execute" triangle in the algorithm properties window
the usage column in the Spectra window
will show which spectra are designated as calibration samples and which are validation samples
For this type of qualitative analysis, we designate all files as calibration files
Next, we can add pre-processing steps to remove variation in the data set
that is not related to chemical differences
but instead may result from scattering, instrumental variation, spectral noise, or background differences
The Raman data presented were collected using an automatic integration time feature
so the integration times for each spectrum are noticeably different
To normalize the data intensities, we will use a standard normal variate normalization algorithm
Click Pre-Process, then Standard Normal Variate
The normalization now appears in the algorithms window
When building classification models, an important step is to mean center the data
as we are interested in the difference of the data from a centered point
not how far the samples are from that center
Click "Pre-processing" then "Center"
The step is now added to the algorithms window
Manual variable selection enables us to use the entire spectrum for analysis
or alternatively to restrict the analysis to selected regions of the data
Using the whole spectrum for analysis will allow the model to be more sensitive to contaminants
or changes in the samples that introduce signal in other spectral regions
However, it can be helpful to remove non-informative or noisy regions of the data from the analysis
For example, in this data we can see that above 1800 wave numbers
there is little Raman signal
Under Variable, choose Manual Selection
A Manual Variables selection is automatically added to the selected algorithms window
Now that we have finished adding pre-processing steps
we can add a classification algorithm
For this data, we will select a Principal Component Analysis Mahalanobis Distance classification method
Principal Component Analysis, or PCA
is an excellent exploratory chemometric tool
that uses a reduced variable space to define the greatest variance within a data set
With PCA, model scores and loadings are computed
and the 2D scores plot provides a visualization of natural groupings of samples
based on the similarity of the data
With each sample represented by a single point in the new Principal Component space
To classify new samples with the model, we'll use Principal Component Analysis
in combination with the parameter known as the Mahalanobis Distance
which measures the distance of a new sample to the center of each class
The Mahalanobis Distance for each new sample is calculated for every specified class
The sample is classified by the software as the group that corresponds to the shortest calculated Mahalanobis Distance
Choose Classification, and then PCA-MD
Click the blue addition button in the Algorithm Properties window
to add the PCA-MD algorithm to the model
The algorithms that are listed in the selected algorithms window
now make up the steps of our classification model
To save this training file with the extension .train
Go to File, then Save As, and save the file
Training files can be edited and saved at any time in Method Development
To execute the model, click the blue triangle at the top of the window
This will perform the steps in sequence
Because our model includes a Manual Variable Selection
The manual variables window appears upon execution of the model
In this window, we can type in the the spectral regions of interest
or use the cursor to define the analysis regions
Because there is no significant Raman signal above 1800 wave numbers
for any of these amino acids in this data set
we will limit our analysis from 200-1800 wave numbers in these spectra
The data shown have been pre-processed and are now mean centered
with peaks in both directions around the zero line
Upon execution of the method, a new model file is created
The model file is a .cmml file that can be saved and used to classify new samples on-line
while connected to a B&W Tek i-Raman series instrument
or offline using previously collected data
To view the scores plots, click the Scores icon in the toolbar
The scores plot shows the clusters of the samples in principal component space
In this plot, we observe three distinct, well separated clusters
These three clusters correspond to the three different classes of amino acids
Based on how close they fall to these three clusters
we can classify new spectra as belonging to one of these groups
There are several other plots available to view data
including loadings and variance
To classify unknown samples from previously acquired data
click Predict, and load the prediction files in the correct format
The prediction results list the individual spectral file names
along with the calculated Mahalanobis distances from each class in the model
The column to the right labeled Class shows the final classification for the sample
In this case, either one, two, or three
The class showing the lowest calculated Mahalanobis distance
is the class chosen by the BWIQ software to represent that sample
Under Generate PDF Report, you can also save a PDF report
which details the method parameters and the classification predictions
For more information on using the BWIQ software, please refer to the user manual
as well as our other support videos
You can also reach us at www.bwtek.com/support
Thank you for watching!
Không có nhận xét nào:
Đăng nhận xét