On Some Aspects of Variable Selection for Partial Least Squares Regression Models

Wiley - Tập 27 Số 3 - Trang 302-313 - 2008
Partha Pratim Roy1, Kunal Roy1
1Drug Theoretics and Cheminformatics Laboratory, Division of Medicinal and Pharmaceutical Chemistry, Department of Pharmaceutical Technology, Faculty of Engineering and Technology, Jadavpur University, Kolkata 700 032, India

Tóm tắt

AbstractThis paper tries to explore the optimum variable selection strategy for Partial Least Squares (PLS) regression using a model dataset of cytoprotection data. The compounds of the dataset were classified using K‐means clustering technique applied on standardized descriptor matrix and ten combinations of training and test sets were generated based on the obtained clusters. For a particular training set, PLS models were developed with a number of components optimized by leave‐one‐out Q2 and then the developed models were validated (externally) using the test set compounds. For each set, PLS model was initially constructed using all descriptors (variables). The variables having least standardized values of regression coefficients were deleted and the next model was developed with a reduced set of variables. These steps were performed several times until further reduction in number of variables did not improve Q2 value. In each case, statistical parameters like predictive R2 (R2pred), squared correlation coefficient between observed and predicted values with (r2) and without ($\rm{ r_0^{\rm{2}} }$) intercept and Root Mean Square Error of Prediction (RMSEP) were calculated from the test set compounds. In case of all ten sets, Q2 values steadily increase on deletion of variables while R2pred values do not show any specific trend. In no case, the highest Q2 and highest R2pred appear in the same trial, i.e., with the same combinations of variables. This suggests that from the viewpoint of external predictability, choice of variables for PLS based on Q2 value may not be optimum. Moreover, a clear separation of r2 and r02 curves in some sets suggests that such models may not be truly predictive in spite of acceptable R2pred values. Another observation is that coefficient of determination R2 for the training set is more immune to changes on deletion of variables than the validation parameters like Q2 and R2pred. Finally, a new parameter rm2 has been suggested to indicate external predictability of QSAR models.

Từ khóa


Tài liệu tham khảo

10.1002/cem.1180020306

L. Eriksson E. Johansson N. Kettaneh‐Wold S. Wold Multi‐ and Megavariate Data Analysis: Principles and Applications Umetrics Umeå2001.

Wold S., 1995, Chemometric Methods in Molecular Design, 195

Selassie C. D., 2003, Burger's Medicinal Chemistry and Drug Discovery, 1

10.1023/A:1025386326946

Golbraikh A., 2002, Mol. Divers., 5, 231, 10.1023/A:1021372108686

10.1016/S1093-3263(01)00123-1

10.1021/ci980033m

10.1080/10629360412331319808

10.1016/j.jmgm.2005.03.003

10.1021/ci0497511

10.1002/qsar.200510153

10.1002/qsar.200430909

10.1016/j.jmgm.2006.06.005

Gramatica P., 2007, QSAR Comb. Sci.

10.1021/ci025626i

10.1002/qsar.200510161

10.1007/s10822-004-5202-8

10.1021/jm030584q

10.1023/A:1025366721142

Wold S., 1995, Chemometric Methods in Molecular Design, 312

Debnath A. K., 2001, Combinatorial Library design and Evaluation, 73

10.1007/s10822-007-9102-6

10.1021/jm0005151

10.1021/jm049252r

Cerius2version 4.8 is a product of Accelrys Inc. San Diego USA http://www.accelrys.com/cerius2.

B. S. Everitt S. Landau M. Leese Cluster Analysis Edward Arnold London2001.

Kowalski R. B., 1982, Handbook of Statistics

Downs G. M., 1995, Advanced Computer Assisted Techniques in Drug Discovery, 111

MINITAB is a statistical software of Minitab Inc. USA.