Handling Multicollinearity in Logistic Regression via PCA, ICA, and GPCA: Theory, Simulation and Application
DOI:
https://doi.org/10.67868/wzz4d456Parole chiave:
Principal Component Analysis, Independent Component Analysis, Generalized Principal Component Analysis, Multicollinearity, Logistic regression, Dimension reductionAbstract
Multicollinearity poses a serious challenge in logistic regression analysis, leading to inflated standard errors, unstable coefficient estimates, and unreliable statistical inference. This study investigates the comparative performance of three dimension reduction techniques — Principal Component Analysis (PCA), Independent Component Analysis (ICA), and Generalized Principal Component Analysis (GPCA) — as remedies for multicollinearity in binary logistic regression. The theoretical framework establishes the asymptotic properties of maximum likelihood estimators under transformed predictors, including consistency, Fisher information, and asymptotic normality via the Central Limit Theorem. A Monte Carlo simulation study was conducted using eight predictor variables across five sample sizes (n = 50, 100, 500, 1000, 10000) and three levels of inter-predictor correlation (ρ = 0.50, 0.75, 0.95), with 2, 3, and 4 retained components evaluated under each scenario. Model performance was assessed using the Akaike Information Criterion (AIC), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and Mean Absolute Deviation (MAD). The methods were further validated on a real-life dataset of automobile characteristics from the STATA repository, comprising nine predictor variables exhibiting substantial multicollinearity as confirmed by Variance Inflation Factor (VIF) analysis. Simulation results demonstrate that GPCA consistently achieves superior predictive accuracy over PCA and ICA across varying sample sizes, correlation levels, and numbers of retained components, as evidenced by lower MSE, RMSE, and MAD values. PCA performs competitively under the AIC criterion and is recommended when small samples constrain model complexity. Real-life data analysis corroborates these findings, with GPCA outperforming competing methods on three out of four selection criteria. These results affirm GPCA as the preferred dimension reduction technique for logistic regression in the presence of multicollinearity, with PCA serving as a viable alternative in small-sample settings.
Riferimenti bibliografici



