Machine Learning for Deep Eutectic Solvent Density: Impact of Feature Representations and Dataset Complexity on Predictive Reliability

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Accurately predicting the density of deep eutectic solvents (DESs) is crucial for optimizing green separation processes. This study investigated the impact of feature representations (ChemBERTa, hybrid, and critical property models) and data partitioning (binary, ternary, and comprehensive datasets) on machine learning predictions. Evaluating RF, XGBoost, CatBoost, and ANN models revealed that tree-based ensembles were highly robust, consistently achieving R > 0.93 on limited datasets. Conversely, ANNs required explicit physical descriptors or massive datasets (>12,000 points) to prevent overfitting. Cross-domain validations demonstrated that extrapolating from simple to complex systems failed due to restricted thermodynamic diversity, whereas specializing from a comprehensive dataset ensured excellent transferability. These findings established that combining large, diverse datasets with ensemble algorithms or physics-informed features were essential for the reliable computational design of multicomponent DES properties.

키워드

Deep eutectic solventDensityMachine learningEnsembleFeatureLIQUID-LIQUID EQUILIBRIA
제목
Machine Learning for Deep Eutectic Solvent Density: Impact of Feature Representations and Dataset Complexity on Predictive Reliability
저자
Park, YoonKook
DOI
10.9713/kcer.2026.64.2.105163
발행일
2026-05
유형
Article
저널명
Korean Chemical Engineering Research(HWAHAK KONGHAK)
64
2
페이지
1 ~ 7