Hybrid Support Vector Regression Model and K-Fold Cross Validation for Water Quality Index Prediction in Langat River, Malaysia
Mamat, N.; Hamzah, F. M.; Jaafar, O.
Show abstract
Water quality analysis is an important step in water resources management and needs to be managed efficiently to control any pollution that may affect the ecosystem and to ensure the environmental standards are being met. The development of water quality prediction model is an important step towards better water quality management of rivers. The objective of this work is to utilize a hybrid of Support Vector Regression (SVR) modelling and K-fold cross-validation as a tool for WQI prediction. According to Department of Environment (DOE) Malaysia, a standard Water Quality Index (WQI) is a function of six water quality parameters, namely Ammoniacal Nitrogen (AN), Biochemical Oxygen Demand (BOD), Chemical Oxygen Demand (COD), Dissolved Oxygen (DO), pH, and Suspended Solids (SS). In this research, Support Vector Regression (SVR) model is combined with K-fold Cross Validation (CV) method to predict WQI in Langat River, Kajang. Two monitoring stations i.e., L15 and L04 have been monitored monthly for ten years as a case study. A series of results were produced to select the final model namely Kernel Function performance, Hyperparameter Kernel value, K-fold CV value and sets of prediction model value, considering all of them undergone training and testing phases. It is found that SVR model i.e., Nu-RBF combined with K-fold CV i.e., 5-fold has successfully predicted WQI with efficient cost and timely manner. As a conclusion, SVR model and K-fold CV method are very powerful tools in statistical analysis and can be used not limited in water quality application only but in any engineering application.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Spatio-temporal modelling for the evaluation of an altered Indian saline Ramsar site and its drivers for ecosystem management and restoration 96%
- Impact of Grain for Green Project on Water Resources and Ecological Water Stress in Yanhe River Basin 96%
- Prediction of Direct Carbon Emissions of Chinese Provincial Residents under Artificial Neural Networks in Deep Learning Environment 96%
Similar papers in this journal
- Assessment of Environmental Factors Associated with Antibiotic Resistance Genes (ARGs) in the Yangtze Delta, China 94%
- Comparing protein-protein interaction networks of SARS-CoV-2 and (H1N1) influenza using topological features 93%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 93%
Similar papers in this journal
- Physiological properties and tailored feeds to support aquaculture of marbled crayfish in closed systems 89%
- Disrupting the biodiversity - ecosystem function relationship: response of shredders and leaf breakdown to urbanization in Andean streams 87%
- Assessment of microphytobenthos communities in the Kinzigcatchment using photosynthesis-related traits, digital light microscopy and 18S-V9 amplicon sequencing 87%
Similar papers in this journal
- A computational study on the role of parameters for identification of thyroid nodules by infrared images (and its comparison with real data) 93%
- Measuring Repositioning in Home Care for Pressure Injury Prevention and Management 91%
- On the prediction of arginine glycation using artificial neural networks 90%
Similar papers in this journal
- Urban Vulnerability Assessment for Pandemic surveillance: The COVID-19 case in Bogotá, Colombia 93%
- The effects of biodegradable mulch film on the growth, yield, and water use efficiency of cotton and maize in an arid region 93%
- Taranto’s long shadow? Cancer mortality shows alarming peaks for specific types in the most polluted city of Italy but also in surrounding towns 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.