Journal of Korean Society of Agricultural Engineers. 2021. 51–60
https://doi.org/10.5389/KSAE.2021.63.2.051

ABSTRACT


MAIN

Ⅰ. Introduction

The temporal variability of water quality parameters is information critical for efficient watershed management (Montgomery et al., 2007; Horsburgh et al., 2009). NPS pollutants are loaded to waterbodies with rainfall events that may last several days. In addition, the transport of NPS pollutants is complicated with hydrological processes that greatly vary over time, even in an event. Thus, a frequent monitoring is often required when assessing NPS pollution and developing water quality management plans (Scholefield et al., 2005; Houser et al., 2006; Jordan et al., 2007). However, frequent NPS monitoring is often prohibitive in terms of time and effort.

The use of high-frequency water quality sensors can be one of the solutions for the cost issue of NPS monitoring, but many of water quality parameters such as suspended solids (SS) and total phosphorus (T-P) still require extensive equipment and lab tests to be quantified (Jones et al., 2011, Villa et al., 2019). In this context, surrogate monitoring can be an efficient alternative. Previous studies have tried to predict the water quality parameters with accessible parameters such as pH, dissolved oxygen (DO), temperature using traditional statistical techniques such as multiple linear regression (MLR) (Çamdevýren et al., 2005; Chenini and Khemiri, 2009; DeForest et al., 2018). Some of them showed that SS and T-P is closely associated with turbidity and discharge rates of a river (Jones et al., 2011; Ziegler et al., 2011). It is a relative measure of light diffraction in water caused by particulate form, and the use of turbidity as a surrogate parameter to estimate T-P is based on that T-P transported in streams predominantly particulate form (Settle et al., 2007; Ruegner et al., 2013). Moreover, Johanna et al. (2014) predicted SS, particulate organic carbon, and particulate nitrogen from turbidity, flow discharge rate, and rainfall depth using MLR. Although, MLR is a simple statistical method to use (Jones et al., 2011; Ziegler et al., 2011; Johanna et al., 2014), and its application is limited to linear relationships, and it is known to be very sensitive to outliers (Chatterjee and Hadi, 1988).

The use of machine learning techniques such as random forest (RF) became affordable with recent advances made in computing resources and techniques, and they have been employed as an analysis tool for hydrologic research (Rodriguez-Galiano et al., 2014; Singh et al., 2017). The RF is a computational algorithm that can predict the values of continuous variables from continuous and/or categorical predictor variables by constructing a decision tree or classifying data (Razi et al., 2005). The algorithm is known not to overfit and efficient in predicting irregular variables that exhibit little periodicity (Diaz-Uriate and de Andres, 2006).

This study explored the potential of the RF algorithm as a tool to efficiently identify primary variables that can serve as surrogate for NPS pollutants. We developed prediction models using MLR (traditional statistical model) and RF (the latest statistical data mining technique) and compared their performance to demonstrate their capacities and potentials. This study focused on SS and T-P that are common NPS pollutants in Korea.

Ⅱ. Material and methods

1. Study areas and NPS monitoring

This study has monitored SS and T-P at the outlet of a catchment, Wol-jeong (WJ), located within the Pungyeongjeongcheon watershed in Korea for two years, from 01/01/2017 and 12/31/2018 (Fig. 1). WJ is mostly covered by agricultural land uses (86%) including rice paddy fields and upland. The mean annual temperature and precipitation of the study catchment are 13.8°C and 1,391 mm, respectively.

https://cdn.apub.kr/journalsite/sites/jksae/2021-063-02/N0740630206/images/PIC95DA.png
Fig. 1

Locations of the study watersheds and NPS pollutant monitoring stations

A set of water pressure sensor and logger (OTT Orpheus mini, Germany) was installed to monitor the level of streamflow, and the velocity (VALEPORT model 002, UK) and cross section of flow were measured six times during the monitoring period. The flow level measurements were converted to discharge data using a flow rating curve developed for the outlet of the study catchment. A digital turbidity sensor (FTS DTS-12, Canada) was installed to monitor the turbidity of flow at the monitoring site (or the outlet). The flow duration curve was then developed from flow observations and the flow conditions was seperated using Table 1. Then, we separated the flow conditions into high-flow periods (the exceedance probability equal to or less than 10%) and low-flow ones (the exceedance probability greater than 10%) so that the flow-dependent behavior of NPS pollutant transport can be examined in the study. Stormwater samples were taken every hour in the first 24 hours of a storm event using an automatic sampler (ISCO portable sampler 6712, USA), and the sampling interval was increased to 6 hours for the rest of the event. The concentrations of SS and T-P were analyzed in a laboratory according to the standard water pollution test method (APHA, 2001). The monitored water quality variables and predictors such as turbidity and runoff discharge rates are summarized in Table 2.

Table 1

Classifications of hydrologic conditions

Hydrologic condition classFlow duration interval (%)
High flows0~10
Moist conditions10~40
Mid-range conditions40~60
Dry conditions60~90
Low flows90~100
⁕ Source: USEPA, 2007
Table 2

Summary statistics of predictors and variables to predict SS and T-P

VariablesUnitMin.MedianMeanMax.Standard deviationNumber of samples
SSmg/L1.126.535.3376.441.7134
T-Pmg/L0.0670.1280.1430.7700.083134
TurbidityNTU6.629.142.9383.048.5134
Runoffm3/s0.40.80.910.80.7134

2. Surrogate model development

The potential relationships between the pollutants of interest (SS and T-P) and the primary environmental variables (discharge rates and turbidity) were explored using the MLR and RF methods. RF is the combination of tree-based classification methods such as a regression tree developed by Breiman, (2001). The processes of growing regression trees can be summarized into two steps. First, data of interest are split into multiple tree branches in the way toward minimizing least-square deviations between observed and predicted variables, which is called the optimal split (OPs) (Hasanipanah et al., 2017). The split is calculated as the difference of variance between the mother node and the left child node, and the right child node is maximized by using Eqs. 1 and 2 (Ließ et al., 2012; Granata et al., 2017).

(1)
Rt=1/nt×i=1nyit-y¯t2
(2)
OPs=Rt-Rtl+Rtr

where R(t) is the variance in a node t, which is divided into the left and right child nodes, tl and tr. n is the number of observations, and https://cdn.apub.kr/journalsite/sites/jksae/2021-063-02/N0740630206/images/PIC960A.gif is the mean value of the predicted variable.

Once tree branches are constructed in the first step, RF prunes the trees by removing nodes that have low explanatory power. Here, the existence of nodes that may be removed imply overfitting. An overfitted model does not sensitively respond to newly added data, and it creates additional uncertainty to the prediction (Breiman et al., 1984). Thus, the pruning process is necessary to reduce prediction uncertainty. The overall procedures of predicting SS and T-P using the MLR and RF models are presented in Fig. 2.

https://cdn.apub.kr/journalsite/sites/jksae/2021-063-02/N0740630206/images/PIC962C.png
Fig. 2

Overall procedures of predicting SS and TP using random forest and multiple linear regression

The RF uses a bagging technique with a randomized subset of variables, which is mathematically described using Eq. 3 (Prasad et al., 2006).

(3)
f^x=1Kk=1KTx

where https://cdn.apub.kr/journalsite/sites/jksae/2021-063-02/N0740630206/images/PIC961A.gif is a RF prediction, K is the number of trees (or branches), and T(x) is the result of each regression tree.

The node size (five) was adapted from literature (Wang et al., 2015) and the’mtry’ is the number of predictors to select at random for each split in the tree model. However, we used only two variables in this study, so we used’mtry’ of two. The RF model calculates the relative importance of each parameter using the Gini impurity (GI) and mean square error (MSE) statistics. GI quantifies the quality of each split, and MSE calculates the mean decrease of prediction accuracy (Breiman, 2001; Ouedraogo et al., 2018). The GI method is used to calculate the quality of each split for the variables in a tree model and calculate the mean decrease using MSE due to splits on every predictor (Breiman, 2001). In this study, we found that K of 100 is enough to stabilize the MSE. However, the running time of RF slightly increased until K value hit 1,000. We selected K value of 1,000 while still providing reasonably short computing time (Fig. 3). The relative importance was calculated by running the RF algorithm 100 times and then averaging the importance rates calculated from the iterations.

https://cdn.apub.kr/journalsite/sites/jksae/2021-063-02/N0740630206/images/PIC961B.png
Fig. 3

Responses of MSE to the number of trees (K)

It is known that one of the considerations in the machine learning technique is that training variables with different range affects the learning rates and it can be crucial problems with prediction performance (Ioffe and Szegedy, 2015). For example, turbidity ranges from 6.6 to 383 NTU, and runoff ranges from 0.6 to 3.2 m3/s in this study. Then, the turbidity and runoff discharge rate data were normalized into the range from 0 to 1 using Eq. 4 to assure that each variable can have the similar level of explanatory power in the analysis.

(4)
x`=x-minxmaxx-minx

where https://cdn.apub.kr/journalsite/sites/jksae/2021-063-02/N0740630206/images/PIC965C.gif is the normalized value of the data set (range of 0 to 1) and https://cdn.apub.kr/journalsite/sites/jksae/2021-063-02/N0740630206/images/PIC966D.gif is the original value.

The flow observations made from 2017 to 2018 showed that the flow discharge rate of 1.08 m3/s correspond to the 10% discharge of the study catchment (Fig. 4). The threshold rate of 1.08 m3/s was used to separate discharge into high and low flow, and then the models were separately applied to the two flow regimes.

https://cdn.apub.kr/journalsite/sites/jksae/2021-063-02/N0740630206/images/PIC966E.png
Fig. 4

Daily runoff and flow duration curve of the study site

The performance of the MLR and RF models was evaluated with the split-sample scheme of 80/20 (80% for training and 20% for validation; Afendras and Markatou, 2019). To validate the models, we used four different statistics and criteria: the Nash-Sutcliffe efficiency (NSE) coefficient, the coefficients of determination (R2), p-value (P), and the Akaike information criterion (AIC) (Nash and Sutcliffe, 1970; Akaike, 1973; Snipes and Taylor, 2014).

(5)
NSE=1-i=1nXi-Yi2i=1nYi-Y¯2
(6)
R2=i=1nXi-Y¯2i=1nYi-Y¯2
(7)
AIC=2k-2lnL
(8)
P=PrT-tH+PrT+tH

where Xi and Yi is the predicted value and observed values from site i and n is the number of samples. In the AIC method, k represents the number of parameters, and L is the likelihood function, which represent the model fit and it should represented by log-value to limit the maximum value of model fit.

A lower AIC score is interpreted as a more accurate model. The p-value is defined as the probability (Pr), under the null htpothesis H about the unknown distribution of the T. If the p-value shows very small value, then the statistical significance is thought to be very large. On the other hand, the criteria commonly set to 0.05, 0.01, 0.005, or 0.001 and we choose 0.001 as criteria in this study.

Ⅲ. Results and discussion

The comparison between observation and prediction result was conducted with normalized value, and the normalized SS and T-P prediction results of the MLR and RF models are compared with the normalized observation values in Fig. 5 and Table 3. Overall, the MLR model prediction of SS and T-P in the low and high flow conditions were acceptable with R2 and NSE greater than 0.5, except T-P in low flow condition. The T-P prediction for the low flow condition shows acceptable R2 but low NSE. The reason for this result is that T-P in agricultural area shows high concentrations for high flow conditions during farming period (Villa et al., 2019). Thus, high flow conditions shows acceptable prediction performance with high concentrations of T-P which shows high fluctuations and low flow condition shows unacceptable prediction performance due to low concentrations of T-P. In general, SS and T-P were more accurately predicted by the MLR models when runoff discharge is relatively large (or the high flow condition). The RF model provided accuracy and its patterns similar to that of the MLR model.

Table 3

Performance evaluation of the MLR and RF models in predicting SS and T-P concentrations

ModelIndicatorLow flow conditionHigh flow condition
SST-PSST-P
MLRR2 0.74 0.61 0.83 0.72
NSE 0.68 0.48 0.81 0.70
AIC-35.8-48.6-49.2-58.3
P-value< 0.001< 0.001< 0.001< 0.001
RFR2 0.77 0.69 0.96 0.85
NSE 0.73 0.42 0.95 0.81
AIC-67.4-48.6-73.9-53.2
P-value< 0.001< 0.001< 0.001< 0.001
https://cdn.apub.kr/journalsite/sites/jksae/2021-063-02/N0740630206/images/PIC967E.png
Fig. 5

Comparison between the NPS pollutants observed and simulated using the MLR and RF models (a) SS using MLR in low flow condition; (b) SS using RF in low flow condition; (c) T-P using MLR in low flow condition; (d) T-P using RF in low flow condition. (e) SS using MLR in high flow condition; (f) SS using RF in high flow condition; (g) T-P using MLR in high flow condition; (h) T-P using RF in high flow condition

As both models provided similar accuracy when predicting SS and T-P from turbidity and runoff discharge rates, their performance was further investigated in terms of AIC values (Table 3). The AIC values of the RF model are lower than those of the MLR model, indicating that the RF model is more suitable for predicting SS and T-P that the MLR model in this study.

The relative importance of the primary environmental variables, runoff discharge rates and turbidity, to predict SS and T-P was quantified as part of the RF modeling (Fig. 6). The analysis of relative importance showed that turbidity was more closely associated with the SS and T-P concentrations than discharge rates with importance rates ranging from 52% to 86% depending on the flow conditions. As mentioned above, the runoff in high flow condition shows higher relative importance than low flow condition, which is related to the high concentrations of T-P in farming period.

https://cdn.apub.kr/journalsite/sites/jksae/2021-063-02/N0740630206/images/PIC96AE.png
Fig. 6

Relative importance rates of runoff discharge rates and turbidity in predicting SS and T-P concentrations

Ⅳ. Conclusions

This study explored the potential of the MLR and RF model as a tool to identify surrogate environmental variables in two different flow conditions for the concentrations of NPS pollutants such as SS and T-P in agricultural sub-watershed. The performance of the MLR and RF models was evaluated with the 80% for training set and 20% for validation set. Both MLR and RF models provided acceptable performance when predicting SS and T-P in the high-flow condition. The relative importance rate analysis showed that turbidity had relatively high explanatory power (52% to 86%) than runoff discharge rates. Moreover, the runoff in high flow condition shows higher relative importance than low flow condition, which is related to the high concentrations of T-P in farming period. This study demonstrated the machine learning techniques could help to improve the efficiency of NPS pollutant monitoring and prediction by identifying fundamental environmental variables that are relatively easily obtained but still closely related to the NPS pollutants. The results also showed that turbidity could serve as a surrogate water quality parameter for SS and T-P concentrations at acceptable accuracy levels.

감사의 글

본 연구는 BK21 기후지능형간척지농업 교육연구팀 장학금 지원과 정부(교육부)의 재원으로 한국연구재단의 지원을 받아 수행된 기초연구사업임 (No. 2017R1D1A3B03029713).

REFERENCES

1
G. Afendras and M. Markatou, Journal of Statistical Planning and Inference, Optimality of training/test size and resampling effectiveness in cross-validation, 199; 286-301 (2019)10.1016/j.jspi.2018.07.005
2
H. Akaike, In Second International Symposium on Information Theory, Information theory and an extension of the maximum likelihood principle; 267-281 (1973)
3
American Public Health Association (APHA), 21st ed. Standard Methods for the Examination of Water and Waste Water, USA. Washington DC. (2001)
4
L. Breiman, Machine Learning, Random forests, 45; 5-32 (2001)
5
L. Breiman, J. Friedman, R. A. Olshen and C. J. Stone, Classification and Regression Trees, Wadsworth. CRC press. (1984)
6
H. Camdevyren, N. Demyr, A. Kanik and S. Keskyn, Ecological Modelling, Use if principal component scores in multiple linear regression models for prediction of Chlorophyll-a in reservoirs, 181(4); 581-589 (2005)10.1016/j.ecolmodel.2004.06.043
7
S. Chattergee and A. S. Hadi, Sensitivity Analysis in Linear Regression, USA. Wiley. (1988)
8
I. Chenini and S. Khemiri, International Journal of Environmental Science and Technology, Evaluation of ground water quality using multiple linear regression and structural equation modeling, 6(3); 509-519 (2009)10.1007/bf03326090
9
D. K. DeForest, K. V. Brix, L. M. Tear and W. J. Adams, Environmental Toxicology and Chemistry, Multiple linear regression models for predicting chronic aluminum toxicity to freshwater aquatic organisms and developing water quality guidelines, 37(1); 80-90 (2018)10.1002/etc.3922
10
R. Diaz-Uriate and S. A. de Andrés, BMC Bioinformatics, Gene selection and classification of microarray data using random forest, 7; 3 (2006)10.1186/1471-2105-7-3
11
F. Granata, S. Papirio, G. Esposito, R. Gargano and G. Marinis, Water, Machine learning algorithms for the forecasting of wastewater quality indicators, 9(2); 105 (2017)10.3390/w9020105
12
M. Hasanipanah, R. S. Faradonbeh, H. B. Amnieh, D. J. Armaghani and M. Monjezi, Engineering with Computers, Forecasting blast-induced ground vibration developing a CART model, 33; 307-316 (2017)10.1007/s00366-016-0475-9
13
J. S. Horsburgh, A. S. Jones, D. K. Stevens, D. G. Tarboton and N. O. Mesner, Environmental Modelling and Software, A sensor network for high frequency estimation of water quality constituent fluxes using surrogates, 25(9); 1031-1044 (2009)10.1016/j.envsoft.2009.10. 012
14
J. N. Houser, P. J. Mulholland and K. O. Maloney, Journal of Environmental Quality, Upland disturbance affects headwater stream nutrients and suspended sediments during baseflow and stormflow, 35; 352-365 (2006)10.2134/jeq2005.0102
15
I. F. Johanna and S. Petra, Journal of Hydrology, A turbidity-based method to continuously monitor sediment, carbon and nitrogen flows in mountainous watersheds, 513; 45-57 (2014)10.1016/j.jhydrol.2014.03.034
16
A. S. Jones, K. S. David, S. H. Jefery and O. Nancy, Journal of the American Water Resources Association, Surrogate measures for providing high frequency estimates of total suspended solids and total phosphorus concentrations, 47(2); 239-253 (2011)10.1111/j.1752-1688.2010.00505.x
17
P. Jordan, A. Arnscheidt and H. Mcgrogan, Hydrology and Earth System Science, Characterising phosphorus transfers in rural catchments using a continuous bankside analyser, 11; 372-381 (2007)10.5194/hess-11-372-2007
18
M. Ließ, B. Glaser and B. Huwe, Geoderma, Uncertainty in the spatial prediction of soil texture comparison of regression tree and random forest models, 170; 70-79 (2012)10.1016/j.geoderma.2011.10.010
19
D. N. Moriasi, J. G. Arnold, M. W. Van Liew, R. L. Bingner, R. D. Harmel and T. L. Veith, Transactions of the ASABE, Model evaluation guidelines for systematic quantification of accuracy in watershed simulations, 50(3); 885-900 (2007)10.13031/2013.23153
20
J. L. Montgomery, T. C. Harmon, C. N. Haas, R. Hooper, N. L. Clesceri, W. Graham, W. J. Kaiser, A. Snaderson, B. Minsker, J. Schnoor and P. Brezonik, Environmental Science and Technology, The waters network: an integrated environmental observatory network for water research, 41(19); 6642-6647 (2007)10.1021/es072618f
21
J. E. Nash and J. V. Sutcliffe, Journal of Hydrology, River flow forecasting through conceptual models part I – a discussion of principles, 10(3); 282-290 (1970)10.1016/0022-1694(70)90255-6
22
I. Ouedraogo, P. Defourny and M. Vanclooster, Hydrogeology Journal, Application of random forest regression and comparison of its performance to multiple linear regression in modeling groundwater nitrate concentration at the African continent scale, 27(3); 1-18 (2018)10.1007/s10040-018-1900-5
23
A. M. Prasad, L. R. Iverson and A. Liaw, Ecosystems, Newer classification and regression tree techniques: bagging and random forests for ecological prediction, 9; 181-199 (2006)10.1007/s10021-005-0054-1
24
M. A. Razi and K. A. Athappilly, Expert Systems with Applications, A comparative predictive analysis of neural networks (NNs), nonlinear regression and classification and regression tree (CART) models, 29(1); 65-74 (2005)10.1016/j.eswa.2005.01.006
25
V. Rodriguez-Galiano, M. P. Mendes, M. J. Garcia-soldado, M. Chica-Olmo and L. Ribeiro, Science of the Total Environment, Predictive modeling of groundwater nitrate pollution using random forest and multisource variables related to intrinsic and specific vulnerability: a case study in an agricultural setting (Southern Spain), 476-477; 189-206 (2014)10.1016/j.scitotenv.2014.01.001
26
H. Ruegner, M. Schwientek, B. Beckingham, B. Kuch and P. Grathwohl, Environmental Earth Sciences, Turbidity as a proxy for total suspended solids (TSS) and particle facilitated pollutant transport in catchments, 69; 373-380 (2013)10.1007/s12665-013-2307-1
27
D. Scholefield, T. L. Goff, J. Braven, L. Dbdon, T. Long and M. Butler, Science of the Total Environment, Concerted diurnal patterns in riverine nutrient concentrations and physical conditions, 344; 201-210 (2005)10.1016/ j.scitotenv.2005.02.014
28
S. Settle, A. Goonetilleke and G. Ayoko, Water, Air, and Soil Pollution, Determination of surrogate indicators for phosphorus and solids in urban stormwater: application of multivariate data analysis techniques, 182; 149-161 (2007)10.1007/s11270-006-9328-2
29
B. Singh, P. Sihag and K. Singh, Modeling Earth Systems and Environment, Modeling of impact of water quality on infiltration rate of soil by random forest regression, 3; 999-1004 (2017)10.1007/s40808-017-0347-3
30
M. Snipes and C. D. Taylor, Wine Economics and Policy, Model selection and Akaike information criteria: an example from wine ratings and prices, 3(1); 3-9 (2014)10.1016/j.wep.2014.03.001
31
United States Environmental Protection Agency (USEPA), ” 841-B-07-006, United States Environmentl Protection Agency, “An approach for using load duration curves in the development of TMDLs; 1-68 (2007)
32
J. Verzani, 2nd ed. for the text “Using R for introductory statistics”, Data sets, etc, Version 2.0-6 (2018)
33
A. Villa, J. Fölster and K. Kyllmar, Environmental Monioring and Assessment, Determining suspended solids and total phosphorus from turbidity: comparison of high-frequency sampling with conventional monitoring methods, 191; 605 (2019)10.1007/s10661-019-7775-7
34
H. Wang, Y. Zhao, R. L. Pu and Z. Z. Zhang, Remote Sensing, Mapping Robinia pseudoacacia forest health conditions by using combined spectral, spatial, and textural information extracted from IKONOS imagery and random forest classifier, 7(7); 9020-9044 (2015)10.3390/rs70709020
35
M. Zambrano-Bigiarini, Goodness-of-fit functions for comparison of simulated and observed hydrological time series, Version 0.3-10 (2017)
36
A. D. Ziegler, L. X. Xi and C. Tantasarin, IAHS-AISH publication, Sediment load monitoring in the Mae Sa catchment in Northern Thailand; 86-91 (2011)
페이지 상단으로 이동하기