Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 https://doi.org/10.1007/s00477-021-02018-9 (0123456789().,-volV)(0123456789(). ,- volV) ORIGINAL PAPER Artificial Intelligence models for prediction of the tide level in Venice Francesco Granata1 • Fabio Di Nunno1 Accepted: 2 April 2021 / Published online: 9 April 2021 The Author(s), under exclusive licence to Springer-Verlag GmbH Germany, part of Springer Nature 2021 Abstract The city of Venice is an extraordinary architectural, artistic and cultural heritage. Unfortunately, its conservation is increasingly threatened by particularly significant high tides. Predicting the tide level in Venice, especially the high waters, is an essential task for the protection of the city and the lagoon. Complex statistical or hydrodynamic models, which require a large amount of input data, are currently used for this purpose. An effective alternative can be provided by models based on Artificial Intelligence algorithms. In this study, several different forecasting models were developed and each model was built in three variants, varying the implemented machine learning algorithm: M5P Regression Tree, Random Forest and Multilayer Perceptron. Until now, regression tree models had never been used to forecast tide levels. All the proposed models proved to be able to forecast the tide level in Venice with good accuracy. The M5P algorithm provided the best performance in most cases. All the models based on M5P were characterized by a coefficient of determination between 0.924 and 0.996, while the Relative Absolute Error was between 5.98 and 26.84%. In addition, good predictions were achieved by neglecting meteorological factors, even in the case of exceptionally high waters. Finally, satisfactory outcomes were also obtained with a forecast horizon of several hours, while a further specific comparison showed that the models based on the considered Machine Learning algorithms are able to outperform the AutoRegressive Integrated Moving Average models with exogenous input variables in forecasting high water. Keywords Tide prediction Machine learning M5P Random Forest Neural networks 1 Introduction Tide prediction is often a matter of prime importance for the peoples living along the coasts. Tides can strongly affect navigation, fishing and recreational activities. Tidal fluctuations have a major influence on coastal ecosystems. The design of coastal engineering works requires knowledge of the extreme values of tidal oscillations. Additionally, many coastal cities are at risk of being flooded when extreme high tide events occur and need to be properly protected. One of the most famous cases is undoubtedly represented by the city of Venice, which is in the middle of the largest lagoon (Fig. 1) in the Mediterranean Sea (about 550 & Francesco Granata f.granata@unicas.it 1 Department of Civil and Mechanical Engineering, University of Cassino and Southern Lazio, Via G. Di Biasio 43, 03043 Cassino, FR, Italy km2), located in the upper Adriatic Sea. The lagoon is connected to the Adriatic Sea by three narrow inlets. Its complex morphology changes continuously from the mainland to the sea (Carniello et al. 2009). Life in the Venetian Lagoon, and especially in the city of Venice, has always been strongly conditioned by the high tide (Camuffo 1993). The observed tide in Venice is the sum of two contributions: the astronomical tide, caused by the gravitational attraction of the Moon and the Sun, and the change in sea level induced by weather conditions (Franco et al. 1982; Fagherazzi et al. 2005). In ordinary conditions, the meteorological effects are small and the observed water level is little different from that induced by the astronomical tide. The latter is basically semidiurnal: two maximum heights and two minimum heights occur within 24 h. On the days of new moon and full moon the highest tidal fluctuations are observed. The predictability of the astronomical tide may be altered, even significantly, by weather factors. In the case of particular weather conditions, with low pressure and strong winds from the South-East (Scirocco), 123 2538 Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 Fig. 1 The Venice Lagoon, with the locations of the measuring stations the meteorological effects become relevant: if they are in phase with a maximum astronomical tide, they can lead to the phenomenon of high water (‘‘acqua alta’’), conventionally defined as a tide level greater than 80 cm with reference to the Punta della Salute tide gauge. At this tide level, the flooding of the lower part of the city begins. When the sea level exceeds 110 cm large areas begin to be flooded. When a strong wind from North-East (Bora) in the northern Adriatic blows together with Scirocco in middle Adriatic, the convergence of wind-induced marine currents may further enhance the phenomenon. An even worse situation occurs in the event of a storm with severe low pressure on the upper Adriatic and contemporary high pressure on the lower Adriatic. When a storm surge occurs on the Adriatic Sea, the latter behaves like a resonant cavity due to its shape, originating seiches of progressively decreasing amplitude (Franco et al. 1982). Of course, this phenomenon also 123 affects the tide height. In recent decades the high waters have become more frequent, also due to eustatism and subsidence (Carminati et al. 2003; Tosi et al. 2009). Moreover, the frequency and magnitude of high tide are correlated to interdecadal climatic oscillations (Fagherazzi et al. 2005). Venice is a unique cultural, artistic and architectural heritage in the world. Its defense is a matter of significant national interest in Italy. In order to protect Venice from high waters, the MOSE system is nearing completion. It consists of 4 arrays of barriers that will be raised and isolate the lagoon from the Adriatic Sea, when the tide height will exceed 110 cm (Umgiesser 2020). Therefore, the prediction of the tide level in the city of Venice, with a particular focus on high waters, is a matter of considerable practical interest, discussed for a long time (Finizio et al. 1972; Zaldivar et al. 2000; Carniello et al. 2005; Wei and Billings 2006; Zampato et al. 2007). Currently, in order to predict tide levels, both statistical models are used, such as Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 the BIGSUMDP expert model (Tosoni and Canestrelli 2011), and hydrodynamic models, such as the SHYFEM model (Umgiesser et al. 2014). These are very complex models, which require a very large number of input data. In predicting complex natural phenomena, models deriving from artificial intelligence studies have become increasingly popular in recent years (e.g. Kişi 2007; Wang et al. 2009; Nourani et al. 2014; Najafzadeh et al. 2016, 2018; Granata et al. 2018, 2020; Choubin et al. 2019; Najafzadeh and Saberi-Movahed 2019; Ghorbani et al. 2019; Di Nunno and Granata 2020; Malik et al. 2020; Najafzadeh and Oliveto 2020; Saberi-Movahed et al. 2020; Aghelpour and Varshavian 2021). So far, these models have been infrequently applied to tide prediction problems. Deo and Chaudhari (1998) used Artificial Neural Networks to predict tidal fluctuations in three different locations on the coast of India. The networks were trained using three different algorithms: error backpropagation, cascade correlation, and conjugate gradient. The authors found that continuous prediction of the tidal levels over 15-min and 30-min intervals was good despite complex tidal propagation in the considered areas. Okwuashi and Ndehedehe (2017) used Support Vector Regression for tide prediction in Lagos, Nigeria. In particular, they compared the results obtained with different kernel functions and obtained the highest accuracy with a polynomial kernel. Imani et al. (2018) applied Extreme Learning Machine and Relevance Vector Machine models to forecast tide levels in Chiayi, Taiwan. They showed that these algorithms outperformed Support Vector Machines and Radial Basis Function networks in predicting daily mean, maximum and minimum sea levels. Riazi (2020) proposed a deep neural networkbased model to forecast tide level in different locations of Queensland, Australia. The model was developed considering 4 input neurons, 3 neurons for astronomical forces and 1 neuron for geology and geomorphology factors, and showed good accuracy. In modeling problems involving the use of machine learning algorithms, the main challenging aspects are the selection of the algorithm, the selection of adequately representative variables, and the availability of suitable data sets. In this study, several different machine learning-based models were built for the forecasting of the tide level in Venice. Models differ in input variables. Three variants of each model were developed, changing the implemented algorithm: M5P Regression Tree, Random Forest, and Multilayer Perceptron, used as a benchmark. Until now, M5P and Random Forest had never been used to predict tide levels. These algorithms are generally very accurate and fast. Furthermore, they are particularly efficient in dealing with highly non-linear problems. At the same time, studies focusing on the prediction of high waters in Venice 2539 using Artificial Intelligence algorithms are not available so far in the literature. After a comparative analysis of the accuracy of the different models and algorithms, the study focused in particular on the ability of the simplest models to predict the phenomenon of high water, in order to provide useful operational tools. In addition, in order to further validate the effectiveness of the above models, a comparison with the classic AutoRegressive Integrated Moving Average with exogenous inputs models was proposed. 2 Materials and methods 2.1 M5P algorithm The M5P algorithm (Quinlan 1992) develops a regression tree to get predictions. A Regression tree is a decision tree in which the target variables are real numbers. In a regression tree the following elements can be identified: a root node, which contains all the data, internal nodes, which impose conditions on the input variables, and leaf nodes, represented by linear regression models on the target values of the variables. During the development process the input dataset is recursively divided into sub-domains and a multivariable linear regression model is built in each of them. At the initial step of the procedure, the entire dataset is divided into two subsets, evaluating all the possible binary split on every field. In the next steps, each partition is divided into smaller subsets, so as the tree branches out. At each stage, the procedure selects the splitting in two distinct partitions that maximizes a function of the Least-Squared Deviation (LSD), namely the within variance R(t) in the node t: 1 X RðtÞ ¼ ðyi ym ðtÞÞ2 ð1Þ NðtÞ i2t where N(t) is the number of subset units in the node t, yi is the target variable value for the i-th unit and ym is the mean of the target variable in the node t. LSD also represents the impurity of the node. The function to be maximized is /ðsp ; tÞ ¼ RðtÞ pL RðtL Þ pR RðtR Þ ð2Þ in which tL and tR are respectively the left and right nodes created by the split sp, while pL and pR are the portion units allocated to the left and right child node. The split sp that maximizes the value of U(sp, t) is selected. The algorithm halts if a stopping rule occurs. The stopping rules may refer to a minimum impurity level, a minimum variation of the impurity provided by new splits, a minimum number of elements in each node, or a maximum tree depth. 123 2540 Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 A regression tree model may be suffering from overfitting when the tree development is complete. This may reduce the model forecasting capability. This problem can be solved by using a pruning technique, which reduces the size of the tree removing the branches that do not significantly contribute to the predictive ability. 2.2 Random Forest A Random Forest is an ensemble of simple regression trees (Breiman 2001), whose predictions are combined to evaluate the final output. Several training subsets of the same size are randomly selected, with replacement, from the initial dataset. Then, a regression tree is built for each subset. These trees will be different and will not lead to the same results. However, Random Forest algorithm uses a modified tree learning process (Granata and de Marinis 2017), which introduces an additional step of randomization. While in standard regression trees each node is partitioned using the best splitting among all features (i.e. input variables), in a random forest each node is divided considering a subset of the input variables, randomly chosen at that node. The number of features to be considered for the division of each node is kept constant during the tree growth process. Each tree is developed as far as possible. There is no pruning process. Random Forest is usually robust to outliers and is very stable. In this study, forests consisting of 100 regression trees were considered. The preliminary analyzes carried out for the optimization of the models showed that increasing the number of trees beyond this value did not lead to an improvement in performance, but only in the calculation times. 2.3 Multilayer Perceptron A Multilayer Perceptron (MLP) is a class of feedforward Artificial Neural Network (Ruck et al. 1990), which can perform regression operations. A MLP contains at least three layers of nodes: an input layer, a hidden layer and an output layer. The input layer is made up of a set of neurons representing the input variables. Each neuron in the hidden layer processes the values from the previous layer with a weighted linear summation and a non-linear activation function. The values from the last hidden layer are transformed into output values by the output layer. The backpropagation algorithm is used for training. In this study were used neural networks with 1 layer of hidden neurons, whose number of neurons was equal to (number of input variables ? 1)/2. Sigmoid was selected as activation function. The assumed learning rate was 0.3, while momentum rate for the backpropagation algorithm was 0.2. These models’ structures proved to be optimal in the addressed cases. 123 2.4 Input dataset, efficiency criteria, and implementation The tide level data were obtained from the measurements carried out by the tide gauge of Punta della Salute (Fig. 1). Instead, the weather data were acquired from measurements carried out at the CNR Platform, outside the Venetian Lagoon (Fig. 1). In the first phase of the study, all data collected, at 30-min intervals, between January 1, 2012 and December 31, 2015 were considered. The choice of the dataset was imposed by the simultaneous availability of the tide and weather data, provided by the Institute for Environmental Protection and Research (ISPRA). The gravitational attraction of the celestial bodies was taken into account by introducing the height of the astronomical tide among the input data. The astronomical tide at the time of interest can be easily calculated, through harmonic analysis, by the equation (Ferla et al. 2007): AT ¼ A0 þ N X An cosðrn t jn Þ ð3Þ n¼1 where An is the amplitude, rn the angular frequency, jn the phase delay of component n. The values of these parameters can be easily found on the Venice Municipality website (www.comune.venezia.it). Four different well-known criteria were employed to assess the effectiveness of the developed models: • • • • the coefficient of determination R2, the Mean Absolute Error (MAE), the Root Mean Squared Error (RMSE), the Relative Absolute Error (RAE). These metrics are defined below: ! Pm 2 2 i¼1 ðfi yi Þ R ¼ 1 Pm 2 i¼1 ðya yi Þ ð4Þ In which m is the total number of experimental data, fi is the forecasted value for data point i, and yi is the observed value for data point i; Pm jfi yi j ð5Þ MAE ¼ i¼1 m sffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi Pm 2 i¼1 ðfi yi Þ RMSE ¼ ð6Þ m Pm jfi yi j RAE ¼ Pmi¼1 ð7Þ j i¼1 ya yi j where ya is the averaged value of the observed data. R2 can be interpreted as the proportion of the variance in the dependent variable that is predictable from the independent variable: it is a goodness of fit measure. MAE evaluates the Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 average magnitude of the errors in a set of predictions, without considering their sign. RMSE is the standard deviation of the residuals, i.e. the prediction errors. RAE is a normalization of the total absolute error, which is divided by the absolute error of a simple predictor: the arithmetic mean of the experimental values. The four adopted metrics allow to fully characterize the efficiency of the models, while alternative metrics generally do not alter the comparison (Granata 2019). The considered algorithms were executed in MATLAB environment. The optimization of the models was carried out by means of a grid search procedure, which changed the values of the parameters and checked the values of the above defined metrics, in order to maximize coefficient of determination and minimize errors. Of the initial dataset, made up of over 70,000 vectors, 85% was used for model training and 15% for testing and validation. A preliminary analysis showed that the results are completely equivalent to those obtainable with a k-fold cross validation. The raw data underwent a normalization preprocessing, i.e. all attributes were rescaled to the range 0 to 1. Normalization proved useful in improving the efficiency of algorithms because the input data had significantly varying scales. 3 Results and discussion The previously described algorithms were used to develop nineteen different tide level prediction models. They mainly differ in the input variables. In addition, three variants were developed for each model, changing the implemented algorithm. The characteristics of all the considered models are summarized in Table 1, together with the values of the metrics referring to the testing phase. For summary reasons, only the results of the models considered to be more interesting are shown and discussed below. Model A is characterized by the more complex structure. It has 28 input variables: the astronomical tide AT, the wind speed WS, the wind direction WD, the barometric pressure BP, and the previously observed tide levels, every hour from 24 h before (Z-24) to 1 h before (Z-1) the time of the forecast (Table 1). Model G has 9 input variables: the astronomical tide AT, the weather parameters WS, WD, and BP, the tide levels observed 24 h before Z-24, 6 h before Z-6, 3 h before Z-3, 2 h before Z-2, and 1 h before Z1. Model I has 8 input variables and differs from Model G only because Z-6 is missing among the input variables. Model N has 7 input variables: AT, WS, WD, BP, Z-24, Z-2, Z-1. Model Q has 6 input variables and differs from Model N because Z-24 is missing among the input variables. In the previous models, the meteorological variables refer to 2 h before the forecast time. Model T is the simplest, being 2541 characterized by only 3 input variables: AT, Z-2, Z-1. This model completely neglects weather variables. Model U has 24 input variables: the astronomical tide AT, the meteorological parameters WS, WD, and BP, and the previously observed tide levels, every hour from 24 h before (Z-24) to 5 h before (Z-5) the time of the forecast. As an example, Fig. 2 shows the comparisons between the predicted and measured values of the tide level, for Model A and the three different algorithms. The comparison for all the test values is shown on the left, while a comparison in the form of a time series is shown on the right, for a short period of time characterized by high waters. Model A (Fig. 2) led to very accurate results. M5P (R2 = 0.9958, RAE = 5.98%) proved to be the best performing algorithm of the three considered. RF showed a significant decrease in performance (R2 = 0.9902, RAE = 9.41%), contrary to expectations. MLP provided results (R2 = 0.9958, RAE = 6.17%) comparable to those obtained with M5P. The significant drawback of using model A lies in the large number of observed sea level values to be included in the input data. The forecasting capabilities of Models G and I were very similar. In addition, the three different algorithms provided similar results. M5P was a little more accurate than the other two procedures. All algorithms predicted high waters with good accuracy, while RF, similarly to what was shown by model A (Fig. 2c), exhibited a slight tendency to overestimate low tide levels. Model N, which differs from Model I in that it does not include Z-3 among the input variables, was slightly less accurate than the latter. Again, M5P outperformed both RF and MLP. In any case, the differences in the performance of the three algorithms were minimal. Model Q, which includes among the input variables Z-2 and Z-1, in addition to the astronomical tide and meteorological factors, showed a further slight decrease in accuracy. M5P confirmed to be the most performing algorithm (R2 = 0.9882, RAE = 10.54%). Finally, Model T, based only on theastronomical tide and the observed values Z-2 and Z-1, also led to good predictions. M5P algorithm provided again the best outcomes (R2 = 0.9855, RAE = 11.78%). Finally, Model U, which will be discussed at the end of this section, was the least accurate, though still with satisfactory results (R2 = 0.9341, RAE = 24.45% for Random Forest). It should be noted that, unexpectedly, the M5P algorithm outperformed RF in most cases. This result did not depend on the number of trees considered in the random forest. The developed models showed different levels of accuracy in different estimation ranges of tide height. For a deeper understanding of this issue, the entire range of observed tide values was divided into three subintervals: - 123 2542 Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 Table 1 Summary of models and results (N.I.V. = number of input variables) MODEL N.I.V INPUT VARIABLES ALGORITHM R2 MAE [cm] RMSE [cm] RAE A 28 AT, WS, WD, BP, Z-24, Z-23, Z-22, Z-21, Z-20, Z-19, Z-18, Z-17, Z-16, Z15, Z-14, Z-13, Z-12, Z-11, Z-10, Z-9, Z-8, Z-7, Z-6, Z-5, Z-4, Z-3, Z-2, Z-1 M5P 0.9958 1.298 1.746 5.98% RF 0.9902 2.016 2.651 9.41% MLP 0.9958 1.341 1.791 6.17% M5P 0.9843 2.557 3.391 11.77% RF 0.9732 3.341 4.366 15.59% MLP 0.9857 2.399 3.228 11.19% M5P RF 0.9952 0.9924 1.337 1.767 1.831 2.324 6.24% 8.25% MLP 0.9950 1.619 2.119 7.56% M5P 0.9948 1.405 1.911 6.56% RF 0.9918 1.832 2.399 8.55% MLP 0.9944 1.844 2.361 8.61% M5P 0.9936 1.583 2.104 7.39% RF 0.9916 1.872 2.435 8.74% MLP 0.9926 1.76 2.295 8.14% M5P 0.9724 3.459 4.409 16.15% RF 0.9736 3.355 4.306 15.66% MLP 0.9728 3.784 4.753 17.66% M5P 0.9936 1.608 2.129 7.50% RF 0.9920 1.823 2.36 8.51% MLP 0.9928 1.749 2.271 8.16% M5P RF 0.9720 0.9720 3.483 3.451 4.436 4.431 16.26% 16.11% MLP 0.9716 3.679 4.627 17.17% M5P 0.9928 1.704 2.275 7.95% RF 0.9902 2.029 2.621 9.47% MLP 0.9920 1.941 2.523 8.98% B 27 AT, WS, WD, BP, Z-24, Z-23, Z-22, Z-21, Z-20, Z-19, Z-18, Z-17, Z-16, Z15, Z-14, Z-13, Z-12, Z-11, Z-10, Z-9, Z-8, Z-7, Z-6, Z-5, Z-4, Z-3, Z-2 C 16 AT, WS, WD, BP, Z-23, Z-21, Z-19, Z-17, Z-15, Z-13, Z-11, Z-9, Z-7, Z-5, Z-3, Z-1 D 12 AT, WS, WD, BP, Z-22, Z-19, Z-16, Z-13, Z-10, Z-7, Z-4, Z-1 E F G 10 9 9 AT, WS, WD, BP, Z-21, Z-17, Z-13, Z-9, Z-5, Z-1 AT, WS, WD, BP, Z-24, Z-18, Z-12, Z-6, Z-1 AT, WS, WD, BP, Z-24, Z-6, Z-3, Z-2, Z-1 H 8 AT, WS, WD, BP, Z-24, Z-18, Z-6, Z-1 I 8 AT, WS, WD, BP, Z-24, Z-3, Z-2, Z-1 L M N O P Q R 7 7 7 7 6 6 5 123 AT, WS, WD, BP, Z-24, Z-3, Z-2 AT, WS, WD, BP, Z-24, Z-12, Z-1 AT, WS, WD, BP, Z-24, Z-2, Z-1 AT, WS, WD, BP, Z-3, Z-2, Z-1 AT, WS, WD, BP, Z-24, Z-1 AT, WS, WD, BP, Z-2, Z-1 AT, WS, WD, Z-2, Z-1 M5P 0.9692 3.582 4.681 16.72% RF 0.9657 3.77 4.907 17.59% MLP 0.9685 4.051 5.223 18.90% M5P 0.9677 3.787 4.788 17.68% RF 0.9679 3.748 4.771 17.49% MLP 0.9675 4.084 5.166 19.07% M5P 0.9908 1.971 2.573 9.19% RF 0.9892 2.123 2.742 9.91% MLP 0.9902 2.076 2.687 9.60% M5P 0.9918 1.802 2.419 8.41% RF MLP 0.9880 0.9902 2.294 2.076 2.923 2.687 10.61% 9.60% M5P 0.9661 3.901 4.897 18.21% RF 0.9649 3.949 4.969 18.44% MLP 0.9641 4.312 5.379 20.13% M5P 0.9882 2.277 2.926 10.54% RF 0.9862 2.456 3.137 11.46% MLP 0.9872 2.365 3.001 10.94% M5P 0.9859 2.496 3.179 11.65% RF 0.9841 2.652 3.369 12.38% MLP 0.9837 3.043 3.778 14.20% Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 2543 Table 1 (continued) R2 MODEL N.I.V INPUT VARIABLES ALGORITHM S 4 AT, BP, Z-2, Z-1 M5P 0.9882 2.261 2.928 10.55% RF 0.9859 2.483 3.190 11.59% MLP 0.9866 2.505 3.160 11.69% M5P RF 0.9855 0.9809 2.524 2.907 3.213 3.702 11.78% 13.57% MLP 0.9835 2.980 3.719 13.91% M5P 0.9242 5.749 7.319 26.84% RF 0.9341 5.237 6.823 24.45% MLP 0.9361 5.425 6.992 25.32% T 3 AT, Z-2, Z-1 U 24 AT, WS, WD, BP, Z-24, Z-23, Z-22, Z-21, Z-20, Z-19, Z-18, Z-17, Z-16, Z15, Z-14, Z-13, Z-12, Z-11, Z-10, Z-9, Z-8, Z-7, Z-6, Z-5 60–0, 0–80, and 80–120 cm. The RAE was assessed in each subinterval, for each model. The values are shown in Fig. 3a. All models showed the lowest RAE values in the range 0–80 cm, and the highest RAE values in the range 80–120 cm. These higher RAE values for high waters was due to the lower number of data available for models’ training that are in the range 80–120 cm. Of the three algorithms, M5P showed the lowest RAE value in all subintervals. However, the increase of RAE in the range Z [ 80 cm did not involve a significant reduction in accuracy, as can be seen from the box-plots of the residuals (i.e. the difference between expected and measured values) reported in Fig. 3b for the different considered M5P-based models. For each given model, the residuals were comparable in the different subintervals. In addition, The distribution of residuals is rather symmetrical for all models. Furthermore, Fig. 3b shows that in the case of high waters, Model T was characterized by errors comparable to those of most other models. Since Model T does not require meteorological data as input, it was possible to retrain the model with a longer time series (January 1, 2009–December 31, 2018) and verify its ability to predict high waters in a much larger number of cases. Figure 4 shows on the left the comparison between the expected and observed values of the high waters that occurred between 2 July 2017 and 31 December 2018, on the right the comparison in the form of a time series between expected and observed values during the exceptional event started on October 29, 2018. The results of the comparisons are satisfactory. Model T, especially in the M5P-based variant, proved to be able to adequately predict even the most severe high water phenomena. In some cases, RF and MLP slightly underestimated severe events MAE [cm] RMSE [cm] RAE (Fig. 4c and Fig. 4e), while M5P did not show this tendency (Fig. 4a). As for the October 29, 2018 event, both M5P (Fig. 4b) and MLP (Fig. 4f) were able to accurately predict all high tide peaks, while they were not as effective with the first two low tide peaks. RF, on the other hand, appreciably underestimated the more severe high tide peak (Fig. 4d). It is interesting to note that, from a preliminary analysis, it was found that the results of the different models showed negligible variations if the weather data referred to 1 h or 3 h before the forecast time, instead of 2 h. This result also confirms a poor sensitivity of forecast models based on machine learning to meteorological parameters. Model L, which was not considered in detail in the discussion of the results, shows that as the forecast horizon increases, there is a slight reduction in accuracy. Model U further confirms this trend. However, Model U responds to the need to provide reliable predictions even with wider forecast horizons, in this case 5 h. This forecast horizon should, for example, enable the procedure for lifting the MOSE gates to be activated. Even if the accuracy of Model U is lower than that of the models previously discussed, while still requiring a large number of input parameters, it demonstrates how an approach based on regression tree algorithms, or on MLP neural networks, is adequate to get good forecasts even several hours in advance. Figure 5 shows the predictions obtained with the three Model U variants over the period from 15 to 18 October 2015. The three variants of the model provided good results, accurately predicting the first peak of 111 cm, even if they then appreciably overestimated the third and the fourth peak, which were lower. The same Fig. 5 shows a comparison with 2 different AutoRegressive Integrated Moving Average with exogenous inputs (ARIMAX) models, used for the prediction of 123 2544 Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 Fig. 2 Model A: comparisons between the predicted and measured values the same tide event and with the same forecast horizon, equal to 5 h. Autoregressive integrated moving average 123 models, in different variants, has been widely used in hydrology for time series predictions (Box and Jenkins Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 2545 Fig. 3 Accuracy assessment in different subintervals 1968; Jalalkamali et al. 2015; Papacharalampous et al. 2018; Aghelpour et al. 2021). However, a detailed description of an ARIMAX model is not reported in this article because it is beyond the scope of this study. It can be found, for example, in Fan et al. (2009). The ARIMAX-1 model, which requires both the astronomical tide and all the meteorological parameters indicated above as exogenous variables, led to an underestimation of the first peak, to an overestimation of the second and fourth peaks, while correctly predicted the third and fifth peaks. The ARIMAX-2 model, which requires only the astronomical tide as exogenous input variable, provided poor predictions over the entire investigated period. This result also highlights a greater sensitivity of ARIMAX based models to meteorological parameters in comparison with the Machine Learning based-models. Overall, ARIMAX models appear to be less effective than models based on Machine Learning algorithms in predicting the wide oscillations observed during high water phenomena. A more in-depth comparison between forecasting models based on Machine Learning algorithms and ARIMAX models, as well as the development of hybrid models based on both (Moeeni and Bonakdari 2017), could be the subject of future developments of this research. Further future developments of this research aim to obtain models characterized by remarkable accuracy even with long-term forecast horizons. 123 2546 Fig. 4 Prediction of high waters by Model T 123 Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 2547 140 RF M5P MLP Experimental data ARIMAX-1 ARIMAX-2 120 100 Z [cm] 80 60 40 20 0 -20 Fig. 5 Forecast of tide fluctuations with a 5 h advance: comparison between ML-based models and ARIMAX models 4 Conclusions Authors’ contributions Both authors contributed equally to this work. The prediction of the tide level in Venice, especially of the most severe high waters, is an essential task for the protection of the city and the lagoon. In this study, several different forecasting models were developed and each model was built in three variants, changing the adopted machine learning algorithm: M5P Regression Tree, Random Forest and Multilayer Perceptron. The main findings of this study can be summarized as follows: Funding This research received no external funding. • all the proposed models, in all their variants, predicted the tide height in Venice with good accuracy; • M5P outperformed the other two algorithms in most cases; • good forecasts can also be obtained by completely neglecting the weather features, even in the case of exceptionally high waters; • it is possible to develop models based on M5P, RF and MLP that provide accurate predictions even with forecast horizons of several hours; • M5P, Random Forest, and MLP appear to be able to outperform ARIMAX models in predicting the wide oscillations observed during high water phenomena. A constantly updated training dataset will also allow to take into account changes in sea level due to climate change and subsidence. Data availability The data are freely available online on the website: https://www.venezia.isprambiente.it/index.php?folder_id=20&lang_ id=2 Declarations Conflicts of interest The authors declare that they have no conflicts of interest. References Aghelpour P, Varshavian V (2020) Evaluation of stochastic and artificial intelligence models in modeling and predicting of river daily flow time series. Stoch Environ Res Risk Assess 34:33–50 Aghelpour P, Bahrami-Pichaghchi H, Varshavian V (2021) Hydrological drought forecasting using multi-scalar streamflow drought index, stochastic models and machine learning approaches, in northern Iran. Stoch Environ Res Risk Assess 26(4):1–21 Box GEP, Jenkins GM (1968) Some recent advances in forecasting and control. J R Stat Soc Ser C (Appl Stat) 17(2):91–109 Breiman L (2001) Random forests. Mach Learn 45(1):5–32 Camuffo D (1993) Analysis of the sea surges at Venice from AD 782 to 1990. Theor Appl Climatol 47(1):1–14 Carminati E, Doglioni C, Scrocca D (2003) Apennines subductionrelated subsidence of Venice (Italy). Geophys Res Lett 30(13): 1–4 Carniello L, Defina A, Fagherazzi S, D’Alpaos L (2005) A combined wind wave–tidal model for the Venice lagoon, Italy. J Geophys Res Earth Surf 110:1–15 123 2548 Stochastic Environmental Research and Risk Assessment (2021) 35:2537–2548 Carniello L, Defina A, D’Alpaos L (2009) Morphological evolution of the Venice lagoon: evidence from the past and trend for the future. J Geophys Res Earth Surf 114:1–10 Choubin B, Mosavi A, Alamdarloo EH, Hosseini FS, Shamshirband S, Dashtekian K, Ghamisi P (2019) Earth fissure hazard prediction using machine learning models. Environ Res 179:108770 Deo MC, Chaudhari G (1998) Tide prediction using neural networks. Comput Aided Civ Infrastruct Eng 13(2):113–120 Di Nunno F, Granata F (2020) Groundwater level prediction in Apulia region (Southern Italy) using NARX neural network. Environ Res 190:110062 Fagherazzi S, Fosser G, D’Alpaos L, D’Odorico P (2005) Climatic oscillations influence the flooding of Venice. Geophys Res Lett 32:1–10 Fan J, Shan R, Cao X (2009) The analysis to Tertiary-industry with ARIMAX model. J Math Res 1(2):156–163 Ferla M, Cordella M, Michielli L, Rusconi A (2007) Long-term variations on sea level and tidal regime in the lagoon of Venice. Estuar Coast Shelf Sci 75(1–2):214–222 Finizio C, Palmieri S, Riccucci A (1972) A numerical model of the Adriatic for the prediction of high tides at Venice. Q J R Meteorol Soc 98(415):86–104 Franco P, Jeftic L, Rizzoli PM, Michelato A, Orlic M (1982) Descriptive model of the Northern Adriatic. Oceanol Acta 5(3):379–389 Ghorbani MA, Deo RC, Karimi V, Kashani MH, Ghorbani S (2019) Design and implementation of a hybrid MLP-GSA model with multi-layer perceptron-gravitational search algorithm for monthly lake water level forecasting. Stoch Environ Res Risk Assess 33:125–147 Granata F (2019) Evapotranspiration evaluation models based on machine learning algorithms—a comparative study. Agric Water Manag 217:303–315 Granata F, de Marinis G (2017) Machine learning methods for wastewater hydraulics. Flow Meas Instrum 57:1–9 Granata F, Saroli M, de Marinis G, Gargano R (2018) Machine learning models for spring discharge forecasting. Geofluids 2018:8328167 Granata F, Gargano R, de Marinis G (2020) Artificial intelligence based approaches to evaluate actual evapotranspiration in wetlands. Sci Total Environ 703:135653 Imani M, Kao HC, Lan WH, Kuo CY (2018) Daily sea level prediction at Chiayi coast, Taiwan using extreme learning machine and relevance vector machine. Glob Planet Change 161:211–221 Jalalkamali A, Moradi M, Moradi N (2015) Application of several artificial intelligence models and ARIMAX model for forecasting drought using the Standardized Precipitation Index. Int J Environ Sci Technol 12:1201–1210 Kişi Ö (2007) Streamflow forecasting using different artificial neural network algorithms. J Hydrol Eng 12(5):532–539 Malik A, Tikhamarine Y, Souag-Gamane D, Kisi O, Pham QB (2020) Support vector regression optimized by meta-heuristic algorithms for daily streamflow prediction. Stoch Environ Res Risk Assess 34:1755–1773 Moeeni H, Bonakdari H (2017) Forecasting monthly inflow with extreme seasonal variation using the hybrid SARIMA-ANN model. Stoch Environ Res Risk Assess 31:1997–2010 Najafzadeh M, Oliveto G (2020) Riprap incipient motion for overtopping flows with machine learning models. J Hydroinf 22(4):749–767 123 Najafzadeh M, Saberi-Movahed F (2019) GMDH-GEP to predict free span expansion rates below pipelines under waves. Mar Georesour Geotechnol 37(3):375–392 Najafzadeh M, Etemad-Shahidi A, Lim SY (2016) Scour prediction in long contractions using ANFIS and SVM. Ocean Eng 111:128–135 Najafzadeh M, Saberi-Movahed F, Sarkamaryan S (2018) NFGMDH-Based self-organized systems to predict bridge pier scour depth under debris flow effects. Mar Georesour Geotechnol 36(5):589–602 Nourani V, Baghanam AH, Adamowski J, Kisi O (2014) Applications of hybrid wavelet–artificial intelligence models in hydrology: a review. J Hydrol 514:358–377 Okwuashi O, Ndehedehe C (2017) Tide modelling using support vector machine regression. J Spat Sci 62(1):29–46 Papacharalampous GA, Tyralis H, Koutsoyiannis D (2018) Comparison of stochastic and machine learning methods for multi-step ahead forecasting of hydrological processes. Stoch Environ Res Risk Assess 33(2):481–514 Quinlan JR (1992) Learning with continuous classes. In: 5th Australian joint conference on artificial intelligence, vol 92, pp 343–348 Riazi A (2020) Accurate tide level estimation: a deep learning approach. Ocean Eng 198:107013 Ruck DW, Rogers SK, Kabrisky M (1990) Feature selection using a multilayer perceptron. J Neural Netw Comput 2(2):40–48 Saberi-Movahed F, Najafzadeh M, Mehrpooya A (2020) Receiving more accurate predictions for longitudinal dispersion coefficients in water pipelines: training group method of data handling using extreme learning machine conceptions. Water Resour Manag 34(2):529–561 Tosi L, Rizzetto F, Zecchin M, Brancolini G, Baradello L (2009) Morphostratigraphic framework of the Venice Lagoon (Italy) by very shallow water VHRS surveys: evidence of radical changes triggered by human-induced river diversions. Geophys Res Lett 36:L09406 Tosoni A, Canestrelli P (2011) Il modello stocastico per la previsione di marea a Venezia. Atti Ist Veneto Sci Lett Arti 169:2010–2011 Umgiesser G (2020) The impact of operating the mobile barriers in Venice (MOSE) under climate change. J Nat Conserv 54:125783 Umgiesser G, Ferrarin C, Cucco A, De Pascalis F, Bellafiore D, Ghezzo M, Bajo M (2014) Comparative hydrodynamics of 10 Mediterranean lagoons by means of numerical modeling. J Geophys Res Oceans 119(4):2212–2226 Wang WC, Chau KW, Cheng CT, Qiu L (2009) A comparison of performance of several artificial intelligence methods for forecasting monthly discharge time series. J Hydrol 374(3–4):294–306 Wei HL, Billings SA (2006) An efficient nonlinear cardinal B-spline model for high tide forecasts at the Venice Lagoon. Nonlinear Process Geophys 13(5):577–584 Zaldivar JM, Gutiérrez E, Galván IM, Strozzi F, Tomasin A (2000) Forecasting high waters at Venice Lagoon using chaotic time series analysis and nonlinear neural networks. J Hydroinf 2(1):61–84 Zampato L, Umgiesser G, Zecchetto S (2007) Sea level forecasting in Venice through high resolution meteorological fields. Estuar Coast Shelf Sci 75(1–2):223–235 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )