Preview

Mining Science and Technology (Russia)

Advanced search

Neural network models for short-term electricity consumption forecasting at a mining enterprise participating in the wholesale electricity market

https://doi.org/10.17073/2500-0632-2026-01-1070

Contents

Scroll to:

Abstract

Electricity consumption at mining enterprises is markedly nonstationary because of intermittent equipment operation and the rapidly varying loads associated with electric-arc steelmaking. Deviations between actual consumption and day-ahead bids in the Wholesale Electricity and Capacity Market (WECM) entail substantial penalties, making conventional statistical methods insufficiently accurate and creating a need for adaptive neural network approaches. This study developed and evaluated a neural network model for short-term electricity consumption forecasting designed to minimize WECM penalties. Six architectures (RNN, GRU, LSTM, BiLSTM, CNN–LSTM, and attention-based models) were compared under two multi-step forecasting strategies: direct and recursive. Hyperparameters were selected using Hyperband. The study used 2022–2023 industrial data from a large mining and processing enterprise (17,520 hourly measurements). The results showed that the GRU architecture with the direct strategy provided the highest accuracy (MAPE of 5.44%) and the greatest reduction in penalty payments relative to the seasonal naïve model (21.25%), equivalent to RUB 38.6 million in quarterly savings during the test period (Q4 2023). Evidence of excessive model complexity was observed: simpler recurrent networks outperformed attention-based models on a limited dataset. The recursive strategy was inferior to the direct strategy because of cumulative error propagation and the inability to use stochastic exogenous PCA factors. The findings support the transition from expert judgment and conventional statistical methods to neural network approaches at mining enterprises with rapidly varying loads.

For citations:


Ematin E.A., Karpenko S.M. Neural network models for short-term electricity consumption forecasting at a mining enterprise participating in the wholesale electricity market. Mining Science and Technology (Russia). https://doi.org/10.17073/2500-0632-2026-01-1070

Neural network models for short-term electricity consumption forecasting at a mining enterprise participating in the wholesale electricity market

Introduction

Short-term electricity consumption forecasting for industrial enterprises is essential for participation in the Wholesale Electricity and Capacity Market (WECM). Under the current market rules, participants1 must submit hourly bids for the following day’s planned consumption. Deviations of actual consumption from the submitted bids entail penalties, whose amount is determined by the imbalance volume and the electricity price in the respective hour.

Electricity consumption at mining and processing enterprises is markedly nonstationary because of the composition and operating modes of their process equipment. At such enterprises, electrical loads operate either continuously (heating furnaces, pumps, fans, and compressors) or on intermittent periodic duty (rolling equipment, electric-arc furnaces, welding equipment, and hoisting and conveying systems). Most electricity use follows a cyclical pattern within a continuously operating, round-the-clock production process. Electric-arc steelmaking is a major source of rapidly varying loads and nonstationarity in the resulting time series: the nonlinear electrical resistance of the arc during melting causes consumption surges and load drops over short intervals [1]. Together, intermittent periodic operation and the nonlinear arc resistance of electric-arc furnaces shape the load time series. The resulting switching patterns differ from the profiles of aggregated power systems. These characteristics complicate the use of conventional forecasting methods and make the search for adaptive approaches particularly relevant.

Classical statistical methods for time-series forecasting – autoregressive integrated moving average (ARIMA) models, autoregressive distributed lag (ADL) models, and Holt–Winters exponential smoothing – are based on assumptions of process stationarity and linearity [2]. For aggregated regional power systems with smoothed load profiles and high inertia, these methods provide high accuracy (MAPE of 1.5–2.0%). However, at the level of an individual industrial enterprise, where abrupt changes in operating mode and event-driven dynamics prevail, the error of classical methods increases to 4–7%, which is unacceptable for minimizing WECM penalty payments.

Deep learning methods, particularly recurrent neural networks (RNNs), can approximate complex nonlinear relationships without explicitly specifying physical equations. Long short-term memory (LSTM) [3] and gated recurrent unit (GRU) [4] architectures mitigate the vanishing-gradient problem in conventional RNNs [5]. These architectures can effectively model long-term dependencies in time series. Studies using regional power-system data show that LSTM achieves a MAPE of 0.3–0.7% over horizons of 1–7 days, outperforming gradient-boosting methods at longer horizons [1].

Attention-based architectures represent another line of research [6], including the Temporal Fusion Transformer [7] and specialized modifications such as Informer [8] and Autoformer [9]. In theory, the attention mechanism enables a model to use information from the entire input sequence by assigning different weights to elements according to their relevance. In practice, however, the use of transformers is limited by their high data requirements: training datasets covering fewer than 3–5 years of hourly measurements tend to lead to overfitting [10].

Existing comparative studies generally use aggregated power-system data [11] and do not account for the rapidly varying loads specific to mining enterprises [1].

Several research gaps remain in electricity consumption forecasting for large continuously operating mining and processing enterprises. First, comparative studies of architectures generally use aggregated power-system data and do not consider the stochastic nature of composite PCA factors derived from production variables. Second, the literature on multi-step forecasting strategies largely focuses on comparisons between Direct and Recursive [12, 13] and relies on the theoretical framework [14, 15] and analyses of exposure bias in recursive RNNs [16], whereas MIMO (one model with a multidimensional output, corresponding to Direct Multi-Output in Section 1.3) and DirRec (a hybrid of Direct and Recursive [17]) have been studied without combining them with exogenous PCA factors. Third, forecasting approaches have not been systematically compared for the rapidly varying loads characteristic of metallurgical facilities with continuously operating core production: available studies either focus on individual operations (electric-arc furnaces [13]) or do not extend the strategy comparison to operational WECM metrics.

The contribution of this study is the adaptation of a set of recurrent architectures for use at large, continuously operating mining and processing enterprises and the empirical evaluation of a multi-step forecasting strategy in the presence of stochastic exogenous PCA factors. The results are intended for WECM planning tasks rather than comparisons on aggregated benchmarks.

This study aims to develop and experimentally evaluate a neural network model for short-term electricity consumption forecasting at a mining enterprise that minimizes WECM penalties.

To achieve this aim, the study addressed the following objectives, which are presented sequentially in Section 1, Methods (Objectives 1–3), and Section 2, Results (Objective 4):

  1. Neural network architectures (RNN, GRU, LSTM, BiLSTM, CNN–LSTM, and attention-based models) were compared for forecasting nonstationary loads at a mining enterprise.
  2. A multi-step forecasting strategy was selected and evaluated, taking into account constraints on the availability of exogenous factors over a 24 h forecasting horizon at hourly resolution.
  3. Model hyperparameters were selected using Hyperband, and the model was validated on industrial data.
  4. The economic benefit of implementing the model was quantified in terms of reduced WECM penalty payments.

Section 1 describes the source data, model-development methodology, and quality metrics. Section 2 presents the results of the experimental comparison of architectures. Section 3 discusses the findings and study limitations.

1 Resolution of the Government of the Russian Federation No. 1172 dated December 27, 2010, “On the Rules of the Wholesale Electricity and Capacity Market” (as amended and supplemented). URL: http://pravo.gov.ru/proxy/ips/?docbody=&nd=102144342

1. Methods

1.1. Source data

The study used data from a large mining enterprise. The dataset comprised hourly electricity consumption measurements from January 2022 through December 2023, totaling 17,520 time points. The target variable was the aggregate balance of electricity consumption by the production shops and on-site electricity generation, obtained from the enterprise’s automated commercial electricity metering system.

The case study involved a large mining and processing enterprise with integrated metallurgical operations. Its main process stages, grouped by production category, included coke production, sintering, blast-furnace ironmaking, electric-arc steelmaking, basic oxygen steelmaking, and sheet rolling, as well as on-site power-generating facilities that partially met demand from internal sources (combined heat and power plants and energy recovery units). Core production operated continuously around the clock; therefore, the hourly electricity consumption profile had no pronounced seasonal component. By load pattern, electricity consumers were divided into those operating continuously (heating furnaces, pumps, fans, and compressors) and intermittently (rolling equipment, electric-arc furnaces, welding equipment, and hoisting and conveying systems). Electric-arc steelmaking was the primary source of nonstationarity in the resulting electricity consumption series: the melting process involved nonlinear electrical resistance of the arc and rapidly varying furnace loads. Identifying details of the enterprise are not reported in this study.

The inputs to the models using the direct forecasting strategy (Direct; see Section 1.3) comprised six channels. The target variable was the enterprise’s historical hourly electricity consumption. The exogenous variables were a composite equipment downtime factor (the first principal component of four downtime indicators; 86.35% explained variance), a composite on-site generation factor (the fourth component of four on-site generation units; 15.05% variance; correlation with the target, 0.34), outdoor air temperature, a composite calendar factor (the fifth component of seven calendar variables; 14.22% variance; correlation with the target, −0.35), and a binary public-holiday indicator. WECM market indicators and the loads of individual process stages (including the electric arc furnace shop) were considered uninformative at the feature-engineering stage because of their low absolute correlations with the target variable and were therefore excluded from the final set of exogenous model inputs.

Composite PCA factors were calculated using a uniform procedure. The source features in each group were first standardized by z-scoring, with mean centering and scaling by the standard deviation. PCA was performed by singular value decomposition of the feature matrix. Factor rotation (Varimax, Promax, etc.) and component whitening were not applied. The component retained from each group was selected by the maximum absolute correlation with the target variable from among solutions with one, two, or an automatically selected number of components at a 95% explained-variance threshold. Three PCA factors were included in the final set of exogenous model inputs: Equipment Downtime – four equipment downtime indicators, component 1 selected, 86.35% explained variance; Generation Output – four on-site generation units, component 4 selected, 15.05% explained variance (correlation with the target, 0.34); and Time – seven calendar features, component 5 selected, 14.22% explained variance (correlation with the target, −0.35). The last components in the Generation Output and Time groups were selected using the maximum |corr| with the target variable rather than the proportion of explained variance.

Missing values were handled at different data-preparation stages according to the nature of the source. When data were merged with the process downtime log, dates were forward-filled. Missing downtime aggregates (total duration and incident count) were set to zero, reflecting the absence of an event. The meteorological series was resampled at hourly intervals, and missing values were time-interpolated. During feature engineering, target-variable outliers identified by the three-standard-deviation rule were replaced with the mean of the nearest valid neighboring values or, if those values were unavailable, with the series median. Principal components within a group were calculated using complete cases, followed by alignment to the hourly grid. Median feature imputation was used when multicollinearity was assessed with variance inflation factors. Before the data were passed to the model, any remaining missing valueswere first forward-filled and then backward-filled over time as a safeguard, ensuring continuity of the final hourly series from January 1, 2022, through December 31, 2023 (17,520 observations).

The dataset was split chronologically without shuffling: the complete hourly series from January 1, 2022, through December 31, 2023 (17,520 observations) was divided into a training set (January 1, 2022–June 30, 2023; 13,104 observations), a validation set (July 1–September 30, 2023; 2,208 observations), and a test set (October 1–December 31, 2023; 2,208 observations), representing approximately 74.8%, 12.6%, and 12.6%, respectively. With a 168 h context window (seven days of history) and a 24 h forecasting horizon, the resulting numbers of training, validation, and test windows were 12,913, 2,017, and 2,017 for the direct-strategy configurations and 12,936, 2,040, and 2,017 for the recursive-strategy configurations (one-step training and 24-step testing).

1.2. Neural network architecture

Gated recurrent units (GRUs) [4] were selected as the baseline architecture. A GRU is a simplified variant of an LSTM [18] that combines the forget and input gates into a single update gate. The mathematical model of a GRU cell is described by the following system of equations:

                                                         zt = σ(Wz ∙ [ht – 1, xt] + bz),

                                                         rt = σ(Wr ∙ [ht – 1, xt] + br),

                                                         h̃t = tanh (Wh ∙ [rt ⊙ ht – 1, xt] + bh),

                                                         ht = (1 − zt) ⊙ ht – 1 + zt ⊙ h̃t,

where zt is the update gate; rt is the reset gate; ht is the hidden state; and ⊙ denotes element-wise multiplication.

The architecture of the developed model (Fig. 1) consists of the following layers:

  1. Input layer with shape (batch, 168, 6).
  2. First GRU layer: 64 units, tanh activation, full-sequence output, and a recurrent dropout rate of 0.2.
  3. Dropout regularization layer with a rate of 0.2.
  4. Second GRU layer: 64 units, tanh activation, and full-sequence output.
  5. Dropout regularization layer with a rate of 0.2.
  6. Third GRU layer: 128 units, tanh activation, and context-vector output.
  7. Dropout regularization layer with a rate of 0.15.
  8. Dense output layer: 24 units (forecast horizon) and linear activation.

The input shape (batch, 168, 6) corresponds to a 168 h context window (one week) and the six input channels listed in Section 1.1 (the target variable and five exogenous variables). Recursive-strategy configurations (see Section 1.3) use a separate univariate input of shape (N, 168, 1), containing only the target-variable history and no exogenous variables.

Fig. 1. Topology of the GRU neural network model (Direct). Numbers in the blocks indicate the number of units; d is the dropout rate; r is the recurrent dropout rate

The number of trainable model parameters (Table 1) is appropriate for the size of the available dataset and balances approximation accuracy against generalization capability.

Table 1

Summary results of the experimental model comparison

Model

MAPE, %

RMSE

Variance

ratio

Penalty, RUB million

Savings, %

Trainable parameters*

GRU (Direct)

5.44

59.77

0.402

143.2

21.25

116,376

Attention-GRU

6.12

73.56

0.237

149.8

17.58

47,328

RNN (Direct)

6.42

76.42

0.201

158.4

12.88

53,272

LSTM (Direct)

6.49

75.15

0.283

161.2

11.32

207,896

BiLSTM (Direct)

6.60

76.49

0.238

165.5

8.94

83,736

RNN (Recursive)

6.85

83.86

0.179

166.3

8.52

13,569

CNN-LSTM (Direct)

6.68

76.74

0.192

166.3

8.52

86,136

Ridge AR

6.78

82.19

0.249

167.0

8.14

–

BiLSTM (Recursive)

6.88

84.35

0.184

167.9

7.65

75,457

LSTM (Recursive)

6.94

85.31

0.183

168.8

7.14

111,841

GRU (Recursive)

6.97

87.37

0.247

169.0

7.04

115,425

Ridge ADL

6.82

80.13

0.292

169.4

6.81

–

CNN-LSTM (Recursive)

7.23

87.59

0.186

176.2

3.05

82,689

S-Naive (Daily)

7.35

96.52

0.750

181.8

0.00

–

LSTM-Attention (Direct)

7.48

83.53

0.090

186.0

−2.33

50,457

S-Naive (Weekly)

9.09

104.22

0.735

231.6

−27.42

–

Enterprise plan

14.32

144.26

0.082

399.2

−119.60

–

*For neural network models, the total number of trainable parameters is reported. For the other configurations (regularized Ridge AR and Ridge ADL autoregressions, seasonal naïve S-Naive models, and the enterprise’s current plan), the number of trainable parameters is not reported: the Ridge models have linear coefficients that are not meaningfully comparable with neural network weights, whereas the naïve models and the current plan have no trainable parameters. A dash is shown in the respective cells.

1.3. Rationale for the multi-step forecasting strategy

The short-term forecasting problem is formulated as learning a mapping from L = 168 h of historical data (a one-week context window) to an H = 24 h forecast vector.

Two multi-step forecasting strategies were considered [19]: Direct Multi-Output, or MIMO [17], and Recursive [17, 20]. Hereafter, “Direct,” “direct strategy,” and “direct method” are used as synonyms for Direct Multi-Output, whereas “recursive scheme,” “recursive configuration,” and “recursive method” are used as synonyms for Recursive. The direct strategy generates the entire forecast vector in a single pass through the neural network using an output layer of size 24. The recursive strategy generates the forecast iteratively: each prediction is fed back to the model as input for the next step in place of the unknown observation. An application-specific consideration in selecting the strategy is the stochastic nature of the PCA factors (equipment downtime and on-site generation mode): in a recursive scheme, their values at each future forecasting step are unavailable as deterministic quantities and would require a separate model to extrapolate the exogenous factors.

Consequently, the Direct and Recursive configurations in this comparison differ in input structure: the direct strategy simultaneously uses a 168 h window for the target variable and the PCA factors and generates a 24 h forecast in a single pass, whereas the recursive configuration in Table 1 is implemented in univariate mode (the input contains only a window of past target-variable values with a shape of N × 168 × 1) to isolate error accumulation over the 24 h horizon. The Recursive values in Table 1 provide a reference for cases in which exogenous information is unavailable over the future horizon and are used to analyze the contribution of the recursive scheme to error accumulation (Section 3).

The strategy was selected through a comparative analysis of historical data (see Section 2).

1.4. Hyperparameter selection method

Hyperparameters were selected using the Hyperband algorithm [21], which adaptively allocates computational resources among competing configurations. The algorithm is based on successive halving: poorly performing configurations are eliminated early in training, reducing the overall search time by a factor of 5–10 compared with a standard random search [22].

The hyperparameter search space included the network architecture (RNN, LSTM, GRU, BiLSTM, CNN–LSTM, and attention-based models), the number of recurrent layers (1–3), hidden-layer size (32–192 neurons), dropout rate (0.1–0.3), learning rate (10⁻⁴–10⁻², logarithmic scale), and batch size (32, 64, or 128).

The model was trained using the Adam optimizer with a learning rate of 2 × 10⁻³; ReduceLROnPlateau dynamically reduced the learning rate when the monitored metric plateaued.Mean squared error (MSE) was used as the loss function. Early stopping was applied when the validation metric did not improve for 10 epochs. The final model converged at epoch 38 of a maximum of 50 epochs.

The training dynamics of the GRU model are shown in Fig. 2.

Fig. 2. Learning curves of the GRU model (Direct): MSE dynamics for the training and validation sets. Best epoch: 25

1.5. Model performance metrics

Model performance was evaluated on the test set (Q4 2023; 85 days, 2,208 h) in a daily forecasting mode. The following metrics were used:

  • MAPE (mean absolute percentage error), the primary accuracy metric.
  • RMSE (root mean squared error), which is sensitive to outliers.
  • Variance ratio, the ratio of forecast variance to observed variance, which characterizes the model’s ability to reproduce process volatility.
  • Total penalty, the WECM penalty calculated as the sum of the products of the absolute error and the hourly electricity price2.
  • Savings, the reduction in penalty payments relative to the baseline forecast – the seasonal naïve model (S-Naive Daily), which uses the previous day’s profile.

The seasonal naïve model (S-Naive Daily) was used as the methodological baseline for estimating savings in penalty payments: it provides a minimum performance benchmark for any nontrivial model [2, 19].

2 Resolution of the Government of the Russian Federation No. 1172 dated December 27, 2010, “On the Rules of the Wholesale Electricity and Capacity Market” (as amended and supplemented). URL: http://pravo.gov.ru/proxy/ips/?docbody=&nd=102144342

2. Results

2.1. Summary results of the experimental comparison

The study compared 17 configurations forming a full factorial design: 12 neural network configurations (six architectures – RNN, GRU, LSTM, BiLSTM, CNN–LSTM [23, 24], and an attention-based model – each using two multi-step forecasting strategies, Direct Multi-Output and Recursive), two regularized autoregressions (Ridge AR and Ridge ADL), two seasonal naïve models (S-Naive Daily and S-Naive Weekly), and the enterprise’s current plan as a reference. The test results for Q4 2023 are presented in Table 1.

The experimental design comprised a full factorial set of 17 configurations, enabling the separate effects of architecture selection, multi-step forecasting strategy, and mapping linearity to be evaluated.

2.2. Analysis of the leading model

The GRU model using the direct forecasting strategy performed best in the experiments. This configuration achieved the highest accuracy (MAPE of 5.44%) and the greatest economic benefit (savings of 21.25%), considerably outperforming both the naïve baselines and more complex attention-based architectures.

The variance ratio of the GRU model using the direct strategy was 0.402, indicating that the model reproduced approximately 40% of the natural volatility of the electricity consumption process. This value was higher than those of the recursive-strategy models (0.179–0.247), which exhibited pronounced forecast smoothing because of error accumulation.

The GRU architecture outperformed the more complex LSTM architecture (MAPE of 6.49%) while using fewer trainable parameters.

The model’s ability to track abrupt load changes is illustrated in Fig. 3 using the load drop from October 30 to November 1, 2023, as an example. Over the 48 h forecasting horizon, the MAPE was 17.22% for the GRU (Direct) model and 43.70% for the linear multivariate Ridge ADL (168, 6) model using the same six-channel input. The 95% prediction intervals shown in the figure were calculated as ±1.96σ̂ using the empirical standard deviation of the error for each horizon hour in the test set.

Fig. 3. Comparison of the GRU (Direct) and Ridge ADL (168, 6) model forecasts during an abrupt load decrease (October 30–November 1, 2023). The main plot shows actual consumption, point forecasts from both models, and 95% prediction intervals based on the empirical error estimate for the test set; the MAPE for each 24 h forecast segment is shown within each forecast day. The lower panel shows the hourly absolute percentage error (APE), with horizontal lines indicating the mean MAPE over the 48 h horizon: GRU (Direct), 17.22%; Ridge ADL (168, 6), 43.70%

2.3. Comparison with statistical baseline models

To assess the contribution of nonlinear relationships to forecasting accuracy, the model was compared with regularized statistical models. Ridge AR (168), a univariate model with a 168 h lag window, achieved a MAPE of 6.78%, whereas Ridge ADL (168, 6), with a six-channel input, achieved a MAPE of 6.82%. Ridge AR (168) was used as a univariate linear benchmark that separates the contribution of nonlinearity from that of the exogenous variables. Ridge ADL (168, 6) is shown in Fig. 3 (see Section 2.2) as a linear model with the same six-channel input as the GRU.

The GRU neural network model reduced MAPE by 1.34 percentage points relative to Ridge AR and increased penalty-payment savings by 14.3%. This difference quantitatively characterizes the contribution of nonlinear effects: according to [11], neural network models account for nonlinear seasonality and trends more accurately than their linear counterparts, as confirmed here by the recurrent GRU architecture.

The variance ratios of the forecasts from the different models are compared in Fig. 4.

Fig. 4. Variance ratio (VR) of model forecasts: VR = 1.0 indicates that the forecast reproduces the variance of the observed series in full

2.4. Economic performance

The economic benefit was evaluated as the reduction in penalties in the Wholesale Electricity and Capacity Market (WECM). Under the current market rules3, the penalty is calculated as the product of the absolute forecasting error and the hourly electricity price.

When the GRU model using the direct strategy replaced the best-performing naïve benchmark – the daily seasonal naïve model (S-Naive Daily) – MAPE decreased by 1.91 percentage points (from 7.35% to 5.44%). Over the test period (Q4 2023), penalty-payment savings amounted to RUB 38.6 million, equivalent to a reduction of 21.25%. The same calculation can be performed for subsequent periods if the operating mode remains unchanged.

3 Resolution of the Government of the Russian Federation No. 1172 dated December 27, 2010, “On the Rules of the Wholesale Electricity and Capacity Market” (as amended and supplemented). URL: http://pravo.gov.ru/proxy/ips/?docbody=&nd=102144342

3. Discussion

3.1. Effects of excessive model complexity

The study revealed a “complexity paradox”: the simpler GRU architecture outperformed attention-based models (Attention-GRU and LSTM-Attention), despite the latter’s theoretical advantages. This result is consistent with the critical review in [10], which showed that transformer models tend to overfit on small and medium-sized datasets containing fewer than 3–5 years of hourly data.

For small and medium-sized industrial datasets, the training signal is insufficient to support the full parametric capacity of the attention mechanism. This increases sensitivity to noise and suppresses forecast volatility to the point that the model effectively reproduces a smoothed average profile. The results obtained on the limited industrial dataset are consistent with the review of recurrent and transformer models [10] (see Table 1).

The GRU architecture uses two control gates (update and reset) rather than the three used by LSTM (input, forget, and output), thereby reducing the number of trainable parameters for a given hidden-state size. For industrial series with recurring missing values and outliers, this reduces the tendency to overfit and increases robustness to outliers and missing data, as observed in the summary results (see Table 1).

3.2. Critical analysis of the recursive strategy

The comparison of strategies showed that the direct method outperformed the recursive method in terms of MAPE and variance ratio for all architectures studied (the numerical differences are reported in Table 1).

The accuracy of the recursive strategy deteriorates because errors accumulate: the forecast at step k contains an error that is passed to the input when the forecast for step k + 1 is generated. Over a 24 h horizon, the resulting forecasts systematically underestimate volatility, as reflected by the variance ratios of the recursive-strategy configurations (see Table 1).

An additional limitation of the recursive strategy in the present problem is its inability to use exogenous PCA factors representing stochastic process modes. The direct strategy accounts for the multivariate nature of electricity consumption without requiring the factors themselves to be forecast.

3.3. Comparison with the current planning system

The enterprise’s current plan is produced using its established procedure for preparing and submitting planned bids to the WECM. Its numerical indicators are provided in Table 1 as a reference; this study does not qualitatively evaluate the current planning system.

3.4. Study limitations and directions for further research

Several limitations of the study should be noted. First, the model was developed and validated using data from one enterprise, and the transferability of the results to other mining and processing enterprises therefore requires further evaluation. The training sample contained approximately 13,000 hourly observations (see Section 1.1) from a single site, limiting generalizability and precluding a comparative analysis of multiple enterprises in this study. Second, the model produces deterministic point forecasts and does not estimate uncertainty, limiting its applicability to risk management in the balancing market.

Promising directions for further research include expanding the study to a cluster of enterprises with different process cycles, developing a probabilistic version of the model based on the DeepAR architecture [25] to generate prediction intervals, and integrating the model into an automated system for submitting WECM bids.

Conclusions

A comparison of six recurrent and hybrid architectures using data from a mining and processing enterprise with continuously operating core production showed that GRU provided the highest electricity consumption forecasting accuracy among the models considered. Attention-based architectures underperformed simpler recurrent models on the limited dataset; LSTM-Attention performed worst, producing an excessively smoothed average forecast profile (see Table 1). The direct multi-step forecasting strategy was substantiated as the only applicable approach when using stochastic composite PCA factors whose future values are not deterministically known. It preserved the natural process volatility: the GRU model’s variance ratio exceeded that of the enterprise’s current plan (see Table 1). Comparison with regularized Ridge AR confirmed the significance of nonlinear relationships in the data structure. Hyperparameters were selected using Hyperband, and validation was performed on 2022–2023 industrial data (~17,500 hourly measurements). Over the test period (Q4 2023), the developed model yielded savings of RUB 38.6 million – a 21.25% reduction in penalty payments relative to the S-Naive Daily baseline.

References

1. Klyuev R. V., Bosikov I. I., Gavrina O. A. et al. Improving the energy efficiency of technological equipment at mining enterprises. Advances in Intelligent Systems and Computing. 2021;1258:262–271. https://doi.org/10.1007/978-3-030-57450-5_24

2. Hyndman R. J., Koehler A. B. Another look at measures of forecast accuracy. International Journal of Forecasting. 2006;22(4):679–688. https://doi.org/10.1016/j.ijforecast.2006.03.001

3. Hochreiter S., Schmidhuber J. Long short-term memory. Neural Computation. 1997;9(8):1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735

4. Cho K., van Merrienboer B., Gulcehre C. et al. Learning phrase representations using RNN encoder-decoder for statistical machine translation. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP). Doha: ACL; 2014. Pp. 1724–1734. https://doi.org/10.3115/v1/D14-1179

5. Bengio Y., Simard P., Frasconi P. Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks. 1994;5(2):157–166. https://doi.org/10.1109/72.279181

6. Vaswani A., Shazeer N., Parmar N. et al. Attention is all you need. In: Advances in Neural Information Processing Systems. 2017;30:5998–6008. https://doi.org/10.48550/arXiv.1706.03762

7. Lim B., Arik S. O., Loeff N., Pfister T. Temporal Fusion Transformers for interpretable multi-horizon time series forecasting. International Journal of Forecasting. 2021;37(4):1748–1764. https://doi.org/10.1016/j.ijforecast.2021.03.012

8. Zhou H., Zhang S., Peng J. et al. Informer: Beyond efficient transformer for long sequence time-series forecasting. In: Proceedings of the AAAI Conference on Artificial Intelligence. 2021;35(12):11106–11115. https://doi.org/10.1609/aaai.v35i12.17325

9. Wu H., Xu J., Wang J., Long M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. In: Advances in Neural Information Processing Systems. 2021;34:22419–22430. https://doi.org/10.48550/arXiv.2106.13008

10. Zeng A., Chen M., Zhang L., Xu Q. Are transformers effective for time series forecasting? In: Proceedings of the AAAI Conference on Artificial Intelligence. 2023;37(9):11121–11128. https://doi.org/10.1609/aaai.v37i9.26317

11. Zhang G. P., Qi M. Neural network forecasting for seasonal and trend time series. European Journal of Operational Research. 2005;160(2):501–514. https://doi.org/10.1016/j.ejor.2003.08.037

12. Ben Taieb S., Atiya A. F. A bias and variance analysis for multistep-ahead time series forecasting. IEEE Transactions on Neural Networks and Learning Systems. 2016;27(1):62–76. https://doi.org/10.1109/TNNLS.2015.2411629

13. Karpenko S. M., Karpenko N. V., Bezginov G. Y. Forecasting of power consumption at mining enterprises using statistical methods. Russian Mining Industry. 2022;(1):82–88. (In Russ.) https://doi.org/10.30686/1609-9192-2022-1-82-88

14. Hewamalage H., Bergmeir C., Bandara K. Recurrent neural networks for time series forecasting: current status and future directions. International Journal of Forecasting. 2021;37(1):388–427. https://doi.org/10.1016/j.ijforecast.2020.06.008

15. Sangiorgio M., Dercole F. Robustness of LSTM neural networks for multi-step forecasting of chaotic time series. Chaos, Solitons & Fractals. 2020;139:110045. https://doi.org/10.1016/j.chaos.2020.110045

16. Zawodnik V., Schwaiger F., Sorger M., Kienberger T. Tackling uncertainty: forecasting the energy consumption and demand of an electric arc furnace with limited knowledge on process parameters. Energies. 2024;17(6):1326. https://doi.org/10.3390/en17061326

17. Ben Taieb S., Bontempi G., Atiya A. F., Sorjamaa A. A review and comparison of strategies for multi-step ahead time series forecasting based on the NN5 forecasting competition. Expert Systems with Applications. 2012;39(8):7067–7083. https://doi.org/10.1016/j.eswa.2012.01.039

18. Li H., Li S., Wu Y. et al. Short-term power load forecasting for integrated energy system based on a residual and attentive LSTM-TCN hybrid network. Frontiers in Energy Research. 2024;12:1384142. https://doi.org/10.3389/fenrg.2024.1384142

19. Bergstra J., Bengio Y. Random search for hyper-parameter optimization. Journal of Machine Learning Research. 2012;13(1):281–305. URL: https://www.jmlr.org/papers/v13/bergstra12a.html

20. Marcellino M., Stock J. H., Watson M. W. A comparison of direct and iterated multistep AR methods for forecasting macroeconomic time series. Journal of Econometrics. 2006;135(1–2):499–526. https://doi.org/10.1016/j.jeconom.2005.07.020

21. Li L., Jamieson K., DeSalvo G., Rostamizadeh A., Talwalkar A. Hyperband: A novel bandit-based approach to hyperparameter optimization. Journal of Machine Learning Research. 2018;18(185):1–52. https://doi.org/10.48550/arXiv.1603.06560

22. Chung J., Gulcehre C., Cho K., Bengio Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. In: NIPS 2014 Workshop on Deep Learning. 2014. https://doi.org/10.48550/arXiv.1412.3555

23. Bai S., Kolter J. Z., Koltun V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint: 1803.01271. 2018. https://doi.org/10.48550/arXiv.1803.01271

24. Alhussein M., Aurangzeb K., Haider S. I. Hybrid CNN-LSTM model for short-term individual household load forecasting. IEEE Access. 2020;8:180544–180557. https://doi.org/10.1109/ACCESS.2020.3028281

25. Salinas D., Flunkert V., Gasthaus J., Januschowski T. DeepAR: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting. 2020;36(3):1181–1191. https://doi.org/10.1016/j.ijforecast.2019.07.001


About the Authors

E. A. Ematin
University of Science and Technology MISIS
Russian Federation

Evgenii A. Ematin – PhD-Student, Department of Energy and Energy Efficiency of the Mining Industry

Moscow

Scopus ID 59012304200

SPIN 2464-1230



S. M. Karpenko
University of Science and Technology MISIS
Russian Federation

Sergey M. Karpenko – Cand. Sci. (Eng.), Associate Professor of the Department of Energy and Energy Efficiency of the Mining Industry

Moscow

Scopus ID 57225144783

SPIN 3950-0362



Review

For citations:


Ematin E.A., Karpenko S.M. Neural network models for short-term electricity consumption forecasting at a mining enterprise participating in the wholesale electricity market. Mining Science and Technology (Russia). https://doi.org/10.17073/2500-0632-2026-01-1070

Views: 21

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 2500-0632 (Online)