Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Forecasting techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

HAIMOS Ensemble Forecasts for Intra-day and Day- Ahead GHI, DNI and Ramps

The objective of this research is to develop a hybrid physics-based/data-driven forecast model to improve direct normal and global horizontal irradiance (DNI and GHI) prediction for horizons ranging from 1 to 72 hours. Project objectives also address key gaps in state-of-the-art solar forecasting: accurate probabilistic solar forecasts and the forecasting of large irradiance ramps (ramp onset and magnitude). The proposed model ensembles Numerical Weather Prediction (NWP) forecasts, determinist physics-based algorithms, and new-generation cloud cover products (high-resolution rapid refresh satellite images and Large Eddy Simulations). The result is the Hybrid Adaptive Input Model Objective Selection (HAIMOS) ensemble model. HAIMOS blends state of the art machine learning methodologies with physics-based models for cloud cover and cloud optical depth forecasts. The technical activities followed a two-pronged strategy. First, the preprocessing of data, the selection of inputs to the nonlinear approximators, the type of approximator and objective functions, and post-processing ensembling techniques included in HAIMOS were all optimized adaptively to find the best model for a specific goal (reduce DNI/GHI forecast error, improve the prediction of ramp onset, etc.). Second, a large effort was put in improving cloud identification and the forecast of cloud cover and cloud optical depth. To this end, new-generation cloud parametrization products were developed in this work. These include improved algorithms to assist in cloud identification, cloud classification and cloud parametrization from satellite images – three key factors in the accuracy of 1 to 6-hours irradiance forecasts and prediction of ramp onset. Furthermore, we also included cloud information extracted from high resolution rapid refresh satellite images (GOES-16) and Large Eddy Simulations (LES). LES was used to model the atmosphere in detail over locations of interest and produce cloud optical depth forecasts. Once these data streams were validated, they were used as input data to the HAIMOS forecast. The model was developed using data from several climatologically distinct locations with potential for high solar penetration. In the last year of the project, we conducted a validation campaign according to the guidelines stipulated by the Topic Area 1 project as described in the FOA. This effort brings, for the first time, proven machine-learning methodologies for generating state-of-the-art solar forecasts interweaved with detailed physics-based models for cloud detection, and cloud optical depth forecasts. HAIMOS will generate accurate irradiance probabilistic forecast to assist in reducing solar generation prediction error. Globally optimized solar forecast models are more likely to impact solar energy stakeholders. The goal of this project was to increase the state-of-the-art forecast skill from their present values of 10 to 35%. At the end of the project, we achieved between 30% and 50% forecast skill across a wide range of horizons for both GHI and DNI.

14 SOLAR ENERGY↗

GenAI4UQ: A software for forward and inverse uncertainty quantification using conditional generative AI

We introduce GenAI4UQ, a software package for forward and inverse uncertainty quantification in model calibration, parameter estimation, and ensemble forecasting. GenAI4UQ leverages a generative AI-based conditional modeling framework to address limitations of traditional inverse modeling techniques, such as Markov Chain Monte Carlo (MCMC) methods. By replacing computationally intensive iterative processes with a direct, learned mapping, GenAI4UQ enables efficient calibration of input parameters and generation of predictions directly from observations. The software supports rapid ensemble forecasting with robust uncertainty quantification while maintaining computational and storage efficiency. Built-in auto-tuning of hyperparameters simplifies model training, ensuring accessibility for users with varying expertise. Its versatile conditional generative framework is applicable across diverse scientific domains. While GenAI4UQ offers significant advantages in flexibility and efficiency, users should interpret its uncertainty estimates with caution in data-sparse scenarios, as the model may overestimate uncertainty—an effect common to all surrogate-based approaches including MCMC with surrogate models. Despite this, GenAI4UQ transforms inverse modeling by providing a fast, reliable, and user-friendly solution. It empowers researchers and practitioners to quickly estimate parameter distributions and generate model predictions for new observations, facilitating efficient decision-making and advancing the state of uncertainty quantification in computational modeling.

97 MATHEMATICS AND COMPUTING↗

The large scale polarization explorer (LSPE) for CMB measurements: performance forecast

The measurement of the polarization of the Cosmic Microwave Background (CMB) radiation is one of the current frontiers in cosmology. In particular, the detection of the divergence-free component of the polarization field, the B-mode component, reveals the presence of gravitational waves in the early Universe. The detection of such component is at the moment the most promising technique to probe the inflationary theory describing the very early evolution of the Universe. The measurement of the polarization of the Cosmic Microwave Background (CMB) radiation is one of the current frontiers in cosmology. In particular, the detection of the primordial divergence-free component of the polarization field, the B-mode, could reveal the presence of gravitational waves in the early Universe. The detection of such a component is at the moment the most promising technique to probe the inflationary theory describing the very early evolution of the Universe. We present the updated performance forecast of the Large Scale Polarization Explorer (LSPE), a program dedicated to the measurement of the CMB polarization. LSPE is composed of two instruments: LSPE-Strip, a radiometer-based telescope on the ground in Tenerife-Teide observatory, and LSPE-SWIPE (Short-Wavelength Instrument for the Polarization Explorer) a bolometer-based instrument designed to fly on a winter arctic stratospheric long-duration balloon. The program is among the few dedicated to observation of the Northern Hemisphere, while most of the international effort is focused into ground-based observation in the Southern Hemisphere. Measurements are currently scheduled in Winter 2022/23 for LSPE-SWIPE, with a flight duration up to 15 days, and in Summer 2022 with two years observations for LSPE-Strip. In this work, we describe the main features of the two instruments, identifying the most critical aspects of the design, in terms of impact on the performance forecast. We estimate the expected sensitivity of each instrument and propagate their combined observing power to the sensitivity to cosmological parameters, including the effect of scanning strategy, component separation, residual foregrounds and partial sky coverage. We also set requirements on the control of the most critical systematic effects and describe techniques to mitigate their impact. LSPE will reach a sensitivity in tensor-to-scalar ratio of σr < 0.01, set an upper limit r < 0.015 at 95% confidence level, and improve constraints on other cosmological parameters.

79 ASTRONOMY AND ASTROPHYSICS↗

A Data-Agnostic, Continuous Machine Learning Framework for Application in High Energy Physics and Beyond: Phase 1 Final Scientific/Technical Report

This Phase 1 effort has focused on the development of continual learning frameworks for use in machine learning, specifically in the applied context of High Energy Physics (HEP). Machine learning (ML) is a transformative technology by which computers, typically through the use of neural networks, are able to perform tasks with proficiency that rivals or surpasses that of human users. Model Degradation & Catastrophic Forgetting are two undesired phenomena which can occur in ML where the performance of a model degrades when either deployed on novel data streams, or trained on novel data which are sufficiently different than the data the models were initially trained on. A natural example where these sorts of effects can be observed is in the performance of detectors in harsh environments, where the detector signature may change over the lifetime of the detector as it ages and deteriorates — precisely what occurs in the experiments conducted in HEP. Real world HEP data is therefore an excellent test-ground and use-case for Continual Learning paradigms, which are techniques used in ML to counteract these problems. Ensemble learning is one such technique, where multiple smaller models are trained on subsets of the overall data and are ensembled together during inference. The intuition behind this technique is that, although there are shifts in the distributions which govern the incoming data streams, these shifts are not expected to be homogeneous or global. If a sufficient diversity in solutions within the various sub-models has been achieved, then at least one sub-model is expected to retain its performance within the overall ensemble. One further strength of this approach is that the architectures of the various models do not need to be identical, and in fact even different modalities of data can naturally be combined in this way. This work focused on applying ensemble learning techniques to derive results using two main datasets, anomaly detection in HEP data & time-series forecasting in semiconductor manufacturing data. Semiconductor manufacturing involves data with surprising similarity to that of HEP (e.g. wafer maps look very similar to digi-occupancy maps) and Cerium Lab’s prominence within the semiconductor industry makes semiconductor manufacturing a natural opportunity for commercialization of this work. Our efforts have led to two strong results. The first is that we evaluated the proposed ensembling techniques using previously proposed machine learning architectures for use in anomaly detection, namely AutoEncoder based models and their derivatives. We also developed new architectures which have not been evaluated in this context before. In fact, this work marks the first use of Vision Transformers for anomaly detection in HEP. Second, we demonstrated that ensemble learning significantly improves model performance in scenarios prone to degradation, validating its effectiveness across both HEP and semiconductor datasets. These results further support ensemble learning as a powerful strategy for mitigating catastrophic forgetting and maintaining robust performance in evolving data environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Novel Modeling Framework for Computationally Efficient and Accurate Real-Time Ensemble Flood Forecasting With Uncertainty Quantification

A novel modeling framework that simultaneously improves accuracy, predictability, and computational efficiency is presented. It embraces the benefits of three modeling techniques integrated together for the first time: surrogate modeling, parameter inference, and data assimilation. The use of polynomial chaos expansion (PCE) surrogates significantly decreases computational time. Parameter inference allows for model faster convergence, reduced uncertainty, and superior accuracy of simulated results. Ensemble Kalman filters assimilate errors that occur during forecasting. To examine the applicability and effectiveness of the integrated framework, we developed 18 approaches according to how surrogate models are constructed, what type of parameter distributions are used as model inputs, and whether model parameters are updated during the data assimilation procedure. We conclude that (1) PCE must be built over various forcing and flow conditions, and in contrast to previous studies, it does not need to be rebuilt at each time step; (2) model parameter specification that relies on constrained, posterior information of parameters (so-called Selected specification) can significantly improve forecasting performance and reduce uncertainty bounds compared to Random specification using prior information of parameters; and (3) no substantial differences in results exist between single and dual ensemble Kalman filters, but the latter better simulates flood peaks. The use of PCE effectively compensates for the computational load added by the parameter inference and data assimilation (up to ~80 times faster). Therefore, the presented approach contributes to a shift in modeling paradigm arguing that complex, high-fidelity hydrologic and hydraulic models should be increasingly adopted for real-time and ensemble flood forecasting.

54 ENVIRONMENTAL SCIENCES↗

Analytics-at-scale of Sensor Data for Digital Monitoring in Nuclear Plants (3 rd Annual Report)

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This report primarily focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 ± 0.014, 0.0026 ± 0, and 0.063 ± 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure. This report summarizes the Fiscal Year 2021 research progress encompassing the (1) data cleaning and feature selection necessary for ML applications; (2) development of short-term forecasting models to predict future plant process parameters for both single and multiple time steps ahead; and (3) validation of the feature selection methods and short-term forecasting models given new data from different systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

State-of-the-Art Techniques for Large-Scale Stochastic Unit Commitment

Recent advances in deterministic unit commitment, both formulaic and algorithmic, along with modern algorithmic approaches for stochastic programming, have enabled the solution of stochastic unit commitment problems with hundreds of scenarios on large-scale transmission networks. In this presentation, we will give an overview of these methods, including lazy transmission constraint generation, lower-bounding techniques, and heuristics, all of which can be executed in concert with customized decomposition approaches for optimization under uncertainty. We demonstrate the effectiveness of these techniques on the TAMU Texas7K synthetic transmission network, leveraging realistic high-resolution forecasts based on NREL renewable resource availability data. The software leveraged for these demonstrations is available via the open-source software packages EGRET (for electrical grid optimization) and mpi-sppy (for optimization under uncertainty).

27 ARPA - Advanced Research Projects Agency-Energy↗

Evaluating downscaled products with expected hydroclimatic co-variances

Abstract. There has been widespread adoption of downscaled products amongst practitioners and stakeholders to ascertain risk from climate hazards at the local scale (e.g., ∼ 5 km resolution). Such products must nevertheless be consistent with physical laws to be credible and of value to users. Here we evaluate statistically and dynamically downscaled products by examining local co-evolution of downscaled temperature and precipitation during convective and frontal precipitation events (two mechanisms testable with just temperature and precipitation). We find that two widely used statistical downscaling techniques (Localized Constructed Analogs version 2, LOCA2, and Seasonal Trends and Analysis of Residuals Empirical Statistical Downscaling Model, STAR-ESDM) generally preserve expected co-variances during convective precipitation events over the historical and future projected intervals as compared to European Centre for Medium-Range Weather Forecasts Reanalysis v5 (ERA5) and two observation-based data products (Livneh and nClimGrid-Daily). However, both techniques dampen future intensification of frontal precipitation that is otherwise robustly captured in global climate models (i.e., prior to downscaling) and with process-based dynamical downscaling across five different regional climate models. In the case of LOCA2, this leads to appreciable underestimation of future frontal precipitation event intensity. This study is one of the first to quantify a likely ramification of the stationarity assumption underlying statistical downscaling methods and identify a phenomenon where projections of future change diverge depending on data production method employed. Finally, our work proposes expected co-variances during convective and frontal precipitation as useful evaluation diagnostics that can be universally applied to a wide range of statistically downscaled products.

54 ENVIRONMENTAL SCIENCES↗

Non-autoregressive time-series methods for stable parametric reduced-order models

Advection-dominated dynamical systems, characterized by partial differential equations, are found in applications ranging from weather forecasting to engineering design where accuracy and robustness are crucial. There has been significant interest in the use of techniques borrowed from machine learning to reduce the computational expense and/or improve the accuracy of predictions for these systems. These rely on the identification of a basis that reduces the dimensionality of the problem and the subsequent use of time series and sequential learning methods to forecast the evolution of the reduced state. Often, however, machine-learned predictions after reduced-basis projection are plagued by issues of stability stemming from incomplete capture of multiscale processes as well as due to error growth for long forecast durations. To address these issues, we have developed a non-autoregressive time series approach for predicting linear reduced-basis time histories of forward models. In particular, we demonstrate that non-autoregressive counterparts of sequential learning methods such as long short-term memory (LSTM) considerably improve the stability of machine-learned reduced-order models. Further, we evaluate our approach on the inviscid shallow water equations and show that a non-autoregressive variant of the standard LSTM approach that is bidirectional in the principal component directions obtains the best accuracy for recreating the nonlinear dynamics of partial observations. Moreover-and critical for many applications of these surrogates-inference times are reduced by three orders of magnitude using our approach, compared with both the equation-based Galerkin projection method and the standard LSTM approach.

97 MATHEMATICS AND COMPUTING↗

Local-thermal-gradient and large-scale-circulation impacts on turbine-height wind speed forecasting over the Columbia River Basin

Abstract. We investigate the sensitivity of turbine-height wind speed forecast to initial condition (IC) uncertainties over the Columbia River Gorge (CRG) and Columbia River Basin (CRB) for two typical weather phenomena, i.e., local-thermal-gradient-induced marine air intrusion and a cold frontal passage. Four types of turbine-height wind forecast anomalies and their associated IC uncertainties related to local thermal gradients and large-scale circulations are identified using the self-organizing map (SOM) technique. The four SOM types are categorized into two patterns, each accounting for half of the ensemble members. The first pattern corresponds to IC uncertainties that alter the wind forecast through a modulating weather system, which produces the strongest wind anomalies in the CRG and CRB. In the second pattern, the moderate uncertainties in local thermal gradient and large-scale circulation jointly contribute to wind forecast anomaly. We analyze the cross section of wind and temperature anomalies through the gorge to explore the evolution of vertical features of each SOM type. The turbine-height wind anomalies induced by large-scale IC uncertainties are more concentrated near the front. In contrast, turbine-height wind anomalies induced by the local IC thermal uncertainties are found above the surface thermal anomalies. Moreover, the wind forecast accuracy in the CRG and CRB is limited by IC uncertainties in a few specific regions, e.g., the 2 m temperature within the basin and large-scale circulation over the northeast Pacific around 140∘ W.

17 WIND ENERGY↗

PyDDA: A Pythonic Direct Data Assimilation Framework for Wind Retrievals

This software assimilates data from an arbitrary number of weather radars together with other spatial wind fields (eg numerical weather forecasting model data) in order to retrieve high resolution three dimensional wind fields. PyDDA uses NumPy and SciPy’s optimization techniques combined with the Python Atmospheric Radiation Measurement (ARM) Radar Toolkit (Py-ART) in order to create wind fields using the 3D variational technique (3DVAR). PyDDA is hosted and distributed on GitHub at https://github.com/openradar/PyDDA. PyDDA has the potential to be used by the atmospheric science community to develop high resolution wind retrievals from radar networks. These retrievals can be used for the evaluation of numerical weather forecasting models and plume modelling. This paper shows how wind fields from 2 NEXt generation RADar (NEXRAD) WSR-88D radars and the High Resolution Rapid Refresh can be assimilated together using PyDDA to create a high resolution wind field inside Hurricane Florence.

54 ENVIRONMENTAL SCIENCES↗

Peeking into the next decade in large-scale structure cosmology with its Effective Field Theory

After the successful full-shape analyses of BOSS data using the Effective Field Theory of Large-Scale Structure, we investigate what upcoming galaxy surveys might achieve. We introduce a “perturbativity prior” that ensures that loop terms are as large as theoretically expected, which is effective in the case of a large number of EFT parameters. After validating our technique by comparison with already-performed analyses of BOSS data, we provide Fisher forecasts using the one-loop prediction for power spectrum and bispectrum for two benchmark surveys: DESI and MegaMapper. We find overall great improvements on the cosmological parameters. In particular, we find that MegaMapper (DESI) should obtain at least a 12σ (2σ) evidence for non-vanishing neutrino masses, bound the curvature Ω k to 0.0012 (0.012), and primordial inflationary non-Gaussianities as follows: f NL loc. to ± 0.26 (3.3), f NL eq. to ±16 (92), f NL orth. to ± 4.2 (27). Such measurements would provide much insight on the theory of Inflation. We investigate the limiting factor of shot noise and ignorance of the EFT parameters.

cosmological parameters from LSS↗

Technical note: Deep learning for creating surrogate models of precipitation in Earth system models

Abstract. We investigate techniques for using deep neural networks to produce surrogatemodels for short-term climate forecasts. A convolutional neural network istrained on 97 years of monthly precipitation output from the 1pctCO2 run (theCO 2 concentration increases by 1 % per year) simulated by the second-generation Canadian Earth System Model (CanESM2). The neural network clearly outperforms a persistence forecast anddoes not show substantially degraded performance even when the forecast lengthis extended to 120 months. The model is prone to underpredicting precipitationin areas characterized by intense precipitation events. Scheduled sampling(forcing the model to gradually use its own past predictions rather than groundtruth) is essential for avoiding amplification of early forecasting errors.However, the use of scheduled sampling also necessitates preforecasting(generating forecasts prior to the first forecast date) to obtain adequateperformance for the first few prediction time steps. We document the trainingprocedures and hyperparameter optimization process for researchers who wish toextend the use of neural networks in developing surrogate models.

54 ENVIRONMENTAL SCIENCES↗

Detection of Anomalous Events in Electronic Health Records

Over the past decade, Health Information Technology (Health IT) has enabled an explosion in the amount of digital information stored in electronic health records (EHRs). According to recent studies, safety-related issues in healthcare can present themselves as anomalies in EHR data. Motivating examples of anomalous events in EHRs include clinical events related to invalid order cancellations or rejections, which may be initiated by clinical staff or automatic software routines in Health IT systems. Such events may be detected using anomaly detection or change point detection methods. In this paper, we explore the use of a forecasting approach to detect anomalies in EHR data using an online Support Vector Regression technique. Specifically, the proposed approach uses temporal frequency of activities in EHRs, coupled with dynamic robust confidence intervals, to characterize events as normal or anomalous. Once an event is characterized as an anomaly, our approach suppresses its effects in subsequent time intervals. The proposed approach shows encouraging results using real-world EHR data from the Veterans Affairs' corporate data warehouse

Pellett, Jordan J.↗

Forecasting Day-Ahead Solar Irradiance for Puerto Rico Using the WRF Model and NSRDB

Accurately predicting solar energy resources is a major challenge in integrating photovoltaics generation on the electric grid. Numerical weather prediction has been recognized by the solar energy community as a major approach to provide solar resource forecasts at various locations and for a variety of timescales. In this study, as a part of the Puerto Rico Grid Resilience and Transitions to 100% Renewable Energy Study (PR100), we develop day-head solar irradiance forecast data using the Weather Research and Forecasting (WRF) model at 3 km and hourly/5-minute. The global horizontal irradiance (GHI) and direct normal irradiance (DNI) forecasts simulated from the WRF model are postprocessed by a simple optimization method using satellite-derived gridded observations from the National Solar Radiation Data Base (NSRDB) to reduce error and bias of the solar irradiance forecasts covering 2018-2020. The NSRDB contributes to improving the GHI and DNI forecasts and also offers the opportunity for an in-depth analysis to evaluate their accuracy over a wide range of Puerto Rico regions. Preliminary results show overall improvements of GHI forecasts up to 37% (DNI: 15%) for mean absolute error and 97% (DNI: 76%) for mean bias error by applying a postprocessing technique to WRF model output.

data models↗

Situation awareness and dynamic ensemble forecasting of abnormal behavior in cyber-physical system

A plurality of monitoring nodes may each generate a time-series of current monitoring node values representing current operation of a cyber-physical system. A feature-based forecasting framework may receive the time-series of and generate a set of current feature vectors using feature discovery techniques. The feature behavior for each monitoring node may be characterized in the form of decision boundaries that separate normal and abnormal space based on operating data of the system. A set of ensemble state-space models may be constructed to represent feature evolution in the time-domain, wherein the forecasted outputs from the set of ensemble state-space models comprise anticipated time evolution of features. The framework may then obtain an overall features forecast through dynamic ensemble averaging and compare the overall features forecast to a threshold to generate an estimate associated with at least one feature vector crossing an associated decision boundary.

Abbaszadeh, Masoud↗

The WRF-Solar Ensemble Prediction System: Development, Test, and Validation

Providing reliable probabilistic solar radiation information is needed to improve management of the uncertainty and variability of solar generation. Thus, guidance on how to develop skillful and accurate ensemble forecasts is essential and it will ultimately contribute to integration of high amounts of solar energy on the grid. A team from the National Renewable Energy Laboratory and the National Center for Atmospheric Research had been collaborating to develop the WRF-Solar ensemble prediction system (WRF-Solar EPS) in the past three years to produce probabilistic solar irradiance forecasts and better predict solar energy by quantifying forecast uncertainty. The WRF-Solar EPS basically generates ensemble members for solar irradiance based on stochastic perturbations to provide intraday and day-ahead probabilistic forecasts. This study will present main research steps in developing the WRF-Solar EPS including: (a) tangent linear analysis for identifying key input variables of six WRF-Solar modules significantly related to predicting of cloud and solar irradiance, (b) combining stochastic perturbation technique with the WRF-Solar model, and (c) ensemble calibration method to decrease error and uncertainty of ensemble-based solar forecasts. The capability of WRF-Solar EPS is now updated to the most recent version of standard WRF model. This presentation will summarize comprehensive results from the evaluation of forecasts against the National Solar Radiation Data Base as well as ground-measured observations. Moreover, we will introduce the user's guide for WRF-Solar EPS (e.g., parameters to configure stochastic perturbations) and future extension of this research.

day-ahead forecast↗

Automated Framework for Groundwater Monitoring Using DWT with LSTM and Transformers

Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.

Discrete Wavelet Transform (DWT)↗