Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data driven model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Distribution Substation Planning Toolkit (dsp-toolkit) v1.0

The Distribution Substation Planning Toolkit (DSP Toolkit) is a software suite designed to streamline the planning and optimization of distribution substations. This toolkit offers a comprehensive set of tools and APIs for data curation, short-term electric load forecasting, and weather-sensitive load adjustment, making it an essential resource for utility companies, engineers, and researchers. Features • Data Preprocessing and Curation: Efficiently manage and preprocess large datasets to ensure high-quality input for analysis. • Short-Term Load Forecasting: Utilize data-driven models to predict short-term electric loads accurately. • Weather-Sensitive Modeling: Automatically adjust load forecasts based on weather data to predict future peak demands more precisely. Uses The DSP Toolkit is ideal for planning and optimizing distribution substations, providing a user-friendly interface and comprehensive documentation. It is suitable for both novice and experienced users, facilitating efficient and accurate planning processes. Advantages • Efficiency: Automates complex planning tasks, reducing manual effort and minimizing errors. • Scalability: Handles large datasets and complex models, making it suitable for large-scale projects. • Community and Support: Open-source with active community contributions, ensuring continuous improvement and support. • Extensibility: Easily extendable with custom modules and plugins, allowing users to tailor the toolkit to their specific needs. The DSP Toolkit stands out by offering a robust, flexible, and user-friendly solution for distribution substation planning. Public Abstract

Li, Han [Lawrence Berkeley National Laboratory (LB↗

Consistency Between Sun-Induced Chlorophyll Fluorescence and Gross Primary Production of Vegetation in North America

Accurate estimation of the gross primary production (GPP) of terrestrial ecosystems is vital for a better understanding of the spatial-temporal patterns of the global carbon cycle. In this study,we estimate GPP in North America (NA) using the satellite-based Vegetation Photosynthesis Model (VPM), MODIS (Moderate Resolution Imaging Spectrometer) images at 8-day temporal and 500 meter spatial resolutions, and NCEP-NARR (National Center for Environmental Prediction-North America Regional Reanalysis) climate data. The simulated GPP (GPP (sub VPM)) agrees well with the flux tower derived GPP (GPPEC) at 39 AmeriFlux sites (155 site-years). The GPP (sub VPM) in 2010 is spatially aggregated to 0.5 by 0.5-degree grid cells and then compared with sun-induced chlorophyll fluorescence (SIF) data from Global Ozone Monitoring Instrument 2 (GOME-2), which is directly related to vegetation photosynthesis. Spatial distribution and seasonal dynamics of GPP (sub VPM) and GOME-2 SIF show good consistency. At the biome scale, GPP (sub VPM) and SIF shows strong linear relationships (R (sup 2) is greater than 0.95) and small variations in regression slopes ((4.60-5.55 grams Carbon per square meter per day) divided by (milliwatts per square meter per nanometer per square radian)). The total annual GPP (sub VPM) in NA in 2010 is approximately 13.53 petagrams Carbon per year, which accounts for approximately 11.0 percent of the global terrestrial GPP and is within the range of annual GPP estimates from six other process-based and data-driven models (11.35-22.23 petagrams Carbon per year). Among the seven models, some models did not capture the spatial pattern of GOME-2 SIF data at annual scale, especially in Midwest cropland region. The results from this study demonstrate the reliable performance of VPM at the continental scale, and the potential of SIF data being used as a benchmark to compare with GPP models.

photosynthesis model↗

Empirical wind model for the middle and lower atmosphere. Part 1: Local time average

The HWM90 thermospheric wind model was revised in the lower thermosphere and extended into the mesosphere and lower atmosphere to provide a single analytic model for calculating zonal and meridional wind profiles representative of the climatological average for various geophysical conditions. Gradient winds from CIRA-86 plus rocket soundings, incoherent scatter radar, MF radar, and meteor radar provide the data base and are supplemented by previous data driven model summaries. Low-order spherical harmonics and Fourier series are used to describe the major variations throughout the atmosphere including latitude, annual, semiannual, and longitude (stationary wave 1). The model represents a smoothed compromise between the data sources. Although agreement between various data sources is generally good, some systematic differences are noted, particularly near the mesopause. Root mean square differences between data and model are on the order of 15 m/s in the mesosphere and 10 m/s in the stratosphere for zonal wind, and 10 m/s and 4 m/s, respectively, for meridional wind.

Hedin, A. E.↗

Can artificial intelligence and data-driven machine learning models match or even replace process-driven hydrologic models for streamflow simulation?: A case study of four watersheds with different hydro-climatic regions across the CONUS

With recent developments in computational techniques, Data-driven Machine Learning Models (DMLs) have shown great potential in simulating streamflow and capturing the rainfall-runoff relationship in given watersheds, which are traditionally fulfilled by Process-based Hydrologic Models (PHMs). There are debates on whether the DMLs can outperform and possibly replace the classical PHMs for streamflow simulation and river forecasting, but no clear conclusions have been made. This study aims to investigate whether the newer DMLs have any potential in further improving the simulation accuracy of classical PHMs, and vice versa. To do this, we compared a few popular PHMs and DMLs over four watersheds across the Continental US (CONUS) that are associated with different input, climate, and regional conditions. A total of five hydrologic models were chosen, including (1) two classical lumped models, i.e., the Sacramento Soil Moisture Accounting (SAC-SMA) and Xinanjiang (XAJ); (2) one modern distributed model, termed Coupled Routing and Excess Storage (CREST); (3) and two DMLs including an Artificial Neural Networks (ANN) and a deep learning model, termed Long Short Term Memory (LSTM). Our results demonstrated that the DMLs still significantly biased when using the baseline input scenario with the PHMs. However, the DMLs fed with delayed input scenarios had great potential and can reach high simulation accuracy. The DMLs, especially the ANN, outperformed other employed models under the rainfall-runoff relationship in which rainfall dominantly drives. Furthermore, the DMLs also showed better performance in the high-flow regime, while the PHMs had a better performance for the low-flow regime, implying both PHMs and DMLs have their own merits and are worthy of joint development. In general, our study indicated a great potential of using DMLs to simulate streamflow, but further studies are still needed to verify the transferability and scalability of DMLs in large-scale experiments, such as the Distributed Model Intercomparison Projects 1&2 conducted by National Weather Services but to compare modern DMLs and PHMs.

58 GEOSCIENCES↗

Physics-Informed Learning Machines for Multiscale and Multiphysics Problems (PHILMS) (Technical Report)

The research work at University of California Santa Barbara (UCSB) resulted in several new developments in the areas of scientific machine learning, numerical analysis, and practical methods for data-driven modeling, prediction, reductions, and simulation. Many of the projects were carried out in collaboration with members of the national laboratories at Sandia National Laboratories (SNL), Pacific Northwestern National Laboratories (PNNL), and other institutions. Results included developing new scientific machine learning methods, related theory and mathematical frameworks for analysis and training, data-driven numerical solvers, and related tools and software for scientific computation. During the support period, over 16+ papers were submitted for publication, and 4 open-source software packages were developed and released (available at http://atzberger.org/). In addition, 7+ students and 2 post-docs were mentored in collaboration with the laboratory staff for future careers in academia, government labs, and industry.

97 MATHEMATICS AND COMPUTING↗

Explainable discrepancy checker and diagnosis for digital Twin-based supervisory control system

By virtually representing a physical object and process, a digital twin (DT) enables optimal autonomous operations by combining classical and novel frameworks in sensors, state predictions, and multi-input/multi-output systems. A DT’s values depend on how well models estimate quantities of interest and on how uncertainty is handled. Moreover, DTs often combine physics-based and data-driven models with mixed fidelities, where classical uncertainty quantification (UQ) struggles with many sources of uncertainty and real-time constraints. Here, this work presents a UQ-based discrepancy checking and diagnosis tool for a DT-based supervisory control system. The tool is developed using metadata from an automated DT development process to learn correlations between sources of uncertainties and outcomes. During operation, it compares predictions with measurements, attributes discrepancies to dominant sources, and recommends parameter and configuration updates. We verify the workflow on a synthetic temperature-control problem and deploy it on a virtual Thermal Energy Delivery System, reducing mismatch and improving control robustness.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

The Future of Sensitivity Analysis: An essential discipline for systems modeling and policy support

Sensitivity analysis (SA) is en route to becoming an integral part of mathematical modeling. The tremendous potential benefits of SA are, however, yet to be fully realized, both for advancing mechanistic and data-driven modeling of human and natural systems, and in support of decision making. In this perspective paper, a multidisciplinary group of researchers and practitioners revisit the current status of SA, and outline research challenges in regard to both theoretical frameworks and their applications to solve real-world problems. Six areas are discussed that warrant further attention, including (1) structuring and standardizing SA as a discipline, (2) realizing the untapped potential of SA for systems modeling, (3) addressing the computational burden of SA, (4) progressing SA in the context of machine learning, (5) clarifying the relationship and role of SA to uncertainty quantification, and (6) evolving the use of SA in support of decision making. An outlook for the future of SA is provided that underlines how SA must underpin a wide variety of activities to better serve science and society.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS-GUI: a web application for global survey of surface water metabolites

Background The Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) is a consortium that aims to understand complex hydrologic, biogeochemical, and microbial connections within river corridors experiencing perturbations such as dam operations, floods, and droughts. For one ongoing WHONDRS sampling campaign, surface water metabolite and microbiome samples are collected through a global survey to generate knowledge across diverse river corridors. Metabolomics analysis and a suite of geochemical analyses have been performed for collected samples through the Environmental Molecular Sciences Laboratory (EMSL). The obtained knowledge and data package inform mechanistic and data-driven models to enhance predictions of outcomes of hydrologic perturbations and watershed function, one of the most critical components in model-data integration. To support efforts of the multi-domain integration and make the ever-growing data package more accessible for researchers across the world, a Shiny/R Graphical User Interface (GUI) called WHONDRS-GUI was created. Results The web application can be run on any modern web browser without any programming or operational system requirements, thus providing an open, well-structured, discoverable dataset for WHONDRS. Together with a context-aware dynamic user interface, the WHONDRS-GUI has functionality for searching, compiling, integrating, visualizing and exporting different data types that can easily be used by the community. The web application and data package are available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1484811 , which enables users to simultaneously obtain access to the data and code and to subsequently run the web app locally. The WHONDRS-GUI is also available for online use at Shiny Server ( https://xmlin.shinyapps.io/whondrs/ ).

59 BASIC BIOLOGICAL SCIENCES↗

Resilient Observer Design for Cyber-Physical Systems with Data-Driven Measurement Pruning

Resilient observer design for Cyber-Physical Systems (CPS) in the presence of adversarial false data injection attacks (FDIA) is an active area of research. The existing state-of-the-art algorithms tend to break down as more and more knowledge of the system is built into the attack model; also as the percentage of attacked nodes increases. From the view of optimization theory, the problem is often cast as a classical error correction problem for which a theoretical limit of has been established as the maximum percentage attacked nodes for which state recovery is guaranteed. Beyond this limit, the performance of -minimization based schemes, for instance, deteriorates rapidly. Similar performance degradation occurs for other types of resilient observers beyond certain percentages of attacked nodes. In order to increase the corresponding percentage of attacked nodes for which state recoveries can be guaranteed, researchers have begun to incorporate prior information into the underlying resilient observer design framework. For the most pragmatic cases, this prior information is often obtained through a data-driven machine learning process. Existing results have shown a strong positive correlation between the maximum attacked percentages that can be tolerated and the accuracy of the data-driven model. Motivated by these results, this chapter examines the case for pruning algorithms designed to improve the Positive Prediction Value (PPV) of the resulting prior information, given stochastic uncertainty characteristics of the underlying machine learning model. Theoretical quantification of the achievable improvement is given. Simulation results show that the pruning algorithm significantly increases the maximum correctable percentage of attacked nodes, even for machine learning model whose prediction power is comparable to the random flip of a coin.

Resilient Observer, Cyber-physical Systems, Data-D↗

A multiscale recurrent neural network model for predicting energy production from geothermal reservoirs

Optimization of energy production from geothermal reservoirs requires reliable prediction of energy production performance under alternative operation and development scenarios. Traditionally, reservoir simulation models are used for the evaluation and screening of alternative production and development plans. However, simulation models require extensive data collection and modeling efforts and are time-consuming to build, run, and update. Data-driven predictive models, on the other hand, can serve as efficient prediction tools that can be used for decision support and management of daily operations and surveillance activities. Data-driven models become particularly attractive when a reservoir simulation model for a field does not exist and/or is difficult to build. Machine learning (ML)-based data-driven models that have recently become popular in several fields exploit statistical patterns and relations in training data to generate predictions. As such, they tend to perform better in interpolation problems (that is, prediction within the training data range) than when they are used to extrapolate beyond the training data. Production data from geothermal reservoirs tend to exhibit short-term variabilities as well as long-term trends, such as monotonically declining production temperatures. Capturing both short-term features and long-term trends with ML-based models is not trivial. We evaluate the use of recurrent neural networks (RNN) for the prediction of energy production from geothermal reservoirs. RNN is a class of ML architectures that are used to represent and predict sequential/dynamic data. Thus, it can be challenging to apply RNN to problems where long-term trends must be captured and extrapolation beyond the training data range is needed. We introduce the multiscale RNN architecture to extend the application of RNN to detect and predict both short-term variabilities and long-term trends in geothermal data. The developed architecture consists of a long-term component to only capture low-frequency data patterns, and a short-term component to detect features with higher frequency and more nonlinearity. The final prediction is obtained by combining the long-term and short-term predictions. Both synthetic and field data are used to evaluate the presented multiscale RNN model. The prediction performance of the multiscale RNN is compared against those obtained from the regular RNN and the autoregressive (AR) model. The results suggest that the multiscale architecture improves the long-term prediction performance of the regular RNN and enhances its robustness against noise.

15 GEOTHERMAL ENERGY↗

Data-Driven RANS Turbulence Closures for Forced Convection Flow in Reactor Downcomer Geometry

Recent progress in data-driven turbulence modeling has shown its potential to enhance or replace traditional equation-based Reynolds-averaged Navier-Stokes (RANS) turbulence models. Here, this work utilizes invariant neural network (NN) architectures to model Reynolds stresses and turbulent heat fluxes in forced convection flows (when the models can be decoupled). As the considered flow is statistically one dimensional, the invariant NN architecture for the Reynolds stress model reduces to the linear eddy viscosity model. To develop the data-driven models, direct numerical and RANS simulations in vertical planar channel geometry mimicking a part of the reactor downcomer are performed. Different conditions and fluids relevant to advanced reactors (sodium, lead, unitary-Prandtl-number fluid, and molten salt) constitute the training database. The models enabled accurate predictions of velocity and temperature, and compared to the baseline k–τ turbulence model with the simple gradient diffusion hypothesis, do not require tuning of the turbulent Prandtl number. The data-driven framework is implemented in the open-source graphics processing unit–accelerated spectral element solver nekRS and has shown the potential for future developments and consideration of more complex mixed convection flows.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Water resource recovery modelling 2021 (WRRmod2021 conference)

Our society is transitioning fast into the digital age, spurred by development of cheap and new sensing technology, breakthroughs in computing, and development of efficient algorithms for optimization. This transition is also visible in the field of wastewater treatment and is driving new model developments, especially by exploiting the large data sets available with many utilities. Not surprisingly, WRRmod2021 had featured a strong session on ‘data-driven models and digitalization’ focused on this hot topic. At the same time, engineering practice calls for more robust models for performance evaluation and optimization of both conventional facilities and innovative processes. As a result, the WRRmod2021 program also exhibited sessions on modelling of new process units (e.g., aerobic granular sludge), modelling of the nitrogen cycle, and integrated/plant-wide modelling.

54 ENVIRONMENTAL SCIENCES↗

Multimodal Data Fusion via Entropy Minimization

The use of gradient-based data-driven models to solve a range of real-world remote sensing problems can in practice be limited by the uniformity of available data. Use of data from disparate sensor types, resolutions, and qualities typically requires compromises based on assumptions that are made prior to model training and may not necessarily be optimal given over-arching objectives. For example, while deep neural networks (NNs) are state-of-the-art in a variety of target detection problems, training them typically requires either limiting the training data to a subset over which uniformity can be enforced or training independent models which subsequently require additional score fusion. The method we introduce here seeks to leverage the benefits of both approaches by allowing correlated inputs from different data sources to co-influence preferred model solutions, while maintaining flexibility over missing and mismatching data. In this work we propose a new data fusion technique for gradient updated models based on entropy minimization and experimentally validate it on a hyperspectral target detection dataset. We demonstrate superior performance compared to currently available techniques using a range of realistic data scenarios, where available data has limited spacial overlap and resolution.

97 MATHEMATICS AND COMPUTING↗

Robotic Planning under Uncertainty in Spatiotemporal Environments in Expeditionary Science

In the expeditionary sciences, spatiotemporally varying environments -- hydrothermal plumes, algal blooms, lava flows, or animal migrations -- are ubiquitous. Mobile robots are uniquely well-suited to study these dynamic, mesoscale natural environments. We formalize expeditionary science as a sequential decision-making problem, modeled using the language of partially-observable Markov decision processes (POMDPs). Solving the expeditionary science POMDP under real-world constraints requires efficient probabilistic modeling and decision-making in problems with complex dynamics and observational models. Previous work in informative path planning, adaptive sampling, and experimental design have shown compelling results, largely in static environments, using data-driven models and information-based rewards. However, these methodologies do not trivially extend to expeditionary science in spatiotemporal environments: they generally do not make use of scientific knowledge such as equations of state dynamics, they focus on information gathering as opposed to scientific task execution, and they make use of decision-making approaches that scale poorly to large, continuous problems with long planning horizons and real-time operational constraints. In this work, we discuss these and other challenges related to probabilistic modeling and decision-making in expeditionary science, and present some of our preliminary work that addresses these gaps. We ground our results in a real expeditionary science deployment of an autonomous underwater vehicle (AUV) in the deep ocean for hydrothermal vent discovery and characterization. Our concluding thoughts highlight remaining work to be done, and the challenges that merit consideration by the reinforcement learning and decision-making community.

Preston, Victoria↗

Data-Driven Radiative Hydrodynamic Modeling of the 2014 March 29 X1.0 Solar Flare

Spectroscopic observations of solar flares provide critical diagnostics of the physical conditions in the flaring atmosphere. Some key features in observed spectra have not yet been accounted for in existing flare models. Here we report a data-driven simulation of the well-observed X1.0 flare on 2014 March 29 that can reconcile some well-known spectral discrepancies. We analyzed spectra of the flaring region from the Interface Region Imaging Spectrograph (IRIS) in Mg II hk, the Interferometric BIdimensional Spectropolarimeter at the Dunn Solar Telescope (DSTIBIS) in H(alpha) 6563A and Ca II 8542A, and the Reuven Ramaty High Energy Solar Spectroscope Imager (RHESSI) in hard X-rays. We constructed a multithreaded flare loop model and used the electron flux inferred from RHESSI data as the input to the radiative hydrodynamic code RADYN to simulate the atmospheric response. We then synthesized various chromospheric emission lines and compared them with the IRIS and IBIS observations. In general, the synthetic intensities agree with the observed ones, especially near the northern footpoint of the flare. The simulated Mg II line profile has narrower wings than the observed one. This discrepancy can be reduced by using a higher microturbulent velocity (27 km/s) in a narrow chromospheric layer. In addition, we found that an increase of electron density in the upper chromosphere within a narrow height range of approx. 800 km below the transition region can turn the simulated Mg II line core into emission and thus reproduce the single peaked profile, which is a common feature in all IRIS flares.

Da Costa, Fatima Rubio↗

Neural Network‐Based Methods for Ocean Surface Wave Measurement Using Submarine Distributed Acoustic Sensing (DAS)

Two new data-driven models for estimating ocean surface waves from distributed acoustic sensing (DAS) submarine cable strain rate are developed using supervised machine learning on a 10-day data set collected offshore of Oliktok Point, Alaska. The new models were trained on target data from seafloor pressure moorings at three sites spaced evenly along 27.1 km of cable and were benchmarked against an empirical transfer function method previously used to estimate waves from DAS. A model which uses convolutional neural networks to transform 2-km frequency-wavenumber strain spectra to seafloor pressure spectra outperforms the benchmark in wave height prediction (RMSE of 0.15 vs. 0.41 m) and period prediction (0.29 vs. 0.37 s) when evaluated on a held-out test data set. When applied to a DAS data set collected on the same cable 2 years prior, the CNN-based model maintained similar significant wave height performance (RMSE = 0.23 m) relative to available satellite altimetry data. A two-hidden-layer, fully connected neural network which transforms 1-D strain spectra to seafloor pressure spectra also outperforms the benchmark in wave height prediction (RMSE of 0.19 vs. 0.41 m), but does not generalize as well to the prior data. Regression-based machine learning is useful for estimating waves from DAS data when the pressure-strain relationship varies temporally and spatially across different wave conditions. Models can be applied to DAS data to measure waves with higher spatial resolution and longer temporal coverage than traditional methods, which often measure waves only at a single point.

Davis, Jacob R. [Univ. of Washington, Seattle, WA ↗

Quantification of neural networks uncertainties with applications to SAFARI-1 axial neutron flux profiles

Deep Neural Networks (DNNs) have been widely used as a data-driven modelling tool in nuclear engineering. However, as a Machine Learning model, Artificial Neural Network (ANN) predictions are subjected to uncertainties originating from the noise in training data, incomplete coverage of the domain, and imperfect neural network architectures. In this work, we target at quantifying the prediction/approximation uncertainties of ANNs using Monte Carlo Dropout (MCD), as well as Bayesian Neural Networks (BNNs) which are solved by variational inference. With a demonstration problem in which neural networks are used to predict the assembly axial neutron flux profiles, the results have shown that the three different neural network models (regular DNNs, DNNs solved with MCD and BNNs) can produce results that agree very well among each other and with the measurement data, on cycles that are not used in the training process. Besides the excellent generalization capability, the uncertainty bands produced by MCD and BNN agree very well, and in general, they can fully envelop the noisy measurement data points. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Statistical analysis and degradation pathway modeling of photovoltaic minimodules with varied packaging strategies

Degradation pathway models constructed using network structural equation modeling (netSEM) are used to study degradation modes and pathways active in photovoltaic (PV) system variants in exposure conditions of high humidity and temperature. This data-driven modeling technique enables the exploration of simultaneous pairwise and multiple regression relationships between variables in which several degradation modes are active in specific variants and exposure conditions. Durable and degrading variants are identified from the netSEM degradation mechanisms and pathways, along with potential ways to mitigate these pathways. A combination of domain knowledge and netSEM modeling shows that corrosion is the primary cause of the power loss in these glass/backsheet PV minimodules. We show successful implementation of netSEM to elucidate the relationships between variables in PV systems and predict a specific service lifetime. The results from pairwise relationships and multiple regression show consistency. This work presents a greater opportunity to be expanded to other materials systems.

electrical measurements↗