Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance Modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Exploring Flood Predictability in Taiwan through Coupled Atmospheric–Hydrological and High-Performance Hydrodynamic Models

Effective flood simulation capabilities can tremendously support early warning and disaster prevention. To examine the applicability of a fully physics-based and high-performance flood simulation and forecasting modeling framework for a flood-prone region in Taiwan, we conduct a numerical experiment that couples the Weather Research and Forecasting (WRF) Model, WRF-Hydrological modeling system (WRF-Hydro), and the Two-Dimensional Runoff Inundation Toolkit for Operational Needs (TRITON) to perform integrated rainfall, streamflow, and flood simulations. Furthermore, we first use the coupled WRF and WRF-Hydro (WWH) to predict rainfall and streamflow and then drive TRITON with the predicted streamflow hydrographs to simulate flood depth and inundation area. With the refined spatial resolution and parameterization, this framework can better predict rainfall with reasonable spatial patterns. Although WWH could overestimate the amount of rainfall in some areas, the uncertain rainfall–streamflow predictions produce reasonable flood maps able to pinpoint regions at risk of flooding. In terms of model efficiency, the graphics processing unit–based computation can yield a speed-up factor as high as ∼13 compared to the central processing unit–based computation, promoting the efficacy of the coupled modeling framework in practical real-time flood forecasting.

Coupled models↗

Comparison of linear regression, k-nearest neighbour and random forest methods in airborne laser-scanning-based prediction of growing stock

Abstract In this study, for five sites around the world, we look at the effects of different model types and variable selection approaches on forest yield modelling performances in an area-based approach (ABA). We compared ordinary least squares regression (OLS), k-nearest neighbours (kNN) and random forest (RF). Our objective was to test if there are systematic differences in accuracy between OLS, kNN and RF in ABA predictions of growing stock volume. The analyses are based on a 5-fold cross-validation at five study sites: an eucalyptus plantation, a temperate forest and three different boreal forests. Two completely independent validation datasets were also available for two of the boreal sites. For the kNN, we evaluated multiple measures of distance including Euclidean, Mahalanobis, most similar neighbour (MSN) and an RF-based distance metric. The variable selection approaches we examined included a heuristic approach (for OLS, kNN and RF), exhaustive search among all combinations (OLS only) and all variables together (RF only). Performances varied by model type and variable selection approaches among sites. OLS and RF had similar accuracies and were more efficient than any of the kNN variants. Variable selection did not affect RF performance. Heuristic and exhaustive variable selection performed similarly for OLS. kNN fared the poorest amongst model types, and kNN with RF distance was prone to overfitting when compared with a validation dataset. Additional caution is therefore required when building kNN models for volume prediction though ABA, being preferable instead to opt for models based on OLS with some variable selection, or RF with all variables together.

Cosenza, Diogo N.↗

At Risk Population Estimates for Belarus, Poland and Slovakia with Machine Learning

High-resolution gridded population modeling is crucial for various applications, including disaster response planning, infectious disease spread modeling, climate change impact estimation, policy development, and more. Multiple gridded population datasets have been developed, each tailored to meet specific objectives. Among them, LandScan Global dataset is designed to represent ambient and unwarned population distributions. However, this dataset relies on a statistical approach that requires manual adjustments, making it time consuming and labour intensive. Existing machine learning (ML) methods often train and test at different spatial resolutions, potentially leading to inflated results, and they rely on Census population totals for disaggregation. To address these limitations, in this study we developed population estimates using ML models trained and tested at a consistent 30 arc-second resolution (≈1 square kilometer), specifically using Random Forest (RF) and XGBoost. These models were trained on 2020 datum to predict for 2021 for three countries: Belarus, Poland, and Slovakia. Our findings show that both RF (MAE varies from 5.75 to 13.25) and XGBoost (MAE varies from 8.15 to 23.44) model performance is close to LandScan Global estimates. Furthermore, neither of the models performed the best across all grid cells: the RF model was more effective in areas with lower populations, while XGBoost excelled in more densely populated regions. The proposed approach can be used for countries where the Census data is not available.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

Classification of Nuclear Reactor Operations Using Spatial Importance and Multisensor Networks

Distributed multisensor networks record multiple data streams that can be used as inputs to machine learning models designed to classify operations relevant to proliferation at nuclear reactors. The goal of this work is to demonstrate methods to assess the importance of each node (a single multisensor) and region (a group of proximate multisensors) to machine learning model performance in a reactor monitoring scenario. This, in turn, provides insight into model behavior, a critical requirement of data-driven applications in nuclear security. Using data collected at the High Flux Isotope Reactor at Oak Ridge National Laboratory via a network of Merlyn multisensors, two different models were trained to classify the reactor’s operational state: a hidden Markov model (HMM), which is simpler and more transparent, and a feed-forward neural network, which is less inherently interpretable. Traditional wrapper methods for feature importance were extended to identify nodes and regions in the multisensor network with strong positive and negative impacts on the classification problem. These spatial-importance algorithms were evaluated on the two different classifiers. The classification accuracy was then improved relative to baseline models via feature selection from 0.583 to 0.839 and from 0.811 ± 0.005 to 0.884 ± 0.004 for the HMM and feed-forward neural network, respectively. While some differences in node and region importance were observed when using different classifiers and wrapper methods, the nodes near the facility’s cooling tower were consistently identified as important—a conclusion further supported by studies on feature importance in decision trees. Node and region importance methods are model-agnostic, inform feature selection for improved model performance, and can provide insight into opaque classification models in the nuclear security domain.

Tibbetts, Jake↗

Rapid measurement of soluble xylo-oligomers using near-infrared spectroscopy (NIRS) and multivariate statistics: calibration model development and practical approaches to model optimization

Rapid monitoring of biomass conversion processes using techniques such as near-infrared (NIR) spectroscopy can be substantially quicker and less labor-, resource-, and energy-intensive than conventional measurement techniques such as gas or liquid chromatography (GC or LC) due to the lack of solvents and preparation methods, as well as removing the need to transfer samples to an external lab for analytical evaluation. The purpose of this study was to determine the feasibility of rapid monitoring of a biomass conversion process using NIR spectroscopy combined with multivariate statistical modeling, and to examine the impact of (1) subsetting the samples in the original dataset by process location and (2) reducing the spectral range used in the calibration model on model performance. We develop multivariate calibration models for the concentrations of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids at multiple points in a biomass conversion process which produces and then purifies XOS compounds from sugar cane bagasse. A single model using samples from multiple locations in the process stream showed acceptable performance as measured by standard statistical measures. However, compared to the single model, we show that separate models built by segregating the calibration samples according to process location show improved performance. We also show that combining an understanding of the sample spectra with simple multivariate analysis tools can result in a calibration model with a substantially smaller spectral range that provides essentially equal performance to the full-range model. We demonstrate that real-time monitoring of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids concentration at multiple points in a process stream using NIR spectroscopy coupled with multivariate statistics is feasible. Segregation of sample populations by process location improves model performance. Models using a reduced spectral range containing the most relevant spectral signatures show very similar performance to the full-range model, reinforcing the importance of performing robust exploratory data analysis before beginning multivariate modeling.

09 BIOMASS FUELS↗

Assessing mechanisms for microbial taxa and community dynamics using process models

Disentangling the assembly mechanisms controlling community composition, structure, distribution, functions, and dynamics is a central issue in ecology. Although various approaches have been proposed to examine community assembly mechanisms, quantitative characterization is challenging, particularly in microbial ecology. Here, we present a novel approach for quantitatively delineating community assembly mechanisms by combining the consumer–resource model with a neutral model in stochastic differential equations. Using time-series data from anaerobic bioreactors that target microbial 16S rRNA genes, we tested the applicability of three ecological models: the consumer–resource model, the neutral model, and the combined model. Our results revealed that model performances varied substantially as a function of population abundance and/or process conditions. The combined model performed best for abundant taxa in the treatment bioreactors where process conditions were manipulated. In contrast, the neutral model showed the best performance for rare taxa. Our analysis further indicated that immigration rates decreased with taxa abundance and competitions between taxa were strongly correlated with phylogeny, but within a certain phylogenetic distance only. The determinism underlying taxa and community dynamics were quantitatively assessed, showing greater determinism in the treatment bioreactors that aligned with the subsequent abnormal system functioning. Given its mechanistic basis, the framework developed here is expected to be potentially applicable beyond microbial ecology.

59 BASIC BIOLOGICAL SCIENCES↗

2021 Q1 Project Report: Optimized Bifacial PV Systems (Q1 FY2021 Project Report)

This project has four main technical objectives: 1) Develop and improve bifacial performance models by adding the capability to evaluate electrical behavior and performance of bifacial modules and arrays under realistic field conditions including irradiance variability caused by racking, module frame, and position in the array. 2) Instrument and monitor performance of fielded bifacial systems to validate performance models and to measure, analyze and publish on bifacial energy gain. These should include both research and commercial bifacial systems and cover a variety of deployment applications. 3) Evaluate optimal bifacial system designs using simulations leveraging high-performance computing, and also using full sized and miniaturized experimental field deployments. 4) Establish and contribute to international test standards for bifacial system performance, testing, and safety, and work with the community to establish installation and siting best practices.

42 ENGINEERING↗

Uncertainty Aware Deep Learning for Fault Prediction Using Multivariate Time Series Signals

The superconducting radio-frequency cavities are a crucial component of the Continuous Electron Beam Accelerator Facility (CEBAF) at Jefferson Lab. When a cavity faults, beam delivery to experimental end users is disrupted. Prediction of cavity faults prior to onset is essential to reduce operation and maintenance costs. In this work, a parallel long short-term memory (LSTM)-convolution neural network (CNN)-based deep learning (DL) model is proposed to predict impending faults using pre-fault signals. Further, we introduce an uncertainty quantification approach using Monte Carlo dropout with the LSTM-CNN model to ascertain confidence in the prediction. The model was tested using multivariate time series signals from stable cavity operations and before faults. Initial results show that on the test dataset, the model can identify impending faults before their onset with an average 10-fold cross validation accuracy of 97.39% and a standard deviation of 0.12% using a 100-ms time window. It is also observed that the model performs better as the prediction time moves closer to the fault onset. For additional context, we compare the performance of the model with three machine-learning-based (ML) fault prediction models. Our proposed parallel LSTM-CNN-based DL method shows better performance than the ML-based methods.

Rahman, Md Monibor↗

Chemical signature characterization with hyperspectral imagery: novel deep learning model architectures and physically-motivated data augmentation techniques

The high spectral resolution afforded by Hyperspectral Imaging (HSI) sensors is poised to bring unprecedented advancements to signature characterization applications. Thus far, much of the research in the machine learning field devoted to HSI applications has focused on a few specific tasks like land-use land-cover classification. In land classification tasks, spatial information is very important, and model architectures are often designed to leverage spatial contexts. However, it is unclear how well these spatially-tuned models will translate to tasks where spectral information is critical, like the detection and characterization of chemicals. In this work, we compare spectral models (inputs are 1D spectra) and spatial-spectral models (inputs are 3D cubes) in the context of predicting chemical concentration maps. We find that spatial-spectral models perform the best, though we find a wide range in performance across the different architectures tested. Additionally, we find that model performance is impacted by the availability of training data, particularly in scenarios where the training data doesn't fully capture the true variance of real-world conditions. We find that data augmentation can help mitigate sparse coverage of observed parameter space (e.g., seasonal or geographic variability in ground cover), and present augmentation strategies that are tailored to hyperspectral data.

• Artificial intelligence (AI) / machine learning ↗

Comparison of temporal resolution selection approaches in energy systems models

Capacity expansion models for the power sector are used to project future decisions over the coming decades by simulating investment and operation decisions for the use of electricity. Due to model performance constraints, these models typically do not explicitly simulate every hour within a year, but instead simulate representative time segments (groups of hours). This paper evaluates different approaches for selecting time segments across three methods: sequential, categorical, and clustering, across a wide range of time-segment quantities, for a total of 204 temporal profiles. To measure the performance of each profile's ability to accurately represent data, the root-mean-square-error of each profile's time segments are compared to the data's original hourly data. The temporal alignment across regions is also measured (i.e., how often windy days align across regions). Different spatial resolutions were applied for a subset of the temporal selection methods to investigate the impact spatial resolution has on performance. This paper provides a framework for measuring the value of different temporal selection methods and of adding more granular data to energy system models. Overall, multi-criteria clustering yields the lowest root-mean-square-error across all datasets evaluated and provides a holistic view of the intertwined relationships between renewable generation and electricity demand.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Location generalizability of image-based air quality models

This paper is to be submitted at the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Computer Vision for Earth Observation workshop. The full paper abstract is below: The ability to rapidly quantify atmospheric pollutants is important both for global emissions monitoring and for mitigating the adverse effects that follow a hazardous chemical release. In the aftermath of a chemical release, imagery is often the only available resource to assess local conditions. Recent work has demonstrated initial success in predicting particulate matter pollution from imagery; however, these results are tied to a specific site and do not generalize to new geographic locations. In this work, we seek to understand how easily deep learning models generalize to new locations in the context of image-based air quality assessments, targeting two distinct tasks: (1) broad measures of particulate matter pollution, and (2) the mass of a given chemical released in hazardous plumes. For the latter, we focus on sulfur dioxide, a toxic aerosol and a major component of particulate matter pollution caused by industrial fossil fuel consumption. To develop a model that operates in the widest possible range of environments, we test different training strategies, including the use of new geolocation foundation models. The best performing models achieve >80% accuracy when evaluating unseen imagery at previously seen sites, but we find significant drops in performance when evaluating imagery from unseen sites, at best 65%. Additionally, we present the public release of the National Parks Air Quality Index Dataset, a new medium-sized dataset that pairs imagery with sensor-based air quality measurements at 15 different national parks.

Byler, Eleanor B. [BATTELLE (PACIFIC NW LAB)]↗

Analysis of Photovoltaic Energy: CRADA Number CRD-21-17862 (Cooperative Research and Development Final Report)

For this project, NREL will work with the Participant to model phase change material performance, model photovoltaic thermodynamics performance, and predict photovoltaic module durability based on environmental conditions to help develop optimized photovoltaic roofing designs. This project will develop thermodynamic models of integrated photovoltaic roofing with phase change materials. These models will be used to optimize the materials and design selections based on reducing photovoltaic operating temperatures while balancing cost effective integration of phase change materials.

14 SOLAR ENERGY↗

Forte: An Interactive Visual Analytic Tool for Trust-Augmented Net-Load Forecasting

Accurate net-load forecasting is vital for energy planning, aiding decisions on trade and load distribution. However, assessing the performance of forecasting models across diverse input variables, like temperature and humidity, remains challenging, particularly for eliciting a high degree of trust in the model outcomes. In this context, there is a growing need for data-driven technological interventions to aid scientists in comprehending how models react to both noisy and clean input variables, thus shedding light on complex behaviors and fostering confidence in the outcomes. In this paper, we present Forte, a visual analytics-based application to explore deep probabilistic net-load forecasting models across various input variables and understand the error rates for different scenarios. With carefully designed visual interventions, this web-based interface empowers scientists to derive insights about model performance by simulating diverse scenarios, facilitating an informed decision-making process. We discuss observations made using Forte and demonstrate the effectiveness of visualization techniques to provide valuable insights into the correlation between weather inputs and net-load forecasts, ultimately advancing grid capabilities by improving trust in forecasting models.

Bhattacharjee, Kaustav↗

Quantifying Uncertainty in PV Energy Estimates Final Report

Uncertainty in PV energy estimates is "one of the most critical areas of lack of understanding" according to independent engineers, financiers, PV model developers, and other industry stakeholders. The primary problem is a lack of rigorous, transparent, widely accepted methods for quantifying uncertainty in energy production estimates. Uncertainty in energy production estimates arises from variability of the solar resource, inexact PV performance models and their parameters, and system reliability considerations. Uncertainty in annual energy production is frequently calculated for larger projects in order to quantify financial risk. Key statistics for energy, such as the P-values "P50" and "P90" (the annual energy values that are exceeded in future years with 50\% and 90\% probability, respectively) are used by financing institutions to calculate the repayment risk for the project. The current methods to estimate these statistics are typically proprietary, specialized, and involve significant post-processing of commercial performance model results. This black-box approach leads to inconsistent P-value estimates from different parties, which reduces investors' confidence in the results. Since the financial community bases its risk assessment on these estimates, reduced confidence increases perceived project risk, and consequently financing costs. The goal of this project was to establish a set of best practices for quantifying uncertainty in energy production estimates, including identifying what sources of uncertainty must be considered with clear definitions and metrics, determining which sources are the biggest drivers of uncertainty, and providing a computationally efficient framework for combining different sources of uncertainty that is flexible enough to accommodate substitutions of data or methods when better information is available. We engaged a wide set of stakeholders to ensure industry endorsement and adoption, and leveraged complementary projects investigating individual sources of uncertainty in great detail, as well as others' work that started down this path.

ENERGY PLANNING, POLICY, AND ECONOMY,SOLAR ENERGY↗

Causality guided machine learning model on wetland CH 4 emissions across global wetlands

Wetland CH 4 emissions are among the most uncertain components of the global CH 4 budget. The complex nature of wetland CH 4 processes makes it challenging to identify causal relationships for improving our understanding and predictability of CH 4 emissions. In this study, we used the flux measurements of CH 4 from eddy covariance towers (30 sites from 4 wetlands types: bog, fen, marsh, and wet tundra) to construct a causality-constrained machine learning (ML) framework to explain the regulative factors and to capture CH 4 emissions at sub-seasonal scale. We found that soil temperature is the dominant factor for CH 4 emissions in all studied wetland types. Ecosystem respiration (CO 2 ) and gross primary productivity exert controls at bog, fen, and marsh sites with lagged responses of days to weeks. Integrating these asynchronous environmental and biological causal relationships in predictive models significantly improved model performance. More importantly, modeled CH 4 emissions differed by up to a factor of 4 under a +1°C warming scenario when causality constraints were considered. These results highlight the significant role of causality in modeling wetland CH 4 emissions especially under future warming conditions, while traditional data-driven ML models may reproduce observations for the wrong reasons. Our proposed causality-guided model could benefit predictive modeling, large-scale upscaling, data gap-filling, and surrogate modeling of wetland CH 4 emissions within earth system land models.

54 ENVIRONMENTAL SCIENCES↗

Comparative investigations of multi-fidelity modeling on performance of electrostatically-actuated cracked micro-beams

Silicon is a commonly used material for the fabrication of beams for use in micro-electrical-mechanical systems (MEMS). Although silicon is a brittle material, it has been shown to accumulate fatigue damage at the micro-scale. Understanding the effect this has on the overall device performance is critical to the design of reliable devices. Analytical methods for modeling damage provide expedient results but are limited by broad modeling assumptions. Numerical models account for more detailed physical phenomena but can be computationally intensive. In this work, two different crack scenarios are modeled using both analytical techniques and 3D computational simulations. First, the effects of a single surface crack on the static deflection and natural frequency of an electrostatically actuated micro-beam are formulated and compared. Then, a new method for approximating damage associated with realistic distributed crack networks is formulated for use in an analytical model and numerical simulations. A method for utilizing experimentally derived crack statistics to inform the analytical and numerical distributed crack models is developed. Good agreement between the analytical and numerical models is obtained for both crack scenarios. Altogether, these models can be used to effectively simulate a variety of damage and fatigue behaviors in silicon-based MEMS devices.

42 ENGINEERING↗

Intensified Soil Moisture Extremes Decrease Soil Organic Carbon Decomposition: A Mechanistic Modeling Analysis

Earth system models have predicted that there will be more frequent and severe precipitation and drought events in terrestrial ecosystems. Microbially mediated decomposition of soil organic carbon (SOC) tends to increase as soils wet and decrease as soils dry. However, the long-term SOC change under intensified moisture extremes remains poorly known as it depends on the frequency and intensity of soil drying and wetting. In this study, we explored long-term SOC dynamics under scenarios of alternating drying-wetting cycles using the Microbial-ENzyme Decomposition model, a mechanistic microbial model. The model was parameterized with 11 years of observations from a temperate deciduous broadleaf forest site, showing satisfactory model performance in both model calibration (R 2 = 0.67) and validation (R 2 = 0.69) against heterotrophic respiration. We then used the model to simulate the long-term SOC dynamics under five scenarios of alternating drying-wetting cycles with different frequencies and severities over a period of 100 years. Results showed that the changes in active microbial biomass C and the corresponding turnover rates of SOC pools were more sensitive to soil drying than soil wetting. As a result, the cumulative soil carbon emission from microbial respiration decreased by 433.7 g C m -2 after the 100-year simulation in the highest frequency and intensity moisture scenario, but was not significantly affected by the lowest frequency and intensity scenario. This study emphasizes the nonlinear response of SOC decomposition to soil moisture changes, which causes decreased decomposition by microbes under drying that is, not compensated by increased decomposition under wetting conditions.

58 GEOSCIENCES↗

Image-based novel fault detection with deep learning classifiers using hierarchical labels

One important characteristic of modern fault classification systems is the ability to flag the system when faced with previously unseen fault types. This work considers the unknown fault detection capabilities of deep neural network-based fault classifiers. Specifically, we propose a methodology on how, when available, labels regarding the fault taxonomy can be used to increase unknown fault detection performance without sacrificing model performance. To achieve this, we propose to utilize soft label techniques to improve the state-of-the-art deep novel fault detection techniques during the training process and novel hierarchically consistent detection statistics for online novel fault detection. Lastly, we demonstrated increased detection performance on novel fault detection in inspection images from the hot steel rolling process, with results well replicated across multiple scenarios and baseline detection methods.

42 ENGINEERING↗