Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Predictive Data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Multifidelity Neural Network Formulations for Prediction of Reactive Molecular Potential Energy Surfaces

Here, this paper focuses on the development of multifidelity modeling approaches using neural network surrogates, where training data arising from multiple model forms and resolutions are integrated to predict high-fidelity response quantities of interest at lower cost. We focus on the context of quantum chemistry and the integration of information from multiple levels of theory. Important foundations include the use of symmetry function-based atomic energy vector constructions as feature vectors for representing structures across families of molecules and single-fidelity neural network training capabilities that learn the relationships needed to map feature vectors to potential energy predictions. These foundations are embedded within several multifidelity topologies that decompose the high-fidelity mapping into model-based components, including sequential formulations that admit a general nonlinear mapping across fidelities and discrepancy-based formulations that presume an additive decomposition. Methodologies are first explored and demonstrated on a pair of simple analytical test problems and then deployed for potential energy prediction for C 5 H 5 using B2PLYP-D3/6-311++G(d,p) for high-fidelity simulation data and Hartree–Fock 6-31G for low-fidelity data. For the common case of limited access to high-fidelity data, our computational results demonstrate that multifidelity neural network potential energy surface constructions achieve roughly an order of magnitude improvement, either in terms of test error reduction for equivalent total simulation cost or reduction in total cost for equivalent error.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

COVID19 Disease Map, a computational knowledge repository of virus–host interaction mechanisms

We need to effectively combine the knowledge from surging literature with complex datasets to propose mechanistic models of SARS-CoV-2 infection, improving data interpretation and predicting key targets of intervention. Here, we describe a large-scale community effort to build an open access, interoperable and computable repository of COVID-19 molecular mechanisms. The COVID-19 Disease Map (C19DMap) is a graphical, interactive representation of disease-relevant molecular mechanisms linking many knowledge sources. Notably, it is a computational resource for graph-based analyses and disease modelling. To this end, we established a framework of tools, platforms and guidelines necessary for a multifaceted community of biocurators, domain experts, bioinformaticians and computational biologists. The diagrams of the C19DMap, curated from the literature, are integrated with relevant interaction and text mining databases. We demonstrate the application of network analysis and modelling approaches by concrete examples to highlight new testable hypotheses. This framework helps to find signatures of SARS-CoV-2 predisposition, treatment response or prioritisation of drug candidates. Such an approach may help deal with new waves of COVID-19 or similar pandemics in the long-term perspective.

59 BASIC BIOLOGICAL SCIENCES↗

Establishing performance metrics for quantitative non-targeted analysis: a demonstration using per- and polyfluoroalkyl substances

Abstract Non-targeted analysis (NTA) is an increasingly popular technique for characterizing undefined chemical analytes. Generating quantitative NTA (qNTA) concentration estimates requires the use of training data from calibration “surrogates,” which can yield diminished predictive performance relative to targeted analysis. To evaluate performance differences between targeted and qNTA approaches, we defined new metrics that convey predictive accuracy, uncertainty (using 95% inverse confidence intervals), and reliability (the extent to which confidence intervals contain true values). We calculated and examined these newly defined metrics across five quantitative approaches applied to a mixture of 29 per- and polyfluoroalkyl substances (PFAS). The quantitative approaches spanned a traditional targeted design using chemical-specific calibration curves to a generalizable qNTA design using bootstrap-sampled calibration values from “global” chemical surrogates. As expected, the targeted approaches performed best, with major benefits realized from matched calibration curves and internal standard correction. In comparison to the benchmark targeted approach, the most generalizable qNTA approach (using “global” surrogates) showed a decrease in accuracy by a factor of ~4, an increase in uncertainty by a factor of ~1000, and a decrease in reliability by ~5%, on average. Using “expert-selected” surrogates ( n = 3) instead of “global” surrogates ( n = 25) for qNTA yielded improvements in predictive accuracy (by ~1.5×) and uncertainty (by ~70×) but at the cost of further-reduced reliability (by ~5%). Overall, our results illustrate the utility of qNTA approaches for a subclass of emerging contaminants and present a framework on which to develop new approaches for more complex use cases. Graphical Abstract

Pu, Shirley (ORCID:0000000201223797)↗

AI-Constrained Bottom-Up Ecohydrology and Improved Prediction of Seasonal, Interannual, and Decadal Flood and Drought Risks

Focal Areas: (2) Predictive modeling through the use of AI techniques and AI-derived model components; the use of AI and other tools to design a prediction system composed of a hierarchy of models (3)Insight gleaned from complex data (both observed & simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI

54 ENVIRONMENTAL SCIENCES↗

Predicting battery capacity from impedance at varying temperature and state of charge using machine learning

Prediction of battery health from electrochemical impedance spectroscopy (EIS) data can enable rapid measurement of battery state in real-world applications without using additional sensors or time-consuming performance measurements. However, deconvoluting the effect of capacity, state of charge, and temperature on EIS response is complicated analytically. Here, various machine-learning models, such as linear, Gaussian process, random forest, and artificial neural network regression, are utilized to predict capacity from EIS using hundreds of capacity, direct current (DC) resistance, and EIS measurements recorded under varying conditions of health, temperature, and state of charge (SOC). Several feature extraction and selection methods from traditional electrochemical analysis and statistical modeling are explored using machine-learning pipelines. EIS data from just two frequencies can accurately predict capacity, and interrogation shows that the optimal set of frequencies is not usually intuitive. Best results are achieved with an ensemble model, which predicts battery capacity with a mean absolute error of 1.9% on data from unobserved cells.

25 ENERGY STORAGE↗

Process prediction and detection of faults using probabilistic bidirectional recurrent neural networks on real plant data

Attaining Industry 4.0 for manufacturing operations requires advanced monitoring systems and real-time data analytics of plant data, among other topics. We propose a Probabilistic Bidirectional Recurrent Network (PBRN) for industrial process monitoring for the early detection of faults. The model is based on a Gated Recurrent Unit (GRU) neural network that allows the model to retain long-term dependencies between sensor data along a time horizon, hence learning the dynamic behavior of the process. To reduce the false-positive detection rate of the model, we compel the model to learn from a highly noisy sensor reading while outputting noise-free sensor outputs. The performance of the proposed model is compared to other data-driven statistical process monitoring schemes using real plant data from an industrial Air Separations Unit (ASU) containing noisy sensor readings. We show that the model can learn from noisy data without reducing its performance. Using two different fault cases, we demonstrate the model’s ability to carry out early fault detection with average false-positive rates of 2.9% and 4.9% for both fault cases. The missed detection rates are 0.1% and 0.2%, respectively.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Elucidating and predicting the dynamic evolution of water and land systems due to natural and energy-related forcings

Focal Area(s): 3. Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI; & 1. Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: Interactions between water, land, and energy systems are complex and occur on a variety of scales, ranging from local to basinal to regional. Accurately predicting the behavior of ground water and surface water systems for 5-10 years and beyond requires an understanding of the current system and the ability to model both the natural system at scale and human-induced forcings related to energy and other activities. Artificial intelligence and machine learning (AI/ML) combined with modern compilation and integration efforts for U.S. groundwater and surface water systems present potential solutions to bolstering detailed physics-based models of these systems. Big data tied with ML and physics-based modeling can drive breakthroughs in understanding the earth system, but research is often impeded by data access (e.g., privacy issues), quality, formats, gaps, multi-source, multi-scale, integration, and spatiotemporal challenges. Effective integration of real data and simulated (synthetic) data that fill gaps is critical. Overcoming these complex data and model integration challenges will enable a transformational approach to acquiring enhanced understanding of environmental systems.

54 ENVIRONMENTAL SCIENCES↗

Forecasting of in situ electron energy loss spectroscopy

Abstract Forecasting models are a central part of many control systems, where high-consequence decisions must be made on long latency control variables. These models are particularly relevant for emerging artificial intelligence (AI)-guided instrumentation, in which prescriptive knowledge is needed to guide autonomous decision-making. Here we describe the implementation of a long short-term memory model (LSTM) for forecasting in situ electron energy loss spectroscopy (EELS) data, one of the richest analytical probes of materials and chemical systems. We describe key considerations for data collection, preprocessing, training, validation, and benchmarking, showing how this approach can yield powerful predictive insight into order-disorder phase transitions. Finally, we comment on how such a model may integrate with emerging AI-guided instrumentation for powerful high-speed experimentation.

36 MATERIALS SCIENCE↗

Can a shock-induced phonon up-pumping model relate to impact sensitivity of molecular crystals, polymorphs and cocrystals?

Impact sensitivity engineering of high-energy molecular crystals requires accurate predictive models. For this purpose, the promising multi-phonon based approach is selected, assessing a bit more its strengths and weaknesses. Presently used with high-quality phonon calculations of 22 molecular crystals, using a physics-based criterion to determine the phonon bath extent, the resulting intrinsic shock sensitivity index (SSI) is compared to the most common marker of impact sensitivity, h 50 , as determined from drop-weight impact tests. Selecting a data subset from experiments performed under very similar conditions (2.5 kg hammer with grit and 30–40 mg samples), the model can predict h 50 values for mono-molecular crystals with very good accuracy, including the ability to discriminate the polymorphs of HMX and CL20. This very good agreement validates an initial indirect up-pumping mechanism occurring under these conditions, where the doorway modes also interact with the phonon bath. However, the phonon bath criterion for mono-molecular crystals does not transfer well to cocrystals. Owing to the vibrational coupling of the co-molecules, it seems a broader phonon bath should be considered. Additionally recalling experimental uncertainty and various experimental factors affecting h 50 values for a given compounds, we recommend that the density of the sample, granularity and morphology be systematically considered and reported along with measurements, which will in turn allow for more systematic data and predictive capabilities for sensitivity models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improve wildfire predictability driven by extreme water cycle with interpretable physically-guided ML/AI

Focal Area(s): Predictive modeling using AI techniques and AI-derived model components; use of AI and other tools to design a prediction system comprising of a hierarchy of models (Primary); insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics or knowledge-guided AI (Secondary). Science Challenge: Wildfires modify land surface characteristics, such as vegetation composition, soil and litter carbon stocks, and surface albedo, with significant consequences for the regional carbon cycle. For example, tropical regions (i.e., African and South America) are particularly vulnerable to wildfire, account for more than 80% of the global burned area, and emit ~1.4 PgC y -1 into the atmosphere together with other dust and aerosols that strongly affect regional climate.

54 ENVIRONMENTAL SCIENCES↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

Power System Waveform Datasets for Machine Learning

The desire for increased visibility across the electricity grid will necessarily increase the deployment of sensing and measurement devices and associated data management needs to unprecedented levels. For the existing sensing and measurement infrastructure, there remains a great amount of “value” yet to be extracted through advanced data management and analytics. Availability of more data will not, by itself, lead to changes in grid visibility, security, and resiliency. To create the predictive and prescriptive environment required to enable new markets and transactions for customer revenue and a reliable grid, the data must be collected, organized, evaluated, and analyzed using sophisticated algorithms to provide actionable information allowing operators and customers to reliably manage an increasingly complex grid. Progress in artificial intelligence (AI) has been largely driven by large, publicly available datasets that can be used to train AI algorithms such as MNIST, a database of handwritten images of digits, and ImageNet, an image database of everyday objects. These types of publicly available databases of real-world training datasets have been largely credited for advancement of image processing, computer vision, and deep learning algorithms that these use cases deploy. However, in the power systems industry to date, there are few databases with proper event labeling, and data access to a publicly available collection of power system event waveforms that will allow users to interact with grid signature data. Publicly available datasets of power system event waveforms, such as the DOE/EPRI dataset, often lack critical metadata or contain limited examples of each event type, and data formats vary widely across these datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Correlations for the specific heat capacity of ( U x Pu 1 - x ) 1 - y Gd y O 2 - z derived from molecular dynamics

We report UO 2 is the primary conventional fuel used in most nuclear reactors with Gd 2 O 3 commonly added as a burnable absorber to produce a more level power distribution in the reactor core at the beginning of operation. It can also be mixed with other actinide oxides to produce mixed oxide (MOx) fuel. In this study, molecular dynamics simulations were used to predict the specific heat capacity of Gd-doped PuO 2 , UO 2 and (U, Pu)O 2 MOx accommodating Gd 3+ substituted at cation sites via two charge compensation mechanisms - oxygen vacancy formation and the oxidation of U 4+ to U 5+ . The specific heat capacity values for PuO 2 and UO 2 are in good agreement with other studies showing a distinct peak at high temperatures - above 1800 K. As Gd 3+ is added, the peak height reduces for each composition considered. An analytical fit was applied to the data where Gd 3+ was fully charge compensated by either oxygen vacancies or U 5+ . The expression was then validated by predicting the specific heat capacity for three compositions of (Ux Pu 1-x ) 1-y Gd y O 2-z containing both oxygen vacancies and U 5+ , and compared to molecular dynamics data.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

HPC Analytics of Fused Thermal Plants Data to Optimize Operating Envelope

In this project, ORNL extensively reviewed the ORAP RAM data, and it guided us to develop machine learning models that can predict time to next failures and forecast failure trends, which will be useful for optimizing power plant operation strategies. More specifically, we trained multiple random forest models and evaluated the model accuracy to validate with 10+ years of historical data. In addition, we implemented a web-based graphical user interface system for the models to show how our models can be used in more intuitive ways. This proof of concept allowed exploration of model use with power plant operators in mind. Developed machine learning models will be helpful for managing risks, planning maintenance and operation, ultimately reducing the down time and increasing the service hours. For future work, there are several interesting research topics including but not limited to model enhancement, creating synergy with traditional failure modeling approaches, and data-driven actionable recommendation and suggestions.

20 FOSSIL-FUELED POWER PLANTS↗

Worldwide Physics-Based Lifetime Prediction of c-Si Modules Due to Solder-Bond Failure

Lifetime prediction of the fielded c-Si solar modules due to location-specific weather conditions has been an important topic of photovoltaic research and the economic viability of solar energy. Data analytic techniques such as the performance ratio method, Statistical clear sky model, and Suns-Vmp methods quantify the degradation from measured data of a solar farm, however, the nonlinear time-dependence and correlated degradations make it difficult to use the empirical degradation rates for ultimate lifetime projection. In this article, we propose a complementary physics-based model to predict the solder bond failure caused by mechanical stress associated with the variations of the temperature. Integrating the worldwide weather information from NASA/NSRDB databases, the model predicts the location-specific output-power degradation and the lifetime of a module due to solder bond failure. The model parameters are calibrated against qualification tests involving thermal cycling of specific batches of modules from a specific technology/manufacturer. The results may be summarized as: 1) Modules installed at higher latitudes show a longer lifetime due to reduced damage accumulation. 2) The reduction of temperature fluctuation close to large bodies of water, such as seashores, increases solder bond lifetime significantly. 3) Relatively speaking, modules installed close to the Tropic of Cancer/Capricorn (23.5 degrees North/South) suffer from a higher solder bond damage and have a shorter lifetime, suggesting a conservative design. This model should serve as a building block of a comprehensive reliability framework that can predict the lifetime of a module that experiences simultaneous and correlated degradation mechanisms involving yellowing, corrosion, and potential-induced degradation.

14 SOLAR ENERGY↗

Machine learning models inaccurately predict current and future high-latitude C balances

The high-latitude carbon (C) cycle is a key feedback to the global climate system, yet because of system complexity and data limitations, there is currently disagreement over whether the region is a source or sink of C. Recent advances in big data analytics and computing power have popularized the use of machine learning (ML) algorithms to upscale site measurements of ecosystem processes, and in some cases forecast the response of these processes to climate change. Due to data limitations, however, ML model predictions of these processes are almost never validated with independent datasets. To better understand and characterize the limitations of these methods, we develop an approach to independently evaluate ML upscaling and forecasting. We mimic data-driven upscaling and forecasting efforts by applying ML algorithms to different subsets of regional process-model simulation gridcells, and then test ML performance using the remaining gridcells. In this study, we simulate C fluxes and environmental data across Alaska using ecosys, a process-rich terrestrial ecosystem model, and then apply boosted regression tree ML algorithms to training data configurations that mirror and expand upon existing AmeriFLUX eddy-covariance data availability. We first show that a ML model trained using ecosys outputs from currently-available Alaska AmeriFLUX sites incorrectly predicts that Alaska is presently a modeled net C source. Increased spatial coverage of the training dataset improves ML predictions, halving the bias when 240 modeled sites are used instead of 15. However, even this more accurate ML model incorrectly predicts Alaska C fluxes under 21st century climate change because of changes in atmospheric CO 2 , litter inputs, and vegetation composition that have impacts on C fluxes which cannot be inferred from the training data. Our results provide key insights to future C flux upscaling efforts and expose the potential for inaccurate ML upscaling and forecasting of high-latitude C cycle dynamics.

54 ENVIRONMENTAL SCIENCES↗

Predictive Analytics for Hydropower Fleet Intelligence

A primary challenge in hydropower industry is the ability to maintain cost-competitiveness, reliability, and security of hydropower assets through evolving power system contexts and aging of the fleet. Maintaining cost-effective and reliable operations under these conditions is expected to require new modernization and maintenance paradigms for changing contexts. Changes in existing practices for O&M will require an understanding of the current state and health of hydropower assets, and the impact of changing paradigms on asset health and reliability. The Hydropower Fleet Intelligence project is developing and evaluating standardized methodologies and analysis tools for data-driven asset reliability and management technologies for hydropower, leading to eventual predictive maintenance planning, repair/replacement decision making, and asset-reliability and cost-optimized operations. A key question is the feasibility of using existing data sets at hydropower facilities to perform assessments of asset reliability. This document uses data from hydropower facilities to assess the potential for using available analytics methods for asset reliability estimates. In addition to reliability assessments, the feasibility of using existing analytics techniques for several other potential applications is discussed. Finally, a case study that a data-driven model is trained to learn nominal operations via vibration data from an asset of a certain plant, and then utilized to identify anomalies on a similar asset from a different plant, highlighting the generic use of proposed Prognostics and Health Management (PHM) approaches.

Yucesan, Yigit↗

Integration and validation of some modules for modelling of high-speed chemically reactive flows in two-phase gas-droplet mixtures

Three modules are integrated into the built-in OpenFOAM rhoCentralFoam solver towards accurate and efficient modelling of high-speed chemically reactive flows in two-phase gas-droplet mixtures within the OpenFOAM 10.0 framework. The first module is the mixture-averaged diffusion model. The second module is the built-in OpenFOAM Lagrangian solver coupled with optimised droplet drag coefficient and convective heat transfer coefficient sub-models. The last module is a sparse stiff chemistry solver based on dynamic adaptive hybrid integration (AHI-S). The optimised droplet sub-models are first verified in correct implementation for subsequent simulations in this work. Further, they show good accuracy against experimental and analytical data in the modelling of ammonia droplet acceleration and cooling in the flowing and/or low-temperature air. The accuracy and efficiency gains related to the mixture-averaged diffusion model and the AHI-S chemistry solver are examined by simulating 1-D detonation propagation in ammonia droplet-free/laden ammoniaoxygen mixtures. Numerical results of detonation propagation speed, gaseous temperature, density, and species distributions around the induction zone show good agreement with experimental data and analytical solutions. Compared to the built-in OpenFOAM diffusion model, the mixture-averaged diffusion model provides different numerical predictions of pulsating instabilities in detonation propagation. It shows better accuracy in depicting the detonation structure within the droplet-free section attributed to improved multi-component diffusion modelling. Compared to the built-in OpenFOAM solver EulerImplicit (backward Euler), the AHI-S chemistry solver reduces the computational cost by around 50%. It achieves satisfactory accuracy in calculating detonation propagation speed within the droplet-free section with the optimal efficiency when the safety factor, β, equals 0.5.

42 ENGINEERING↗