Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Predictive Data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Lithium-Ion Battery Diagnostics Using Electrochemical Impedance via Machine-Learning

Diagnosing battery states such as health, state-of-charge, or temperature is crucial for ensuring the safety and reliability of electrochemical energy storage systems. While some states, such as temperature, may be measured using cheap sensors, accurate diagnosis of battery health metrics usually requires time-consuming performance measurements, making them infeasible for use in real-world operation. These health metrics can be measured during lab-testing and then estimated on-line using predictive life models or via state observer algorithms such as Kalman filters, but these predictive methods should be supplemented by actual measurement of battery health whenever possible to ensure reliability. Rapid measurement of battery health may be done by various types of fast diagnostic techniques such as electrochemical impedance spectroscopy (EIS), which can be performed in only a few minutes and require only a fraction of the energy and power needed for a full charge and discharge measurement. But there is a substantial challenge for estimating battery health using EIS data, as EIS is sensitive to cell temperature, state-of-charge, current, and resting time in addition to health. Thus, utilizing EIS data to predict battery capacity requires correcting for all these additional variables, a task that is extremely difficult to handle analytically. This talk utilizes machine-learning methods to estimate the effectiveness of battery capacity prediction from EIS data, leveraging a data set of hundreds of EIS measurements recorded at varying temperature and state-of-charge throughout a 500-day aging study of 32 commercial, large-format NMC-Graphite lithium-ion batteries. Using EIS as input to machine-learning models is complicated by the nonlinear response of impedance to battery health, temperature, and state-of-charge, as well as the collinearity between the impedance response at neighboring frequencies, which can easily lead to overfit models. To train robust models, features from EIS data need to be extracted from the data or some subset of critical frequencies selected. Many approaches for extracting and selecting features from EIS data from electrochemical analysis and machine-learning fields were identified for analysis: using the entire raw spectra; selection of one, two, or many frequencies from the entire spectra; selecting interesting points from the EIS measurement using domain knowledge; fitting EIS with an equivalent-circuit model; calculating statistics on the raw impedance values; and reducing the dimensionality of the data using unsupervised linear (principal component analysis) and non-linear (uniform manifold approximation and projection) methods. These approaches were rigorously compared using a machine-learning pipeline approach, training linear, Gaussian process, and random forest regression models and quantifying performance using cross-validation as well as a held-out test set. An artificial neural network model trained on the raw spectra was also tested. Promising pipelines were fine-tuned via Bayesian hyperparameter optimization using cross-validation loss and training with class-specific weights to counter data set imbalance. The most reliable method for utilizing impedance in this work was the selection of two optimal frequencies through an exhaustive search, resulting in about 2% mean absolute error on test data for both Gaussian process and random forest model architectures. Interrogation of a variety of models reveals critical frequencies of 100 Hz and 103 Hz for this data set, though the optimal set of frequencies is not necessarily intuitive, i.e., the best performing models are not simply those that use impedance at frequencies that have the highest correlation to the relative discharge capacity. The best performing model is an ensemble model, which is able to predict battery capacity with 1.9% mean absolute error for unseen cells using impedance recorded at a variety of temperatures and states-of-charge.

battery↗

A structured framework for predicting sustainable aviation fuel properties using liquid-phase FTIR and machine learning

Sustainable aviation fuels have the potential to improve efficiency, reduce emissions, and enhance energy security. To help identify viable sustainable aviation fuels and accelerate research, machine learning models have been developed to predict relevant physicochemical properties. However, many models have limited applicability, leverage data from complex analytical techniques with confined spectral ranges, or use feature decomposition methods that offer limited interpretability. Using liquid-phase Fourier Transform Infrared (FTIR) spectra, this study presents a structured method for creating accurate and interpretable property prediction models for neat molecules, aviation fuels, and blends. Liquid FTIR spectra can be collected quickly and consistently, offering high reliability, sensitivity, and component specificity using less than 2 ml of sample. The method first decomposes FTIR spectra into fundamental building blocks using non-negative matrix factorization (NMF) to enable scientific analysis of FTIR spectra attributes and fuel properties. The NMF features are then used to create five ensemble models for predicting final boiling point, flash point, freezing point, density at 15°C, and kinematic viscosity at -20°C. All models were trained using experimental property data from neat molecules, aviation fuels, and blends. The models accurately predict key properties across a broad range of neat molecules and representative fuels and blends, while enabling interpretation of relationships between compositional elements, such as functional groups or chemical classes, and their resulting properties. This demonstrates strong potential to support sustainable aviation fuel research and development. The models and data are available on an interactive web tool.

Fourier transform infrared spectroscopy↗

A graph signal processing‐based multiple model Kalman filter ( GSP‐MMKF ) tool for predictive analytics: An air separation unit process application

Abstract The industrial Air Separations Unit (ASU) is a complicated and tightly operated process. The use of dynamic process analytics is also a key element of safe and economic operation of these processes, with increasing focus on predictive analytics to take preemptive actions. With the availability of real‐time data from hundreds of sensors, the data analysis process should also consider the topology of the data, as seen in sensor networks. In this paper, a novel tool is presented that considers the complex connectivity patterns in the sensor network and uses local adaptive disturbance estimations to predict global network‐scale trends. The paper introduces the emerging field of Graph Signal Processing (GSP) and presents a rigorous derivation of the tool starting from the extraction of the sensor‐network (in a graph theoretical sense) from the data. This network, which is in the form of a matrix, is then used to derive a Kalman‐filter type of state‐space model driven by input disturbances. Multiple disturbance models (e.g., step, ramp, periodic) are included to allow the model to have different kinds of disturbance propagation. Each graph node (representing the sensors used) dynamically adapts to the most recent detected disturbance individually. These estimated disturbances are propagated to the global network using the graph. Modifications to ensure stability are also discussed. The fidelity of the tool is tested on certain downtime events and the paper concludes by discussing the advantages of the method and planned future improvements.

Ghosh, Sambit↗

Regression Analysis with the Directed Infusion of Data

Integrating artificial intelligence and machine learning tools into industry necessitates large-scale collaborative efforts that ensure the robust and accurate execution of downstream analytics such as time series prediction, uncertainty quantification, grid optimization, and condition monitoring. However, concerns related to data privacy pervade the nuclear industry due to the proprietary nature of its data and the possibility of data leakage. Legacy techniques such as encryption often require the explicit transmission of data to trustworthy parties, thereby inviting data leakage concerns. The ideal collaboration scenario avoids the explicit dissemination of data/code while maintaining experimental fidelity, which is currently accomplished using various techniques such as trusted execution environments, homomorphic encryption, differential privacy, and multimatrix masking. These techniques, however, often necessitate a trade-off between trust, efficiency, and utility. This article extends a previously proposed technique called the directed infusion of data (DIOD) that ensures data privacy, allows for scalable obfuscation, and combats the risk of data leakage without compromising utility. The experiments discussed in this article examine a regression-type scenario using DIOD with the goal of preserving the inferential link between two variables. Using the point-kinetics equations, regression experiments compare the performance of a model trained using the original data to that of a model trained using the obfuscated data, which produced identical results. Our claim is further strengthened by an information theoretic proof and experiment, which showed that the inferential content between variables remains the same after obfuscation, thereby avoiding the required communication of the proprietary data.

47 - OTHER INSTRUMENTATION↗

Mesoscale modeling and semi-analytical approach for the microstructure-aware effective thermal conductivity of porous polygranular materials

Here we established a comprehensive modeling approach for investigating the microstructure-aware effective thermal conductivity ($κ_{eff}$) for porous microstructures containing solid particles and gaseous pores. Our approach combines the mesoscale computational modeling framework and the semi-analytical method, allowing for efficient prediction of $κ_{eff}$ for realistic porous microstructures, while considering complicated microstructural thermal conduction pathways effectively in the prediction. We used the diffuse-interface mesoscale computational model to generate extensive simulated $κ_{eff}$ data for realistic digital representations of microstructures with wide ranges of porosity ($f_p$), thermal conductivity of the gas phase ($κ_g$), and thermal conductivity of the solid phase ($κ_s$). From the simulated data, we identified two property variation regimes for $κ_{eff}$: (1) a slow $κ_{eff}$ increase for $κ_s ~ κ_g$; and (2) a faster $κ_{eff}$ increase for $κ_s \gg κ_g$. To capture the key features of the relationship between the microstructure and $κ_{eff}$, we derived a semi-analytical model by introducing structure and intensification factors. The two new factors incorporate the calibrated effective contribution of the solid volume with $κ_s$ and additional interfacial effects into the prediction of $κ_{eff}$, respectively, allowing for consideration of parallel, serial, and interfacial conduction mechanisms effectively. Using the selected simulation data, we quantified key model parameters within the semi-analytical model and verified that the parameterized model exhibits excellent agreement with simulated $κ_{eff}$ for the entire range of the parameter space.

36 MATERIALS SCIENCE↗

Energy Management Information Systems Technical Resources Report

Guide supports federal facility staff in understanding, designing, procuring, and implementing Energy Management Information Systems (EMIS) as a valuable component of their portfolio-level energy and water planning and management strategies. As a broad and rapidly evolving family of tools that monitor, analyze, and control building energy use and system performance, EMIS tools present significant opportunities for federal sector energy savings and improved operational performance. EMIS are at the forefront of transforming energy management best practices by providing building owners and operators with well-organized building performance and energy consumption data, enabling a host of analytic capabilities. These capabilities include portfolio-wide energy benchmarking, data visualization, and key performance indicator tracking; automated fault detection and diagnostics (AFDD); artificial intelligence for predictive analytics and control; automated measurement and verification of energy conservation measures; and supervisory control enabling automated system optimization and demand management.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Assessing Teleconnections-Induced Predictability of Regional Water Cycle on Seasonal to Decadal Timescales Using Machine Learning Approaches

Focal Area: (3) Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI. Science Challenge: Predictive understanding of the regional water cycle based on teleconnections of climate modes of variability using AI methods.

54 ENVIRONMENTAL SCIENCES↗

Smart Mobility in the Cloud: Enabling Real-Time Situational Awareness and Cyber-Physical Control Through a Digital Twin for Traffic

This article presents the design, implementation, and use cases of the Chattanooga Digital Twin (CTwin) towards the vision for next-generation smart city applications for urban mobility management. CTwin is an end-to-end web-based platform that incorporates various aspects of the decision-making process for optimizing urban transportation systems in Chattanooga, Tennessee, to reduce traffic congestion, incidents, and vehicle fuel consumption. The platform serves as a cyberinfrastructure to collect and integrate multi-domain urban mobility data from various online repositories and Internet of Things (IoT) sensors, covering multiple urban aspects (e.g., traffic, natural hazards, weather, and safety) that are relevant to urban mobility management. The platform enables advanced capabilities for: (a) real-time situational awareness on traffic and infrastructure conditions on highways and urban roads, (b) cyber-physical control for optimizing traffic signal timing, and (c) interactive visual analytics on big urban mobility data and various metrics for traffic prediction and transportation performance evaluation. The platform is designed using a multi-level componentization paradigm and is implemented using modular and adaptive architecture, rendering it as a generalizable and extendable prototype for other urban management applications. We present several use cases to demonstrate CTwin's core capabilities for supporting decision-making in smart urban mobility management.

33 ADVANCED PROPULSION SYSTEMS↗

Spatiotemporal measurements of striations in a glow discharge’s positive column using laser-collisional induced fluorescence

Here we have observed the behavior of striations caused by ionization waves propagating in low-pressure helium DC discharges using the non-invasive laser-collision induced fluorescence (LCIF) diagnostic. To achieve this, we developed an analytic fit of collisional radiative model (CRM) predictions to interpret the LCIF data and recover quantitative two-dimensional spatial maps of the electron density, n e , and the ratios of LCIF emission states that can be correlated with T e with the use of accurate distribution functions at localized positions within striated helium discharges at 500 mTorr, 750 mTorr, and 1 Torr. To our knowledge, these are the first spatiotemporal, laser-based, experimental measurements of n e in DC striations. The n e and 447:588 ratio distributions align closely with striation theory. Constriction of the positive column appears to occur with decreased gas pressure, as shown by the radial n e distribution. We identify a transition from a slow ionization wave to a fast ionization wave between 750 mTorr and 1 Torr. These experiments validate our analytic fit of n e , allowing the implementation of an LCIF diagnostic in helium without the need to develop a CRM.

42 ENGINEERING↗

An AI-Enabled MODEX Framework for Improving Predictability of Subsurface Water Storage across Local and Continental Scales

Focal Area: (2) Predictive modeling through the use of AI techniques and AI-derived model components. (3) Insight gleaned from complex data using AI, big data analytics, and other advanced methods. We propose an AI-enabled model-experiment (MODEX) framework to improve the predictability of subsurface water storage (SWS) from local to conus scales in a changing environment by taking advantage of DOE’s observation and simulation capabilities, as well as to inform the model and the observation development.

54 ENVIRONMENTAL SCIENCES↗

Method for simultaneous characterization and expansion of reference libraries for small molecule identification

A variational autoencoder (VAE) has been developed to learn a continuous numerical, or latent, representation of molecular structure to expand reference libraries for small molecule identification. The VAE has been extended to include a chemical property decoder, trained as a multitask network, to shape the latent representation such that it assembles according to desired chemical properties. The approach is unique in its application to metabolomics and small molecule identification, focused on properties that are obtained from experimental measurements (m/z, CCS) paired with its training paradigm, which involves a cascade of transfer learning iterations. First, molecular representation is learned from a large dataset of structures with m/z labels. Next, in silico property values are used to continue training. Finally, the network is further refined by being trained with the experimental data. The trained network is used to predict chemical properties directly from structure and generate candidate structures with desired chemical properties. The network is extensible to other training data and molecular representations, and for use with other analytical platforms, for both chemical property and feature prediction as well as molecular structure generation.

Colby, Sean M.↗

Data for A Hybrid Biophysical-Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81–0.94) and H (R2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

AI/ML↗

Learning to simulate high energy particle collisions from unlabeled data

In many scientific fields which rely on statistical inference, simulations are often used to map from theoretical models to experimental data, allowing scientists to test model predictions against experimental results. Experimental data is often reconstructed from indirect measurements causing the aggregate transformation from theoretical models to experimental data to be poorly-described analytically. Instead, numerical simulations are used at great computational cost. We introduce Optimal-Transport-based Unfolding and Simulation (OTUS), a fast simulator based on unsupervised machine-learning that is capable of predicting experimental data from theoretical models. Without the aid of current simulation information, OTUS trains a probabilistic autoencoder to transform directly between theoretical models and experimental data. Identifying the probabilistic autoencoder’s latent space with the space of theoretical models causes the decoder network to become a fast, predictive simulator with the potential to replace current, computationally-costly simulators. Here, we provide proof-of-principle results on two particle physics examples, Z-boson and top-quark decays, but stress that OTUS can be widely applied to other fields.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Temperature Dependence of Band Gap Renormalization in High-T Sensor Materials via First-Principles and Experimental Corroboration

Understanding the temperature dependence of functional properties of high-T gas sensing materials is vital for their applications in combustion environments. The electron-phonon coupling that derives the electronic structure change with temperatures is a key property of interest as it affects other sensing responses. Herein, we assess the temperature dependence of band gap renormalization in metal oxides and perovskites by employing Allen-Heine-Cardona theory with first-principles simulations and corroborate with experimental observation. The calculated temperature-dependent band gap changes of these materials studied are in good agreement with in-house experimental data, proving that the theory can adequately predict renormalization on the band gap in the system of interest. The predicted and measured band gap variations are characterized using an analytical model, which can provide useful insights on the simulated zero-temperature band gaps. Based on the available data, a set of 53 metal oxides and perovskites were identified as potential high-T gas sensors. A machine learning model has been developed to predict the band-gap change by capturing the overall trend of the empirical parameters with respect to a reduced feature obtained by transforming the set of available physical features.

Park, Jongwoo↗

Facilitating better and faster simulations of aerosol-cloud interactions in Earth system models

Focal Area(s): 1. Predictive modeling through the use of AI techniques and AI-derived model components; the use of AI and other tools to design a prediction system comprising a hierarchy of models. 2. Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI. Science Challenge: One major challenge that Earth system models (ESMs) face in providing credible prediction of the Earth system and its water cycle characteristics (e.g., mean state, variability, and extreme events) is to accurately simulate aerosol-cloud interactions (ACI). The physical, chemical, and dynamical processes affecting ACI are extremely complex and they range from nanoscale to planetary scale. In each model development cycle, scientists spend significant efforts investigating model deficiencies and uncertainties associated with aerosols (e.g., emissions, chemical processes, aerosol microphysics, and transport) and clouds (e.g., macrophysics, microphysics, turbulence, and large-scale circulation) in order to develop improved treatments. However, despite decades of active research, ACI is still a major source of uncertainty in climate projections, even though great progress has been made. Specific scientific challenges include: (i) Parameterizations are developed based on limited data; (ii) The complexity of a parameterization required for accurate predictions is not understood; (iii) Incomplete and unknown physics leads to errors in the fully coupled Earth system; and (iv) Complex physics is computationally too expensive to employ in ESMs.

54 ENVIRONMENTAL SCIENCES↗

Can machine learning accelerate process understanding and decision‐relevant predictions of river water quality?

Abstract The global decline of water quality in rivers and streams has resulted in a pressing need to design new watershed management strategies. Water quality can be affected by multiple stressors including population growth, land use change, global warming, and extreme events, with repercussions on human and ecosystem health. A scientific understanding of factors affecting riverine water quality and predictions at local to regional scales, and at sub‐daily to decadal timescales are needed for optimal management of watersheds and river basins. Here, we discuss how machine learning (ML) can enable development of more accurate, computationally tractable, and scalable models for analysis and predictions of river water quality. We review relevant state‐of‐the art applications of ML for water quality models and discuss opportunities to improve the use of ML with emerging computational and mathematical methods for model selection, hyperparameter optimization, incorporating process knowledge into ML models, improving explainablity, uncertainty quantification, and model‐data integration. We then present considerations for using ML to address water quality problems given their scale and complexity, available data and computational resources, and stakeholder needs. When combined with decades of process understanding, interdisciplinary advances in knowledge‐guided ML, information theory, data integration, and analytics can help address fundamental science questions and enable decision‐relevant predictions of riverine water quality.

54 ENVIRONMENTAL SCIENCES↗

Leveraging Hydropower Multi-Sensor Data for Inference and Age-Informed Modeling

Increased demand of operational flexibility such as faster ramp up/down in generation, and more frequent start/stops are putting hydropower plants and their associated components in unprecedented stress. Consequently, these plants are at the high risk of extended and more frequent outage to accommodate unscheduled, and unexpected maintenance. Therefore, hydropower plants are in critical need of data driven and age-informed analysis for their regular and unscheduled operation. Yet not all hydropower plants are exhaustively equipped with sensors and/or measurement streams for their respective components – demanding solutions on how to detect, identify, and locate the cause of any event from the unobservable. Idaho National Laboratory (INL) analyzed the anonymized measurements and event records from the Hydropower Research Institute (HRI) to address this issue, as part of the Water Power Technologies Office (WPTO) funded one year multi-lab project. First, we investigated how time series of multiple sensor measurements can be leveraged to identify an event “root cause” as well as to develop an inference (i.e., estimate the unobservable) problem. INL also investigated how individual hydropower components’ reaction or response times vary across the pre-event, during event, and post-event conditions – enabling the hydropower dynamic models to be age-informed. Finally, the impact of clustering multi-sensor time series on short-term vibration prediction is analyzed. INL will present key findings from these analyses and recommend next steps for stakeholder adoption.

13 HYDRO ENERGY↗

In Situ Inference for Earth System Predictability

Focal Area: Focal Area 3: Insight gleaned from complex simulated data using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI. Science Challenge: An understanding of future evolution in precipitation extremes is critical to numerous DOE mission questions. Extreme events are by nature short time-scale events that are difficult to diagnose in available model data. Accurate modeling of extreme events necessarily requires high spatial resolution at the storm scale locally. However, the environment in which storms grow is dependent on global, remote, processes. These complex spatiotemporal relationships are impossible to diagnose at resolutions required to accurately model storms responsible for extreme precipitation. At exascale, climate simulations will produce results at fine enough resolution to investigate these relationships. However, the resulting data from these simulations will be far too large to save for post-simulation analysis. We advocate for fitting statistical models inside the simulations as they run, a context known as in situ, which will facilitate scientific investigations using the full fine-scale data stream. Figure 1 shows an example of the type of model we could consider, a Bayesian hierarchical spatial regression model. Precipitation extremes at each grid cell are modeled using extreme value distributions. Since extremes are rare, fitting models to individual grid cells can result in high variance and poor estimates. Instead, the model can be made more robust by smoothing the parameters of the extreme value model across space. Additionally, the parameters themselves can be functionally linked to other variables elsewhere in the simulation. Thus, we can use the fine-scale data to build more robust models for extremes that link extreme behavior to other climate patterns.

54 ENVIRONMENTAL SCIENCES↗