Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data driven model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Determination of Barometric Altimeter Errors for the Orion Exploration Flight Test-1 Entry

The EFT-1 mission is the unmanned flight test for the upcoming Multi-Purpose Crew Vehicle (MPCV). During entry, the EFT-1 vehicle will trigger several Landing and Recovery System (LRS) events, such as parachute deployment, based on onboard altitude information. The primary altitude source is the filtered navigation solution updated with GPS measurement data. The vehicle also has three barometric altimeters that will be used to measure atmospheric pressure during entry. In the event that GPS data is not available during entry, the altitude derived from the barometric altimeter pressure will be used to trigger chute deployment for the drogues and main parachutes. Therefore it is important to understand the impact of error sources on the pressure measured by the barometric altimeters and on the altitude derived from that pressure. There are four primary error sources impacting the sensed pressure: sensor errors, Analog to Digital conversion errors, aerodynamic errors, and atmosphere modeling errors. This last error source is induced by the conversion from pressure to altitude in the vehicle flight software, which requires an atmosphere model such as the US Standard 1976 Atmosphere model. There are several secondary error sources as well, such as waves, tides, and latencies in data transmission. Typically, for error budget calculations it is assumed that all error sources are independent, normally distributed variables. Thus, the initial approach to developing the EFT-1 barometric altimeter altitude error budget was to create an itemized error budget under these assumptions. This budget was to be verified by simulation using high fidelity models of the vehicle hardware and software. The simulation barometric altimeter model includes hardware error sources and a data-driven model of the aerodynamic errors expected to impact the pressure in the midbay compartment in which the sensors are located. The aerodynamic model includes the pressure difference between the midbay compartment and the free stream pressure as a function of altitude, oscillations in sensed pressure due to wake effects, and an acoustics model capturing fluctuations in pressure due to motion of the passive vents separating the barometric altimeters from the outside of the vehicle.

Brown, Denise L.↗

Genetic programming for interpretable, data-driven continuum damage models.

The damage mechanisms that lead to failure in engineering alloys have been studied extensively, but converting this knowledge into constitutive models that are suitable for engineering-scale analysis remains a challenge. Evolution laws for continuum damage have been developed in the past and have proven effective but suffer from many non-physical assumptions that inhibit the overall accuracy of the model. Further, the assumptions inherent in these existing models prevent them from being applicable to a broad class of materials. At the same time, computational models of fine-scale damage mechanisms continue to advance making it tractable to generate large training data sets through computer simulation. Data-driven machine learning approaches can leverage these data sets to avoid making limiting assumptions, and instead produce models directly from the results of microstructural simulations and/or experiments. Many of these machine learning approaches are rapid and accurate, but they offer little to no insight into the underlying relationships among state variables being discovered. Conversely, genetic programming symbolic regression (GPSR) is a machine learning method that produces analytic expressions relating the state variables, allowing maximal insight and interpretability. To that end, we propose using GPSR as a data-driven method of obtaining microstructurally informed continuum damage models. Data is generated using microstructural simulations of damage evolution, parameterized over microstructural statistics (i.e., pore shape) and nominally applied deformations. Analytic expressions for damage evolution are obtained from the data using GPSR, and these expressions are then utilized within a continuum constitutive model. Overall, this approach is a promising method of automatically obtaining analytic relations describing constitutive phenomena in a material.

Buche, Michael Robert↗

Learning the structure of wind: A data-driven nonlocal turbulence model for the atmospheric boundary layer

In this work, we develop a novel data-driven approach to modeling the atmospheric boundary layer. This approach leads to a nonlocal, anisotropic synthetic turbulence model which we refer to as the deep rapid distortion (DRD) model. Our approach relies on an operator regression problem that characterizes the best fitting candidate in a general family of nonlocal covariance kernels parameterized in part by a neural network. This family of covariance kernels is expressed in Fourier space and is obtained from approximate solutions to the Navier–Stokes equations at very high Reynolds numbers. Each member of the family incorporates important physical properties such as mass conservation and a realistic energy cascade. The DRD model can be calibrated with noisy data from field experiments. After calibration, the model can be used to generate synthetic turbulent velocity fields. To this end, we provide a new numerical method based on domain decomposition which delivers scalable, memory-efficient turbulence generation with the DRD model as well as others. We demonstrate the robustness of our approach with both filtered and noisy data coming from the 1968 Air Force Cambridge Research Laboratory Kansas experiments. Using these data, we witness exceptional accuracy with the DRD model, especially when compared to the International Electrotechnical Commission standard.

17 WIND ENERGY↗

DATA DRIVEN APPROACH TO ANALYZING THE IMPACT OF POWER PLANT CYCLING ON AIR PREHEATER DEGRADATION AND REMAINING USEFUL LIFE

Due to the increased penetration of renewable energy generation sources, and fluctuations of the oil and gas prices, modern coal burning power plants deal with increased variability in the demand for power generation. These varying demands result in their intermittent under-capacity operation (cycling). Periodical ramping down and back up to follow the daily power demands causes damages to the plant components reducing its operational life. In this paper we analyze the impact of cycling on a rotary Ljungstrom air preheater (APH) unit installed at a coal fire power plant in the US. An inefficient air preheater can significantly impact boiler performance. Due to the repeated boiler’s hot-cold start, the APH experiences fluctuating operating conditions that result in accelerated degradation mechanisms, such as dew-point corrosion, fouling/deposition plugging, and air heater leakage.The analysis in this paper utilizes field data related to APH basket replacement, and the number of cycles experienced by the boiler to model the life expectancy of the baskets. The data-driven model enables preventive maintenance strategies for the APH by predicting how long the APH baskets will last in a probabilistic sense. The analysis showed that an increase in cycling for a fixed operation time can reduce the APH basket remaining useful life by about 30%.

Sharma, Himanshu↗

Data Driven Approach to Analyzing The Impact of Power Plant Cycling on Air Preheater Degradation and Remaining Useful Life

Due to the increased penetration of renewable energy generation sources, and fluctuations of the oil and gas prices, modern coal burning power plants deal with increased variability in the demand for power generation. These varying demands result in their intermittent under-capacity operation (cycling). Periodical ramping down and back up to follow the daily power demands causes damages to the plant components reducing its operational life. In this paper we analyze the impact of cycling on a rotary Ljungstrom air preheater (APH) unit installed at a coal fire power plant in the US. An inefficient air preheater can significantly impact boiler performance. Due to the repeated boiler’s hot-cold start, the APH experiences fluctuating operating conditions that result in accelerated degradation mechanisms, such as dew-point corrosion, fouling/deposition plugging, and air heater leakage. The analysis in this paper utilizes field data related to APH basket replacement, and the number of cycles experienced by the boiler to model the life expectancy of the baskets. The data-driven model enables preventive maintenance strategies for the APH by predicting how long the APH baskets will last in a probabilistic sense. Furthermore, the analysis showed that an increase in cycling for a fixed operation time can reduce the APH basket remaining useful life by about 30%.

20 FOSSIL-FUELED POWER PLANTS↗

Streamlining Ocean Dynamics Modeling with Fourier Neural Operators: A Multiobjective Hyperparameter and Architecture Optimization Approach

Training an effective deep learning model to learn ocean processes involves careful choices of various hyperparameters. We leverage DeepHyper’s advanced search algorithms for multiobjective optimization, streamlining the development of neural networks tailored for ocean modeling. The focus is on optimizing Fourier neural operators (FNOs), a data-driven model capable of simulating complex ocean behaviors. Selecting the correct model and tuning the hyperparameters are challenging tasks, requiring much effort to ensure model accuracy. DeepHyper allows efficient exploration of hyperparameters associated with data preprocessing, FNO architecture-related hyperparameters, and various model training strategies. We aim to obtain an optimal set of hyperparameters leading to the most performant model. Moreover, on top of the commonly used mean squared error for model training, we propose adopting the negative anomaly correlation coefficient as the additional loss term to improve model performance and investigate the potential trade-off between the two terms. The numerical experiments show that the optimal set of hyperparameters enhanced model performance in single timestepping forecasting and greatly exceeded the baseline configuration in the autoregressive rollout for long-horizon forecasting up to 30 days. Utilizing DeepHyper, we demonstrate an approach to enhance the use of FNO in ocean dynamics forecasting, offering a scalable solution with improved precision.

97 MATHEMATICS AND COMPUTING↗

Individual Data Sparsity in Smart Thermostat Big Data: Impacts on Modeling Thermostat Use Behavior Dynamics

This study explores the impacts of the sparsity of individual thermostat interaction data on modeling thermostat use behavior dynamics using a dataset of over 100,000 smart thermostats. In developing a data-driven model of Thermal Frustration Theory (TFT), we investigate the challenges and trade-offs in clustering occupant data to enhance predictive accuracy. Our findings reveal that a single, aggregated model fails to capture the diversity of occupant behaviors, resulting in extremely poor prediction performance. Conversely, excessive clustering exacerbates data sparsity, undermining model reliability. By identifying an optimal clustering strategy, we achieve a balance that significantly improves the prediction of manual setpoint changes during demand response (DR) events, enhancing energy management and occupant comfort

Fannon, David↗

Data-Efficient Dimensionality Reduction and Surrogate Modeling of High-Dimensional Stress Fields

Tensor datatypes representing field variables like stress, displacement, velocity, etc., have increasingly become a common occurrence in data-driven modeling and analysis of simulations. Numerous methods [such as convolutional neural networks (CNNs)] exist to address the meta-modeling of field data from simulations. As the complexity of the simulation increases, so does the cost of acquisition, leading to limited data scenarios. Modeling of tensor datatypes under limited data scenarios remains a hindrance for engineering applications. Here, in this article, we introduce a direct image-to-image modeling framework of convolutional autoencoders enhanced by information bottleneck loss function to tackle the tensor data types with limited data. The information bottleneck method penalizes the nuisance information in the latent space while maximizing relevant information making it robust for limited data scenarios. The entire neural network framework is further combined with robust hyperparameter optimization. We perform numerical studies to compare the predictive performance of the proposed method with a dimensionality reduction-based surrogate modeling framework on a representative linear elastic ellipsoidal void problem with uniaxial loading. The data structure focuses on the low-data regime (fewer than 100 data points) and includes the parameterized geometry of the ellipsoidal void as the input and the predicted stress field as the output. The results of the numerical studies show that the information bottleneck approach yields improved overall accuracy and more precise prediction of the extremes of the stress field. Additionally, an in-depth analysis is carried out to elucidate the information compression behavior of the proposed framework.

artificial intelligence↗

Predictive Analytics for Hydropower Fleet Intelligence

A primary challenge in hydropower industry is the ability to maintain cost-competitiveness, reliability, and security of hydropower assets through evolving power system contexts and aging of the fleet. Maintaining cost-effective and reliable operations under these conditions is expected to require new modernization and maintenance paradigms for changing contexts. Changes in existing practices for O&M will require an understanding of the current state and health of hydropower assets, and the impact of changing paradigms on asset health and reliability. The Hydropower Fleet Intelligence project is developing and evaluating standardized methodologies and analysis tools for data-driven asset reliability and management technologies for hydropower, leading to eventual predictive maintenance planning, repair/replacement decision making, and asset-reliability and cost-optimized operations. A key question is the feasibility of using existing data sets at hydropower facilities to perform assessments of asset reliability. This document uses data from hydropower facilities to assess the potential for using available analytics methods for asset reliability estimates. In addition to reliability assessments, the feasibility of using existing analytics techniques for several other potential applications is discussed. Finally, a case study that a data-driven model is trained to learn nominal operations via vibration data from an asset of a certain plant, and then utilized to identify anomalies on a similar asset from a different plant, highlighting the generic use of proposed Prognostics and Health Management (PHM) approaches.

Yucesan, Yigit↗

A Data System for a Rapid Evaluation Class of Subscale Aerial Vehicle

A low cost, rapid evaluation, test aircraft is used to develop and test airframe damage diagnosis algorithms at Langley Research Center as part of NASA's Aviation Safety Program. The remotely operated subscale aircraft is instrumented with sensors to monitor structural response during flight. Data is collected for good and compromised airframe configurations to develop data driven models for diagnosing airframe state. This paper describes the data acquisition system (DAS) of the rapid evaluation test aircraft. A PC/104 form factor DAS was developed to allow use of Matlab, Simulink simulation code in Langley's existing subscale aircraft flight test infrastructure. The small scale of the test aircraft permitted laboratory testing of the actual flight article under controlled conditions. The low cost and modularity of the DAS permitted adaptation to various flight experiment requirements.

Hogge, Edward F.↗

Learning thermodynamic master equations for open quantum systems

The characterization of Hamiltonians and other components of open quantum dynamical systems plays a crucial role in quantum computing and other applications. Scientific machine learning techniques have been applied to this problem in a variety of ways, including by modeling with deep neural networks. However, the majority of mathematical models describing open quantum systems are linear, and the natural nonlinearities in learnable models have not been incorporated using physical principles. We present a data-driven model for open quantum systems that includes learnable, thermodynamically consistent terms. The trained model is interpretable, as it directly estimates the system Hamiltonian and linear components of coupling to the environment. We validate the model on synthetic two and three-level data, as well as experimental two-level data collected from a quantum device at Lawrence Livermore National Laboratory.

Mathematics and Computing↗

Vapor–liquid equilibrium estimation of n-alkane/nitrogen mixtures using neural networks

Understanding fluid phase behavior, like VLE, in high P&T conditions is crucial for developing high-fidelity simulations of chemically reacting flows in liquid-fueled combustion systems and also forms an integral part of the design-modeling of the control processes in chemical industries. Two data-driven models have been proposed here in this study, each of which was competent in estimating VLE for the Type III binary systems of C 10 /N 2 and C 12 /N 2 , at pressures ranging up to 50–60 MPa. Both models showed better performance in predicting equilibrium pressure as compared to VLE modeled using PR-EOS. A modified model has also been proposed, capable of estimating the full phase envelope for the binary systems of C 10 /N 2 and C 12 /N 2 across a wide range of temperatures, and thus exhibit the mixture critical pressure at the concerned temperature. The diverse applicability of the proposed network architecture was further exhibited while estimating the VLE of a ternary system of C 1 /C 10 /N 2 .

97 MATHEMATICS AND COMPUTING↗

Assimilation of Multiscale Data into Multifidelity Biogeochemical Models (Final Report)

Quantitative predictions of subsurface processes rely on computational models that capture, with different degrees of fidelity, complex interactions between hydrologic and biogeochemical processes. Molecular- and pore-scale models provide a high-fidelity representation of these processes but are impractical at the field scale. Reduced complexity (e.g., field-scale or data-driven) models sacrifice some degree of fidelity in favor of computational efficiency. Uncertainty and assimilation of data into model predictions, pose a question of model selection: Given a significant difference in computational cost, can a lower-fidelity model be preferable to its higher-fidelity counterpart? Since a predictive model must be computable in reasonable time it has to operate at the field scale, with molecular- and pore-scale data (obtained either from simulations or measurements) determining both the model's structure and parameters. Availability of such multi-resolution data raises a question hitherto undressed in hydrogeology: Can a coarse model dynamically ``learn'' its own structure (e.g., adjust the reaction pathways in its transport module) as more fine-scale data become available during simulations? Our results in multifidelity simulations, data assimilation, and machine learning led us to hypothesize that the answer to these questions is ``yes''. The overarching goal of this project was to confirm this hypothesis by developing scalable, computationally efficient tools for assimilation of data into multiscale models, in which fine-scale data and simulations dynamically inform and autonomously modify coarse-scale models. Our research led to nine manuscripts, three of which have been published and the other six are currently under review.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Developing Data-Driven Synthetic Infrastructure Models for Resilience Analysis

Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Building thermal load prediction through shallow machine learning and deep learning

Building thermal load prediction informs the optimization of cooling plant and thermal energy storage. Physics-based prediction models of building thermal load are constrained by the model and input complexity. In this study, we developed 12 data-driven models (7 shallow learning, 2 deep learning, and 3 heuristic methods) to predict building thermal load and compared shallow machine learning and deep learning. The 12 prediction models were compared with the measured cooling demand. It was found XGBoost (Extreme Gradient Boost) and LSTM (Long Short Term Memory) provided the most accurate load prediction in the shallow and deep learning category, and both outperformed the best baseline model, which uses the previous day's data for prediction. Then, we discussed how the prediction horizon and input uncertainty would influence the load prediction accuracy. Major conclusions are twofold: first, LSTM performs well in short-term prediction (1 h ahead) but not in long term prediction (24 h ahead), because the sequential information becomes less relevant and accordingly not so useful when the prediction horizon is long. Second, the presence of weather forecast uncertainty deteriorates XGBoost's accuracy and favors LSTM, because the sequential information makes the model more robust to input uncertainty. Training the model with the uncertain rather than accurate weather data could enhance the model's robustness. Finally, our findings have two implications for practice. First, LSTM is recommended for short-term load prediction given that weather forecast uncertainty is unavoidable. Second, XGBoost is recommended for long term prediction, and the model should be trained with the presence of input uncertainty.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Predictive Modeling and Uncertainty Quantification in Condition Monitoring of Active Components: A Reactor Coolant Pump Use Case

This work develops data-driven models for onset of thermal barrier leakage in reactor coolant pumps. It incorporates uncertainty quantification to enhance the reliability and robustness of pre- dictions. Using synthetic data generated by the Generic Pressurized Water Reactor simulator, realistic degradation scenarios were simulated across lifecycle stages—beginning, middle, and end of life. Key variables, including differential pressure, flow rate, vibration, and temperatures, were analyzed using machine learning framework. The fully connected neural network models demonstrated exceptional performance, achieving R2 scores exceeding 0.99 and root mean square errors as low as around 8.23 × 10-2 gallon per minute (gpm) for the three stages of the lifecy- cle. UQ analysis further validated the model’s robustness, with narrow uncertainty bounds during steady-state operations and appropriately wider bounds during transitional phases, reflecting the physical behavior of the system. This work addresses important gaps in real-time condition moni- toring and regulatory compliance by integrating advanced condition monitoring technologies with UQ into IST programs. The ability to detect thermal barrier leakage early and quantify prediction reliability supports optimizing maintenance strategies while ensuring nuclear power plants’ safe and reliable operation.

99 - GENERAL AND MISCELLANEOUS↗

Physics-Informed Machine Learning for Epidemiological Models

One challenge of using compartmental SEIR models for public health planning is the difficulty in manually tuning parameters to capture behavior reflected in the real-world data. This team conducted initial, exploratory analysis of a novel technique to use physics-informed machine learning tools to rapidly develop data-driven models for physical systems. This machine learning approach may be used to perform data assimilation of compartment models which account for unknown interactions between geospatial domains (i.e. diffusion processes coupling across neighborhoods/counties/states/etc.). Results presented here are early, proof-of-concept ideas that demonstrate initial success in using a physically informed neural network (PINN) model to assimilate data in a compartmental epidemiology model. The results demonstrate initial success and warrant further research and development.

60 APPLIED LIFE SCIENCES↗

Physics-Reinforced Machine Learning Algorithms for Multiscale Closure Model Discovery

The central objective of this project was to address the challenge of modeling and simulating complex multiscale turbulence phenomena by leveraging physics-guided machine learning (PGML) and hybrid modeling approaches. By integrating physics-based methods with data-driven models, the research focused on achieving robust and scalable solutions for geophysical turbulence, enhancing numerical weather prediction and climate research tools. The project resulted in significant advancements in computational modeling paradigms, predictive tools for reduced-order modeling, and innovative algorithms for fluid dynamics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗