Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Domain knowledge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Ultra-Short-Term Spatiotemporal Forecasting of Renewable Resources: An Attention Temporal Convolutional Network Based Approach

The rapid increase in the penetration of renewable energy resources characterized by high variability and uncertainty is bringing new challenges to the power system operation. To ensure the efficient and reliable operation of electric grid, an accurate and general short-term forecasting algorithm with interpretability is desired. Moreover, the extensive off-site information provided by the proliferation of new renewable plants stimulates the interests in the spatiotemporal forecasting. In this paper, an attention temporal convolutional network, which is built on stacked dilated causal convolutional networks and attention mechanisms, is proposed to perform the ultra-short-term spatiotemporal forecasting of renewable resources. Compared with the existing spatiotemporal forecasting methods, the presented model needs no domain knowledge and can be applied to different forecasting tasks such as solar generation and wind speed forecasting. Here, the attention mechanism improves the interpretability. The algorithm can be used to produce both point and probabilistic forecasts. Numerical results on the data sets from National Renewable Energy Laboratory show superior performance over five baselines, in terms of skill scores. Compared with the baselines, the average improvements of accuracy introduced by the proposed method for the point and probabilistic forecasting are 15.08% and 15.85%, respectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Efficient Generalized Boundary Detection Using a Sliding Information Distance

In this work, we present a general machine learning algorithm for boundary detection within general signals based on an efficient, accurate, and robust approximation of the universal normalized information distance. Our approach uses an adaptive sliding information distance (SLID) combined with a wavelet-based approach for peak identification to locate the boundaries. Special emphasis is placed on developing an adaptive formulation of SLID to handle general signals with multiple unknown and/or drifting section lengths. Although specialized algorithms may outperform SLID when domain knowledge is available, these algorithms are limited to specific applications and do not generalize. SLID excels in these cases. We demonstrate the versatility and efficacy of SLID on a variety of signal types, including synthetically generated sequences of tokens, binary executables for reverse engineering applications, and time series of seismic events.

42 ENGINEERING↗

Correlating Time-Resolved Pressure Measurements With Rim Sealing Effectiveness for Real-Time Turbine Health Monitoring

Purge flow is bled from the upstream compressor and supplied to the under-platform region to prevent hot main gas path ingress that damages vulnerable under-platform hardware components. A majority of turbine rim seal research has sought to identify methods of improving sealing technologies and understanding the physical mechanisms that drive ingress. While these studies directly support the design and analysis of advanced rim seal geometries and purge flow systems, the studies are limited in their applicability to real-time monitoring required for condition-based operation and maintenance. As operational hours increase for in-service engines, this lack of rim seal performance feedback results in progressive degradation of sealing effectiveness, thereby leading to reduced hardware life. To address this need for rim seal performance monitoring, this study utilizes measurements from a one-stage turbine research facility operating with true-scale engine hardware at engine-relevant conditions. Time-resolved pressure measurements collected from the rim seal region are regressed with sealing effectiveness through the use of common machine learning techniques to provide real-time feedback of sealing effectiveness. Two modeling approaches are presented that use a single sensor to predict sealing effectiveness accurately over a range of two turbine operating conditions. Here, the results show that an initial purely data-driven model can be further improved using domain knowledge of relevant turbine operations, which yields sealing effectiveness predictions within 3% of measured values.

42 ENGINEERING↗

Correlating Time-Resolved Pressure Measurements With Rim Sealing Effectiveness for Real-Time Turbine Health Monitoring

Purge flow is bled from the upstream compressor and supplied to the under-platform region to prevent hot main gas path ingress that damages vulnerable under-platform hardware components. A majority of turbine rim seal research has sought to identify methods of improving sealing technologies and understanding the physical mechanisms that drive ingress. While these studies directly support the design and analysis of advanced rim seal geometries and purge flow systems, the studies are limited in their applicability to real-time monitoring required for condition-based operation and maintenance. As operational hours increase for in-service engines, this lack of rim seal performance feedback results in progressive degradation of sealing effectiveness, thereby leading to reduced hardware life. To address this need for rim seal performance monitoring, the present study utilizes measurements from a one-stage turbine research facility operating with true-scale engine hardware at engine-relevant conditions. Time-resolved pressure measurements collected from the rim seal region are regressed with sealing effectiveness through the use of common machine learning techniques to provide real-time feedback of sealing effectiveness. Two modelling approaches are presented that use a single sensor to predict sealing effectiveness accurately over a range of two turbine operating conditions. Results show that an initial purely data-driven model can be further improved using domain knowledge of relevant turbine operations, which yields sealing effectiveness predictions within three percent of measured values.

Compressors↗

Power Grid Behavioral Patterns and Risks of Generalization in Applied Machine Learning

Recent years have seen a rich literature of data-driven approaches designed for power grid applications. However, insufficient consideration of domain knowledge can impose a high risk to the practicality of the methods. Specifically, ignoring the grid-specific spatiotemporal patterns (in load, generation, and topology, etc.) can lead to outputting infeasible, unrealizable, or completely meaningless predictions on new inputs. To address this concern, this paper investigates real-world operational data to provide insights into power grid behavioral patterns, including the time-varying topology, load, and generation, as well as the spatial differences (in peak hours, diverse styles) between individual loads and generations. Then based on these observations, we evaluate the generalization risks in some existing ML works caused by ignoring these grid-specific patterns in model design and training.

Li, Shimiao↗

Digital Twin Technology (“Morpheus”) for Optimized Building Operations [SWR-22-74]

The electrification of buildings is an important step to reducing greenhouse gas emissions across all industries. The management of increasingly electrified buildings is a complex pursuit, and there remains a need for cost-effective software capable of handling the computational burden required of such complexity. Through a partnership with Dallas Fort Worth (DFW) Airport, researchers at NREL have developed a digital twin modeling framework to optimize building operations, called Morpheus. Pairing predictive control with automatic fault detection and diagnostics, Morpheus decreases energy expenditures, costs, and faults for large facilities. Additionally, Morpheus employs artificial intelligence to continuously improve its performance using information provided by sensor systems, human experts with deep industry domain knowledge, and even from other similar machines or fleets of machines. Coupling this novel energy-management software with other digital twins, such as NREL’s Athena software for mobility operations, enables robust decision-making for asset and space management. The implementation of Morpheus at DFW has resulted in significantly improved HVAC system operations and reduced both peak power and overall energy consumption. This enhanced functionality comes at a more affordable price than previously developed digital twins and can be customized for other facilities’ geometries to provide optimal, individualized control of a facility’s energy consumption.

Chinde, Venkatesh↗

EXPLAINABLE AND TRUSTWORTHY DIAGNOSTICS ACHIEVABLE THROUGH PROCESS-BASED AUTOMATED REASONING

An approach has been developed that incorporates domain knowledge to obtain a more explainable and trustworthy equipment health monitoring diagnosis than might otherwise be obtained from a purely data-driven method. Physics-based models serve to constrain the realizable solution space and render a more trusted diagnosis. An automated reasoning algorithm performs backward chaining to infer a diagnosis that is consistent with logic statements that have been evaluated as true. This diagnosis is made explainable to an operator by providing the forward chaining path that elucidates for inspection and validity testing those truths implied by the diagnosis.

automated reasoning↗

graphenv: a Python library for reinforcement learning on graph search spaces

Many important and challenging problems in combinatorial optimization (CO) can be expressed as graph search problems, in which graph vertices represent full or partial solutions and edges represent decisions that connect them. Graph structure not only introduces strong relational inductive biases for learning (Battaglia et al., 2018) - in this context, by providing a way to explicitly model the value of transitioning (along edges) between one search state (vertex) and the next - but lends itself to problems both with and without clearly defined algebraic structure. For example, classic CO problems on graphs such as the Traveling Salesman Problem (TSP) can be expressed as either pure graph search or integer programs. Other problems, however, such as molecular optimization, do no have concise algebraic formulations and yet are readily implemented as a graph search (V. et al., 2022; Zhou et al., 2019). Such "model-free" problems constitute a large fraction of modern reinforcement learning (RL) research owing to the fact that it is often much easier to write a forward simulation that expresses all of the state transitions and rewards, than to write down the precise mathematical expression of the full optimization problem. In the case of molecular optimization, for example, one can use domain knowledge alongside existing software libraries to model the effect of adding a single bond or atom to an existing but incomplete molecule, and let the RL algorithm build a model of how good a given decision is by "experiencing" the simulated environment many times through. In contrast, a model-based mathematical formulation that fully expresses all the chemical and physical constraints is intractable. In recent years, RL has emerged as an effective paradigm for optimizing searches over graphs and led to state-of-the-art heuristics for games like Go and chess, as well as for classical CO problems such as the TSP. This combination of graph search and RL, while powerful, requires non-trivial software to execute, especially when combining advanced state representations such as Graph Neural Networks (GNN) with scalable RL algorithms.

97 MATHEMATICS AND COMPUTING↗

Advanced stationary and nonstationary kernel designs for domain-aware Gaussian processes

Gaussian process regression is a widely-applied method for function approximation and uncertainty quantification. The technique has gained popularity recently in the machine learning community due to its robustness and interpretability. The mathematical methods we discuss in this paper are an extension of the Gaussian-process framework. We are proposing advanced kernel designs that only allow for functions with certain desirable characteristics to be elements of the reproducing kernel Hilbert space (RKHS) that underlies all kernel methods and serves as the sample space for Gaussian process regression. These desirable characteristics reflect the underlying physics; two obvious examples are symmetry and periodicity constraints. In addition, non-stationary kernel designs can be defined in the same framework to yield flexible multi-task Gaussian processes. We will show the impact of advanced kernel designs on Gaussian processes using several synthetic and two scientific data sets. The results of our research show that including domain knowledge, communicated through advanced kernel designs, has a significant impact on the accuracy and relevance of the function approximation.

97 MATHEMATICS AND COMPUTING↗

Development of a Framework for Data Integration, Assimilation, and Learning for Geological Carbon Sequestration (DIAL-GCS) (Final Report)

This project aimed to develop and demonstrate a Data Integration, Assimilation, and Learning framework for geologic carbon sequestration projects (DIAL-GCS). DIAL-GCS is an intelligence monitoring system (IMS) for automating GCS closed-loop management by leveraging recent developments in machine learning technologies, complex event processing (CEP), and reduced-order modeling. The safe and efficient operation of GCS repositories requires integrated monitoring to track the injected CO¬2 as it moves within a storage reservoir. GCS projects are data intensive, as a result of proliferation of digital instrumentation and smart-sensing technologies. GCS projects are also resource intensive, often requiring multidisciplinary teams performing different monitoring, verification, accounting (MVA) tasks throughout the lifecycle of a project to ensure secure containment of injected CO2. The success of GCS thus depends in a large part on our ability to access, assimilate, and analyze heterogeneous data and information sources in a timely manner. This project included a number of meaningful and necessary tasks to transform the human domain knowledge into machine-interpretable rules for automating knowledge extraction and discovery in GCS. The specific technical objectives of the proposed DIAL-GCS project were to develop an ontology-driven GCS data management module for storing, querying, and exchanging GCS data (both historic and live sensor data) from multiple sources and in heterogeneous formats. Incorporate a CEP engine for detecting abnormal situations by seamlessly combining expert knowledge, rule-based reasoning, and machine learning. Enable uncertainty quantification and predictive analytics using a combination of coupled-process modeling, AI/ML methods, and reduced-order modeling, and integrate and demonstrate the system’s capabilities with both real and simulated data. As far as we know, this is one of the first projects aimed to develop intelligent monitoring systems (IMS) targeting the GCS. Under this project, the team had developed a large number of web applications and scientific algorithms that contribute the main theme of intelligent monitoring. The team has published more than a dozen peer reviewed papers and disseminated the research results at multiple technical meetings.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Predictive Data-driven Platform for Subsurface Energy Production

Subsurface energy activities such as unconventional resource recovery, enhanced geothermal energy systems, and geologic carbon storage require fast and reliable methods to account for complex, multiphysical processes in heterogeneous fractured and porous media. Although reservoir simulation is considered the industry standard for simulating these subsurface systems with injection and/or extraction operations, reservoir simulation requires spatio-temporal “Big Data” into the simulation model, which is typically a major challenge during model development and computational phase. In this work, we developed and applied various deep neural network-based approaches to (1) process multiscale image segmentation, (2) generate ensemble members of drainage networks, flow channels, and porous media using deep convolutional generative adversarial network, (3) construct multiple hybrid neural networks such as convolutional LSTM and convolutional neural network-LSTM to develop fast and accurate reduced order models for shale gas extraction, and (4) physics-informed neural network and deep Q-learning for flow and energy production. We hypothesized that physicsbased machine learning/deep learning can overcome the shortcomings of traditional machine learning methods where data-driven models have faltered beyond the data and physical conditions used for training and validation. We improved and developed novel approaches to demonstrate that physics-based ML can allow us to incorporate physical constraints (e.g., scientific domain knowledge) into ML framework. Outcomes of this project will be readily applicable for many energy and national security problems that are particularly defined by multiscale features and network systems.

58 GEOSCIENCES↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine

03 NATURAL GAS↗

Deep Analysis Net with Causal Embedding for Coal-fired Power Plant Fault Detection and Diagnosis (DANCE4CFDD)

Fault detection and diagnosis is critical to power plant operation to ensure attaining high reliability while reducing operation cost. As more renewable power is introduced to the power grid, traditional fossil power plants take on the extra burden of excessive load cycling to compensate the generation variability from renewable power. Such load cycling will pose more reliability challenges to power plant operation. There are a number of challenges faced by today’s asset health management system in coal- fired (or gas) power plants: 1) high-dimensional nonlinear interaction among multiple time series measurements; 2) high measurement variance induced by operational conditions/modes; 3) variation among asset types and plant configurations; and 4) a small number of faulty events to learn from. To cope with these challenges, today’s fielded asset health management systems rely heavily on manual efforts from domain experts and hand-crafted features or rules based on domain knowledge. Despite its role in plant reliability, such a practice is costly and hinders its scalability and sustainability, particularly when a plant undergoes modifications. The objective of this project is to develop a novel end-to-end AI learning system that is trainable (i.e., the AI representation of a complex system behavior can be directly learned from properly labeled data) for accurate fault detection and root cause analysis. The ability to create a fault detection model directly from time series could alleviate the efforts associated with today’s asset management solution development. In the course of this project, we have achieved the following: Created an AI model development environment incorporating state-of-the-art neural network architectures for rapid model development and evaluation; Developed novel learning strategies for training of fault detection model; Developed special-purpose neural network architecture embedded with variable association graph aiming for better interpretability; Developed a learning strategy to leverage a small number of faulty events for enhanced fault detection capability; Conducted detailed experimental study based on public benchmark datasets and demonstrated the effectiveness of the proposed solution; and Validated the developed system with data from both a coal-fired plant boiler dynamic simulation model and real-world coal-fired power plant covering multiple asset and fault types. Overall, the project attained a technology readiness level of TRL 5 from TRL 2 at the beginning of the project.

20 FOSSIL-FUELED POWER PLANTS↗

Predictive Battery Lifetime Modeling at NREL [Slides]

Battery lifetime models are used to extrapolate data from accelerated aging tests to simulate degradation in real-world applications such as electric vehicles and battery energy storage systems. Methods developed at NREL utilize both expert domain-knowledge and machine-learning to identify models, using statistical methods such as cross-validation and bootstrap resampling to interrogate model performance and quantify uncertainty. These models can be utilized in systems level simulations to predict battery performance or technoeconomic models to estimate the lifetime cost of battery systems.

25 ENERGY STORAGE↗

Integrated Design of Ultradurable, Low CO 2 Alternative Binder Systems via Machine Learning

This ARPA-E project developed a machine learning tool to use in formulation design of cementitious binders for concrete having 50% less embodied CO 2 and possessing twice the durability compared to concrete based on ordinary portland cement (OPC) binders. The technical focus was on limestone/calcined clay cement (LC3), the leading replacement for OPC. Here, hierarchical machine learning (HML) was used to model the flowability, set time, strength, and durability of LC3 concrete. This methodology identifies latent variables derived from domain knowledge and empirical models that develop an accurate model for a response surface from small datasets. For the flowability metric, particle packing was a dominant factor, while strength and durability were both strongly determined by the fraction of metakaolin and the water:solids ratio. Under constraints of water:binder ratio, material performance metrics, embodied CO 2 , and cost per tonne of OPC, multi-objective optimization was used to design binders parameterized by the mineral composition replacing OPC, particle size distributions, and water:solids ratio. The trained algorithm was able to predict multiple mixes met these performance criteria, and experimental testing validated the predictions. The model demonstrated here is relevant for North America, where pure kaolin deposits are found broadly. The approach is being taken forward into commercial application by Ansatz AI, a materials informatics company founded by PI Washburn and co-PI Poczos. Through collaborations with the cement and concrete industry, and funding from SBIR programs, a commercial software will be developed in future research.

36 MATERIALS SCIENCE↗

Mirostructure Characterization of Friction Consolidated Copper-Nickel using a Machine Learning Approach: Developing Process to Microstructure Associations

Friction consolidation (FC) is a solid phase processing approach where discrete material forms such as powders, chips, nuggets, etc. are densified via shear deformation. The precursors are placed in a billet container and brought in contact with a rotating tool that applying the desirable amount of normal force. Under the combined action of the rotation and normal pressure, the discrete precursor is consolidated through porosity reduction and shear deformation. FC is increasingly being studied as an attractive approach to manufacturing fully dense parts from powder forms owing to its ability to mix, alloy and consolidate difficult-to-process precursors in minimal number of process steps. Material consolidation and deformation in shear consolidation processes have been studied extensively previously for different material combinations previously. However, despite the extensive research in this area, understanding of the mechanistic processes in pore consolidation, deformation-induced mixing and material solubility during FC is still evolving. Material development using solid phase processing approaches such as FC is often performed based on research experience/education, which can be biased. Conventional analysis and simulation tools in this area tend to be successful only when material thermodynamic pathways and microstructural evolution sequences resulting from processing are clearly defined or known. They are not as effective for emerging advanced manufacturing technologies where material evolution pathways are not well established. The ability to predict optimal process parameters based on material chemistry and bulk properties is essential to accelerate materials design and processing, as are an understanding of the relevant structure-processing-property relationships. These structure-processing-property-performance relationships are at the core of materials science research. Microstructure characterization provides the link to these four core areas, often through visualizing material microstructure using imaging techniques. However, linking microstructure image data (i.e., micrographs) to variables of interest (e.g., processing parameters, material chemistry) in a reproducible, generalizable, and quantitative manner is a significant challenge. Typically, quantitatively linking image data to processing history relies on significant domain knowledge and manual or subject matter expert (SME)-heuristic based image analysis. Such an approach to image analysis has the potential to be biased, inefficient, and difficult to replicate.

36 MATERIALS SCIENCE↗

NuGraph2: A Graph Neural Network for Neutrino Event Reconstruction

Neutrino experiments are set to probe some of the most important open questions in physics, from CP violation and the nature of dark matter. The technology of choice for many of these experiments is the liquid argon time projection chamber (LArTPC). In current LArTPC experiments, reconstruction performance often represents a limiting factor for the sensitivity. New developments are therefore needed to unlock the full potential of LArTPC experiments. NuGraph2 is a state of the art Graph Neural Network for reconstruction of data in LArTPC experiments. NuGraph2 utilizes a heterogeneous graph structure, with separate subgraphs of 2D nodes (hits in each plane) connected across planes via 3D nodes (space points). The model provides a consistent description of the neutrino interaction across all planes. NuGraph2 is a multi-purpose network, with a common message-passing attention engine connected to multiple decoders with different classification or regression tasks. These include the classification of detector hits according to the particle type that produced them (semantic segmentation) and the separation of hits from the neutrino interaction from hits due to noise or cosmic-ray background. Additional decoders are being developed, performing tasks such as the regression of the neutrino interaction vertex position. Performance results will be presented based on publicly available samples from MicroBooNE. These include both physics performance metrics, achieving 95% accuracy for semantic segmentation and 98% classification of neutrino hits, as well as computational metrics for training and for inference on CPU or GPU. The status of the NuGraph integration in the LArSoft software framework will be presented, as well as initial studies about model interpretability and injection of domain knowledge.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Augmented Human Analysis (AHA)

Radio frequency (RF) signal monitoring generally emphasizes intentionally generated signals, such as WiFi, Bluetooth, or cellular transmissions. However, electronic devices also produce unintended radiated emissions (UREs), which could also be useful in RF spectrum analysis. In either case, deriving intelligence from RF signals is typically a human-intensive process requiring significant domain knowledge. In the Augmented Human Analysis (AHA) project, we investigate the utility of dimensionally aligned signal projection (DASP) and machine learning (ML) algorithms for accelerating RF analysis workflows. We find that while DASP algorithms can indeed highlight signal characteristics relevant for classification tasks, the choice of algorithmic hyperparameters greatly affects performance. To address this challenge, we evaluate the quality of DASP outputs using the silhouette score, which measures how well data points cluster; high silhouette scores indicate good clustering, and thus good hyperparameter values. This approach is critical for machine learning pipelines as the DASP parameters cannot be directly optimized during model training. By identifying good DASP parameters, and thus good DASP outputs, as a preprocessing step, we can decrease the amount of effort required for downstream ML model training. We demonstrate our workflow using a dataset of UREs from common household devices, showing that even without the aid of ML, proper selection of DASP parameters enables clustering by device type.

42 ENGINEERING↗