Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

174 records · Page 10

Multi Sensor Approach to Address Sustainable Development

The main objectives of Earth Science research are many folds: to understand how does this planet operates, can we model her operation and eventually develop the capability to predict such changes. However, the underlying goals of this work are to eventually serve the humanity in providing societal benefits. This requires continuous, and detailed observations from many sources in situ, airborne and space. By and large, the space observations are the way to comprehend the global phenomena across continental boundaries and provide credible boundary conditions for the mesoscale studies. This requires a multiple sensors, look angles and measurements over the same spot in accurately solving many problems that may be related to air quality, multi hazard disasters, public health, hydrology and more. Therefore, there are many ways to address these issues and develop joint implementation, data sharing and operating strategies for the benefit of the world community. This is because for large geographical areas or regions and a diverse population, some sound observations, scientific facts and analytical models must support the decision making. This is crucial for the sustainability of vital resources of the world and at the same time to protect the inhabitants, endangered species and the ecology. Needless to say, there is no single sensor, which can answer all such questions effectively. Due to multi sensor approach, it puts a tremendous burden on any single implementing entity in terms of information, knowledge, budget, technology readiness and computational power. And, more importantly, the health of planet Earth and its ability to sustain life is not governed by a single country, but in reality, is everyone's business on this planet. Therefore, with this notion, it is becoming an impractical problem by any single organization/country to bear this colossal responsibility. So far, each developed country within their means has proceeded along satisfactorily in implementing their Earth observing needs but it has left a big void in the developing world who have very limited resources to invest in the space measurements. This paper gives some serious thoughts in what options are there in undertaking this tremendous challenge. The problem is multi-dimensional in terms of budget, technology availability, environmental legislations, public awareness, and communication limitations. Some of these issues are introduced, discussed and possible implementation strategies are provided in this paper to move out of this predicament. A strong emphasis is placed on international cooperation and collaboration to see a collective benefit for this effort

Habib, Shahid↗

Big Data Analysis and Technical Review of Regeneration for Carbon Capture Processes

Carbon capture remains an integral technology to mitigate pollution from one of the most prevalent greenhouse gases. CO 2 desorption/absorbent regeneration for both solid- and liquid-based systems is widely recognized as an energy-intensive and costly process operation. Consequently, tremendous work was devoted towards developing new absorbents and regeneration processes to promote their economic feasibility for extensive implementation. In this review, we broadly and deeply review more than 10,000 papers and extract the hidden trends of carbon capture and absorbents regeneration in the past few decades, using a novel data-mining analysis technique. We comprehensively analyzed an array of recent absorbent regeneration methods utilized in post-combustion, pre-combustion, carbon capture from industrial point sources, and direct air carbon capture, with an emphasis on sorbent and solvent-based techniques. In conclusion, advanced regeneration methods in these techniques were illustrated and discussed, followed by recommendations for further research efforts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Explainable and trustworthy artificial intelligence for correctable modeling in chemical sciences

Data science has primarily focused on big data, but for many physics, chemistry, and engineering applications, data are often small, correlated and, thus, low dimensional, and sourced from both computations and experiments with various levels of noise. Typical statistics and machine learning methods do not work for these cases. Expert knowledge is essential, but a systematic framework for incorporating it into physics-based models under uncertainty is lacking. Here, we develop a mathematical and computational framework for probabilistic artificial intelligence (AI)–based predictive modeling combining data, expert knowledge, multiscale models, and information theory through uncertainty quantification and probabilistic graphical models (PGMs). We apply PGMs to chemistry specifically and develop predictive guarantees for PGMs generally. Our proposed framework, combining AI and uncertainty quantification, provides explainable results leading to correctable and, eventually, trustworthy models. The proposed framework is demonstrated on a microkinetic model of the oxygen reduction reaction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Thermochemical Data Fusion Using Graph Representation Learning

Large databases are required for “Big Data” applications in catalysis and materials science. Thermochemical databases can be created by combining data from various sources and by correcting low-fidelity datasets to higher accuracy with minimal computation. To achieve this “data fusion”, thermochemical quantities of interest, calculated at various levels of density functional theory (DFT), need to be mapped to the same, high levels of theory. In this work, a graph theoretical, statistical framework is proposed for such tasks. Subgraph frequencies are shown to provide a natural representation for learning these fusion maps. The maps are linear and are learnt with automated descriptor selection. Using a dataset of as few as ~1% from the QM9 database of 133,885 molecules, these models can predict multiple thermochemical quantities at a higher level of theory with an accuracy of 1 kcal/mol. Here, the method is explainable, generalizable, and provides a diagnostic tool for outlier identification

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multi-fidelity information fusion with concatenated neural networks

Recently, computational modeling has shifted towards the use of statistical inference, deep learning, and other data-driven modeling frameworks. Although this shift in modeling holds promise in many applications like design optimization and real-time control by lowering the computational burden, training deep learning models needs a huge amount of data. This big data is not always available for scientific problems and leads to poorly generalizable data-driven models. This gap can be furnished by leveraging information from physics-based models. Exploiting prior knowledge about the problem at hand, this study puts forth a physics-guided machine learning (PGML) approach to build more tailored, effective, and efficient surrogate models. For our analysis, without losing its generalizability and modularity, we focus on the development of predictive models for laminar and turbulent boundary layer flows. In particular, we combine the self-similarity solution and power-law velocity profile (low-fidelity models) with the noisy data obtained either from experiments or computational fluid dynamics simulations (high-fidelity models) through a concatenated neural network. We illustrate how the knowledge from these simplified models results in reducing uncertainties associated with deep learning models applied to boundary layer flow prediction problems. The proposed multi-fidelity information fusion framework produces physically consistent models that attempt to achieve better generalization than data-driven models obtained purely based on data. While we demonstrate our framework for a problem relevant to fluid mechanics, its workflow and principles can be adopted for many scientific problems where empirical, analytical, or simplified models are prevalent. In line with grand demands in novel PGML principles, this work builds a bridge between extensive physics-based theories and data-driven modeling paradigms and paves the way for using hybrid physics and machine learning modeling approaches for next-generation digital twin technologies.

42 ENGINEERING↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗

A Vision for Coupling Operation of US Fusion Facilities with HPC Systems and the Implications for Workflows and Data Management

The operation of large US Department of Energy (DOE) research facilities, like the DIII-D National Fusion Facility, results in the collection of complex multi-dimensional scientific datasets, both experimental and model-generated. In the future, it is envisioned that integrated data analysis coupled with large-scale high performance computing (HPC) simulations will be used to improve experimental planning and operation. Practically, massive data sets from these simulations provide the physics basis for generation of both reduced semi-analytic and machine-learning-based models. Storage of both HPC simulation datasets (generated from US DOE leadership computing facilities) and experimental datasets presents significant challenges. In this paper, we present a vision for a DOE-wide data management workflow that integrates US DOE fusion facilities with leadership computing facilities. Data persistence and long-term availability beyond the length of allocated projects is essential, particularly for verification and recalibration of artificial intelligence and machine learning (AI/ML) models. Because these data sets are often generated and shared among hundreds of users across multiple leadership computing facility centers, they would benefit from cross-platform accessibility, persistent identifiers (e.g. DOI, or digital object identifier), and provenance tracking. Here, the ability to handle different data access patterns suggests that a combination of low cost, high latency (e.g. for storing ML training sets) and high cost, low latency systems (e.g. for real-time, integrated machine control feedback) may be needed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Automated exploitation of the big configuration space of large adsorbates on transition metals reveals chemistry feasibility

Mechanistic understanding of large molecule conversion and the discovery of suitable heterogeneous catalysts have been lagging due to the combinatorial inventory of intermediates and the inability of humans to enumerate all structures. Here, we introduce an automated framework to predict stable configurations on transition metal surfaces and demonstrate its validity for adsorbates with up to 6 carbon and oxygen atoms on 11 metals, enabling the exploration of ~10 8 potential configurations. It combines a graph enumeration platform, force field, multi-fidelity DFT calculations, and first-principles trained machine learning. Clusters in the data reveal groups of catalysts stabilizing different structures and expose selective catalysts for showcase transformations, such as the ethylene epoxidation on Ag and Cu and the lack of C-C scission chemistry on Au. Deviations from the commonly assumed atom valency rule of small adsorbates are also manifested. This library can be leveraged to identify catalysts for converting large molecules computationally.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Analytics-at-scale of Sensor Data for Digital Monitoring in Nuclear Plants (3 rd Annual Report)

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This report primarily focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 ± 0.014, 0.0026 ± 0, and 0.063 ± 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure. This report summarizes the Fiscal Year 2021 research progress encompassing the (1) data cleaning and feature selection necessary for ML applications; (2) development of short-term forecasting models to predict future plant process parameters for both single and multiple time steps ahead; and (3) validation of the feature selection methods and short-term forecasting models given new data from different systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Oxidation and pyrolysis of methyl propyl ether

The ignition, oxidation, and pyrolysis chemistry of methyl propyl ether (MPE) was probed experimentally at several different conditions, and a comprehensive chemical kinetic model was constructed to help understand the observations, with many of the key parameters computed using quantum chemistry and transition state theory. Experiments were carried out in a shock tube measuring time variation of CO concentrations, in a flow tube measuring product concentrations, and in a rapid compression machine (RCM) measuring ignition delay times. The detailed reaction mechanism was constructed using the Reaction Mechanism Generator software. Sensitivity and flux analyses were used to identify key rate and thermochemical parameters, which were then computed using quantum chemistry to improve the mechanism. Validation of the final model against the 1-20 bar 600-1500 K experimental data is presented with a discussion of the kinetics. The model is in excellent agreement with most of the shock tube and RCM data. Strong non-monotonic variation in conversion and product distribution is observed in the flow-tube experiments as the temperature is increased, and unusually strong pressure dependence and significant heat release during the compression stroke is observed in the RCM experiments. These observations are largely explained by a close competition between radical decomposition and addition to O2 at different sites in MPE; this causes small shifts in conditions to lead to big shifts in the dominant reaction pathways. The validated mechanism was used to study the chemistry occurring during ignition in a diesel engine, simulated using Ignition Quality Test (IQT) conditions. At the IQT conditions, where the MPE concentration is higher, bimolecular reactions of peroxy radicals are much more important than in the RCM.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

EVA Performance Prediction

Astronaut physical performance capabilities in micro gravity EV A or on planetary surfaces when encumbered by a life support suit and debilitated by a long exposure to micro gravity will be less than unencumbered pre flight capabilities. The big question addressed by human factors engineers is: what can the astronaut be expected to do on EVA or when we arrive at a planetary surface? A second question is: what aids to performance will be needed to enhance the human physical capability? These questions are important for a number of reasons. First it is necessary to carry out accurate planning of human physical demands to ensure that time and energy critical tasks can be carried out with confidence. Second it is important that the crew members (and their ground or planetary base monitors) have a realistic picture of their own capabilities, as excessive fatigue can lead to catastrophic failure. Third it is important to design appropriate equipment to enhance human sensory capabilities, locomotion, materials handling and manipulation. The evidence from physiological research points to musculoskeletal, cardiovascular and neurovestibular degradation during long duration exposure to micro gravity . The evidence from the biomechanics laboratory (and the Neutral Buoyancy Laboratory) points to a reduction in range of motion, strength and stamina when encumbered by a pressurized suit. The evidence from a long history of EVAs is that crewmembers are indeed restricted in their physical capabilities. There is a wealth of evidence in the literature on the causes and effects of degraded human performance in the laboratory, in sports and athletics, in industry and in other physically demanding jobs. One approach to this challenge is through biomechanical and performance modeling. Such models must be based on thorough task analysis, reliable human performance data from controlled studies, and functional extrapolations validated in analog contexts. The task analyses currently carried out for EVA activities are based more on extensive domain experience than any formal analytic structure. Conversely, physical task analysis for industrial and structured evidence from training and EV A contexts. Again on earth there is considerable evidence of human performance degradation due to encumbrance and fatigue. These industrial models generally take the form of a discounting equation. The development of performance estimates for space operations, such as timeline predictions for EVA is generally based on specific input from training activity, for example in the NBL or KC135. uniformed services tasks on earth are much more formalized. Human performance data in the space context has two sources: first there is the micro analysis of performance in structured tasks by the space physiology community and second there is the less structured evidence from training and EV A contexts.

Peacock, Brian↗

Securing the Modern Grid: Federal Investments, Digitization, and Supply Chain Strategy

Across the United States (U.S.) grid expansion and modernization is underway, paving the way for accelerated load growth and intelligent resource management. Digitization of the grid is supported by several state and federal programs, providing support for utilities installing advanced metering infrastructure (AMI), AI-powered analytics systems, battery energy storage systems (BESS), and distributed energy resource management systems (DERMS) to transform the grid from a one-way power delivery system into an intelligent, responsive network that will enable faster load growth and power expansion of data centers for advanced artificial intelligence (AI) applications. The digital transformation of America's grid presents opportunity for increased efficiency and resiliency but also introduces new digital risks that require careful management. Digital equipment often contains several vulnerabilities such as unencrypted communication protocols, and persistent remote access capabilities that could be exploited to manipulate device settings, coordinate service disruptions, or inject false data into grid operations. These digital risks become particularly important as the grid must rapidly scale to support AI-driven data centers, which the administration has identified as essential for maintaining U.S. technological leadership and economic competitiveness. These vulnerabilities are compounded by supply chain realities: Chinese manufacturers currently produce 70-90% of essential grid components including inverters, batteries, and control systems, with the U.S. lacking domestic manufacturing capacity for critical assets like extra-high voltage transformers. Recent federal legislation has established Foreign Entity of Concern (FEOC) restrictions to address these risks, requiring projects to achieve escalating thresholds of non-FEOC content to receive tax credits while utilities work to expand sourcing channels for their supply chains and strengthen security measures. These restrictions arrive precisely when utilities face unprecedented electricity demand growth driven by the rapid growth in data centers, creating a considerable challenge: rapidly expanding infrastructure while navigating complex compliance requirements while lacking viable alternatives for many critical components. Idaho National Laboratory (INL) and its partners have developed practical approaches to help utilities navigate these intersecting challenges as they leverage federal investment to strengthen and grow the grid. These solutions include Cyber-Informed Engineering (CIE) principles that build resilience directly into systems, the Cirrus tool for secure cloud migration, and enhanced procurement guidance that embeds security requirements throughout equipment lifecycles. Federal initiatives, such as the Technical Assistance for Digital Assurance (TADA) project, provide direct support to utilities implementing these approaches while facilitating knowledge sharing across the industry. While these tools and frameworks cannot eliminate all risks inherent in foreign supply chain dependencies, they offer pragmatic pathways for strengthening security posture without sacrificing the deployment momentum essential to meeting surging electricity demand. Ultimately, securing America's digital energy infrastructure demands dedicated coordination across multiple fronts: building domestic supply chains, implementing robust digital assurance practices, and maintaining the aggressive modernization timeline necessary for reliability, resilience, and energy independence.

24 POWER TRANSMISSION AND DISTRIBUTION↗