Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Data Science and Computation for Rapid and Dynamic Compression Experiment Workflows at Experimental Facilities, September 8-11, 2020. Workshop Report

The application of high pressure to materials has enabled discoveries in scientific fields such as planetary science, materials science, and materials synthesis. Recent advances in X-ray user light sources and other facilities, co-location and integration of user facilities with high-pressure drivers, availability of high-performance computing (HPC) platforms, and the development of new data science techniques have created opportunities for, and challenges in, advancing data analytics for rapid and dynamic compression experiments. To address these challenges, harness the emerging technology now available, and expedite scientific discovery, Los Alamos National Laboratory (LANL) hosted a virtual workshop entitled “Data Science and Computation for Rapid and Dynamic Compression Workflows at Experimental Facilities” from September 8 to 11, 2020. The workshop included 95 registered scientists and analytics experts from 15 universities, 9 United States (US) national laboratories, 5 US and European X-ray light sources, neutron sources such as the Los Alamos Neutron Science Center (LANSCE), other big science facilities such as the National Ignition Facility (NIF), and an industry representative. The workshop included 31 invited talks and 4 lightning talks by students and postdocs.

36 MATERIALS SCIENCE↗

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information

Detecting and anticipating global proliferation expertise and capability evolution from unstructured, noisy, and incomplete public data streams is a highly desired, but extremely challenging task. Here, in this article, we present our pioneering data-driven approach to support the non-proliferation mission to detect and explain the evolution of proliferation expertise and capability development globally from terabytes of publicly available information (PAI), focusing on our knowledge extraction pipeline and descriptive analytics. We first discuss how we fuse nine open-source data streams, including multilingual data, to convert 4 TB of unstructured data to structured knowledge and encode dynamically evolving proliferation expertise representations—content and context graphs. For this, we rely on natural language processing (NLP) and deep learning (DL) models to perform information extraction, topic modeling, and distributed text representation (aka embedding) learning. We then present interactive, usable, and explainable descriptive analytics to refine domain knowledge and present it in a human-understandable form. Finally, we introduce future work avenues that will leverage our dynamic knowledge representations and descriptive analytics to enable predictive and prescriptive inferences to achieve real-time domain understanding and contextual reasoning about global proliferation expertise and capability evolution.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Exploration of Domain Aware Machine Learning for Grid Analytics: Transfer-Learnt Energy Models to Assist Buildings Control with Sparse Field Data

Buildings are a primary consumer of energy in the United States and are also increasingly being perceived as providers of grid services such as load shifting, shedding and modulation. High fidelity models of building energy consumption are needed to set appropriate baselines for measurement and verification (M&V) of controllers designed for energy efficient operation of buildings and to enable buildings to provide grid services via. participation in demand response programs. State-of-the-art building energy modeling techniques either rely on Physics based models, or extensive instrumentation of the building envelope to gather “big” data to train machine learning based models such as deep neural networks. While Physics based models are often limited by their accuracy, it is not always feasible to gather a significant amount of field data required to train machine learning based models with sufficient accuracy. In this paper, we explore the use of transfer learning-based strategies to address unsatisfactory accuracy of models for estimating building energy consumption when available field data for training is sparse or of unacceptable quality. In particular, we transfer knowledge in the form of data and parameters, from Physics based simulation frameworks to the field to improve the model accuracy, thus resulting in a Physics-informed Machine Learning framework. We evaluated the efficacy of our approach on field data collected from six commercial buildings and our results indicate that the proposed transfer learning based models provide comparative (and in some cases better) accuracy than state-of-the-art machine learning and deep learning solutions, with just one month of field data.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Emerging Strategies for Modifying Lignin Chemistry to Enhance Biological Lignin Valorization

Biological lignin valorization represents a promising approach contributing to sustainable and economic biorefineries. Here, the low level of valuable lignin–derived products remains a major challenge hindering the implementation of microbial lignin conversion. Lignin's properties play a significant role in determining the efficiency of lignin bioconversion. To date, despite significant progress in the development of biomass pretreatment, lignin fractionation, and fermentation over the last few decades, little efforts have gone into identifying the ideal lignin substrates for an efficient microbial metabolism. In this Minireview, emerging and state–of–the–art strategies for biomass pretreatment and lignin fractionation are summarized to elaborate their roles in modifying lignin structure for bioconversion. Fermentation strategies aimed at enhancing lignin depolymerization for microbial utilization are systematically reviewed as well. With an improved understanding of the ideal lignin structure elucidated by comprehensive metabolic pathways and/or big data analysis, modifying lignin chemistry could be more directional and effective. Ultimately, together with the progress of fermentation process optimization, biological lignin valorization will become more competitive in biorefineries.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Discovery of Signatures, Anomalies, and Precursors in Synchrophasor Data with Matrix Profile and Deep Recurrent Neural Networks (Final Project Report)

The widespread deployment of phasor measurement unit (PMU) across the U.S. together with the burgeoning machine learning technology made it possible to develop data-driven PMU data analytics to improve grid security and reliability in a more insightful and effective manner. Although PMU applications have been explored for over a decade, the representative PMU usage is limited to the bulk power system monitoring mainly due to the data integrity issues associated with PMUs (typically missing, fragmented, and wrongly amplified data). To forge a breakthrough on this stalemate and embrace PMUs for power system control and protection as well, we applied various advanced machine learning and big data analysis technology to the power system event detection and classification as the first step toward the power system control and protection pertaining to grid security enhancement.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Water narratives in local newspapers within the United States

Sustainable use of water resources continues to be a challenge across the globe. This is in part due to the complex set of physical and social behaviors that interact to influence water management from local to global scales. Analyses of water resources have been conducted using a variety of techniques, including qualitative evaluations of media narratives. This study aims to augment these methods by leveraging computational and quantitative techniques from the social sciences focused on text analyses. Specifically, we use natural language processing methods to investigate a large corpus (approx. 1.8M) of newspaper articles spanning approximately 35 years (1982–2017) for insights into human-nature interactions with water. Focusing on local and regional United States publications, our analysis demonstrates important dynamics in water-related dialogue about drinking water and pollution to other critical infrastructures, such as energy, across different parts of the country. Our assessment, which looks at water as a system, also highlights key actors and sentiments surrounding water. Extending these analytical methods could help us further improve our understanding of the complex roles of water in current society that should be considered in emerging activities to mitigate and respond to resource conflicts and climate change.

54 ENVIRONMENTAL SCIENCES↗

Signatures of primordial energy injection from axion strings

Axion strings are horizon-size topological defects that may be produced in the early Universe. Ultralight axion-like particles may form strings that persist to temperatures below that of big bang nucleosynthesis. Such strings have been considered previously as sources of gravitational waves and cosmic microwave background (CMB) polarization rotation. In this work we show, through analytic arguments and dedicated adaptive mesh refinement cosmological simulations, that axion strings deposit a subdominant fraction of their energy into high-energy Standard Model (SM) final states, for example, by the direct production of heavy radial modes that subsequently decay to SM particles. This high-energy SM radiation is absorbed by the primordial plasma, leading to novel signatures in precision big bang nucleosynthesis, the CMB power spectrum, and gamma-ray surveys. In particular, we show that CMB power spectrum data constrains axion strings with decay constants f a ≲ 10 12 GeV , up to model dependence on the ultraviolet completion, for axion masses m a ≲ 10 − 29 eV ; future CMB surveys could find striking evidence of axion strings with lower decay constants. Published by the American Physical Society 2024

79 ASTRONOMY AND ASTROPHYSICS↗

Spatiotemporal Pattern Recognition in the PMU Signals in the WECC system

Phasor measurement unit (PMU) data has been used by multiple power system applications, including state estimation, post event analysis, oscillation detection, model validation, and many others. Still, due to its big data nature and availability to general research institutions, comprehensive understanding of the spatiotemporal patterns and underlying mechanisms are incomplete. This study applies a set of signal processing and machine learning approaches aiming at deciphering the characteristic behaviors of multiple phasor measurement units (PMUs) attributes (e.g., voltage, frequency, rate of change of frequency, phase angle), including their auto-correlation, cross-dependence, similarities and discrepancies across units and temporal scales, and distributions of anomalies and their linkages to potential external factors such as weather events. Data analytics are applied to PMUs from the U.S. Western Electricity Coordinating Council (WECC) system. The PMU measurements, recorded events, outages, and weather extremes are all from real world datasets. The findings from the study and mechanistic understanding of the PMU dynamics help provide guidance on system control or preventing blackouts. The derived metrics can be directly used for adjusting or filtering simulated PMU data used for advanced algorithm development.

Hou, Zhangshuan↗

Lattice Disorder and Oxygen Migration Pathways in Pyrochlore and Defect-Fluorite Oxides

Atomic-scale disorder plays an important role in the chemical and physical properties of oxide materials. The structural flexibility of pyrochlore-type oxides allows for crystal-chemical engineering of these properties. Compositional modification can push pyrochlore oxides toward a disordered defect-fluorite structure with anion Frenkel pair defects that facilitate oxygen migration. The local structure of the long-range average cubic defect-fluorite was recently claimed to consist of randomly arranged orthorhombic weberite-type domains. Here, we show, using low-temperature neutron total-scattering experiments, that this is not the case for Zr-rich defect-fluorites. By analyzing data from the pyrochlore/defect-fluorite Y 2 Sn 2–x Zr x O 7 series using a combination of neutron pair distribution function and big-box modelling, we have differentiated and quantified the relationship between anion sub-lattice disorder and Frenkel defects. These details directly influence the energy landscape for oxygen migration and are crucial for simulations and design of new materials with improved properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Big Data Analysis and Technical Review of Regeneration for Carbon Capture Processes

Carbon capture remains an integral technology to mitigate pollution from one of the most prevalent greenhouse gases. CO 2 desorption/absorbent regeneration for both solid- and liquid-based systems is widely recognized as an energy-intensive and costly process operation. Consequently, tremendous work was devoted towards developing new absorbents and regeneration processes to promote their economic feasibility for extensive implementation. In this review, we broadly and deeply review more than 10,000 papers and extract the hidden trends of carbon capture and absorbents regeneration in the past few decades, using a novel data-mining analysis technique. We comprehensively analyzed an array of recent absorbent regeneration methods utilized in post-combustion, pre-combustion, carbon capture from industrial point sources, and direct air carbon capture, with an emphasis on sorbent and solvent-based techniques. In conclusion, advanced regeneration methods in these techniques were illustrated and discussed, followed by recommendations for further research efforts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Explainable and trustworthy artificial intelligence for correctable modeling in chemical sciences

Data science has primarily focused on big data, but for many physics, chemistry, and engineering applications, data are often small, correlated and, thus, low dimensional, and sourced from both computations and experiments with various levels of noise. Typical statistics and machine learning methods do not work for these cases. Expert knowledge is essential, but a systematic framework for incorporating it into physics-based models under uncertainty is lacking. Here, we develop a mathematical and computational framework for probabilistic artificial intelligence (AI)–based predictive modeling combining data, expert knowledge, multiscale models, and information theory through uncertainty quantification and probabilistic graphical models (PGMs). We apply PGMs to chemistry specifically and develop predictive guarantees for PGMs generally. Our proposed framework, combining AI and uncertainty quantification, provides explainable results leading to correctable and, eventually, trustworthy models. The proposed framework is demonstrated on a microkinetic model of the oxygen reduction reaction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Thermochemical Data Fusion Using Graph Representation Learning

Large databases are required for “Big Data” applications in catalysis and materials science. Thermochemical databases can be created by combining data from various sources and by correcting low-fidelity datasets to higher accuracy with minimal computation. To achieve this “data fusion”, thermochemical quantities of interest, calculated at various levels of density functional theory (DFT), need to be mapped to the same, high levels of theory. In this work, a graph theoretical, statistical framework is proposed for such tasks. Subgraph frequencies are shown to provide a natural representation for learning these fusion maps. The maps are linear and are learnt with automated descriptor selection. Using a dataset of as few as ~1% from the QM9 database of 133,885 molecules, these models can predict multiple thermochemical quantities at a higher level of theory with an accuracy of 1 kcal/mol. Here, the method is explainable, generalizable, and provides a diagnostic tool for outlier identification

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multi-fidelity information fusion with concatenated neural networks

Recently, computational modeling has shifted towards the use of statistical inference, deep learning, and other data-driven modeling frameworks. Although this shift in modeling holds promise in many applications like design optimization and real-time control by lowering the computational burden, training deep learning models needs a huge amount of data. This big data is not always available for scientific problems and leads to poorly generalizable data-driven models. This gap can be furnished by leveraging information from physics-based models. Exploiting prior knowledge about the problem at hand, this study puts forth a physics-guided machine learning (PGML) approach to build more tailored, effective, and efficient surrogate models. For our analysis, without losing its generalizability and modularity, we focus on the development of predictive models for laminar and turbulent boundary layer flows. In particular, we combine the self-similarity solution and power-law velocity profile (low-fidelity models) with the noisy data obtained either from experiments or computational fluid dynamics simulations (high-fidelity models) through a concatenated neural network. We illustrate how the knowledge from these simplified models results in reducing uncertainties associated with deep learning models applied to boundary layer flow prediction problems. The proposed multi-fidelity information fusion framework produces physically consistent models that attempt to achieve better generalization than data-driven models obtained purely based on data. While we demonstrate our framework for a problem relevant to fluid mechanics, its workflow and principles can be adopted for many scientific problems where empirical, analytical, or simplified models are prevalent. In line with grand demands in novel PGML principles, this work builds a bridge between extensive physics-based theories and data-driven modeling paradigms and paves the way for using hybrid physics and machine learning modeling approaches for next-generation digital twin technologies.

42 ENGINEERING↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗

A Vision for Coupling Operation of US Fusion Facilities with HPC Systems and the Implications for Workflows and Data Management

The operation of large US Department of Energy (DOE) research facilities, like the DIII-D National Fusion Facility, results in the collection of complex multi-dimensional scientific datasets, both experimental and model-generated. In the future, it is envisioned that integrated data analysis coupled with large-scale high performance computing (HPC) simulations will be used to improve experimental planning and operation. Practically, massive data sets from these simulations provide the physics basis for generation of both reduced semi-analytic and machine-learning-based models. Storage of both HPC simulation datasets (generated from US DOE leadership computing facilities) and experimental datasets presents significant challenges. In this paper, we present a vision for a DOE-wide data management workflow that integrates US DOE fusion facilities with leadership computing facilities. Data persistence and long-term availability beyond the length of allocated projects is essential, particularly for verification and recalibration of artificial intelligence and machine learning (AI/ML) models. Because these data sets are often generated and shared among hundreds of users across multiple leadership computing facility centers, they would benefit from cross-platform accessibility, persistent identifiers (e.g. DOI, or digital object identifier), and provenance tracking. Here, the ability to handle different data access patterns suggests that a combination of low cost, high latency (e.g. for storing ML training sets) and high cost, low latency systems (e.g. for real-time, integrated machine control feedback) may be needed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Automated exploitation of the big configuration space of large adsorbates on transition metals reveals chemistry feasibility

Mechanistic understanding of large molecule conversion and the discovery of suitable heterogeneous catalysts have been lagging due to the combinatorial inventory of intermediates and the inability of humans to enumerate all structures. Here, we introduce an automated framework to predict stable configurations on transition metal surfaces and demonstrate its validity for adsorbates with up to 6 carbon and oxygen atoms on 11 metals, enabling the exploration of ~10 8 potential configurations. It combines a graph enumeration platform, force field, multi-fidelity DFT calculations, and first-principles trained machine learning. Clusters in the data reveal groups of catalysts stabilizing different structures and expose selective catalysts for showcase transformations, such as the ethylene epoxidation on Ag and Cu and the lack of C-C scission chemistry on Au. Deviations from the commonly assumed atom valency rule of small adsorbates are also manifested. This library can be leveraged to identify catalysts for converting large molecules computationally.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Analytics-at-scale of Sensor Data for Digital Monitoring in Nuclear Plants (3 rd Annual Report)

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This report primarily focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 ± 0.014, 0.0026 ± 0, and 0.063 ± 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure. This report summarizes the Fiscal Year 2021 research progress encompassing the (1) data cleaning and feature selection necessary for ML applications; (2) development of short-term forecasting models to predict future plant process parameters for both single and multiple time steps ahead; and (3) validation of the feature selection methods and short-term forecasting models given new data from different systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Oxidation and pyrolysis of methyl propyl ether

The ignition, oxidation, and pyrolysis chemistry of methyl propyl ether (MPE) was probed experimentally at several different conditions, and a comprehensive chemical kinetic model was constructed to help understand the observations, with many of the key parameters computed using quantum chemistry and transition state theory. Experiments were carried out in a shock tube measuring time variation of CO concentrations, in a flow tube measuring product concentrations, and in a rapid compression machine (RCM) measuring ignition delay times. The detailed reaction mechanism was constructed using the Reaction Mechanism Generator software. Sensitivity and flux analyses were used to identify key rate and thermochemical parameters, which were then computed using quantum chemistry to improve the mechanism. Validation of the final model against the 1-20 bar 600-1500 K experimental data is presented with a discussion of the kinetics. The model is in excellent agreement with most of the shock tube and RCM data. Strong non-monotonic variation in conversion and product distribution is observed in the flow-tube experiments as the temperature is increased, and unusually strong pressure dependence and significant heat release during the compression stroke is observed in the RCM experiments. These observations are largely explained by a close competition between radical decomposition and addition to O2 at different sites in MPE; this causes small shifts in conditions to lead to big shifts in the dominant reaction pathways. The validated mechanism was used to study the chemistry occurring during ignition in a diesel engine, simulated using Ignition Quality Test (IQT) conditions. At the IQT conditions, where the MPE concentration is higher, bimolecular reactions of peroxy radicals are much more important than in the RCM.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗