Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data (Final Report)

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components: (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine.

20 FOSSIL-FUELED POWER PLANTS↗

Parallel I/O Evaluation Techniques and Emerging HPC Workloads: A Perspective

Emerging workloads such as artificial intelligence, big data analytics and complex multi-step workflows alongside future exascale applications are anticipated future HPC workloads, which will result in a more diverse I/O system workload and even less predictable I/O behavior and access patterns. Along with the ever increasing gap between the compute and storage performance capabilities, the in-depth understanding of extreme-scale I/O behavior and the I/O performance modeling and prediction are essential tools of the large-scale I/O evaluation process for addressing the needs of extreme-scale hybrid workloads. In this survey article, we focus on the state-of-the-art of the I/O behavior and performance analysis process for HPC systems in a 5-year time window and identify future research challenges.

Neuwirth, Sarah↗

Strategies for Integrating Deep Learning Surrogate Models with HPC Simulation Applications

The emerging trend of the convergence of high performance computing (HPC), machine learning/deep learning (ML/DL), and big data analytics presents a host of challenges for large-scale computing campaigns that seek best practices to interleave traditional scientific simulation-based workloads with ML/DL models. A portfolio of systematic approaches to incorporate deep learning into modeling and simulation serves a vital need when we support AI for science at a computing facility. In this paper, we evaluate several strategies for deploying deep learning surrogate models in a representative physics application on supercomputers at the Oak Ridge Leadership Computing Facility (OLCF). We discuss a set of recommended deployment architectures and implementation approaches. We analyze and evaluate these alternatives and show their performance and scalability up to 1000 GPUs on two mainstream platforms equipped with different deep learning hardware and software stacks.

Yin, Junqi↗

Topological Optimization with Big Steps

Using persistent homology to guide optimization has emerged as a novel application of topological data analysis. Existing methods treat persistence calculation as a black box and backpropagate gradients only onto the simplices involved in particular pairs. We show how the cycles and chains used in the persistence calculation can be used to prescribe gradients to larger subsets of the domain. In particular, we show that in a special case, which serves as a building block for general losses, the problem can be solved exactly in linear time. This relies on another contribution of this paper, which eliminates the need to examine a factorial number of permutations of simplices with the same value. Here, we present empirical experiments that show the practical benefits of our algorithm: the number of steps required for the optimization is reduced by an order of magnitude.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Big Data Meets Geothermal Exploration (CRADA Final Report)

As part of the Cyclotron Road program, Zanskar Geothermal & Minerals, Inc. investigated the application of micro-earthquake and ambient noise seismology methods to imaging and characterizing the structural characteristics and hydrothermal flux of subsurface faults. Significant advances in what could be resolved were enabled by two major developments in seismology: 1) the availability of large-n arrays of low-cost seismometers, and 2) the availability of increased computational power and semi-automated data reduction algorithms. In tandem, these advances may improve the signal-to-noise ratio and spatial precision of the data collected and enable higher-resolution characterization of subsurface fracture systems and their spatio-temporal evolution. These tools supported efforts to reduce dry-hole risk and to improve wellfield productivity for geothermal resource development. In particular, two applications of these advances were evaluated: 1) fracture-seismic imaging, which was used to detect ambient emissions from fluid-filled fractures, and 2) reservoir tomography, which used information about travel paths, source locations, and source parameters of micro-earthquakes to identify areas of enhanced permeability. Integration of these methods provided guidance for siting wells and served as prior constraints for reservoir models, informing forecasts of power potential and production and injection strategies aimed at minimizing temperature decline and improving overall resource productivity.

15 GEOTHERMAL ENERGY↗

An overview of the operation architectures and energy management system for multiple microgrid clusters

The emerging novel energy infrastructures, such as energy communities, smart building-based microgrids, electric vehicles enabled mobile energy storage units raise the requirements for a more interconnective and interoperable energy system. It leads to a transition from simple and isolated microgrids to relatively large-scale and complex interconnected microgrid systems named multi-microgrid clusters. In order to efficiently, optimally, and flexibly control multi-microgrid clusters, cross-disciplinary technologies such as power electronics, control theory, optimization algorithms, information and communication technologies, cyber-physical, and big-data analysis are needed. This paper introduces an overview of the relevant aspects for multi-microgrids, including the outstanding features, architectures, typical applications, existing control mechanisms, as well as the challenges.

24 POWER TRANSMISSION AND DISTRIBUTION↗

How Bayesian methods can improve R -matrix analyses of data: The example of the d t reaction

The 3 H(d, n) 4 He reaction is of significant interest in nuclear astrophysics and nuclear applications. It is an important, early step in big-bang nucleosynthesis and a key process in nuclear fusion reactors. We use one- and two-level R-matrix approximations to analyze data on the cross section for this reaction at center-of-mass energies below 215 keV. We critically examine the data sets using a Bayesian statistical model that allows for both common-mode and additional point-to-point un- certainties. We use Markov Chain Monte Carlo sampling to evaluate this R-matrix-plus-statistical model and find two-level R-matrix results that are stable with respect to variations in the channel radii. The S factor at 40 keV evaluates to 25.36(19) MeV b (68% credibility interval). We discuss our Bayesian analysis in detail and provide guidance for future applications of Bayesian methods to R-matrix analyses. We also discuss possible paths to further reduction of the S-factor uncertainty.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Electricity use in big area additive manufacturing of fiber-reinforced polymer composites

In recent years, additive manufacturing (AM), especially large-format additive manufacturing (LFAM), has gained momentum in the manufacturing industry. While LFAM offers benefits over conventional manufacturing processes, such as minimizing material waste and providing vast geometric freedom, assessing its sustainability remains challenging due to limited data, particularly on energy consumption. Most existing data pertain to small-scale or desktop AM and are not directly applicable to LFAM. In this study, we conducted real-time measurements of electricity usage for a type of LFAM known as big area additive manufacturing (BAAM), which typically uses fiber-reinforced polymer pellets as feedstock. We collected electricity usage data from fifteen printing jobs over two months in an industrial production setting. These data fill the existing gap and can be reused to enhance the community’s understanding of LFAM electricity usage, support further research, and promote sustainable development in advanced manufacturing technologies.

ecology↗

Transfer-Learnt Energy Models for Predicting Electricity Consumption in Buildings with Limited and Sparse Field Data

Modeling energy consumption is critical for energy-efficient utilization of the electric appliances in a building, smart grid programs (like demand-response), and many other smart home applications. State-of-the-art energy modeling techniques either rely on theoretical models, or extensive instrumentation of the building envelope to gather ``big" data to train a deep neural network. While theoretical models are often limited by their estimation accuracy, it is not always feasible to gather a significant amount of field data. In this paper, we explore transfer learning-based strategies to train much more accurate model for energy estimation when using a sparse field data. We transferred knowledge, in the form of data and parameters, from the simulation framework to the field data. We evaluated the efficacy of our approach on field data collected from six commercial buildings and our results indicate that transfer learning-based models trained over one month data can perform comparative (and in some cases better) than the state-of-the-art machine learning and deep learning solutions.

Jain, Milan↗

Primordial nucleosynthesis with non-extensive statistics

The conventional Big Bang model successfully anticipates the initial abundances of 2 H(D), 3 He, and 4 He, aligning remarkably well with observational data. However, a persistent challenge arises in the case of 7 Li, where the predicted abundance exceeds observations by a factor of approximately three. Despite numerous efforts employing traditional nuclear physics to address this incongruity over the years, the enigma surrounding the lithium anomaly endures. In this context, we embark on an exploration of Big Bang nucleosynthesis (BBN) of light element abundances with the application of Tsallis non-extensive statistics. A comparison is made between the outcomes obtained by varying the non-extensive parameter q away from its unity value and both observational data and abundance predictions derived from the conventional big bang model. Here, a good agreement is found for the abundances of 4 He, 3 He and 7 Li, implying that the lithium abundance puzzle might be due to a subtle fine-tuning of the physics ingredients used to determine the BBN. However, the deuterium abundance deviates from observations.

Bertulani, Carlos A.↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

7th World Congress on Integrated Computational Materials Engineering (ICME 2023) (Final Technical Report)

Integrated Computational Materials Engineering (ICME) has received international attention due to its potential to shorten product development time, while lowering cost and improving design and manufacturing outcomes. ICME is an approach to designing materials solutions for specific applications that use computer modeling programs to predict the behavior of materials and integrate this information into the overall materials, processing, and manufacturing design cycle. The 7th World Congress on Integrated Computational Materials Engineering (ICME 2023) was held in Orlando, Florida from May 21–25, 2023 with the goal to convene stakeholders from across all areas of modeling and simulation, experimental specialization, and design, as well as from across academia, government, and industry, to address ICME tools and techniques and their integration, as well as to examine their application in engineering. This atmosphere facilitated rich interactions between the experimentalists, modelers, and computational and design, from academia, government, and industry, to discuss ICME tools and techniques and their application in engineering.

36 MATERIALS SCIENCE↗

Hardening DOE R&D Software Tools for Web-based Visualization SBIR Phase I Final Report

Ubiquitous web-based visualization is essential to delivering large-scale data visualization to various stakeholders, from the scientist to the board member. These stakeholders will not tolerate a stalled application or a pop-up window asking them to wait for the processing to complete. They require a responsive and interactive visualization environment with high-quality imagery suitable for detailed analysis and boardroom presentations. At Kitware, Inc., we have accomplished web visualization to this point, leveraging state-of-the-art tools like HTML5, CSS3, SVG, Canvas, and WebGL. Solutions that leverage a combination of these technologies are necessary to handle workloads that vary significantly in data size efficiently. However, it is not always practical to move large data to the web client for visualization. Kitware's ParaView as a Service combines client-side visualization using both distributed processing and remote rendering on big data impractical to move. Existing distributed processing and remote rendering solution's interactivity is below the expectations of web-based applications. Our project examined proposed solutions to the areas outlined above in ParaView as a Service. We have investigated concurrent pipelines, streaming images, progressive rendering, and optimization of algorithms and data movement to address these concerns. For the Phase I project, we completed the proposed work plan. As a result, the project produced three prototypes of essential importance for web visualization and the ParaView as a Service community. We created a simple desktop application for an interactive streamline placement prototype, a web-based interactive streamline placement prototype, and a web-based progressive rendering utilizing raytracing prototype. These prototypes relied on the hardening of emerging software toolkits funded by the Department of Energy (DOE) Advanced Scientific Computing Research (ASCR) program (such as ParaView, VTK-m, and Mochi). We blended these components into web-based visualization prototypes that meet the industry's expectations for interactivity and responsiveness. The Phase I project had four essential focus areas: 1. Develop prototype ParaView as a Service backend server using asynchronous, non-blocking design principles. 2. Develop a prototype web application that uses the ParaView as a Service backend server for remote data visualization. 3. Implement image streaming with encoding/compression and progressive rendering capabilities in the proposed platform. 4. Evaluate the prototype developed and summarize observations, including the challenges and pitfalls of our approach. After our successful completion of Phase I, we are strongly positioned to propose a successful Phase II project.

Geveci, Berk↗

Through the lens of bioenergy crops: advances, bottlenecks, and promises of plant engineering

Advances in engineering of bioenergy crops were driven over the past years by adapting technological breakthroughs and accelerating conventional applications but also exposed intriguing challenges. New tools revealed rich interconnectivity in the exponentially growing and dynamic 'big' omics data' of metabolomes, transcriptomes, and genomes at previously inaccessible magnitude (global, cross-species, meta-) and resolution (single cell). Insights enabled fresh hypotheses and stimulated disciplines such as functional genomics with discovery of broad regulatory networks and their determinants, that is, DNA parts, including promoters, regulatory elements, and transcription factors. Their rational design, assembly into increasingly complex blueprints, and installation into diverse chassis is an existing frontier that may benefit from emerging technologies to address bottlenecks. Interweaving nature-inspired to fully synthetic parts has already allowed building of fine-tuned regulatory circuits, or new-to-nature metabolic routes insulated from the biological context of the chassis species. Similarly, developments and the evolving need for unifying principles in plant transformation and species-agnostic technologies highlight future opportunities for engineering the next generation of bioenergy plants.

60 APPLIED LIFE SCIENCES↗

Analytics-at-scale of Sensor Data for Digital Monitoring in Nuclear Plants (3 rd Annual Report)

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This report primarily focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 ± 0.014, 0.0026 ± 0, and 0.063 ± 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure. This report summarizes the Fiscal Year 2021 research progress encompassing the (1) data cleaning and feature selection necessary for ML applications; (2) development of short-term forecasting models to predict future plant process parameters for both single and multiple time steps ahead; and (3) validation of the feature selection methods and short-term forecasting models given new data from different systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Development of Short-Term Forecasting Models Using Plant Asset Data and Feature Selection

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This paper focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 +- 0.014, 0.0026 +- 0, and 0.063 +- 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

DLIO: A DATA-CENTRIC BENCHMARK FOR DEEP LEARNING APPLICATIONS

SF-22-136 Deep learning has been shown as a successful method for various tasks, and its popularity results in numerous open-source deep learning software tools. Deep learning has been applied to a broad spectrum of scientific domains such as cosmology, particle physics, computer vision, fusion, and astrophysics. Scientists have performed a great deal of work to optimize the computational performance of deep learning frameworks. However, the same cannot be said for I/O performance. As deep learning algorithms rely on big-data volume and variety to effectively train neural networks accurately, I/O is a significant bottleneck on large-scale distributed deep learning training. DLIO, is a novel representative benchmark suite built based on the I/O profiling of the selected workloads. DLIO can be utilized to accurately emulate the I/O behavior of modern deep learning applications. Using DLIO, application developers and system software solution architects can identify potential I/O bottlenecks in their applications and guide optimizations to boost the I/O performance leading to lower training times. The storage vendor can also use DLIO as a guide for designing and optimize the storage and filesystem targeting at deep learning application.

ZHENG, HUIHUO↗

A causal data fusion method for the general exposure and outcome

Abstract With the advent of the big data era, the need to combine multiple individual data sets to draw causal effects arises naturally in many medical and biological applications. Especially each data set cannot measure enough confounders to infer the causal effect of an exposure on an outcome. In this article, we extend the method proposed by a previous study to causal data fusion of more than two data sets without external validation and to a more general (continuous or discrete) exposure and outcome. Theoretically, we obtain the condition for identifiability of exposure effects using multiple individual data sources for the continuous or discrete exposure and outcome. The simulation results show that our proposed causal data fusion method has unbiased causal effect estimate and higher precision than traditional regression, meta‐analysis and statistical matching methods. We further apply our method to study the causal effect of BMI on glucose level in individuals with diabetes by combining two data sets. Our method is essential for causal data fusion and provides important insights into the ongoing discourse on the empirical analysis of merging multiple individual data sources.

Li, Hongkai↗