Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

DLIO: A DATA-CENTRIC BENCHMARK FOR DEEP LEARNING APPLICATIONS

SF-22-136 Deep learning has been shown as a successful method for various tasks, and its popularity results in numerous open-source deep learning software tools. Deep learning has been applied to a broad spectrum of scientific domains such as cosmology, particle physics, computer vision, fusion, and astrophysics. Scientists have performed a great deal of work to optimize the computational performance of deep learning frameworks. However, the same cannot be said for I/O performance. As deep learning algorithms rely on big-data volume and variety to effectively train neural networks accurately, I/O is a significant bottleneck on large-scale distributed deep learning training. DLIO, is a novel representative benchmark suite built based on the I/O profiling of the selected workloads. DLIO can be utilized to accurately emulate the I/O behavior of modern deep learning applications. Using DLIO, application developers and system software solution architects can identify potential I/O bottlenecks in their applications and guide optimizations to boost the I/O performance leading to lower training times. The storage vendor can also use DLIO as a guide for designing and optimize the storage and filesystem targeting at deep learning application.

ZHENG, HUIHUO↗

A causal data fusion method for the general exposure and outcome

Abstract With the advent of the big data era, the need to combine multiple individual data sets to draw causal effects arises naturally in many medical and biological applications. Especially each data set cannot measure enough confounders to infer the causal effect of an exposure on an outcome. In this article, we extend the method proposed by a previous study to causal data fusion of more than two data sets without external validation and to a more general (continuous or discrete) exposure and outcome. Theoretically, we obtain the condition for identifiability of exposure effects using multiple individual data sources for the continuous or discrete exposure and outcome. The simulation results show that our proposed causal data fusion method has unbiased causal effect estimate and higher precision than traditional regression, meta‐analysis and statistical matching methods. We further apply our method to study the causal effect of BMI on glucose level in individuals with diabetes by combining two data sets. Our method is essential for causal data fusion and provides important insights into the ongoing discourse on the empirical analysis of merging multiple individual data sources.

Li, Hongkai↗

SEAFORML (Smart Exploration and Analysis For Optimal and Robust Machine Learning)

The poster discusses data analysis of the WAVgraph database and applied machine learning methods for it. The database is a long-term project that seeks to be a comprehensive repository of information on cyber threats and is updated regularly. It was previously unanalyzed and unexplored. The goal was to learn more about it and its contents in order to have a better understanding and enable better use. The data analysis and discovery enabled further exploration through natural language processing, similarity, and clustering methods. The poster shows some of the insights from the analysis and explains the methods used for the machine learning applications.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Toward a Smart Metaverse City: Immersive Realism and 3D Visualization of Digital Twin Cities

Metaverse and its related extended reality technologies can enable immersive, realistic, and participatory visualization of 3D data, and their use and the potential within a smart city can be effective for supporting urban research and urban operations management. This book chapter describes a vision for prototyping a “Smart Metaverse City” to combine the unique advantage of the Metaverse technology with the two-way connectivity of a digital twin city application. This synergy aims to create a virtual environment for immersive geovisualization to help researchers and the public understand the complex urban system through science-based and data-driven approaches. This book chapter selectively reviews past technological and paradigm advancements for collecting, analyzing, and visualizing 3D urban big data. Then we present a prototyping Geographic Information System (GIS) framework, together with some relevant data sources and open-source web technologies, to help researchers create a smart Metaverse city. We demonstrate our vision and discuss its application opportunities through a real-world example, a digital twin city developed at the Oak Ridge National Laboratory, to facilitate participatory, smart, and sustainable campus management.

Xu, Haowen↗

A Digital Twin Framework Utilizing Machine Learning for Robust Predictive Maintenance: Enhancing Tire Health Monitoring

We introduce a novel digital twin (DT) framework for the predictive maintenance of long-term physical systems. Using monitoring tire health as an application, we show how the DT framework can be used to enhance automotive safety and efficiency, and how the technical challenges can be overcome using a three-step approach. First, to manage the data complexity over a long operation span, we employ data reduction techniques to concisely represent physical tires using historical performance and usage data. Relying on these data, for fast real-time prediction, we train a transformer-based model offline on our concise dataset to predict future tire health over time, represented as remaining casing potential (RCP). Based on our architecture, our model quantifies both epistemic and aleatoric uncertainties, providing reliable confidence intervals around predicted RCP. Second, to incorporate real-time data, we update the predictive model in the DT framework, ensuring its accuracy throughout its lifespan with the aid of hybrid modeling and the use of the discrepancy function. Third, to assist decision-making in predictive maintenance, we implement a tire state decision algorithm, which strategically determines the optimal timing for tire replacement based on RCP forecasted by our transformer model. This approach ensures that our DT accurately predicts system health, continually refines its digital representation, and supports predictive maintenance decisions. Furthermore, our framework effectively embodies a physical system, leveraging big data and machine learning (ML) for predictive maintenance, model updates, and decision-making.

advanced computing infrastructure↗

The Challenges of Human-Autonomy Teaming

Machine intelligence is improving rapidly based on advances in big data analytics, deep learning algorithms, networked operations, and continuing exponential growth in computing power (Moores Law). This growth in the power and applicability of increasingly intelligent systems will change the roles humans, shifting them to tasks where adaptive problem solving, reasoning and decision-making is required. This talk will address the challenges involved in engineering autonomous systems that function effectively with humans in aeronautics domains.

artificial intelligence↗

The Challenges of Human-Autonomy Teaming

Machine intelligence is improving rapidly based on advances in big data analytics, deep learning algorithms, networked operations, and continuing exponential growth in computing power (Moores Law). This growth in the power and applicability of increasingly intelligent systems will change the roles humans, shifting them to tasks where adaptive problem solving, reasoning and decision-making is required. This talk will address the challenges involved in engineering autonomous systems that function effectively with humans in aeronautics domains.

Human-Autonomy teaming↗

Multi-fidelity information fusion with concatenated neural networks

Recently, computational modeling has shifted towards the use of statistical inference, deep learning, and other data-driven modeling frameworks. Although this shift in modeling holds promise in many applications like design optimization and real-time control by lowering the computational burden, training deep learning models needs a huge amount of data. This big data is not always available for scientific problems and leads to poorly generalizable data-driven models. This gap can be furnished by leveraging information from physics-based models. Exploiting prior knowledge about the problem at hand, this study puts forth a physics-guided machine learning (PGML) approach to build more tailored, effective, and efficient surrogate models. For our analysis, without losing its generalizability and modularity, we focus on the development of predictive models for laminar and turbulent boundary layer flows. In particular, we combine the self-similarity solution and power-law velocity profile (low-fidelity models) with the noisy data obtained either from experiments or computational fluid dynamics simulations (high-fidelity models) through a concatenated neural network. We illustrate how the knowledge from these simplified models results in reducing uncertainties associated with deep learning models applied to boundary layer flow prediction problems. The proposed multi-fidelity information fusion framework produces physically consistent models that attempt to achieve better generalization than data-driven models obtained purely based on data. While we demonstrate our framework for a problem relevant to fluid mechanics, its workflow and principles can be adopted for many scientific problems where empirical, analytical, or simplified models are prevalent. In line with grand demands in novel PGML principles, this work builds a bridge between extensive physics-based theories and data-driven modeling paradigms and paves the way for using hybrid physics and machine learning modeling approaches for next-generation digital twin technologies.

42 ENGINEERING↗

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗

Artificial intelligence in cancer research, diagnosis and therapy

Artificial intelligence and machine learning techniques are breaking into biomedical research and health care, which importantly includes cancer research and oncology, where the potential applications are vast. These include detection and diagnosis of cancer, subtype classification, optimization of cancer treatment and identification of new therapeutic targets in drug discovery. While big data used to train machine learning models may already exist, leveraging this opportunity to realize the full promise of artificial intelligence in both the cancer research space and the clinical space will first require significant obstacles to be surmounted. In this Viewpoint article, we asked four experts for their opinions on how we can begin to implement artificial intelligence while ensuring standards are maintained so as transform cancer diagnosis and the prognosis and treatment of patients with cancer and to drive biological discovery.

60 APPLIED LIFE SCIENCES↗

Applications of remote sensor data by state and Federal user agencies in Arizona

The use of NASA high altitude aerial photography of south eastern Arizona to develop a natural resources information system for Federal lands is discussed. The data are to be used by local, State, and Federal agencies in connection with geologic mapping projects, water resources investigations, and land use studies to determine the alignment of a proposed major aqueduct. In addition, the data are used to confirm land ownership boundaries, detect changes in land use, and legislative reappointment mapping. Other applications include mapping vegetive cover, evaluation of changes in wildlife habitat, location of deer kills, and as a base for recording telemetry data from radio-collared big game animals.

Schumann, H. H.↗

A Machine-Learning Approach to Assess Aircraft Engine System Performance

Artificial intelligence (AI)/machine learning, and big data are transforming the global business environment. They have become the most disruptive technologies for organizations to improve workplace efficiency and productivity. This work explored the application of machine learning-based predictive analytics that would enable aircraft engine designers to estimate engine system performance quickly during the conceptual design stage. Supervised machine-learning algorithm was employed to study patterns in an existing database of production and research turbofan engines, and built predictive analytics for use in predicting system performance of new turbofan designs. Specifically, the author developed deep-learning analytics to predict turbofan system weight, using turbofan design parameters as the input. The predictive analytics were trained and deployed in Keras, an open-source neural networks API (application program interface) written in Python, with TensorFlow (an open-source artificial AI library developed by Google) serving as the backend engine. The current engine-weight prediction results, together with those for the TSFC (thrust specific fuel consumption) and core-size predictions that were studied previously by the author, show that machine learning-based predictive analytics can be an effective, time-saving tool for aircraft engine design-space exploration during the conceptual design stage. It would enable expeditious identification of the best engine design amongst several candidates.

Michael T Tong↗

Defining and applying an electricity demand flexibility benchmarking metrics framework for grid-interactive efficient commercial buildings

Building demand flexibility (DF) research has recently gained attention. To unlock building DF as a predictable grid resource, we must establish a quantitative understanding of the resource size, performance variability, and predictability based on large empirical datasets. Researchers have proposed various sets of theoretical metrics to measure this performance. Some metrics have been applied to simulation results, but most fall short of exploring the complexities in real building applications. There are practical metrics used in individual demand response field studies but they alone cannot fulfil the job of DF benchmarking across a diverse group of buildings. The electrical grid's geographically diverse and changing nature presents challenges to comparing building DF performance measured under different conditions (i.e., benchmarking DF). To address this challenge, a novel DF benchmarking framework focused on load shedding and shifting is presented; the foundation is a set of simple, proven single-event metrics with attributes describing event conditions. These enable benchmarking and visualization in different dimensions for identifying trends that represent how these attributes influence DF. To test its feasibility and scalability, the DF framework was applied to two case studies of 11 office buildings and 121 big-box retail buildings with demand response participation data. Furthermore, these examples provided a pathway for using both building level benchmarking and aggregation to extract insights into building DF about magnitude, consistency, and influential factors. Potential applications of the framework and real-world values have been identified for grid and building stakeholders.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Climate Analytics as a Service

Climate science is a big data domain that is experiencing unprecedented growth. In our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS). CAaaS combines high-performance computing and data-proximal analytics with scalable data management, cloud computing virtualization, the notion of adaptive analytics, and a domain-harmonized API to improve the accessibility and usability of large collections of climate data. MERRA Analytic Services (MERRA/AS) provides an example of CAaaS. MERRA/AS enables MapReduce analytics over NASA's Modern-Era Retrospective Analysis for Research and Applications (MERRA) data collection. The MERRA reanalysis integrates observational data with numerical models to produce a global temporally and spatially consistent synthesis of key climate variables. The effectiveness of MERRA/AS has been demonstrated in several applications. In our experience, CAaaS is providing the agility required to meet our customers' increasing and changing data management and data analysis needs.

big data↗

Using Likwid and Byfl to Benchmark Hardware Performance

This paper outlines a benchmarking study conducted during my internship at LANL, focusing on CPU (Computer Processing Unit) and program performance assessment. The primary goal was to gather memory access data using three methods across five polybench kernels The data gathered would then be used to compare and contrast to one another and calculate operational intensity for performance comparisons. Benchmarking tools like Byfl and Likwid were employed, with Byfl offering hardware-independent data through LLVM compiler communication and Likwid directly interacting with computer hardware. The study considered various benchmarking factors, including optimization levels, Big O notation ((n)), CPU diversity and specific kernel equations. Big O notation was utilized to simplify code complexity, with detailed breakdwons of operations and memory components for each polybench application. Specific O(n) equations enabled nuanced kernel compariosns, facilitating the identification of performance variations. CPU efficiency assessments were conducted using Likwid tests on two CPUs. The central focus on code optimization aimed at achieving higher speeds and reduced memory usage through streamlined code. Future work propsoes creating a roofline model, synthesizing benchmarking data into a comprehensive data graph to assist in optimizing code and improving hardware performance. The potential impact on the laboratory or national mission was underscored, emphasizing the importance of optimizing applications and hardware to conserve resources and accelerate program execution. The specific relevance to LANL’s operations in math-intensive fields such as Nuclear Fission, Space Exploration, and Nanotechnology highlights the necessity of efficient benchmarking for resource conservation and proram speed. Overall, this study contributes to the understanding of CPU and program performance, providing insights for future optimization efforts in a laboratory setting

97 MATHEMATICS AND COMPUTING↗

Communication Network Awareness Machine System Phase I Development: The Intelligent Party-Line Schema

As NextGen continues toward the full implementation of a Net-Centric Architecture (N-CA)it will inherently provide a continuous increase to the Three-Vs components (Volume, Velocity, and Variety) of big data . This will create an insurmountable environment for direct-action aviation personnel (DAAP)as the DAAP’s natural abilities to manage and process data into actionable information will be overmatched by the Three-Vs. Therefore, conducting operations within a N-CA requires that new tools and applications be researched and developed to aid the DAAP’s ability to understand and manage data, mitigate non-normals, create contingency plans and actions. This paper will describe a research area at NASA Langley Research Center known as the Intelligent Party-Line (IPL).

Intelligent Party-Line↗

Smart Mobility in the Cloud: Enabling Real-Time Situational Awareness and Cyber-Physical Control Through a Digital Twin for Traffic

This article presents the design, implementation, and use cases of the Chattanooga Digital Twin (CTwin) towards the vision for next-generation smart city applications for urban mobility management. CTwin is an end-to-end web-based platform that incorporates various aspects of the decision-making process for optimizing urban transportation systems in Chattanooga, Tennessee, to reduce traffic congestion, incidents, and vehicle fuel consumption. The platform serves as a cyberinfrastructure to collect and integrate multi-domain urban mobility data from various online repositories and Internet of Things (IoT) sensors, covering multiple urban aspects (e.g., traffic, natural hazards, weather, and safety) that are relevant to urban mobility management. The platform enables advanced capabilities for: (a) real-time situational awareness on traffic and infrastructure conditions on highways and urban roads, (b) cyber-physical control for optimizing traffic signal timing, and (c) interactive visual analytics on big urban mobility data and various metrics for traffic prediction and transportation performance evaluation. The platform is designed using a multi-level componentization paradigm and is implemented using modular and adaptive architecture, rendering it as a generalizable and extendable prototype for other urban management applications. We present several use cases to demonstrate CTwin's core capabilities for supporting decision-making in smart urban mobility management.

33 ADVANCED PROPULSION SYSTEMS↗