Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

Dimensionality reduction using elastic measures

With the recent surge in big data analytics for hyperdimensional data, there is a renewed interest in dimensionality reduction techniques. In order for these methods to improve performance gains and understanding of the underlying data, a proper metric needs to be identified. This step is often overlooked, and metrics are typically chosen without consideration of the underlying geometry of the data. Here, in this paper, we present a method for incorporating elastic metrics into the t-distributed stochastic neighbour embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP). We apply our method to functional data, which is uniquely characterized by rotations, parameterization and scale. If these properties are ignored, they can lead to incorrect analysis and poor classification performance. Through our method, we demonstrate improved performance on shape identification tasks for three benchmark data sets (MPEG-7, Car data set and Plane data set of Thankoor), where we achieve 0.77, 0.95 and 1.00 F1 score, respectively.

97 MATHEMATICS AND COMPUTING↗

Aggregate attack surface management for network discovery of operational technology

Interconnectivity has become a substratum of technology as the benefits of data-driven functionality are being realized in nearly all industries. Increased connectivity of Operational Technology (OT) exacerbates cyber risks because Industrial Control Systems (ICS) are becoming exposed to the Internet. These exposures are often done inadvertently through misconfigurations as additional network devices come online. Attack surface management (ASM) platforms can be used to identify vulnerabilities by performing external network discovery over the Internet using web spiders. These web spiders enable big data analytics of Internet of Things (IoT) devices as identifiable information of Internet-exposed equipment are archived in searchable databases that are made publicly available. There are a multitude of ASM service providers on the market. Here, this study was conducted to evaluate several commonly known tools to determine the aggregate attack surface of control systems. Queries were crafted by targeting commonly known manufacturers and communication protocols found in OT networks. Identified devices were that categorized based on technology types. Each query was replicated between several tools to target identical ICS equipment. Findings in this paper suggested a significant variance in the exposures discovered by each tool, but unique contributions were identified for each tool when a merged attack surface was derived. Therefore, all tools should be used in aggregate.

97 MATHEMATICS AND COMPUTING↗

Machine learning models inaccurately predict current and future high-latitude C balances

The high-latitude carbon (C) cycle is a key feedback to the global climate system, yet because of system complexity and data limitations, there is currently disagreement over whether the region is a source or sink of C. Recent advances in big data analytics and computing power have popularized the use of machine learning (ML) algorithms to upscale site measurements of ecosystem processes, and in some cases forecast the response of these processes to climate change. Due to data limitations, however, ML model predictions of these processes are almost never validated with independent datasets. To better understand and characterize the limitations of these methods, we develop an approach to independently evaluate ML upscaling and forecasting. We mimic data-driven upscaling and forecasting efforts by applying ML algorithms to different subsets of regional process-model simulation gridcells, and then test ML performance using the remaining gridcells. In this study, we simulate C fluxes and environmental data across Alaska using ecosys, a process-rich terrestrial ecosystem model, and then apply boosted regression tree ML algorithms to training data configurations that mirror and expand upon existing AmeriFLUX eddy-covariance data availability. We first show that a ML model trained using ecosys outputs from currently-available Alaska AmeriFLUX sites incorrectly predicts that Alaska is presently a modeled net C source. Increased spatial coverage of the training dataset improves ML predictions, halving the bias when 240 modeled sites are used instead of 15. However, even this more accurate ML model incorrectly predicts Alaska C fluxes under 21st century climate change because of changes in atmospheric CO 2 , litter inputs, and vegetation composition that have impacts on C fluxes which cannot be inferred from the training data. Our results provide key insights to future C flux upscaling efforts and expose the potential for inaccurate ML upscaling and forecasting of high-latitude C cycle dynamics.

54 ENVIRONMENTAL SCIENCES↗

Strategies for Integrating Deep Learning Surrogate Models with HPC Simulation Applications

The emerging trend of the convergence of high performance computing (HPC), machine learning/deep learning (ML/DL), and big data analytics presents a host of challenges for large-scale computing campaigns that seek best practices to interleave traditional scientific simulation-based workloads with ML/DL models. A portfolio of systematic approaches to incorporate deep learning into modeling and simulation serves a vital need when we support AI for science at a computing facility. In this paper, we evaluate several strategies for deploying deep learning surrogate models in a representative physics application on supercomputers at the Oak Ridge Leadership Computing Facility (OLCF). We discuss a set of recommended deployment architectures and implementation approaches. We analyze and evaluate these alternatives and show their performance and scalability up to 1000 GPUs on two mainstream platforms equipped with different deep learning hardware and software stacks.

Yin, Junqi↗

Enabling HPC Scientific Workflows for Serverless

The convergence of edge computing, big data analytics, and AI with traditional scientific calculations is increasingly being adopted in HPC workflows. Workflow management systems are crucial for managing and orchestrating these complex computational tasks. However, it is difficult to identify patterns within the growing population of HPC workflows. Serverless has emerged as a novel computing paradigm, offering dynamic resource allocation, quick response time, fine-grained resource management and auto-scaling. In this paper, we propose a framework to enable HPC scientific workflows on serverless. Our approach integrates a widely used traditional HPC workflow generator with an HPC serverless workflow management system to create benchmark suites of scientific workflows with diverse characteristics. These workflows can be executed on different serverless platforms. We comprehensively compare executing workflows on traditional local containers and serverless computing platforms. Our results show that serverless can reduce CPU and memory usage respectively by 78.11% and 73.92% without compromising performance.

Andrei da silva, Anderson↗

Dimensionally reduced machine learning model for predicting single component octanol–water partition coefficients

Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Big Data Analysis of Synchrophasor Data: Outcomes of Research Activities Supported by DOE FOA 1861

This report describes the key outcomes of research activities sponsored by the Department of Energy’s Funding Opportunity Announcement (FOA) number 1861 that was aimed at advancing the state-of-the-art in big data analytics applied to transmission-level synchrophasor measurements. The FOA resulted in eight research grants where the awardees developed machine learning and artificial intelligence tools and approaches. The commonalities in tools and approaches used by the awardees are explored, and insights gained from how the project outcomes might be operationalized are discussed. This report does not seek to comprehensively summarize all research supported by the FOA, rather it focuses on enabling the fast dissemination of major findings to the broader power systems community.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

EDX ClaiMM: Digital Resources for the Critical Minerals and Materials Community

Securing critical mineral supply chains is essential for transitioning to a clean energy economy and for maintaining national security. Big-data analytics can serve as a cost-effective means of identifying new domestic critical mineral resources but only if data can be easily located and digested. Using ArcGIS Enterprise Sites, EDX ClaiMM was developed to increase the accessibility of critical minerals data, reducing time spent on data collection and integration. Hosted tools provide rapid visualization and exploration of key datasets, unlocking insights to support resource assessments.

Yesenchak, Rachel↗

Emerging Materials Technologies That Matter to Manufacturers

A brief overview of emerging materials technologies. Exploring the weight reduction benefit of replacing Carbon Fiber with Carbon Nanotube (CNT) in Polymer Composites. Review of the benign purification method developed for CNT sheets. The future of manufacturing will include the integration of computational material design and big data analytics, along with Nanomaterials as building blocks.

nanomaterials↗

Utilizing HDF4 File Content Maps for the Cloud

We demonstrate a prototype study that HDF4 file content map can be used for efficiently organizing data in cloud object storage system to facilitate cloud computing. This approach can be extended to any binary data formats and to any existing big data analytics solution powered by cloud computing because HDF4 file content map project started as long term preservation of NASA data that doesn't require HDF4 APIs to access data.

Elastic Search↗

The Challenges of Human-Autonomy Teaming

Machine intelligence is improving rapidly based on advances in big data analytics, deep learning algorithms, networked operations, and continuing exponential growth in computing power (Moores Law). This growth in the power and applicability of increasingly intelligent systems will change the roles humans, shifting them to tasks where adaptive problem solving, reasoning and decision-making is required. This talk will address the challenges involved in engineering autonomous systems that function effectively with humans in aeronautics domains.

artificial intelligence↗

The Challenges of Human-Autonomy Teaming

Machine intelligence is improving rapidly based on advances in big data analytics, deep learning algorithms, networked operations, and continuing exponential growth in computing power (Moores Law). This growth in the power and applicability of increasingly intelligent systems will change the roles humans, shifting them to tasks where adaptive problem solving, reasoning and decision-making is required. This talk will address the challenges involved in engineering autonomous systems that function effectively with humans in aeronautics domains.

Human-Autonomy teaming↗

Introduction to NASA Goddard Workshop on Artificial Intelligence

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few.This workshop will be investigating how AI technologies can be adapted or developed to address the following challenges: Discover events of interest and correlations in large amounts of science data; improve the outcomes of science modeling and data assimilation using improved data processing, integration, and analysis. Design advisors for mission planning and operations, including anomaly detection and spacecraft health monitoring. Develop tools for engineering support, including advanced manufacturing, orbit determination, new component design and system engineering. Customize intelligent user interfaces, including visual analytics and natural language processing.

Le Moigne, Jacqueline↗

Overview of Artificial Intelligence (AI) at NASA Goddard

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few. During the Tour, we will show a few examples of the current AI activities at NASA Goddard.

Le Moigne, Jacqueline↗

Air Traffic Management TestBed Simulation Architect: User's Guide

The Air Traffic Management (ATM) TestBed is a Platform as a Service that is being developed by the National Aeronautics and Space Administration (NASA) to help design, configure, integrate, run, and monitor air traffic simulations. The platform provides cloud services including back-end big-data analytics tools, on-demand computing resource management, data storage, and communication middleware. The ATM TestBed reduces the time to test concepts and technologies, supports interactions among various concepts such as human-in-the-loop and automation-in-the-loop simulations, and enables collaborative simulations by sharing technologies and tools in the ATM community. The Simulation Architect application provides a graphical user interface tool for designing traffic scenarios and simulations using blocks representing components and links representing message channels linking them. This guide describes a high-level user interface design of Simulation Architect and provides information for a new user to compose traffic scenarios and simulations.

Software User Guide↗

Air Traffic Management TestBed Traffic Viewer: Developer's Guide

The Air Traffic Management (ATM) TestBed is a Platform as a Service that is being developed by the National Aeronautics and Space Administration (NASA) to help design, configure, integrate, run, and monitor air traffic simulations. The platform is designed to provide cloud services including back-end, big-data analytics tools, on-demand computing resource management, data storage, and communication middleware. The ATM TestBed reduces the time to test concepts and technologies, supports interactions among various methods such as human-in-the-loop and automation-in-the-loop simulations, and enables collaborative simulations by sharing technologies and tools in the ATM community. The Traffic Viewer application provides a graphical user interface tool for visualizing real and simulated air traffic as well as airspace definition in two-dimensional space. This guide describes a high-level design and implementation of Traffic Viewer and provides information for a new developer or a user to add new capabilities by following the software design and leveraging existing capabilities.

Lai, Chok Fung↗

TPSAS-NF1676L-20739-DND

Overview of LaRC's Comprehensive Digital Transformation strategy. Includes foundational items such as modeling & simulation, big data analytics, high performance computing, and advanced information technology. Builds upon the foundation to recommend advanced "virtual capabilities," such as a virtual flight test capability to complement flight testing and wind tunnels.

Ed McLarney↗