Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Using Likwid and Byfl to Benchmark Hardware Performance

This paper outlines a benchmarking study conducted during my internship at LANL, focusing on CPU (Computer Processing Unit) and program performance assessment. The primary goal was to gather memory access data using three methods across five polybench kernels The data gathered would then be used to compare and contrast to one another and calculate operational intensity for performance comparisons. Benchmarking tools like Byfl and Likwid were employed, with Byfl offering hardware-independent data through LLVM compiler communication and Likwid directly interacting with computer hardware. The study considered various benchmarking factors, including optimization levels, Big O notation ((n)), CPU diversity and specific kernel equations. Big O notation was utilized to simplify code complexity, with detailed breakdwons of operations and memory components for each polybench application. Specific O(n) equations enabled nuanced kernel compariosns, facilitating the identification of performance variations. CPU efficiency assessments were conducted using Likwid tests on two CPUs. The central focus on code optimization aimed at achieving higher speeds and reduced memory usage through streamlined code. Future work propsoes creating a roofline model, synthesizing benchmarking data into a comprehensive data graph to assist in optimizing code and improving hardware performance. The potential impact on the laboratory or national mission was underscored, emphasizing the importance of optimizing applications and hardware to conserve resources and accelerate program execution. The specific relevance to LANL’s operations in math-intensive fields such as Nuclear Fission, Space Exploration, and Nanotechnology highlights the necessity of efficient benchmarking for resource conservation and proram speed. Overall, this study contributes to the understanding of CPU and program performance, providing insights for future optimization efforts in a laboratory setting

97 MATHEMATICS AND COMPUTING↗

Communication Network Awareness Machine System Phase I Development: The Intelligent Party-Line Schema

As NextGen continues toward the full implementation of a Net-Centric Architecture (N-CA)it will inherently provide a continuous increase to the Three-Vs components (Volume, Velocity, and Variety) of big data . This will create an insurmountable environment for direct-action aviation personnel (DAAP)as the DAAP’s natural abilities to manage and process data into actionable information will be overmatched by the Three-Vs. Therefore, conducting operations within a N-CA requires that new tools and applications be researched and developed to aid the DAAP’s ability to understand and manage data, mitigate non-normals, create contingency plans and actions. This paper will describe a research area at NASA Langley Research Center known as the Intelligent Party-Line (IPL).

Intelligent Party-Line↗

Smart Mobility in the Cloud: Enabling Real-Time Situational Awareness and Cyber-Physical Control Through a Digital Twin for Traffic

This article presents the design, implementation, and use cases of the Chattanooga Digital Twin (CTwin) towards the vision for next-generation smart city applications for urban mobility management. CTwin is an end-to-end web-based platform that incorporates various aspects of the decision-making process for optimizing urban transportation systems in Chattanooga, Tennessee, to reduce traffic congestion, incidents, and vehicle fuel consumption. The platform serves as a cyberinfrastructure to collect and integrate multi-domain urban mobility data from various online repositories and Internet of Things (IoT) sensors, covering multiple urban aspects (e.g., traffic, natural hazards, weather, and safety) that are relevant to urban mobility management. The platform enables advanced capabilities for: (a) real-time situational awareness on traffic and infrastructure conditions on highways and urban roads, (b) cyber-physical control for optimizing traffic signal timing, and (c) interactive visual analytics on big urban mobility data and various metrics for traffic prediction and transportation performance evaluation. The platform is designed using a multi-level componentization paradigm and is implemented using modular and adaptive architecture, rendering it as a generalizable and extendable prototype for other urban management applications. We present several use cases to demonstrate CTwin's core capabilities for supporting decision-making in smart urban mobility management.

33 ADVANCED PROPULSION SYSTEMS↗

Using Machine Learning to Predict Core Sizes of High-Efficiency Turbofan Engines

With the rise in big data and analytics, machine learning is transforming many industries. It is being increasingly employed to solve a wide range of complex problems, producing autonomous systems that support human decision-making. For the aircraft engine industry, machine learning of historical and existing engine data could provide insights that help drive for better engine design. This work explored the application of machine learning to engine preliminary design. Engine core-size prediction was chosen for the first study because of its relative simplicity in terms of number of input variables required (only three). Specifically, machine-learning predictive tools were developed for turbofan engine core-size prediction, using publicly available data of two hundred manufactured engines and engines that were studied previously in NASA aeronautics projects. The prediction results of these models show that, by bringing together big data, robust machine-learning algorithms and automation, a machine learning-based predictive model can be an effective tool for turbofan engine core-size prediction. The promising results of this first study paves the way for further exploration of the use of machine learning for aircraft engine preliminary design.

Tong, Michael T.↗

Using Machine Learning to Predict Core Sizes of High-Efficiency Turbofan Engines

With the rise in big data and analytics, machine learning is transforming many industries. It is being increasingly employed to solve a wide range of complex problems, producing autonomous systems that support human decision-making. For the aircraft engine industry, machine learning of historical and existing engine data could provide insights that help drive for better engine design. This work explored the application of machine learning to engine preliminary design. Engine core-size prediction was chosen for the first study because of its relative simplicity in terms of number of input variables required (only three). Specifically, machine-learning predictive tools were developed for turbofan engine core-size prediction, using publicly available data of two hundred manufactured engines and engines that were studied previously in NASA aeronautics projects. The prediction results of these models show that, by bringing together big data, robust machine-learning algorithms and data science, a machine learning-based predictive model can be an effective tool for turbofan engine core-size prediction. The promising results of this first study paves the way for further exploration of the use of machine learning for aircraft engine preliminary design.

Core Size↗

GOOML - Finding Optimization Opportunities for Geothermal Operations: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach. We have used this framework to develop digital twins that provide steamfield operators with an operational environment to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management for real world applications. The GOOML modeling software is built on a generic component-based systems framework that allows for both historical and forecast analysis. A GOOML model can perform historical data-assimilation using first-principal thermodynamics to create a meaningful data model. Historical production data can then be coupled with a forecast framework to train machine-learning models of steamfield components to predict future outputs. This modeling environment enables digital exploration of steamfield design configurations and operational scenarios. GOOML digital twins have been developed for steamfields in New Zealand and the United States representing differing power generation and field conditions. These digital twins have been validated by comparing hindcast predictions against historical production data. Reinforcement learning experiments were conducted to demonstrate the ability to programmatically explore the operations space using machine learning agents. Our initial results are compelling; two to five percent increases in annual energy production were demonstrated by the GOOML models with no additional infrastructure build required. GOOML offers a new approach to geothermal operations by applying state-of-the-art machine learning algorithms, comprehensive data analytics, and interaction with digital twins. Through application of these tools, operators will realize greater availability and higher net generation which will increase the cost effectiveness of geothermal energy projects.

access↗

Investigating Access Performance of Long Time Series with Restructured Big Model Data

Data sets generated by models are substantially increasing in volume, due to increases in spatial and temporal resolution, and the number of output variables. Many users wish to download subsetted data in preferred data formats and structures, as it is getting increasingly difficult to handle the original full-size data files. For example, application research users such as those involved with wind or solar energy, or extreme weather events are likely only interested in daily or hourly model data at a single point (or for a small area) for a long time period, and prefer to have the data downloaded in a single file. With native model file structures, such as hourly data from NASA Modern-Era Retrospective analysis for Research and Applications Version-2 (MERRA-2), it may take over 10 hours for the extraction of parameters-of-interest at a single point for 30 years. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) is exploring methods to address this particular user need. One approach is to create value-added data by reconstructing the data files. Taking MERRA-2 data as an example, we have tested converting hourly data from one-day-per-file into different data cubes, such as one-month, or one-year. Performance is compared for reading local data files and accessing data through interoperable services, such as OPeNDAP. Results show that, compared to the original file structure, the new data cubes offer much better performance for accessing long time series. We have noticed that performance is associated with the cube size and structure, the compression method, and how the data are accessed. An optimized data cube structure will not only improve data access, but also may enable better online analysis services

reanalysis↗

Unleashing the Power of Industrial Big Data through Scalable Manual Labeling

Big Data plays a central role in the remarkable results achieved by Machine Learning (ML) and especially Deep Learning (DL) in the recent years. However, the difficulty in obtaining a reasonable amount of labeled samples limits ML/DL application in various domains, including industrial equipment and system monitoring. In this paper the need for methods that turn manual labeling into a scalable process is highlighted. A real world problem is analyzed for which weak supervision methods, successfully employed in other domains, did not produce acceptable results. An alternative approach based on clustering ensembles is described and tested, achieving good performance.

Paes Leao, Bruno↗

Anticipated Changes in Conducting Scientific Data-Analysis Research in the Big-Data Era

A Big-Data environment is one that is capable of orchestrating quick-turnaround analyses involving large volumes of data for numerous simultaneous users. Based on our experiences with a prototype Big-Data analysis environment, we anticipate some important changes in research behaviors and processes while conducting scientific data-analysis research in the near future as such Big-Data environments become the mainstream. The first anticipated change will be the reduced effort and difficulty in most parts of the data management process. A Big-Data analysis environment is likely to house most of the data required for a particular research discipline along with appropriate analysis capabilities. This will reduce the need for researchers to download local copies of data. In turn, this also reduces the need for compute and storage procurement by individual researchers or groups, as well as associated maintenance and management afterwards. It is almost certain that Big-Data environments will require a different "programming language" to fully exploit the latent potential. In addition, the process of extending the environment to provide new analysis capabilities will likely be more involved than, say, compiling a piece of new or revised code.We thus anticipate that researchers will require support from dedicated organizations associated with the environment that are composed of professional software engineers and data scientists. A major benefit will likely be that such extensions are of higherquality and broader applicability than ad hoc changes by physical scientists. Another anticipated significant change is improved collaboration among the researchers using the same environment. Since the environment is homogeneous within itself, many barriers to collaboration are minimized or eliminated. For example, data and analysis algorithms can be seamlessly shared, reused and re-purposed. In conclusion, we will be able to achieve a new level of scientific productivity in the Big-Data analysis environments.

Kuo, Kwo-Sen↗

Investigation and Evaluation of Advanced Spectrum Management Concepts for Aeronautical Communications

With the emergence of new aerial vehicles into the airspace and the continued growth of aviation operations, there will be an increasing demand for voice and data communications within the National Airspace System (NAS). The continued use of existing VHF and UHF frequency allocations is not a sustainable approach, and as a result, a new spectrum management solution is required to support future mission needs. The proposed spectrum management concepts leverage modern advancements such as artificial intelligence (AI) and big data to dynamically optimize the spectrum utilization based on the predicted communications demand throughout the airspace. This technical investigation considers both air-ground and air-air communications networks, and can be applied to both existing applications such as the air traffic control (ATC) system, as well as future applications, such as the emerging Advanced Air Mobility (AAM). To support the evaluation of the proposed concepts and technologies, a modeling and simulation capability is currently under development and will continue to evolve to support new and advanced airspace applications. It is anticipated that this proposed spectrum concept will better serve the spectrum needs of future NAS applications.

Eric J Knoblock↗

Investigation and Evaluation of Advanced Spectrum Management Concepts for Aeronautical Communications

With the emergence of new aerial vehicles into the airspace and the continued growth of aviation operations, there will be an increasing demand for voice and data communications within the National Airspace System (NAS). The continued use of existing VHF and UHF frequency allocations is not a sustainable approach, and as a result, a new spectrum management solution is required to support future mission needs. The proposed spectrum management concepts leverage modern advancements such as artificial intelligence (AI) and big data to dynamically optimize the spectrum utilization based on the predicted communications demand throughout the airspace. This technical investigation considers both air-ground and air-air communications networks, and can be applied to both existing applications such as the air traffic control (ATC) system, as well as future applications, such as the emerging Advanced Air Mobility (AAM). To support the evaluation of the proposed concepts and technologies, a modeling and simulation capability is currently under development and will continue to evolve to support new and advanced airspace applications. It is anticipated that this proposed spectrum concept will better serve the spectrum needs of future NAS applications.

Eric J Knoblock↗

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes↗

Special Issue: Geostatistics and Machine Learning

Abstract Recent years have seen a steady growth in the number of papers that apply machine learning methods to problems in the earth sciences. Although they have different origins, machine learning and geostatistics share concepts and methods. For example, the kriging formalism can be cast in the machine learning framework of Gaussian process regression. Machine learning, with its focus on algorithms and ability to seek, identify, and exploit hidden structures in big data sets, is providing new tools for exploration and prediction in the earth sciences. Geostatistics, on the other hand, offers interpretable models of spatial (and spatiotemporal) dependence. This special issue on Geostatistics and Machine Learning aims to investigate applications of machine learning methods as well as hybrid approaches combining machine learning and geostatistics which advance our understanding and predictive ability of spatial processes.

58 GEOSCIENCES↗

Small Angle Scattering Data Analysis Assisted by Machine Learning Methods

Small angle scattering (SAS) is a widely used technique for characterizing structures of wide ranges of materials. For such wide ranges of applications of SAS, there exist a large number of ways to model the scattering data. While such analysis models are often available from various suites of SAS data analysis software packages, selecting the right model to start with poses a big challenge for beginners to SAS data analysis. Here, we present machine learning (ML) methods that can assist users by suggesting scattering models for data analysis. A series of one-dimensional scattering curves have been generated by using different models to train the algorithms. The performance of the ML method is studied for various types of ML algorithms, resolution of the dataset, and the number of the dataset. The degree of similarities among selected scattering models is presented in terms of the confusion matrix. The scattering model suggestions with prediction scores provide a list of scattering models that are likely to succeed. Therefore, if implemented with extensive libraries of scattering models, this method can speed up the data analysis workflow by reducing search spaces for appropriate scattering models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Modeling and Simulation of an Integrated Gate Turnaround Management Concept

An Integrated Gate Turnaround Management (IGTM) prototype was developed at NASA Ames Simulation Laboratories (SimLabs) using Dallas Ft. Worth International Airport (DFW) to demonstrate the IGTM concepts feasibility and benefits. The simulation architecture includes: the IGTM controller, an Airline Operations Control (AOC) application, Big DataAnalytics Input (BAI) application, a terminal traffic simulation or known as NASA-developed Surface Operation Simulator and Scheduler (SOSS), and a Database Server. ActiveMQ, a Java messaging service, was used to emulate the System Wide Information Management (SWIM) data network messaging. This paper describes the modeling and simulation of the IGTM concept, and illustrates selected use cases to demonstrate the feasibility and benefits of the IGTM concept for optimizing gate turnaround operations.

gate turnaround management↗

Machine and Deep Learning: Artificial Intelligence Application in Biotic and Abiotic Stress Management in Plants

Biotic and abiotic stresses significantly affect plant fitness, resulting in a serious loss in food production. Biotic and abiotic stresses predominantly affect metabolite biosynthesis, gene and protein expression, and genome variations. However, light doses of stress result in the production of positive attributes in crops, like tolerance to stress and biosynthesis of metabolites, called hormesis. Advancement in artificial intelligence (AI) has enabled the development of high-throughput gadgets such as high-resolution imagery sensors and robotic aerial vehicles, i.e., satellites and unmanned aerial vehicles (UAV), to overcome biotic and abiotic stresses. These High throughput (HTP) gadgets produce accurate but big amounts of data. Significant datasets such as transportable array for remotely sensed agriculture and phenotyping reference platform (TERRA-REF) have been developed to forecast abiotic stresses and early detection of biotic stresses. For accurately measuring the model plant stress, tools like Deep Learning (DL) and Machine Learning (ML) have enabled early detection of desirable traits in a large population of breeding material and mitigate plant stresses. In this review, advanced applications of ML and DL in plant biotic and abiotic stress management have been summarized.

59 BASIC BIOLOGICAL SCIENCES↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗