Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Using Machine Learning to Predict Core Sizes of High-Efficiency Turbofan Engines

With the rise in big data and analytics, machine learning is transforming many industries. It is being increasingly employed to solve a wide range of complex problems, producing autonomous systems that support human decision-making. For the aircraft engine industry, machine learning of historical and existing engine data could provide insights that help drive for better engine design. This work explored the application of machine learning to engine preliminary design. Engine core-size prediction was chosen for the first study because of its relative simplicity in terms of number of input variables required (only three). Specifically, machine-learning predictive tools were developed for turbofan engine core-size prediction, using publicly available data of two hundred manufactured engines and engines that were studied previously in NASA aeronautics projects. The prediction results of these models show that, by bringing together big data, robust machine-learning algorithms and automation, a machine learning-based predictive model can be an effective tool for turbofan engine core-size prediction. The promising results of this first study paves the way for further exploration of the use of machine learning for aircraft engine preliminary design.

Tong, Michael T.↗

Using Machine Learning to Predict Core Sizes of High-Efficiency Turbofan Engines

With the rise in big data and analytics, machine learning is transforming many industries. It is being increasingly employed to solve a wide range of complex problems, producing autonomous systems that support human decision-making. For the aircraft engine industry, machine learning of historical and existing engine data could provide insights that help drive for better engine design. This work explored the application of machine learning to engine preliminary design. Engine core-size prediction was chosen for the first study because of its relative simplicity in terms of number of input variables required (only three). Specifically, machine-learning predictive tools were developed for turbofan engine core-size prediction, using publicly available data of two hundred manufactured engines and engines that were studied previously in NASA aeronautics projects. The prediction results of these models show that, by bringing together big data, robust machine-learning algorithms and data science, a machine learning-based predictive model can be an effective tool for turbofan engine core-size prediction. The promising results of this first study paves the way for further exploration of the use of machine learning for aircraft engine preliminary design.

Core Size↗

GOOML - Finding Optimization Opportunities for Geothermal Operations: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach. We have used this framework to develop digital twins that provide steamfield operators with an operational environment to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management for real world applications. The GOOML modeling software is built on a generic component-based systems framework that allows for both historical and forecast analysis. A GOOML model can perform historical data-assimilation using first-principal thermodynamics to create a meaningful data model. Historical production data can then be coupled with a forecast framework to train machine-learning models of steamfield components to predict future outputs. This modeling environment enables digital exploration of steamfield design configurations and operational scenarios. GOOML digital twins have been developed for steamfields in New Zealand and the United States representing differing power generation and field conditions. These digital twins have been validated by comparing hindcast predictions against historical production data. Reinforcement learning experiments were conducted to demonstrate the ability to programmatically explore the operations space using machine learning agents. Our initial results are compelling; two to five percent increases in annual energy production were demonstrated by the GOOML models with no additional infrastructure build required. GOOML offers a new approach to geothermal operations by applying state-of-the-art machine learning algorithms, comprehensive data analytics, and interaction with digital twins. Through application of these tools, operators will realize greater availability and higher net generation which will increase the cost effectiveness of geothermal energy projects.

access↗

Investigating Access Performance of Long Time Series with Restructured Big Model Data

Data sets generated by models are substantially increasing in volume, due to increases in spatial and temporal resolution, and the number of output variables. Many users wish to download subsetted data in preferred data formats and structures, as it is getting increasingly difficult to handle the original full-size data files. For example, application research users such as those involved with wind or solar energy, or extreme weather events are likely only interested in daily or hourly model data at a single point (or for a small area) for a long time period, and prefer to have the data downloaded in a single file. With native model file structures, such as hourly data from NASA Modern-Era Retrospective analysis for Research and Applications Version-2 (MERRA-2), it may take over 10 hours for the extraction of parameters-of-interest at a single point for 30 years. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) is exploring methods to address this particular user need. One approach is to create value-added data by reconstructing the data files. Taking MERRA-2 data as an example, we have tested converting hourly data from one-day-per-file into different data cubes, such as one-month, or one-year. Performance is compared for reading local data files and accessing data through interoperable services, such as OPeNDAP. Results show that, compared to the original file structure, the new data cubes offer much better performance for accessing long time series. We have noticed that performance is associated with the cube size and structure, the compression method, and how the data are accessed. An optimized data cube structure will not only improve data access, but also may enable better online analysis services

reanalysis↗

Unleashing the Power of Industrial Big Data through Scalable Manual Labeling

Big Data plays a central role in the remarkable results achieved by Machine Learning (ML) and especially Deep Learning (DL) in the recent years. However, the difficulty in obtaining a reasonable amount of labeled samples limits ML/DL application in various domains, including industrial equipment and system monitoring. In this paper the need for methods that turn manual labeling into a scalable process is highlighted. A real world problem is analyzed for which weak supervision methods, successfully employed in other domains, did not produce acceptable results. An alternative approach based on clustering ensembles is described and tested, achieving good performance.

Paes Leao, Bruno↗

Anticipated Changes in Conducting Scientific Data-Analysis Research in the Big-Data Era

A Big-Data environment is one that is capable of orchestrating quick-turnaround analyses involving large volumes of data for numerous simultaneous users. Based on our experiences with a prototype Big-Data analysis environment, we anticipate some important changes in research behaviors and processes while conducting scientific data-analysis research in the near future as such Big-Data environments become the mainstream. The first anticipated change will be the reduced effort and difficulty in most parts of the data management process. A Big-Data analysis environment is likely to house most of the data required for a particular research discipline along with appropriate analysis capabilities. This will reduce the need for researchers to download local copies of data. In turn, this also reduces the need for compute and storage procurement by individual researchers or groups, as well as associated maintenance and management afterwards. It is almost certain that Big-Data environments will require a different "programming language" to fully exploit the latent potential. In addition, the process of extending the environment to provide new analysis capabilities will likely be more involved than, say, compiling a piece of new or revised code.We thus anticipate that researchers will require support from dedicated organizations associated with the environment that are composed of professional software engineers and data scientists. A major benefit will likely be that such extensions are of higherquality and broader applicability than ad hoc changes by physical scientists. Another anticipated significant change is improved collaboration among the researchers using the same environment. Since the environment is homogeneous within itself, many barriers to collaboration are minimized or eliminated. For example, data and analysis algorithms can be seamlessly shared, reused and re-purposed. In conclusion, we will be able to achieve a new level of scientific productivity in the Big-Data analysis environments.

Kuo, Kwo-Sen↗

Investigation and Evaluation of Advanced Spectrum Management Concepts for Aeronautical Communications

With the emergence of new aerial vehicles into the airspace and the continued growth of aviation operations, there will be an increasing demand for voice and data communications within the National Airspace System (NAS). The continued use of existing VHF and UHF frequency allocations is not a sustainable approach, and as a result, a new spectrum management solution is required to support future mission needs. The proposed spectrum management concepts leverage modern advancements such as artificial intelligence (AI) and big data to dynamically optimize the spectrum utilization based on the predicted communications demand throughout the airspace. This technical investigation considers both air-ground and air-air communications networks, and can be applied to both existing applications such as the air traffic control (ATC) system, as well as future applications, such as the emerging Advanced Air Mobility (AAM). To support the evaluation of the proposed concepts and technologies, a modeling and simulation capability is currently under development and will continue to evolve to support new and advanced airspace applications. It is anticipated that this proposed spectrum concept will better serve the spectrum needs of future NAS applications.

Eric J Knoblock↗

Investigation and Evaluation of Advanced Spectrum Management Concepts for Aeronautical Communications

With the emergence of new aerial vehicles into the airspace and the continued growth of aviation operations, there will be an increasing demand for voice and data communications within the National Airspace System (NAS). The continued use of existing VHF and UHF frequency allocations is not a sustainable approach, and as a result, a new spectrum management solution is required to support future mission needs. The proposed spectrum management concepts leverage modern advancements such as artificial intelligence (AI) and big data to dynamically optimize the spectrum utilization based on the predicted communications demand throughout the airspace. This technical investigation considers both air-ground and air-air communications networks, and can be applied to both existing applications such as the air traffic control (ATC) system, as well as future applications, such as the emerging Advanced Air Mobility (AAM). To support the evaluation of the proposed concepts and technologies, a modeling and simulation capability is currently under development and will continue to evolve to support new and advanced airspace applications. It is anticipated that this proposed spectrum concept will better serve the spectrum needs of future NAS applications.

Eric J Knoblock↗

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes↗

Special Issue: Geostatistics and Machine Learning

Abstract Recent years have seen a steady growth in the number of papers that apply machine learning methods to problems in the earth sciences. Although they have different origins, machine learning and geostatistics share concepts and methods. For example, the kriging formalism can be cast in the machine learning framework of Gaussian process regression. Machine learning, with its focus on algorithms and ability to seek, identify, and exploit hidden structures in big data sets, is providing new tools for exploration and prediction in the earth sciences. Geostatistics, on the other hand, offers interpretable models of spatial (and spatiotemporal) dependence. This special issue on Geostatistics and Machine Learning aims to investigate applications of machine learning methods as well as hybrid approaches combining machine learning and geostatistics which advance our understanding and predictive ability of spatial processes.

58 GEOSCIENCES↗

Modeling and Simulation of an Integrated Gate Turnaround Management Concept

An Integrated Gate Turnaround Management (IGTM) prototype was developed at NASA Ames Simulation Laboratories (SimLabs) using Dallas Ft. Worth International Airport (DFW) to demonstrate the IGTM concepts feasibility and benefits. The simulation architecture includes: the IGTM controller, an Airline Operations Control (AOC) application, Big DataAnalytics Input (BAI) application, a terminal traffic simulation or known as NASA-developed Surface Operation Simulator and Scheduler (SOSS), and a Database Server. ActiveMQ, a Java messaging service, was used to emulate the System Wide Information Management (SWIM) data network messaging. This paper describes the modeling and simulation of the IGTM concept, and illustrates selected use cases to demonstrate the feasibility and benefits of the IGTM concept for optimizing gate turnaround operations.

gate turnaround management↗

Machine and Deep Learning: Artificial Intelligence Application in Biotic and Abiotic Stress Management in Plants

Biotic and abiotic stresses significantly affect plant fitness, resulting in a serious loss in food production. Biotic and abiotic stresses predominantly affect metabolite biosynthesis, gene and protein expression, and genome variations. However, light doses of stress result in the production of positive attributes in crops, like tolerance to stress and biosynthesis of metabolites, called hormesis. Advancement in artificial intelligence (AI) has enabled the development of high-throughput gadgets such as high-resolution imagery sensors and robotic aerial vehicles, i.e., satellites and unmanned aerial vehicles (UAV), to overcome biotic and abiotic stresses. These High throughput (HTP) gadgets produce accurate but big amounts of data. Significant datasets such as transportable array for remotely sensed agriculture and phenotyping reference platform (TERRA-REF) have been developed to forecast abiotic stresses and early detection of biotic stresses. For accurately measuring the model plant stress, tools like Deep Learning (DL) and Machine Learning (ML) have enabled early detection of desirable traits in a large population of breeding material and mitigate plant stresses. In this review, advanced applications of ML and DL in plant biotic and abiotic stress management have been summarized.

59 BASIC BIOLOGICAL SCIENCES↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

Introduction to the special issue on smart transportation

Transportation is getting smarter and smarter, with the prominence of connected automated vehicle technologies in the global auto industry’s near-term growth strategies, of big data analytics and unprecedented access to sensing data of mobility, and of integration of this analytics into the optimization of mobility and transport. Further, these developments are setting off a wave of smart transportation innovations, which are featured by new methods and applications driven by various forms of sensor data such as GPS, CAN bus, LIDA, images, etc. At the same time, complexities surrounding the use, conflation, and processing of disparate data in near real-time is of essence for the design and development of futuristic smart transportation.

33 ADVANCED PROPULSION SYSTEMS↗

A Spatiotemporal Indexing Approach for Efficient Processing of Big Array-Based Climate Data with MapReduce

Climate observations and model simulations are producing vast amounts of array-based spatiotemporal data. Efficient processing of these data is essential for assessing global challenges such as climate change, natural disasters, and diseases. This is challenging not only because of the large data volume, but also because of the intrinsic high-dimensional nature of geoscience data. To tackle this challenge, we propose a spatiotemporal indexing approach to efficiently manage and process big climate data with MapReduce in a highly scalable environment. Using this approach, big climate data are directly stored in a Hadoop Distributed File System in its original, native file format. A spatiotemporal index is built to bridge the logical array-based data model and the physical data layout, which enables fast data retrieval when performing spatiotemporal queries. Based on the index, a data-partitioning algorithm is applied to enable MapReduce to achieve high data locality, as well as balancing the workload. The proposed indexing approach is evaluated using the National Aeronautics and Space Administration (NASA) Modern-Era Retrospective Analysis for Research and Applications (MERRA) climate reanalysis dataset. The experimental results show that the index can significantly accelerate querying and processing (10 speedup compared to the baseline test using the same computing cluster), while keeping the index-to-data ratio small (0.0328). The applicability of the indexing approach is demonstrated by a climate anomaly detection deployed on a NASA Hadoop cluster. This approach is also able to support efficient processing of general array-based spatiotemporal data in various geoscience domains without special configuration on a Hadoop cluster.

big data↗

PanDA: Production and Distributed Analysis System

The Production and Distributed Analysis (PanDA) system is a data-driven workload management system engineered to operate at the LHC data processing scale. The PanDA system provides a solution for scientific experiments to fully leverage their distributed heterogeneous resources, showcasing scalability, usability, flexibility, and robustness. The system has successfully proven itself through nearly two decades of steady operation in the ATLAS experiment, addressing the intricate requirements such as diverse resources distributed worldwide at about 200 sites, thousands of scientists analyzing the data remotely, the volume of processed data beyond the exabyte scale, dozens of scientific applications to support, and data processing over several billion hours of computing usage per year. PanDA’s flexibility and scalability make it suitable for the High Energy Physics community and wider science domains at the Exascale. Beyond High Energy Physics, PanDA’s relevance extends to other big data sciences, as evidenced by its adoption in the Vera C. Rubin Observatory and the sPHENIX experiment. As the significance of advanced workflows continues to grow, PanDA has transformed into a comprehensive ecosystem, effectively tackling challenges associated with emerging workflows and evolving computing technologies. The paper discusses PanDA’s prominent role in the scientific landscape, detailing its architecture, functionality, deployment strategies, project management approaches, results, and evolution into an ecosystem.

97 MATHEMATICS AND COMPUTING↗