Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Machine Learning Lifecycle for Earth Science Application: A Practical Insight into Production Deployment

Earth science domain presents unique sets of problems that are increasingly being solved using data driven approaches. The availability of big Earth science data offers immense potential for Machine learning (ML) as evident from numerous research publications lately. However, many of these publications are not ending up as production applications mainly because the data scientists who develop the ML models are now expected to complete the ML lifecycle by deploying and scaling the models in production. We introduce ML lifecycle to the Earth science community including the opportunities and challenges that lie ahead in each phase of the lifecycle. We demonstrate the lifecycle using an Earth science problem that we used ML to address and transitioned to production.

Maskey, Manil↗

Topological Optimization with Big Steps

Using persistent homology to guide optimization has emerged as a novel application of topological data analysis. Existing methods treat persistence calculation as a black box and backpropagate gradients only onto the simplices involved in particular pairs. We show how the cycles and chains used in the persistence calculation can be used to prescribe gradients to larger subsets of the domain. In particular, we show that in a special case, which serves as a building block for general losses, the problem can be solved exactly in linear time. This relies on another contribution of this paper, which eliminates the need to examine a factorial number of permutations of simplices with the same value. Here, we present empirical experiments that show the practical benefits of our algorithm: the number of steps required for the optimization is reduced by an order of magnitude.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Big Data Meets Geothermal Exploration (CRADA Final Report)

As part of the Cyclotron Road program, Zanskar Geothermal & Minerals, Inc. investigated the application of micro-earthquake and ambient noise seismology methods to imaging and characterizing the structural characteristics and hydrothermal flux of subsurface faults. Significant advances in what could be resolved were enabled by two major developments in seismology: 1) the availability of large-n arrays of low-cost seismometers, and 2) the availability of increased computational power and semi-automated data reduction algorithms. In tandem, these advances may improve the signal-to-noise ratio and spatial precision of the data collected and enable higher-resolution characterization of subsurface fracture systems and their spatio-temporal evolution. These tools supported efforts to reduce dry-hole risk and to improve wellfield productivity for geothermal resource development. In particular, two applications of these advances were evaluated: 1) fracture-seismic imaging, which was used to detect ambient emissions from fluid-filled fractures, and 2) reservoir tomography, which used information about travel paths, source locations, and source parameters of micro-earthquakes to identify areas of enhanced permeability. Integration of these methods provided guidance for siting wells and served as prior constraints for reservoir models, informing forecasts of power potential and production and injection strategies aimed at minimizing temperature decline and improving overall resource productivity.

15 GEOTHERMAL ENERGY↗

Machine Learning Technologies and Their Applications for Science and Engineering Domains Workshop -- Summary Report

The fields of machine learning and big data analytics have made significant advances in recent years, which has created an environment where cross-fertilization of methods and collaborations can achieve previously unattainable outcomes. The Comprehensive Digital Transformation (CDT) Machine Learning and Big Data Analytics team planned a workshop at NASA Langley in August 2016 to unite leading experts the field of machine learning and NASA scientists and engineers. The primary goal for this workshop was to assess the state-of-the-art in this field, introduce these leading experts to the aerospace and science subject matter experts, and develop opportunities for collaboration. The workshop was held over a three day-period with lectures from 15 leading experts followed by significant interactive discussions. This report provides an overview of the 15 invited lectures and a summary of the key discussion topics that arose during both formal and informal discussion sections. Four key workshop themes were identified after the closure of the workshop and are also highlighted in the report. Furthermore, several workshop attendees provided their feedback on how they are already utilizing machine learning algorithms to advance their research, new methods they learned about during the workshop, and collaboration opportunities they identified during the workshop.

Ambur, Manjula↗

An overview of the operation architectures and energy management system for multiple microgrid clusters

The emerging novel energy infrastructures, such as energy communities, smart building-based microgrids, electric vehicles enabled mobile energy storage units raise the requirements for a more interconnective and interoperable energy system. It leads to a transition from simple and isolated microgrids to relatively large-scale and complex interconnected microgrid systems named multi-microgrid clusters. In order to efficiently, optimally, and flexibly control multi-microgrid clusters, cross-disciplinary technologies such as power electronics, control theory, optimization algorithms, information and communication technologies, cyber-physical, and big-data analysis are needed. This paper introduces an overview of the relevant aspects for multi-microgrids, including the outstanding features, architectures, typical applications, existing control mechanisms, as well as the challenges.

24 POWER TRANSMISSION AND DISTRIBUTION↗

How Bayesian methods can improve R -matrix analyses of data: The example of the d t reaction

The 3 H(d, n) 4 He reaction is of significant interest in nuclear astrophysics and nuclear applications. It is an important, early step in big-bang nucleosynthesis and a key process in nuclear fusion reactors. We use one- and two-level R-matrix approximations to analyze data on the cross section for this reaction at center-of-mass energies below 215 keV. We critically examine the data sets using a Bayesian statistical model that allows for both common-mode and additional point-to-point un- certainties. We use Markov Chain Monte Carlo sampling to evaluate this R-matrix-plus-statistical model and find two-level R-matrix results that are stable with respect to variations in the channel radii. The S factor at 40 keV evaluates to 25.36(19) MeV b (68% credibility interval). We discuss our Bayesian analysis in detail and provide guidance for future applications of Bayesian methods to R-matrix analyses. We also discuss possible paths to further reduction of the S-factor uncertainty.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Electricity use in big area additive manufacturing of fiber-reinforced polymer composites

In recent years, additive manufacturing (AM), especially large-format additive manufacturing (LFAM), has gained momentum in the manufacturing industry. While LFAM offers benefits over conventional manufacturing processes, such as minimizing material waste and providing vast geometric freedom, assessing its sustainability remains challenging due to limited data, particularly on energy consumption. Most existing data pertain to small-scale or desktop AM and are not directly applicable to LFAM. In this study, we conducted real-time measurements of electricity usage for a type of LFAM known as big area additive manufacturing (BAAM), which typically uses fiber-reinforced polymer pellets as feedstock. We collected electricity usage data from fifteen printing jobs over two months in an industrial production setting. These data fill the existing gap and can be reused to enhance the community’s understanding of LFAM electricity usage, support further research, and promote sustainable development in advanced manufacturing technologies.

ecology↗

Transfer-Learnt Energy Models for Predicting Electricity Consumption in Buildings with Limited and Sparse Field Data

Modeling energy consumption is critical for energy-efficient utilization of the electric appliances in a building, smart grid programs (like demand-response), and many other smart home applications. State-of-the-art energy modeling techniques either rely on theoretical models, or extensive instrumentation of the building envelope to gather ``big" data to train a deep neural network. While theoretical models are often limited by their estimation accuracy, it is not always feasible to gather a significant amount of field data. In this paper, we explore transfer learning-based strategies to train much more accurate model for energy estimation when using a sparse field data. We transferred knowledge, in the form of data and parameters, from the simulation framework to the field data. We evaluated the efficacy of our approach on field data collected from six commercial buildings and our results indicate that transfer learning-based models trained over one month data can perform comparative (and in some cases better) than the state-of-the-art machine learning and deep learning solutions.

Jain, Milan↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Open Source Visualization and Analysis Platform for 3D Reconstructions of Materials by Transmission Electron Microscopy

Three-dimensional characterization of materials at the nano- and meso-scale has become possible with transmission and scanning transmission electron microscopes (S/TEM). Its importance has extended to a wide class of nanomaterials such as hydrogen fuel cells, solar cells, industrial catalysts, new battery materials and semiconductor devices, as well as spanning high-tech industry, universities, and national labs. While capable instrumentation is abundant, this rapidly expanding demand for high-resolution tomography is bottlenecked by software that is instead tailored for lower-dose, biological applications and not optimized for higher-resolution materials applications. Existing tomography fails to utilize the chemical information provided by S/TEM spectrometers. To address this problem, this project delivered a scalable, fully functional, freely-distributable, open Source materials tomography package with a modern user interface that enables automated acquisition, alignment, and real-time reconstruction of raw tomography data, and provides advanced segmentation, three-dimensional chemical visualization and analysis optimized for materials applications. It has established an extendable framework capable of automation for high-throughput for the tomography of materials from data acquisition to visualization. Phase I and II delivered a full-featured, cross-platform, clean and integrated application for S/TEM materials tomography. It can read projection data from the microscope, with graphical tools for alignment of data, tomographic reconstruction, segmentation, and visualization of the reconstructed 3D volume. The entire pipeline can be saved to an XML state file, enabling fully reproducible data processing and analysis with a full Python environment for the development of custom algorithms, processing, and analysis requirements. Phase IIB extended the application to provide real-time tomographic reconstruction, "live updates" as data is processed, analysis of big data, and multi-channel tomography. Here, quantitative assessment of nanomaterials can occur as data is being recorded on a 3D visualization platform that accommodates additional information from multiple channels. We provide unmatched high-throughput tomography, where the complete tomographic pipeline from measurement to 3D visualization can occur rapidly. The project was extended to offer capabilities for micro-CT, X-ray, neutron, focused ion beam, atomic electron tomography and atom probe tomography. With over 600 transmission electron microscopes worldwide and approximately 50 coming online each year, the demand and impact of an open-source tomography tool is large. Significant opportunities exist in high-tech industry, universities, and national labs to enable or enhance three-dimensional imaging at the nanoscale and bring automated high-throughput approaches that will accelerate progress in materials characterization and metrology. The project supports a service based business model by enabling lab-specific acquisition and processing customization and integration-support and development that will be provided into Phase III and beyond.

36 MATERIALS SCIENCE↗

7th World Congress on Integrated Computational Materials Engineering (ICME 2023) (Final Technical Report)

Integrated Computational Materials Engineering (ICME) has received international attention due to its potential to shorten product development time, while lowering cost and improving design and manufacturing outcomes. ICME is an approach to designing materials solutions for specific applications that use computer modeling programs to predict the behavior of materials and integrate this information into the overall materials, processing, and manufacturing design cycle. The 7th World Congress on Integrated Computational Materials Engineering (ICME 2023) was held in Orlando, Florida from May 21–25, 2023 with the goal to convene stakeholders from across all areas of modeling and simulation, experimental specialization, and design, as well as from across academia, government, and industry, to address ICME tools and techniques and their integration, as well as to examine their application in engineering. This atmosphere facilitated rich interactions between the experimentalists, modelers, and computational and design, from academia, government, and industry, to discuss ICME tools and techniques and their application in engineering.

36 MATERIALS SCIENCE↗

Chapter 4: Packings, simulation, and big data--artificial intelligence emulation of soft matter

At the suggestion of NASA’s Physical Science Research Program in the Space Life and Physical Science Research and Application Division, Paul Chaikin, Noel Clark, and Sidney Nagel organized a focus session and workshop for the 2020 American Physical Society (APS) March meeting under the auspices of the Division of Soft Matter. Three overarching themes emerged from the workshop and are presented with additional details: • Machines made out of machines • Scalable self-sustaining ecosystems • Active materials and metamaterials This report lays out only some of the potential directions for soft matter dynamics over the next two decades. It also lays out the role that gravity plays in the organization of the basic building blocks of matter. Not only will research on soft matter have tremendous application towards understanding its behavior in our terrestrial environment, but also potentially in other NASA programs such as planetary science, exploration, robotics, etc. Attached is a White Paper for the Decadal Survey that consists of an extended Title along with the previous Introduction and Chapter 2.4 from NASA/CP-20205010493.

Soft matter↗

Applications of digital image analysis capability in Idaho

The use of digital image analysis of LANDSAT imagery in water resource assessment is discussed. The data processing systems employed are described. The determination of urban land use conversion of agricultural land in two southwestern Idaho counties involving estimation and mapping of crop types and of irrigated land is described. The system was also applied to an inventory of irrigated cropland in the Snake River basin and establishment of a digital irrigation water source/service area data base for the basin. Application of the system to a determination of irrigation development in the Big Lost River basin as part of a hydrologic survey of the basin is also described.

Johnson, K. A.↗

A Survey of Singular Value Decomposition Methods for Distributed Tall/Skinny Data

The Singular Value Decomposition (SVD) is one of the most important matrix factorizations, enjoying a wide variety of applications across numerous application domains. In statistics and data analysis, the common applications of SVD inclue Principal Components Analysis (PCA) and regression. Usually these applications arise on data that has far more rows than columns, so-called "tall/skinny" matrices. In the big data analytics context, this may take the form of hundreds of millions to billions of rows with only a few hundred columns. There is a need, therefore, for fast, accurate, and scalable tall/skinny SVD implementations which can fully utilize modern computing resources. To that end, we present a survey of three different algorithms for computing the SVD for these kinds of tall/skinny data layouts using MPI for communication. We contextualize these with common big data analytics techniques. Finally, we present both CPU and GPU timing results from the Summit supercomputer, and discuss possible alternative approaches.

Schmidt, Drew↗

An Integrated Gate Turnaround Management Concept Leveraging Big Data Analytics for NAS Performance Improvements

"Gate Turnaround" plays a key role in the National Air Space (NAS) gate-to-gate performance by receiving aircraft when they reach their destination airport, and delivering aircraft into the NAS upon departing from the gate and subsequent takeoff. The time spent at the gate in meeting the planned departure time is influenced by many factors and often with considerable uncertainties. Uncertainties such as weather, early or late arrivals, disembarking and boarding passengers, unloading/reloading cargo, aircraft logistics/maintenance services and ground handling, traffic in ramp and movement areas for taxi-in and taxi-out, and departure queue management for takeoff are likely encountered on the daily basis. The Integrated Gate Turnaround Management (IGTM) concept is leveraging relevant historical data to support optimization of the gate operations, which include arrival, at the gate, departure based on constraints (e.g., available gates at the arrival, ground crew and equipment for the gate turnaround, and over capacity demand upon departure), and collaborative decision-making. The IGTM concept provides effective information services and decision tools to the stakeholders, such as airline dispatchers, gate agents, airport operators, ramp controllers, and air traffic control (ATC) traffic managers and ground controllers to mitigate uncertainties arising from both nominal and off-nominal airport gate operations. IGTM will provide NAS stakeholders customized decision making tools through a User Interface (UI) by leveraging historical data (Big Data), net-enabled Air Traffic Management (ATM) live data, and analytics according to dependencies among NAS parameters for the stakeholders to manage and optimize the NAS performance in the gate turnaround domain. The application will give stakeholders predictable results based on the past and current NAS performance according to selected decision trees through the UI. The predictable results are generated based on analysis of the unique airport attributes (e.g., runway, taxiway, terminal, and gate configurations and tenants), and combined statistics from past data and live data based on a specific set of ATM concept-of-operations (ConOps) and operational parameters via systems analysis using an analytic network learning model. The IGTM tool will then bound the uncertainties that arise from nominal and off-nominal operational conditions with direct assessment of the gate turnaround status and the impact of a certain operational decision on the NAS performance, and provide a set of recommended actions to optimize the NAS performance by allowing stakeholders to take mitigation actions to reduce uncertainty and time deviation of planned operational events. An IGTM prototype was developed at NASA Ames Simulation Laboratories (SimLabs) to demonstrate the benefits and applicability of the concept. A data network, using the System Wide Information Management (SWIM)-like messaging application using the ActiveMQ message service, was connected to the simulated data warehouse, scheduled flight plans, a fast-time airport simulator, and a graphic UI. A fast-time simulation was integrated with the data warehouse or Big Data/Analytics (BAI), scheduled flight plans from Aeronautical Operational Control AOC, IGTM Controller, and a UI via a SWIM-like data messaging network using the ActiveMQ message service, illustrated in Figure 1, to demonstrate selected use-cases showing the benefits of the IGTM concept on the NAS performance.

Efficent ATM systems↗

Hardening DOE R&D Software Tools for Web-based Visualization SBIR Phase I Final Report

Ubiquitous web-based visualization is essential to delivering large-scale data visualization to various stakeholders, from the scientist to the board member. These stakeholders will not tolerate a stalled application or a pop-up window asking them to wait for the processing to complete. They require a responsive and interactive visualization environment with high-quality imagery suitable for detailed analysis and boardroom presentations. At Kitware, Inc., we have accomplished web visualization to this point, leveraging state-of-the-art tools like HTML5, CSS3, SVG, Canvas, and WebGL. Solutions that leverage a combination of these technologies are necessary to handle workloads that vary significantly in data size efficiently. However, it is not always practical to move large data to the web client for visualization. Kitware's ParaView as a Service combines client-side visualization using both distributed processing and remote rendering on big data impractical to move. Existing distributed processing and remote rendering solution's interactivity is below the expectations of web-based applications. Our project examined proposed solutions to the areas outlined above in ParaView as a Service. We have investigated concurrent pipelines, streaming images, progressive rendering, and optimization of algorithms and data movement to address these concerns. For the Phase I project, we completed the proposed work plan. As a result, the project produced three prototypes of essential importance for web visualization and the ParaView as a Service community. We created a simple desktop application for an interactive streamline placement prototype, a web-based interactive streamline placement prototype, and a web-based progressive rendering utilizing raytracing prototype. These prototypes relied on the hardening of emerging software toolkits funded by the Department of Energy (DOE) Advanced Scientific Computing Research (ASCR) program (such as ParaView, VTK-m, and Mochi). We blended these components into web-based visualization prototypes that meet the industry's expectations for interactivity and responsiveness. The Phase I project had four essential focus areas: 1. Develop prototype ParaView as a Service backend server using asynchronous, non-blocking design principles. 2. Develop a prototype web application that uses the ParaView as a Service backend server for remote data visualization. 3. Implement image streaming with encoding/compression and progressive rendering capabilities in the proposed platform. 4. Evaluate the prototype developed and summarize observations, including the challenges and pitfalls of our approach. After our successful completion of Phase I, we are strongly positioned to propose a successful Phase II project.

Geveci, Berk↗

Through the lens of bioenergy crops: advances, bottlenecks, and promises of plant engineering

Advances in engineering of bioenergy crops were driven over the past years by adapting technological breakthroughs and accelerating conventional applications but also exposed intriguing challenges. New tools revealed rich interconnectivity in the exponentially growing and dynamic 'big' omics data' of metabolomes, transcriptomes, and genomes at previously inaccessible magnitude (global, cross-species, meta-) and resolution (single cell). Insights enabled fresh hypotheses and stimulated disciplines such as functional genomics with discovery of broad regulatory networks and their determinants, that is, DNA parts, including promoters, regulatory elements, and transcription factors. Their rational design, assembly into increasingly complex blueprints, and installation into diverse chassis is an existing frontier that may benefit from emerging technologies to address bottlenecks. Interweaving nature-inspired to fully synthetic parts has already allowed building of fine-tuned regulatory circuits, or new-to-nature metabolic routes insulated from the biological context of the chassis species. Similarly, developments and the evolving need for unifying principles in plant transformation and species-agnostic technologies highlight future opportunities for engineering the next generation of bioenergy plants.

60 APPLIED LIFE SCIENCES↗

Analytics-at-scale of Sensor Data for Digital Monitoring in Nuclear Plants (3 rd Annual Report)

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This report primarily focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 ± 0.014, 0.0026 ± 0, and 0.063 ± 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure. This report summarizes the Fiscal Year 2021 research progress encompassing the (1) data cleaning and feature selection necessary for ML applications; (2) development of short-term forecasting models to predict future plant process parameters for both single and multiple time steps ahead; and (3) validation of the feature selection methods and short-term forecasting models given new data from different systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗