Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

Introduction to the special issue on smart transportation

Transportation is getting smarter and smarter, with the prominence of connected automated vehicle technologies in the global auto industry’s near-term growth strategies, of big data analytics and unprecedented access to sensing data of mobility, and of integration of this analytics into the optimization of mobility and transport. Further, these developments are setting off a wave of smart transportation innovations, which are featured by new methods and applications driven by various forms of sensor data such as GPS, CAN bus, LIDA, images, etc. At the same time, complexities surrounding the use, conflation, and processing of disparate data in near real-time is of essence for the design and development of futuristic smart transportation.

33 ADVANCED PROPULSION SYSTEMS↗

PanDA: Production and Distributed Analysis System

The Production and Distributed Analysis (PanDA) system is a data-driven workload management system engineered to operate at the LHC data processing scale. The PanDA system provides a solution for scientific experiments to fully leverage their distributed heterogeneous resources, showcasing scalability, usability, flexibility, and robustness. The system has successfully proven itself through nearly two decades of steady operation in the ATLAS experiment, addressing the intricate requirements such as diverse resources distributed worldwide at about 200 sites, thousands of scientists analyzing the data remotely, the volume of processed data beyond the exabyte scale, dozens of scientific applications to support, and data processing over several billion hours of computing usage per year. PanDA’s flexibility and scalability make it suitable for the High Energy Physics community and wider science domains at the Exascale. Beyond High Energy Physics, PanDA’s relevance extends to other big data sciences, as evidenced by its adoption in the Vera C. Rubin Observatory and the sPHENIX experiment. As the significance of advanced workflows continues to grow, PanDA has transformed into a comprehensive ecosystem, effectively tackling challenges associated with emerging workflows and evolving computing technologies. The paper discusses PanDA’s prominent role in the scientific landscape, detailing its architecture, functionality, deployment strategies, project management approaches, results, and evolution into an ecosystem.

97 MATHEMATICS AND COMPUTING↗

Operation and performance of VRF systems: Mining a large-scale dataset

The energy consumption of air-conditioning systems has gained increasing attention as it contributes significantly to the global building energy use. The variable refrigerant flow (VRF) system is a common air-conditioning system applied widely in residential and office buildings in China. Understanding the actual operation and performance of VRF systems is fundamental for the energy-efficient design and operation of VRF systems. Previous research on VRF system operation used either limited field data covering certain building types and climate zones or used a questionnaire to obtain a larger dataset. However, they did not capture the wide applications of VRF systems quantitatively across all building types, climate zones, and operating conditions. To fill this gap, statistical and clustering analysis was conducted on the newly proposed key performance indicators of approximately 287,000 VRF systems for residential and commercial buildings in all five climate zones in China. In this work, the main findings are: (1) VRF systems are mainly used for cooling in all climate zones in China; (2) among all building types, the duration of use is lowest in residential buildings and highest in hotels and medical buildings; (3) the distribution of the ideal VRF cooling coefficient of performance (COP) is similar across all climate zones and building types; whereas the COPs of ideal VRF heating in the Severe Cold region and Cold regions are lower than those in other climate zones; and (4) partial load operations for VRF systems are common in residential buildings and office buildings due to the part-time-part-space operation mode. These findings can inform the actual application of VRF systems in China, supporting the design, operation, industry standard development, and performance optimization of VRF systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

GOOML Big Kahuna Forecast Modeling and Genetic Optimization Files

This submission includes example files associated with the Geothermal Operational Optimization using Machine Learning (GOOML) Big Kahuna fictional power plant, which uses synthetic data to model a fictional power plant. A forecast was produced using the GOOML data model framework and fictional input data, and a genetic optimization is included which determines optimal flash plant parameters. The inputs and outputs associated with the forecast and genetic optimization are included. The input and output files consist of data, configuration files, and plots. A link to the Physics-Guided Neural Networks (phygnn) GitHub repository is also included, which augments a traditional neural network loss function with a generic loss term that can be used to guide the neural network to learn physical or theoretical constraints. phygnn is used by the GOOML framework to help integrate its machine learning models into the relevant physics and engineering applications. Note that the data included in this submission are intended to provide a demonstration of GOOML's capabilities. Additional files that have not been released to the public are needed for users to run these models and reproduce these results. Units can be found in the readme data resource.

15 GEOTHERMAL ENERGY↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗

Raman Microscopy Analysis of Wyoming CarbonSAFE Pilot Well Thin Sections for Mineralogy and Organic Matter Characterization

Scanning confocal Raman microspectroscopy (RMS) was used to analyze thin sections for inorganic mineral content and for evaluation of organic material (OM). Thin sections were selected from a stratigraphic test well core in the Powder River Basin, Wyoming under the Wyoming CarbonSAFE project. The well was analyzed with a variety of rock and fluid characterization techniques to determine the feasibility of a commercial-scale CO2 storage site. RMS complements other analyses, including traditional petrography, SEM, porosity and permeability, and XRD For this aspect of the study, RMS is especially important in the evaluation of OM in sealing lithologies. At prospective geologic CO2 storage sites, it is imperative to assess the unconventional oil and gas potential of seals to ensure that any future development would not compromise the integrity of the seals. Whole-slide mineralogical surveys were performed on thin sections from various shale and sandstone formations. Surveys were analyzed with Direct Classical Least Squares to identify and quantify minerals. The location and concentration of minerals was color-coded and overlaid on optical images for visualization of the distribution of minerals. Dense hyperspectral Raman mapping of OM was performed on five thin sections. Eleven spectral parameters diagnostic of organic type and thermal maturity were used to train a Partial Least Squares (PLS) calibration against a set of artificially matured samples spanning the pre- to mid-oil window. The PLS was applied to the study set and a post-mature set. Additionally, the PLS was applied to each point in hyperspectral maps for visualization of trends in maturity across sample sets and discrimination of organic matter types within a given map. In inorganic surveys on thin sections, a total of 14 unique inorganic minerals were identified in Raman spectra including quartz, dolomite, calcite, hematite and anhydrite. Shale thin sections tended to be dominated by organic material. OM was often observed mixed with inorganic minerals. Sand- and mudstones were dominated by inorganic minerals. The PLS extrapolated the post-mature set to reflectances >1.2%. The study set ranged from very immature to postmature in the median of map fit-peak parameters. However, point maturity maps indicate that matrix OM in all study samples is immature and that discreet organic particles selected for mapping, which may be inertinites, bias medians towards more-mature. The work demonstrates the capabilities of RMS to perform both whole-slide mineralogy and OM analysis with applications to formation evaluation in oil & gas, carbon sequestration and mining. Here, analysis of sealing formations in the well indicates high levels of immature organic matter that would not be a viable target for future oil production that could compromise the CO2 storage site. The work brings together diverse disciplines from geology and petrography to analytical chemistry, big data and microscopy.

Myers, Grant↗

Application of Quantum Machine Learning to High Energy Physics Analysis at LHC Using Quantum Computer Simulators and Quantum Computer Hardware

Machine learning enjoys widespread success in High Energy Physics (HEP) analyses at LHC. However the ambitious HL-LHC program will require much more computing resources in the next two decades. Quantum computing may offer speed-up for HEP physics analyses at HL-LHC, and can be a new computational paradigm for big data analyses in High Energy Physics.We have successfully employed three methods (1) Variational Quantum Classifier (VQC) method, (2) Quantum Support Vector Machine Kernel (QSVM-kernel) method and (3) Quantum Neural Network (QNN) method for two LHC flagship analyses: ttH (Higgs production in association with two top quarks) and H->mumu (Higgs decay to two muons, the second generation fermions). We shall address the progressive improvements in performance from method (1) to method (3).We will present our experiences and results of a study on LHC High Energy Physics data analyses with IBM Quantum Simulator and Quantum Hardware (using IBM Qiskit framework), Google Quantum Simulator (using Google Cirq framework), and Amazon Quantum Simulator (using Amazon Braket cloud service). The work is in the context of a Qubit platform (a gate-model quantum computer). Taking into account the present limitation of hardware access, different quantum machine learning methods are studied on simulators and the results are compared with classical machine learning methods (BDT, classical Support Vector Machine and classical Neural Network). Furthermore, we do apply quantum machine learning on IBM quantum hardware to compare performance between quantum simulator and quantum hardware. The work is performed by an international and interdisciplinary collaboration with the Department of Physics and Department of Computer Sciences of University of Wisconsin, CERN Quantum Technology Initiative, IBM Research Zurich, IBM T.J. Watson Research Center, Fermilab Quantum Institute, BNL Computational Science Initiative, State University of New York at Stony Brook, and Quantum Computing and AI Research of Amazon Web Services. This work pioneers a close collaboration of academic institutions with industrial corporations in the High Energy Physics analyses effort. Though the size of event samples in future HL-LHC physics and the limited number of qubits pose some challenges to the Quantum Machine learning studies for High Energy Physics, more advanced quantum computers with larger number of qubits, reduced noise and improved running time (as envisioned by IBM and Google) may outperform classical machine learning in both classification power and in speed.Although the era of efficient quantum computing may still be years away, we have made promising progress and obtained preliminary results in applying quantum machine learning to High Energy Physics. A PROOF OF PRINCIPLE.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Discovery of Signatures, Anomalies, and Precursors in Synchrophasor Data with Matrix Profile and Deep Recurrent Neural Networks (Final Project Report)

The widespread deployment of phasor measurement unit (PMU) across the U.S. together with the burgeoning machine learning technology made it possible to develop data-driven PMU data analytics to improve grid security and reliability in a more insightful and effective manner. Although PMU applications have been explored for over a decade, the representative PMU usage is limited to the bulk power system monitoring mainly due to the data integrity issues associated with PMUs (typically missing, fragmented, and wrongly amplified data). To forge a breakthrough on this stalemate and embrace PMUs for power system control and protection as well, we applied various advanced machine learning and big data analysis technology to the power system event detection and classification as the first step toward the power system control and protection pertaining to grid security enhancement.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Secondary Ion Mass Spectral Imaging of Metals and Alloys

Secondary Ion Mass Spectrometry (SIMS) is an outstanding technique for Mass Spectral Imaging (MSI) due to its notable advantages, including high sensitivity, selectivity, and high dynamic range. As a result, SIMS has been employed across many domains of science. In this review, we provide an in-depth overview of the fundamental principles underlying SIMS, followed by an account of the recent development of SIMS instruments. The review encompasses various applications of specific SIMS instruments, notably static SIMS with time-of-flight SIMS (ToF-SIMS) as a widely used platform and dynamic SIMS with Nano SIMS and large geometry SIMS as successful instruments. We particularly focus on SIMS utility in microanalysis and imaging of metals and alloys as materials of interest. Additionally, we discuss the challenges in big SIMS data analysis and give examples of machine leaning (ML) and Artificial Intelligence (AI) for effective MSI data analysis. Finally, we recommend the outlook of SIMS development. It is anticipated that in situ and operando SIMS has the potential to significantly enhance the investigation of metals and alloys by enabling real-time examinations of material surfaces and interfaces during dynamic transformations.

36 MATERIALS SCIENCE↗

GeoThermalCloud framework for fusion of big data and multi-physics models in Nevada and Southwest New Mexico

Our GeoThermalCloud framework is designed to process geothermal datasets using a novel toolbox for unsupervised and physics-informed machine learning called SmartTensors. More information about GeoThermalCloud can be found at the GeoThermalCloud GitHub Repository. More information about SmartTensors can be found at the SmartTensors Github Repository and the SmartTensors page at LANL.gov. Links to these pages are included in this submission. GeoThermalCloud.jl is a repository containing all the data and codes required to demonstrate applications of machine learning methods for geothermal exploration. GeoThermalCloud.jl includes: - site data - simulation scripts - jupyter notebooks - intermediate results - code outputs - summary figures - readme markdown files GeoThermalCloud.jl showcases the machine learning analyses performed for the following geothermal sites: - Brady: geothermal exploration of the Brady geothermal site, Nevada - SWNM: geothermal exploration of the Southwest New Mexico (SWNM) region - GreatBasin: geothermal exploration of the Great Basin region, Nevada Reports, research papers, and presentations summarizing these machine learning analyses are also available and will be posted soon.

15 GEOTHERMAL ENERGY↗

Software-defined network for end-to-end networked science at the exascale

Domain science applications and workflow processes are currently forced to view the network as an opaque infrastructure into which they inject data and hope that it emerges at the destination with an acceptable Quality of Experience. There is little ability for applications to interact with the network to exchange information, negotiate performance parameters, discover expected performance metrics, or receive status/troubleshooting information in real time. The work presented here is motivated by a vision for a new smart network and smart application ecosystem that will provide a more deterministic and interactive environment for domain science workflows. The Software-Defined Network for End-to-end Networked Science at Exascale (SENSE) system includes a model-based architecture, implementation, and deployment which enables automated end- to-end network service instantiation across administrative domains. An intent based interface allows applications to express their high-level service requirements, an intelligent orchestrator and resource control systems allow for custom tailoring of scalability and real-time responsiveness based on individual application and infrastructure operator requirements. This allows the science applications to manage the network as a first-class schedulable resource as is the current practice for instruments, compute, and storage systems. Deployment and experiments on production networks and testbeds have validated SENSE functions and performance. Emulation based testing verified the scalability needed to support research and education infrastructures. Key contributions of this work include an architecture definition, reference implementation, and deployment. This provides the basis for further innovation of smart network services to accelerate scientific discovery in the era of big data, cloud computing, machine learning and artificial intelligence.

47 OTHER INSTRUMENTATION↗

Software-Defined Network for End-to-end Networked Science at the Exascale

Domain science applications and workflow processes are currently forced to view the network as an opaque infrastructure into which they inject data and hope that it emerges at the destination with an acceptable Quality of Experience. There is little ability for applications to interact with the network to exchange information, negotiate performance parameters, discover expected performance metrics, or receive status/troubleshooting information in real time. The work we presen here is motivated by a vision for a new smart network and smart application ecosystem that will provide a more deterministic and interactive environment for domain science workflows. The Software-Defined Network for End-to-end Networked Science at Exascale (SENSE) system includes a model-based architecture, implementation, and deployment which enables automated end-to-end network service instantiation across administrative domains. An intent based interface allows applications to express their high-level service requirements, an intelligent orchestrator and resource control systems allow for custom tailoring of scalability and real-time responsiveness based on individual application and infrastructure operator requirements. This allows the science applications to manage the network as a first-class schedulable resource as is the current practice for instruments, compute, and storage systems. Deployment and experiments on production networks and testbeds have validated SENSE functions and performance. Emulation based testing verified the scalability needed to support research and education infrastructures. Key contributions of this work include an architecture definition, reference implementation, and deployment. This provides the basis for further innovation of smart network services to accelerate scientific discovery in the era of big data, cloud computing, machine learning and artificial intelligence.

97 MATHEMATICS AND COMPUTING↗

Release of ENDF81SaB: ENDF/B-VIII.1-Based ACE Data Files for Thermal Scattering

On August 30, 2024, the National Nuclear Data Center (NNDC) released the ENDF/B-VIII.1 nuclear data library. The library was released in the standard Evaluated Nuclear Data File (ENDF) format. These files can be accessed on the NNDC's website (www.nndc.bnl.gov). The files provided in the thermal neutron scattering sublibrary were processed into A Compact ENDF (ACE)-formatted files, verified, and validated by the XCP-5 Nuclear Data Team, resulting in the ENDF81SaB application library. This report details the processing of these files and the quality assurance approach taken. This is not intended to be a full validation effort; rather, this library is intended to simply reproduce the released files for further validation testing by the community. The validation basis and details of the evaluations are documented in the forthcoming ``Big Paper''.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Modelling urban-scale occupant behaviour, mobility, and energy in buildings: A survey

The proliferation of urban sensing, IoT, and big data in cities provides unprecedented opportunities for a deeper understanding of occupant behaviour and energy usage patterns at the urban scale. This enables data-driven building and energy models to capture the urban dynamics, specifically the intrinsic occupant and energy use behavioural profiles that are not usually considered in traditional models. Although there are related reviews, none have investigated urban data for use in modelling occupant behaviour and energy use at multiple scales, from buildings to neighbourhood to city. This survey paper aims to fill this gap by providing a critical summary and analysis of the works reported in the literature. We present the different sources of occupant-centric urban data that are useful for data-driven modelling and categorise the range of applications and recent data-driven modelling techniques for urban behaviour and energy modelling, along with the traditional stochastic and simulation-based approaches. Finally, we present a set of recommendations for future directions in data-driven modelling of occupant behaviour and energy in buildings at the urban scale.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Leveraging generative adversarial networks to create realistic scanning transmission electron microscopy images

Abstract The rise of automation and machine learning (ML) in electron microscopy has the potential to revolutionize materials research through autonomous data collection and processing. A significant challenge lies in developing ML models that rapidly generalize to large data sets under varying experimental conditions. We address this by employing a cycle generative adversarial network (CycleGAN) with a reciprocal space discriminator, which augments simulated data with realistic spatial frequency information. This allows the CycleGAN to generate images nearly indistinguishable from real data and provide labels for ML applications. We showcase our approach by training a fully convolutional network (FCN) to identify single atom defects in a 4.5 million atom data set, collected using automated acquisition in an aberration-corrected scanning transmission electron microscope (STEM). Our method produces adaptable FCNs that can adjust to dynamically changing experimental variables with minimal intervention, marking a crucial step towards fully autonomous harnessing of microscopy big data.

77 NANOSCIENCE AND NANOTECHNOLOGY↗