Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data tools”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

4th Big Data for Nuclear Power Plants Workshop 2023

The Ohio State University and Idaho National Laboratory organized the 4 th Big Data for Nuclear Power Plants Workshop in November, 2023 in Columbus, Ohio. Workshop topics were chosen to understand the challenges and gaps that need to be addressed to maximize the impact of data on the nuclear industry, as well as the associated applications and risks. Discussions were focused around six specific application areas: Operation and Maintenance; Machine Learning in Nuclear Materials and Advanced Manufacturing; Cybersecurity; High-Performance Computing and Massive Computation; Big Data and Digital Twins; and Nuclear Non-Proliferation. The opportunities, challenges, and risks identified in the six focus areas explored in this workshop are diverse, but some common themes emerge, such as the importance of data integrity, quality, coverage, privacy, and traceability. Big data and AI/ML tools can be leveraged to reduce costs, optimize human tasking, and reduce human error across various application areas. In order for the nuclear industry to benefit from big data and advanced analytic capabilities, it is essential to address challenges and risks, such as data privacy, model reliability, and computational resource availability. Learning from other industries that have successfully implemented big data and AI/ML technologies, like the aerospace industry, can help the nuclear industry successfully integrate these technologies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗

Through the lens of bioenergy crops: advances, bottlenecks, and promises of plant engineering

Advances in engineering of bioenergy crops were driven over the past years by adapting technological breakthroughs and accelerating conventional applications but also exposed intriguing challenges. New tools revealed rich interconnectivity in the exponentially growing and dynamic 'big' omics data' of metabolomes, transcriptomes, and genomes at previously inaccessible magnitude (global, cross-species, meta-) and resolution (single cell). Insights enabled fresh hypotheses and stimulated disciplines such as functional genomics with discovery of broad regulatory networks and their determinants, that is, DNA parts, including promoters, regulatory elements, and transcription factors. Their rational design, assembly into increasingly complex blueprints, and installation into diverse chassis is an existing frontier that may benefit from emerging technologies to address bottlenecks. Interweaving nature-inspired to fully synthetic parts has already allowed building of fine-tuned regulatory circuits, or new-to-nature metabolic routes insulated from the biological context of the chassis species. Similarly, developments and the evolving need for unifying principles in plant transformation and species-agnostic technologies highlight future opportunities for engineering the next generation of bioenergy plants.

60 APPLIED LIFE SCIENCES↗

High‐Resolution National‐Scale Water Modeling Is Enhanced by Multiscale Differentiable Physics‐Informed Machine Learning

Abstract The National Water Model (NWM) is a key tool for flood forecasting, planning, and water management. Key challenges facing the NWM include calibration and parameter regionalization when confronted with big data. We present two novel versions of high‐resolution (∼37 km 2 ) differentiable models (a type of hybrid model): one with implicit, unit‐hydrograph‐style routing and another with explicit Muskingum‐Cunge routing in the river network. The former predicts streamflow at basin outlets whereas the latter presents a discretized product that seamlessly covers rivers in the conterminous United States (CONUS). Both versions use neural networks to provide a multiscale parameterization and process‐based equations to provide a structural backbone, which were trained simultaneously (“end‐to‐end”) on 2,807 basins across the CONUS and evaluated on 4,997 basins. Both versions show great potential to elevate future NWM performance for extensively calibrated as well as ungauged sites: the median daily Nash‐Sutcliffe efficiency of all 4,997 basins is improved to around 0.68 from 0.48 of NWM3.0. As they resolve spatial heterogeneity, both versions greatly improved simulations in the western CONUS and also in the Prairie Pothole Region, a long‐standing modeling challenge. The Muskingum‐Cunge version further improved performance for basins >10,000 km 2 . Overall, our results show how neural‐network‐based parameterizations can improve NWM performance for providing operational flood predictions while maintaining interpretability and multivariate outputs. The modeling system supports the Basic Model Interface (BMI), which allows seamless integration with the next‐generation NWM. We also provide a CONUS‐scale hydrologic data set for further evaluation and use.

Song, Yalan [Civil and Environmental Engineering T↗

7th World Congress on Integrated Computational Materials Engineering (ICME 2023) (Final Technical Report)

Integrated Computational Materials Engineering (ICME) has received international attention due to its potential to shorten product development time, while lowering cost and improving design and manufacturing outcomes. ICME is an approach to designing materials solutions for specific applications that use computer modeling programs to predict the behavior of materials and integrate this information into the overall materials, processing, and manufacturing design cycle. The 7th World Congress on Integrated Computational Materials Engineering (ICME 2023) was held in Orlando, Florida from May 21–25, 2023 with the goal to convene stakeholders from across all areas of modeling and simulation, experimental specialization, and design, as well as from across academia, government, and industry, to address ICME tools and techniques and their integration, as well as to examine their application in engineering. This atmosphere facilitated rich interactions between the experimentalists, modelers, and computational and design, from academia, government, and industry, to discuss ICME tools and techniques and their application in engineering.

36 MATERIALS SCIENCE↗

Macromolecules & Manufacturing Science

Outline • SRNL Overview • Mission overview • Polymers enabling the mission • R&D Highlights • Polymers in radiation environments • Tooling in shielded cells • Packaging for nuclear material shipments • Polymers supporting tank waste remediation • Ref electrode • Epoxy and polymer grout • Polymers for fusion energy • Deuterium labelling • Polymers for additive manufacturing • Coalescence and blends: experimental and predictive • Process modelling and sorting through big data (peregrine and latticeJ)

Chatham, Camden [Savannah River National Laborator↗

EDX ClaiMM: Digital Resources for the Critical Minerals and Materials Community

Securing critical mineral supply chains is essential for transitioning to a clean energy economy and for maintaining national security. Big-data analytics can serve as a cost-effective means of identifying new domestic critical mineral resources but only if data can be easily located and digested. Using ArcGIS Enterprise Sites, EDX ClaiMM was developed to increase the accessibility of critical minerals data, reducing time spent on data collection and integration. Hosted tools provide rapid visualization and exploration of key datasets, unlocking insights to support resource assessments.

Yesenchak, Rachel↗

Integrating Intelligent Hydro-informatics into an effective Early Warning System for risk-informed urban flood management

The urban drainage system constantly facing flooding issues in coastal and urban areas. Robust and accurate urban flood management, particularly considering fast-moving compound floods, is crucial to minimize the impact of flood disasters in coastal cities. Till now, Ho Chi Minh City (HCMC) lacks an effective means of urban flood management because of flood risk communication among residents. Existing flood risk communication tools rely on post-disaster flood model outcomes and data. Therefore, this research proposes a real-time Early Urban Flooding Warning System (EUFWS) integrated with a user-friendly web and app interface. The backbone of this system consists of flood models developed using machine learning (ML) algorithms, combined with big data and Web-GIS visualization, with ML serving as the core for constructing the EUFWS. EUFWS offer several key advantages: they are available at all times, accessible from anywhere, and provide a real-time, multi-user working platform. Additionally, the system is flexible, allowing for the easy addition of components and services and scalable, adjusting to workload demands. EUFWS have been successfully deployed in Thu Duc City, Vietnam, as a case study and are operating effectively. EUFWS have been successfully deployed in Thu Duc City, Vietnam, as a case study and are operating effectively. Research results indicate that EUFWS supported decision-makers to be effectively risk informed and make intelligent decisions during urban flood emergencies. Finally, this underscores the significant potential of integrating ML and information technology to enhance the management of smart urban drainage systems in flood-prone cities worldwide.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning-Enabled Image Classification for Automated Electron Microscopy

Abstract Traditionally, materials discovery has been driven more by evidence and intuition than by systematic design. However, the advent of “big data” and an exponential increase in computational power have reshaped the landscape. Today, we use simulations, artificial intelligence (AI), and machine learning (ML) to predict materials characteristics, which dramatically accelerates the discovery of novel materials. For instance, combinatorial megalibraries, where millions of distinct nanoparticles are created on a single chip, have spurred the need for automated characterization tools. This paper presents an ML model specifically developed to perform real-time binary classification of grayscale high-angle annular dark-field images of nanoparticles sourced from these megalibraries. Given the high costs associated with downstream processing errors, a primary requirement for our model was to minimize false positives while maintaining efficacy on unseen images. We elaborate on the computational challenges and our solutions, including managing memory constraints, optimizing training time, and utilizing Neural Architecture Search tools. The final model outperformed our expectations, achieving over 95% precision and a weighted F-score of more than 90% on our test data set. This paper discusses the development, challenges, and successful outcomes of this significant advancement in the application of AI and ML to materials discovery.

Materials Science↗

BiG-SCAPE 2.0 and BiG-SLiCE 2.0: scalable, accurate and interactive sequence clustering of metabolic gene clusters

Microbial metabolic gene clusters encode the biosynthesis or catabolism of metabolites that facilitate ecological specialization, mediate microbiome interactions and constitute a major source of medicines and crop protection agents. Here, we present BiG-SCAPE and BiG-SLiCE 2.0, next-generation methods that facilitate scalable, accurate and interactive gene cluster analyses. BiG-SCAPE 2.0 updates its classification, alignment methods, and visualizations, enabling more accurate analysis, up to 8x faster runtimes and halved memory requirements. BiG-SLiCE 2.0 updates its distance metric, pHMM database, and classification logic, resulting in increased sensitivity nearing that of BiG-SCAPE. Analysis of 260,630 biosynthetic gene clusters from publicly available genomes reveals that both tools generate concurring estimates of gene cluster diversity, thus providing significantly extended methodological support for recent evidence indicating that the vast majority of natural product diversity remains unexplored. Together, these updates will facilitate global genome mining efforts for natural product discovery and microbiome analyses scalable with current data sizes.

Draisma, Arjan [Wageningen University & Research (↗

Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale. This has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. In this work, we implement ML-based geothermal exploration and an enhanced geothermal systems (EGS) design tool to achieve the above goals. Our exploration tool is GeoThermalCloud (GTC) EGS design tool is GeoDT-ML. GTC (github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. It enables the identification of critical measurements needed to identify geothermal resource signatures. GeoDT-ML (github.com/SmartTensors/GeoThermalCloud.jl/tree/master/) adds coupling to GeoDT (https://github.com/GeoDesignTool/GeoDT.git) for stochastic EGS design optimization and performance prediction. GeoDT-ML leverages recent advances in deep learning and high-performance computing. Contributors to this effort include LANL, PNNL, Google, Stanford, and Julia Computing.

15 GEOTHERMAL ENERGY↗

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES↗

EDX ClaiMM

EDX ClaiMM is a centralized data & analytical platform designed to revolutionize U.S. critical minerals and materials (CMM) activities. By providing a robust digital infrastructure, ClaiMM will accelerate the combination, leveraging, and rapid utilization of vital data, advanced tools, and cutting-edge research advancements in CMM. This adaptive digital research hub connects the CMM community to essential knowledge products and offers access to interoperable datasets, databases, models, software, and tools from the National Energy Technology’s (NETL’s) Energy Data eXchange (EDX) and other authoritative sources, serving both public and private sectors. EDX ClaiMM delivers AI-informed solutions to address fundamental knowledge gaps and fosters the innovation of new techniques for enhanced characterization and recovery of CMMs within the U.S. By leveraging cloud-hosted, scalable digital infrastructure, ClaiMM meets public–private applied energy needs. It equips the CMM community with priority digital resources that harness on-site and cloud compute capabilities, enabling big data storage, advanced processing, analytics, and visualization.

Critical Materials; Critical Minerals; Rare Earth ↗

Big Data Meets Geothermal Exploration (CRADA Final Report)

As part of the Cyclotron Road program, Zanskar Geothermal & Minerals, Inc. investigated the application of micro-earthquake and ambient noise seismology methods to imaging and characterizing the structural characteristics and hydrothermal flux of subsurface faults. Significant advances in what could be resolved were enabled by two major developments in seismology: 1) the availability of large-n arrays of low-cost seismometers, and 2) the availability of increased computational power and semi-automated data reduction algorithms. In tandem, these advances may improve the signal-to-noise ratio and spatial precision of the data collected and enable higher-resolution characterization of subsurface fracture systems and their spatio-temporal evolution. These tools supported efforts to reduce dry-hole risk and to improve wellfield productivity for geothermal resource development. In particular, two applications of these advances were evaluated: 1) fracture-seismic imaging, which was used to detect ambient emissions from fluid-filled fractures, and 2) reservoir tomography, which used information about travel paths, source locations, and source parameters of micro-earthquakes to identify areas of enhanced permeability. Integration of these methods provided guidance for siting wells and served as prior constraints for reservoir models, informing forecasts of power potential and production and injection strategies aimed at minimizing temperature decline and improving overall resource productivity.

15 GEOTHERMAL ENERGY↗

Evaluating the Impact of Power Outages on Occupancy Patterns During the 2021 Texas Power Crisis

Large-scale power outages, such as those caused by extreme weather events, have a big impact on human behavior. A short power outage is merely a nuisance for most, and may not change people's locations. An outage that lasts for a few hours can result in spoiled food and medical supplies, and people will have to restock spoiled items. Long outages result in temperatures outside tolerable levels in homes, and may prompt people to acquire supplies, such as generators and gas, or change location. The long outages during Winter Storm Uri in Texas resulted in millions of dollars in property damage due to freezing pipes. This level of damage is expected to result in a sharp increase in supply runs and contractor activity. In this paper, we present a tool to explore differences in visiting patterns before, during, and after power outages. It allows to compare different points of interest like medical facilities, grocery stores, hardware stores, and other types of businesses.

big data↗