Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data enhancement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

FY24 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

Algorithms for Machine Learning (ML) and data analysis for the 3013 Surveillance Program have been developed in an ongoing collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). The objective of the algorithms is to automate the identification of corrosion and crack formation in the Inner Container Closure Weld Region (ICCWR) of the canister system used to store Pu-bearing material. Data for corrosion and cracking is collected from large binary files generated by a Laser Confocal Microscope (LCM), the Wide Area 3D Measurement System (WAMS), or,in a recent proposal, by a Scanning Electron Microscope (SEM). The ML software uses the physical attributes in the data files (e.g., one or more of: height, color, and 16-bit grayscale values as functions of position in a plane projection) to detect signs of surface corrosion and cracking after being trained on similar data, with the features to be detected. Although the initial scope included screening for broader indicators of corrosion, e.g., pitting, identification of potential cracks was prioritized for the past several years at the request of program leadership. Labeled training data is essential to developing the ML algorithm, and enhancements to data labeling capability have been developed to address this essential precursor to application of ML routines. Efficient labeling is particularly important in view of the large volume of data required to train ML algorithms and the relative rarity of cracks in the ICCWR data set. The updated program will read binary data from either LCM, WAMS or SEM files, interrogate data attributes, facilitate user labeling of data for training ML algorithms, execute ML algorithms, output parameters from trained ML algorithms, report ML model accuracy with respect to labeled data, and generate graphical representations for various analyses. In FY24, hourglass neural networks (HNNs) that were initiated in FY22 were further developed and tested using available LCM data, and their performance was tested against that of the alternative U-Net Neural Network algorithm structure. HNNs along with previously developed Convolutional Neural Networks (CNNs) and Deep Neural Networks (DNNs) comprise a suite of ML tools for identification of cracks in the ICCWR

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Deep learning-driven super-resolution in Raman hyperspectral imaging: Efficient high-resolution reconstruction from low-resolution data

Deep learning (DL) has become an indispensable tool in hyperspectral data analysis, automatically extracting valuable features from complex, high-dimensional datasets. Super-resolution reconstruction, an essential aspect of hyperspectral data, involves enhancing spatial resolution, particularly relevant to low-resolution hyperspectral data. Yet, the pursuit of super-resolution in hyperspectral analysis is fraught with challenges, including acquiring ground truth high-resolution data for training, generalization, and scalability. The pressing issue of extended spectral acquisition times, notably for high-resolution scans, is a significant roadblock in hyperspectral imaging. Super-resolution methods offer a promising solution by providing higher spatial resolution data to expedite data collection and yield more efficient outcomes. This paper delves into a practical application of these concepts using Raman imaging, where spectral acquisition times can be prohibitively long. In this context, DL-based super-resolution models demonstrate their efficacy by predicting and reconstructing high-resolution Raman data from low-resolution input, eliminating the need for resource-intensive high-resolution scans. While previous work often relied on substantial high-resolution datasets, this study showcases the ability to achieve similar outcomes even with limited data, presenting a more practical and cost-effective approach. In conclusion, the results offer a glimpse into the transformative potential of this technology to streamline hyperspectral imaging applications by saving valuable time and resources through the successful generation of high-resolution data from low-resolution inputs.

42 ENGINEERING↗

Two and three-dimensional near infrared subcutaneous structure imager using realtime nonlinear video processing

A method includes placing an imaging device in contact with a sample, irradiating the sample via an illuminator such that radiation scatters off of the structure, and providing the scattered radiation to a detector to generate a digital video signal. A processor then performs various operations including subtracting a sum of a black state value and a user-defined black level clamping value from the digital video signal to create subtracted image data, substituting negative values in the subtracted image data with a reference value to create modified image data, and amplifying the modified image data by a user-defined gain to create contrast enhanced image data.

Krasowski, Michael J.↗

Improving Text Classification with Large Language Model-Based Data Augmentation

Large Language Models (LLMs) such as ChatGPT possess advanced capabilities in understanding and generating text. These capabilities enable ChatGPT to create text based on specific instructions, which can serve as augmented data for text classification tasks. Previous studies have approached data augmentation (DA) by either rewriting the existing dataset with ChatGPT or generating entirely new data from scratch. However, it is unclear which method is better without comparing their effectiveness. This study investigates the application of both methods to two datasets: a general-topic dataset (Reuters news data) and a domain-specific dataset (Mitigation dataset). Our findings indicate that: 1. ChatGPT generated new data consistently enhanced model’s classification results for both datasets. 2. Generating new data generally outperforms rewriting existing data, though crafting the prompts carefully is crucial to extract the most valuable information from ChatGPT, particularly for domain-specific data. 3. The augmentation data size affects the effectiveness of DA; however, we observed a plateau after incorporating 10 samples. 4. Combining the rewritten sample with new generated sample can potentially further improve the model’s performance.

97 MATHEMATICS AND COMPUTING↗

Thermodynamics-guided machine learning model for predicting convective boundary layer height and its multi-site applicability

Accurate estimation of convective boundary layer height (CBLH) is vital for weather, climate, and air quality modeling. Machine learning (ML) shows promise in CBLH prediction, but input parameter selection often lacks physical grounding, limiting generalizability. This study introduces a novel ML framework for CBLH prediction, integrating thermodynamic constraints and the diurnal CBLH cycle as an implicit physical guide. Boundary layer growth is modeled as driven by surface heat fluxes and atmospheric heat absorption represented with the low tropospheric stability, using the diurnal cycle as input and output. TPOT and AutoKeras are employed to select optimal models, validated against Doppler lidar-derived CBLH data, achieving an R 2 of 0.84 across untrained years. Comparisons of eddy covariance (ECOR) and energy balance Bowen ratio (EBBR) flux measurements show the same prediction capability. Models trained on the ARM SGP C1 site with ECOR data and tested at E37 and E39 yield R 2 values of 0.79 and 0.81, respectively, demonstrating their adaptability. The ML model trained with all sites' data slightly enhances the performance compared with ML models trained over single-site data. The interquartile range for predicted CBLH is consistently narrower than that for DL-derived CBLH, reflecting lower variability in predicted CBLH compared to DL-derived CBLH, which is influenced by additional factors, which are not well represented with the model inputs. The model's generalizability across multiple sites at the ARM SGP site demonstrates its potential for transfer to greater distances, offering a scalable approach for enhancing boundary layer parameterization in atmospheric models.

Chu, Yufei [Stony Brook Univ., NY (United States)]↗

Integrated Risk-Informed Condition Based Maintenance Capability and Automated Platform: Technical Report 3

This project is a collaborative research effort between PKMJ Technical Services LLC, Idaho National Laboratory, and Public Service Enterprise Group (PSEG) Nuclear, LLC. The collaboration, led by PKMJ Technical Services LLC, is part of the industry Funding Opportunity Announcement (FOA) award under Advanced Nuclear Technology Development FOA #DE-FOA-0001817. The pilot demonstration focuses on the Circulating Water System (CWS), an important non-safety-related system that impacts the power generation capability of the plant site. Achieving riskinformed condition-based Predictive Maintenance (PdM) on the CWS will result in significant economic benefits, and the developed methodologies can also be applied to other plant systems. This approach supports an industry goal of ensuring that nuclear power generation remains a viable, economically competitive option in the energy market. Operation and Maintenance (O&M) costs include labor-intensive Preventive Maintenance (PM) programs that involve manually performed inspection, calibration, testing, and maintenance of plant assets at periodic frequencies as well as time-based replacement of assets, irrespective of condition. This project offers an alternative by focusing on riskinformed condition-based maintenance to reduce O&M costs while still maintaining plant health and safety. This report summarizes the progress made toward achieving a risk-informed condition-based maintenance approach. The research and development (R&D) activities presented in this report are associated with development of a nuclear digital platform application, integration of fault signature models, and automated work management processes. The fault signatures and Machine Learning (ML) models are key components in predictive analytics and are heavily leveraged to improve the insights received by existing plant process data sources. Availability of the analysis results within a centralized digital platform enhances efficiency by enabling automation of activities otherwise performed manually. Personnel are presented with enhanced information that can be used to evaluate plant status and risks. Utilizing the enhancements to data analytics supports automated responses, (i.e. issuance of work orders) to address developing equipment faults and thus preventing forced, unplanned shutdowns of components or systems. The R&D activities described within this report lay the foundation for developing and demonstrating a digital automated platform to centralize the implementation of condition monitoring and response to equipment faults. The digital automated platform is cloud-based and designed to enable improved efficiency of plant processes. The digital platform includes content related to maintenance optimization, fault signature analysis, and plant records, which can all be used to support efficiencies when located within a centralized digital platform. These efficiencies could be further enhanced when deployed through industry-wide deployment of the technology to improve insights and processes based upon economies of scale.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower (DIVERS-H)

U.S. hydropower plants face potential threats from shrinking water supply, rising demands, and warmer stream temperatures from various causes. Power plant owners, operators, and regulators require new tools to take advantage of and interpret the diverse range of scientific data being produced by both observational methods (for example, satellite, radar, stream gauges) and computer modeling methods that evaluate and predict how earth's dynamic systems (atmosphere, oceans, land surface, and sea ice) are changing and interacting. Combining datasets such as these with AI-based analyses introduces a novel decision support system to help users anticipate and address potential impacts on power generation stations. This new technology has been named DIVERS-H for "Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower." In Phase I, technical feasibility was established with the development and demonstration of all the new technologies that are required. Most notably, DIVERS-H will use new artificial intelligence (AI) methods to capture the complex dynamics of water availability, demand, and environmental changes. In addition, new data management software was developed, and a prototype user interface was implemented as the precursor to a full scale decision support system. With technical research complete, the project focus now shifts to development of a commercial software product to provide users with actionable insight into water availability and the risk/resilience of critical systems at their locations of interest. Although DIVER-H was originally conceived as a tool for hydroelectric power applications, the same underlying technology can be readily applied to other water-consuming systems including coal, natural gas, oil, and nuclear power plants.

Chaudhary, Aashish [Kitware, Inc., Clifton Park, N↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Data-Driven Energy Resilience Assessment and Enhancement in Urban Communities: A Case Study in Detroit

This paper presents a data-driven framework for assessing and enhancing energy resilience in urban communities. The resilience assessment is based on two datasets: 1) annual aggregated power outage data and 2) 15-minute interval outage data. High-impact, low-probability (HILP) events are identified within these datasets to evaluate community resilience under extreme conditions. To enhance resilience, an optimization framework utilizing mixed integer linear programming is developed to determine the optimal sizing and placement of solar photovoltaic (PV) systems and battery energy storage systems (BESS). This method offers a cost-effective and practical solution for improving energy resilience in vulnerable communities. Furthermore, a case study of the City of Detroit in Michigan demonstrates the effectiveness of the framework through simulation and validation.

Energy resilience assessment↗

Coincident learning for beam-based rf station fault identification using phase information at the SLAC linac coherent light source

Anomalies in radio-frequency (rf) stations can result in unplanned downtime and performance degradation in linear accelerators such as SLAC’s Linac Coherent Light Source (LCLS). Detecting these anomalies is challenging due to the complexity of accelerator systems, high data volume, and scarcity of labeled fault data. Prior work identified faults using beam-based detection, combining rf amplitude and beam position monitor data. Due to the simplicity of the rf amplitude data, classical methods are sufficient to identify faults, but the recall is constrained by the low-frequency and asynchronous characteristics of the data. In this work, we leverage high-frequency, time-synchronous rf phase data to enhance anomaly detection in the LCLS accelerator. Due to the complexity of phase data, classical methods fail, and we instead train deep neural networks within the Coincident Anomaly Detection (CoAD) framework. We find that applying CoAD to phase data detects nearly 3 times as many anomalies as when applied to amplitude data, while achieving broader coverage across rf stations. Furthermore, the rich structure of phase data enables us to cluster anomalies into distinct physical categories. Through the integration of auxiliary system status bits, we link clusters to specific fault signatures, providing additional granularity for uncovering the root cause of faults. We also investigate interpretability via Shapley values, confirming that the learned models focus on the most informative regions of the data and providing insight for cases where the model makes mistakes. This work demonstrates that phase-based anomaly detection for rf stations improves both diagnostic coverage and root cause analysis in accelerator systems and that deep neural networks are essential for effective analysis.

Accelerator Physics (physics.acc-ph)↗

Bioenergy Research Centers Data Sharing Portal

The bioenergy.org website is the end product of the Data Sharing Shared Research objective for the Bioenergy Centers. The objective is to Enhance BRC data legacy through the development and use of shared software tools to make previously published datasets more findable and accessible through the Inter-BRC Data Products Portal, and to explore the use of generative AI to assist in the exploration of published datasets. The software repository is released per the license information below.

Thrower, Nicholas↗

Unlocking hidden information in sparse small-angle neutron scattering measurements

Hypothesis Small-Angle Neutron Scattering (SANS) is a powerful technique for studying soft matter systems such as colloids, polymers, and lyotropic phases, providing nanoscale structural insights. However, its effectiveness is limited by low neutron flux, leading to long acquisition times and noisy data. Here, we hypothesize that Bayesian statistical inference using Gaussian Process Regression (GPR) can reconstruct high-fidelity scattering data from sparse measurements by leveraging intensity smoothness and continuity. Experiments and Simulations The method was benchmarked computationally and validated through SANS experiments on various soft matter systems, including wormlike micelles, colloidal suspensions, polymeric structures, and lyotropic phases. GPR-based inference was applied to both experimental and synthetic data to evaluate its effectiveness in noise reduction and intensity reconstruction. Findings GPR significantly enhances SANS data quality and therefore reducing measurement times by up to two orders of magnitude. This cost-effective approach maximizes experimental efficiency, enabling high-throughput studies and real-time monitoring of dynamic systems. It is particularly beneficial for weakly scattering and time-sensitive studies. Beyond SANS, this framework applies to other low-SNR techniques, including laboratory-based small-angle X-ray scattering and various dynamical scattering methods. Furthermore, it offers transformative potential for compact neutron sources, enhancing their viability for structural analysis in resource-limited settings.

Small angle neutron scattering↗

Evaluation of Multi-Fidelity Soil Moisture Products Across the Continental United States

We have aggregated the most recent soil moisture datasets from a diverse range of sources, encompassing the Continental United States (CONUS). These sources encompass gridded data from remote sensing products, reanalysis products, machine learning-based projects, and land surface modeling products. Additionally, we have obtained and processed in-situ soil moisture observations from the International Soil Moisture Network. The collected datasets exhibit variations in both temporal and spatial resolutions. Among the 20 datasets, six are available at a spatial resolution of 0.25 degrees, while three are at a coarser spatial resolution of 25 km. To minimize spatial interpolation, we conducted data uncertainty evaluations at the 0.25-degree spatial resolution. For our data evaluations, we maintained a monthly temporal resolution, which effectively captures soil moisture seasonality and interannual variability. Our data processing strategy preserves the raw data and interpolated data at their original temporal resolutions. Datasets with higher temporal resolutions, including daily, three-hourly, and hourly datasets, are set aside for subsequent analyses. These analyses will delve into topics such as soil moisture changes and recovery during extreme weather events. Furthermore, we have processed auxiliary data to enhance our evaluation, leveraging tools such as Google Earth Engine. This includes incorporating topography data, land use land cover data, Köppen-Geiger climate classification, and more to provide a comprehensive assessment from multiple sources.

Li, Lingcheng↗

Data from: "Warming of alpine tundra enhances belowground production and shifts community towards resource acquisition traits"

This archive contains data used to draw conclusions in “Warming of alpine tundra enhances belowground production and shifts community towards resource acquisition traits”, by Yang et al. 2020. Data were collected on Niwot Ridge, in an alpine meadow within the Alpine Treeline Warming Experiment (ATWE) field sites in Colorado, USA. Samples were also processed in the U.S. Geological Survey Forest and Rangeland Ecosystem Science Center, in Boise, Idaho. File formats in this archive include comma-separated values (.csv), portable document format (.pdf), Microsoft Excel (.xlsx), and two types of geospatial files: keyhole markup language (.kml), and ESRI shapefiles (.shp). Leaf scans are .jpg images, and root scans are .tiff/.tif images.The .csv files can be opened using R, Microsoft Excel, or any simple text-editing software such as TextEdit and Notepad. Microsoft Excel files can be opened using Microsoft Excel, and .pdf files can be opened with Adobe Acrobat Reader, Preview, or other compatible programs. Scanned images can be opened using any photo and/or picture viewing software.The .kml file can be opened using Google Earth and Google Maps, and the shapefiles can be opened by any programs compatible with shapefiles, such as the ArcGIS Desktop suite, and QGIS.------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------Measures of belowground net primary productivity (BNPP) are required to understand whether aboveground net primary production (ANPP) changes reflect changes in allocation or are indicative of a whole plant NPP response. Plant functional traits provide a key way to scale from the individual plant to the community level, and provide insight into drivers of NPP responses to environmental change. We used infrared heaters to warm an alpine plant community at Niwot Ridge, Colorado, and applied supplemental water to compensate for soil water loss induced by warming. We measured ANPP, BNPP, and leaf and root functional traits across treatments after 5 years of continuous warming. Community-level ANPP and total NPP (ANPP + BNPP) did not respond to heating or watering, but BNPP increased in response to heating. Heating decreased community-level leaf dry matter content and increased total root length, indicating a shift in strategy from resource conservation to acquisition in response to warming.

13C/12C isotope ratio↗

Performance Characterization and Provenance of Distributed Task-based Workflows on HPC Platforms

Understanding performance and provenance of task-based workflows poses significant challenges, particularly in distributed configurations where resources are shared by multiple applications. Task-based workflow management systems further complicate performance predictability because of their dynamicity that subtly alters task execution order from run to run. In this paper we propose a layered characterization framework for performance and task provenance for Dask.distributed workflows running on high-performance computing (HPC) platforms. It collects data from jobs, the workflow management system, and the operating system to aid in understanding the performance of these workflows. Our approach encompasses three main contributions: first, an extension of Dask.distributed to capture high-fidelity task provenance using Mochi data services; second, the adaptation of the established HPC I/O characterization tool Darshan to gather high-fidelity I/O data, thereby enhancing the granularity of our analysis; and third, a framework to combine and process the collected data and provide helpful insights into performance characterization and reproducibility, alongside our lessons learned.

Dask↗

AL-ASMR: Active Learning of Atomistic Surrogate Models for Rare Events

Atomistic simulation with artificial intelligence (AI) is an emerging tool for understanding materials' properties and behaviors and predicting novel materials with optimized/targeted properties. Neural network potentials (NNPs) are outstanding in this field as they have shown a comparable accuracy to ab initio electronic structure calculations for reproducing potential energy surfaces while being several orders of magnitude faster. However, such NNPs can perform poorly outside of their training domain and typically fail catastrophically in the prediction of rare events in molecular dynamics (MD) simulations. For effective AL loops to distinguish the informative data from enhanced sampled configurations, we developed a decision engine by configurational similarity and uncertainty quantification (UQ) with data augmentation.

Jung, Gang Seob↗

Innovative control mechanism for research and test reactors using mandrel-shaped control rods

Research and test reactors have historically played a pivotal role in supporting the initial development of nuclear reactors. They continue to provide essential data for enhancing fuel designs and material knowledge. However, with many such reactors aging and the growing demand for data to bolster advanced reactor development, it is more necessary to research potential design attributes of the next generation of research and test reactors. For test reactors dedicated to fuel and material testing, the design of control mechanisms significantly influences the stabilization of neutron flux levels in irradiation positions while sustaining criticality. This study presents an innovative control mechanism for potential research and test reactor designs. It employs small absorber rods that move in opposite axial directions to maintain axial symmetry of power and neutron flux during burnup cycles. These rods maximize reactivity worth while also offering flexibility to flatten the radial power distribution. An axial translation of the control mechanisms’ absorbers, as compared to the rotational movement of absorbers in control cylinders, also provides a benefit to available excess reactivity and cycle length. Additionally, this work utilizes a simplified core model of the Advanced Test Reactor to assess the performance of this control mechanism. Compared to the current control system based on rotating control cylinders, the new control mechanism has the potential to enhance, or at least maintain, neutronic performance parameters in this reactor design.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗