Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data enhancement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Outcomes of PAX sapiens-Supported Global Wildlife Data Sharing Conferences for Enhanced One Health Security (GWDSC)

Across two consecutive Global Wildlife Data Sharing Conferences supported by PAX sapiens—Year 1 (May 2024) at Pacific Northwest National Laboratory and Year 2 (2025) in Ciudad Real, Spain—the initiative converted wildlife data sharing from aspiration into operational reality, producing measurable impacts in platform development, data mobilization, standards harmonization, and international partnership formation. The conferences addressed a critical gap in global health security: while 75% of emerging infectious diseases affect both humans and animals and over 60% originate in wildlife, wildlife health surveillance has historically lagged behind human and agricultural sectors due to fragmented databases, inconsistent terminology, uneven capacity, and limited cross-border coordination. By convening practitioners, government agencies, international organizations, academic institutions, and NGOs, the GWDSC catalyzed trust-based relationships and practical workflows that enable earlier detection, better risk assessment, and more effective prevention of threats at the wildlife–domestic animal–human–environment interface.

54 ENVIRONMENTAL SCIENCES↗

Machine-Learning-aided Approach for Predicting the Thermal Expansion Behaviors in Advanced Test Reactor Capsules (NURETH-20 full paper)

Instrumented experiments at test reactors are essential to deploying new advanced reactor systems. Designing new experiments and generating data on specific conditions require both time and cost investment. A high-fidelity model of the experiment environment can be created using finite element analysis software to support the actual experiments, but computation time is still a concern in applying outcomes to real-time usage (e.g., a digital twin). This research proposes a machine-learning-aided approach to temperature and displacement predictions, based on the thickness of the outer gas gap on the experimental capsule used for the in-pile demonstration of a novel thermal conductivity probe in the Advanced Test Reactor. The capsule consisted of U10Zr fuel, a rodlet, sodium, and inner and outer capsules. There were gas gaps between the fuel and rodlet and between the inner and outer capsule. The learning data consisted of an experimental capsule’s radial distributions of temperature and displacement, as obtained from Abaqus and the physical features. For the first step, temperature was predicted using three positional parameters. Then the displacement was predicted using six different positional parameters. Each physical feature was normalized to be both nondimensional and standardized. The temperature and displacement predictions showed good agreement in all cases involving interpolation and extrapolation. Also, data similarity enhancement increased the similarity between training and target data increasing the predictive accuracy of machine-learning models. In some cases of extrapolation, the accuracy of the machine-learning model showed limited performance, but still data similarity enhancement improved the accuracy.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Enhancing Discoverability and Management of Atmospheric Data at Scale: Solutions from the ARM Data Center

The Atmospheric Radiation Measurement (ARM) is a multi-laboratory and multi-institutional U.S. Department of Energy (DOE) Office of Science National User Facility. The ARM Data Center (ADC), located at Oak Ridge National Laboratory, collects, archives, and shares vast atmospheric data crucial for climate research. The ADC manages over 7 PB of data from 460 instruments worldwide, processing it into more than 11,000 diverse data products using the Network Common Data Form (NetCDF) for machine-independent accessibility. The primary challenge addressed in this paper is the efficient management and distribution of vast and diverse datasets essential for the climate research community, enhancing accessibility through advanced tools like Data Discovery. The ADC has developed advanced infrastructure and software architecture to handle the continuous influx of heterogeneous data to enhance data discoverability, resulting in increased scientific collaboration. In 2023, users from over 34 countries downloaded and utilized ARM data, resulting in 1,455 publications. The ADC’s efforts have significantly improved the discoverability and usability of atmospheric data, fostering extensive scientific research and collaboration. This paper details the solutions implemented by the ADC team for efficient data discovery and distribution, and it demonstrates ARM’s capability of staging processed data for scientific analysis.

Shah, Chirag [ORNL] (ORCID:0000000203145737)↗

Remote sensing detection enhancement

Big Data in the area of Remote Sensing has been growing rapidly. Remote sensors are used in surveillance, security, traffic, environmental monitoring, and autonomous sensing. Real-time detection of small moving targets using a remote sensor is an ongoing, challenging problem. Since the object is located far away from the sensor, the object often appears too small. The object’s signal-to-noise-ratio (SNR) is often very low. Occurrences such as camera motion, moving backgrounds (e.g., rustling leaves), low contrast and resolution of foreground objects makes it difficult to segment out the targeted moving objects of interest. Due to the limited appearance of the target, it is tough to obtain the target’s characteristics such as its shape and texture. Without these characteristics, filtering out false detections can be a difficult task. Detecting these targets, would often require the detector to operate under a low detection threshold. However, lowering the detection threshold could lead to an increase of false alarms. In this paper, the author will introduce a new method that improves the probability to detect low SNR objects, while decreasing the number of false alarms as compared to using the traditional baseline detection technique.

47 OTHER INSTRUMENTATION↗

Improving the Quality of Geothermal Data Through Data Standards and Pipelines Within the Geothermal Data Repository: Preprint

For machine learning outputs to be applicable to real world problems, high quality data are needed to ensure high quality results. With the more recent emphasis on machine learning in geothermal, there is an increasing need for greater focus on the quality of the data available for use in these projects. For example, Geothermal Operational Optimization Using Machine Learning (GOOML) utilized large quantities of geothermal power plant operational data to inform power plant operational configurations to maximize power generation. High quality datasets result from dependable sensors or devices collecting data, high frequency of measurements, sufficient data points, adequate metadata, reliable storage of data, and sufficient data curation. Another component that contributes to high quality data is reusability, which can be enhanced through data standardization. Data Standardization creates consistency in formatting and contents of like datasets, lessening preprocessing requirements and ensuring adequate information provided by a given dataset. The Geothermal Data Repository (GDR) aims to help improve data quality through automated data standardization for high-value datasets through the implementation of data pipelines alongside reliable and accessible long-term storage for datasets. As such, the GDR has decided to shift away from recommending the use of Excel-based content models and towards the implementation of automated data pipelines. This takes the burden of data standardization off the user and project team and will increase the availability of standardized geothermal data available through the GDR. A set of recommendations, or a data standard for each data type will exist with each data pipeline in order to advise data collection for maximum usability for future research. This paper serves to describe the GDR's proposed transition towards data standardization through automated data pipelines, to discuss the need for and value of such a shift, and to call for suggestions from the community regarding the most useful data standards and pipelines.

data↗

Multiphysics FISPACT-II and TENDL-2019 simulation: neutron-induced damage metrics

Nuclear interactions can be the source of atomic displacement, irradiation-induced defects and transmutation in structural materials. Such quantities are derived from, or can be correlated to, nuclear kinematic simulations of primary atomic energy distributions spectra and the quantification of the numbers of secondary defects produced per primary as a function of the available recoils, residual and emitted, energies. Recoil kinematics of neutral, residual, gas, proton, alpha particle emissions are now more rigorously treated based on recent, complete and enhanced nuclear data parsed in state-of-the-art processing tools. Defect production metrics are the starting point in the complex problem of correlating and simulating the behaviour of materials under irradiation, as experimental information is rare or scattered. Detailed, segregated primary knock-on-atom metrics are now becoming available as the starting point of further simulation processes of isolated and clustered defects in material lattices. This allows more materials, lattices, neutron incident energy ranges, and irradiation conditions to be explored with sufficient data to adequately cover both standard and novel applications and materials: the broader reactor applications landscape. The damage metrics of of materials are systematically explored under typical but different reactor's type environment. The inventory code FISPACT-II combined with the enhanced nuclear data forms of the TENDL-2019 libraries allow one to not only calculate dpa from mostly scattering events but to also properly predict gas production, nuclear heating and transmutation under the same conditions.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Cross-Layered Distributed Data-Driven Framework for Enhanced Smart Grid Cyber-Physical Security

Smart Grid (SG) research and development has drawn much attention from academia, industry and government due to the great impact it will have on society, economics and the environment. Securing the SG is a considerably significant challenge due the increased dependency on communication networks to assist in physical process control, exposing them to various cyber-threats. In addition to attacks that change measurement values using False Data Injection (FDI) techniques, attacks on the communication network may disrupt the power system's real-time operation by intercepting messages, or by flooding the communication channels with unnecessary data. Addressing these attacks requires a cross-layer approach. In this paper a cross-layered strategy is presented, called Cross-Layer Ensemble CorrDet with Adaptive Statistics(CECD-AS), which integrates the detection of faulty SG measurement data as well as inconsistent network inter-arrival times and transmission delays for more reliable and accurate anomaly detection and attack interpretation. Numerical results show that CECD-AS can detect multiple False Data Injections, Denial of Service (DoS) and Man In The Middle (MITM) attacks with a high F1-score compared to current approaches that only use SG measurement data for detection such as the traditional physics-based State Estimation, Ensemble CorrDet with Adaptive Statistics strategy and other machine learning classification-based detection schemes.

cyber-physical security↗

Data-driven modeling to enhance municipal water demand estimates in response to dynamic climate conditions

Altered precipitation and temperature patterns from a changing climate will affect supply, demand, and overall municipal water system operations throughout the arid western U.S. While supply forecasts leverage hydrological models to connect climate influences with surface water availability, demand forecasts typically estimate water use independent of climate and other externalities. Stemming from an increased focus on seasonal water demand management, we use the Salt Lake City, Utah municipal water system as a test bed to assess model accuracy versus complexity trade-offs between simple climate-independent econometric-based models and complex climate-sensitive data-driven models to average to extreme wet and dry climate conditions—representative of a new climate normal. Here, the climate-independent model displayed low performance during extreme dry conditions with predictions exceeding 90% and 40% of the observed monthly and seasonal volumetric demands, respectively, which we attribute to insufficient model complexity. The climate-sensitive models displayed greater accuracy in all conditions, with an ordinary least squares model demonstrating a measurable reduction in prediction bias (3.4% vs. -27.3%) and RMSE (74.0 lpcd vs. 294 lpcd) compared to the climate-independent model. The climate-sensitive workflow increased model accuracy and characterized climate-demand interactions, demonstrating a novel tool to enhance water system management.

54 ENVIRONMENTAL SCIENCES↗

Bragg Coherent Diffraction Imaging for In Situ Studies in Electrocatalysis

Electrocatalysis is at the heart of a broad range of physicochemical applications that play an important role in the present and future of a sustainable economy. Among the myriad of different electrocatalysts used in this field, nanomaterials are of ubiquitous importance. An increased surface area/volume ratio compared to bulk makes nanoscale catalysts the preferred choice to perform electrocatalytic reactions. Bragg coherent diffraction imaging (BCDI) was introduced in 2006 and since has been applied to obtain 3D images of crystalline nanomaterials. BCDI provides information about the displacement field, which is directly related to strain. Lattice strain in the catalysts impacts their electronic configuration and, consequently, their binding energy with reaction intermediates. Even though there have been significant improvements since its birth, the fact that the experiments can only be performed at synchrotron facilities and its relatively low resolution to date (~10 nm spatial resolution) have prevented the popularization of this technique. Herein, we will briefly describe the fundamentals of the technique, including the electrocatalysis relevant information that we can extract from it. Subsequently, we review some of the computational experiments that complement the BCDI data for enhanced information extraction and improved understanding of the underlying nanoscale electrocatalytic processes. We next highlight success stories of BCDI applied to different electrochemical systems and in heterogeneous catalysis to show how the technique can contribute to future studies in electrocatalysis. Finally, we outline current challenges in spatiotemporal resolution limits of BCDI and provide our perspectives on recent developments in synchrotron facilities as well as the role of machine learning and artificial intelligence in addressing them.

bragg coherent diffraction imaging↗

Heterogeneous Machine Learning on High Performance Computing for End to End Driving of Autonomous Vehicles

Current artificial intelligence techniques for end to end driving of autonomous vehicles typically rely on a single form of learning or training processes along with a corresponding dataset or simulation environment. Relatively speaking, success has been shown for a variety of learning modalities in which it can be shown that the machine can successfully “drive” a vehicle. However, the realm of real-world driving extends significantly beyond the realm of limited test environments for machine training. This creates an enormous gap in capability between these two realms. With their superior neural network structures and learning capabilities, humans can be easily trained within a short period of time to proceed from limited test environments to real world driving. For machines though, this gap is guarded by at least two challenges: 1) machine learning techniques remain brittle and unable to generalize to a wide range of scenarios, and 2) effective training data that enhances generalization and generates the desired driving behavior. Further, each challenge can be computationally intensive on its own thereby exasperating the gap. Moreover, is has not yet been shown that a single form of learning or training is capable of addressing a large range of scenarios. As a result, solving the first challenge does not inherently solve the second and vice versa. The work described here discusses an approach to address the first challenge that would also provide a foundation for solving the second. Our approach utilizes a combination of conditional imitation learning with a static dataset, reinforcement learning with a simulation environment, and high-performance computing to train a neural network. As a result, this reduces the “time to solution” from to the existing techniques for autonomous driving and provides an extensible framework to address the second key challenge.

autonomous vehicles↗

Data for Creating Yellow Seed Camelina sativa with Enhanced Oil Accumulation by CRISPR-Mediated Disruption of Transparent Testa 8

Camelina ( Camelina sativa L.), a hexaploid member of the Brassicaceae family, is an emerging oilseed crop being developed to meet the increasing demand for plant oils as biofuel feedstocks. In other Brassicas, high oil content can be associated with a yellow seed phenotype, which is unknown for camelina. We sought to create yellow seed camelina using CRISPR/Cas9 technology to disrupt its Transparent Testa 8 (TT8) transcription factor genes and to evaluate the resulting seed phenotype. We identified three TT8 genes, one in each of the three camelina subgenomes, and obtained independent CsTT8 lines containing frameshift edits. Disruption of TT8 caused seed coat colour to change from brown to yellow reflecting their reduced flavonoid accumulation of up to 44%, and the loss of a well-organized seed coat mucilage layer. Transcriptomic analysis of CsTT8-edited seeds revealed significantly increased expression of the lipid-related transcription factors LEC1, LEC2, FUS3, and WRI1 and their downstream fatty acid synthesis-related targets. These changes caused metabolic remodelling with increased fatty acid synthesis rates and corresponding increases in total fatty acid (TFA) accumulation from 32.4% to as high as 38.0% of seed weight, and TAG yield by more than 21% without significant changes in starch or protein levels compared to parental line. These data highlight the effectiveness of CRISPR in creating novel enhanced-oil germplasm in camelina. The resulting lines may directly contribute to future net-zero carbon energy production or be combined with other traits to produce desired lipid-derived bioproducts at high yields.

Biofuels↗

Pooling Data Improves Multimodel IDF Estimates over Median-Based IDF Estimates: Analysis over the Susquehanna and Florida

Traditional multimodel methods for estimating future changes in precipitation intensity, duration, and frequency (IDF) curves rely on mean or median of models’ IDF estimates. Such multimodel estimates are impaired by large estimation uncertainty, shadowing their efficacy in planning efforts. Here, assuming that each climate model is one representation of the underlying data generating process, i.e., the Earth system, we propose a novel extension of current methods through pooling model data: (i) evaluate performance of climate models in simulating the spatial and temporal variability of the observed annual maximum precipitation (AMP), (ii) bias-correct and pool historical and future AMP data of reasonably performing models, and (iii) compute IDF estimates in a nonstationary framework from pooled historical and future model data. Pooling enhances fitting of the extreme value distribution to the data and assumes that data from reasonably performing models represent samples from the “true” underlying data generating distribution. Through Monte Carlo simulations with synthetic data, we show that return periods derived from pooled data have smaller biases and lesser uncertainty than those derived from ensembles of individual model data. We apply this method to NA-CORDEX models to estimate changes in 24-h precipitation intensity–frequency (PIF) estimates over the Susquehanna watershed and Florida peninsula. Our approach identifies significant future changes at more stations compared to median-based PIF estimates. The analysis suggests that almost all stations over the Susquehanna and at least two-thirds of the stations over the Florida peninsula will observe significant increases in 24-h precipitation for 2–100-yr return periods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗