Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data needs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Effect of spacer grids on high-burnup fuel fragmentation, relocation, and dispersal

Increasing the fuel burnup limit in light-water reactors to improve fuel cycle economics requires a strong technical foundation. Experimental observations from the Halden and Studsvik programs have revealed severe fuel fragmentation during loss-of-coolant accident (LOCA) conditions, highlighting the need for additional technical evaluation. Consequently, further LOCA test data are needed to complement existing findings and improve the understanding of fuel fragmentation, relocation, and dispersal (FFRD) behavior. Oak Ridge National Laboratory’s Severe Accident Test Station has played a significant role in advancing the understanding of high-burnup fuel fragmentation, relocation, and dispersal phenomena. One remaining gap in the available experimental database is the effect of fuel assembly structural features on cladding deformation behavior during a LOCA, and more specifically, their impact on the fuel’s ability to fragment, relocate, and disperse. Recent analyses using the BISON fuel performance code suggest that cladding deformation near grid spacers will remain below the 3% threshold that has been reported in the NRC Research Information Letter, indicating that the cladding could remain mechanically constrained during the LOCA event. This paper builds upon the BISON analyses to design and conduct a series of out-of-cell tests aimed at further evaluating cladding deformation in and around grid spacers. In addition, these tests were used to assess local cladding temperature conditions and compare them against analytical predictions in order to better replicate expected in-reactor behavior. Finally, an in-cell high-burnup LOCA test was designed and performed to evaluate the effects of a grid spacer, or cladding restraint, on fuel fragmentation, relocation, and dispersal susceptibility. The high-burnup test results differed from those of historical LOCA experiments, with a recorded rupture temperature of 861°C. Two ballooned regions and corresponding rupture openings were observed, with rupture widths of approximately 0.64 mm for both ruptures and rupture lengths of 4.8 mm and 5.6 mm, respectively.

Capps, Nathan [ORNL]↗

FunMC^2: A Filter for Uncertainty Visualization of Marching Cubes on Multi-Core Devices

Visualization is an important tool for scientists to extract understanding from complex scientific data. Scientists need to understand the uncertainty inherent in all scientific data in order to interpret the data correctly. Uncertainty visualization has been an active and growing area of research to address this challenge. Algorithms for uncertainty visualization can be expensive, and research efforts have been focused mainly on structured grid types. Further, support for uncertainty visualization in production tools is limited. In this paper, we adapt an algorithm for computing key metrics for visualizing uncertainty in Marching Cubes (MC) to multi-core devices and present the design, implementation, and evaluation for a Filter for uncertainty visualization of Marching Cubes on Multi-Core devices (FunMC2). FunMC2 accelerates the uncertainty visualization of MC significantly, and it is portable across multi-core CPUs and GPUs. Evaluation results show that FunMC2 based on OpenMP runs around 11× to 41× faster on multi-core CPUs than the corresponding serial version using one CPU core. FunMC2 based on a single GPU is around 5× to 9× faster than FunMC2 running by OpenMP. Moreover, FunMC2 is flexible enough to process ensemble data with both structured and unstructured mesh types. Furthermore, we demonstrate that FunMC2 can be seamlessly integrated as a plugin into ParaView, a production visualization tool for post-processing.

Wang, Jay↗

An Approach to Data Management Planning for Protected Data Projects

This report describes what is required in a data management plan for data that needs to be protected in some fashion. Here we provide an overview of what a data manager should consider, including the data ingestion and various extract-transform-load processes, metadata considerations and documentation, to what might need to be accounted for in the event of data loss. In addition, this report includes two appendices: forms that, when filled out, make the user compliant with DOE data management requirements as well as additional requirements for handling protected data at ORNL.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Homomorphic Encryption for Electrical Metering Aggregation: Protecting the Privacy of Building Tenants

Electrical meters are devices that measure consumer electricity usage. The data collected by these meters is necessary for utility billing and electrical grid management but can also be used to assess the environmental impact of buildings. Prior research has found that unprotected metering data could potentially be used to infer some information about the behaviors of building tenants by detecting changes in electricity usage. For example, a period of low electricity usage could suggest that the tenants are not in the building. As smart metering becomes more common, there is a growing need for data privacy protections for metering data that do not negatively impact the quality and availability of data used for energy management and billing applications. To identify potential solutions, we developed a Python-based data aggregation platform to analyze the potential efficacy of privacy-enhancing technologies for energy metering applications. This platform aggregates groups of metering sites into virtual buildings, which could potentially detach changes in electrical activity from individual tenants, making it more difficult to track the activity of a specific tenant. To further protect data during analysis, this project utilizes homomorphic encryption as part of its initial approach. Homomorphic encryption offers a means of protecting energy consumption data while permitting mathematical operations to be performed without the need to know the data contents. This allows for data to be processed into usable statistics without revealing energy consumption information. A series of homomorphic encryption libraries were evaluated to determine their applicability and limitations in the context of metering data. The use of these techniques may help to reassure consumers and encourage further adoption of smart grid infrastructure.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗

TURBO: Terabits/s Using Reconfigurable Bandwidth Optics (Final Report)

Large scale, big science applications are generating petabytes of data that need to be shared among multiple locations. Moving a petabyte of data at 100 Gb/s takes 22 hours. Reducing this to single digit hours or below is clearly attractive and with future applications anticipating exabyte scale transfers, multi-terabit per second capacities are needed. This scale of capacity is available using optical systems. Core optical systems have aggregate per fiber capacities on the order of 10 Tb/s and more than 100 Tb/s system experiments have been realized in the lab. But large optical capacities are not just reserved for the core as similar system capacities are available in metro networks right up to the enterprise, campus, or data center. In general, the same technology is used today in the metro area as in the long-haul networks—reconfigurable optical add drop multiplexing (ROADM) node-based wavelength division multiplexing (WDM) systems. Data centers already support massive internal capacities on the scale of 100’s of Tb/s. As new photonic integrated technologies mature, for example through initiatives such as the Integrated Photonics Innovative Manufacturing Institute (IP-IMI), the cost of optical interfaces is expected to decrease, and higher capacity and higher performance interfaces will become more affordable. This set of circumstances creates the potential that Tb/s capacities will be available at the enterprise and campus level to support large scale science network applications.

42 ENGINEERING↗

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the FECM NETL Carbon Management Program Review Meeting 2024.

Creason, Christopher↗

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the Geological Society of America Connects 2024 Annual Meeting in Anaheim, California, 22-25 September 2024.

Creason, Christopher↗

Recent advances in integrated hydrologic models: Integration of new domains

Over the past several decades, hydrologic models have advanced from independent models of the surface and subsurface to integrated models that can capture the terrestrial hydrologic cycle within one framework. In recent years, these coupled frameworks have seen the inclusion of biogeochemical processes, ecohydrology, sedimentation and erosion, cold region hydrology, anthropogenic activities, and atmospheric processes. This expansion is the result of increased computational, data, and modeling capabilities and capacities, as well as improved understanding of the processes that drive these integrated systems. Here, in this study, we review these recent advances to integrate new processes and systems into existing terrestrial hydrologic models and highlight the significant challenges and opportunities that remain. We identify that with so many models currently available and in development, selecting the most appropriate model is difficult, and we suggest a path for new or novice modelers to find the most appropriate code based on their needs. In addition, data required to parameterize and calibrate these models can often constrain their applicability and usefulness. However, advances in environmental sensors and measurement technology, in addition to data assimilation of non-traditional data (e.g. remote sensing, qualitative data) are providing new ways of addressing this issue. As we expand hydrologic models to integrate more processes and systems, our computational demands also increase. Recent and emerging advances in computational platforms, including cloud and quantum computing, in addition to the use of machine learning to capture some processes, will continue to support the use of increasingly larger and more complex, process-based models. Finally, we highlight that it is critical to develop state-of-the-science models that are accessible to all model users, not just those applied for research and development. We encourage continued development of diverse modeling platforms, considering the user needs, data availability, and computational resources.

54 ENVIRONMENTAL SCIENCES↗

Feature Analysis, Tracking, and Data Reduction: An Application to Multiphase Reactor Simulation MFiX-Exa for In-Situ Use Case

As we enter the exascale computing regime, powerful supercomputers continue to produce much higher amounts of data than what can be stored for offline data processing. To utilize such high compute capabilities on these machines, much of the data processing needs to happen in situ, when the full high-resolution data is available at the supercomputer memory. In this article, we discuss our MFiX-Exa simulation, which models multiphase flow by tracking a very large number of particles through the simulation domain. In one of the use cases, the carbon particles interact with air to produce carbon dioxide bubbles from the reactor. These bubbles are of primary interest to the domain experts for these simulations. For this particle-based simulation, we propose a streaming technique that can be deployed in situ to efficiently identify the bubbles, track them over time, and use them to down-sample the data with minimal loss in these features.

97 MATHEMATICS AND COMPUTING↗

High-performance data format for scientific data storage and analysis

Here, in this article, we present the High-Performance Output (HiPO) data format developed at Jefferson Laboratory for storing and analyzing data from Nuclear Physics experiments. The format was designed to efficiently store large amounts of experimental data, utilizing modern fast compression algorithms. The purpose of this development was to provide organized data in the output, facilitating access to relevant information within the large data files. The HiPO data format has features that are suited for storing raw detector data, reconstruction data, and the final physics analysis data efficiently, eliminating the need to do data conversions through the lifecycle of experimental data. The HiPO data format is implemented in C++ and JAVA, and provides bindings to FORTRAN, Python, and Julia, providing users with the choice of data analysis frameworks to use. In this paper, we will present the general design and functionalities of the HiPO library and compare the performance of the library with more established data formats used in data analysis in High Energy and Nuclear Physics (such as ROOT and Parquete). In columnar data analysis, HiPO surpasses established data formats in performance and can be effectively applied to data analysis in other scientific fields.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Unsupervised multimodal fusion of in-process sensor data for advanced manufacturing process monitoring

Effective monitoring of manufacturing processes is crucial for maintaining product quality and operational efficiency. Modern manufacturing environments often generate vast amounts of complementary multimodal data, including visual imagery from various perspectives and resolutions, hyperspectral data, and machine health monitoring information such as actuator positions, accelerometer readings, and temperature measurements. However, fusing and interpreting this complex, high-dimensional data presents significant challenges, particularly when labeled datasets are unavailable or impractical to obtain. This paper presents a novel approach to multimodal sensor data fusion in manufacturing processes, inspired by the Contrastive Language-Image Pre-training (CLIP) model. We leverage contrastive learning techniques to correlate different data modalities without the need for labeled data, overcoming limitations of traditional supervised machine learning methods in manufacturing contexts. Our proposed method demonstrates the ability to handle and learn encoders for five distinct modalities: visual imagery, audio signals, laser position (x and y coordinates), and laser power measurements. By compressing these high-dimensional datasets into low-dimensional representational spaces, our approach facilitates downstream tasks such as process control, anomaly detection, and quality assurance. The unsupervised nature of our method makes it broadly applicable across various manufacturing domains, where large volumes of unlabeled sensor data are common. We evaluate the effectiveness of our approach through a series of experiments, demonstrating its potential to enhance process monitoring capabilities in advanced manufacturing systems. This research contributes to the field of smart manufacturing by providing a flexible, scalable framework for multimodal data fusion that can adapt to diverse manufacturing environments and sensor configurations. The proposed method paves the way for more robust, data-driven decision-making in complex manufacturing processes.

Contrastive Learning↗

ggtaxplot v 0.0.1

ggtaxplot is an R package designed to process and visualize taxonomic data through a taxonomic river plot. This package is ideal for researchers and data scientists who need to visualize taxonomic data. ggtaxplot function processes data and generates a taxonomic river plot, allowing users to visualize the distribution of taxa across different samples.

Coclet, Clement [Lawrence Berkeley National Labora↗

Development of Prognostic Models Using Plant Asset Data

The recent growth of machine learning and artificial intelligence technologies provides opportunities for leveraging data-driven algorithms to address the problems of diagnostics and prognostics in the nuclear power industry. The use of machine learning and other statistical methods as prognostic models is of particular interest in the nuclear industry to accurately predict future equipment or plant state given a set of measurements. Such predictive capability will enable predictive assessment of component condition and remaining life and allow for condition-based predictive maintenance. The resulting optimization of maintenance scheduling and reduction in unnecessary maintenance activities will lower overall maintenance costs and improve the economics of nuclear power. This report discusses the various aspects of data processing and model development that are likely to influence the performance of prognostic models. Data from a boiling-water reactor was used to evaluate several prognostic models to identify key considerations for developing such models to predict data-driven plant state and equipment degradation condition. Preliminary results indicate the need for data sets that are relevant to the problem at hand and contain signatures that may be correlated to the prediction problem. Assuming such data exist, development of prognostic models using data-driven methods requires an understanding of the various sources of influence on the prediction accuracy (such as the model architecture, data preprocessing approaches, and potentially external factors influencing the equipment or plant system under assessment). Ongoing research is evaluating these factors in greater detail and examining techniques for calculating prediction uncertainty bounds.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Emulation Modeling for Development of Cyber-Defense Capabilities for Satellite Systems

The objective of this project was to develop a novel capability to generate synthetic data sets for the purpose of training Machine Learning (ML) algorithms for the detection of malicious activities on satellite systems. The approach experimented with was to a) generate sparse data sets using emulation modeling and b) enlarge the sparse data using Generative Adversarial Networks (GANs). We based our emulation modeling on the Open Source NASA Operational Simulator for Small Satellites (NOS3) developed by the Katherine Johnson Independent Verification and Validation (IV&V) program in West Virginia. Significant new capabilities on NOS3 had to be developed for our data set generation needs. To expand these data sets for the purpose of training ML, we experimented with a) Extreme Learning Machines (ELMs) and b) Wasserstein-GANs (WGAN-GP).

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Foundations of Molecular 'Isotomics'

The naturally occurring rare isotopes are versions of common elements, such as hydrogen, carbon and oxygen, that contain a larger than usual number of neutrons in their atomic nuclei and therefore are higher in mass than the common atoms of that element. Isotopes exist for most elements and are found in most natural and synthetic materials, but are uneven in their distribution because chemical and physical processes are isotope-selective (e.g., a chemical reaction may proceed more rapidly for one isotope than for another). For this reason, abundances of isotopes in a material of interest can provide a record, or ‘signature’ of various features of that material’s origin and history. These signatures have been used in the geo, life, chemical and physical sciences in a wide variety of ways over close to 8 decades. However, many such applications struggle to reach unique interpretations of isotopic data because multiple factors combine to control a given sample’s overall isotopic content. That is, the factors controlling isotopic content are too numerous and complex to fully constrain from a simple measurement of a material’s isotope abundances. However, the distribution of isotopes within materials, at molecular scales potentially provides a vastly larger number and diversity of constraints on the chemical and physical processes that comprise a material’s history. The rare isotopes may be concentrated into one atomic position in a molecule relative to another, some proportion of molecules in a sample may contain two or more rare isotopes, and those multiply-isotope-substituted forms of molecules may also have uneven distributions of those isotopes across individual atomic sites. For these reasons, even small, seemingly simple molecules, such as sugars, amino acids or drug compounds, actually exist in a vast number of isotopically unique forms (often millions or more), and each one of those forms is in some sense an independent ‘vote’ on that sample’s history. This project has focused on opening this rich archive of information by enabling the creation of routinely and widely applicable ways of measuring and interpreting isotopic structures of molecules. This work has included the development of core technologies and analytical methods, advancing fundamental understanding of the physical and chemical properties of isotopic versions of molecules, and conducting proof of concept studies of illustrative geochemical, cosmochemical and forensic problems in order to show how these technologies, methods and principles come together to solve problems in new ways. A key to the success of this project was the adaptation of ‘Fourier transform mass spectrometry’ (FTMS) to the task of precisely measuring proportions of the rare, naturally occurring isotopic forms of molecules. FTMS is a highly specialized form of mass spectrometry that traps ions within magnetic or electrostatic cavities and, effectively, ‘listens’ (through registering of subtle electrical signals) to the harmonic signals they make while rapidly orbiting within those cavities. These signals have periods that are a function of their mass and strength (or ‘loudness’) that is proportional to their abundances. Thus, these signals constrain relative amounts of molecules that differ in their mass due to various isotopic substitutions. This technology has been essential to the identification of organic molecules in the life, chemical and environmental sciences for over 4 decades, but generally has lacked the control, stability and precision to meaningfully measure rare isotope forms of molecules. This project’s most fundamental contribution has been to modify FTMS, both in terms of hardware and methods, to enable such measurements. The raw data of molecular isotopic structure is tremendously voluminous and complex, so another important activity of this project has been developing the theoretical and data-science tools needed to interpret the data generated by this new form of isotopic measurement. A particularly challenging part of this task has been predicting molecular isotopic structure, as only through the comparison of measurements with predictions can we make progress on hypothesis driven research questions. We have attacked this this prediction task through a combination of first-principles chemical-physics models of the effects of isotope substitution on molecule properties and data-science models that permit us to generalize that chemical physics to cases that have not yet been studied by detailed chemical physics theory. The proof of concept applications we have pursued over the course of this study include biological reactions of amino acids and other biomolecules, non-biological synthesis of organic molecules in extra-terrestrial settings such as meteorites, petroleum geoscience questions concerning the origin and evolution of natural gas, oil and kerogen compounds, and forensic questions such as the sourcing of chemical weapons. The successes of these applications have laid the groundwork for the next phase of this field’s development, which will include larger scale and more ambitious studies of molecular isotopic structure as a means of diagnosing human diseases, such as cancer, and reconstructing detailed interpretations of the origin and evolution of organic molecules in modern and geological environments.

Cesar, Jaime↗

Validation of the Cloud_CCI (Cloud Climate Change Initiative) cloud products in the Arctic

The role of clouds in the Arctic radiation budget is not well understood. Ground-based and airborne measurements provide valuable data to test and improve our understanding. However, the ground-based measurements are intrinsically sparse, and the airborne observations are snapshots in time and space. Passive remote sensing measurements from satellite sensors offer high spatial coverage and an evolving time series, having lengths potentially of decades. However, detecting clouds by passive satellite remote sensing sensors is challenging over the Arctic because of the brightness of snow and ice in the ultraviolet and visible spectral regions and because of the small brightness temperature contrast to the surface. Consequently, the quality of the resulting cloud data products needs to be assessed quantitatively. In this study, we validate the cloud data products retrieved from the Advanced Very High Resolution Radiometer (AVHRR) post meridiem (PM) data from the polar-orbiting NOAA-19 satellite and compare them with those derived from the ground-based instruments during the sunlit months. The AVHRR cloud data products by the European Space Agency (ESA) Cloud Climate Change Initiative (Cloud_CCI) project uses the observations in the visible and IR bands to determine cloud properties. The ground-based measurements from four high-latitude sites have been selected for this investigation: Hyytiälä (61.84°N, 24.29°E), North Slope of Alaska (NSA; 71.32°N, 156.61°W), Ny-Ålesund (Ny-Å; 78.92°N, 11.93°E), and Summit (72.59°N, 38.42°W). The liquid water path (LWP) ground-based data are retrieved from microwave radiometers, while the cloud top height (CTH) has been determined from the integrated lidar–radar measurements. The quality of the satellite products, cloud mask and cloud optical depth (COD), has been assessed using data from NSA, whereas LWP and CTH have been investigated over Hyytiälä, NSA, Ny-Å, and Summit.

54 ENVIRONMENTAL SCIENCES↗