Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data needs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

A Ceph S3 Object Data Store for HEP

We present a novel data format design that obviates the need for data tiers by storing individual event data products in column objects. The objects are stored and retrieved through Ceph S3 technology, with a layout designed to minimize metadata volume and maximize data processing parallelism. Performance benchmarks of data storage and retrieval are presented.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Final Design for Thermal/Epithermal eXperiments using High 240 Pu Content Plutonium/Aluminum Zero Power Research Reactor Plates with Polyethylene Moderator (IER 520 Final Design CED-2 Report)

The US Department of Energy Nuclear Criticality Safety Program (NCSP) convened a multinational Thermal Epithermal eXperiments (TEX) meeting in July of 2011 to discuss the data and experimental needs of criticality safety practitioners. The number one and two priority integral experiment data needs were for 239 Pu and 240 Pu, with special emphasis on cross section performance in the intermediate energy range (from 0.625 eV to 100 keV). LLNL measured five critical configurations with LANL for the plutonium test bed (IER-184) and published the experiments as International Criticality Safety Benchmark Evaluation Project evaluation PUMET-MIXED-002. Modeling of the benchmark configurations using ENDF/B-VIII.0 nuclear data showed significant overprediction of reactivity for configurations that had a large percentage of fissions in the intermediate energy regime. This report documents a variation on the TEX plutonium test bed to provide a test of 240 Pu cross sections, with sensitivity of the configuration to 240 Pu radiative capture and fission cross sections a priority for the design.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Maintaining Trust in Reduction: Preserving the Accuracy of Quantities of Interest for Lossy Compression

As the growth of data sizes continues to outpace computational resources, there is a pressing need for data reduction techniques that can significantly reduce the amount of data and quantify the error incurred in compression. Compressing scientific data presents many challenges for reduction techniques since it is often on non-uniform or unstructured meshes, is from a high-dimensional space, and has many Quantities of Interests (QoIs) that need to be preserved. To illustrate these challenges, we focus on data from a large scale fusion code, XGC. XGC uses a Particle-In-Cell (PIC) technique which generates hundreds of PetaBytes (PBs) of data a day, from thousands of timesteps. XGC uses an unstructured mesh, and needs to compute many QoIs from the raw data, f.One critical aspect of the reduction is that we need to ensure that QoIs derived from the data (density, temperature, flux surface averaged momentums, etc.) maintain a relative high accuracy. We show that by compressing XGC data on the high-dimensional, nonuniform grid on which the data is defined, and adaptively quantizing the decomposed coefficients based on the characteristics of the QoIs, the compression ratios at various error tolerances obtained using a multilevel compressor (MGARD) increases more than ten times. We then present how to mathematically guarantee that the accuracy of the QoIs computed from the reduced f is preserved during the compression. We show that the error in the XGC density can be kept under a user-specified tolerance over 1000 timesteps of simulation using the mathematical QoI error control theory of MGARD, whereas traditional error control on the data to be reduced does not guarantee the accuracy of the QoIs.

Gong, Qian↗

MSR Salt Spill Accident Testing Using Eutectic NaCl-UCl 3

Assessing the potential consequences of identified accident scenarios is an essential part of the licensing process for a new nuclear reactor and is achieved by using accident progression models. A likely accident scenario that will be evaluated by developers of molten salt reactors (MSRs) is a spill of radionuclide-bearing fuel salt onto the reactor containment floor (i.e., a salt spill accident). Models to determine the consequences of a molten salt spill accident need to be parameterized and validated using experimental data, but little experimental effort has been dedicated to fill these data gaps to date. Specifically, data is needed to enable model development for individual processes (e.g., spreading, heat transfer, corrosion, radionuclide vaporization, and aerosol generation) that quantifies the sensitivities towards the initial conditions of the spill, the ambient environment, and the features of the containment. In addition, integrated experiments that simulate salt spill accidents will need to be conducted to provide insight into coupled processes and data for model validation, but these experiments will require the use of proven methods to quantify the processes under evaluation.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Call to Action for Global Access to and Harmonization of Quality Information of Individual Earth Science Datasets

Knowledge about the quality of data and metadata is important to support informed decisions on the (re)use of individual datasets and is an essential part of the ecosystem that supports open science. Quality assessments reflect the reliability and usability of data. They need to be consistently curated, fully traceable, and adequately documented, as these are crucial for sound decision- and policy-making efforts that rely on data. Quality assessments also need to be consistently represented and readily integrated across systems and tools to allow for improved sharing of information on quality at the dataset level for individual quality attribute or dimension. Although the need for assessing the quality of data and associated information is well recognized, methodologies for an evaluation framework and presentation of resultant quality information to end users may not have been comprehensively addressed within and across disciplines. Global interdisciplinary domain experts have come together to systematically explore needs, challenges and impacts of consistently curating and representing quality information through the entire lifecycle of a dataset. This paper describes the findings of that effort, argues the importance of sharing dataset quality information, calls for community action to develop practical guidelines, and outlines community recommendations for developing such guidelines. Practical guidelines will allow for global access to and harmonization of quality information at the level of individual Earth science datasets, which in turn will support open science.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Digitalization of an experimental electrochemical reactor via the smart manufacturing innovation platform

The exponential increase in data produced over the last two decades has revolutionized the way we collect, store, process, analyze, model, and interpret information to improve profitability. Manufacturing is no exception. How- ever, Smart Manufacturing, the digital practice, organization, workforce, and infrastructure transformation for collection and deployment of data and models at scale and at all levels of manufacturing, is a complex, costly, and labor-intensive journey that is still seeing slow adoption. The Clean Energy Smart Manufacturing Innovation Institute (CESMII), a national Manufacturing USA public-private partnership sponsored by the Department of Energy, is addressing this scaled use of data and modeling in manufacturing. CESMII has focused on how to col- lect and use operating data for numerous applications that improve productivity, precision, and performance of manufacturing operations from factory floor to supply chain using process simulation, predictive analytics, mon- itoring and control, and real-time optimization. Because contextualized data are key, CESMII has developed the Smart Manufacturing Innovation Platform (SMIP) to lower the barriers to the data that are needed to accelerate data-based model building, improve data visualization, and more quickly gain insights. Reusable, standards-based ways of doing data collection, ingestion, and contextualization are particularly important for scaling access and use of data. The SMIP uses a standards-based definition and construct for reusable information models called an SM Profile. When an SM Profile is used in conjunction with the SMIP, the SMIP ensures the availability of contextualized, operational data for model building. The present work demonstrates Smart Manufacturing and the application of the SMIP for building several data-centered models for the operation and control of an ex- perimental electrochemical reactor that reduces carbon dioxide (CO 2 ) gas to valuable liquid and gas chemicals, such as alcohols, olefins, and syngas. We describe how the SMIP plays a central role in more effective model building and we demonstrate how the electochemical reactor can be controlled and optimized for the desired products. Use of the SMIP involves the transmission of real-time sensor measurements to a cloud resource so that the operating data are available to all model building experts. The data collection and transmission process is fully automated to greatly reduce the need for manual manipulation of the data. Data-driven machine learning models are used for advanced real-time state estimation, real-time optimization, and model-based feedback control for the reactor. The application models are implemented as a system to monitor the data flow and control the electrochemical reactor with a single visualization interface. SM Profiles are used to demonstrate reusability of the information models for the reactor and the instrumentation. The application packages, algorithms, and user interfaces developed are cast as Docker images in a library to facilitate reusability of the application models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Nuclear data resources and initiatives for nuclear astrophysics

Research into the cosmic synthesis of the elements, the evolution and explosion of stars, the nature of the early Universe, and other important topics in nuclear astrophysics are at the forefront of nuclear science. These studies are motivating laboratory measurements and theoretical calculations that, after significant investments, are pushing the boundaries of what is possible. The latest nuclear results, however, must be specially prepared before they can be used to advance our knowledge of the cosmos. This processing requires a set of resources unique to nuclear astrophysics, and an impressive collection of nuclear reaction and nuclear structure datasets, processing codes, thermonuclear reaction rate libraries, and simulation codes and services have been developed for the field. There are, however, some serious challenges to these efforts that will only worsen in the future, making it important to develop strategies and act now to ensure a sustainable future for this work. After detailing the specific data types needed for nuclear astrophysics and the available data resources, the major challenges in this work and their implications are discussed. A set of initiatives are proposed to meet those challenges along with suggested implementations and possible ways that they may advance our understanding of the Universe and strengthen the field of nuclear astrophysics.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Considerations for using Privacy Preserving Machine Learning Techniques for Safeguards

In international nuclear safeguards, the International Atomic Energy Agency (IAEA) is tasked with inspecting and verifying nuclear facilities and their activities. Data analytics and machine learning to support inspections require large amounts of data that nuclear facility operators may consider proprietary or sensitive, so the IAEA may not have full access. Allowing computation over private data without compromising its security therefore has value for safeguards inspections and analysis. Privacy-preserving machine learning (PPML) consists of security-focused techniques that allow data analytics and machine learning algorithms to run on sensitive data without revealing it. This includes ideas like homomorphic encryption (HE), secure multiparty computation (SMPC), and secure enclaves. HE allows algorithms and mathematical operations to be conducted directly on the encrypted data instead of first decrypting it. With SMPC, multiple entities collaboratively compute over distributed data such that no party is able to directly view any others’ original data. Secure enclaves allow computation to take place in a separate and heavily blocked-off section of a CPU. Techniques like these allow for several potential use cases in which the security of data is essential. With SMPC, machine learning models can be trained over the input data from multiple entities, resulting in a model that all users can benefit from without leaking the input data from any particular entity. With SMPC or a zero-knowledge proof (ZKP), an algorithm returning some single answer or truth value can be run on someone else’s data without ever needing to see that data, potentially allowing for verification or proof of some underlying question. HE can allow for outsourcing computation on data to a hostile or untrusted environment. Although most of the research in this field resides within the health and financial domains, tools from PPML may have similar applications in nuclear safeguards. Allowing the IAEA to compute over proprietary information, such as process models and raw sensor data using PPML techniques, provides the baseline for running complex analytics without needing direct unencrypted access to the underlying data, maintaining its privacy. Important limitations to consider for these techniques include the efficiency and level of security required. The security of HE and SMPC come at the cost of speed—the significant amount of overhead means that algorithms implemented in these protocols and encryption schemes are slower than when run on plaintext. Additionally, several important parameters determine what techniques or protocols are used based on the security requirements. SMPC protocols may need to be selected for resistance against a party that attempts to deviate from the protocol to distort the result or gain access to additional information, and a protocol secure against these attacks may further increase the overhead of the algorithm.

97 MATHEMATICS AND COMPUTING↗

Preventing Reverse Engineering of Critical Industrial Data with DIOD

Business analytics augmented by artificial intelligence and machine learning (AI/ML) have revolutionized the role of data in the modern world. In recent years, businesses have incorporated data into their decision-making process for better prediction, risk-assessment, content creation, etc. While such businesses often seek to leverage the full use of their data through third-party AI/ML services, they are often hampered by the risks of data leaks, reverse-engineering, stolen technology, etc. that often have disastrous consequences for businesses and their stakeholders alike. Thus, there arises a need for data masking prior to its transmission that obfuscates proprietary information while preserving the information relevant for AI/ML applications. In order to meet the needs of industrial data which are significantly different from those of data warehouses, previous work proposed an efficient time and space-scalable data masking paradigm known as the deceptive infusion of data (DIOD) methodology. The present work expands upon this work by leveraging existing reverse-engineering capabilities to facilitate the decomposition of industrial data into its proprietary and AI/ML-relevant parts, referred to as fundamental and inference metadata respectively. Both sets of metadata are further obfuscated in accordance with the DIOD methodology to create the DIOD rendition of the industrial data, which is rendered immune to reverse-engineering by discarding proprietary information and only preserving AI/ML-relevant information. Additionally, constraints of the original DIOD manuscript are relaxed using mutual information by configuring the methodology to the target AI/ML application to unlock the full potential of the DIOD methodology. As an example, data from a nuclear reactor is transformed into that from a nonlinear spring-mass system with different levels of data masking as required by the generic system and the target application.

97 MATHEMATICS AND COMPUTING↗

Darshan for HEP applications

Modern HEP workflows must manage increasingly large and complex data collections. HPC facilities may be employed to help meet these workflows’ growing data processing needs. However, a better understanding of the I/O patterns and underlying bottlenecks of these workflows is necessary to meet the performance expectations of HPC systems.Darshan is a lightweight I/O characterization tool that captures concise views of HPC application I/O behavior. It intercepts application I/O calls at runtime, records file access statistics for each process, and generates log files detailing application I/O access patterns.Typical HEP workflows include event generation, detector simulation, event reconstruction, and subsequent analysis stages. A study of the I/O behavior of the ATLAS simulation and filtering stage, and the CMS simulation workflow using Darshan is presented, including insights into the I/O operations and data access size.

Wang, Rui↗

A model to assess Zircaloy’s mechanical property changes following a transient beyond critical heat flux

Maintaining the integrity of nuclear fuel rods is essential for ensuring public health and safety in nuclear power generation. During reactor operation, this integrity is confirmed by demonstrating compliance with established regulatory acceptance criteria. For moderate-frequency events, such as limiting transients and anticipated operational occurrences (AOOs), the current fuel integrity criterion is based on preventing boiling transition. This criterion assumes that prevention of boiling transition will prevent excessive cladding heating and, thus, fuel failure during normal operations. While conservative, this approach places significant constraints on core design, fuel cycle economics, and a plant’s ability to perform major power uprates, leading to suboptimal fuel utilization and inefficient carbon-free energy production. A more efficient approach could be achieved by revising the failure criterion to a material-specific limit rather than strictly preventing the boiling transition, since boiling transition per se is not a cause of fuel cladding failure. Here, as a result, a new licensing framework based on material properties, termed time-at-temperature (t@T), is needed. This approach would allow for brief periods of post–critical heat flux operation during an AOO without compromising safety. Implementing the t@T licensing strategy requires a robust technical foundation in material properties, which must be established through comprehensive data collection on both unirradiated and irradiated fuel and cladding materials. This foundation would enable the development of a safety basis that ensures safe operation while providing greater flexibility and efficiency for reactor operation. This paper documents a thorough review of the available data to establish a baseline knowledge that can inform the development of cladding mechanical models, as well as identify experimental data gaps that need to be addressed in future research. Machine learning and data informatics were utilized to extract the importance of parameters on the t@T parameter. Industry tools were used to perform baseline analyses to define the relevant transient conditions for data analysis. The subsequent review successfully identified applicable experimental data, as well as sufficient data to evaluate changes in cladding mechanical properties following an AOO transient. Rather than developing new models, this work coupled existing irradiation annealing and recrystallization models to calculate changes in hardness, yield stress, and ultimate tensile stress following an AOO event. The findings from this review were summarized to highlight the experimental data needs required to fill remaining gaps and support the development of future t@T licensing methodologies.

Cladding performance↗

Opportunities for Earth Observation to Inform Risk Management for Ocean Tipping Points

Abstract As climate change continues, the likelihood of passing critical thresholds or tipping points increases. Hence, there is a need to advance the science for detecting such thresholds. In this paper, we assess the needs and opportunities for Earth Observation (EO, here understood to refer to satellite observations) to inform society in responding to the risks associated with ten potential large-scale ocean tipping elements: Atlantic Meridional Overturning Circulation; Atlantic Subpolar Gyre; Beaufort Gyre; Arctic halocline; Kuroshio Large Meander; deoxygenation; phytoplankton; zooplankton; higher level ecosystems (including fisheries); and marine biodiversity. We review current scientific understanding and identify specific EO and related modelling needs for each of these tipping elements. We draw out some generic points that apply across several of the elements. These common points include the importance of maintaining long-term, consistent time series; the need to combine EO data consistently with in situ data types (including subsurface), for example through data assimilation; and the need to reduce or work with current mismatches in resolution (in both directions) between climate models and EO datasets. Our analysis shows that developing EO, modelling and prediction systems together, with understanding of the strengths and limitations of each, provides many promising paths towards monitoring and early warning systems for tipping, and towards the development of the next generation of climate models.

Wood, Richard A. (ORCID:0000000239609513)↗

The composition of gases from a diffusion flame above longleaf pine needle fuel beds

The gas and tar composition of a diffusion flame from longleaf pine needles is currently poorly understood and more data are needed to fill in the gap between pyrolysis data and smoke plume data, thus improving physical and chemical modeling of wildland smoke formation. A pilot experiment to measure light gas and tar composition of such a flame is described for three flame regions: persistent flame (flame base), intermittent flame, and smoke plume. Flame gases from 24 experimental fires were collected in canisters and analyzed using EPA method TO-14A for CO 2 , CO, H 2 , CH 4 , and C 2 to C 7 hydrocarbon gases. Condensed gas (tar) samples were collected and analyzed using GC/MS. Other light gases were measured using FTIR spectroscopy. Results from compositional data analysis suggest significant differences in (relative) concentration of compounds detected in the three regions of the flame. Statistical tests for differences in flame zones were performed using the canister data: Concentration of hydrocarbons relative to CO and CO 2 decreased from the persistent flame zone above the pyrolyzing needles through the intermittent flame region into the flame-free plume. This was likely due to both chemical reactions (oxidation) occurring in the flame as well as the introduction of air into the flame/plume by entrainment.

Biomass↗

IM3 Phase 2 Official Simulations: GCAM-Demeter-SELECT Annualized Land Use and Land Cover, Wood Harvest and Fertilization Data with Dynamic Urbanization Harmonized to CLM Land Definitions at 0.125 Degrees

Annualized land use land cover data, including wood harvest and fertilizer use data from the Global Change Analysis Model (GCAM) downscaled to 0.125 degrees for couping with the Community Land Model (CLM). GCAM here refers to GCAM-USA v5.3.im3 which has an enhanced electricity sector and an updated data system needed to represent regional to local scale dynamics. Data is also harmonized with future urbanization projections from the Spatially-Explicit, Long-term, Empirical City developmenT (SELECT) model. Projections/Data are generated using the demeter land use and land cover downscaling model. Original projections were generated at 0.05 degrees before being aggregared to 0.125 degrees. Projections are available for 8 alternative scenarios. Two versions of final data are included- one with managed forests or harvested forest area per pixel broken out and one with the same aggregated into total forests. Following folders are included: demeter_78_PFT_output:This is the final output of dynamic land use land cover change for 78 PFTs as required by CLM raw_outputs_incl_managed_forest: This is the final output but with managed forests broken out as a different PFT. Essentially a 79th PFT is added. wood_harvest_outputs: Wood harvest output per pixel in gC/m2 fertilization_outputs: Fertilizer use per pixel in gN/m2 Each NetCDF file in each folder represents a projection for a separate year, scenario. Land use outputs are organized as PFT level data saved as subdata. Link to GCAM version used- https://data.msdlive.org/records/yb23g-44274 Link to SELECT documentation -https://www.sciencedirect.com/science/article/pii/S1364815219301707 Link to CLM documentation- - https://www.cesm.ucar.edu/models/clm In case of questions contact- kanishka.narayan@pnnl.gov

GCAM-USA↗

Developing an Interactive Landscape for Mobility Resources: Preprint

As the world continues to be increasingly driven by data, the ways researchers and professionals sort and collect this data is critical. In the world of mobility data, new levels of data from public transportation systems, location services, and other means are being lost due to how little organization exists. Much of the data is proprietary, and there are few if any de jure or even de facto standards connecting data. There is also little knowledge about the gaps that exist in the data. In this project, we created an interactive landscape where mobility resources are categorized and organized in an easy to use, living document. We made this landscape with open-source code from the CNCF Cloud Native Landscape and repurposed it to the mobility data's needs. Additionally, unlike previous sources that organize mobility data, this document can be updated through GitHub by those in the field to keep its sources relevant. Following the creation of a beta version of the landscape, we conducted several interviews with industry researchers and professionals to ensure the landscape would be useful. The result is an online hub where mobility researchers and resource creators can easily access research and collaborate.

ADVANCED PROPULSION SYSTEMS↗

Energy and AI: Evaluating Future Grid and Water Stress Due to Data Centers

Projections of the need for new data centers to support Artificial Intelligence (AI) are large but highly uncertain. Recent projections indicate up to a 15% annual growth rate in data center electricity demand within the next 5-10 years. Given that most electric utilities are required to have a reserve margin of roughly the same magnitude as the projected growth in demand, these new data center loads could soon threaten resource adequacy and reliability unless data centers build their own generation, interruptible loads are negotiated, commensurate new capacity and/or transmission is built, or some combination of these options. Similarly, depending on the cooling technology and geographic location of new data centers, they could threaten water adequacy in water scarce regions. This presentation highlights the grid and water implications of new data centers to support AI.

Mongird, Kendall (ORCID:0000000328077088)↗

Scaling SQL to the Supercomputer for Interactive Analysis of Simulation Data

AI and simulation workloads consume and generate large amounts of data that need to be searched, transformed and merged with other data. With the goal of treating data as a first-class citizen inside a traditionally compute-centric HPC environment, we explore how the use of accelerators and high-speed interconnects can speed up tasks which otherwise constitute bottlenecks in computational discovery workflows. BlazingSQL is SQL engine that runs natively on NVIDIA GPUs and supports internode communication for fast analytics on terabyte-scale tabular data sets. We show how a fast interconnect improves query performance if leveraged through the Unified Communication X (UCX) middleware. We envision that future computing platforms will integrate accelerated database query capabilities for immediate and interactive analysis of large simulation data.

Glaser, Jens↗