Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Pando

SAND2025-02006O Pando is a distributed data analysis software tool. It is designed to handle large-scale graph analysis problems, often with a specific focus on blockchain/cryptocurrency data. Pando handles scalability by running on a distributed cluster of servers. Users can customize the output using the program’s plugin/extension design methodology. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Gabert, Kasimir↗

Nuclear Physics Exascale Requirements Review: An Office of Science Review sponsored jointly by Advanced Scientific Computing Research and Nuclear Physics, June 15 - 17, 2016, Gaithersburg, Maryland

Imagine being able to predict — with unprecedented accuracy and precision — the structure of the proton and neutron, and the forces between them, directly from the dynamics of quarks and gluons, and then using this information in calculations of the structure and reactions of atomic nuclei and of the properties of dense neutron stars (NSs). Also imagine discovering new and exotic states of matter, and new laws of nature, by being able to collect more experimental data than we dream possible today, analyzing it in real time to feed back into an experiment, and curating the data with full tracking capabilities and with fully distributed data mining capabilities. Making this vision a reality would improve basic scientific understanding, enabling us to precisely calculate, for example, the spectrum of gravity waves emitted during NS coalescence, and would have important societal applications in nuclear energy research, stockpile stewardship, and other areas. This review presents the components and characteristics of the exascale computing ecosystems necessary to realize this vision.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Nuclear Physics Network Requirements Review (Final Report)

The Energy Sciences Network (ESnet) is the high-performance network user facility for the US Department of Energy (DOE) Office of Science (SC) and delivers highly reliable data transport capabilities optimized for the requirements of data-intensive science. In essence, ESnet is the circulatory system that enables the DOE science mission by connecting all its laboratories and facilities in the US and abroad. ESnet is funded and stewarded by the Advanced Scientific Computing Research (ASCR) program and managed and operated by the Scientific Networking Division at Lawrence Berkeley National Laboratory (LBNL). ESnet is widely regarded as a global leader in the research and education networking community. ESnet interconnects DOE national laboratories, user facilities, and major experiments so that scientists can use remote instruments and computing resources as well as share data with collaborators, transfer large datasets, and access distributed data repositories. ESnet is specifically built to provide Between July 2023 and October 2023, ESnet and the Nuclear Physics program (NP) of the DOE SC organized an ESnet requirements review of NP-supported activities. Preparation for these events included identification of key stakeholders: program and facility management, research groups, and technology providers. Each stakeholder group was asked to prepare formal case study documents about its relationship to the NP program to build a complete understanding of the current, near-term, and long-term status, expectations, and processes that will support the science going forward.

97 MATHEMATICS AND COMPUTING↗

High Energy Physics Network Requirements Review: Two-Year Update

The Energy Sciences Network (ESnet) is the high-performance network user facility for the US Department of Energy (DOE) Office of Science (SC) and delivers highly reliable data transport capabilities optimized for the requirements of data-intensive science. In essence, ESnet is the circulatory system that enables the DOE science mission by connecting all its laboratories and facilities in the US and abroad. ESnet is funded and stewarded by the Advanced Scientific Computing Research (ASCR) program and managed and operated by the Scientific Networking Division at Lawrence Berkeley National Laboratory (LBNL). ESnet is widely regarded as a global leader in the research and education networking community. ESnet interconnects DOE national laboratories, user facilities, and major experiments so that scientists can use remote instruments and computing resources as well as share data with collaborators, transfer large datasets, and access distributed data repositories. ESnet is specifically built to provide a range of network services tailored to meet the unique requirements of the DOE’s data-intensive science. In July 2023, the Energy Sciences Network (ESnet) and the High Energy Physics program (HEP) of the DOE SC organized an interim ESnet requirements review of HEP-supported activities, to follow up on the work started during the 2020 HEP Network Requirements Review. Preparation for these events included checking back with the key stakeholders: program and facility management, research groups, and technology providers. Each stakeholder group was asked to prepare updates to their previously submitted case study documents, so that ESnet could update the understanding of any changes to the current, near-term, and long-term status, expectations, and processes that will support the science activities of the program.

97 MATHEMATICS AND COMPUTING↗

Transforming Energy Through Computational Excellence: Advanced Scientific Visualization Reveals Energy Insights

The National Renewable Energy Laboratory's world-class researchers and analysts, along with the Insight Center (our state-of-the-art scientific visualization facility) make data immersion a reality, allowing users to step into and explore their data. With the rise of large, diverse, and distributed data sets, scientific visualization is now critical to the process of scientific discovery and to managing and analyzing data and extracting insights. NREL provides visualization capabilities and facilities that are supported by state-of-the-art equipment, leading-edge techniques, and expert staff.

data science↗

Strength distributions of laminated FeNi-based metal amorphous nanocomposite ribbons

Metal Amorphous Nanocomposite (MANC) materials offer low losses at high magnetic switching frequency, enabling high power density motors with increased rotational speed. While MANCs have high strength, they are brittle. The use of motor components such as a rotor consisting of brittle material presents a reliability concern. Here, a promising MANC alloy is subjected to tensile tests and failure is observed with high-speed photography. A method is developed to prepare tensile specimens of laminated MANC and epoxy layers, simulating the stacking of an epoxy-impregnated tape-wound core. Tensile tests are conducted for single layer ribbon and for five- and ten-layer stacks of laminated material with thin layers of thermosetting epoxy. Failure distributions are shown to have increasing Weibull modulus with increasing layer count. The composite MANC material system is modeled using chain-of-bundles models. Using a k-failure model, we show that single ribbon strength distribution data can be used to predict well the failure distribution of laminated stacks. The agreement occurs when the assumed ineffective length, over which load is recovered in a failed layer, is comparable to the observed interlaminar separation length.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Uncertainty Quantification via Stable Distribution Propagation

We propose a new approach for propagating stable probability distributions through neural networks. Our method is based on local linearization, which we show to be an optimal approximation in terms of total variation distance for the ReLU non-linearity. This allows propagating Gaussian and Cauchy input uncertainties through neural networks to quantify their output uncertainties. To demonstrate the utility of propagating distributions, we apply the proposed method to predicting calibrated confidence intervals and selective prediction on out-of-distribution data. The results demonstrate a broad applicability of propagating distributions and show the advantages of our method over other approaches such as moment matching.

Artificial Intelligence (cs.AI)↗

A data science approach for analysis and reconstruction of spinodal-like composition fields in irradiated FeCrAl alloys

A statistical method for the analysis of continuously distributed data representative of composition fluctuations in irradiated FeCrAl alloys acquired using Energy Dispersive X-ray Spectroscopy (EDS) method is presented. Using probability distribution functions, direct and cross-covariances between the elemental compositions, the effects of alloy composition and irradiation dose were investigated on the spatial distribution and length scale of composition fluctuations at the nanoscale. We have observed that, for neutron-irradiated FeCrAl alloys, the distribution of Fe and Cr followed a left-skewed and right-skewed distribution, respectively for all (average) alloy compositions and irradiation doses. The analysis also revealed enhanced spatial gradients in the elemental compositions at higher irradiation dose. Direct and cross-covariance estimates of the experimental data were also utilized for reconstruction of composition data through fitting it to a parametric form of the covariance functions. Linear Model of Coregionalization was used to determine the parameters of the covariance functions. Subsequently, a spectral method was utilized for simulating a realization of the alloy compositions. Close correspondence was observed between the experimental and the reconstructed data which was analyzed using probability distribution functions and covariance functions. Composition space of the experimental and reconstructed data and dislocation velocities as a function of applied stress and line directions over the entire composition maps were also examined.

36 MATERIALS SCIENCE↗

HEPnOS: a Specialized Data Service for High Energy Physics Analysis

In this paper, we present HEPnOS, a distributed data service for managing data produced by high-energy physics (HEP) experiments. Using HEPnOS, HEP applications can use HPC resources more effciently than traditional fle-based applications. The fle-based model leads to a rigid, chunk-based allocation of computational resources and limits the number of cores that can be used concurrently by an HEP application. The fundamental problem is that organizing domain-specifc data into fles inadvertently introduces a single, artifcial, confated tuning parameter that puts key optimization goals into confict: larger fle sizes reduce metadata overhead and thus improve I/O effciency, but smaller fle sizes provide more opportunity for workfow parallelism and load balancing. In this work, we introduce a domain-specifc data service that decouples that constraint so that data can be accessed and processed in its natural granularity while still maintaining I/O effciency. By removing the constraints introduced by fle handling we are able to obtain better scaling and make effcient use of more cores for processing a fxed-sized data sample. We demonstrate the improved scalability by using an application developed in the fle-based paradigm and comparing it to a version modifed to use HEPnOS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Risk-Informed Approach to Trustworthiness Assessment in Digital Twins-Based Autonomous Control

In autonomous control systems, digital twins (DTs) are used to perform diagnostic and prognostic functions. The trustworthiness of these DTs is dependent on quality and coverage of the training data, model accuracy and integrity of sensor data. This work introduces a methodology to determine the trustworthiness of a DT system given faulty sensor data using a risk informed approach. Bayesian Belief Networks (BBNs) are used to propagate uncertainties and determine the probability of trustable recommendations. The decision to trust the control action provided by the DT is based on the DT output, expert opinion, and severity of problems. The performance of DTs is reliant on the data they are trained on. When they encounter out of distribution data, the trustworthiness of the recommendations decreases. To address this issue, we include an expert component that provides input on sensor degradation. For this, we utilize a generative artificial intelligence (AI) model, such as Generative Pretrained Transformer (GPT). The GPT functions as an expert with broad knowledge. The GPT is fine-tuned to understand and discriminate sensor degradation scenarios using manufactured data. This methodology is demonstrated through a case study on a Nearly Autonomous Management and Control System (NAMAC) during a steady state scenario. Various sensor degradation types with different severity levels are considered. Degraded sensor data is processed by the DT system and the fine-tuned GPT. Finally, using the BBN, we combine the GPT information and the DT output with its sources of uncertainty. This provides an output regarding the trustworthiness of the DT recommendation.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

A Multi-criteria CCUS Screening Evaluation of the Gulf of Mexico, USA - Supplementary Data

The Wendt et al. study aims to incorporate multiple and disparate carbon capture, utilization, and storage (CCUS) decision-making criteria into a systematic, quantitative analytical approach to help identify areas with potentially high suitability to serve as offshore CO2 storage or EOR regions. Spatially-distributed data from publicly-available sources within the Gulf of Mexico (GOM) study area (limited to federal waters; state waters were not evaluated) was compiled using the U.S. Department of Energy's (DOE) National Energy Technology Laboratory (NETL) Cumulative Spatial Impact Layers™ (CSIL) tool to easily aggregate data based on evenly-distributed grids across the study region set at a resolution of approximately 25 square miles (65 square kilometers). The data included in this Microsoft Excel™ workbook provide the aggregated scores for each grid point across the study domain as well the weighting for each criterion under the four scenarios evaluated.

CCUS↗

HydraGNN v3.0

New or improved capabilities included in v3.0 release are as follows: 1. Enhancement in message passing layers through generalization of the class inheritance to enable the inclusion of a broader set of message passing policies Inclusion of equivariant message passing layers from the original implementations of: SchNet (https://pubs.aip.org/aip/jcp/article/148/24/241722/962591/SchNet-A-deep-learning-architecture-for-molecules); DimeNet++ (https://arxiv.org/abs/2011.14115); EGNN models (https://arxiv.org/pdf/2102.09844.pdf) 2. Restructuring of class inheritance for data management 3. Support of DDStore https://github.com/ORNL/DDStore capabilities for improved distributed data parallelism on large volumes of data that cannot fit on intra-node memory capacities 4. Large-scale system support for OLCF-Crusher and OLCF-Frontier

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

A kinetic-based regularization method for data science applications

We propose a physics-based regularization technique for function learning, inspired by statistical mechanics. By drawing an analogy between optimizing the parameters of an interpolator and minimizing the energy of a system, we introduce corrections that impose constraints on the lower-order moments of the data distribution. This minimizes the discrepancy between the discrete and continuum representations of the data, in turn allowing to access more favorable energy landscapes, thus improving the accuracy of the interpolator. Our approach improves performance in both interpolation and regression tasks, even in high-dimensional spaces. Unlike traditional methods, it does not require empirical parameter tuning, making it particularly effective for handling noisy data. We also show that thanks to its local nature, the method offers computational and memory efficiency advantages over Radial Basis Function interpolators, especially for large datasets.

97 MATHEMATICS AND COMPUTING↗

MetallData

MetallData is an HPC platform for interactive data science applications at HPC-scales. It provides an ecosystem for persistent distributed data structures, including algorithms, interactivity and storage.

Pearce, RogerA↗

Considerations for using Privacy Preserving Machine Learning Techniques for Safeguards

In international nuclear safeguards, the International Atomic Energy Agency (IAEA) is tasked with inspecting and verifying nuclear facilities and their activities. Data analytics and machine learning to support inspections require large amounts of data that nuclear facility operators may consider proprietary or sensitive, so the IAEA may not have full access. Allowing computation over private data without compromising its security therefore has value for safeguards inspections and analysis. Privacy-preserving machine learning (PPML) consists of security-focused techniques that allow data analytics and machine learning algorithms to run on sensitive data without revealing it. This includes ideas like homomorphic encryption (HE), secure multiparty computation (SMPC), and secure enclaves. HE allows algorithms and mathematical operations to be conducted directly on the encrypted data instead of first decrypting it. With SMPC, multiple entities collaboratively compute over distributed data such that no party is able to directly view any others’ original data. Secure enclaves allow computation to take place in a separate and heavily blocked-off section of a CPU. Techniques like these allow for several potential use cases in which the security of data is essential. With SMPC, machine learning models can be trained over the input data from multiple entities, resulting in a model that all users can benefit from without leaking the input data from any particular entity. With SMPC or a zero-knowledge proof (ZKP), an algorithm returning some single answer or truth value can be run on someone else’s data without ever needing to see that data, potentially allowing for verification or proof of some underlying question. HE can allow for outsourcing computation on data to a hostile or untrusted environment. Although most of the research in this field resides within the health and financial domains, tools from PPML may have similar applications in nuclear safeguards. Allowing the IAEA to compute over proprietary information, such as process models and raw sensor data using PPML techniques, provides the baseline for running complex analytics without needing direct unencrypted access to the underlying data, maintaining its privacy. Important limitations to consider for these techniques include the efficiency and level of security required. The security of HE and SMPC come at the cost of speed—the significant amount of overhead means that algorithms implemented in these protocols and encryption schemes are slower than when run on plaintext. Additionally, several important parameters determine what techniques or protocols are used based on the security requirements. SMPC protocols may need to be selected for resistance against a party that attempts to deviate from the protocol to distort the result or gain access to additional information, and a protocol secure against these attacks may further increase the overhead of the algorithm.

97 MATHEMATICS AND COMPUTING↗

A Probabilistic Autoencoder for Type Ia Supernova Spectral Time Series

We construct a physically parameterized probabilistic autoencoder (PAE) to learn the intrinsic diversity of Type Ia supernovae (SNe Ia) from a sparse set of spectral time series. The PAE is a two-stage generative model, composed of an autoencoder that is interpreted probabilistically after training using a normalizing flow. We demonstrate that the PAE learns a low-dimensional latent space that captures the nonlinear range of features that exists within the population and can accurately model the spectral evolution of SNe Ia across the full range of wavelength and observation times directly from the data. By introducing a correlation penalty term and multistage training setup alongside our physically parameterized network, we show that intrinsic and extrinsic modes of variability can be separated during training, removing the need for the additional models to perform magnitude standardization. We then use our PAE in a number of downstream tasks on SNe Ia for increasingly precise cosmological analyses, including the automatic detection of SN outliers, the generation of samples consistent with the data distribution, and solving the inverse problem in the presence of noisy and incomplete data to constrain cosmological distance measurements. We find that the optimal number of intrinsic model parameters appears to be three, in line with previous studies, and show that we can standardize our test sample of SNe Ia with an rms of 0.091 ± 0.010 mag, which corresponds to 0.074 ± 0.010 mag if peculiar velocity contributions are removed.

79 ASTRONOMY AND ASTROPHYSICS↗

The Rucio File Catalog in DIRAC implemented for Belle II

DIRAC and Rucio are two standard pieces of software widely used in the HEP domain. DIRAC provides Workload and Data Management function- alities, among other things, while Rucio is a dedicated, advanced Distributed Data Management system. Many communities that already use DIRAC have expressed their interest in using DIRAC for Workload Management in combi- nation with Rucio for Data Management. In this paper, we describe the integra- tion of the Rucio File Catalog into DIRAC that was initially developed for the Belle II collaboration.

97 MATHEMATICS AND COMPUTING↗

Securely Aggregated Coded Matrix Inversion

Coded computing is a method for mitigating straggling workers in a centralized computing network, by using erasure-coding techniques. Federated learning is a decentralized model for training data distributed across client devices. In this work we propose approximating the inverse of an aggregated data matrix, where the data is generated by clients; similar to the federated learning paradigm, while also being resilient to stragglers. To do so, we propose a coded computing method based on gradient coding. We modify this method so that the coordinator does not access the local data at any point; while the clients access the aggregated matrix in order to complete their tasks. Here, the network we consider is not centrally administrated, and the communications which take place are secure against potential eavesdroppers.

97 MATHEMATICS AND COMPUTING↗