Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Improving Enzyme Optimum Temperature Prediction with Resampling Strategies and Ensemble Learning

Accurate prediction of the optimal catalytic temperature ( T opt ) of enzymes is vital in biotechnology, as enzymes with high T opt values are desired for enhanced reaction rates. Recently, a machine learning method (temperature optima for microorganisms and enzymes, TOME) for predicting T opt was developed. TOME was trained on a normally distributed data set with a median T opt of 37 °C and less than 5% of T opt values above 85 °C, limiting the method’s predictive capabilities for thermostable enzymes. Due to the distribution of the training data, the mean squared error on T opt values greater than 85 °C is nearly an order of magnitude higher than the error on values between 30 and 50 °C. Here, we apply ensemble learning and resampling strategies that tackle the data imbalance to significantly decrease the error on high T opt values (>85 °C) by 60% and increase the overall R 2 value from 0.527 to 0.632. The revised method, temperature optima for enzymes with resampling (TOMER), and the resampling strategies applied in this work are freely available to other researchers as Python packages on GitHub.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Distribution Transformer Health Monitoring using Smart Meter Data

The distribution electric grid has become a highly complex and intelligent network with changing load and customer types. This has generated unprecedented challenges and opportunities for utility companies–opportunities especially in the area of asset health/performance management. Moreover, several utilities are increasingly moving from the traditional reactive and time-based asset monitoring approach to a more proactive condition based method. However, this needs to be done in a low-cost and efficient manner. Here, this paper explores how existing sensor infrastructure such as smart meters can be utilized to provide utility operators with more visibility into the health and operation of their assets. The paper focuses primarily on service transformers.

24 POWER TRANSMISSION AND DISTRIBUTION↗

HydraGNN v3.0

New or improved capabilities included in v3.0 release are as follows: 1. Enhancement in message passing layers through generalization of the class inheritance to enable the inclusion of a broader set of message passing policies Inclusion of equivariant message passing layers from the original implementations of: SchNet (https://pubs.aip.org/aip/jcp/article/148/24/241722/962591/SchNet-A-deep-learning-architecture-for-molecules); DimeNet++ (https://arxiv.org/abs/2011.14115); EGNN models (https://arxiv.org/pdf/2102.09844.pdf) 2. Restructuring of class inheritance for data management 3. Support of DDStore https://github.com/ORNL/DDStore capabilities for improved distributed data parallelism on large volumes of data that cannot fit on intra-node memory capacities 4. Large-scale system support for OLCF-Crusher and OLCF-Frontier

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

A kinetic-based regularization method for data science applications

We propose a physics-based regularization technique for function learning, inspired by statistical mechanics. By drawing an analogy between optimizing the parameters of an interpolator and minimizing the energy of a system, we introduce corrections that impose constraints on the lower-order moments of the data distribution. This minimizes the discrepancy between the discrete and continuum representations of the data, in turn allowing to access more favorable energy landscapes, thus improving the accuracy of the interpolator. Our approach improves performance in both interpolation and regression tasks, even in high-dimensional spaces. Unlike traditional methods, it does not require empirical parameter tuning, making it particularly effective for handling noisy data. We also show that thanks to its local nature, the method offers computational and memory efficiency advantages over Radial Basis Function interpolators, especially for large datasets.

97 MATHEMATICS AND COMPUTING↗

MetallData

MetallData is an HPC platform for interactive data science applications at HPC-scales. It provides an ecosystem for persistent distributed data structures, including algorithms, interactivity and storage.

Pearce, RogerA↗

Considerations for using Privacy Preserving Machine Learning Techniques for Safeguards

In international nuclear safeguards, the International Atomic Energy Agency (IAEA) is tasked with inspecting and verifying nuclear facilities and their activities. Data analytics and machine learning to support inspections require large amounts of data that nuclear facility operators may consider proprietary or sensitive, so the IAEA may not have full access. Allowing computation over private data without compromising its security therefore has value for safeguards inspections and analysis. Privacy-preserving machine learning (PPML) consists of security-focused techniques that allow data analytics and machine learning algorithms to run on sensitive data without revealing it. This includes ideas like homomorphic encryption (HE), secure multiparty computation (SMPC), and secure enclaves. HE allows algorithms and mathematical operations to be conducted directly on the encrypted data instead of first decrypting it. With SMPC, multiple entities collaboratively compute over distributed data such that no party is able to directly view any others’ original data. Secure enclaves allow computation to take place in a separate and heavily blocked-off section of a CPU. Techniques like these allow for several potential use cases in which the security of data is essential. With SMPC, machine learning models can be trained over the input data from multiple entities, resulting in a model that all users can benefit from without leaking the input data from any particular entity. With SMPC or a zero-knowledge proof (ZKP), an algorithm returning some single answer or truth value can be run on someone else’s data without ever needing to see that data, potentially allowing for verification or proof of some underlying question. HE can allow for outsourcing computation on data to a hostile or untrusted environment. Although most of the research in this field resides within the health and financial domains, tools from PPML may have similar applications in nuclear safeguards. Allowing the IAEA to compute over proprietary information, such as process models and raw sensor data using PPML techniques, provides the baseline for running complex analytics without needing direct unencrypted access to the underlying data, maintaining its privacy. Important limitations to consider for these techniques include the efficiency and level of security required. The security of HE and SMPC come at the cost of speed—the significant amount of overhead means that algorithms implemented in these protocols and encryption schemes are slower than when run on plaintext. Additionally, several important parameters determine what techniques or protocols are used based on the security requirements. SMPC protocols may need to be selected for resistance against a party that attempts to deviate from the protocol to distort the result or gain access to additional information, and a protocol secure against these attacks may further increase the overhead of the algorithm.

97 MATHEMATICS AND COMPUTING↗

A Probabilistic Autoencoder for Type Ia Supernova Spectral Time Series

We construct a physically parameterized probabilistic autoencoder (PAE) to learn the intrinsic diversity of Type Ia supernovae (SNe Ia) from a sparse set of spectral time series. The PAE is a two-stage generative model, composed of an autoencoder that is interpreted probabilistically after training using a normalizing flow. We demonstrate that the PAE learns a low-dimensional latent space that captures the nonlinear range of features that exists within the population and can accurately model the spectral evolution of SNe Ia across the full range of wavelength and observation times directly from the data. By introducing a correlation penalty term and multistage training setup alongside our physically parameterized network, we show that intrinsic and extrinsic modes of variability can be separated during training, removing the need for the additional models to perform magnitude standardization. We then use our PAE in a number of downstream tasks on SNe Ia for increasingly precise cosmological analyses, including the automatic detection of SN outliers, the generation of samples consistent with the data distribution, and solving the inverse problem in the presence of noisy and incomplete data to constrain cosmological distance measurements. We find that the optimal number of intrinsic model parameters appears to be three, in line with previous studies, and show that we can standardize our test sample of SNe Ia with an rms of 0.091 ± 0.010 mag, which corresponds to 0.074 ± 0.010 mag if peculiar velocity contributions are removed.

79 ASTRONOMY AND ASTROPHYSICS↗

The Rucio File Catalog in DIRAC implemented for Belle II

DIRAC and Rucio are two standard pieces of software widely used in the HEP domain. DIRAC provides Workload and Data Management function- alities, among other things, while Rucio is a dedicated, advanced Distributed Data Management system. Many communities that already use DIRAC have expressed their interest in using DIRAC for Workload Management in combi- nation with Rucio for Data Management. In this paper, we describe the integra- tion of the Rucio File Catalog into DIRAC that was initially developed for the Belle II collaboration.

97 MATHEMATICS AND COMPUTING↗

Securely Aggregated Coded Matrix Inversion

Coded computing is a method for mitigating straggling workers in a centralized computing network, by using erasure-coding techniques. Federated learning is a decentralized model for training data distributed across client devices. In this work we propose approximating the inverse of an aggregated data matrix, where the data is generated by clients; similar to the federated learning paradigm, while also being resilient to stragglers. To do so, we propose a coded computing method based on gradient coding. We modify this method so that the coordinator does not access the local data at any point; while the clients access the aggregated matrix in order to complete their tasks. Here, the network we consider is not centrally administrated, and the communications which take place are secure against potential eavesdroppers.

97 MATHEMATICS AND COMPUTING↗

Aerosol responses to precipitation along North American air trajectories arriving at Bermuda

North American pollution outflow is ubiquitous over the western North Atlantic Ocean, especially in winter, making this location a suitable natural laboratory for investigating the impact of precipitation on aerosol particles along air mass trajectories. We take advantage of observational data collected at Bermuda to seasonally assess the sensitivity of aerosol mass concentrations and volume size distributions to accumulated precipitation along trajectories (APT). The mass concentration of particulate matter with aerodynamic diameter less than 2.5 µm normalized by the enhancement of carbon monoxide above background (PM 2.5 /ΔCO) at Bermuda was used to estimate the degree of aerosol loss during transport to Bermuda. Results for December–February (DJF) show that most trajectories come from North America and have the highest APTs, resulting in a significant reduction (by 53 %) in PM 2.5 /ΔCO under high-APT conditions (> 13.5 mm) relative to low-APT conditions (< 0.9 mm). Moreover, PM 2.5 /ΔCO was most sensitive to increases in APT up to 5 mm (–0.044 µg m –3 ppbv –1 mm –1 ) and less sensitive to increases in APT over 5 mm. While anthropogenic PM 2.5 constituents (e.g., black carbon, sulfate, organic carbon) decrease with high APT, sea salt, in contrast, was comparable between high- and low-APT conditions owing to enhanced local wind and sea salt emissions in high-APT conditions. The greater sensitivity of the fine-mode volume concentrations (versus coarse mode) to wet scavenging is evident from AErosol RObotic NETwork (AERONET) volume size distribution data. A combination of GEOS-Chem model simulations of the 210 Pb submicron aerosol tracer and its gaseous precursor 222 Rn reveals that (i) surface aerosol particles at Bermuda are most impacted by wet scavenging in winter and spring (due to large-scale precipitation) with a maximum in March, whereas convective scavenging plays a substantial role in summer; and (ii) North American 222 Rn tracer emissions contribute most to surface 210 Pb concentrations at Bermuda in winter (~75 %–80 %), indicating that air masses arriving at Bermuda experience large-scale precipitation scavenging while traveling from North America. A case study flight from the ACTIVATE field campaign on 22 February 2020 reveals a significant reduction in aerosol number and volume concentrations during air mass transport off the US East Coast associated with increased cloud fraction and precipitation. These results highlight the sensitivity of remote marine boundary layer aerosol characteristics to precipitation along trajectories, especially when the air mass source is continental outflow from polluted regions like the US East Coast.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Density estimation via measure transport: Outlook for applications in the biological sciences

Abstract One among several advantages of measure transport methods is that they allow or a unified framework for processing and analysis of data distributed according to a wide class of probability measures. Within this context, we present results from computational studies aimed at assessing the potential of measure transport techniques, specifically, the use of triangular transport maps, as part of a workflow intended to support research in the biological sciences. Scenarios characterized by the availability of limited amount of sample data, which are common in domains such as radiation biology, are of particular interest. We find that when estimating a distribution density function given limited amount of sample data, adaptive transport maps are advantageous. In particular, statistics gathered from computing series of adaptive transport maps, trained on a series of randomly chosen subsets of the set of available data samples, leads to uncovering information hidden in the data. As a result, in the radiation biology application considered here, this approach provides a tool for generating hypotheses about gene relationships and their dynamics under radiation exposure.

gene expression data↗

Machine Learning for Distributed Acoustic Sensing data (MLDAS) v1.0.1

MLDAS is a Python-written package for exploratory data analysis and deep learning training on Distributed Acoustic Sensing data. The machine learning tools are powered by the PyTorch library and designed to work efficiently on large scale datasets using parallel computing. Various SLURM scripts as well as a tutorial have also been made available to allow geophysicists to quickly and easily implement the available tools in their analysis workflow on supercomputer facilities.

Dumont, Vincent↗

Updates and Validation for the n+ 63,65 Cu Cross Sections [Abstract]

The neutron induced total, elastic, and capture cross sections of 63,65 Cu isotopes were selected for evaluation in the resolved and unresolved resonance energy ranges by the National Criticality Safety Program to resolve discrepancies related to benchmark performance. This is especially evident for the series of ZEUS benchmarks in which copper is used as a reflector. Because copper is also used as structural material in both fission and fusion reactors, the need to address benchmark discrepancies linked to nuclear data deficiencies is a task of primary importance. The aim of this work is to describe the steps of evaluation work towards a consistent improvement of the benchmark performance. The R-matrix analysis with the SAMMY code focused on the 63 Cu(n,γ) reaction channel between 100-300 keV coupled to unresolved resonance region parameters up to 650 keV to fit average cross section data from a recent experiment. Due to the high sensitivity of many benchmarks to elastic scattering angular distribution data, especially for the 65 Cu isotope, the impact of these data was tested by generating Legendre coefficients from both resonance parameters and the Hauser-Feshbach model. Guided by the findings of Shaw et al., the performance of the current evaluation for 65 Cu was compared to that of ENDF/B-VII.1 and ENDF/B-VIII.0 by testing the reactivity coefficients corresponding to the validation suite of experimental criticality benchmarks for thermal, intermediate, and fast systems taken from the International Criticality Safety Benchmark Experiments Project Handbook. The benchmark performance is especially sensitive to 63 Cu(n,γ) and 65 Cu elastic scattering for neutron energies in the 100–500 keV region, whereas 100 keV is the upper limit of the resolved resonance region in the ENDF/B-VIII.0 evaluations for 63,65 Cu. The results highlight the need to handle the transition from the resolved resonance region to the high energy region carefully.

07 ISOTOPE AND RADIATION SOURCES↗

Regularizing INR with Diffusion Prior for Self-Supervised 3D Reconstruction OF Neutron Computed Tomography Data

Recently, generative diffusion priors have made huge strides as inverse problem solvers, including the ability to be adapted for inference on out-of-distribution data. Concurrently, implicit neural representations (INRs) have emerged as fast and lightweight inverse imaging solvers that are amenable to hybrid approaches that combine learned priors with traditional inverse problem formulations. In this paper, we present a diffusive computed tomography (CT) inversion framework for regularizing INRs called Diffusive INR (DINR), designed to enable high-quality reconstruction from sparse-view neutron CT. Pretrained purely on synthetic data, DINR is evaluated on simulated and experimentally obtained observations of concrete microstructures, where traditional reconstruction methods suffer substantial degradation when the number of views is reduced. Our approach delivers superior performance, reduces reconstruction artifacts, and achieves gains in PSNR and SSIM, enabling accurate micro-structural characterization even under extreme data limitations compared to state-of-the-art sparse-view reconstruction techniques.

Hossain, Maliha [ORNL]↗

Uncertainty Quantification of Global Net Methane Emissions From Terrestrial Ecosystems Using a Mechanistically Based Biogeochemistry Model

Quantification of methane (CH4) emissions from wetlands and its sinks from uplands is still fraught with large uncertainties. Here, a methane biogeochemistry model was revised, parameterized, and verified for various wetland ecosystems across the globe. The model was then extrapolated to the global scale to quantify the uncertainty induced from four different types of uncertainty sources including parameterization, wetland type distribution, wetland area distribution, and meteorological input. We found that global wetland emissions are 212 ± 62 and 212 ± 32 Tg CH4 year -1 (1Tg = 10 12 g) due to uncertain parameters and wetland type distribution, respectively, during 2000–2012. Using two wetland distribution data sets and three sets of climate data, the model simulations indicated that the global wetland emissions range from 186 to 212 CH 4 year -1 for the same period. The parameters were the most significant uncertainty source. After combining the global methane consumption in the range of -34 to -46 Tg CH 4 year -1 , we estimated that the global net land methane emissions are 149–176 Tg CH 4 year -1 due to uncertain wetland distribution and meteorological input. Spatially, the northeast United States and Amazon were two hotspots of methane emission, while consumption hotspots were in the Eastern United States and eastern China. During 1950–2016, both wetland emissions and upland consumption increased during El Niño events and decreased during La Niña events. This study highlights the need for more in situ methane flux data, more accurate wetland type, and area distribution information to better constrain the model uncertainty.

58 GEOSCIENCES↗

Exploring DAOS as a Burst Buffer for a 100 Gbps DAQ Real-Time Streaming System

We present an experimental evaluation of a burst buffer for a real-time DAQ streaming system designed to transmit instrument data to remote data centers. The system is based on EJ-FAT, a load balancing system capable of Nx 100Gbps streams, distributing data from event sources to processing nodes. We explore applying the DAOS system as a burst buffer to serve a number of purposes: improve resiliency, elasticity and add new functions into the processing pipeline. In the evaluation a sender transmits events over a 100Gbps network to a receiver integrated with DAOS to store the reassembled events using DAOS APIs. We evaluate the system for possible bottlenecks and provide end-to-end evaluation with a burst buffer using DAOS storage abstractions. We show that a receiver node can support 38.1 Gbps. This proves the viability of our approach and allows us to extend this work to investigate scale-out properties and new streaming optimizations.

Mei, Xinxin↗

Applications of Federated Learning in Semiconductor Manufacturing [Poster]

As semiconductor manufacturers explore advanced data analytics and modeling techniques and data hungry machine learning models increase in popularity due to their accuracy in solving generalized problems and ability to learn complex relationships, federated learning emerges as a privacy preserving machine learning technique for preserving data privacy and ensuring intellectual property protection. Federated Learning is a machine learning technique focused on training models using distributed data that never needs to be centrally stored, allowing the use of advanced machine learning techniques without compromising data privacy, and in the semiconductor manufacturing industry advanced machine learning techniques can reduce cost and time, but maintaining data privacy is essential to maintaining a competitive advantage. This paper systematically reviews existing literature on applications of federated learning in the semiconductor manufacturing industry with a focus on identifying common themes, algorithms, and gaps within the literature to drive future research directions. The findings reveal five key themes, including improvements in quality assurance, virtual models, privacy preservation, reliable data practices, and emerging trends and developments. By identifying key themes in literature on federated learning and semiconductor manufacturing and analyzing gaps and discussed methodologies, this study highlights several potential future research directions to expand the application of federated learning techniques in the semiconductor manufacturing domain.

42 ENGINEERING↗