Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Scalable statistical inference of photometric redshift via data subsampling

Handling big data has largely been a major bottleneck in traditional statistical models. Consequently, when accurate point prediction is the primary target, machine learning models are often preferred over their statistical counterparts for bigger problems. But full probabilistic statistical models often outperform other models in quantifying uncertainties associated with model predictions. We develop a data-driven statistical modeling framework that combines the uncertainties from an ensemble of statistical models learned on smaller subsets of data carefully chosen to account for imbalances in the input space. We demonstrate this method on a photometric redshift estimation problem in cosmology, which seeks to infer a distribution of the redshift—the stretching effect in observing the light of far-away galaxies—given multivariate color information observed for an object in the sky. Our proposed method performs balanced partitioning, graph-based data subsampling across the partitions, and training of an ensemble of Gaussian process models.

data subsampling↗

Physics-Informed Gaussian Process Inference of Liquid Structure from Scattering Data

We present a nonparametric Bayesian framework to infer radial distribution functions from experimental scattering measurements with uncertainty quantification using nonstationary Gaussian processes. The Gaussian process prior mean and kernel functions are designed to mitigate well-known numerical challenges with the Fourier transform, including discrete measurement binning and detector windowing, while encoding fundamental yet minimal physical knowledge of the liquid structure. We demonstrate uncertainty propagation of the Gaussian process posterior to unmeasured quantities of interest. Experimental radial distribution functions of liquid argon and water with uncertainty quantification are provided as both a proof of principle for the method and a benchmark for molecular models.

Chemical structure↗

Cinematic reflectometry using QIKR, the quite intense kinetics reflectometer

The Quite Intense Kinetics Reflectometer (QIKR) will be a general-purpose, horizontal-sample-surface neutron reflectometer. Reflectometers measure the proportion of an incident probe beam reflected from a surface as a function of wavevector (momentum) transfer to infer the distribution and composition of matter near an interface. The unique scattering properties of neutrons make this technique especially useful in the study of soft matter, biomaterials, and materials used in energy storage. Exploiting the increased brilliance of the Spallation Neutron Source Second Target Station, QIKR will collect specular and off-specular reflectivity data faster than the best existing such machines. It will often be possible to collect complete specular reflectivity curves using a single instrument setting, enabling “cinematic” operation, wherein the user turns on the instrument and “films” the sample. Samples in time-dependent environments (e.g., temperature, electrochemical, or undergoing chemical alteration) will be observed in real time, in favorable cases with frame rates as fast as 1 Hz. Cinematic data acquisition promises to make time-dependent measurements routine, with time resolution specified during post-experiment data analysis. This capability will be deployed to observe such processes as in situ polymer diffusion, battery electrode charge–discharge cycles, hysteresis loops, and membrane protein insertion into lipid layers.

47 OTHER INSTRUMENTATION↗

Neural network denoising of x-ray images from high-energy-density experiments

Noise is a consistent problem for x-ray transmission images of High-Energy-Density (HED) experiments because it can significantly affect the accuracy of inferring quantitative physical properties from these images. We consider experiments that use x-ray area backlighting to image a thin layer of opaque material within a physics package to observe its hydrodynamic evolution. The spatial variance of the x-ray transmission across the system due to changing opacity serves as an analog for measuring density in this evolving layer. The noise in these images adds nonphysical variations in measured intensity, which can significantly reduce the accuracy of our inferred densities, particularly at small spatial scales. Denoising these images is thus necessary to improve our quantitative analysis, but any denoising method also affects the underlying information in the image. In this paper, we present a method for denoising HED x-ray images via a deep convolutional neural network model with a modified DenseNet architecture. In our denoising framework, we estimate the noise present in the real (data) images of interest and apply the inferred noise distribution to a set of natural images. These synthetic noisy images are then used to train a neural network model to recognize and remove noise of that character. We show that our trained denoiser network significantly reduces the noise in our experimental images while retaining important physical features.

47 OTHER INSTRUMENTATION↗

Saccharomycotina yeasts defy long-standing macroecological patterns

The Saccharomycotina yeasts (“yeasts” hereafter) are a fungal clade of scientific, economic, and medical significance. Yeasts are highly ecologically diverse, found across a broad range of environments in every biome and continent on earth; however, little is known about what rules govern the macroecology of yeast species and their range limits in the wild. Here, we trained machine learning models on 12,816 terrestrial occurrence records and 96 environmental variables to infer global distribution maps at ~1 km2 resolution for 186 yeast species (~15% of described species from 75% of orders) and to test environmental drivers of yeast biogeography and macroecology. We found that predicted yeast diversity hotspots occur in mixed montane forests in temperate climates. Diversity in vegetation type and topography were some of the greatest predictors of yeast species richness, suggesting that microhabitats and environmental clines are key to yeast diversity. We further found that range limits in yeasts are significantly influenced by carbon niche breadth and range overlap with other yeast species, with carbon specialists and species in high-diversity environments exhibiting reduced geographic ranges. Finally, yeasts contravene many long-standing macroecological principles, including the latitudinal diversity gradient, temperature-dependent species richness, and a positive relationship between latitude and range size (Rapoport’s rule). These results unveil how the environment governs the global diversity and distribution of species in the yeast subphylum. These high-resolution models of yeast species distributions will facilitate the prediction of economically relevant and emerging pathogenic species under current and future climate scenarios.

59 BASIC BIOLOGICAL SCIENCES↗

Physics-inspired spatiotemporal-graph AI ensemble for the detection of higher order wave mode signals of spinning binary black hole mergers

We present a new class of AI models for the detection of quasi-circular, spinning, non-precessing binary black hole mergers whose waveforms include the higher order gravitational wave modes ($\ell$, |m|) = {(2,2), (2,1), (3,3), (3,2), (4,4)}, and mode mixing effects in the $\ell$ = 3, |m| = 2 harmonics. These AI models combine hybrid dilated convolution neural networks to accurately model both short- and long-range temporal sequential information of gravitational waves; and graph neural networks to capture spatial correlations among gravitational wave observatories to consistently describe and identify the presence of a signal in a three detector network encompassing the Advanced LIGO and Virgo detectors. We first trained these spatiotemporal-graph AI models using synthetic noise, using 1.2 million modeled waveforms to densely sample this signal manifold, within 1.7 h using 256 NVIDIA A100 GPUs in the Polaris supercomputer at the Argonne Leadership Computing Facility. This distributed training approach exhibited optimal classification performance, and strong scaling up to 512 NVIDIA A100 GPUs. With these AI ensembles we processed data from a three detector network, and found that an ensemble of 4 AI models achieves state-of-the-art performance for signal detection, and reports two misclassifications for every decade of searched data. We distributed AI inference over 128 GPUs in the Polaris supercomputer and 128 nodes in the Theta supercomputer, and completed the processing of a decade of gravitational wave data from a three detector network within 3.5 h. Finally, we fine-tuned these AI ensembles to process the entire month of February 2020, which is part of the O3b LIGO/Virgo observation run, and found 6 gravitational waves, concurrently identified in Advanced LIGO and Advanced Virgo data, and zero false positives. This analysis was completed in one hour using one NVIDIA A100 GPU.

79 ASTRONOMY AND ASTROPHYSICS↗

Gaussian processes meet NeuralODEs: a Bayesian framework for learning the dynamics of partially observed systems from scarce and noisy data

We present a machine learning framework (GP-NODE) for Bayesian model discovery from partial, noisy and irregular observations of nonlinear dynamical systems. The proposed method takes advantage of differentiable programming to propagate gradient information through ordinary differential equation solvers and perform Bayesian inference with respect to unknown model parameters using Hamiltonian Monte Carlo sampling and Gaussian Process priors over the observed system states. This allows us to exploit temporal correlations in the observed data, and efficiently infer posterior distributions over plausible models with quantified uncertainty. The use of the Finnish Horseshoe as a sparsity-promoting prior for free model parameters also enables the discovery of parsimonious representations for the latent dynamics. A series of numerical studies is presented to demonstrate the effectiveness of the proposed GP-NODE method including predator–prey systems, systems biology and a 50-dimensional human motion dynamical system. This article is part of the theme issue ‘Data-driven prediction in dynamical systems’.

Science & Technology - Other Topics↗

Dark Energy Survey Year 6 results: Redshift calibration of the MagLim++ lens sample

In this work, we derive and calibrate the redshift distribution of the MagLim++ lens galaxy sample used in the Dark Energy Survey Year 6 (DES Y6) 3 x 2pt cosmology analysis. The 3 x 2pt analysis combines galaxy clustering from the lens galaxy sample and weak gravitational lensing. The redshift distributions are inferred using the SOMPZ method - a Self-Organizing Map framework that combines deep-field multi-band photometry, wide-field data, and a synthetic source injection ( B alrog) catalog. Key improvements over the DES Year 3 (Y3) calibration include a noise-weighted SOM metric, an expanded Balrog catalogue, and an improved scheme for propagating systematic uncertainties, which allows us to generate O(10 8 ) redshift realizations that collectively span the dominant sources of uncertainty. These realizations are then combined with independent clustering-redshift measurements via importance sampling. The resulting calibration achieves typical uncertainties on the mean redshift of 1-2%, corresponding to a 20-30% average reduction relative to DES Y3. We compress the n(z) uncertainties into a small number of orthogonal modes for use in cosmological inference. Marginalizing over these modes leads to only a minor degradation in cosmological constraints. Here, this analysis establishes the MagLim++ sample as a robust lens sample for precision cosmology with DES Y6 and provides a scalable framework for future surveys.

dark energy↗

Bayesian mixture model approach to quantifying the empirical nuclear saturation point

The equation of state (EOS) in the limit of infinite symmetric nuclear matter exhibits an equilibrium density, $n_0 \approx 0.16 \, \mathrm{fm}^{-3}$, at which the pressure vanishes and the energy per particle attains its minimum, $E_0 \approx -16 \, \mathrm{MeV}$. Although not directly measurable, the nuclear saturation point $(n_0,E_0)$ can be extrapolated by density functional theory (DFT), providing tight constraints for microscopic interactions derived from chiral effective field theory (EFT). However, when considering several DFT predictions for $(n_0,E_0)$ from Skyrme and Relativistic Mean Field (RMF) models together, a discrepancy between these model classes emerges at high confidence levels that each model prediction's uncertainty cannot explain. How can we leverage these DFT constraints to rigorously benchmark nuclear saturation properties of chiral interactions? To address this question, we present a Bayesian mixture model that combines multiple DFT predictions for $(n_0,E_0)$ using an efficient conjugate prior approach. The inferred posterior distribution for the saturation point's mean and covariance matrix follows a Normal-inverse-Wishart class, resulting in posterior predictives in the form of correlated, bivariate $t$-distributions. The DFT uncertainty reports are then used to mix these posteriors using an ordinary Monte Carlo approach. At the 95\% credibility level, we estimate $n_0 \approx 0.157 \pm 0.010 \, \mathrm{fm}^{-3}$ and $E_0 \approx -15.97 \pm 0.40 \, \mathrm{MeV}$ for the marginal (univariate) $t$-distributions. Combined with chiral EFT calculations of the pure neutron matter EOS, we obtain bivariate normal distributions for the nuclear symmetry energy and its slope parameter evaluated at $n_0$: $S_v \approx 32.0 \pm 1.1 \, \mathrm{MeV}$ and $L\approx 52.6\pm 8.1 \, \mathrm{MeV}$ (95\%), respectively. Furthermore, our Bayesian framework is publicly available, so practitioners can readily use and extend our results.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Reinforcement Learning via Gaussian Processes with Neural Network Dual Kernels

While deep neural networks (DNNs) and Gaussian Processes (GPs) are both popularly utilized to solve problems in reinforcement learning, both approaches feature undesirable drawbacks for challenging problems. DNNs learn complex non-linear embeddings, but do not naturally quantify uncertainty and are often data-inefficient to train. GPs infer posterior distributions over functions, but popular kernels exhibit limited expressivity on complex and high-dimensional data. Fortunately, recently discovered conjugate and neural tangent kernel functions encode the behavior of overparameterized neural networks in the kernel domain. We demonstrate that these kernels can be efficiently applied to regression and reinforcement learning problems by analyzing a baseline case study.We apply GPs with neural network dual kernels to solve reinforcement learning tasks for the first time. We demonstrate, using the well understood mountain-car problem, that GPs empowered with dual kernels perform at least as well as those using the conventional radial basis function kernel. Finally, we conjecture that by inheriting the probabilistic rigor of GPs and the powerful embedding properties of DNNs, GPs using NN dual kernels will empower future reinforcement learning models on difficult domains.

97 MATHEMATICS AND COMPUTING↗

Resilient Distributed Field Estimation

We study resilient distributed field estimation under measurement attacks. A network of agents or devices measures a large, spatially distributed physical field parameter. An adversary arbitrarily manipulates the measurements of some of the agents. Each agent's goal is to process its measurements and information received from its neighbors to estimate only a few specific components of the field. Here, we present SAFE, the saturating adaptive field estimator, a consensus + innovations distributed field estimator that is resilient to measurement attacks. Under sufficient conditions on the compromised measurement streams, the physical coupling between the field and the agents' measurements, and the connectivity of the cyber communication network, SAFE guarantees that each agent's estimate converges almost surely to the true value of the components of the parameter in which the agent is interested. Finally, we illustrate the performance of SAFE through numerical examples.

97 MATHEMATICS AND COMPUTING↗

"Sintering" Models and In-Situ Experiments: Data Assimilation for Microstructure Prediction in SLS Additive Manufacturing of Nylon Components

Selective laser sintering methods are workhorses for additively manufacturing polymer-based components. The ease of rapid prototyping also means it is easy to produce illicit components. It is necessary to have a data-calibrated in-situ physical model of the build process in order to predict expected and defective microstructure characteristics that inform component provenance. Toward this end, sintering models are calibrated and characteristics such as component defects are explored. This is accomplished by assimilating multiple data streams, imaging analysis, and computational model predictions in an adaptive Bayesian parameter estimation algorithm. From these data sources, along with a phase-field model, bulk porosity distributions are inferred. Model parameters are constrained to physically-relevant search directions by sensitivity analysis, and then matched to predictions using adaptive sampling. Using this feedback loop, data-constrained estimates of sintering model parameters along with uncertainty bounds are obtained.

3D printing, Additive Manufacturing, polymer, sint↗

Uncovering the basis of protein-protein interaction specificity with a combinatorially complete library

Protein-protein interaction specificity is often encoded at the primary sequence level. However, the contributions of individual residues to specificity are usually poorly understood and often obscured by mutational robustness, sequence degeneracy, and epistasis. Using bacterial toxin-antitoxin systems as a model, we screened a combinatorially complete library of antitoxin variants at three key positions against two toxins. This library enabled us to measure the effect of individual substitutions on specificity in hundreds of genetic backgrounds. These distributions allow inferences about the general nature of interface residues in promoting specificity. We find that positive and negative contributions to specificity are neither inherently coupled nor mutually exclusive. Further, a wild-type antitoxin appears optimized for specificity as no substitutions improve discrimination between cognate and non-cognate partners. By comparing crystal structures of paralogous complexes, we provide a rationale for our observations. Collectively, this work provides a generalizable approach to understanding the logic of molecular recognition.

59 BASIC BIOLOGICAL SCIENCES↗

Dark Energy Survey Year 6 Results: Redshift Calibration of the MagLim++ Lens Sample

In this work, we derive and calibrate the redshift distribution of the MagLim++ lens galaxy sample used in the Dark Energy Survey Year 6 (DES Y6) 3x2pt cosmology analysis. The 3x2pt analysis combines galaxy clustering from the lens galaxy sample and weak gravitational lensing. The redshift distributions are inferred using the SOMPZ method - a Self-Organizing Map framework that combines deep-field multi-band photometry, wide-field data, and a synthetic source injection (Balrog) catalog. Key improvements over the DES Year 3 (Y3) calibration include a noise-weighted SOM metric, an expanded Balrog catalogue, and an improved scheme for propagating systematic uncertainties, which allows us to generate O($10^8$) redshift realizations that collectively span the dominant sources of uncertainty. These realizations are then combined with independent clustering-redshift measurements via importance sampling. The resulting calibration achieves typical uncertainties on the mean redshift of 1-2%, corresponding to a 20-30% average reduction relative to DES Y3. We compress the $n(z)$ uncertainties into a small number of orthogonal modes for use in cosmological inference. Marginalizing over these modes leads to only a minor degradation in cosmological constraints. This analysis establishes the MagLim++ sample as a robust lens sample for precision cosmology with DES Y6 and provides a scalable framework for future surveys.

Giannini, G. [Chicago U., Astron. Astrophys. Ctr.;↗

Data Center High-Temperature Liquid Cooling and Heat Reuse Techno-Economic Study: Preprint

Data centers are energy-intensive facilities with growing demands for efficiency and cost-effective operations. Smaller, more distributed edge inference data centers are expected to proliferate as AI applications require low latency closer to the user of AI tools, which presents a growing opportunity to explore the systems implications of liquid cooling on water and energy use. This study analyzes the implementation of high-temperature liquid cooling systems in a prototypical inference 1-MW data center and explores the potential for heat reuse across varying climates with a goal to optimize energy efficiency, reduce capital and operational costs, and identify opportunities for high-performance cooling and water use reduction infrastructure. This analysis evaluated configurations utilizing a peak day hourly sizing and systems performance spreadsheet to evaluate design and operational conditions from which component sizes, installed cost, operational cost, and performance metrics were determined for the Base case and the Elevated case. The techno-economic analysis included heat reuse applications across a range of heat recovery temperatures and heat rejection options. The analysis shows that high-temperature liquid cooling allows for improved energy efficiency, lower water consumption, and lower capital costs compared to traditional cooling approaches. Transitioning to elevated water inlet/outlet temperatures (50 degrees C/60 degrees C) eliminates the need for chillers, cooling towers, and heat recovery equipment in many scenarios across three distinct climate zones. This results in up to 75% capital cost savings for the cooling and heat recovery equipment, and with significantly reduced water consumption, especially in non-heat reuse applications. Heat generated from data centers can also be repurposed for space heating, domestic hot water, and other applications, and is most cost-effective when data center outlet temperatures exceed 55-60 degrees C.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Risk-Informed Condition Evaluation of Solar-centered Energy Generation and Distribution Networks through Bayesian Learning and Inference

We develop a methodology based on Bayesian inference over Probabilistic Graphical Models (PGMs) to understand and quantify risk in solar-centered grids using targeted measurements and learned system behavior. Being non-prescriptive but, rather, able to infer system behavior and, ultimately, address risk queries from data, our machine learning-type paradigm is tailored for diverse topologies and threat scenarios often associated with distributed energy generation and photovoltaic distributed energy resources (PV-DERs) in particular. We describe algorithmic processes for: (i) learning the structure of PGMs that result from attack-prone PV-DER-proliferated distribution systems, (ii) quantifying cause-effect relationships, and (iii) evaluating risk queries based on diverse evidence. The contributions are illustrated on a residential grid subject to output impairment attacks on its PV-DER infrastructure.

24 POWER TRANSMISSION AND DISTRIBUTION↗