Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “space computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A velocity space hybridization-based Boltzmann equation solver

In the present research, a new method for simulation of rarefied gas flows is proposed, a velocity-space hybrid of both a DSMC representation of particles and a discrete velocity quasi-particle representation of the distribution function. The hybridization scheme is discussed in detail, and is numerically verified for two test-cases: the BKW relaxation problem and a stationary Maxwellian distribution. It is demonstrated that such a velocity-space hybridization can provide computational benefits when compared to a pure discrete velocity method or pure DSMC approach, while retaining some of the more attractive properties of discrete velocity methods. Further possible improvements to the velocity-space hybrid approach are discussed.

97 MATHEMATICS AND COMPUTING↗

Differentiable Multiphysics Codes: A Breakthrough Technology for Simulation and Computing

This document summarizes the findings of a strategic planning exercise commissioned by the Weapons Simulation and Computing, Computational Physics (WSC/CP) program at the Lawrence Livermore National Laboratory (LLNL) in FY24. During the year, the committee met with multiple stakeholder communities to gather input, opinions, suggestions and concerns which have been incorporated throughout this document. The key findings from this exercise are summarized: • The development of multiphysics modelling and simulation (mod/sim) codes and software technologies, their deployment on exascale compute platforms, and their broad adoption across the NNSA is a major success of the Advanced Simulation and Computing (ASC) program and the Exascale Computing Project (ECP). Sustained investment in these core technologies is essential. • Today’s state of the art involves running ensembles of O(100K) simulations to perform uncertainty quantification (UQ) and design studies using multiple statistical methods such as Bayesian optimization to understand sensitivities of our models and explore parameterized design spaces. Even with exascale computing, we are practically limited to O(10) parameters in these studies since the number of simulations required to sample the space scales exponentially with the number of design parameters. • The data from these simulation ensembles is increasingly being used to train machine learned (ML) surrogates (or reduced order models, ROMs) which can then be used for optimization or real time design exploration. However, the trained surrogates are still limited in the number of parameters they can represent due to the sampling limitations previously noted. • Augmenting our suite of integrated multiphysics simulation codes, both current and emerging, with the ability to compute gradients (solution derivatives) of arbitrary simulation outputs with respect to (some or all) simulation inputs would be a breakthrough technology, opening the door to a new era of efficient and automated inverse design based on verified and validated mod/sim capabilities. • This capability, which we refer to as differentiable multiphysics codes (DMCs), would revolutionize both UQ and optimization studies by breaking the curse of dimensionality that presently limits our “gradient-free” ensemble based computing approach. A similar breakthrough occurred in the AI/ML community once the ability to compute gradients of arbitrary loss functions using back-propagation became commonplace. Gradient information from the multiphysics codes can also be used to dramatically improve the efficiency and scale of training of ML/ROM surrogates for rapid assessments. • Achieving this in our suite of codes will be a grand challenge, similar to the amount of effort that was required to transition from CPU to GPU computing. It will require buy-in from the entire WSC/CP program and beyond, including all integrated codes, physics and engineering models, third-party library dependencies and performance portability abstractions. It will also require investment in research and development of numerical methods for computing adjoints of coupled physics across multiple adaptively refined moving meshes and of stochastic (Monte Carlo) and mesh free (SPH) methods. • New software and numerical techniques, largely pioneered by the AI/ML community, make this feasible. Chief among these is automatic differentiation (AD), the ability to employ AD at point-wise locations in a physics calculation (instead of traditional black-box approaches) and the ability to perform “back-propagation in time” (or reverse mode AD) for non-linear partial differential equations (PDEs). Fundamentally, the conclusion of this strategic planning exercise is that the time is right to undertake a large scale effort in WSC, centered on the existing integrated codes, to continue the natural evolution of mod/sim in the age of AI/ML. Instead of attempting to replace mod/sim with purely data driven AI/ML models, we believe the key to success is to integrate AI/ML by building on top of the decades of hard-won knowledge and the verified/validated multiphysics modelling capability that is the hallmark of the ASC program.

97 MATHEMATICS AND COMPUTING↗

An accurate and efficient fragmentation approach via the generalized many-body expansion for density matrices

With relevant chemical space growing larger and larger by the day, the ability to extend computational tractability over that larger space is of paramount importance in virtually all fields of science. The solution we aim to provide here for this issue is in the form of the generalized many-body expansion for building density matrices (GMBE-DM) based on the set-theoretical derivation with overlapping fragments, through which the energy can be obtained by a single Fock build. In combination with the purification scheme and the truncation at the one-body level, the DM-based GMBE(1)-DM-P approach shows both highly accurate absolute and relative energies for medium-to-large size water clusters with about an order of magnitude better than the corresponding energy-based GMBE(1) scheme. Simultaneously, GMBE(1)-DM-P is about an order of magnitude faster than the previously proposed MBE-DM scheme [F. Ballesteros and K. U. Lao, J. Chem. Theory Comput. 18, 179 (2022)] and is even faster than a supersystem calculation without significant parallelization to rescue the fragmentation method. For even more challenging systems including ion–water and ion–pair clusters, GMBE(1)-DM-P also performs about 3 and 30 times better than the energy-based GMBE(1) approach, respectively. In addition, this work provides the first overlapping fragmentation algorithm with a robust and effective binning scheme implemented internally in a popular quantum chemistry software package. Thus, GMBE(1)-DM-P opens a new door to accurately and efficiently describe noncovalent clusters using quantum mechanics.

Chemistry↗

Origin of the magnetic couplings for the weak ferromagnet Li + [TCNE] •- (TCNE = Tetracyanoethylene)

Based upon the 16-K crystal structure, the nearest-neighbor spin couplings (J; H = –2JS a •S b ) for the weak ferromagnet (=canted antiferromagnet) Li + [TCNE] •- (TCNE = tetracyanoethylene) (T c = 21 K), which possesses two interpenetrating diamondoid sublattices with two layers of parallel [TCNE] •- s canted with respect to each other by ~60°, are computed at the B3LYP/aug-cc-pV-TZ-level. Li[TCNE] has computed interlattice, through-space interactions that exceed the intralattice, through-bonding to Li + interactions, that are antiferromagnetic. The interlattice interactions dominate the spin coupling leading to the observed weak ferromagnetic behavior. Finally, the computed J's based upon the 50-K structure have the strongest ferromagnetic intralayer interlattice interactions increase by 2.9 ± 0.6% while the antiferromagnetic interlayer interlattice interaction increases by 23% with respect to the 16K results and are in accord with the observed initial increase and subsequent decrease in the temperature dependence of the canting angle that is a consequence of the these changing opposed interactions arising from the change in structure as a function of temperature.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Generative diffusion model surrogates for mechanistic agent-based biological models

Mechanistic, multicellular, agent-based models are commonly used to investigate tissue, organ, and organism-scale biology at single-cell resolution. The Cellular-Potts Model (CPM) is a powerful and popular framework for developing and interrogating these models. CPMs become computationally expensive at large space- and time- scales making application and investigation of developed models difficult. Surrogate models may allow for the accelerated evaluation of CPMs of complex biological systems. However, the stochastic nature of these models means each set of parameters may give rise to different model configurations, complicating surrogate model development. In this work, we leverage denoising diffusion probabilistic models (DDPMs) to train a generative AI surrogate of a CPM used to investigate in vitro vasculogenesis. We describe the use of an image classifier to learn the characteristics that define unique areas of a 2-dimensional parameter space. We then apply this classifier to aid in surrogate model selection and verification. Our CPM model surrogate generates model configurations 20,000 timesteps ahead of a reference configuration and demonstrates approximately a 22x reduction in computational time as compared to native code execution. Our work represents a step towards the implementation of DDPMs to develop digital twins of stochastic biological systems.

97 MATHEMATICS AND COMPUTING↗

DMTN-118: Review of Timeseries Features

Rubin Observatory will compute timeseries variability features on lightcurves to aid users in identifying objects of interest, both during realtime Alert Production as well as in the annual Data Releases. The Data Products Definition Document (DPDD; LSE-163) allocates space for pre-computed timeseries features, and a sample set is baselined in LDM-151. However, in the subsequent decade the scientfic community has made a great deal of further progress in this area. This technote reviews the relevant literature, grouping related features where possible; discusses potential concerns and open questions; and proposes a new baseline feature set.

79 ASTRONOMY AND ASTROPHYSICS↗

Form factors and spectral densities from Lightcone Conformal Truncation

We use the method of Lightcone Conformal Truncation (LCT) to obtain form factors and spectral densities of local operators $\mathcal{O}$ in $\phi^4$ theory in two dimensions. We show how to use the Hamiltonian eigenstates from LCT to obtain form factors that are matrix elements of a local operator $\mathcal{O}$ between single-particle bra and ket states, and we develop methods that significantly reduce errors resulting from the finite truncation of the Hilbert space. We extrapolate these form factors as a function of momentum to the regime where, by crossing symmetry, they are form factors of $\mathcal{O}$ between the vacuum and a two-particle asymptotic scattering state. We also compute the momentum-space time-ordered two-point functions of local operators in LCT. These converge quickly at momenta away from branch cuts, allowing us to indirectly obtain the time-ordered correlator and the spectral density at the branch cuts. We focus on the case where the local operator $\mathcal{O}$ is the trace Θ of the stress tensor.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Performance Characteristics of the BlueField-2 SmartNIC

High-performance computing (HPC) researchers have long envisioned scenarios where application workflows could be improved through the use of programmable processing elements embedded in the network fabric. Recently, vendors have introduced programmable Smart Network Interface Cards (SmartNICs) that enable computations to be offloaded to the edge of the network. There is great interest in both the HPC and high-performance data analytics (HPDA) communities in understanding the roles these devices may play in the data paths of upcoming systems. This paper focuses on characterizing both the networking and computing aspects of NVIDIA’s new BlueField-2 SmartNIC when used in a 100Gb/s Ethernet environment. For the networking evaluation we conducted multiple transfer experiments between processors located at the host, the SmartNIC, and a remote host. These tests illuminate how much effort is required to saturate the network and help estimate the processing headroom available on the SmartNIC during transfers. For the computing evaluation we used the stress-ng benchmark to compare the BlueField-2 to other servers and place realistic bounds on the types of offload operations that are appropriate for the hardware. Our findings from this work indicate that while the BlueField-2 provides a flexible means of processing data at the network’s edge, great care must be taken to not overwhelm the hardware. While the host can easily saturate the network link, the SmartNIC’s embedded processors may not have enough computing resources to sustain more than half the expected bandwidth when using kernel-space packet processing. From a computational perspective, encryption operations, memory operations under contention, and on-card IPC operations on the SmartNIC perform significantly better than the general-purpose servers used for comparisons in our experiments. Therefore, applications that mainly focus on these operations may be good candidates for offloading to the SmartNIC.

97 MATHEMATICS AND COMPUTING↗

Reconfigurable Framework for Resilient Semantic Segmentation for Space Applications

Deep learning (DL) presents new opportunities for enabling spacecraft autonomy, onboard analysis, and intelligent applications for space missions. However, DL applications are computationally intensive and often infeasible to deploy on radiation-hardened (rad-hard) processors, which traditionally harness a fraction of the computational capability of their commercial-off-the-shelf counterparts. Commercial FPGAs and system-on-chips present numerous architectural advantages and provide the computation capabilities to enable onboard DL applications; however, these devices are highly susceptible to radiation-induced single-event effects (SEEs) that can degrade the dependability of DL applications. In this article, we propose Reconfigurable ConvNet (RECON), a reconfigurable acceleration framework for dependable, high-performance semantic segmentation for space applications. In RECON, we propose both selective and adaptive approaches to enable efficient SEE mitigation. In our selective approach, control-flow parts are selectively protected by triple-modular redundancy to minimize SEE-induced hangs, and in our adaptive approach, partial reconfiguration is used to adapt the mitigation of dataflow parts in response to a dynamic radiation environment. Combined, both approaches enable RECON to maximize system performability subject to mission availability constraints. We perform fault injection and neutron irradiation to observe the susceptibility of RECON and use dependability modeling to evaluate RECON in various orbital case studies to demonstrate a 1.5–3.0× performability improvement in both performance and energy efficiency compared to static approaches.

97 MATHEMATICS AND COMPUTING↗

Ergodic and nonergodic many-body dynamics in strongly nonlinear lattices

The study of nonlinear oscillator chains in classical many-body dynamics has a storied history going back to the seminal work of Fermi et al. [Los Alamos Scientific Laboratory Report No. LA-1940, 1955 (unpublished)]. Here, we introduce a family of such systems which consist of chains of N harmonically coupled particles with the nonlinearity introduced by confining the motion of each individual particle to a box or stadium with hard walls. The stadia are arranged on a one-dimensional lattice but they individually do not have to be one dimensional, thus permitting the introduction of chaos already at the lattice scale. For the most part we study the case where the motion is entirely one dimensional. We find that the system exhibits a mixed phase space for any finite value of N . Computations of Lyapunov spectra at randomly picked phase space locations and a direct comparison between Hamiltonian evolution and phase space averages indicate that the regular regions of phase space are not significant at large system sizes. While the continuum limit of our model is itself a singular limit of the integrable sinh Gordon theory, we do not see any evidence for the kind of nonergodicity famously seen in the work of Fermi et al. Finally, we examine the chain with particles confined to two-dimensional stadia where the individual stadium is already chaotic and find a much more chaotic phase space at small system sizes.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Homomorphic data compression for real time photon correlation analysis

The construction of highly coherent X-ray sources, combined with next-generation detectors that are larger and faster, has enabled new research opportunities across the scientific landscape. Among the techniques that benefit most from these advancements is X-ray photon correlation spectroscopy (XPCS), where faster acquisition unlocks the ability to study faster dynamics within samples. However, faster acquisition on larger detectors also introduces unprecedented challenges for online data processing and offline data storage. Such challenges are particularly prominent for XPCS, where real time analyses require simultaneous calculation of all the previously acquired data in the time series. We present a homomorphic compression scheme to effectively reduce the computational time and memory space required for XPCS analysis. Leveraging similarities in the mathematical expression between a matrix-based compression algorithm and the correlation calculation, our approach allows direct operation on the compressed data without their decompression. The offline compression scheme extends storage capacity by a factor of 40 while preserving key features in the lossy compressed data. Meanwhile, the online compression scheme reduces the computational time to below 1 ms, enabling real time calculation of the correlation functions at kHz framerate. Our demonstration of a homomorphic compression of scientific data provides an effective solution to the big data challenge at coherent light sources. Beyond the example shown in this work, the framework can be extended to facilitate real-time operations directly on a compressed data stream for other techniques.

36 MATERIALS SCIENCE↗

Multiscale Neural Networks for Approximating Green’s Functions

Neural networks (NNs) have been widely used to solve partial differential equations (PDEs) in the applications of physics, biology, and engineering. One effective approach for solving PDEs with a fixed differential operator is learning Green’s functions. However, Green’s functions are notoriously difficult to learn due to their poor regularity, which typically requires larger NNs and longer training times. In this work, we address these challenges by leveraging multiscale NNs to learn Green’s functions. Through theoretical analysis using multiscale Barron space methods and experimental validation, we show that the multiscale approach significantly reduces the necessary NN size and accelerates training.

97 MATHEMATICS AND COMPUTING↗

INL Senior Project

What did my team set out to accomplish:? Can I put a custom Machine Learning Model on FPGA?? Can I analyze network traffic in real time?? Can a QSFP port be used with an FPGA?? Does a visual representation of the latent space enhance our understanding of network traffic?? What is QSFP QSFP (Quad Small Form-Factor Pluggable)? QSFP supports transfer speeds generally up to 100Gb/s? Runs 4 parallel lines running up to 28 Gb/s? Why the latent space is important to our project ?Latent space is the compressed mapping of data points in a non-linear fashion? Create an understanding of the relationship of data collected? Can represent that relationship of a single network packet in 3 points (X, Y, Z) Project Outline FPGA? Custom Xilinx Petalinux Image for the operating system? Python program to collect packets and run them through the DPU (Data Processing Unit)? The program then sends the information over a socket to a computer? Display Program? Python Program that collects the information sent from the FPGA and display it in a graph

99 GENERAL AND MISCELLANEOUS↗

Organic carbon enables the biotic engineering of beneficial soil structure in Profundihumic and Haplic Ferralsols

We investigated how organic matter may, directly and indirectly, modify the porosity of Ferralsols, that is, deeply weathered soils of the tropics and subtropics. Although empirical and anecdotal evidence suggests that organic matter accumulation may increase porosity, a mechanistic understanding of the processes underlying this beneficial effect is lacking, especially so for Ferralsols. To achieve our end, we leveraged the fact that the Profundihumic qualifier of Ferralsols (PF) is distinguished from Haplic Ferralsols (HF) by both a much larger average carbon content in the first 1 m of soil depth (19 kg C m -3 in PF vs. 10 kg C m -3 in HF) and a significantly lower bulk density (1.05 ± 0.08 kg L -1 in PF vs. 1.21 ± 0.05 kg L -1 in HF). Through exhaustive modelling of carbon – bulk density relationships, we demonstrate that the lower bulk density of PF cannot be satisfactorily explained by a simple dilution effect. Rather, we found that bulk density correlated with carbon content when combined with carbon: nitrogen ratio (r 2 = 0.51), black carbon content (r 2 = 0.75), and Δ 14 C (r 2 =0.81). Total pore space was greater in PF (61± 3%) than in HF (55 ± 2%), but x-ray computed tomography revealed that pore space inside soil aggregates of 4–5 mm diameter does not vary between the studied Ferralsols. We further observed nearly twice as many roots and burrows in PF compared with HF. We thus infer that the mechanism responsible for the increase in porosity is most likely an enhancement of resource availability (e.g., energy, carbon, and nutrients) for the organisms (earthworms, ants, termites, etc.) that physically displace soil particles and promote soil aggregation. As a result of increased resource availability, soil organisms can create especially the mesoscale structural soil features necessary for unrestricted water flow and rapid gas exchange. In conclusion, this insight paves the way for the development of land management technologies to optimize the physical shape and capacity of the soil bioreactor.

54 ENVIRONMENTAL SCIENCES↗

Two-loop master integrals for leading-color $$ pp\to t\overline{t}H $$ amplitudes with a light-quark loop

Abstract We compute the two-loop master integrals for leading-color QCD scattering amplitudes including a closed light-quark loop in$$ t\overline{t}H $$ t t ¯ H production at hadron colliders. Exploiting numerical evaluations in modular arithmetic, we construct a basis of master integrals satisfying a system of differential equations inϵ-factorized form. We present the analytic form of the differential equations in terms of a minimal set of differential one-forms. We explore properties of the function space of analytic solutions to the differential equations in terms of iterative integrals which can be exploited for studying the analytic form of related scattering amplitudes. Finally, we solve the differential equations using generalized series expansions to numerically evaluate the master integrals in physical phase space. As the first computation of a set of two-loop seven-scale master integrals, our results provide valuable input for analytic studies of scattering amplitudes in processes involving massive particles and a large number of kinematic scales.

Physics↗

Tailored computational approaches to interrogate heavy element chemistry and structure in condensed phase

In this chapter we are presenting a brief review of the challenges encountered in the study of 4f and 5f block elements in the condensed phase. Their recovery, use in molten salt reactors and other interesting applications necessitate the use of molecular dynamics and large-scale models that take into account both the electronic structure and relativistic corrections. Sampling the multitude of electronic and atomic configurational states is at the heart of reliable predictions of structure, reactivity, dynamics and transport of heavy metals in complex environments. We present three examples that combine lanthanide elements with large scale models: i) our recent developments of a versatile adaptive learning method that enables global optimization in high dimensional spaces, ii) results of computed pKa values of lanthanide aqua complexes, and iii) structure and computed EXAFS of heavy elements in molten salts. The latter two employ our recently-optimized lanthanide pseudo-potentials and companion basis sets for condensed phase ab initio molecular dynamics.

Nguyen, Manh Thuong↗

A conservative phase-space moving-grid strategy for a 1D-2V Vlasov–Fokker–Planck Solver

In this work, we develop a conservative configuration- and velocity-space (i.e., phase-space) moving-grid strategy for the Vlasov–Fokker–Planck (VFP) equation in a planar geometry. The velocity-space grid is normalized and shifted in terms of the thermal speed and the bulk-fluid velocity, respectively. The configuration-space grid is moved according to a mesh-motion-partial-differential equation (MMPDE), which equidistributes a monitor function that is inversely proportional to the gradient-length scales of the macroscopic plasma quantities. The resulting inertial terms in the transformed VFP equations are discretized to ensure the discrete conservation of mass, momentum, and energy. To satisfy the discrete conservation theorems in the presence of phase-space mesh motion, we employ the method of discrete nonlinear constraints – explored in previous studies – but the underlying symmetries are determined in a much more efficient manner than before. The conservative grid-adaptivity strategy provides an efficient scheme that resolves important physical structures in the phase-space while controlling the computational complexity at all times. We demonstrate the favorable features of the proposed algorithm through a set of test cases of increasing complexity. The problems test independent components of the algorithms, as well as the integrated capability on settings relevant to inertial confinement fusion.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗