Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “linear scaling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

On the statistical theory of self-gravitating collisionless dark matter flow: Scale and redshift variation of velocity and density distributions

The statistics of velocity and density fields are crucial for cosmic structure formation and evolution. Here, this paper extends our previous work on the two-point second-order statistics for the velocity field [Phys. Fluids 35, 077105 (2023)] to one-point probability distributions for both density and velocity fields. The scale and redshift variation of density and velocity distributions are studied by a halo-based non-projection approach. First, all particles are divided into halo and out-of-halo particles so that the redshift variation can be studied via generalized kurtosis of distributions for halo and out-of-halo particles, respectively. Second, without projecting particle fields onto a structured grid, the scale variation is analyzed by identifying all particle pairs on different scales $r$. We demonstrate that: (i) Delaunay tessellation can be used to reconstruct the density field. The density correlation, spectrum, and dispersion functions were obtained, modeled, and compared with the N-body simulation; (ii) the velocity distributions are symmetric on both small and large scales and are non-symmetric with a negative skewness on intermediate scales due to the inverse energy cascade on small scales with a constant rate $\varepsilon_u$; (iii) On small scales, the even order moments of pairwise velocity $\Delta u_L$ follow a two-thirds law $\propto{(-\varepsilon_ur)}^{2/3}$, while the odd order moments follow a linear scaling $\langle(\Delta u_L)^{2n+1}\rangle=(2n+1)\langle(\Delta u_L)^{2n}\rangle\langle\Delta u_L\rangle\propto{r}$; (iv) The scale variation of the velocity distributions was studied for longitudinal velocities $u_L$ or $u_L^{'}$, pairwise velocity (velocity difference) $\Delta u_L$=$u_L^{'}$-$u_L$ and velocity sum $\Sigma u_L$=$u^{'}_L$+$u_L$. Fully developed velocity fields are never Gaussian on any scale, despite that they can initially be Gaussian; (v) On small scales, $u_L$ and $\Sigma u_L$ can be modeled by a $X$ distribution to maximize the entropy of the system. The distribution of $\Delta u_L$ can be different; (vi) On large scales, $\Delta u_L$ and $\Sigma u_L$ can be modeled by a logistic or a $X$ distribution, while $u_L$ has a different distribution; (vii) the redshift variation of the velocity distributions follows the evolution of the $X$ distribution involving a shape parameter $\alpha(z)$ decreasing with time.

79 ASTRONOMY AND ASTROPHYSICS↗

Scalable Gaussian Processes, GPyTorch Application Benchmarking, and Targeted Adaptive Design (TAD) on ThetaGPU

We aim at showcasing the scalability of Gaussian Process (GP). The naive GP implementation scales cubically with data size, which can be prohibitive, so GP has not heretofore been considered suitable for very large-scale problem settings. We take advantage of GPyTorch, a library for scalable GPs built on top of PyTorch that incorporates GPU acceleration. With GPyTorch, one can achieve nearly linear scaling with structured kernel interpolation (SKI) and constant-time predictive covariances computation with LanczOs Variance Estimates (LOVE) while preserving accuracy. We also take advantage of the computational power of ThetaGPU, a supercomputer of Argonne Leadership Computing Facility (ALCF). In addition, we implement a scalable, GPU-ready version of Targeted Adaptive Design (TAD), a GP-based data-driven algorithm that efficiently searches the control space of an advanced manufacturing experiment for settings capable of producing a required design within a specified tolerance, despite the poorly known mapping from control settings to design. We finally show our benchmarking for GPyTorch and TAD performance on CPU vs. ThetaGPU and discuss the results and implications.

97 MATHEMATICS AND COMPUTING↗

DDStore: Distributed Data Store for Scalable Training of Graph Neural Networks on Large Atomistic Modeling Datasets

Graph neural networks (GNNs) are a class of Deep Learning models used in designing atomistic materials for effective screening of large chemical spaces. To ensure robust prediction, GNN models must be trained on large volumes of atomistic data on leadership class supercomputers. Even with the advent of modern architectures that consist of multiple storage layers that include node-local NVMe devices in addition to device memory for caching large datasets, extreme-scale model training faces I/O challenges at scale.We present DDStore, an in-memory distributed data store designed for GNN training on large-scale graph data. DDStore provides a hierarchical, distributed, data caching technique that combines data chunking, replication, low-latency random access, and high throughput communication. DDStore achieves near-linear scaling for training a GNN model using up to 1000 GPUs on the Summit and Perlmutter supercomputers, and reaches up to a 6.15x reduction in GNN training time compared to state-of-the-art methodologies.

Choi, Jong Youl↗

SPARC-X: Quantum simulations at extreme scale - reactive dynamics from first principles

We have developed the massively parallel electronic structure code SPARC-X: a computational framework for performing Kohn-Sham Density Functional Theory (DFT) calculations that can scale linearly with the number of atoms in the system, while being able to leverage petascale and emerging exascale parallel computers to study chemical phenomena at unprecedented length and time scales. SPARC-X exploits a recent breakthrough in electronic structure methodologies: systematically improvable, strictly local, orthonormal, discontinuous real-space bases that efficiently and systematically capture the local chemistry of the system. With further adaptation using new machine-learning techniques and the use of the massively parallel Spectral Quadrature (SQ) electronic structure method, the algorithmic complexity and prefactor associated with DFT calculations involving semilocal as well as hybrid functionals are dramatically reduced. Using petascale computational resources, SPARC-X enables quantum mechanical simulations at length and time scales previously accessible only by empirical approaches, e.g., 1,000,000 atoms for a few picoseconds using semilocal functionals or 1,000 atoms for a few picoseconds using hybrid functionals. Using exascale resources, the sizes and times targeted are two orders of magnitude larger. Such a capability has applications in a wide variety of chemical sciences, including reactive interfaces where large length- and/or long time-scales are needed and traditional force fields fail. This is particularly important in dynamic catalysis, where bond breaking and formation must be understood in detail. We developed, tested, and employed the SPARC-X framework to understand the photocatalytic properties of TiO 2 nanoparticles, revealing finite size effects that cannot be captured with standard model systems or functionals. This integrated development and application strategy ensures that SPARC-X remains a robust, efficient, and scalable software package for quantum simulations on current petascale and emerging exascale computing resources.

97 MATHEMATICS AND COMPUTING↗

Growth morphology with anisotropic surface kinetics

The morphological evolution of crystals growing from an incongruent vapor phase is studied using a Monte Carlo model, and the full range of growth morphologies is recovered. The diffusion in the bulk nutrient and the anisotropy in the interface kinetics are morphologically destabilizing and stabilizing, respectively. For a given set of simulation parameters and lattice symmetries there is a critical size, which scales linearly with the mean free path in the vapor, beyond which a crystal cannot retain its stable, macroscopically faceted growth shape. Surface diffusion stabilizes faceted growth on the shorter scale of the mean surface diffusion length. In simulations with a uniform drift superimposed on the random walk nutrient transport, crystal faces oriented toward the drift show enhanced morphological stability compared to the purely diffusive situation. Rotational drifts with periodic reversal of direction are morphologically stabilizing for all crystal facets.

Xiao, Rong-Fu↗

Changes in photochemically significant solar UV spectral irradiance as estimated by the composite Mg II index and scale factors

Quantitative assessment of the impact of solar ultraviolet irradiance variations on stratospheric ozone abundances currently requires the use of proxy indicators. The Mg II core-to-wing index has been developed as an indicator of solar UV activity between 175-400 nm that is independent of most instrument artifacts, and measures solar variability on both rotational and solar cycle time scales. Linear regression fits have been used to merge the individual Mg II index data sets from the Nimbus-7, NOAA-9, and NOAA-11 instruments onto a single reference scale. The change in 27-dayrunning average of the composite Mg II index from solar maximum to solar minimum is approximately 8 percent for solar cycle 21, and approximately 9 percent for solar cycle 22 through January 1992. Scaling factors based on the short-term variations in the Mg II index and solar irradiance data sets have been developed to estimate solar variability at mid-UV and near-UV wavelengths. Near 205 nm, where solar irradiance variations are important for stratospheric photo-chemistry and dynamics, the estimated change in irradiance during solar cycle 22 is approximately 10 percent using the composite Mg II index and scale factors.

Deland, Matthew T.↗

Statistics of galaxy orientations - Morphology and large-scale structure

Using the Uppsala General Catalog of bright galaxies and the northern and southern maps of the Lick counts of galaxies, statistical evidence of a morphology-orientation effect is found. Major axes of elliptical galaxies are preferentially oriented along the large-scale features of the Lick maps. However, the orientations of the major axes of spiral and lenticular galaxies show no clear signs of significant nonrandom behavior at a level of less than about one-fifth of the effect seen for ellipticals. The angular scale of the detected alignment effect for Uppsala ellipticals extends to at least theta of about 2 deg, which at a redshift of z of about 0.02 corresponds to a linear scale of about 2/h Mpc.

Lambas, Diego G.↗

Small-scale microwave background anisotropies implied by large-scale data

In the absence of reheating microwave background radiation (MBR) anisotropies on arcminute scales depend uniquely on the amplitude and the coherence length of the primordial density fluctuations (PDFs). These can be determined from the recent data on galaxy correlations, xi(r), on linear scales (APM survey). We develop here expressions for the MBR angular correlation function, C(theta), on arcminute scales in terms of the power spectrum of PDFs and demonstrate their accuracy by comparing with detailed calculations of MBR anisotropies. We then show how to evaluate C(theta) directly in terms of the observed xi(r) and show that the APM data give information on the amplitude, C(O), and the coherence angle of MBR anisotropies on small scales.

Kashlinsky, A.↗

Bias correcting regional scale Earth system model projections: novel approach using empirical mode decomposition

Bias correction is a crucial step in using Earth system model outputs for assessments, as it adjusts systematic errors by comparing the model to observations. However, standard methods – ranging from mean-based linear scaling to distribution-based quantile mapping typically treat bias correction as a single-scale process, overlooking the fact that biases can manifest differently across daily, seasonal, and annual timescales. In this study, we propose a novel, timescale-aware bias-correction approach built on Empirical Mode Decomposition. By decomposing the meteorological signal into multiple oscillatory components and aggregating them to represent distinct timescales, we apply targeted corrections to each component, thereby preserving both short- and long-term structure in the data. Experimental illustrations show that the timescale-aware EMDBC framework matches the performance of conventional quantile-delta mapping (QDM) at the native daily scale and achieves progressively larger bias reductions at bi-weekly, seasonal, and annual scales. As a result, the proposed approach offers a more robust path to accurate and reliable Earth system projections, strengthening their utility for resilience and adaptation planning.

Ganguli, Arkaprabha [Argonne National Laboratory (↗

Probing the galaxy–halo connection with total satellite luminosity

ABSTRACT We demonstrate how the total luminosity in satellite galaxies is a powerful probe of dark matter haloes around central galaxies. The method cross-correlates central galaxies in spectroscopic galaxy samples with fainter galaxies detected in photometric surveys. Using models, we show that the total galaxy luminosity, Lsat, scales linearly with host halo mass, making Lsat an excellent proxy for Mh. Lsat is also sensitive to the formation time of the halo. We demonstrate that probes of galaxy large-scale environment can break this degeneracy. Although this is an indirect probe of the halo, it yields a high signal-to-noise ratio measurement for galaxies expected to occupy haloes at <1012 M⊙, where other methods suffer from larger errors. In this paper, we focus on observational and theoretical systematics in the Lsat method. We test the robustness of our method of finding central galaxies and our methods of estimating the number of background galaxies. We implement this method on galaxies in the Sloan Digital Sky Survey (SDSS) data, with satellites identified in fainter imaging data. We find excellent agreement between our theoretical predictions and the observational measurements. Finally, we compare our Lsat measurements to weak lensing estimates of Mh for red and blue subsamples. In the stellar mass range where the measurements overlap, we find consistent results, where red galaxies live in larger haloes. However, the Lsat approach allows us to probe significantly lower mass galaxies. At these masses, the Lsat values are equivalent. This example shows the potential of Lsat as a probe of dark haloes.

79 ASTRONOMY AND ASTROPHYSICS↗

Generative Representations for Automated Design of Robots

A method of automated design of complex, modular robots involves an evolutionary process in which generative representations of designs are used. The term generative representations as used here signifies, loosely, representations that consist of or include algorithms, computer programs, and the like, wherein encoded designs can reuse elements of their encoding and thereby evolve toward greater complexity. Automated design of robots through synthetic evolutionary processes has already been demonstrated, but it is not clear whether genetically inspired search algorithms can yield designs that are sufficiently complex for practical engineering. The ultimate success of such algorithms as tools for automation of design depends on the scaling properties of representations of designs. A nongenerative representation (one in which each element of the encoded design is used at most once in translating to the design) scales linearly with the number of elements. Search algorithms that use nongenerative representations quickly become intractable (search times vary approximately exponentially with numbers of design elements), and thus are not amenable to scaling to complex designs. Generative representations are compact representations and were devised as means to circumvent the above-mentioned fundamental restriction on scalability. In the present method, a robot is defined by a compact programmatic form (its generative representation) and the evolutionary variation takes place on this form. The evolutionary process is an iterative one, wherein each cycle consists of the following steps: 1. Generative representations are generated in an evolutionary subprocess. 2. Each generative representation is a program that, when compiled, produces an assembly procedure. 3. In a computational simulation, a constructor executes an assembly procedure to generate a robot. 4. A physical-simulation program tests the performance of a simulated constructed robot, evaluating the performance according to a fitness criterion to yield a figure of merit that is fed back into the evolutionary subprocess of the next iteration. In comparison with prior approaches to automated evolutionary design of robots, the use of generative representations offers two advantages: First, a generative representation enables the reuse of components in regular and hierarchical ways and thereby serves a systematic means of creating more complex modules out of simpler ones. Second, the evolved generative representation may capture intrinsic properties of the design problem, so that variations in the representations move through the design space more effectively than do equivalent variations in a nongenerative representation. This method has been demonstrated by using it to design some robots that move, variously, by walking, rolling, or sliding. Some of the robots were built (see figure). Although these robots are very simple, in comparison with robots designed by humans, their structures are more regular, modular, hierarchical, and complex than are those of evolved designs of comparable functionality synthesized by use of nongenerative representations.

Homby, Gregory S.↗

Accelerating Collective Communication in Data Parallel Training across Deep Learning Frameworks

This work develops new techniques within Horovod, a generic communication library supporting data parallel training across deep learning frameworks. In particular, we improve the Horovod control plane by implementing a new coordination scheme that takes advantage of the characteristics of the typical data parallel training paradigm, namely the repeated execution of collectives on the gradients of a fixed set of tensors. Using a caching strategy, we execute Horovod’s existing coordinator-worker logic only once during a typical training run, replacing it with a more efficient decentralized orchestration strategy using the cached data and a global intersection of a bitvector for the remaining training duration. Next, we introduce a feature for end users to explicitly group collective operations, enabling finer grained control over the communication buffer sizes. To evaluate our proposed strategies, we conduct experiments on a world-class supercomputer — Summit. We compare our proposals to Horovod’s original design and observe 2x performance improvement at a scale of 6000 GPUs; we also compare them against tf.distribute and torch.DDP and achieve 12% better and comparable performance, respectively, using up to 1536 GPUs; we compare our solution against BytePS in typical HPC settings and achieve about 20% better performance on a scale of 768 GPUs. Finally, we test our strategies on a scientific application (STEMDL) using up to 27,600 GPUs (the entire Summit) and show that we achieve a near-linear scaling of 0.93 with a sustained performance of 1.54 exaflops (with standard error +- 0.02) in FP16 precision.

Romero, Joshua↗

Small-Scale Drop-Size Variability: Empirical Models for Drop-Size-Dependent Clustering in Clouds

By analyzing aircraft measurements of individual drop sizes in clouds, it has been shown in a companion paper that the probability of finding a drop of radius r at a linear scale l decreases as l(sup D(r)), where 0 less than or equals D(r) less than or equals 1. This paper shows striking examples of the spatial distribution of large cloud drops using models that simulate the observed power laws. In contrast to currently used models that assume homogeneity and a Poisson distribution of cloud drops, these models illustrate strong drop clustering, especially with larger drops. The degree of clustering is determined by the observed exponents D(r). The strong clustering of large drops arises naturally from the observed power-law statistics. This clustering has vital consequences for rain physics, including how fast rain can form. For radiative transfer theory, clustering of large drops enhances their impact on the cloud optical path. The clustering phenomenon also helps explain why remotely sensed cloud drop size is generally larger than that measured in situ.

Marshak, Alexander↗

Core States of Neutron Stars from Anatomizing Their Scaled Structure Equations

Given an Equation of State (EOS) for neutron star (NS) matter, there is a unique mass–radius sequence characterized by a maximum mass ${M}_{\mathrm{NS}}^{\max}$ at radius $R$ max . We first show analytically that the ${M}_{\mathrm{NS}}^{\max}$ and $R$ max scale linearly with two different combinations of the NS central pressure $P$ $c$ and energy density $ε$ $c$ , by dissecting perturbatively the dimensionless Tolman–Oppenheimer–Volkoff (TOV) equations governing NS internal variables. The scaling relations are then verified via 87 widely used and rather diverse phenomenological as well as 17 microscopic NS EOSs with/without considering hadron–quark phase transitions and hyperons, by solving numerically the original TOV equations. The EOS of the densest NS matter allowed before it collapses into a black hole is then obtained. Using the universal ${M}_{\mathrm{NS}}^{\max}$ and $R$ max scalings and Neutron Star Interior Composition Explorer and XMM-Newton mass–radius observational data for PSR J0740+6620, a very narrow constraining band on the NS central EOS is extracted directly from the data for the first time, without using any specific input EOS model.

79 ASTRONOMY AND ASTROPHYSICS↗

Small-scale signatures of primordial non-Gaussianity in k-nearest neighbour cumulative distribution functions

ABSTRACT Searches for primordial non-Gaussianity in cosmological perturbations are a key means of revealing novel primordial physics. However, robustly extracting signatures of primordial non-Gaussianity from non-linear scales of the late-time Universe is an open problem. In this paper, we apply k-Nearest Neighbour cumulative distribution functions, kNN-CDFs, to the quijote-png simulations to explore the sensitivity of kNN-CDFs to primordial non-Gaussianity. An interesting result is that for halo samples with $M_\mathrm{ h}\langle 10^{14}$ M$_\odot$ $h^{-1}$, the kNN-CDFs respond to equilateral PNG in a manner distinct from the other parameters. This persists in the galaxy catalogues in redshift space and can be differentiated from the impact of galaxy modelling, at least within the halo occupation distribution (HOD) framework considered here. kNN-CDFs are related to counts-in-cells and, through mapping a subset of the kNN-CDF measurements into the count-in-cells picture, we show that our results can be modelled analytically. A caveat of the analysis is that we only consider the HOD framework, including assembly bias. It will be interesting to validate these results with other techniques for modelling the galaxy–halo connection, e.g. (hybrid) effective field theory or semi-analytical methods.

Coulton, William R. (ORCID:0000000212973673)↗

Correlations in cosmic density fields

A method is proposed to place constraints on the functional form of the high-order correlation functions zeta(sub n) that arise in cosmic density fields at large scales. This technique is based on a mass-in-cell statistic and a difference of mass in partitions of a cell. The relationship between these measures is sensitive to the formal structure of the zeta(sub n) as well as their amplitudes. This relationship is quantified in several theoretical models of structure, based on the hierarchical clustering paradigm. The results lead to a test for specific types of hierarchical clustering that is sensitive to correlations of all orders. The method is applied to examples of simulated large-scaled structure dominated by cold dark matter. In the preliminary study, the hierarchical paradigm appears to be a realistic approximation over a broad range of the scales. Furthermore, there is evidence that graphs of low-order vertices are dominant. On the basis of simulated data a phenomological model is specified that gives a good representation of clustering from linear scales to the strongly clustered regime (zeta(sub 2) approximately 500).

Bromley, B. C.↗

Investigation of pedestal parameters and divertor heat fluxes in small ELM regimes in DIII-D

Abstract Divertor heat flux and its correlation with pedestal parameters within various small edge localized mode (ELM) regimes, including high beta poloidal, type-II and ELMs with negative triangularity H-modes were investigated in DIII-D. The parallel energy fluences of type-II and high beta poloidal small ELM regimes fall below the linear scaling with pedestal electron pressure for type-I ELMs put forward in Eich et al 2017 ( Nucl. Mater. Energy 12 84–90). The negative triangularity of H-mode ELMs follow the Eich scaling for type-I ELMs. The parallel heat flux and total heat loads to the divertor were determined using high-time resolution infrared thermography, while pedestal parameters were obtained through self-consistent kinetic equilibrium reconstructions. Linear regressions for the type-II and high beta poloidal regimes demonstrate that an equivalent 7.5 MA small ELM scenario in ITER would fall below the ~5 MJ m − 2 leading edge melting limit for tungsten (Gunn et al 2017 Nucl. Fusion 57 046025). Utilizing fast thermography, the scrape-off layer power fall-off length for both inter-ELM and intra-ELM was determined and compared to the Eich scaling with poloidal magnetic field in Eich et al (ASDEX Upgrade Team and JET EFDA Contributors 2013 Nucl. Fusion 53 093031). Except for the high beta poloidal scenario, all the small ELM regimes during both inter- and intra-ELM periods had power fall-off lengths ( λ q ) larger then would be expected from the B pol , MP − 1 scaling associated with type-I ELMs, signifying their potential in managing heat loads and offering a solution for core–edge integration.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Quantum Nonlinear Acoustic Hall Effect and Inverse Acoustic Faraday Effect in Dirac Insulators

Here, we propose to realize the quantum nonlinear Hall effect and the inverse Faraday effect through the acoustic wave in a time-reversal invariant but inversion broken Dirac insulator. We focus on the acoustic frequency much lower than the Dirac gap such that the interband transition is suppressed and these effects arise solely from the intrinsic valley-contrasting band topology. The corresponding acoustoelectric conductivity and magnetoacoustic susceptibility are both proportional to the quantized valley Chern number and independent of the quasiparticle lifetime. The linear and nonlinear components of the longitudinal and transverse topological currents can be tuned by adjusting the polarization and propagation directions of the surface acoustic wave. The static magnetization generated by a circularly polarized acoustic wave scales linearly with the acoustic frequency as well as the strain-induced charge density. Our results unveil a quantized nonlinear topological acoustoelectric response of gapped Dirac materials, like hexagonal boron nitride and transition-metal dichalcogenide, paving the way toward room-temperature acoustoelectric devices due to their large band gaps.

36 MATERIALS SCIENCE↗