Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “applied computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

A Perspective on Sustainable Computational Chemistry Software Development and Integration

The power of quantum chemistry to predict the ground and excited state properties of complex chemical systems has driven the development of computational quantum chemistry software, integrating advances in theory, applied mathematics, and computer science. The emergence of new computational paradigms associated with exascale technologies also poses significant challenges that require a flexible forward strategy to take full advantage of existing and forthcoming computational resources. In this context, the sustainability and interoperability of computational chemistry software development are among the most pressing issues. In this perspective, we discuss software infrastructure needs and investments with an eye to fully utilize exascale resources and provide unique computational tools for next-generation science problems and scientific discoveries.

36 MATERIALS SCIENCE↗

Modern chemical graph theory

Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures

Leite, Leonardo S. G.↗

Vertex protein PduN tunes encapsulated pathway performance by dictating bacterial metabolosome morphology

Engineering subcellular organization in microbes shows great promise in addressing bottlenecks in metabolic engineering efforts; however, rules guiding selection of an organization strategy or platform are lacking. Here, we study compartment morphology as a factor in mediating encapsulated pathway performance. Using the 1,2-propanediol utilization microcompartment (Pdu MCP) system from Salmonella enterica serovar Typhimurium LT2, we find that we can shift the morphology of this protein nanoreactor from polyhedral to tubular by removing vertex protein PduN. Analysis of the metabolic function between these Pdu microtubes (MTs) shows that they provide a diffusional barrier capable of shielding the cytosol from a toxic pathway intermediate, similar to native MCPs. However, kinetic modeling suggests that the different surface area to volume ratios of MCP and MT structures alters encapsulated pathway performance. Finally, we report a microscopy-based assay that permits rapid assessment of Pdu MT formation to enable future engineering efforts on these structures.

59 BASIC BIOLOGICAL SCIENCES↗

HQ-Sim: High-performance State Vector Simulation of Quantum Circuits on Heterogeneous HPC Systems

Quantum circuit simulations are applied in more and more circum- stances as the quantum computing community becomes broader. It helps researchers to evaluate quantum algorithms and relieve the burden of limited quantum computing resources. However, most of the state-of-the-art quantum simulators utilize either CPU or GPU to store and calculate the state vector, which results in resource starvation. Moreover, the maximum number of qubits supported by the simulator is bounded by the memory, since the memory utilization increases exponentially with the number of qubits. In this study, we leverage Heterogeneous computing to utilize both CPU and GPU to store and update state vectors. We also integrate lossy data compression to reduce memory requirements. Specifically, we develop a heterogeneous framework that has a dynamic scheduler to fully utilize the computing resources. We apply lossy compression to chunked state vector to make the maximum number of qubits higher than the regular simulators, the compression also benefits the data movement between CPU and GPU.

Zhang, Boyuan↗

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING↗

A scalable, open-source implementation of a large-scale mechanistic model for single cell proliferation and death signaling

Mechanistic models of how single cells respond to different perturbations can help integrate disparate big data sets or predict response to varied drug combinations. However, the construction and simulation of such models have proved challenging. Here, we developed a python-based model creation and simulation pipeline that converts a few structured text files into an SBML standard and is high-performance- and cloud-computing ready. We applied this pipeline to our large-scale, mechanistic pan-cancer signaling model (named SPARCED) and demonstrate it by adding an IFNγ pathway submodel. We then investigated whether a putative crosstalk mechanism could be consistent with experimental observations from the LINCS MCF10A Data Cube that IFNγ acts as an anti-proliferative factor. The analyses suggested this observation can be explained by IFNγ-induced SOCS1 sequestering activated EGF receptors. This work forms a foundational recipe for increased mechanistic model-based data integration on a single-cell level, an important building block for clinically-predictive mechanistic models.

60 APPLIED LIFE SCIENCES↗

Physics-constrained machine learning for electrodynamics without gauge ambiguity based on Fourier transformed Maxwell’s equations

We utilize a Fourier transformation-based representation of Maxwell’s equations to develop physics-constrained neural networks for electrodynamics without gauge ambiguity, which we label the Fourier–Helmholtz–Maxwell neural operator method. In this approach, both of Gauss’s laws and Faraday’s law are built in as hard constraints, as well as the longitudinal component of Ampère–Maxwell in Fourier space, assuming the continuity equation. An encoder–decoder network acts as a solution operator for the transverse components of the Fourier transformed vector potential, $\hat{A}_⟂(k,t)$, whose two degrees of freedom are used to predict the electromagnetic fields. This method was tested on two electron beam simulations. Among the models investigated, it was found that a U-Net architecture exhibited the best performance as it trained quicker, was more accurate and generalized better than the other architectures examined. We demonstrate that our approach is useful for solving Maxwell’s equations for the electromagnetic fields generated by intense relativistic charged particle beams and that it generalizes well to unseen test data, while being orders of magnitude quicker than conventional simulations. We show that the model can be re-trained to make highly accurate predictions in as few as 20 epochs on a previously unseen data set.

97 MATHEMATICS AND COMPUTING↗

UMap: An application-oriented user level memory mapping library

Exploiting the prominent role of complex memories in exascale node architecture, the UMap page fault handler offers new capabilities to access large memory-mapped data sets directly. UMap provides flexible configuration options to customize page handling to each application, including analysis of massive observational and simulation data sets. The high-performance design features I/O decoupling, dynamic load balancing, and application-level controls. Page faults triggered by application threads and processes accessing data mapped to a UMapp’ed region are handled via the Linux userfaultfd protocol, an asynchronous message-oriented kernel-user communication mechanism that avoids the context switch penalty of traditional signal fault handlers. UMap is fully open source. In this paper, we give an overview of the UMap library architecture, its extensible plugin architecture, and the use/performance of UMap in emerging heterogeneous memory hierarchies such as near-node Non-volatile Memory (NVM) and network attached memories. We highlight new capabilities in two pagefault management plugins, the NetworkStore and SparseStore. We demonstrate the integration between UMap and multiple ECP products including Caliper, Metall, ZFP, Mochi, and Ripples.

97 MATHEMATICS AND COMPUTING↗

Memcomputing the Spectrum of Correlated Quantum Systems

The goals and objectives of the grant DE‐SC0020892 were to apply a new computing paradigm, MemComputing, to efficiently simulate properties of correlated quantum systems. The method has been applied to a wide set of problems ranging from quantum state tomography to finding the ground state of correlated systems. In all cases, substantial advantages compared to state-of-the-art approaches have been obtained. The project has also led to the suggestion of the transformer architecture (used nowadays in large-language models) as an efficient quantum state representation, and a better understanding of the role of memory in the generation of long-range order in neural networks. This grant has supported the work of a PhD student, inspired a new class on unconventional computing taught at the University of California, San Diego and has generated several peer-reviewed papers.

97 MATHEMATICS AND COMPUTING↗

Landau Singularities Revisited: Computational Algebraic Geometry for Feynman Integrals

We reformulate the analysis of singularities of Feynman integrals in a way that can be practically applied to perturbative computations in the standard model in dimensional regularization. After highlighting issues in the textbook treatment of Landau singularities, we develop an algorithm for classifying and computing them using techniques from computational algebraic geometry. We introduce an algebraic variety called the principal Landau determinant, which captures the singularities even in the presence of massless particles or UV/IR divergences. We illustrate this for 114 example diagrams, including a cutting-edge 2-loop 5-point nonplanar QCD process with multiple mass scales. Published by the American Physical Society 2024

Physics↗

Accelerating Scientific Simulations with Bi-Fidelity Weighted Transfer Learning

High-fidelity modeling is an essential design tool for many engineering applications. However, for complex systems, computational cost can be a limiting factor. Analyzing parameter sensitivity, uncertainty quantification, and design optimization require many model evaluations. Surrogate models are often used to develop the relationship between model parameters and quantities of interest. However, in the case of complex systems, surrogate models require several degrees of freedom and, thus, a large number of data points to determine the correct dependencies. For many applications, this may be prohibitively expensive. The reduction of computational requirements can be achieved by leveraging low-fidelity models. Low-fidelity models represent the system at a coarser resolution with the advantage of computational efficiency. Therefore, a bi-fidelity modeling paradigm, which augments the accuracy of a low-fidelity model in a computationally efficient manner by invoking limited runs of a high-fidelity model, can be leveraged to sufficiently balance the accuracy and computational requirements. In this work, a bi-fidelity weighted transfer learning method using neural networks was applied to a computational fluid dynamics heat transfer modeling problem. The transfer learning advantage was investigated as a function of hyperparameters. Our main finding is that the use of a bi-fidelity modeling paradigm achieves accuracy close to that of a high-fidelity Gaussian process model while significantly reducing computational cost. The bi-fidelity model achieves comparable performance with 90 high-fidelity samples-that is, 60% less than the samples needed to achieve similar accuracy without the use of bi-fidelity modeling,

Borowiec, Katarzyna↗

Cloud droplet diffusional growth in homogeneous isotropic turbulence: bin microphysics versus Lagrangian super-droplet simulations

The increase in the spectral width of an initially monodisperse population of cloud droplets in homogeneous isotropic turbulence is investigated by applying a finite-difference fluid flow model combined with either Eulerian bin microphysics or a Lagrangian particle-based scheme. The turbulence is forced applying a variant of the so-called linear forcing method that maintains the mean turbulent kinetic energy (TKE) and the TKE partitioning between velocity components. The latter is important for maintaining the quasi-steady forcing of the supersaturation fluctuations that drive the increase in the spectral width. We apply a large computational domain (64 3 m 3 ), one of the domains considered in Thomas et al. (2020). The simulations apply 1 m grid length and are in the spirit of the implicit large eddy simulation (ILES), that is, with small-scale dissipation provided by the model numerics. This is in contrast to the scaled-up direct numerical simulation (DNS) applied in Thomas et al. (2020). Two TKE intensities and three different droplet concentrations are considered. Analytic solutions derived in Sardina et al. (2015), valid for the case when the turbulence integral timescale is much larger than the droplet phase relaxation timescale, are used to guide the comparison between the two microphysics simulation techniques. The Lagrangian approach reproduces the scalings relatively well. Representing the spectral width increase in time is more challenging for the bin microphysics because appropriately high resolution in the bin space is needed. The bin width of 0.5 µm is only sufficient for the lowest droplet concentration (26 cm -3 ). For the highest droplet concentration (650 cm -3 ), an order of magnitude smaller bin size is barely sufficient. The scalings are not expected to be valid for the lowest droplet concentration and the high-TKE case, and the two microphysics schemes represent similar departures. Finally, because the fluid flow is the same for all simulations featuring either low or high TKE, one can compare point-by-point simulation results. Such a comparison shows very close temperature and water vapor point-by-point values across the computational domain and larger differences between simulated mean droplet radii and spectral width. The latter are explained by fundamental differences in the two simulation methodologies, numerical diffusion in the Eulerian bin approach and a relatively small number of Lagrangian particles that are used in the particle-based microphysics.

54 ENVIRONMENTAL SCIENCES↗

The LSBmax algorithm for boosting resilience of electric grids post (N‐2) contingencies

Abstract A computationally improved algorithm is presented to find the best transmission switching (TS) candidate for boosting resilience of electricity grids subject to ( N ‐2) contingencies. Here, resilience is computed as the reduction in load shed after the above‐mentioned ( N‐ ) contingencies. TS is a planned line outage, and past research shows that changing the transmission system's topology changes the power flow and removes post contingency violations. Finding the best TS candidate in a computationally suitable time for effectively boosting resilience is a challenge. The best TS candidate is found using a novel heuristic method by decreasing the search space based on proximity to the bus with the maximum load shedding (LSB). The LSB algorithm is faster than existing algorithms in the literature; and, it is compatible with both the AC and DC optimal power flow formulations. To validate the authors' claims of speedup and accuracy, two metrics are used to analyze the results from the IEEE 39‐bus and 118‐bus systems. Finally, the inherent parallelism of the LSB algorithm is leveraged on a high‐performance computing platform and applied to the large‐scale Polish 2383‐bus test system to validate scalability in both size and speedup in computation time.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reduced scaling formulation of CASPT2 analytical gradients using the supporting subspace method

We present a reduced scaling and exact reformulation of state specific complete active space second-order perturbation (CASPT2) analytical gradients in terms of the MP2 and Fock derivatives using the supporting subspace method. This work follows naturally from the supporting subspace formulation of the CASPT2 energy in terms of the MP2 energy using dressed orbitals and Fock builds. For a given active space configuration, the terms corresponding to the MP2-gradient can be evaluated with O(N5) operations, while the rest of the calculations can be computed with O(N3) operations using Fock builds, Fock gradients, and linear algebra. When tensor-hyper-contraction is applied simultaneously, the computational cost can be further reduced to O(N4) for a fixed active space size. The new formulation enables efficient implementation of CASPT2 analytical gradients by leveraging the existing graphical processing unit (GPU)-based MP2 and Fock routines. We present benchmark results that demonstrate the accuracy and performance of the new method. Example applications of the new method in ab initio molecular dynamics simulation and constrained geometry optimization are given.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unifying Combinatorial and Graphical Methods in Artificial Intelligence

Recently, a new graph Laplacian, called the inner product Laplacian, was introduced which generalizes many existing Laplacians, including the normalized and combinatorial Laplacian and their weighted variants. The key observation behind the inner product Laplacian is that by defining appropriate inner product spaces on the vertices and edges, the standard Laplacians can be recovered as Hodge Laplacians over the simplicial complex formed by the edges and vertices. These inner product spaces form a natural way to incorporate non-combinatorial information into the definition of a domain-specific Laplacian. In particular, in contrast to current domain-specific weighting schemes which rely solely on edge weights, information regarding the similarity of non-adjacent vertices and arbitrary pairs of edges can be effectively incorporated into the Laplacian. In order to illustrate this approach we consider the problem of calculating the potential energy of an atomistic configuration using Graph Neural Networks. In comparison with start-of-the-art approaches, such as SchNet, our approach replaces a learned (via auto-encoder) representation of the atom types with an inner product space on atoms based on scientific knowledge (e.g., electronegativity). We will illustrate how this approach captures key chemical properties of the molecules and compare the energy calculations with state-of-the-art neural network approaches. However, to compute the resulting Laplacian involves a mixture of sparse and dense matrix computation and yields a dense matrix as the basis for the graph convolution. This dense convolutional kernel necessitates moving away from the standard message passing framework for graph neural networks and increases the computational cost of applying the kernel. In order to mitigate these costs we investigate means of leveraging the mixed sparse and dense computations to reduce the overall computational cost and how these approaches can be automatically transferred to energy efficient hardware (e.g., field programmable gate arrays (FPGAs)).

97 MATHEMATICS AND COMPUTING↗

(U) A General-Purpose Code for Correlated Sampling Using Batch Statistics with MCNP6 for Fixed-Source Problems

Correlated sampling can be used to reduce the uncertainty of a difference of tallies by taking advantage of the negative covariance term in the sandwich formula. Booth first showed how correlated sampling can be applied with batch statistics using MCNP’s tally fluctuation chart (TFC) to reduce the uncertainty of a difference of tallies in fixed-source problems. Booth presented a problem in which a 1273% uncertainty in a difference was reduced to 8% by accounting for correlations. Researchers He and Su recently studied correlated sampling using the TFC in MCNP version 5. They determined that the code did not print enough digits in the TFC tally means for accurate batch statistics in some cases. After modifying the source code, they concluded that “correlated sampling can yield a standard deviation of about one magnitude smaller than that predicted by the direct, un-correlated simulation when the changes in system response are small (say about 1%), which is equivalent to saving in CPU time by a factor of 100. Such saving [sic] becomes less significant as the change in system response becomes larger.” He and Su provided the formulas needed to apply batch statistics to compute the correlated uncertainty of a difference of tallies. In this report, we follow up on their work by providing the formulas needed to apply batch statistics to compute the correlated uncertainty of a ratio of tallies and of a difference of two tallies divided by a third tally. We extend these formulas to differences and ratios of ratios. These formulas are applied to reduce the uncertainty associated with calculating a relative sensitivity. He and Su did not investigate the accuracy of their correlated sampling uncertainty estimates. We use their test problems and evaluate the accuracy of the uncertainty estimates by comparing with results obtained from random sampling, and, in simple cases, with theoretical values of the “exact” uncertainties. We find that the uncertainties obtained from batch statistics are accurate as long as at least 100 batches are used. We present a new computer code, COSUBS (COrrelated Sampling Using Batch Statistics), that reads MCNP6 TFCs and applies correlated sampling using batch statistics for the tally combinations that the user specifies. COSUBS is a very general tool that compares all TFCs for a base case and one or two perturbed cases. It computes uncertainties for ratios if given only a base case. This report is organized as follows. The equations to apply batch statistics to the difference of random tallies are reviewed in Sec. II. Section III presents the equations for applying batch statistics to a ratio of random tallies; this is useful for computing relative sensitivities using a one-sided finite difference and the relative sensitivity using the differential operator method. Section IV presents the equations for applying batch statistics to a difference of two random tallies divided by a third; this is useful for computing a relative sensitivities using a central difference. Section V presents the equations for applying batch statistics to a difference of two ratios with four random tallies. Section VI presents the equations for applying batch statistics to a one-sided finite difference estimate of the relative sensitivity of a ratio (this uses four random tallies). Section VII presents the equations for applying batch statistics to a central difference estimate of the relative sensitivity of a ratio (this uses six random tallies). Section VIII presents the equations for applying batch statistics to a sum of random tallies. Section IX discusses how to apply batch statistics using MCNP6. Section X presents COSUBS, describing its command-line options and logic. Sections XI through XVI present numerical results for various test problems. Section XVII is a summary and conclusions. Appendix A derives the theoretical Monte Carlo tally variance given certain assumptions; these variances are used to verify the batch statistics for some of the problems. Appendix B lists the MCNP6 input for the unperturbed example problem. Appendix C presents modifications made to MCNP6.3 to support this work.

97 MATHEMATICS AND COMPUTING↗

Computational study of runaway electrons in MST tokamak discharges with applied resonant magnetic perturbation

A numerical study of magnetohydrodynamics (MHD) and tracer-particle evolution investigates the effects of resonant magnetic perturbations (RMPs) on the confinement of runaway electrons (REs) in tokamak discharges conducted in the Madison Symmetric Torus. In computational results of applying RMPs having a broad toroidal spectrum but a single poloidal harmonic, m = 1 RMP does not suppress REs, whereas m = 3 RMP achieves significant deconfinement but not the complete suppression obtained in the experiment. MHD simulations with the NIMROD code produce sawtooth oscillations, and the associated magnetic reconnection can affect the trajectory of REs starting in the core region. Simulations with m = 3 RMP produce chaotic magnetic topology over the outer region, but the m = 1 RMP produces negligible changes in field topology, relative to applying no RMP. Using snapshots of the MHD simulation fields, full-orbit relativistic electron test particle computations with KORC show ≈50% loss from the m = 3 RMP compared to the 10%–15% loss from the m = 1 RMP. Here, test particle computations of the m = 3 RMP in the time-evolving MHD simulation fields show correlation between MHD activity and late-time particle losses, but total electron confinement is similar to computations using magnetic-field snapshots.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Modeling, analysis, and optimization of complex nuclear processes and facilities via computational methods: The HALEU process case study

Improving and adapting industrial systems to timely meet changing programmatic and market demands is an important goal to achieve, including when operating and maintaining complex nuclear processes and facilities. However, changes to these complex systems are costly, particularly when they are already in place and bounded to stringent requirements and constraints such as when handling radioactive material and contaminated equipment. These conditions often exist when treating spent nuclear fuel remotely within shielded nuclear radiation chambers, commonly referred as hot cells, to condition nuclear material and/or fabricate products for utilization in other nuclear enterprises such as in the manufacture of advanced nuclear fuel. The illustrative case considered here is the production of high assay low enriched uranium (HALEU) products supporting the deployment of advanced nuclear reactors. For the HALEU program, resources invested were and are being systematically analyzed so that these investments are maximized in a facility that is nearly 60 years old. A methodology that has effectively enabled optimized and improvements in the Spent Fuel Treatment (SFT) program, and consequently the HALEU program, involves discrete event simulation as addressed in this article. Here, the quantification of multiple productivity metrics, including material processing rates, cycle times, bottlenecks, number of material transfers as well as equipment, workstation, and material handling utilization, has resulted in a myriad of diverse discoveries and data-informed decisions regarding process layout and constituent, labor levels and schedules, selection of new process units, storage needs, and other critical process configurations. This article describes such a computational capability being applied for decision-making, illustrates its application to an actual process and program, provides illustrative results, and argues how computational methods for the modeling, analysis, and optimization of complex processes and facilities does lead to informed decisions derived from data and not only from intuition.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗