Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Relativistic resolution-of-the-identity with Cholesky integral decomposition

In this study, we present an efficient integral decomposition approach called the restricted-kinetic-balance resolution-of-the-identity (RKB-RI) algorithm, which utilizes a tunable RI method based on the Cholesky integral decomposition for in-core relativistic quantum chemistry calculations. The RKB-RI algorithm incorporates the restricted-kinetic-balance condition and offers a versatile framework for accurate computations. Notably, the Cholesky integral decomposition is employed not only to approximate symmetric large-component electron repulsion integrals but also those involving small-component basis functions. In addition to comprehensive error analysis, we investigate crucial conditions, such as the kinetic balance condition and variational stability, which underlie the applicability of Dirac relativistic electronic structure theory. Here we compare the computational cost of the RKB-RI approach with the full in-core method to assess its efficiency. To evaluate the accuracy and reliability of the RKB-RI method proposed in this work, we employ actinyl oxides as benchmark systems, leveraging their properties for validation purposes. This investigation provides valuable insights into the capabilities and performance of the RKB-RI algorithm and establishes its potential as a powerful tool in the field of relativistic quantum chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fast and scalable quantum Monte Carlo simulations of electron-phonon models

We introduce methodologies for highly scalable quantum Monte Carlo simulations of electron-phonon models, and report benchmark results for the Holstein model on the square lattice. The determinant quantum Monte Carlo (DQMC) method is a widely used tool for simulating simple electron-phonon models at finite temperatures, but incurs a computational cost that scales cubically with system size. Alternatively, near-linear scaling with system size can be achieved with the hybrid Monte Carlo (HMC) method and an integral representation of the Fermion determinant. Here, we introduce a collection of methodologies that make such simulations even faster. To combat "stiffness" arising from the bosonic action, we review how Fourier acceleration can be combined with time-step splitting. To overcome phonon sampling barriers associated with strongly-bound bipolaron formation, we design global Monte Carlo updates that approximately respect particle-hole symmetry. To accelerate the iterative linear solver, we introduce a preconditioner that becomes exact in the adiabatic limit of infinite atomic mass. Finally, we demonstrate how stochastic measurements can be accelerated using fast Fourier transforms. Here, these methods are all complementary and, combined, may produce multiple orders of magnitude speedup, depending on model details.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Acceleration of the particle-in-cell code Osiris with graphics processing units

Fully relativistic particle-in-cell (PIC) simulations are crucial for advancing our knowledge of plasma physics. Modern supercomputers based on graphics processing units (GPUs) offer the potential to perform PIC simulations of unprecedented scale, but require robust and feature-rich codes that can fully leverage their computational resources. In this work, this demand is addressed by adding GPU acceleration to the PIC code Osiris. An overview of the algorithm, which features a CUDA extension to the underlying Fortran architecture, is given. Detailed performance benchmarks for thermal plasmas are presented, which demonstrate excellent weak scaling on NERSC's Perlmutter supercomputer and high levels of absolute performance. The robustness of the code to model a variety of physical systems is demonstrated via simulations of Weibel filamentation and laser-wakefield acceleration run with dynamic load balancing. Finally, measurements and analysis of energy consumption are provided that indicate that the GPU algorithm is up to ~14 times faster and ~7 times more energy efficient than the optimized CPU algorithm on a node-to-node basis. The described development addresses the PIC simulation community's computational demands both by contributing a robust and performant GPU-accelerated PIC code and by providing insight into efficient use of GPU hardware.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Multireference Equation-of-Motion Driven Similarity Renormalization Group: Theoretical Foundations and Applications to Ionized States

We present a formulation and implementation of an equation-of-motion (EOM) extension of the multireference driven similarity renormalization group (MR-DSRG) formalism for ionization potentials (IP-EOM-DSRG). The IP-EOM-DSRG formalism results in a Hermitian generalized eigenvalue problem, delivering accurate ionization potentials for strongly correlated systems. The EOM step scales as O(N 5 ) with the basis set size N, allowing for efficient calculation of spectroscopic properties, such as transition energies and intensities. The IP-EOM-DSRG formalism is combined with three truncation schemes of the parent MR-DSRG theory: an iterative nonperturbative method with up to two-body excitations [MR-LDSRG(2)] and second- and third-order perturbative approximations [DSRG-MRPT2/3]. We benchmark these variants by computing (1) the vertical valence ionization potentials of a series of small molecules at both equilibrium and stretched geometries; (2) the spectroscopic constants of several low-lying electronic states of the OH, CN, N 2 + , and CO + radicals; and (3) the binding curves of low-lying electronic states of the CN radical. A comparison with experimental data and theoretical results shows that all three IP-EOM-DSRG methods accurately reproduce the vertical ionization potentials and spectroscopic constants of these systems. Notably, the DSRG-MRPT3 and MR-LDSRG(2) versions outperform several state-of-the-art multireference methods of comparable or higher cost.

Hamiltonians↗

Accurate numerical simulations of open quantum systems using spectral tensor trains

Decoherence between qubits is a major bottleneck in quantum computations. Decoherence results from intrinsic quantum and thermal fluctuations as well as noise in the external fields that perform the measurement and preparation processes. With prescribed colored noise spectra for intrinsic and extrinsic noise, we present a numerical method, Quantum Accelerated Stochastic Propagator Evaluation (Q-ASPEN), to solve the time-dependent noise-averaged reduced density matrix in the presence of intrinsic and extrinsic noise. Q-ASPEN is arbitrarily accurate and can be applied to provide estimates for the resources needed to error-correct quantum computations. We employ spectral tensor trains, which combine the advantages of tensor networks and pseudospectral methods, as a variational ansatz to the quantum relaxation problem and optimize the ansatz using methods typically used to train neural networks. Here, the spectral tensor trains in Q-ASPEN make accurate calculations with tens of quantum levels feasible. We present benchmarks for Q-ASPEN on the spin-boson model in the presence of intrinsic noise and on a quantum chain of up to 32 sites in the presence of extrinsic noise. In our benchmark, the memory cost of Q-ASPEN scales as a low-order polynomial in the size of the system once the number of system states surpasses the number of basis functions used in the spectral expansion.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

The Gutzwiller conjugate gradient minimization method for correlated electron systems

In this report we review our recent work on the Gutzwiller conjugate gradient minimization method, an ab initio approach developed for correlated electron systems. The complete formalism has been outlined that allows for a systematic understanding of the method, followed by a discussion of benchmark studies of dimers, one- and two-dimensional single-band Hubbard models. In the end, we present some preliminary results of multi-band Hubbard models and large-basis calculations of F 2 to illustrate our efforts to further reduce the computational complexity.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Nuclear Responses with Neural-Network Quantum States

We introduce a variational Monte Carlo framework that combines neural-network quantum states with the Lorentz integral transform technique to compute the dynamical properties of self-bound quantum many-body systems in continuous Hilbert spaces. While broadly applicable to various quantum systems, including atoms and molecules, in this initial application we focus on the photoabsorption cross section of light nuclei, where benchmarks against numerically exact techniques are available. Our accurate theoretical predictions are complemented by robust uncertainty quantification, enabling meaningful comparisons with experiments. Here, we demonstrate that a relatively simple nuclear Hamiltonian—based on a leading-order pionless EFT expansion and known to accurately reproduce ground-state energies of nuclei with 𝐴 ≤ 40—also provides a reliable description of the photoabsorption cross section.

Ab initio calculations↗

OReole-FM: successes and challenges toward billion-parameter foundation models for high-resolution satellite imagery

While the pretraining of Foundation Models (FMs) for remote sensing (RS) imagery is on the rise, models remain restricted to a few hundred million parameters. Scaling models to billions of parameters has been shown to yield unprecedented benefits including emergent abilities, but requires data scaling and computing resources typically not available outside industry R&D labs. In this work, we pair high-performance computing resources including Frontier supercomputer, America's first exascale system, and high-resolution optical RS data to pretrain billion-scale FMs. Our study assesses performance of different pretrained variants of vision Transformers across image classification, semantic segmentation and object detection benchmarks, which highlight the importance of data scaling for effective model scaling. Moreover, we discuss construction of a novel TIU pretraining dataset, model initialization, with data and pretrained models intended for public release. By discussing technical challenges and details often lacking in the related literature, this work is intended to offer best practices to the geospatial community toward efficient training and benchmarking of larger FMs.

Ambrozio Dias, Philipe↗

Nuclear responses with neural-network quantum states

We introduce a variational Monte Carlo framework that combines neural-network quantum states with the Lorentz integral transform technique to compute the dynamical properties of self-bound quantum many-body systems in continuous Hilbert spaces. While broadly applicable to various quantum systems, including atoms and molecules, in this initial application we focus on the photoabsorption cross section of light nuclei, where benchmarks against numerically exact techniques are available. Our accurate theoretical predictions are complemented by robust uncertainty quantification, enabling meaningful comparisons with experiments. We demonstrate that a simple nuclear Hamiltonian, based on a leading-order pionless effective field theory expansion and known to accurately reproduce the ground-state energies of nuclei with $A\leq 20$ nucleons also provides a reliable description of the photoabsorption cross section.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Extending the computational reach of a superconducting qutrit processor

Quantum computing with qudits is an emerging approach that exploits a larger, more connected computational space, providing advantages for many applications, including quantum simulation and quantum error correction. Nonetheless, qudits are typically afflicted by more complex errors and suffer greater noise sensitivity which renders their scaling difficult. In this work, we introduce techniques to tailor arbitrary qudit Markovian noise to stochastic Weyl–Heisenberg channels and mitigate noise that commutes with our Clifford and universal two-qudit gate in generic qudit circuits. We experimentally demonstrate these methods on a superconducting transmon qutrit processor, and benchmark their effectiveness for multipartite qutrit entanglement and random circuit sampling, obtaining up to 3× improvement in our results. To the best of our knowledge, this constitutes the first-ever error mitigation experiment performed on qutrits. Our work shows that despite the intrinsic complexity of manipulating higher-dimensional quantum systems, noise tailoring and error mitigation can significantly extend the computational reach of today’s qudit processors.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Energetically consistent model reduction for metriplectic systems

The metriplectic formalism is useful for describing complete dynamical systems which conserve energy and produce entropy. This creates challenges for model reduction, as the elimination of high-frequency information will generally not preserve the metriplectic structure which governs long-term stability of the system. Based on proper orthogonal decomposition, a provably convergent metriplectic reduced-order model is formulated which is guaranteed to maintain the algebraic structure necessary for energy conservation and entropy formation. Further, numerical results on benchmark problems show that the proposed method is remarkably stable, leading to improved accuracy over long time scales at a moderate increase in cost over naive methods.

42 ENGINEERING↗

Thermal Management for FPGA Nodes in HPC Systems

The integration of FPGAs into large-scale computing systems is gaining attention. In these systems, real-time data handling for networking, tasks for scientific computing, and machine learning can be executed with customized datapaths on reconfigurable fabric within heterogeneous compute nodes. At the same time, thermal management, particularly battling the cooling cost and guaranteeing the reliability, is a continuing concern. The introduction of new heterogeneous components into HPC nodes only adds further complexities to thermal modeling and management. The thermal behavior of multi-FPGA systems deployed within large compute clusters is less explored. Here, we first show that the thermal behaviors of different FPGAs of the same generation can vary due to their physical locations in a rack and process variation, even though they are running the same tasks. We present a machine learning–based model to capture the thermal behavior of each individual FPGA in the cluster. We then propose two thermal management strategies guided by our thermal model. First, we mitigate thermal variation and hotspots across the cluster by proactive thermal-aware task placement. Under the tested system and benchmarks, we achieve up to 26.4° C and on average 13.3° C system temperature reduction with no performance penalty. Second, we utilize this thermal model to guide HLS parameter tuning at the task design stage to achieve improved thermal response after deployment.

97 MATHEMATICS AND COMPUTING↗

One- and two-qubit gate infidelities due to motional errors in trapped ions and electrons

In this work, we derive analytic formulas that determine the effect of error mechanisms on one- and two-qubit gates in trapped ions and electrons. First, we analyze and derive expressions for the effect of driving field inhomogeneities on one-qubit gate fidelities. Second, we derive expressions for two-qubit gate errors, including static motional frequency shifts, trap anharmonicities, field inhomogeneities, heating, and motional dephasing. We show that, for small errors, each of our expressions for infidelity converges to its respective numerical simulation; this shows that our formulas are sufficient for determining error budgets for high-fidelity gates, obviating numerical simulations in future projects. All of the derivations are general to any internal qubit state, and any mixed state of the ion crystal's motion that is diagonal in the Fock state basis. Our treatment of static motional frequency shifts, trap anharmonicities, heating, and motional dephasing apply to both laser-based and laser-free gates, while our treatment of field inhomogeneities applies to laser-free systems.

74 ATOMIC AND MOLECULAR PHYSICS↗

Data-driven, structure-preserving approximations to entropy-based moment closures for kinetic equations

In this study, we present a data-driven approach for approximating entropy-based closures of moment systems from kinetic equations. The proposed closure learns the entropy function by fitting the map between the moments and the entropy of the moment system, and thus does not depend on the spacetime discretization of the moment system or specific problem configurations such as initial and boundary conditions. With convex and C 2 approximations, this data-driven closure inherits several structural properties from entropy-based closures, such as entropy dissipation, hyperbolicity, and H-Theorem. We construct convex approximations to the Maxwell–Boltzmann entropy using convex splines and neural networks, test them on the plane source benchmark problem for linear transport in slab geometry, and compare the results to the standard, entropy-based systems which solve a convex optimization problem to find the closure. Numerical results indicate that these data-driven closures provide accurate solutions in much less computation time than that required by the optimization routine.

97 MATHEMATICS AND COMPUTING↗

Workshop on Addressing Rigor and Reproducibility in Thermal, Heterogeneous Catalysis

Heterogeneous catalysis has long served as the bedrock of the manufacturing of energy carriers, fuels and chemicals, and various technologies for pollution abatement. The significant complexity and variability spanning the entire breadth of catalyst material properties, synthesis methods, characterization techniques, and evaluation procedures, has focused attention on the need to establish community-accepted best practices for ensuring high-quality, benchmarked, and reproducible data. In addition, increased societal urgency to transition to clean energy and reduce greenhouse gas concentrations has incentivized interdisciplinary, convergent, and translational approaches to catalysis research in recent years. Research engineers and scientists with expertise cutting broadly across materials science, chemical synthesis, interfacial science, spectroscopy, and methods of data science and computational simulation, all bring diverse and important perspectives to catalysis research, but often with little awareness of the complexity of catalytic systems, especially in their working environment. As has already occurred in other scientific fields, there has been growing recognition and consensus in the heterogeneous catalysis research community that mechanisms are needed to improve the rigor and reproducibility (R&R) of experimental measurements, to ensure alignment of the broader research community with a common core of best practices specific to the realization of high-quality catalysis research. Similarly, the field is moving rapidly toward computationally informed and data science-driven catalyst design, but the success of implementing such predictive tools hinges on model training and validation rooted in rigorously obtained and reproducible experimental data that are benchmarked to common specifications. As such, this workshop was convened to prepare a report summarizing best practices for reporting data and performing experiments that researchers can use to benchmark, validate, and reproduce data in specific sub-fields of thermal, heterogeneous catalysis. Additionally, we discussed recommendations for future actions that may improve R&R in this field. The workshop organizers and participants include a diverse range of catalysis researchers from various employment sectors (e.g., academia, industry, national laboratory), institutional mission and resources (e.g., PhD-granting research universities, non-PhD-granting teaching universities), career stage (e.g., early, mid and late-career), technical expertise, and demographic background. This diverse group was involved in the discussion of workshop agenda items, writing this report, and discussing possible future action items for the community to consider, which helped ensure that a broad range of perspectives were captured in the description of the problems at hand and the creation of actionable solutions that may be effectively adopted by the diverse practitioners in catalysis research. Importantly, this group of workshop participants also included very early career researchers (e.g., senior PhD students, postdoctoral scholars) who will become the next generation of scientific leaders in various sectors, thus capturing emerging perspectives of newcomers to the field to shape its future while positively impacting the development of its future workforce. We envision that this effort will help advance the field of catalysis science by improving the rigor and reproducibility of experimental data collected by current researchers and future newcomers to the field, which is of broad importance to health and vitality of any scientific discipline. Therefore, best practices identified in this endeavor for thermal heterogeneous catalysis can be translated to such efforts in other areas of catalysis and other scientific fields involving the study of materials, and vice versa. We also envision this to be an ongoing effort, with future workshops that are convened to discuss issues of rigor and reproducibility on technical topics that were unable to be covered in this workshop due to its scope limitations, and as emerging methods and materials become more prevalent in the research community.

36 MATERIALS SCIENCE↗

Protein C-GeM: A Coarse-Grained Electron Model for Fast and Accurate Protein Electrostatics Prediction

The electrostatic potential (ESP) is a powerful property for understanding and predicting electrostatic charge distributions that drive interactions between molecules. In this study, we compare various charge partitioning schemes including fitted charges, density-based quantum mechanical (QM) partitioning schemes, charge equilibration methods, and our recently introduced coarse-grained electron model, C-GeM, to describe the ESP for protein systems. Furthermore, when benchmarked against high quality density functional theory calculations of the ESP for tripeptides and the crambin protein, we find that the C-GeM model is of comparable accuracy to ab initio charge partitioning methods, but with orders of magnitude improvement in computational efficiency since it does not require either the electron density or the electrostatic potential as input.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Overcoming Barriers in Electrochemical Toluene Hydrogenation for Efficient Hydrogen Storage by Pt 3 Au Alloy Catalysts

Hydrogen storage and transportation are essential for the hydrogen economy, and liquid organic hydrogen carriers (LOHCs), such as a toluene/methylcyclohexane (TOL/MCH) system, offer significant advantages in terms of safety and efficiency. However, the electrochemical reduction of TOL to MCH (TER) faces challenges from competing with the hydrogen evolution reaction (HER) and catalyst instability. Here, in this study, Pt 3 Au is introduced as a highly effective catalyst for TER. Through density functional theory screening, we identified distinctive properties of Pt 3 Au, including enhanced binding to the TER intermediates and effective HER suppression. Experimental validation confirmed these computational predictions, with Pt 3 Au achieving the highest reported Faradaic efficiency (98%) in proton exchange membrane systems. Moreover, long-term testing demonstrated that Pt 3 Au maintained Faradaic efficiencies of >90% over 9 h, highlighting its robustness and operational stability. By integrating computational modeling and experimental evaluation, this work addresses key limitations in LOHC catalysis. Pt 3 Au establishes a benchmark for selective and stable TER performance, paving the way for advanced hydrogen storage technologies. These findings emphasize the critical role of rational catalyst design in overcoming the challenges associated with scalable and efficient hydrogen storage solutions.

LOHC↗

Convex Relaxations for Quadratic On/Off Constraints and Applications to Optimal Transmission Switching

This paper studies mixed-integer nonlinear programs featuring disjunctive constraints and trigonometric functions and presents a strengthened version of the convex quadratic relaxation of the optimal transmission switching problem. We first characterize the convex hull of univariate quadratic on/off constraints in the space of original variables using perspective functions. Next, we introduce new tight quadratic relaxations for trigonometric functions featuring variables with asymmetrical bounds. These results are used to further tighten recent convex relaxations introduced for the optimal transmission switching problem in power systems. Using the proposed improvements, along with bound propagation, on 23 medium-sized test cases in the PGLib benchmark library with a relaxation gap of more than 1%, we reduce the gap to less than 1% on five instances. The tightened model has promising computational results when compared with state-of-the-art formulations.

97 MATHEMATICS AND COMPUTING↗