Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evolvable hardware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Enhancing quantum annealing accuracy through replication-based error mitigation *

Abstract Quantum annealers like those manufactured by D-Wave Systems are designed to find high quality solutions to optimization problems that are typically hard for classical computers. They utilize quantum effects like tunneling to evolve toward low-energy states representing solutions to optimization problems. However, their analog nature and limited control functionalities present challenges to correcting or mitigating hardware errors. As quantum computing advances towards applications, effective error suppression is an important research goal. We propose a new approach called replication based mitigation (RBM) based on parallel quantum annealing (QA). In RBM, physical qubits representing the same logical qubit are dispersed across different copies of the problem embedded in the hardware. This mitigates hardware biases, is compatible with limited qubit connectivity in current annealers, and is well-suited for currently available noisy intermediate-scale quantum annealers. Our experimental analysis shows that RBM provides solution quality on par with previous methods while being more flexible and compatible with a wider range of hardware connectivity patterns. In comparisons against standard QA without error mitigation on larger problem instances that could not be handled by previous methods, RBM consistently gets better energies and ground state probabilities across parameterized problem sets.

Djidjev, Hristo N. (ORCID:0000000192868824)↗

Exploring Architectural-Aware Affinity Policies in Modern HPC Runtimes

Modern commodity and High-Performance Computing (HPC) systems are evolving with complex CPU architectures. These architectures now feature higher core and NUMA domain counts and implement features such as hyperthreading. When considering significant differences in hardware configurations, library availability, and hardware-tailored system/software stacks, which could substantially vary from one system to another, performance portability is hard to achieve. Throughout the years, this trend resulted in an increasingly high burden on application developers to fine-tune their workloads for each architecture. This work explores how hardware-dependent aspects such as locality/process/thread affinity affect performance in modern CPU architectures. We focus our study on the Global Memory and Threading (GMT) distributed runtime system as a representative of Partitioned Global Address Space (PGAS) software stacks commonly adopted for productivity. In particular, to appreciate performance implications, we evaluate GMT’s thread affinity policies, and, introduce two new ones which exploit architectural awareness. Finally, we explore alternative NUMA configurations via different process bindings and perform a scalability study on three HPC clusters with varying CPU architectures and NUMA layouts. Our analysis indicates that more complex architectures are more affected by affinity and binding policies and highlights the importance of setting proper runtime configurations to achieve superior performance.

Di Dio Lavore, Ian↗

Modeling and Power-Hardware-in-the-Loop Validation of Synchronous Wind: An Inverterless Grid-Forming Wind Power Plant

Grid-forming (GFM) control of Type-3 and Type-4 wind turbine generators (WTGs) has attracted substantial attention in power systems research; however, the limited overcurrent capability of power electronics converters continues to deteriorate the grid strength of the evolving power systems. Synchronous wind, also known as a Type-5 WTG, offers a unique GFM solution to address grid integration and grid strength issues by keeping the grid largely synchronous at very high integration levels of renewable generation. A Type-5 WTG interfaces with the electric grid via a synchronous generator driven by a variable-speed hydraulic torque converter; hence, the wind rotor operates in variable-speed mode for maximum power generation, and the generator shaft remains synchronous to the grid. This paper develops and tests a high-fidelity model of a Type-5 WTG in a power-hardware-in-the-loop (PHIL) testing environment. The PHIL demonstration shows that a Type-5 WTG inherently behaves as a GFM unit and can obtain similar performance in terms of power responses, wind rotor dynamics, and efficiency compared to a Type-3 WTG in high-wind conditions. The developed model provides further insight into how Type-5 WTGs can benefit the smooth transition to power systems with high integration levels of inverter-based resources.

grid strength↗

Modeling and Power-Hardware-in-the-Loop Validation of Synchronous Wind: An Inverterless Grid-Forming Wind Power Plant: Preprint

Grid-forming (GFM) control of Type-3 and Type-4 wind turbine generators has attracted substantial attention in power systems research; however, the limited over-current capability of power electronics converters continues to deteriorate the grid strength of the evolving power systems. Synchronous wind, also known as Type-5 wind turbine generator (WTG), offers a unique GFM solution to address grid integration and grid strength issues by keeping the grid largely synchronous at very high penetration levels of renewable generation. A Type-5 WTG interfaces to the electric grid via a synchronous generator (SG) driven by a variable-speed hydraulic torque converter; hence, the wind rotor operates in variable-speed mode for maximum power generation and the generator shaft remains synchronous to the grid. This paper developed and tested a high-fidelity model of Type-5 WTG under power-hardware-in-the-loop (PHIL) testing environment. The PHIL demonstration showed that a Type-5 WTGs inherently behaves as a GFM unit and can obtain similar performance in terms of power responses, wind rotor dynamics, and efficiency compared to Type-3 WTG in high wind conditions. The developed model also provides further insight on how Type-5 WTGs can benefit the smooth transition to power systems with high integration level of inverter-based resources.

grid strength↗

DS-TIDE: Harnessing Dynamical Systems for Efficient Time-Independent Differential Equation Solving

Time-Independent Differential Equations (TIDEs) are central to modeling equilibrium behavior across a wide range of scientific and engineering domains, from electrostatics to porous media flow. Conventional numerical solvers offer reliable solutions but incur significant computational costs due to fine-grained discretization and iterative procedures. Machine learning-based approaches address this by replacing iterative solving processes with one-time inference; however, their sophisticated models require extensive training resources that often exceed those of traditional solvers. Consequently, designing a TIDE solver that achieves high accuracy, broad applicability, and exceptional computational efficiency remains a fundamental challenge. In this paper, we propose DS-TIDE, a novel hardware solver that is inspired by, and subsequently leverages, the intrinsic connection between Dynamical Systems (DS) and Differential Equations (DEs) to efficiently and accurately solve TIDEs. DS-TIDE employs a CMOS-compatible DS-based processor, whose physical states evolve under carefully designed DE-driven dynamics and naturally converge to equilibrium -- the solution of the target TIDE -- within ~1µs on a ~1-watt DS-TIDE processor. To enhance expressivity, DS-TIDE incorporates Heterogeneous Dynamics with Temporal Layering (HDTL), which solves TIDEs through a three-stage DS evolution -- conditioning, solving, and decoding -- each governed by specialized dynamics. The entire evolution process is analogous to an infinitely deep neural network temporally unrolled, offering the system the capability of representing complex equations. Furthermore, DS-TIDE is equipped with an on-device DS-DE Auto-Alignment mechanism that dynamically adapts intrinsic hardware dynamics within milliseconds, effectively aligning the system’s dynamics to diverse target DEs. Experimental results across TIDEs from a wide range of scientific and engineering domains demonstrate that DS-TIDE achieves ~10^3× speedup, ~10^5× energy savings, and competitive or superior accuracy compared to state-of-the-art numerical and ML-based solvers.

Liu, Chuan↗

High-fidelity dimer excitations using quantum hardware

The quantum simulation of entangled spin systems can play a central role in quantum magnetic materials discovery. Additionally, the simulation of spectroscopic signatures, such as the dynamical structure factor accessed in inelastic neutron scattering (INS), necessitates a long timescale for circuit evolution. This is because the energy resolution is directly related to the time over which the circuit could be meaningfully evolved. However, canonical Trotterization requires deep circuits precluding such long-time evolution—even for a small number of qubits. Here, in this study, we demonstrate “direct” resource efficient fast-forwarding (REFF) measurements with short-depth circuits that can be used to capture longer time dynamics of spin Hamiltonians. We showcase the results of the dynamics of a quantum spin dimer, the basic quantum unit of emergent many-body spin systems, whose density of states we simulate accurately. The long temporal evolution and measurement of the two-spin correlation functions enable the calculation of the dynamical structure factor S⁡(Q = 0, ω) measured in the neutron scattering cross-section. We exhibit the clarity of the triplet gap and the triplet splitting of the quantum dimer with class-leading fidelity that enables comparison to experimental neutron data. Our results on current circuit hardware outline an important workflow to predict and benchmark against the outputs of INS experiments of quantum magnets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Gauge-fixing quantum density operators at scale

We provide a theory, algorithms, and simulations of nonequilibrium quantum systems using a one-dimensional (1D) completely positive (CP), matrix-product (MP) density-operator (𝜌) representation. By generalizing the matrix product state's orthogonality center, to additionally store positive classical mixture correlations, the MP⁢𝜌 factorization naturally emerges. In this setting, we analytically and numerically examine the virtual gauge freedoms associated with the representation of quantum density operators. Based on this perspective, we simplify algorithms in certain limits to speed up the integration of the canonical-form master-equation dynamics. This enables us to quickly evolve under the dynamics of two-body quantum channels without resorting to optimization-based methods. In addition to this technical advance, we also scale up numerical examples and discuss implications for accurately modeling hardware architectures and predicting their performance in the near term. This includes an example of the quantum to classical transition of informationally leaky, i.e., decohering, qubits. In this setting, because of loss from environmental interactions, nonlocal complex coherence correlations are converted into global incoherent classical statistical mixture correlations. Lastly, the representation of both global and local correlations is discussed. We expect this work to have applications in additional nonequilibrium settings, beyond qubit engineering.

Gangapuram, Amit Jamadagni [Oak Ridge National Lab↗

Quantum Time Dynamics Mediated by the Yang–Baxter Equation and Artificial Neural Networks

Quantum computing shows great potential, but errors pose a significant challenge. This study explores new strategies for mitigating quantum errors using artificial neural networks (ANNs) and the Yang–Baxter equation (YBE). Unlike traditional error mitigation methods, which are computationally intensive, we investigate artificial error mitigation. We developed a novel method that combines ANNs for noise mitigation combined with the YBE to generate noisy data. This approach effectively reduces noise in quantum simulations, enhancing the accuracy of the results. The YBE rigorously preserves quantum correlations and symmetries in spin chain simulations in certain classes of integrable lattice models, enabling effective compression of quantum circuits while retaining linear scalability with the number of qubits. This compression facilitates both full and partial implementations, allowing the generation of noisy quantum data on hardware alongside noiseless simulations using classical platforms. By introducing controlled noise through the YBE, we enhance the data set for error mitigation. We train an ANN model on partial data from quantum simulations, demonstrating its effectiveness in mitigating errors in time-evolving quantum states, providing a scalable framework to enhance quantum computation fidelity, particularly in noisy intermediate-scale quantum (NISQ) systems. We demonstrate the efficacy of this approach by performing quantum time dynamics simulations using the Heisenberg XY Hamiltonian on real quantum devices.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Blueprint for DOE Quantum Supercomputing: Ensuring U.S. Leadership in the Quantum Decade

Quantum computing stands at the threshold of a transformative decade, where the field will evolve from small-scale demonstrations toward practical scientific computing at scale. This Blueprint identifies fault-tolerant quantum computers (FTQCs) as a viable, scalable, and broadly applicable path to achieving “quantum scientific utility,” defined as solving scientifically valuable problems beyond the reach of conventional, classical computers. This capability is expected to show scientific demonstrations in the late 2020s and to mature in the early-to-mid 2030s. This Blueprint outlines a strategy to prepare the U.S. Department of Energy (DOE) for FTQCs and their integration into the U.S. national scientific computing infrastructure. Its purpose is to identify the steps, milestones, and research directions necessary for DOE to enable initial deployment of FTQCs in 2028 as a scientific tool for the nation and mature this capability into the 2030s. DOE has a long history of supporting quantum information science and technology, contributing significantly to research advancements, training a quantum-ready workforce, and providing access to early small-scale quantum hardware. Given recent demonstrations of logical operations on error-corrected logical qubits and the advancement of commercial hardware roadmaps, DOE should begin preparations for large-scale, fault-tolerant quantum computing deployment for DOE science missions. This Blueprint proposes that DOE focus on (1) deploying first-generation scientifically relevant quantum computers with at least 100 logical qubits and performing at least 10,000 to 100,000 hard logical operations in scientifically relevant calculations; (2) developing essential FTQC programming competencies, system software, and facility readiness; and (3) investing in cutting edge focused R&D that fosters breakthroughs in scientific applications, algorithms, and logical architectures needed to accelerate the advent of scientific utility. This effort will position DOE to transition to larger systems: production-scale quantum computers that comprise 1,000 to 10,000 logical qubits, perform 1 to 10 billion hard logical operations, and execute scientifically useful computations at scale. Achieving these goals will require DOE facilities to evolve with urgency to support scientific campaigns that integrate quantum and classical computing resources into efficient workflows, novel software and firmware environments for compiling and routing quantum programs on FTQC machines, and suitable infrastructure for quantum hardware. It will also require further development and optimization of scientific applications from the fields of materials science, quantum chemistry, and high-energy and nuclear physics. The Blueprint calls for transformative R&D and collective action to accelerate the advent of scientific quantum utility and bring it within reach by 2028.

97 MATHEMATICS AND COMPUTING↗

Sensitive dependence on initial conditions in a formation of magnetic vortices

The magnetic vortex exhibits promise as a true random number generator for hardware-based encryption and probabilistic computing due to its stochastic formation of energetically equivalent fourfold degenerate states, characterized by two topologies: polarity and chirality. However, a comprehensive understanding of the stochastic formation of magnetic vortices remains elusive. In this work, we show that the magnetization relaxation in asymmetric Permalloy disks evolves along a pitchfork bifurcation, with both bifurcation paths leading to the formation of magnetic vortices with the same chirality. In the bifurcation, one formation path is always chosen under weak in-plane magnetic fields, ultimately determining the final magnetic vortex state. By delaying the in-plane magnetic field, we quantitatively investigate when the final vortex state is determined and find that it is closely associated with the initial conditions rather than the bifurcation point itself. Our findings provide valuable insights into future spintronic-based encryption and probabilistic computing.

Jeong, Suyeong↗

Shifting Between Compute and Memory Bounds: A Compression-Enabled Roofline Model

In the evolving landscape of high-performance computing, especially to fight the end of Moore’s Law and Dennard’s Scaling, the ability to shift between compute-bound and memory-bound states is critical for enhancing adaptability and flexibility to diverse system and domain-specific architectures. Such capability is vital for optimizing performance across distinguished hardware configurations, such as accelerators, memory hierarchies, and cache systems. Despite that ad hoc optimization techniques, such as compressed/approximate computation, have been enabled for compute-/data-intensive computing for improved performance in distinct hardware settings, there lacks an understanding of 1) the rational behind performance improvement; 2) capability of different optimizations; 3) what optimization to respond to specific computational and memory demands. This work proposes a compression-enabled roofline model to facilitate this adaptability with data compression techniques to balance and transform between computational and memory demands. This model enables applications to adjust in response to the specific strengths and limitations of the underlying hardware and system to optimize resource utilization. The effectiveness of this approach is demonstrated with matrix multiplication kernels on different input sizes, with turning on/off various compression techniques, including 1) low-precision floating point; 2) sparse matrix formulation; and 3) compressed arrays with ZFP. By reducing memory transfer volumes and cache misses and increasing data locality and computational intensity through compression, the specific roofline model can transform between compute and memory bounds to align more efficiently with system capabilities. This advancement not only improves overall performance but also maximizes adaptability in diverse computing environments.

Naraparaju, Ramasoumya [University of Washington]↗

Accelerating high-order continuum kinetic plasma simulations using multiple GPUs

Kinetic plasma simulations solve the Vlasov-Poisson or Vlasov-Maxwell equations to evolve scalar-variable distribution functions in position-velocity phase space and vector-variable electromagnetic fields in configuration space. The immense computational cost of evolving high-dimensional variables, and their large number of degrees of freedom, often limits the utility of continuum kinetic simulations and presents a challenge when it comes to accurately simulating real-world physical phenomena. To address this challenge, we present techniques that accelerate and minimize the computational work required for a scalable Vlasov-Poisson solver. We show theoretical hardware compute and communication bounds for solving a fourth-order finite-volume Vlasov-Poisson system. These bounds are then used to inform and evaluate the design of performance portable algorithms for a multiple graphics processing unit (GPU) accelerated version of the Vlasov-Poisson solver VCK-CPU [1]. We demonstrate that the multi-GPU Vlasov solver implementation, VCK-GPU, simultaneously minimizes required inter-process data transfer while also being bounded by the machine network performance limits. This results in an overall strong scaling speedup per timestep of up to 40x in three-dimensional phase space (one position, two velocity coordinates) and 54x in four dimensional phase space (two position, two velocity coordinates) and a 341x increase in simulation throughput of the GPU accelerated code over the existing CPU code. The GPU code is also able to weak scale up to 256 compute nodes and 1024 GPUs. In conclusion, we demonstrate that the improved compute performance enables exploring configurations which were previously computationally infeasible, including resolving fine-scale distribution function filamentation and multi-species dynamics with realistic electron-proton mass ratios.

Continuum kinetics↗

How to Build a Quantum Supercomputer: Scaling from Hundreds to Millions of Qubits

In the span of four decades, quantum computation has evolved from an intellectual curiosity to a potentially realizable technology. Today, small-scale demonstrations have become possible for quantum algorithmic primitives on hundreds of physical qubits and proof-of-principle error-correction on a single logical qubit. Nevertheless, despite significant progress and excitement, the path toward a full-stack scalable technology is largely unknown. There are significant outstanding quantum hardware, fabrication, software architecture, and algorithmic challenges that are either unresolved or overlooked. These issues could seriously undermine the arrival of utility-scale quantum computers for the foreseeable future. Here, we provide a comprehensive review of these scaling challenges. We show how the road to scaling could be paved by adopting existing semiconductor technology to build much higher-quality qubits, employing system engineering approaches, and performing distributed quantum computation within heterogeneous high-performance computing infrastructures. These opportunities for research and development could unlock certain promising applications, in particular, efficient quantum simulation/learning of quantum data generated by natural or engineered quantum systems. To estimate the true cost of such promises, we provide a detailed resource and sensitivity analysis for classically hard quantum chemistry calculations on surface-code error-corrected quantum computers given current, target, and desired hardware specifications based on superconducting qubits, accounting for a realistic distribution of errors. Furthermore, we argue that, to tackle industry-scale classical optimization and machine learning problems in a cost-effective manner, heterogeneous quantum-probabilistic computing with custom-designed accelerators should be considered as a complementary path toward scalability.

Mohseni, Masoud↗

AthenaK: A Performance-portable Version of the Athena++ Adaptive Mesh Refinement Framework

We describe AthenaK: a new implementation of the Athena++ block-based adaptive mesh refinement framework using the Kokkos programming model. Finite volume methods for Newtonian, special relativistic, and general relativistic (GR) hydrodynamics and magnetohydrodynamics (MHD), and GR-radiation hydrodynamics and MHD, as well as a module for evolving Lagrangian tracer or charged test particles (e.g., cosmic rays) are implemented using the framework. In two companion papers, we describe (1) a new solver for the Einstein equations based on the Z4c formalism, and (2) a GRMHD solver in dynamical spacetimes also implemented using the framework, enabling new applications in numerical relativity. By adopting Kokkos, the code can be run on virtually any hardware, including CPUs, GPUs from multiple vendors, and emerging Advanced RISC Machine processors. AthenaK shows excellent performance and weak scaling, achieving over 1 billion cell updates per second for hydrodynamics in three dimensions on a single NVIDIA Grace Hopper processor. It does this with a typical parallel efficiency of 80% on 65,536 AMD GPUs on the OLCF Frontier system. Such performance portability enables AthenaK to leverage modern exascale computing systems for challenging applications in astrophysical fluid dynamics, numerical relativity, and multimessenger astrophysics.

79 ASTRONOMY AND ASTROPHYSICS↗

A Brief Survey on High Performance Computing Systems Power Management

This paper provides a survey of software-based power management techniques in High Performance Computing (HPC) systems. Seven existing power management and monitoring tools and frameworks are discussed. These are: Variorum, dynamic energy-performance optimizer (DEPO), Powersched, Bull Dynamic Power Optimizer (BDPO), Energy Aware Runtime (EAR), Global Extensible Open Power Manager (GEOPM), and PoLiMEr. Each of these tools is evaluated based on hardware abstraction, optimization methods, usability, and experimental validation. This survey highlights the diversity of approaches in managing energy efficiency, from vendor-neutral APIs to algorithm-driven power capping, and dynamic frequency adjustments. Given that energy requirements for large computational systems is increasing quickly, the importance of integrating these tools into existing HPC environments and the need for further research in this rapidly evolving field is also discussed.

97 - MATHEMATICS AND COMPUTING↗

Transmission electron microscopy with in-situ ion irradiation: Facilities and community

Whilst there is a clear scientific and technological need for the technical capabilities of transmission electron microscopes with in-situ ion irradiation, it also requires a collaborative community of international researchers to support such facilities in successfully meeting this demand. Instruments of this type serve to provide fundamental understanding of the mechanisms which drive changes in materials important to nuclear fission and fusion energy, the semiconductor industry, quantum information systems, space travel, astronomy, geology and many more applications. As these areas continue to evolve and the instrumentation possibilities expand, the capacity of in-situ ion irradiation facilities must also develop hand-in-hand with the user community to deliver an ever-greater diversity of high-fidelity extreme-environment experimentation. Future directions for the field, such as miniaturization from MEMS/microfluidic devices and advanced controls with ML-based analysis, continuously emerge to advance both the hardware and software which support the coupling of TEMs with ion beams. This review sets out to provide up-to-date insights into the community and advancement of current, and development of future, facilities which have the potential to further unlock access to the nanoscale exploration of coupled extreme environments crucial to many of the important science and engineering challenges we face today.

In-situ irradiation↗

Empowering Scientific Innovation Through An Integrated Research Infrastructure: The Role of the Advanced Computing Ecosystem

As the landscape of computational science evolves, the Department of Energy (DOE) is reimagining the roles of its large-scale computing facilities to meet emerging research challenges. The Integrated Research Infrastructure (IRI) program aims to transform how experiments are designed, conducted, and shared, with significant impacts on all stakeholders. In response, the Oak Ridge Leadership Computing Facility (OLCF) has established the Advanced Computing Ecosystem (ACE), a strategic framework to prepare its hardware, software, and experimental capabilities for the IRI era. ACE focuses on integrating novel compute environments, orchestrating advanced workflows, and developing foundational technologies, ensuring a seamless transition to IRI while accelerating scientific discovery. This paper outlines ACE's role in advancing OLCF's mission and its impact on the future of computational science.

Widener, Patrick↗

An end-to-end workflow for executing a classically bootstrapped variational quantum algorithm on an academic quantum computer

Academic quantum computing platforms often face unique challenges in executing quantum workloads due to fragmented software environments and limited engineering support. Unlike commercial ecosystems, academic devices typically evolve without full-stack integration in mind, making it difficult to run complex applications—such as variational quantum algorithms (VQA)—reliably and efficiently. Issues such as incompatible software layers and lack of automated job management significantly increase the overhead of theory-experiment collaboration. To address these challenges, we develop a modular, end-to-end workflow that decouples application-layer code from low-level hardware control, automates circuit submission and result collection, and supports fine-grained circuit-level job scheduling and recovery. The architecture employs a dual-end application programming interface (API) design, enabling robust operation across unstable or resource-constrained hardware backends. For practical use, the framework is lightweight and user-friendly, allowing rapid prototyping of full-stack workflows using basic Python tools. We validate this workflow on a high-fidelity trapped-ion quantum computer by demonstrating a variational quantum eigensolver (VQE) experiment with a classically bootstrapped ansatz initialization technique. The system successfully executed over 60,000 circuits across multiple molecular test cases with minimal human intervention, highlighting the framework’s effectiveness in enabling reproducible, resilient quantum experimentation in academic settings.

Clifford↗