Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Exploring temporal community evolution: algorithmic approaches and parallel optimization for dynamic community detection

Abstract Dynamic (temporal) graphs are a convenient mathematical abstraction for many practical complex systems including social contacts, business transactions, and computer communications. Community discovery is an extensively used graph analysis kernel with rich literature for static graphs. However, community discovery in a dynamic setting is challenging for two specific reasons. Firstly, the notion of temporal community lacks a widely accepted formalization, and only limited work exists on understanding how communities emerge over time. Secondly, the added temporal dimension along with the sheer size of modern graph data necessitates new scalable algorithms. In this paper, we investigate how communities evolve over time based on several graph metrics under a temporal formalization. We compare six different algorithmic approaches for dynamic community detection for their quality and runtime. We identify that a vertex-centric (local) optimization method works as efficiently as the classical modularity-based methods. To its advantage, such local computation allows for the efficient design of parallel algorithms without incurring a significant parallel overhead. Based on this insight, we design a shared-memory parallel algorithm DyComPar , which demonstrates between 4 and 18 fold speed-up on a multi-core machine with 20 threads, for several real-world and synthetic graphs from different domains.

97 MATHEMATICS AND COMPUTING↗

High Performance Adaptive Physics Refinement to Enable Large-Scale Tracking of Cancer Cell Trajectory

The ability to track simulated cancer cells through the circulatory system, important for developing a mechanistic understanding of metastatic spread, pushes the limits of today's supercomputers by requiring the simulation of large fluid volumes at cellular-scale resolution. To overcome this challenge, we introduce a new adaptive physics refinement (APR) method that captures cellular-scale interaction across large domains and leverages a hybrid CPU-GPU approach to maximize performance. Through algorithmic advances that integrate multi-physics and multi-resolution models, we establish a finely resolved window with explicitly modeled cells coupled to a coarsely resolved bulk fluid domain. In this work we present multiple validations of the APR framework by comparing against fully resolved fluid-structure interaction methods and employ techniques, such as latency hiding and maximizing memory bandwidth, to effectively utilize heterogeneous node architectures. Collectively, these computational developments and performance optimizations provide a robust and scalable framework to enable system-level simulations of cancer cell transport.

59 BASIC BIOLOGICAL SCIENCES↗

Sheath transitions in a cylindrical filament discharge: Axisymmetric 1D3V PIC-MCC simulations

We present the first nonplanar hot cathode discharge simulations that capture the role of the trapped-ions plasma, elucidating new phenomena unobservable in planar geometric discharges. A discharge struck between a single emitting wire filament cathode and a bounding anode is simulated in cylindrical geometry using an axisymmetric (radial) particle-in-cell Monte-Carlo collisions code. Operating the discharge near its ionization energy threshold can lead to the formation of a two plasma mode (TPM). One plasma forms in the conventional upstream region through electron impact ionization of background neutrals. A second plasma, whose global effect on the discharge was not previously well understood, forms downstream through the trapping of cold ions in the potential well of the filament’s virtual cathode, a process enabled by ion-neutral charge exchange collisions. Three space charge regions intersperse the electrode gap—an emissive sheath between the cathode filament and trapped-ions plasma, a double layer between the two plasmas, and a classical sheath between the upstream plasma and the outer anode. Simulations exhibit mode transitions and quenching instabilities that transform the discharge between the TPM and other single-plasma sheath modes that include classical (temperature-limited), space charge limited, and inverse (anode glow) modes. The transitions are explained via “aid-and-compete” dynamics wherein the growth of one plasma enhances growth in the other while concurrently exhibiting expansion dynamics antagonistic to each other. The system exhibits strong hysteresis memory during the mode transitions. Improved understanding and control of these sheath mode transitions are expected to benefit plasma applications with hot cathodes.

Electrical hysteresis↗

Fast and Flexible Inference Framework for Continuum Reverberation Mapping Using Simulation-based Inference with Deep Learning

Continuum reverberation mapping (CRM) of active galactic nuclei (AGN) monitors multiwavelength variability signatures to constrain accretion disk structure and supermassive black hole (SMBH) properties. The upcoming Vera Rubin Observatory’s Legacy Survey of Space and Time will survey tens of millions of AGN over the next decade, with thousands of AGN monitored with almost daily cadence in the deep drilling fields. However, existing CRM methodologies often require long computation time and are not designed to handle such large amounts of data. In this paper, we present a fast and flexible inference framework for CRM using simulation-based inference (SBI) with deep learning to estimate SMBH properties from AGN light curves. We use a long short-term memory summary network to reduce the high dimensionality of the light curve data and then use a neural density estimator to estimate the posterior of SMBH parameters. Using simulated light curves, we find SBI can produce more accurate SMBH parameter estimation with 10 3 –10 5 times speed up in inference efficiency compared to traditional methods. The SBI framework is particularly suitable for wide-field CRM surveys as the light curves will have identical observing patterns, which can be incorporated into the SBI simulation. We explore the performance of our SBI model on light curves with irregular-sampled, realistic observing cadence and alternative variability characteristics to demonstrate the flexibility and limitation of the SBI framework.

79 ASTRONOMY AND ASTROPHYSICS↗

Report on Alternative Devices to Pyrotechnics on Spacecraft

Pyrotechnics accomplish many functions on today's spacecraft, possessing minimum volume/weight, providing instantaneous operation on demand, and requiring little input energy. However, functional shock, safety, and overall system cost issues, combined with emergence and availability of new technologies question their continued use on space missions. Upon request from the National Aeronautics and Space Administration's (NASA) Program Management Council (PMC), Langley Research Center (LaRC) conducted a survey to identify and evaluate state-of-the-art non-explosively actuated (NEA) alternatives to pyrotechnics, identify NEA devices planned for NASA use, and investigate potential interagency cooperative efforts. In this study, over 135 organizations were contacted, including NASA field centers, Department of Defense (DOD) and other government laboratories, universities, and American and European industrial sources resulting in further detailed discussions with over half, and 18 face-to-face briefings. Unlike their single use pyrotechnic predecessors, NEA mechanisms are typically reusable or refurbishable, allowing flight of actual tested units. NEAs surveyed include spool-based devices, thermal knife, Fast Acting Shockless Separation Nut (FASSN), paraffin actuators, and shape memory alloy (SMA) devices (e.g., Frangibolt). The electro-mechanical spool, paraffin actuator and thermal knife are mature, flight proven technologies, while SMA devices have a limited flight history. There is a relationship between shock, input energy requirements, and mechanism functioning rate. Some devices (e.g., Frangibolt and spool based mechanisms) produce significant levels of functional shock. Paraffin, thermal knife, and SMA devices can provide gentle, shock-free release but cannot perform critically timed, simultaneous functions. The FASSN flywheel-nut release device possesses significant potential for reducing functional shock while activating nearly instantaneously. Specific study recommendations include: (1) development of NEA standards, specifically in areas of material characterization, functioning rates, and test methods; (2) a systems level approach to assure successful NEA technology application; and (3) further investigations into user needs, along with industry/government system-level real spacecraft cost benefit trade studies to determine NEA application foci and performance requirements. Additional survey observations reveal an industry and government desire to establish partnerships to investigate remaining unknowns and formulate NEA standards, specifically those driven by SMAs. Finally, there is increased interest and need to investigate alternative devices for such functions as stage/shroud separation and high pressure valving. This paper summarizes results of the NASA-LaRC survey of pyrotechnic alternatives. State of-the-art devices with their associated weight and cost savings are presented. Additionally, a comparison of functional shock characteristics of several devices are shown, and potentially related technology developments are highlighted.

Lucy, M. H.↗

Computationally Efficient Multiscale Neural Networks Applied to Fluid Flow in Complex 3D Porous Media

Abstract The permeability of complex porous materials is of interest to many engineering disciplines. This quantity can be obtained via direct flow simulation, which provides the most accurate results, but is very computationally expensive. In particular, the simulation convergence time scales poorly as the simulation domains become less porous or more heterogeneous. Semi-analytical models that rely on averaged structural properties (i.e., porosity and tortuosity) have been proposed, but these features only partly summarize the domain, resulting in limited applicability. On the other hand, data-driven machine learning approaches have shown great promise for building more general models by virtue of accounting for the spatial arrangement of the domains’ solid boundaries. However, prior approaches building on the convolutional neural network (ConvNet) literature concerning 2D image recognition problems do not scale well to the large 3D domains required to obtain a representative elementary volume (REV). As such, most prior work focused on homogeneous samples, where a small REV entails that the global nature of fluid flow could be mostly neglected, and accordingly, the memory bottleneck of addressing 3D domains with ConvNets was side-stepped. Therefore, important geometries such as fractures and vuggy domains could not be modeled properly. In this work, we address this limitation with a general multiscale deep learning model that is able to learn from porous media simulation data. By using a coupled set of neural networks that view the domain on different scales, we enable the evaluation of large ( $$>512^3$$ > 512 3 ) images in approximately one second on a single graphics processing unit. This model architecture opens up the possibility of modeling domain sizes that would not be feasible using traditional direct simulation tools on a desktop computer. We validate our method with a laminar fluid flow case using vuggy samples and fractures. As a result of viewing the entire domain at once, our model is able to perform accurate prediction on domains exhibiting a large degree of heterogeneity. We expect the methodology to be applicable to many other transport problems where complex geometries play a central role.

36 MATERIALS SCIENCE↗

Board Level Proton Testing Book of Knowledge for NASA Electronic Parts and Packaging Program

This book of knowledge (BoK) provides a critical review of the benefits and difficulties associated with using proton irradiation as a means of exploring the radiation hardness of commercial-off-the-shelf (COTS) systems. This work was developed for the NASA Electronic Parts and Packaging (NEPP) Board Level Testing for the COTS task. The fundamental findings of this BoK are the following. The board-level test method can reduce the worst case estimate for a board's single-event effect (SEE) sensitivity compared to the case of no test data, but only by a factor of ten. The estimated worst case rate of failure for untested boards is about 0.1 SEE/board-day. By employing the use of protons with energies near or above 200 MeV, this rate can be safely reduced to 0.01 SEE/board-day, with only those SEEs with deep charge collection mechanisms rising this high. For general SEEs, such as static random-access memory (SRAM) upsets, single-event transients (SETs), single-event gate ruptures (SEGRs), and similar cases where the relevant charge collection depth is less than 10 μm, the worst case rate for SEE is below 0.001 SEE/board-day. Note that these bounds assume that no SEEs are observed during testing. When SEEs are observed during testing, the board-level test method can establish a reliable event rate in some orbits, though all established rates will be at or above 0.001 SEE/board-day. The board-level test approach we explore has picked up support as a radiation hardness assurance technique over the last twenty years. The approach originally was used to provide a very limited verification of the suitability of low cost assemblies to be used in the very benign environment of the International Space Station (ISS), in limited reliability applications. Recently the method has been gaining popularity as a way to establish a minimum level of SEE performance of systems that require somewhat higher reliability performance than previous applications. This sort of application of the method suggests a critical analysis of the method is in order. This is also of current consideration because the primary facility used for this type of work, the Indiana University Cyclotron Facility (IUCF) (also known as the Integrated Science and Technology (ISAT) hall), has closed permanently, and the future selection of alternate test facilities is critically important. This document reviews the main theoretical work on proton testing of assemblies over the last twenty years. It augments this with review of reported data generated from the method and other data that applies to the limitations of the proton board-level test approach. When protons are incident on a system for test they can produce spallation reactions. From these reactions, secondary particles with linear energy transfers (LETs) significantly higher than the incident protons can be produced. These secondary particles, together with the protons, can simulate a subset of the space environment for particles capable of inducing single event effects (SEEs). The proton board-level test approach has been used to bound SEE rates, establishing a maximum possible SEE rate that a test article may exhibit in space. This bound is not particularly useful in many cases because the bound is quite loose. We discuss the established limit that the proton board-level test approach leaves us with. The remaining possible SEE rates may be as high as one per ten years for most devices. The situation is actually more problematic for many SEE types with deep charge collection. In cases with these SEEs, the limits set by the proton board-level test can be on the order of one per 100 days. Because of the limited nature of the bounds established by proton testing alone, it is possible that tested devices will have actual SEE sensitivity that is very low (e.g., fewer than one event in 1 × 10(exp 4) years), but the test method will only be able to establish the limits indicated above. This BoK further examines other benefits of proton board-level testing besides hardness assurance. The primary alternate use is the injection of errors. Error injection, or fault injection, is something that is often done in a simulation environment. But the proton beam has the benefit of injecting the majority of actual SEEs without risk of something being missed, and without the risk of simulation artifacts misleading the SEE investigation.

Guertin, Steven M.↗

Demonstration of Automatically-Generated Adjoint Code for Use in Aerodynamic Shape Optimization

Gradient-based optimization requires accurate derivatives of the objective function and constraints. These gradients may have previously been obtained by manual differentiation of analysis codes, symbolic manipulators, finite-difference approximations, or existing automatic differentiation (AD) tools such as ADIFOR (Automatic Differentiation in FORTRAN). Each of these methods has certain deficiencies, particularly when applied to complex, coupled analyses with many design variables. Recently, a new AD tool called ADJIFOR (Automatic Adjoint Generation in FORTRAN), based upon ADIFOR, was developed and demonstrated. Whereas ADIFOR implements forward-mode (direct) differentiation throughout an analysis program to obtain exact derivatives via the chain rule of calculus, ADJIFOR implements the reverse-mode counterpart of the chain rule to obtain exact adjoint form derivatives from FORTRAN code. Automatically-generated adjoint versions of the widely-used CFL3D computational fluid dynamics (CFD) code and an algebraic wing grid generation code were obtained with just a few hours processing time using the ADJIFOR tool. The codes were verified for accuracy and were shown to compute the exact gradient of the wing lift-to-drag ratio, with respect to any number of shape parameters, in about the time required for 7 to 20 function evaluations. The codes have now been executed on various computers with typical memory and disk space for problems with up to 129 x 65 x 33 grid points, and for hundreds to thousands of independent variables. These adjoint codes are now used in a gradient-based aerodynamic shape optimization problem for a swept, tapered wing. For each design iteration, the optimization package constructs an approximate, linear optimization problem, based upon the current objective function, constraints, and gradient values. The optimizer subroutines are called within a design loop employing the approximate linear problem until an optimum shape is found, the design loop limit is reached, or no further design improvement is possible due to active design variable bounds and/or constraints. The resulting shape parameters are then used by the grid generation code to define a new wing surface and computational grid. The lift-to-drag ratio and its gradient are computed for the new design by the automatically-generated adjoint codes. Several optimization iterations may be required to find an optimum wing shape. Results from two sample cases will be discussed. The reader should note that this work primarily represents a demonstration of use of automatically- generated adjoint code within an aerodynamic shape optimization. As such, little significance is placed upon the actual optimization results, relative to the method for obtaining the results.

Green, Lawrence↗

Parallel Computation of the Jacobian Matrix for Nonlinear Equation Solvers Using MATLAB

Demonstrating speedup for parallel code on a multicore shared memory PC can be challenging in MATLAB due to underlying parallel operations that are often opaque to the user. This can limit potential for improvement of serial code even for the so-called embarrassingly parallel applications. One such application is the computation of the Jacobian matrix inherent to most nonlinear equation solvers. Computation of this matrix represents the primary bottleneck in nonlinear solver speed such that commercial finite element (FE) and multi-body-dynamic (MBD) codes attempt to minimize computations. A timing study using MATLAB's Parallel Computing Toolbox was performed for numerical computation of the Jacobian. Several approaches for implementing parallel code were investigated while only the single program multiple data (spmd) method using composite objects provided positive results. Parallel code speedup is demonstrated but the goal of linear speedup through the addition of processors was not achieved due to PC architecture.

Rose, Geoffrey K.↗

Hierarchical Epoxy Structures via Tunable Polymerization-Induced Phase Separation Combined with Additive Manufacturing

Polymerization-induced phase separation (PIPS) allows for the control of thermoset morphologies and properties, enabling the tuning of domain sizes and thermomechanical response. However, its use in generating substructural features in additively manufactured materials has been limited. In this work, we combine epoxy PIPS with UV curable acrylate and rheological modifiers to print nano- to macro-phase separating materials via a two-step, dual-cure approach. This method enables direct ink write printing of hierarchical structures with both controlled morphologies through phase separation and macroscale architecture through print design. We find that formulations for phase-separating materials require judicious incorporation of additives to enable printability and to provide sufficient green strength. Atomic force microscopy-nano infrared mapping reveals tunable, reticulated nano- to micron-scale domains of the resultant multiphase materials and their morphology changes due to additives, resulting in alterations to thermomechanical and tensile properties. Shape memory behavior is also demonstrated through multimaterial additive manufacturing of epoxies with functionally graded internal morphology using active mixing techniques, highlighting this method’s ability to fabricate complex architectures with controlled morphologies and thermomechanical response.

Van Meter, Kylie E [Organic Materials Science, San↗

Dimensionality reduction of the many-body problem using coupled-cluster subsystem flow equations: classical and quantum computing perspective

We discuss reduced-scaling strategies employing recently introduced sub-system embedding sub-algebras coupled-cluster formalism (SES-CC) to describe many-body systems. These strategies utilize properties of the SES-CC formulations where the equations describing certain classes of sub- systems can be integrated into a computational flows composed coupled eigenvalue problems of reduced dimensionality. Additionally, these flows can be defined at the level of the CC Ansatz defined by selected classes of cluster amplitudes, which define the wave function ”memory” of possible partitionings of the many-body system into constituent sub-systems. One of the possible ways of solving these coupled problems is through implementing procedures, where the information is passed between the sub-systems in a self-consistent manner. As a special case, we consider local flow formulations where the so-called local character of correlation effects can be closely related to properties of sub-system embedding sub-algebras employing localized molecular basis. We also generalize flow equations to the time domain and to downfolding methods utilizing double exponential unitary CC Ansatz (DUCC), where reduced dimensionality of constituent sub-problems offer a possibility of efficient utilization of limited quantum resources in modeling realistic systems.

Electron correlation, quantum chemistry, quantum c↗

Surviving and Thriving in Space and on Earth's Oceans, Human Logistics and Sustainability: Comparisons and Considerations

Ocean exploration sailing journeys from hundreds of years ago typically required large vessels and large crews (in comparison with today’s space capsules) to travel between the continents and around the world. Modern sailors of today are able to complete similar distant voyages, in small vessels, with a minimal crew, comparable in size to modern space travel crews. This paper uses a systems engineering approach (e.g. using the NASA Human Integration Design Handbook (HIDH), NASA-SP-2010-3407, 2010 and the “Advanced Life Support Baseline Values and Assumptions Document, (BVAD)” NASA-CR-2004-208941, 2004.), to examine and compare the logistics and sustainability aspects of a small crew traveling on Earth's oceans in sailing vessels versus humans traveling in space. The “Mālama Honua Worldwide Voyage” of the Hōkūleʻa, a replica of an ancient Hawaiian double hulled sailing canoe, will be used as a baseline minimalist case study. This is a good comparison case since the Polynesian exploration of the vast (and virtually empty) Pacific Ocean with limited resources is an analogue to human space travel. A modern sailboat is compared to the ancient Polynesian methods and then a space craft is assessed with similar functional decomposition methods. In 1992 during his second Space Shuttle mission (STS-52, Columbia) Astronaut Lacy Veach received a radio message from a student: "What are the similarities and differences between canoe and space travel?" Astronaut Charles Lacy Veach answered, "Both are voyages of exploration. Hōkūle‘a is in the past, Columbia is in the future." Navigator Nainoa Thompson added from the sailing canoe, "Columbia is the highest achievement of modern technology today, a voyaging canoe was the highest achievement of technology in its day." This paper is dedicated to the memory of two great Hawaiian astronauts: US Air Force Colonel Charles Lacy Veach and US Air Force Colonel Ellison Onizuka and to legendary waterman and Hōkūleʻa crew member Eddie Aikau who was lost at sea in 1978, at the beginning of a 30-day, 2,500-mile (4,000km) journey by the Hōkūleʻa to follow the ancient route of the Polynesian migration between the Hawaiian and Tahitian island chains.

Robert P Mueller↗

A Massively Parallel Implementation of the CCSD(T) Method Using the Resolution-of-the-Identity Approximation and a Hybrid Distributed/Shared Memory Parallelization Model

In this work, a parallel algorithm is described for the coupled-cluster singles and doubles method augmented with a perturbative correction for triple excitations [CCSD(T)] using the resolution-of-the-identity (RI) approximation for two-electron repulsion integrals (ERIs). The algorithm bypasses the storage of four-center ERIs by adopting an integral-direct strategy. The CCSD amplitude equations are given in a compact quasi-linear form by factorizing them in terms of amplitude-dressed three-center intermediates. A hybrid MPI/OpenMP parallelization scheme is employed, which uses the OpenMP-based shared memory model for intranode parallelization and the MPI-based distributed memory model for internode parallelization. Parallel efficiency has been optimized for all terms in the CCSD amplitude equations. Two different algorithms have been implemented for the rate-limiting terms in the CCSD amplitude equations that entail and -scaling computational costs, where N O and N V denote the number of correlated occupied and virtual orbitals, respectively. One of the algorithms assembles the four-center ERIs requiring N V 4 and N O 2 N V 2 -scaling memory costs in a distributed manner on a number of MPI ranks, while the other algorithm completely bypasses the assembling of quartic memory-scaling ERIs and thus largely reduces the memory demand. It is demonstrated that the former memory-expensive algorithm is faster on a few hundred cores, while the latter memory-economic algorithm shows a better strong scaling in the limit of a few thousand cores. The program is shown to exhibit a near-linear scaling, in particular for the compute-intensive triples correction step, on up to 8000 cores. The performance of the program is demonstrated via calculations involving molecules with 24–51 atoms and up to 1624 atomic basis functions. As the first application, the complete basis set (CBS) limit for the interaction energy of the π-stacked uracil dimer from the S66 data set has been investigated. This work reports the first calculation of the interaction energy at the CCSD(T)/aug-cc-pVQZ level without local orbital approximation. The CBS limit for the CCSD correlation contribution to the interaction energy was found to be -8.01 kcal/mol, which agrees very well with the value -7.99 kcal/mol reported by Schmitz, Hättig, and Tew [ Phys. Chem. Chem. Phys. 2014 , 16 , 22167-22178]. The CBS limit for the total interaction energy was estimated to be -9.64 kcal/mol.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Stochastic Vector Techniques in Ground-State Electronic Structure

Herein we review a suite of stochastic vector computational approaches for studying the electronic structure of extended condensed matter systems. These techniques help reduce algorithmic complexity, facilitate efficient parallelization, simplify computational tasks, accelerate calculations, and diminish memory requirements. While their scope is vast, we limit our study to ground-state and finite temperature density functional theory (DFT) and second-order many-body perturbation theory. More advanced topics, such as quasiparticle (charge) and optical (neutral) excitations and higher-order processes, are covered elsewhere. We start by explaining how to use stochastic vectors in computations, characterizing the associated statistical errors. Next, we show how to estimate the electron density in DFT and discuss effective techniques to reduce statistical errors. Finally, we review the use of stochastic vectors for calculating correlation energies within the second-order Møller-Plesset perturbation theory and its finite temperature variational form. Example calculation results are presented and used to demonstrate the efficacy of the methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Simultaneous measurement of the exchange parameter and saturation magnetization using propagating spin waves

The exchange interaction in ferromagnetic ultra thin films is a critical parameter in magnetization-based storage and logic devices, yet the accurate measurement of it remains a challenge. While a variety of approaches are currently used to determine the exchange parameter, each has its limitations, and good agreement among them has not been achieved. To date, neutron scattering, magnetometry, Brillouin light scattering, spin-torque ferromagnetic resonance spectroscopy, and Kerr microscopy have all been used to determine the exchange parameter. Here, we present a method that exploits the wavevector selectivity of Brillouin light scattering to measure the spin wave dispersion in both the backward volume and Damon–Eshbach orientations. The exchange, saturation magnetization, and magnetic thickness are then determined by a simultaneous fit of both dispersion branches with general spin wave theory without any prior knowledge of the thickness of a magnetic “dead layer.” In this study, we demonstrate the strength of this technique for ultrathin metallic films, typical of those commonly used in industrial applications for magnetic random-access memory.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Entangling Quantum Generative Adversarial Networks

Generative adversarial networks (GANs) are one of the most widely adopted machine learning methods for data generation. In this work, we propose a new type of architecture for quantum generative adversarial networks (an entangling quantum GAN, EQ-GAN) that overcomes limitations of previously proposed quantum GANs. Leveraging the entangling power of quantum circuits, the EQ-GAN converges to the Nash equilibrium by performing entangling operations between both the generator output and true quantum data. In the first multiqubit experimental demonstration of a fully quantum GAN with a provably optimal Nash equilibrium, we use the EQ-GAN on a Google Sycamore superconducting quantum processor to mitigate uncharacterized errors, and we numerically confirm successful error mitigation with simulations up to 18 qubits. Finally, we present an application of the EQ-GAN to prepare an approximate quantum random access memory and for the training of quantum neural networks via variational datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

The cost of conservative synchronization in parallel discrete event simulations

The performance of a synchronous conservative parallel discrete-event simulation protocol is analyzed. The class of simulation models considered is oriented around a physical domain and possesses a limited ability to predict future behavior. A stochastic model is used to show that as the volume of simulation activity in the model increases relative to a fixed architecture, the complexity of the average per-event overhead due to synchronization, event list manipulation, lookahead calculations, and processor idle time approach the complexity of the average per-event overhead of a serial simulation. The method is therefore within a constant factor of optimal. The analysis demonstrates that on large problems--those for which parallel processing is ideally suited--there is often enough parallel workload so that processors are not usually idle. The viability of the method is also demonstrated empirically, showing how good performance is achieved on large problems using a thirty-two node Intel iPSC/2 distributed memory multiprocessor.

Nicol, David M.↗

Moment-based adaptive time integration for thermal radiation transport

Here, in this paper we develop a framework for moment-based adaptive time integration of deterministic multifrequency thermal radiation transpot (TRT). We generalize our recent semi-implicit-explicit (IMEX) integration framework for gray TRT to multifrequency TRT, and also introduce a semi-implicit variation that facilitates higher-order integration of TRT, where each stage is implicit in all components except opacities. To appeal to the broad literature on adaptivity with Runge–Kutta methods, we derive new embedded methods for four asymptotic preserving IMEX Runge–Kutta schemes we have found to be robust in our previous work on TRT and radiation hydrodynamics. We then use a moment-based high-order-low-order representation of the transport equations. Due to the high dimensionality, memory is always a concern in simulating TRT. We form error estimates and adaptivity in time purely based on temperature and radiation energy, for a trivial overhead in computational cost and memory usage compared with the base second order integrators. We then test the adaptivity in time on the tophat and Larsen problem, demonstrating the ability of the adaptive algorithm to naturally vary the timestep across 4–5 orders of magnitude, ranging from the dynamical timescales of the streaming regime to the thick diffusion limit.

97 MATHEMATICS AND COMPUTING↗