Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

A Dual Approach in Direct Ink Writing of Thermally Cured Shape Memory Rubber Toughened Epoxy

Bisphenol A-based epoxies are much used in a wide range of composite and coating applications due to their excellent thermomechanical properties. However, their 3D printability remains a challenge with most reported materials suffering from high brittleness and low toughness. In this work, we have described especially modified epoxy resins that enable 3D printing with both fast and slow curing rates. These materials exhibit greatly enhanced toughness, tunable thermomechanical properties, and excellent shape memory behavior. Two different printing systems, including a two-part static mixing printhead and a single extrusion printhead, were developed for fast- and slow-curing epoxies, respectively. The rheology of inks in both systems has been modified into printable thixotropic fluids with the aid of silica nanoparticles and other additives. Epoxide-functionalized telechelic polybutadiene was added into the resins, which are then introduced inside the epoxy network after cross-linking. The addition of polybutadiene rubber significantly improves the toughness (over 135%), fracture strain (over 200%), and shape memory behavior. By adding different amounts of the rubber telechelic, thermomechanical properties, including modulus, elongation, and Tg of epoxy, can be well controlled in a wide range to satisfy different applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Reduced Dynamics of Full Counting Statistics

We present a theory of modified reduced dynamics in the presence of counting fields. Reduced dynamics techniques are useful for describing open quantum systems at long emergent timescales when the memory timescales are short. However, they can be difficult to formulate for observables spanning the system and its environment, such as those characterizing transport properties. A large variety of mixed system--environment observables, as well as their statistical properties, can be evaluated by considering counting fields. Given a numerical method able to simulate the field-modified dynamics over the memory timescale, we show that the long-lived full counting statistics can be efficiently obtained from the reduced dynamics. We demonstrate the utility of the technique by computing the long-time current in the nonequilibrium Anderson impurity model from short-time Monte Carlo simulations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

UPC++ v1.0 Specification, Revision 2020.10.0

UPC++ is a C++11 library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2021.9.0

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

UPC++ v1.0 Specification (Rev. 2023.9.0)

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification (Revision 2022.3.0)

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2023.3.0

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification (Revision 2021.3.0)

UPC++ is a C++11 library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2022.9.0

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2020.3.0

UPC++ is a C++11 library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

Shifting Between Compute and Memory Bounds: A Compression-Enabled Roofline Model

In the evolving landscape of high-performance computing, especially to fight the end of Moore’s Law and Dennard’s Scaling, the ability to shift between compute-bound and memory-bound states is critical for enhancing adaptability and flexibility to diverse system and domain-specific architectures. Such capability is vital for optimizing performance across distinguished hardware configurations, such as accelerators, memory hierarchies, and cache systems. Despite that ad hoc optimization techniques, such as compressed/approximate computation, have been enabled for compute-/data-intensive computing for improved performance in distinct hardware settings, there lacks an understanding of 1) the rational behind performance improvement; 2) capability of different optimizations; 3) what optimization to respond to specific computational and memory demands. This work proposes a compression-enabled roofline model to facilitate this adaptability with data compression techniques to balance and transform between computational and memory demands. This model enables applications to adjust in response to the specific strengths and limitations of the underlying hardware and system to optimize resource utilization. The effectiveness of this approach is demonstrated with matrix multiplication kernels on different input sizes, with turning on/off various compression techniques, including 1) low-precision floating point; 2) sparse matrix formulation; and 3) compressed arrays with ZFP. By reducing memory transfer volumes and cache misses and increasing data locality and computational intensity through compression, the specific roofline model can transform between compute and memory bounds to align more efficiently with system capabilities. This advancement not only improves overall performance but also maximizes adaptability in diverse computing environments.

Naraparaju, Ramasoumya [University of Washington]↗

Structural Simluation Toolkit (SST) v.12.0

The Structural Simulation Toolkit (SST) was developed to explore innovations in highly concurrent computing systems where the instruction set architecture (ISA), micro-architecture, and memory interact with the programming model and communications system. The package provides a fully modular design for extensive exploration of an individual system parameter as well as a parallel simulation environment based on message passing interface (MPI) which enable a high level of performance as well as the ability to look at large systems. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Rodrigues, ArunF.↗

TEAM Project Review, Year 2

This report summarizes our research activities within the TEAM project between December 2020 and December 2021, funded by the ASCR Advanced Research in Quantum Computing program. During the reporting period the LLNL-MSU team has made progress on several fronts. An overarching goal of the team is to provide a comprehensive suite of software tools that can be used for the Characterize-Optimize-Compute loop needed to implement and execute algorithms on quantum devices. We are concurrently developing lightweight solvers that can be used on desktop computers to find optimal control pulses and to characterize small quantum systems (consisting of a few transmons and cavities). However, desktop computers are insufficient for simulating and characterizing larger quantum systems. We have therefore also developed parallel, distributed memory, simulators and optimization solvers, both for open and closed quantum systems. These parallel solvers have, for example, been used to study quantum optimal control for pure-state preparation, utilizing 1000’s of cores on a modern high-performance computing (HPC) platform.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Method and apparatus for two-layer copy-on-write

A system, apparatus and method are provided in which a range of virtual memory addresses and a copy of that range are mapped to the same first system address range in a data processing system until an address in the virtual memory address range, or its copy, is written to. The common system address range includes a number of divisions. Responsive to a write request to an address in a division of the common address range, a second system address range is generated. The second system address range is mapped to the same physical addresses as the first system address range, except that the division containing the address to be written to and its corresponding division in the second system address range are mapped to different physical addresses. First layer mapping data may be stored in a range table buffer and updated when the second system address range is generated.

97 MATHEMATICS AND COMPUTING↗

Variable-porosity panel systems and associated methods

Variable-porosity panel systems and associated methods. A variable-porosity panel system includes a panel assembly with an exterior layer defining a plurality of exterior layer pores and a sliding layer adjacent to the exterior layer and defining a plurality of sliding layer pores. The variable-porosity panel system additionally includes a shape memory alloy (SMA) actuator configured to translate the sliding layer relative to the exterior layer to modulate a porosity of the panel assembly. The SMA actuator includes an SMA element configured to exert an actuation force on the sliding layer and at least partially received within an SMA element receiver of the sliding layer. The SMA element extends out of the sliding layer only at a sliding layer first end. A method of operating the variable-porosity panel system includes assembling the variable-porosity panel system and/or transitioning the panel assembly of the variable-porosity panel system among the plurality of panel configurations.

Calkins, Frederick T.↗

Spin-optomechanical cavity interfaces by deep subwavelength phonon-photon confinement

A central goal of quantum information science is transferring qubits between space, time, and modality. Spin-based systems in solids are promising quantum memories, but high-fidelity transfer of their quantum states to telecom optical fields remains challenging. Here, we introduce a phonon-mediated interface between spins in a diamond nanobeam optomechanical crystal and telecom optical fields by a simultaneous deep-subwavelength confinement of optical and acoustic fields with mode volumes $V_{\textrm{mech}}$$/Λ^3_\textrm{p} ~ 10^{-5}$ and $V_{\textrm{opt}}$$/λ^3 ~ 10^{−3}$, respectively. This confinement boosts the spin-mechanical coupling rate of Group-IV silicon vacancy (SiV − ) centers by an order of magnitude to ~ 32 MHz while retaining high acousto-optical couplings. The optical cavity couples to the spin irrespective of the emitter’s native excited states, avoiding spectral diffusion. Using Quantum Monte Carlo simulations, we estimate heralded entanglement fidelities exceeding 0.96 between two such interfaces. We anticipate broad utility beyond diamond emitter-telecom systems to most solid-state quantum memories.

Raniwala, Hamza [Massachusetts Inst. of Technology↗

Unconventional short-range structural fluctuations in cuprate superconductors

Abstract The interplay between structural and electronic degrees of freedom in complex materials is the subject of extensive debate in physics and materials science. Particularly interesting questions pertain to the nature and extent of pre-transitional short-range order in diverse systems ranging from shape-memory alloys to unconventional superconductors, and how this microstructure affects macroscopic properties. Here we use neutron and X-ray diffuse scattering to uncover universal structural fluctuations in La 2-x Sr x CuO 4 and Tl 2 Ba 2 CuO 6+δ , two cuprate superconductors with distinct point disorder effects and with optimal superconducting transition temperatures that differ by more than a factor of two. The fluctuations are present in wide doping and temperature ranges, including compositions that maintain high average structural symmetry, and they exhibit unusual, yet simple scaling behaviour. The scaling regime is robust and universal, similar to the well-known critical fluctuations close to second-order phase transitions, but with a distinctly different physical origin. We relate this behaviour to pre-transitional phenomena in a broad class of systems with structural and magnetic transitions, and propose an explanation based on rare structural fluctuations caused by intrinsic nanoscale inhomogeneity. We also uncover parallels with superconducting fluctuations, which indicates that the underlying inhomogeneity plays an important role in cuprate physics.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

EXAGRAPH: Graph and combinatorial methods for enabling exascale applications

Combinatorial algorithms in general and graph algorithms in particular play a critical enabling role in numerous scientific applications. However, the irregular memory access nature of these algorithms makes them one of the hardest algorithmic kernels to implement on parallel systems. With tens of billions of hardware threads and deep memory hierarchies, the exascale computing systems in particular pose extreme challenges in scaling graph algorithms. The codesign center on combinatorial algorithms, ExaGraph, was established to design and develop methods and techniques for efficient implementation of key combinatorial (graph) algorithms chosen from a diverse set of exascale applications. Algebraic and combinatorial methods have a complementary role in the advancement of computational science and engineering, including playing an enabling role on each other. In this paper, we survey the algorithmic and software development activities performed under the auspices of ExaGraph from both a combinatorial and an algebraic perspective. In particular, we detail our recent efforts in porting the algorithms to manycore accelerator (GPU) architectures. We also provide a brief survey of the applications that have benefited from the scalable implementations of different combinatorial algorithms to enable scientific discovery at scale. We believe that several applications will benefit from the algorithmic and software tools developed by the ExaGraph team.

97 MATHEMATICS AND COMPUTING↗