Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “program processors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Cache management based on access type priority

Systems, apparatuses, and methods for cache management based on access type priority are disclosed. A system includes at least a processor and a cache. During a program execution phase, certain access types are more likely to cause demand hits in the cache than others. Demand hits are load and store hits to the cache. A run-time profiling mechanism is employed to find which access types are more likely to cause demand hits. Based on the profiling results, the cache lines that will likely be accessed in the future are retained based on their most recent access type. The goal is to increase demand hits and thereby improve system performance. An efficient cache replacement policy can potentially reduce redundant data movement, thereby improving system performance and reducing energy consumption.

Yin, Jieming↗

Controlling accesses to a branch prediction unit for sequences of fetch groups

An electronic device is described that handles control transfer instructions (CTIs) when executing instructions in program code. The electronic device has a processor that includes a branch prediction functional block and a sequential fetch logic functional block. The sequential fetch logic functional block determines, based on a record associated with a CTI, that a specified number of fetch groups of instructions that were previously determined to include no CTIs are to be fetched for execution in sequence following the CTI. When each of the specified number of fetch groups is fetched and prepared for execution, the sequential fetch logic prevents corresponding accesses of the branch prediction functional block for acquiring branch prediction information for instructions in that fetch group.

Yalavarti, Adithya↗

EJFAT: Towards Intelligent Compute Destination Load Balancing

To handle increased data flow, Jefferson Lab (JLab) is partnering with ESnet for development of an AI/ML directed compute work Load Balancer (LB) of UDP streamed data. The LB is FPGA based featuring dynamically configurable, low latency and high throughput destination address switching. The LB provides integration of edge and core computing to support JLab experimental programs, the Electron-Ion Collider, as well as data centers of the future. In the ESnet/JLab FPGA Accelerated Transport (EJFAT) initiative, the function of the LB Data Plane (DP) is to redirect data streams to selectable (but unknown to sender) destination hosts based on current worload and within that host to destination ports as a function of sub- stream id. This effects hierarchical scaling, first across compute machines for processing over a series of events and second, across ports so different data source sub-streams may be assigned to different processors for further parallelization. The LB Control Plane (CP) programs the DP using compute farm telemetry to direct and balance workloads across a compute cluster as the operating conditions require. While Proportional/Integrative/Derivative (PID) controllers are often seen in similar applications, here we investigate the feasibility of a Reinforcement Learning (RL) based schedule manager running in the CP to provide dynamic updates to the DP scheduling policy.

Lawrence, David↗

Systems and methods for tensor scheduling

A technique for efficient scheduling of operations in a program for parallelized execution thereof using a multi-processor runtime environment having two or more processors includes constraining the type or number of loop optimization transforms that may be explored such that memory and processing capacity available for the scheduling task are not exceeded, while facilitating a tradeoff between memory locality, parallelization, and/or data communication between memory modules of the multi-processor runtime environment.

Meister, Benoit J.↗

Toward Performance Portable Programming for Heterogeneous System-on-Chips: Case Study with Qualcomm Snapdragon SoC

Future heterogeneous Domain-Specific System-on-Chips (DSSoC) will be extraordinarily complex in terms of processors, memory hierarchies, and interconnection networks.To manage this complexity, architects, system software designers, and application developers need programming technologies that are flexible, accurate, efficient, and productive. These technologies will need to be as independent of any one specific architecture as is practical, because the sheer dimensionality and scale of the complexity will not allow porting and optimizing applications foreach given DSSoC. To address these issues, we are developing Cosmic Castle, a performance portable programming toolchain for streaming applications on heterogeneous architectures. The primary focus of Cosmic Castle is on enabling efficient and performant code generation through the smart compiler and intelligent runtime system. This paper presents the preliminary evaluation of our ongoing work toward Cosmic Castle. Specifically, we detail our code porting efforts and evaluate various benchmarks on the Qualcomm Snapdragon SoC using tools developed through Cosmic Castle.

Cabrera, Anthony↗

EAP Patterns

EAP Patterns: Memory access and iteration patterns from the EAP code base with the physics removed. This is intended to be a serial app representing memory access patterns. We will populate the data structures from EAP output files to provide representative patterns of face and cell loops within the EAP code base The arguments to the program are the EAP output file and the number of MPI processors to emulate. Currently there is no MPI in the application and the number of processors represents a means of emulating the halo (clone) cells around the domain of the given processor. An optional argument `processor_ID` can be provided to specify the processor to emulate. In the future we plan to include the ability to run multiple MPI ranks to more accurately represent the on-node demands on memory bandwidth.

Swaminarayan, Sriram↗

Composable Programming of Hybrid Workflows for Quantum Simulation

We present a composable design scheme for the development of hybrid quantum/classical algorithms and workflows for applications of quantum simulation. Our object-oriented approach is based on constructing an expressive set of common data structures and methods that enable programming of a broad variety of complex hybrid quantum simulation applications. The abstract core of our scheme is distilled from the analysis of the current quantum simulation algorithms. Subsequently, it allows a synthesis of new hybrid algorithms and workflows via the extension, specialization, and dynamic customization of the abstract core classes defined by our design. We implement our design scheme using the hardware-agnostic programming language QCOR into the QuaSiMo library. To validate our implementation, we test and show its utility on commercial quantum processors from IBM, running some prototypical quantum simulations.

97 MATHEMATICS AND COMPUTING↗

Enabling power measurement and control on Astra: The first petascale Arm supercomputer

Astra, deployed in 2018, was the first petascale supercomputer to utilize processors based on the ARM instruction set. The system was also the first under Sandia's Vanguard program which seeks to provide an evaluation vehicle for novel technologies that with refinement could be utilized in demanding, large-scale HPC environments. In addition to ARM, several other important first-of-a-kind developments were used in the machine, including new approaches to cooling the datacenter and machine. Here we document our experiences building a power measurement and control infrastructure for Astra. While this is often beyond the control of users today, the accurate measurement, cataloging, and evaluation of power, as our experiences show, is critical to the successful deployment of a large-scale platform. While such systems exist in part for other architectures, Astra required new development to support the novel Marvell ThunderX2 processor used in compute nodes. In addition to documenting the measurement of power during system bring up and for subsequent on-going routine use, we present results associated with controlling the power usage of the processor, an area which is becoming of progressively greater interest as data centers and supercomputing sites look to improve compute/energy efficiency and find additional sources for full system optimization.

97 MATHEMATICS AND COMPUTING↗

Universal control of a six-qubit quantum processor in silicon

Future quantum computers capable of solving relevant problems will require a large number of qubits that can be operated reliably. However, the requirements of having a large qubit count and operating with high fidelity are typically conflicting. Spins in semiconductor quantum dots show long-term promise but demonstrations so far use between one and four qubits and typically optimize the fidelity of either single- or two-qubit operations, or initialization and readout. Here, we increase the number of qubits and simultaneously achieve respectable fidelities for universal operation, state preparation and measurement. We design, fabricate and operate a six-qubit processor with a focus on careful Hamiltonian engineering, on a high level of abstraction to program the quantum circuits, and on efficient background calibration, all of which are essential to achieve high fidelities on this extended system. State preparation combines initialization by measurement and real-time feedback with quantum-non-demolition measurements. These advances will enable testing of increasingly meaningful quantum protocols and constitute a major stepping stone towards large-scale quantum computers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Demonstration of the On-the-Fly Shielding Analysis Method: Spent Fuel and Waste Disposition

This report documents work performed supporting the US Department of Energy (DOE) Office of Nuclear Energy (NE) Spent Fuel and Waste Disposition (SFWD) Integrated Waste Management activities under work breakdown structure element 1.08.02.04.01, “Data and Tools Development, Validation, and Maintenance.” In particular, this report fulfills milestone M3SF-21OR020401016, “Implement on-the-fly dose analysis methodology in UNF-ST&DARDS” within work package SF-21OR02040101, “Commercial SNF Characterization - ORNL.” The Used Nuclear Fuel - Storage, Transportation & Disposal Analysis Resource and Data System (UNFST& DARDS) enables automated dose rate calculations for spent nuclear fuel (SNF) transportation packages and storage casks using a Monte Carlo radiation transport code. The explicit method uses a detailed model of the SNF system and its contents. Therefore, a dose rate calculation is required for each as-loaded transportation package or storage cask because the SNF assemblies within a canister typically have unique irradiation characteristics. An alternate method, referred to as the “on-the-fly” shielding analysis method, has been proposed that requires only a set of Monte Carlo dose rate calculations for each transportation packaging/storage cask design. The results of the Monte Carlo dose rate calculations are independent of the SNF assembly irradiation and decay characteristics. The dose rate values may then be combined with the radiation source strength of the SNF assemblies associated with a particular transportation packaging/storage cask design to determine actual dose rates. This report presents on-the-fly dose rate calculations for a representative SNF storage cask and verification of the on-the-fly dose rate calculation results by comparison with reference dose rate calculations using the explicit Monte Carlo dose rate calculation. The on-the-fly shielding analysis method was implemented in UNF-ST&DARDS. A Python program was developed to process the MAVRIC dose rate results obtained by source particle type, energy group, and fuel geometry region. A Python processor created binary files, which were saved as a special UNF-ST&DARDS library for on-the- fly shielding analyses. UNF-ST&DARDS uses the precalculated on-the-fly binary libraries generated by the Python data processor and directly executes the Python code for on-the-fly dose analysis. This Python code unzips the pre-generated binary files mentioned above, reads the data, and combines them with user-specified sources for dose and uncertainty calculations. The Python programs were verified using Excel calculations and by comparison with the values obtained with the MAVRIC post-processing utilities applied to the 3dmap files. This method can currently be used to determine dose rates for as-loaded HI-STORM FW storage casks. The UNF-ST&DARDS analysis wizard for on-the-fly shielding analysis is described in this report.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Traversable wormhole dynamics on a quantum processor

The holographic principle, theorized to be a property of quantum gravity, postulates that the description of a volume of space can be encoded on a lower-dimensional boundary. The anti-de Sitter (AdS)/conformal field theory correspondence or duality is the principal example of holography. The Sachdev–Ye–Kitaev (SYK) model of N >> 1 Majorana fermions has features suggesting the existence of a gravitational dual in AdS 2 , and is a new realization of holography. Here, we invoke the holographic correspondence of the SYK many-body system and gravity to probe the conjectured ER=EPR relation between entanglement and spacetime geometry through the traversable wormhole mechanism as implemented in the SYK model. A qubit can be used to probe the SYK traversable wormhole dynamics through the corresponding teleportation protocol. This can be realized as a quantum circuit, equivalent to the gravitational picture in the semiclassical limit of an infinite number of qubits. Here we use learning techniques to construct a sparsified SYK model that we experimentally realize with 164 two-qubit gates on a nine-qubit circuit and observe the corresponding traversable wormhole dynamics. Despite its approximate nature, the sparsified SYK model preserves key properties of the traversable wormhole physics: perfect size winding, coupling on either side of the wormhole that is consistent with a negative energy shockwave, a Shapiro time delay, causal time-order of signals emerging from the wormhole, and scrambling and thermalization dynamics. Our experiment was run on the Google Sycamore processor. By interrogating a two-dimensional gravity dual system, our work represents a step towards a program for studying quantum gravity in the laboratory. Future developments will require improved hardware scalability and performance as well as theoretical developments including higher-dimensional quantum gravity duals and other SYK-like models.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Programmable simulations of molecules and materials with reconfigurable quantum processors

Simulations of quantum chemistry and quantum materials are believed to be among the most important applications of quantum information processors. However, realizing practical quantum advantage for such problems is challenging because of the prohibitive computational cost of programming typical problems into quantum hardware. Here we introduce a simulation framework for strongly correlated quantum systems represented by model spin Hamiltonians that uses reconfigurable qubit architectures to simulate real-time dynamics in a programmable way. Our approach also introduces an algorithm for extracting chemically relevant spectral properties via classical co-processing of quantum measurement results. We develop a digital–analogue simulation toolbox for efficient Hamiltonian time evolution using digital Floquet engineering and hardware-optimized multi-qubit operations to accurately realize complex spin–spin interactions. As an example, we propose an implementation based on Rydberg atom arrays. In addition, we show how detailed spectral information can be extracted from the dynamics through snapshot measurements and single-ancilla control, enabling the evaluation of excitation energies and finite-temperature susceptibilities from a single dataset. To illustrate the approach, we show how to use the method to compute key properties of a polynuclear transition-metal catalyst and two-dimensional magnetic materials.

74 ATOMIC AND MOLECULAR PHYSICS↗

Parallel Variable Population Multi-Objective Optimizer (pvpmoo) v1.0

This is a parallel variable population multi-objective optimizer with an adaptive unified differential evolution algorithm or a genetic algorithm. It can also be used for single objective optimization. Some features of this code include: 1) The population size varies from generation to generation to save the total # of objective function evaluations. 2) The population is uniformly distributed to a number of parallel processors for simultaneous objective function evaluation. 3) The objective function evaluation can be attained from an external simulation program with control variables in its input file and objectives calculated from its output files. 4) The optimizer includes an adaptive unified differential evolution algorithm and a real value genetic algorithm. The parameters in the unified differential evolution algorithm can be chosen to attain any mutation schemes in the published literature.

Qiang, Ji↗

Tensor Network Quantum Virtual Machine for Simulating Quantum Circuits at Exascale

The numerical simulation of quantum circuits is an indispensable tool for development, verification, and validation of hybrid quantum-classical algorithms intended for near-term quantum co-processors. The emergence of exascale high-performance computing (HPC) platforms presents new opportunities for pushing the boundaries of quantum circuit simulation. Here, we present a modernized version of the Tensor Network Quantum Virtual Machine (TNQVM) that serves as the quantum circuit simulation backend in the eXtreme-scale ACCelerator (XACC) framework. The new version is based on the scalable tensor network processing library ExaTN (Exascale Tensor Networks). It provides multiple configurable quantum circuit simulators that perform either an exact quantum circuit simulation via the full tensor network contraction or an approximate simulation via a suitably chosen tensor factorization scheme. Upon necessity, stochastic noise modeling from real quantum processors is incorporated into the simulations by modeling quantum channels with Kraus tensors. By combining the portable XACC quantum programming frontend and the scalable ExaTN numerical processing backend, we introduce an end-to-end virtual quantum development environment that can scale from laptops to future exascale platforms. We report initial benchmarks of our framework, which include a demonstration of the distributed execution, incorporation of quantum decoherence models, and simulation of the random quantum circuits used for the certification of quantum supremacy on Google’s Sycamore superconducting architecture.

Nguyen, Thien↗

Bringing OpenCL to Commodity RISC-V CPUs

The importance of open-source hardware has been increasing in recent years with the introduction of the RISC-V Open ISA. This has also accelerated the push for support of the open-source software stack from compiler tools to full-blown operating systems. Parallel computing with today’s Application Programming Interfaces such as OpenCL has proven to be effective at leveraging the parallelism in commodity multi-core processors and programmable parallel accelerators. However, to the best of our knowledge, there is currently no publicly available implementation of OpenCL targeting commodity RISC-V processors that is accessible to the open-source community. Besides opening RISC-V to the existing rich variety of scientific parallel applications, OpenCL also provides access to a unique genre of benchmarks useful in computer architecture research. In this work, we extended an Open-source implementation of OpenCL to target RISC-V CPUs. Our work not only cover commodity multi-core RISC-V processors, but also plethora of low- profile embedded RISC-V CPUs that often do not support atomic instructions or multi-threading.

Tine, Blaise↗

Differentiable, Learnable, Regionalized Process-Based Models With Multiphysical Outputs can Approach State-Of-The-Art Hydrologic Prediction Accuracy

Predictions of hydrologic variables across the entire water cycle have significant value for water resources management as well as downstream applications such as ecosystem and water quality modeling. Recently, purely data-driven deep learning models like long short-term memory (LSTM) showed seemingly insurmountable performance in modeling rainfall runoff and other geoscientific variables, yet they cannot predict untrained physical variables and remain challenging to interpret. Here, we show that differentiable, learnable, process-based models (called δ models here) can approach the performance level of LSTM for the intensively observed variable (streamflow) with regionalized parameterization. We use a simple hydrologic model HBV as the backbone and use embedded neural networks, which can only be trained in a differentiable programming framework, to parameterize, enhance, or replace the process-based model's modules. Without using an ensemble or post-processor, δ models can obtain a median Nash-Sutcliffe efficiency of 0.732 for 671 basins across the USA for the Daymet forcing data set, compared to 0.748 from a state-of-the-art LSTM model with the same setup. For another forcing data set, the difference is even smaller: 0.715 versus 0.722. Meanwhile, the resulting learnable process-based models can output a full set of untrained variables, for example, soil and groundwater storage, snowpack, evapotranspiration, and baseflow, and can later be constrained by their observations. Both simulated evapotranspiration and fraction of discharge from baseflow agreed decently with alternative estimates. The general framework can work with models with various process complexity and opens up the path for learning physics from big data.

54 ENVIRONMENTAL SCIENCES↗