Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “program processors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Node Monitoring as a Fault Detection Countermeasure against Information Leakage within a RISC-V Microprocessor

Advanced, superscalar microprocessors (μP) are highly susceptible to wear-out failures because of their highly complex, densely packed circuit structure and extreme operational frequencies. Although many types of fault detection and mitigation strategies have been proposed, none have addressed the specific problem of detecting faults that lead to information leakage events on I/O channels of the μP. Information leakage can be defined very generally as any type of output that the executing program did not intend to produce. In this work, we restrict this definition to output that represents a security concern, and in particular, to the leakage of plaintext or encryption keys, and propose a counter-based countermeasure to detect faults that cause this type of leakage event. Fault injection (FI) experiments are carried out on two RISC-V microprocessors emulated as soft cores on a Xilinx multi-processor System-on-chip (MPSoC) FPGA. The μP designs are instrumented with a set of counters that records the number of transitions that occur on internal nodes. The transition counts are collected from all internal nodes under both fault-free and faulty conditions, and are analyzed to determine which counters provide the highest fault coverage and lowest latency for detecting leakage faults. We show that complete coverage of all leakage faults is possible using only a single counter strategically placed within the branch compare logic of the μPs.

42 ENGINEERING↗

Accelerating Neutrino Event Generation in MARLEY Using CUDA-Based RNG and GPU Parallelization

MARLEY is a simulation tool that helps scientists study how low-energy neutrinos interact with matter. To work properly, MARLEY uses random numbers thousands of times in each simulation. These random numbers are important for modeling things like how neutrinos collide with atoms and what particles they produce. Right now, MARLEY runs on a regular computer processor (CPU) and uses a built-in random number generator called the Mersenne Twister. This setup works, but it can be slow, especially when trying to simulate many events. This research focuses on making MARLEY run faster by moving the random number generation and some of the repetitive calculations from the CPU to a graphics processing unit (GPU), which can handle many tasks at the same time. We use CUDA (a tool for programming NVIDIA GPUs) and cuRAND (a GPU-based random number library) to test faster alternatives to the current random number system. We compare different GPU-based generators, like curand_mtgp32, xorwow, and philox, to see which ones are the quickest and still give reliable results. Early tests show that using the GPU can make MARLEY simulations much faster. This project not only helps improve current simulation performance but also moves closer to a full simulation chain where all stages can run on modern GPU hardware.

Dunkley, Kimieka [Florida A-M]↗

FLASH: FPGA-Accelerated Smart Switches with GCN Case Study

Some communication switches, e.g., the Mellanox SHArP and those in the IBM BlueGene clusters, are augmented to process packets at the application level with fixed-function collectives. This approach, however, lacks flexibility, which limits their applicability in diverse and dynamic workloads. Recently, a new type of programmable packet processor, which uses high-level languages, e.g., P4, has emerged as possible candidates. P4-based switches, however, fall short in certain applications, including machine learning, where capabilities not currently supported by P4 are needed. These include more complex calculation, such as sparse computation and fused multiply-accumulate, data-intensive floating point operations, data reuse, and significant memory. The problem addressed here is that such a switch augmentation needs to support: a large amount of state, significant flexible compute capability, and ease of programming, all while maintaining full functionality, including ensuring high throughput, and demonstrating utility. In this work, we propose a programmable look-aside-type accelerator that can be embedded into, or attached to, existing communication switch pipelines and that is capable of processing packets at line-rate. The proposed in-switch accelerator is based on mixing an ISA (subset of RISC-V instructions) with dataflow graphs (found in CGRAs). To augment performance, vector instructions are also supported. To facilitate usability, we have developed a complete toolchain to compile user-provided C/C++ codes to appropriate back-end instructions for configuring the accelerator. While this approach is flexible enough to support various workloads, in this paper, we consider Graph Convolutional Networks (GCNs) as a case study. Experimental results show that this approach considerably improves the performance of distributed GCN applications.

Haghi, Pouya↗

Programmable photonic integrated meshes for modular generation of optical entanglement links

Abstract Large-scale generation of quantum entanglement between individually controllable qubits is at the core of quantum computing, communications, and sensing. Modular architectures of remotely-connected quantum technologies have been proposed for a variety of physical qubits, with demonstrations reported in atomic and all-photonic systems. However, an open challenge in these architectures lies in constructing high-speed and high-fidelity reconfigurable photonic networks for optically-heralded entanglement among target qubits. Here we introduce a programmable photonic integrated circuit (PIC), realized in a piezo-actuated silicon nitride (SiN)-in-oxide CMOS-compatible process, that implements an N × N Mach–Zehnder mesh (MZM) capable of high-speed execution of linear optical transformations. The visible-spectrum photonic integrated mesh is programmed to generate optical connectivity on up to N = 8 inputs for a range of optically-heralded entanglement protocols. In particular, we experimentally demonstrated optical connections between 16 independent pairwise mode couplings through the MZM, with optical transformation fidelities averaging 0.991 ± 0.0063. The PIC’s reconfigurable optical connectivity suffices for the production of 8-qubit resource states as building blocks of larger topological cluster states for quantum computing. Our programmable PIC platform enables the fast and scalable optical switching technology necessary for network-based quantum information processors.

47 OTHER INSTRUMENTATION↗

The Importance of Scientific Visualization on Novel Hardware

Innovation in HPC hardware and adoption of heterogeneous systems has led to a variety of unique programming models. This has led to a challenge for scientific visualization software (and indeed all HPC software) to take full advantage of recent generations of supercomputing. Edge computing, where hardware is specialized for the needs of the particular application, exacerbates the problem. VTK-m has had many successes on this front by providing device-agnostic algorithms that compare favorably to implementations written to specific devices as shown in Table 1. However, VTK-m has focused mostly on GPU and traditional CPU multicore technology. There are numerous processor technologies, both existing and potential future, that are not being addressed by current R&D efforts.

97 MATHEMATICS AND COMPUTING↗

Profiling the BLAST bioinformatics application for load balancing on high-performance computing clusters

Abstract Background The Basic Local Alignment Search Tool (BLAST) is a suite of commonly used algorithms for identifying matches between biological sequences. The user supplies a database file and query file of sequences for BLAST to find identical sequences between the two. The typical millions of database and query sequences make BLAST computationally challenging but also well suited for parallelization on high-performance computing clusters. The efficacy of parallelization depends on the data partitioning, where the optimal data partitioning relies on an accurate performance model. In previous studies, a BLAST job was sped up by 27 times by partitioning the database and query among thousands of processor nodes. However, the optimality of the partitioning method was not studied. Unlike BLAST performance models proposed in the literature that usually have problem size and hardware configuration as the only variables, the execution time of a BLAST job is a function of database size, query size, and hardware capability. In this work, the nucleotide BLAST application BLASTN was profiled using three methods: shell-level profiling with the Unix “time” command, code-level profiling with the built-in “profiler” module, and system-level profiling with the Unix “gprof” program. The runtimes were measured for six node types, using six different database files and 15 query files, on a heterogeneous HPC cluster with 500+ nodes. The empirical measurement data were fitted with quadratic functions to develop performance models that were used to guide the data parallelization for BLASTN jobs. Results Profiling results showed that BLASTN contains more than 34,500 different functions, but a single function, RunMTBySplitDB, takes 99.12% of the total runtime. Among its 53 child functions, five core functions were identified to make up 92.12% of the overall BLASTN runtime. Based on the performance models, static load balancing algorithms can be applied to the BLASTN input data to minimize the runtime of the longest job on an HPC cluster. Four test cases being run on homogeneous and heterogeneous clusters were tested. Experiment results showed that the runtime can be reduced by 81% on a homogeneous cluster and by 20% on a heterogeneous cluster by re-distributing the workload. Discussion Optimal data partitioning can improve BLASTN’s overall runtime 5.4-fold in comparison with dividing the database and query into the same number of fragments. The proposed methodology can be used in the other applications in the BLAST+ suite or any other application as long as source code is available.

59 BASIC BIOLOGICAL SCIENCES↗

Unified Memory: GPGPU-Sim/UVM Smart Integration

CPU/GPU heterogeneous compute platforms are an ubiquitous element in computing and a programming model specified for this heterogeneous computing model is important for both performance and programmability. A programming model that exposes the shared, unified, address space between the heterogeneous units is a necessary step in this direction as it removes the burden of explicit data movement from the programmer while maintaining performance. GPU vendors, such as AMD and NVIDIA, have released software-managed runtimes that can provide programmers the illusion of unified CPU and GPU memory by automatically migrating data in and out of the GPU memory. However, this runtime support is not included in GPGPU-Sim, a commonly used framework that models the features of a modern graphics processor that are relevant to non-graphics applications. UVM Smart was developed, which extended GPGPU-Sim 3.x to in- corporate the modeling of on-demand pageing and data migration through the runtime. This report discusses the integration of UVM Smart and GPGPU-Sim 4.0 and the modifications to improve simulation performance and accuracy.

97 MATHEMATICS AND COMPUTING↗

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak↗

Massively parallel axisymmetric fluid model for streamer discharges

A highly parallelizable fluid plasma simulation tool based upon the first-order drift-diffusion equations is discussed. Atmospheric pressure plasmas have densities and gradients that require small element sizes in order to accurately simulate the plasm resulting in computational meshes on the order of millions to tens of millions of elements for realistic size plasma reactors. To enable simulations of this nature, parallel computing is required and must be optimized for the particular problem. Here, a finite-volume, electrostatic drift-diffusion implementation for low-temperature plasma is discussed. The implementation is built upon the Message Passing Interface (MPI) library in C++ using Object Oriented Programming. The underlying numerical method is outlined in detail and benchmarked against simple streamer formation from other streamer codes. Electron densities, electric field, and propagation speeds are compared with the reference case and show good agreement. Convergence studies are also performed showing a minimal space step of approximately 4 μm required to reduce relative error to below 1% during early streamer simulation times and even finer space steps are required for longer times. Additionally, strong and weak scaling of the implementation are studied and demonstrate the excellent performance behavior of the implementation up to 100 million elements on 1024 processors. Lastly, different advection schemes are compared for the simple streamer problem to analyze the influence of numerical diffusion on the resulting quantities of interest.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Lithium promotion of Pt/m-ZrO 2 catalysts for low temperature water-gas shift

Low temperature water-gas shift (LTS) is an important reaction occurring in a fuel processor for producing and purifying hydrogen. Platinum supported on m-ZrO 2 belongs to a family of catalysts consisting of metal nanoparticles and an active partially reducible oxide, with the catalysis proposed to occur at the boundary between metal particles and the support. In this investigation, increasing the loading of lithium dopant increased the LTS rate up to 0.54 wt % lithium, where conversion was 2.4 times that of the unpromoted catalyst at 260°C. Further increases in lithium loading up to 1.5 wt % decreased the rate, although it remained higher than that of the unpromoted catalyst. Infrared spectroscopy and CO 2 temperature programmed desorption experiments showed three effects with increasing lithium loading: (1) lithium promoter weakened the C–H bond of formate, the proposed rate limiting step of the interfacial surface formate mechanism; (2) high levels of lithium suppressed the platinum site capacity required for hydrogen transfer; and (3) high levels of lithium increased catalyst basicity. Aspects (2) and (3) tended to inhibit desorption of product CO 2 , an acidic molecule the removal of which is metal-catalyzed. XANES and XPS experiments revealed that electron transfer to enrich Pt nanoparticles is unlikely the root cause of C–H bond weakening in formate. However, other electronic effects (e.g., electrostatic effects or molecular rearrangement due to enhanced basicity) were not ruled out.

08 HYDROGEN↗

New developments in structure-property and amorphous materials research at the upgraded 16-BM-B Paris-Edinburgh press station at HPCAT

Beamline 16-BM-B of the High-Pressure Collaborative Access Team (HPCAT) at the Advanced Photon Source (APS) provides a Paris-Edinburgh press program to probe the structure and properties of crystalline and amorphous materials up to 12 GPa at room temperature or 7 GPa at 2000°C. During the recent APS upgrade, 16-BM-B undertook major improvements to its instrumentation and measurement techniques. The upgrade includes the addition of a 1.2 m horizontal “condenser” mirror and a new variable-sized collimation system, the combination of which decreases the acquisition time for energy dispersive X-ray diffraction measurements by a factor of 10 when compared to pre-upgrade measurements. The addition of a new Ge energy-sensitive detector and DANTE (XGLab) digital pulse processor allows for processing the increased diffraction counts with minimal deadtime. In addition to upgraded beamline components, new techniques are being developed for eventual release to the user community, including electrical resistivity and tomography measurements.

APS upgrade↗

AthenaK: A Performance-portable Version of the Athena++ Adaptive Mesh Refinement Framework

We describe AthenaK: a new implementation of the Athena++ block-based adaptive mesh refinement framework using the Kokkos programming model. Finite volume methods for Newtonian, special relativistic, and general relativistic (GR) hydrodynamics and magnetohydrodynamics (MHD), and GR-radiation hydrodynamics and MHD, as well as a module for evolving Lagrangian tracer or charged test particles (e.g., cosmic rays) are implemented using the framework. In two companion papers, we describe (1) a new solver for the Einstein equations based on the Z4c formalism, and (2) a GRMHD solver in dynamical spacetimes also implemented using the framework, enabling new applications in numerical relativity. By adopting Kokkos, the code can be run on virtually any hardware, including CPUs, GPUs from multiple vendors, and emerging Advanced RISC Machine processors. AthenaK shows excellent performance and weak scaling, achieving over 1 billion cell updates per second for hydrodynamics in three dimensions on a single NVIDIA Grace Hopper processor. It does this with a typical parallel efficiency of 80% on 65,536 AMD GPUs on the OLCF Frontier system. Such performance portability enables AthenaK to leverage modern exascale computing systems for challenging applications in astrophysical fluid dynamics, numerical relativity, and multimessenger astrophysics.

79 ASTRONOMY AND ASTROPHYSICS↗

Qubit Assignment Using Time Reversal

As quantum computers with large numbers of qubits become increasingly available, experiments executed on a given device may not utilize all available qubits. In this case, the outcome of executing a quantum program will depend on the ability to efficiently select a subset of high-performing physical qubits. For any given quantum program and device there are many ways to assign physical qubits for execution of the program, and assignments will differ in performance due to the variability in quality across qubits and entangling operations on a single device. Evaluating the performance of each assignment using fidelity estimation introduces significant experimental overhead and will be infeasible for many applications, while relying on standard device benchmarks provides incomplete information about the performance of any specific program. Furthermore, the number of possible assignments grows combinatorially in the number of qubits on the device and in the program, motivating the use of heuristic optimization techniques. We demonstrate a practical solution to the problem of qubit assignment by using simulated annealing with a cost function based on the Loschmidt echo, a diagnostic that measures the reversibility of a quantum process. We provide theoretical justification for this choice of cost function by demonstrating that the optimal qubit assignment coincides with the optimal qubit assignment based on state fidelity in the weak error limit, and we provide experimental justification using diagnostics performed on Google’s superconducting qubit devices. We then establish the performance of simulated annealing for qubit assignment using classical simulations of noisy devices as well as optimization experiments performed on a quantum processor. Our results demonstrate that the use of Loschmidt echoes and simulated annealing provides a scalable and flexible approach to optimizing qubit assignment on near-term hardware.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

VERAIO Software Management Plan

VERAIO is a set of utility codes used to provide a common set of input and outputs to the Virtual Environment for Reactor Applications (VERA). VERA is a collection of several different computer codes that all have a common input and output. This prevents the need to manage input and output from each individual code, allowing for ease of use and reducing errors associated with code operability. The VERAIO utilities include VERAIn, VERAView, and VERARun. Each of these utilities is described below. VERAIn is an input processor that reads an ASCII input file generated by users, parses the file, performs some error checking, and writes an XML file to be read by other VERA codes. The main purpose of VERAIn is to provide a common input to all of the VERA codes so users only need to learn one input. VERAIn is written in Perl and uses YAML configuration files to provide flexibility. VERAView is a graphical user interface (GUI) that reads a VERA HDF output file and allows users to visualize results. VERAView is written in Python. VERARun is a script that drives the VERA execution in a high performance computing (HPC) environment. Work performed at the code level supports the Quality Assurance Program Plan (QAPP) (VERA-QA-001), and VERA Software Quality Assurance Plan (VERA-QA-002).

97 MATHEMATICS AND COMPUTING↗

VERAIO Software Management Plan

VERAIO is a set of utility codes used to provide a common set of inputs and outputs to the Virtual Environment for Reactor Applications (VERA). VERA is a collection of several different computer codes that all have a common input and output. This prevents the need to manage input and output from each individual code, allowing for ease of use and reducing errors associated with code operability. The VERAIO utilities include VERAIn, VERAView, and VERARun. Each of these utilities is described below. VERAIn is an input processor that reads an ASCII input file generated by users, parses the file, performs some error checking, and writes an XML file to be read by other VERA codes. The main purpose of VERAIn is to provide a common input to all of the VERA codes, so users only need to learn one input. VERAIn is written in Perl and uses YAML configuration files to provide flexibility. VERAView is a graphical user interface (GUI) that reads a VERA hierarchical data format (HDF) output file and allows users to visualize results. VERAView is written in Python. VERARun is a script that drives the VERA execution in a high performance computing (HPC) environment. Work performed at the code level supports the quality assurance program plan (QAPP) (VERA-QA-001) and the VERA Software Quality Assurance Plan (VERA-QA-002).

97 MATHEMATICS AND COMPUTING↗

Reinforcement Learning for Load-balanced Parallel Particle Tracing

We explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors.

Distributed and parallel particle tracing↗

Implementation of McMurchie–Davidson Algorithm for Gaussian AO Integrals Suited for SIMD Processors

We report an implementation of the McMurchie− Davidson evaluation scheme for 1- and 2-particle Gaussian AO integrals designed for processors with Single Instruction Multiple Data (SIMD) instruction sets. Like in our recent MD implementation for graphical processing units (GPUs) [Asadchev, A.; Valeev, E. F.. J. Chem. Phys. 2024, 160, 244109.], variable-sized batches of shellsets of integrals are evaluated at a time. By optimizing for the floating point instruction throughput rather than minimizing the number of operations, this approach achieves up to 50% of the theoretical hardware peak FP64 performance for many common SIMD-equipped platforms (AVX2, AVX512, NEON), which translates to speedups of up to 30 over the state-of-the-art one-shellset-at-a-time implementation of Obara−Saika-type schemes in Libint for a variety of primitive and contracted integrals. As with our previous work, we rely on the standard C++ programming language such as the std::simd standard library feature to be included in the 2026 ISO C++ standard without any explicit code generation to keep the code base small and portable. The implementation is part of the open source LibintX library freely available at https://github.com/ValeevGroup/libintx.

Basis sets↗

Universal Polarization Transformations: Spatial Programming of Polarization Scattering Matrices Using a Deep Learning‐Designed Diffractive Polarization Transformer

Abstract Controlled synthesis of optical fields having nonuniform polarization distributions presents a challenging task. Here, a universal polarization transformer is demonstrated that can synthesize a large set of arbitrarily‐selected, complex‐valued polarization scattering matrices between the polarization states at different positions within its input and output field‐of‐views (FOVs). This framework comprises 2D arrays of linear polarizers positioned between isotropic diffractive layers, each containing tens of thousands of diffractive features with optimizable transmission coefficients. After its deep learning‐based training, this diffractive polarization transformer can successfully implement N i N o = 10 000 different spatially‐encoded polarization scattering matrices with negligible error, where N i and N o represent the number of pixels in the input and output FOVs, respectively. This universal polarization transformation framework is experimentally validated in the terahertz spectrum by fabricating wire‐grid polarizers and integrating them with 3D‐printed diffractive layers to form a physical polarization transformer. Through this set‐up, an all‐optical polarization permutation operation of spatially‐varying polarization fields is demonstrated, and distinct spatially‐encoded polarization scattering matrices are simultaneously implemented between the input and output FOVs of a compact diffractive processor. This framework opens up new avenues for developing novel devices for universal polarization control and may find applications in, e.g., remote sensing, medical imaging, security, material inspection, and machine vision.

Optical neural networks↗