Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance Portability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Streaming Generalized Canonical Polyadic Tensor Decompositions

In this paper, we develop a method which we call OnlineGCP for computing the Generalized Canonical Polyadic (GCP) tensor decomposition of streaming data. GCP differs from traditional canonical polyadic (CP) tensor decompositions as it allows for arbitrary objective functions which the CP model attempts to minimize. This approach can provide better fits and more interpretable models when the observed tensor data is strongly non-Gaussian. In the streaming case, tensor data is gradually observed over time and the algorithm must incrementally update a GCP factorization with limited access to prior data. In this work, we extend the GCP formalism to the streaming context by deriving a GCP optimization problem to be solved as new tensor data is observed, formulate a tunable history term to balance reconstruction of recently observed data with data observed in the past, develop a scalable solution strategy based on segregated solves using stochastic gradient descent methods, describe a software implementation that provides performance and portability to contemporary CPU and GPU architectures and integrates with Matlab for enhanced usability, and demonstrate the utility and performance of the approach and software on several synthetic and real tensor data sets.

97 MATHEMATICS AND COMPUTING↗

Methodology and Application of Physical Security Effectiveness Based on Dynamic Force-on-Force Modeling

This report describes the research and development being performed at INL towards a dynamic modeling and simulation framework to enable physical security optimization at commercial nuclear power plants. The framework is based on the dynamic modeling tool EMRALD and is demonstrated for applications that can result in physical security optimization. Two main applications are presented: 1. Integrating FLEX portable equipment performance with FOF models of a plant’s physical security posture, and 2. Location optimization of bullet resistant enclosure. The generic framework for modeling FLEX portable equipment is described in detail, followed by a case study modeling an adversarial attack aimed at causing a radiological release by sabotaging the plant’s power supply and its ultimate heat sink capabilities at a hypothetical PWR. Two distinct FLEX deployment strategies, series and parallel, are modeled with distinct timelines. The results of the adversarial attack modeled in a commercial FOF tool, AVERT, are integrated with the FLEX deployment model in EMRALD. Monte Carlo simulation is used to model the distribution of the timeline in FLEX deployment strategies. Thermal-hydraulic analysis of FLEX performance is performed in RELAP5 and integrated with the EMRALD simulations to provide more realistic timelines in the models. The results demonstrate that, even in the extreme case of a successful adversarial attack, deployment of FLEX equipment can result in a significantly high likelihood of preventing radiological release. The modeling and simulation framework of integrating FLEX equipment with FOF models enables the NPPs to credit FLEX portable equipment in the plant security posture, resulting in an efficient and optimized physical security. The objective of location optimization of BRE is to determine the best location in the plant for a new BRE being planned by the plant to enhance their physical security effectiveness. The plant physical security FOF model is integrated with EMRALD that performs Monte Carlo simulation to run different attack scenarios and a discrete set of potential BRE locations. Sensitivity analysis is used to determine the most effective location for the BRE. The optimization approach can be extended to wide applications such as location optimization of remotely operated weapons and other strategic fixed assets.

97 MATHEMATICS AND COMPUTING↗

Development of a menu of performance tests self-administered on a portable microcomputer

Eighteen cognitive, motor, and information processing performance subtests were screened for self-administration over 10 trials by 16 subjects. When altered presentation forms of the same test were collectively considered, the battery composition was reduced to 10 distinctly different measures. A fully automated microbased testing system was employed in presenting the battery of subtests. Successful self-administration of the battery provided for the field testing of the automated system and facilitated convenient data collection. Total test administration time was 47.2 minutes for each session. Results indicated that nine of the tests stabilized, but for a short battery of tests only five are recommended for use in repeated-measures research. The five recommended tests include: the Tapping series, Number Comparison, Short-term Memory, Grammatical Reasoning, and 4-Choice Reaction Time. These tests can be expected to reveal three factors: (1) cognition, (2) processing quickness, and (3) motor. All the tests stabilized in 24 minutes, or approximately two 12-minute sessions.

Wilkes, Robert L.↗

Porting the Nonlinear Optimization Library HiOp to Accelerator-Based Hardware Architectures

While interior point method has been the centerpiece of nonlinear programming tools used in science and engineering, its reliance on linear solvers that can tackle sparse symmetric indefinite and highly ill-conditioned problems made it difficult to implement it effectively on hardware accelerators. HiOp optimization package attempts to provide an implementation of the interior point method suitable for hardware accelerators by compressing the original sparse problem to produce an underlying linear problem that is dense and of manageable size. Implementations of dense linear solvers are more mature and utilize hardware accelerators better than their sparse counterparts. There is a number of important domain problems, such as optimal power flow analysis for power grids, where the sparse problem can be effectively compressed and deploying dense linear solver within the interior point method can improve performance. Here we describe a portable implementation of HiOp optimization engine, which uses a linear solver from Magma library and runs entirely on hardware accelerators. To compress the problem, HiOp uses customized mixed dense-sparse linear algebra. All HiOp kernels are implemented using Umpire and RAJA portability libraries. We describe details of the implementation and discuss trade-offs between performance, portability and development cost.

97 MATHEMATICS AND COMPUTING↗

Optically Enhanced Bonding Workstation for Robust Bonding

Process control is one of the methods recommended by the FAA to reduce risk in fabrication of structurally bonded composite joints for aircraft structure based on guidance provided in circular AC-107B for certification of structurally bonded joints. An Optically Enhanced Bonding Workstation is presented here that reduces the risk in bonded joint fabrication. Results will be presented demonstrating the benefits of process monitoring and its ability to reduce risk in performing pre-bond composite surface preparation steps. This supports reduction in the timeline to certification of bonded composite structures through development of a robust bonding process upstream of any part certification steps. Sanding surface preparation has been identified as a high risk process step that is known to impact bond performance. Control of sanding during surface preparation can be performed using portable surface analysis tools previously identified including included gloss, color, Fourier Transform Infrared spectroscopy (FTIR) and optically stimulated electron emissions (OSEE). Threshold limits for the surface analysis tool measurements were determined based on an example objective bonding system utilizing a common EA9394 paste adhesive measured using standard double cantilever beam fracture toughness testing. The patented Optically Enhanced Bonding Workstation (OEBW), was tailored to monitor and control the epoxy composite surface preparation step. Surface analysis tool threshold limits were incorporated into the OEBW to demonstrate improved composite bond performance through process control. The surface analysis tools investigated here can easily be incorporated into an automated system due to their applicability to rapidly quantify the composite sanded surface treatment and their portability.

Kutscha, Eileen O.↗

MatRIS: Addressing the Challenges for Portability and Heterogeneity Using Tasking for Matrix Decomposition (Cholesky)

The ubiquitous in-node heterogeneity of HPC and cloud computing platforms makes software portability and performance optimization extremely challenging. Described here, the MatRIS multilevel math library abstraction framework employs tasking to alleviate these difficulties. MatRIS includes the IRIS task-based runtime on the bottom level and exposes different layers of abstraction to render algorithms architecturally agnostic. MatRIS ensures the decomposition and creation of tasks that represent the necessary encapsulation of the optimized kernels from both vendor and open-source math libraries. Once built, MatRIS can select different combinations of accelerators at runtime, making it portable even on diverse heterogeneous architectures. By leveraging the IRIS runtime’s features for managing heterogeneity, MatRIS deploys algorithms that remove the need to specify orchestration and data transfer. This study describes how the serial task abstraction of a tiled Cholesky factorization is made portable and scalable in the case of multi-device and multi-vendor heterogeneity on a node with NVIDIA and AMD GPUs by using MatRIS. First, we demonstrate that Cholesky in MatRIS provides multi-GPU scalability that offers competitive performance versus cuSolverMG. Then, we present the challenges and opportunities for heterogeneous execution.

Monil, M. A. H.↗

Portable Programming Model Exploration for LArTPC Simulation in a Heterogeneous Computing Environment: OpenMP vs. SYCL

The evolution of the computing landscape has resulted in the proliferation of diverse hardware architectures, with different flavors of GPUs and other compute accelerators becoming more widely available. To facilitate the efficient use of these architectures in a heterogeneous computing environment, several programming models are available to enable portability and performance across different computing systems, such as Kokkos, SYCL, OpenMP and others. As part of the High Energy Physics Center for Computational Excellence (HEP-CCE) project, we investigate if and how these different programming models may be suitable for experimental HEP workflows through a few representative use cases. One of such use cases is the Liquid Argon Time Projection Chamber (LArTPC) simulation which is essential for LArTPC detector design, validation and data analysis. Following up on our previous investigations of using Kokkos to port LArTPC simulation in the Wire-Cell Toolkit (WCT) to GPUs, we have explored OpenMP and SYCL as potential portable programming models for WCT, with the goal to make diverse computing resources accessible to the LArTPC simulations. In this work, we describe how we utilize relevant features of OpenMP and SYCL for the LArTPC simulation module in WCT. We also show performance benchmark results on multi-core CPUs, NVIDIA and AMD GPUs for both the OpenMP and the SYCL implementations. Comparisons with different compilers will also be given where appropriate.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Performance on HPC Platforms Is Possible Without C++

Computing at large scales has become extremely challenging due to increasing heterogeneity in both hardware and software. More and more scientific workflows must tackle a range of scales and use machine learning and AI intertwined with more traditional numerical modeling methods, placing more demands on computational platforms. These constraints indicate a need to fundamentally rethink the way computational science is done and the tools that are needed to enable these complex workflows. The current set of C++-based solutions may not suffice, and relying exclusively upon C++ may not be the best option, especially because several newer languages and boutique solutions offer more robust design features to tackle the challenges of heterogeneity. In June 2023, we held a mini symposium that explored the use of newer languages and heterogeneity solutions that are not tied to C++ and that offer options beyond template metaprogramming and Parallel. For for performance and portability. In conclusion, we describe some of the presentations and discussion from the mini symposium in this article.

97 MATHEMATICS AND COMPUTING↗

KARL: A Knowledge-Assisted Retrieval Language

Data classification and storage are tasks typically performed by application specialists. In contrast, information users are primarily non-computer specialists who use information in their decision-making and other activities. Interaction efficiency between such users and the computer is often reduced by machine requirements and resulting user reluctance to use the system. This thesis examines the problems associated with information retrieval for non-computer specialist users, and proposes a method for communicating in restricted English that uses knowledge of the entities involved, relationships between entities, and basic English language syntax and semantics to translate the user requests into formal queries. The proposed method includes an intelligent dictionary, syntax and semantic verifiers, and a formal query generator. In addition, the proposed system has a learning capability that can improve portability and performance. With the increasing demand for efficient human-machine communication, the significance of this thesis becomes apparent. As human resources become more valuable, software systems that will assist in improving the human-machine interface will be needed and research addressing new solutions will be of utmost importance. This thesis presents an initial design and implementation as a foundation for further research and development into the emerging field of natural language database query systems.

Dominick, Wayne D.↗

INTEGRATION OF FLEX EQUIPMENT AND OPERATOR ACTIONS IN PLANT FORCE-ON-FORCE MODELS WITH DYNAMIC RISK ASSESSMENT

The overall operation and maintenance cost to protect nuclear power plants accounts for approximately 7% of the total cost of power generation, with labor accounting for half of this cost. In the current research, from interaction with utilities and other stakeholders, it was determined that physical security forces account for nearly 20% of the entire workforce at several nuclear power plants. Labor costs continue to rise in the U.S., so any measures to reduce the cost of operating a nuclear power plant will need to include a reduction in labor. The physical security pathway within the DOE’s Light Water Reactor Sustainability program aims to lower the cost of physical security through directed research into modeling and simulation, application of advanced sensors or deployment of advanced weapons. This report presents a modeling and simulation framework for integrating Diverse and Flexible Mitigation Capability (FLEX) portable equipment performance with Force on Force models of a plant’s physical security posture. The generic framework is described in detail, followed by a case study of modeling an adversarial attack aimed at causing a radiological release by sabotaging the plant’s power supply and its ultimate heat sink capabilities at a hypothetical nuclear power plant. Two different FLEX deployment strategies, series and parallel, are modeled with distinct timelines. The results of the adversarial attack modeled in a commercial Force on Force tool are integrated with the FLEX deployment model in INL’s dynamic modeling tool EMRALD. Monte Carlo simulation is used to model the distribution of the timeline in FLEX deployment strategies. The results demonstrate that, even in the extreme case of a successful adversarial attack, deployment of FLEX equipment can result in a significantly high likelihood of preventing radiological release. The modeling and simulation framework integrating FLEX equipment with Force on Force models enables the nuclear power plants to credit FLEX portable equipment in the plant security posture, resulting in an efficient and optimized physical security.

97 MATHEMATICS AND COMPUTING↗

Static Subspace Approximation for Random Phase Approximation Correlation Energies: Implementation and Performance

Developing theoretical understanding of complex reactions and processes at interfaces requires using methods that go beyond semilocal density functional theory to accurately describe the interactions between solvent, reactants and substrates. Methods based on many-body perturbation theory, such as the random phase approximation (RPA), have previously been limited due to their computational complexity. However, this is now a surmountable barrier due to the advances in computational power available, in particular through modern GPU-based supercomputers. In this work, we describe the implementation of RPA calculations within BerkeleyGW and show its favorable computational performance on large complex systems relevant for catalysis and electrochemistry applications. Our implementation builds off of the static subspace approximation which, by employing a compressed representation of the frequency dependent polarizability, enables the evaluation of the RPA correlation energy with significant acceleration and systematically controllable accuracy. We find that the computational cost of calculating the RPA correlation energy scales only linearly with system size for systems containing up to 50 thousand bands, and is expected to scale quadratically thereafter. We also show excellent strong scaling results across several supercomputers, demonstrating the performance and portability of this implementation.

algorithmic development↗

Automatic Code Generation for High-Performance Graph Algorithms

Graph problems are common across fields of scientific computing and social sciences. However, despite their importance, implementing graph algorithms effectively on modern computing systems is a challenging task that requires significant programming effort and generally results in customized implementations. Current computing and memory hierarchies are not architected for irregular computations resulting in challenges for graph algorithms to achieve high performance on those architectures. In this paper, we present GraphX, a novel compiler framework and DSL designed to simplify the development of efficient graph algorithms and achieve high performance on modern computing systems. GraphX consists of a DSL for efficient implementation of graph algorithms, various optimizations, such as support for sparse linear algebra and workspace transformations, optimized graph primitives, including semiring and masking, and a high-performance code generation engine. Using GraphX, users can implement graph algorithms using a semantically-rich language with graph-oriented operators. GraphX uses these semantics to automatically generate efficient code for target architectures, increasing performance and portability across architectures. The composable nature of GraphX makes it possible to extend the set of optimizations and architectures without modifying the source code. We demonstrate GraphX outperforms state-of-the-art graph libraries, such as LAGraph, up to $3.7 speedup in semiring operations, $2.19 speedup in an important sparse computational kernel, and $9.05 speedup in graph processing algorithms.

compiler, graph algorithms, semiring, masking, wor↗

Enabling particle applications for exascale computing platforms

The Exascale Computing Project (ECP) is invested in co-design to assure that key applications are ready for exascale computing. Within ECP, the Co-design Center for Particle Applications (CoPA) is addressing challenges faced by particle-based applications across four “sub-motifs”: short-range particle–particle interactions (e.g., those which often dominate molecular dynamics (MD) and smoothed particle hydrodynamics (SPH) methods), long-range particle–particle interactions (e.g., electrostatic MD and gravitational N-body), particle-in-cell (PIC) methods, and linear-scaling electronic structure and quantum molecular dynamics (QMD) algorithms. Our crosscutting co-designed technologies fall into two categories: proxy applications (or “apps”) and libraries. Proxy apps are vehicles used to evaluate the viability of incorporating various types of algorithms, data structures, and architecture-specific optimizations and the associated trade-offs; examples include ExaMiniMD, CabanaMD, CabanaPIC, and ExaSP2. Libraries are modular instantiations that multiple applications can utilize or be built upon; CoPA has developed the Cabana particle library, PROGRESS/BML libraries for QMD, and the SWFFT and fftMPI parallel FFT libraries. Success is measured by identifiable “lessons learned” that are translated either directly into parent production application codes or into libraries, with demonstrated performance and/or productivity improvement. The libraries and their use in CoPA’s ECP application partner codes are also addressed.

97 MATHEMATICS AND COMPUTING↗

SciDAC ISEP: Integrated Simulation of Energetic Particles in Burning Plasmas

The objective of the SciDAC Center for Integrated Simulation of Energetic Particles in Burning Plasmas (ISEP) is to improve physics understanding of energetic particle (EP) confinement and EP interactions with burning thermal plasmas through large-scale simulations. The ISEP center will develop a multiscale and multiphysics ISEP framework for a predictive capability of EP physics and deliver an EP module incorporating both first-principles simulations and high fidelity reduced transport models to the fusion whole device modeling (WDM) project. The ISEP framework will enable us to perform long time, global kinetic simulations of EP physics in burning plasmas, by utilizing the full power of the next generation supercomputers. Our research and development activities will build on fruitful collaborations with computer scientists and applied mathematicians to offer enabling technologies for performance scalability, portability, solvers, coupling for integration with the fusion WDM project, and long-term preservation of data.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A design for a new catalog manager and associated file management for the Land Analysis System (LAS)

Due to the larger number of different types of files used in an image processing system, a mechanism for file management beyond the bounds of typical operating systems is necessary. The Transportable Applications Executive (TAE) Catalog Manager was written to meet this need. Land Analysis System (LAS) users at the EROS Data Center (EDC) encountered some problems in using the TAE catalog manager, including catalog corruption, networking difficulties, and lack of a reliable tape storage and retrieval capability. These problems, coupled with the complexity of the TAE catalog manager, led to the decision to design a new file management system for LAS, tailored to the needs of the EDC user community. This design effort, which addressed catalog management, label services, associated data management, and enhancements to LAS applications, is described. The new file management design will provide many benefits including improved system integration, increased flexibility, enhanced reliability, enhanced portability, improved performance, and improved maintainability.

Greenhagen, Cheryl↗

ISLE: Intelligent Selection of Loop Electronics. A CLIPS/C++/INGRES integrated application

The Intelligent Selection of Loop Electronics (ISLE) system is an integrated knowledge-based system that is used to configure, evaluate, and rank possible network carrier equipment known as Digital Loop Carrier (DLC), which will be used to meet the demands of forecasted telephone services. Determining the best carrier systems and carrier architectures, while minimizing the cost, meeting corporate policies and addressing area service demands, has become a formidable task. Network planners and engineers use the ISLE system to assist them in this task of selecting and configuring the appropriate loop electronics equipment for future telephone services. The ISLE application is an integrated system consisting of a knowledge base, implemented in CLIPS (a planner application), C++, and an object database created from existing INGRES database information. The embedibility, performance, and portability of CLIPS provided us with a tool with which to capture, clarify, and refine corporate knowledge and distribute this knowledge within a larger functional system to network planners and engineers throughout U S WEST.

Fischer, Lynn↗

Performance of a Bounce-Averaged Global Model of Super-Thermal Electron Transport in the Earth's Magnetic Field

In this paper, we report the results of our recent research on the application of a multiprocessor Cray T916 supercomputer in modeling super-thermal electron transport in the earth's magnetic field. In general, this mathematical model requires numerical solution of a system of partial differential equations. The code we use for this model is moderately vectorized. By using Amdahl's Law for vector processors, it can be verified that the code is about 60% vectorized on a Cray computer. Speedup factors on the order of 2.5 were obtained compared to the unvectorized code. In the following sections, we discuss the methodology of improving the code. In addition to our goal of optimizing the code for solution on the Cray computer, we had the goal of scalability in mind. Scalability combines the concepts of portabilty with near-linear speedup. Specifically, a scalable program is one whose performance is portable across many different architectures with differing numbers of processors for many different problem sizes. Though we have access to a Cray at this time, the goal was to also have code which would run well on a variety of architectures.

McGuire, Tim↗