Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Implementing a neural network interatomic model with performance portability for emerging exascale architectures

The two main thrusts of computational science are increasingly accurate predictions and faster calculations; to this end, the zeitgeist in molecular dynamics (MD) simulations is pursuing machine learned and data driven interatomic models, e.g. neural network potentials, and novel hardware architectures, e.g. GPUs. Current implementations of neural network potentials are orders of magnitude slower than traditional interatomic models and while looming exascale computing offers the ability to run large, accurate simulations with these models, achieving portable performance for MD with new and varied exascale hardware requires rethinking traditional algorithms, using novel data structures, and library solutions. We re-implement a neural network interatomic model in CabanaMD, an MD proxy application, built on libraries developed for performance portability. Our implementation shows significantly improved thread scaling in this complex kernel as compared to a current LAMMPS implementation, across both strong and weak scaling. Our single-source solution enables simulations up to 20 million atoms on a single CPU node and 4 million atoms with improved performance on a single GPU. Furthermore, we also explore parallelism and data layout choices (using flexible data structures called AoSoAs) and their effect on performance, seeing up to ~50% and ~5% improvements in performance on a GPU by choosing the right level of parallelism and data layout respectively.

97 MATHEMATICS AND COMPUTING↗

High-throughput terahertz imaging: progress and challenges

Abstract Many exciting terahertz imaging applications, such as non-destructive evaluation, biomedical diagnosis, and security screening, have been historically limited in practical usage due to the raster-scanning requirement of imaging systems, which impose very low imaging speeds. However, recent advancements in terahertz imaging systems have greatly increased the imaging throughput and brought the promising potential of terahertz radiation from research laboratories closer to real-world applications. Here, we review the development of terahertz imaging technologies from both hardware and computational imaging perspectives. We introduce and compare different types of hardware enabling frequency-domain and time-domain imaging using various thermal, photon, and field image sensor arrays. We discuss how different imaging hardware and computational imaging algorithms provide opportunities for capturing time-of-flight, spectroscopic, phase, and intensity image data at high throughputs. Furthermore, the new prospects and challenges for the development of future high-throughput terahertz imaging systems are briefly introduced.

36 MATERIALS SCIENCE↗

A physical unclonable neutron sensor for nuclear arms control inspections

Abstract Classical sensor security relies on cryptographic algorithms executed on trusted hardware. This approach has significant shortcomings, however. Hardware can be manipulated, including below transistor level, and cryptographic keys are at risk of extraction attacks. A further weakness is that sensor media themselves are assumed to be trusted, and any authentication and encryption is done ex situ and a posteriori. Here we propose and demonstrate a different approach to sensor security that does not rely on classical cryptography and trusted electronics. We designed passive sensor media that inherently produce secure and trustworthy data, and whose honest and non-malicious nature can be easily established. As a proof-of-concept, we manufactured and characterized the properties of non-electronic, physical unclonable, optically complex media sensitive to neutrons for use in a high-security scenario: the inspection of a military facility to confirm the absence or presence of nuclear weapons and fissile materials.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Evolutionary vs imitation learning for neuromorphic control at the edge*

Abstract Neuromorphic computing offers the opportunity to implement extremely low power artificial intelligence at the edge. Control applications, such as autonomous vehicles and robotics, are also of great interest for neuromorphic systems at the edge. It is not clear, however, what the best neuromorphic training approaches are for control applications at the edge. In this work, we implement and compare the performance of evolutionary optimization and imitation learning approaches on an autonomous race car control task using an edge neuromorphic implementation. We show that the evolutionary approaches tend to achieve better performing smaller network sizes that are well-suited to edge deployment, but they also take significantly longer to train. We also describe a workflow to allow for future algorithmic comparisons for neuromorphic hardware on control applications at the edge.

Schuman, Catherine↗

Performance-Aligned LLMs for Generating Fast HPC Code

Optimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. Here, we demonstrate that our fine-tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code.

Computer science↗

Fugu v.0.1

SAND2021-15052 O Fugu provides a common software framework for designing and prototyping algorithms for spiking neuromorphic hardware and compiling to multiple hardware platforms. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Aimone, James↗

Mixed Precision xsdk Project: Final Report for Year Ended March 31, 2022

In the past year, as part of the Mixed Precision xsdk Project we have worked an the development and analysis of algorithms for mixed precision hardware. In the following we briefly discuss and summarize our contributions. The postdoctoral researchers were Srikara Pranesh (to September 30, 2021) and Mantas Mikaitis (from October 1, 2021).

97 MATHEMATICS AND COMPUTING↗

Citadels Final Report (GMLC 2.2.1: Citadels)

This is the final project report for the Grid Modernization Laboratory Consortium (GMLC) Resilient Distribution System (RDS) Citadels project. The primary goal of this GMLC project was to increase the operational flexibly of power systems by engaging microgrids distributedly, coordinated using consensus algorithms. The primary goal was successfully achieved. The primary goal was divided into three areas: Implement peer-to-peer control between microgrid controllers using the Open Field Message Bus (OpenFMB) approach; Develop and implement consensus algorithms on commercially available hardware that allows a group of microgrids to distributedly implement operational controls; Develop the architectures and controls to enable groups of microgrids to coordinate their operations to support the bulk power system during abnormal events, and end-use loads in the event the bulk power systems fail.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Quantum Computing for Energy-Related Applications

Growing interest in quantum computing and simulations have created opportunities for its deployment to improve processes pertaining to energy production, distribution, and consumption. While quantum computing is considered as a paradigm shift in our basic understanding of physical computation, effective implementation of quantum computing in energy applications also depends on progress and development in the dimensions of both quantum computing hardware and quantum computing algorithms. To fully address the status and future challenges of quantum information science (QIS) applied within the energy sector, in this presentation, we firstly summarize recent advancements on the applications of quantum computing to energy infrastructure and materials, complex energy system processes, advanced manufacturing, and energy system security. Then, we will demonstrate the results of quantum computing both on a simulator and a quantum device accessing from OLCF. Our first example is to use the variational quantum eigensolver (VQE) with a unitary coupled cluster with singles and doubles (UCCSD) ansatz to simulate a series of LixHyq molecules (q=-1, 0, +1). The obtained results showed that the quantum computing VQE-UCCSD is comparable to classical CCSD for small systems like LiH with respect to full configuration interaction (FCI). Targeting on CO2 capture application, our second example is to use VQE to quantify molecular vibrational energies and reaction pathways between CO2 and a simplified amine-based solvent model—NH3 to form H2NCOOH. This research showcases quantum computing applications in the study of CO2 capture reactions.

Duan, Yuhua↗

Development of birefringent filters for spaceflight

The critical problem for flight of a birefringent filter is the shock mounting of the calcite. The design presented here bonds the calcite block with silicon rubbers to the calcite holder. The calcite together with its all necessary polarizers and rotating achromatic plates are mounted together in units called a filter module. By using a set of modules containing calcite crystals of differing lengths, a filter can be produced. A description of the modules is given. Also described is a container for the filter modules, which can be used both to hermetically seal the system or contain an index matching oil. The response of a filter element while being controlled by the Lockheed Temperature Control is described and the determination of the wavelength sensitivity to temperature of calcite is explained. Operation of the filter using a software control algorithm instead of a hardware temperature controller is shown. Some radiation considerations of filter systems are given.

Title, A. M.↗

Real-time onboard geometric image correction

A system to perform real time onboard geometric connection of LANDSAT D resolution satellite imagery is described. System requirements, algorithms, sensors, and other hardware components are defined. Feasibility of implementing the correction process is demonstrated using Kalman filter techniques to incorporate information from onboard ephemeris, attitude control, and ground control points. Random access sensor systems, such as charge injected devices and charge coupled devices are used to obtain pixel values at desired ground location, thus greatly reducing the data processing requirements.

Discenza, W.↗

Memory efficient solution of the primitive equations for numerical weather prediction on the CYBER 205

Numerical Weather Prediction (NWP), for both operational and research purposes, requires only fast computational speed but also large memory. A technique for solving the Primitive Equations for atmospheric motion on the CYBER 205, as implemented in the Mesoscale Atmospheric Simulation System, which is fully vectorized and requires substantially less memory than other techniques such as the Leapfrog or Adams-Bashforth Schemes is discussed. The technique presented uses the Euler-Backard time marching scheme. Also discussed are several techniques for reducing computational time of the model by replacing slow intrinsic routines by faster algorithms which use only hardware vector instructions.

Tuccillo, J. J.↗

The science of computing - Parallel computation

Although parallel computation architectures have been known for computers since the 1920s, it was only in the 1970s that microelectronic components technologies advanced to the point where it became feasible to incorporate multiple processors in one machine. Concommitantly, the development of algorithms for parallel processing also lagged due to hardware limitations. The speed of computing with solid-state chips is limited by gate switching delays. The physical limit implies that a 1 Gflop operational speed is the maximum for sequential processors. A computer recently introduced features a 'hypercube' architecture with 128 processors connected in networks at 5, 6 or 7 points per grid, depending on the design choice. Its computing speed rivals that of supercomputers, but at a fraction of the cost. The added speed with less hardware is due to parallel processing, which utilizes algorithms representing different parts of an equation that can be broken into simpler statements and processed simultaneously. Present, highly developed computer languages like FORTRAN, PASCAL, COBOL, etc., rely on sequential instructions. Thus, increased emphasis will now be directed at parallel processing algorithms to exploit the new architectures.

Denning, P. J.↗

High-dynamic GPS tracking

The results of comparing four different frequency estimation schemes in the presence of high dynamics and low carrier-to-noise ratios are given. The comparison is based on measured data from a hardware demonstration. The tested algorithms include a digital phase-locked loop, a cross-product automatic frequency tracking loop, and extended Kalman filter, and finally, a fast Fourier transformation-aided cross-product frequency tracking loop. The tracking algorithms are compared on their frequency error performance and their ability to maintain lock during severe maneuvers at various carrier-to-noise ratios. The measured results are shown to agree with simulation results carried out and reported previously.

Hinedi, S.↗

Euler/Navier-Stokes calculations of transonic flow past fixed- and rotary-wing aircraft configurations

Computational fluid dynamics has an increasingly important role in the design and analysis of aircraft as computer hardware becomes faster and algorithms become more efficient. Progress is being made in two directions: more complex and realistic configurations are being treated and algorithms based on higher approximations to the complete Navier-Stokes equations are being developed. The literature indicates that linear panel methods can model detailed, realistic aircraft geometries in flow regimes where this approximation is valid. As algorithms including higher approximations to the Navier-Stokes equations are developed, computer resource requirements increase rapidly. Generation of suitable grids become more difficult and the number of grid points required to resolve flow features of interest increases. Recently, the development of large vector computers has enabled researchers to attempt more complex geometries with Euler and Navier-Stokes algorithms. The results of calculations for transonic flow about a typical transport and fighter wing-body configuration using thin layer Navier-Stokes equations are described along with flow about helicopter rotor blades using both Euler/Navier-Stokes equations.

Deese, J. E.↗

Computational aspects of multibody dynamics

Computational aspects are addressed which impact the requirements for developing a next generation software system for flexible multibody dynamics simulation which include: criteria for selecting candidate formulation, pairing of formulations with appropriate solution procedures, need for concurrent algorithms to utilize computer hardware advances, and provisions for allowing open-ended yet modular analysis modules.

Park, K. C.↗

Robotic space construction

Research at Langley AFB concerning automated space assembly is reviewed, including a Space Shuttle experiment to test astronaut ability to assemble a repetitive truss structure, testing the use of teleoperated manipulators to construct the Assembly Concept for Construction of Erectable Space Structures I truss, and assessment of the basic characteristics of manipulator assembly operations. Other research topics include the simultaneous coordinated control of dual-arm manipulators and the automated assembly of candidate Space Station trusses. Consideration is given to the construction of an Automated Space Assembly Laboratory to study and develop the algorithms, procedures, special purpose hardware, and processes needed for automated truss assembly.

Mixon, Randolph W.↗