Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Raster-Based Approach to Solar Pressure Modeling

An algorithm has been developed to take advantage of the graphics processing hardware in modern computers to efficiently compute high-fidelity solar pressure forces and torques on spacecraft, taking into account the possibility of self-shading due to the articulation of spacecraft components such as solar arrays. The process is easily extended to compute other results that depend on three-dimensional attitude analysis, such as solar array power generation or free molecular flow drag. The impact of photons upon a spacecraft introduces small forces and moments. The magnitude and direction of the forces depend on the material properties of the spacecraft components being illuminated. The parts of the components being lit depends on the orientation of the craft with respect to the Sun, as well as the gimbal angles for any significant moving external parts (solar arrays, typically). Some components may shield others from the Sun. The purpose of this innovation is to enable high-fidelity computation of solar pressure and power generation effects of illuminated portions of spacecraft, taking self-shading from spacecraft attitude and movable components into account. The key idea in this innovation is to compute results dependent upon complicated geometry by using an image to break the problem into thousands or millions of sub-problems with simple geometry, and then the results from the simpler problems are combined to give high-fidelity results for the full geometry. This process is performed by constructing a 3D model of a spacecraft using an appropriate computer language (OpenGL), and running that model on a modern computer's 3D accelerated video processor. This quickly and accurately generates a view of the model (as shown on a computer screen) that takes rotation and articulation of spacecraft components into account. When this view is interpreted as the spacecraft as seen by the Sun, then only the portions of the craft visible in the view are illuminated. The view as shown on the computer screen is composed of up to millions of pixels. Each of those pixels is associated with a small illuminated area of the spacecraft. For each pixel, it is possible to compute its position, angle (surface normal) from the view direction, and the spacecraft material (and therefore, optical coefficients) associated with that area. With this information, the area associated with each pixel can be modeled as a simple flat plate for calculating solar pressure. The vector sum of these individual flat plate models is a high-fidelity approximation of the solar pressure forces and torques on the whole vehicle. In addition to using optical coefficients associated with each spacecraft material to calculate solar pressure, a power generation coefficient is added for computing solar array power generation from the sum of the illuminated areas. Similarly, other area-based calculations, such as free molecular flow drag, are also enabled. Because the model rendering is separated from other calculations, it is relatively easy to add a new model to explore a new vehicle or mission configuration. Adding a new model is performed by adding OpenGL code, but a future version might read a mesh file exported from a computer-aided design (CAD) system to enable very rapid turnaround for new designs

Wright, Theodore W. II↗

Development of Guidelines for In-Situ Repair of SLS-Class Composite Flight Hardware

The purpose of composite repair development at KSC (John F. Kennedy Space Center) is to provide support to the CTE (Composite Technology for Exploration) project. This is a multi-space center effort with the goal of developing bonded joint technology for SLS (Space Launch System) -scale composite hardware. At KSC, effective and efficient repair processes need to be developed to allow for any potential damage to composite components during transport or launch preparation. The focus of the composite repair development internship during the spring of 2018 was on the documentation of repair processes and requirements for process controls based on techniques developed through hands-on work with composite test panels. Three composite test panels were fabricated for the purpose of repair and surface preparation testing. The first panel included a bonded doubler and was fabricated to be damaged and repaired. The second and third panels were both fabricated to be cut into lap-shear samples to test the strength of bond of different surface preparation techniques. Additionally, jointed composite test panels were impacted at MSFC (Marshall Space Flight Center) and analyzed for damage patterns. The observations after the impact tests guided the repair procedure at KSC to focus on three repair methods. With a finalized repair plan in place, future work will include the strength testing of different surface preparation techniques, demonstration of repair methods, and repair of jointed composite test panels being impacted at MSFC.

Weber, Thomas P., Jr.↗

Inflatable Deployable Space Structures Technology Summary

There has been limited in inflatable deployable space structures since the 1950's due to their potential for low cost flight hardware, exceptionally high mechanical packaging efficiency, deployment reliability and low weight.

inflatable↗

Lunar Power Transmission for Fission Surface Power

This work reviews key considerations and challenges in the implementation of lunar power transmission for the 40 kW Fission Surface Power (FSP) System. The overall FSP electric power flow is presented, and potential bulk power transmission strategies are discussed. Metallic conductors operating at elevated voltage are identified as the most efficient and feasible solution, and hardware implementation challenges for AC and DC voltage step up/down are discussed. This work is intended to drive progress on identifying the most feasible power transmission strategy for landed fission power and does not represent consensus or preference by NASA regarding a particular course of action.

Lunar↗

Efficient Implementation for Unitary Coupled Cluster State Preparation for Near-Term Quantum Computers

Unitary coupled cluster theory (UCC) is a common wave function ansatz for quantum simulation of molecular electronic structure using the variational quantum eigenvalue solver (VQE). Even for small molecules using a double-ζ basis, the number of variational parameters required to minimize the electronic energy (i.e., optimize the circuit) is large and beyond the reach of current quantum computers. For example, a circuit simulating C2 using the UCCSD ansatz and the cc-pVDZ basis set with frozen-core will require over 10,000 variational parameters and a Hilbert space of over 10^8 determinants. To make progress on simulating such molecular systems on near-term quantum computers, we explore how much of the optimization can be approximately prepared with classical simulation while reducing the number of optimization steps performed on a quantum device. Recently, Chen, Cheng, and Freericks [J. Chem. Theory Comput. 2021, 17, 841-847] presented an algorithm for the factorized form of the UCC ansatz that allows for efficient UCC optimizations on classical hardware. We flip the algorithm around and use it to prepare approximate quantum circuits for systems that require a large number of qubits to represent. We will present results from our implementation and discuss strategies for incorporating this implementation for algorithms involving near-term quantum computers.

J Wayne Mullinax↗

Lunar Power Transmission for Fission Surface Power

This work reviews key considerations and challenges in the implementation of lunar power transmission for the 40 kW Fission Surface Power (FSP) System. The overall FSP electric power flow is presented, and potential bulk power transmission strategies are discussed. Metallic conductors operate at elevated voltage are identified as the most efficient and feasible solution, and hardware implementation challenges for AC and DC voltage boost are discussed. This work is intended to drive progress on identifying the most feasible power transmission strategy for landed fission power and does not represent consensus or preference by NASA regarding a particular course of action

Lunar↗

NASA Langley Aerothermodynamics Laboratory: Hypersonic Testing Capabilities

A description of the NASA Langley Research Center’s Langley Aerothermodynamics Laboratory (LAL) will be presented in the paper, along with descriptions and details of the facility test techniques and recent upgrades. The LAL consists of three hypersonic blow-down wind tunnels covering Mach numbers of 6 and 10 and unit Reynolds number ranges of 0.5 to 8.3 million per foot as well as a 60-ft Vacuum Sphere Test Chamber. LAL facilities are used to study and define the aerodynamic performance and aeroheating characteristics of flight vehicle concepts. Data collected in the facilities have been used for design and optimization, anchoring computational predictions, generation of aerodynamic databases and design of Thermal Protection Systems. Over the years modifications and enhancements have been made to the facility hardware and instrumentation to increase efficiency, data quality, capabilities and reliability to better meet the programmatic requirements. Recent utilization information illustrates the need for the capabilities associated with these facilities. Recent test programs include the Space Shuttle Program, Crew Exploration Vehicle/Orion/Multi-Purpose Crew Vehicle, Hypersonic International Flight Research Experimentation (HIFiRE), Mars Science Laboratory, Hypersonic Inflatable Aerodynamic Decelerator System (HIADS) and X-51 among others and usage has been split between NASA, Commercial Crew, Department of Defense and private company programs. Plans for future improvements to the facility infrastructure and instrumentation will also be presented.

Karen Berger↗

Efficient Implementation for Unitary Coupled Cluster State Preparation for Near-Term Quantum Computers

Unitary coupled cluster theory (UCC) is a common wave function ansatz for quantum simulation of molecular electronic structure using the variational quantum eigenvalue solver (VQE). Even for small molecules using a double-ζ basis, the number of variational parameters required to minimize the electronic energy (i.e., optimize the circuit) is large and beyond the reach of current quantum computers. For example, a circuit simulating C2 using the UCCSD ansatz and the cc-pVDZ basis set with frozen-core will require over 10,000 variational parameters and a Hilbert space of over 10^(8) determinants. To make progress on simulating such molecular systems on near-term quantum computers, we explore how much of the optimization can be approximately prepared with classical simulation while reducing the number of optimization steps performed on a quantum device. Recently, Chen, Cheng, and Freericks [J. Chem. Theory Comput. 2021, 17, 841-847] presented an algorithm for the factorized form of the UCC ansatz that allows for efficient UCC optimizations on classical hardware. We flip the algorithm around and use it to prepare approximate quantum circuits for systems that require a large number of qubits to represent. We will present results from our implementation and discuss strategies for incorporating this implementation for algorithms involving near-term quantum computers.

Quantum Computing↗

Efficient Implementation for Unitary Coupled Cluster State Preparation for Near-Term Quantum Computers

Unitary coupled cluster theory (UCC) is a common wave function ansatz for quantum simulation of molecular electronic structure using the variational quantum eigenvalue solver (VQE). Even for small molecules using a double-ζ basis, the number of variational parameters required to minimize the electronic energy (i.e., optimize the circuit) is large and beyond the reach of current quantum computers. For example, a circuit simulating C2 using the UCCSD ansatz and the cc-pVDZ basis set with frozen-core will require over 10,000 variational parameters and a Hilbert space of over 10^(8) determinants. To make progress on simulating such molecular systems on near-term quantum computers, we explore how much of the optimization can be approximately prepared with classical simulation while reducing the number of optimization steps performed on a quantum device. Recently, Chen, Cheng, and Freericks [J. Chem. Theory Comput. 2021, 17, 841-847] presented an algorithm for the factorized form of the UCC ansatz that allows for efficient UCC optimizations on classical hardware. We flip the algorithm around and use it to prepare approximate quantum circuits for systems that require a large number of qubits to represent. We will present results from our implementation and discuss strategies for incorporating this implementation for algorithms involving near-term quantum computers.

Quantum Computing↗

Towards Generic Parallel Programming in Computer Science Education with Kokkos

Parallel patterns, views, and spaces are promising abstractions to capture the programmer's intent as well as the contextual information that can be used by an underlying runtime to efficiently map software to parallel hardware. These abstractions can be valuable in cases where an algorithm must accommodate requirements of code and performance portability across hardware architectures and vendor programming models. Kokkos is a parallel programming model for host- and accelerator architectures that relies on these abstractions and targets these requirements. It consists of a pure C++ interface, a specification, and a programming library. The programming library exposes patterns and types and maps them to an underlying abstract machine model. The abstract machine model offers a generic view of parallel hardware. While Kokkos is gaining popularity in large-scale HPC applications at some DOE laboratories, we believe that the implemented concepts are of interest to a broader audience including academia as they may contribute to a generic, vendor, and architecture-independent education of parallel programming. In this work, we give an insight into the design considerations of this programming model and list important abstractions. Further, we document best practices obtained from giving virtual classes on Kokkos and give pointers to resources that the reader may consider valuable for a lecture on generic parallel programming for students with preexisting knowledge on this matter.

Ciesko, Jan↗

Mastering HPC Runtime Prediction: From Observing Patterns to a Methodological Approach: Preprint

The continual expansion of high-performance computing (HPC) brings with it an increasing need for efficiency. Heavy investment in energy, hardware, and software infrastructure to support peta- and exascale computing requires the optimization of existing systems and, wherever possible, the discernment and adoption of best-practices towards these goals. Such is the case for runtime prediction. When a job is submitted to an HPC system, an estimate of its runtime is provided by the user in the form of "requested wallclock''. Error in this user-provided estimate can lead to jobs being prematurely killed by the scheduler, increased wait time on the queue, and decreased system utilization. More than fifteen years of research has been directed at mitigating these effects by using data-driven runtime predictions. Codified here is a set of commonalities and insights emerging from this body of work, which we present as recommendations and best practices. These practices are combined into a methodological approach described and evaluated on an 11-million-job dataset from the National Renewable Energy Laboratory's petascale HPC system, Eagle. This dataset and the accompanying codebase have been released to the public domain for the benefit of the wider HPC research community.

high performance computing↗

Status and future prospects of using numerical methods to study complex flows at High Reynolds numbers

The calculation of flow fields past aircraft configuration at flight Reynolds numbers is considered. Progress in devising accurate and efficient numerical methods, in understanding and modeling the physics of turbulence, and in developing reliable and powerful computer hardware is discussed. Emphasis is placed on efficient solutions to the Navier-Stokes equations.

Maccormack, R. W.↗

Virtualized Logical Qubits: A 2.5D Architecture for Error-Corrected Quantum Computing

Current, near-term quantum devices have shown great progress in the last several years culminating recently with a demonstration of quantum supremacy. In the medium-term, however, quantum machines will need to transition to greater reliability through error correction, likely through promising techniques like surface codes which are well suited for near-term devices with limited qubit connectivity. We discover quantum memory, particularly resonant cavities with transmon qubits arranged in a 2.5D architecture, can efficiently implement surface codes with substantial hardware savings and performance/fidelity gains. Specifically, we virtualize logical qubits by storing them in layers of qubit memories connected to each transmon. Surprisingly, distributing each logical qubit across many memories has a minimal impact on fault tolerance and results in substantially more efficient operations. Our design permits fast transversal application of CNOT operations between logical qubits sharing the same physical address (same set of cavities) which are 6x faster than standard lattice surgery CNOTs. We develop a novel embedding which saves approximately 10x in transmons with another 2x savings from an additional optimization for compactness. Although qubit virtualization pays a 10x penalty in serialization, advantages in the transversal CNOT and in area efficiency result in fault-tolerance and performance comparable to conventional 2D transmon-only architectures. Our simulations show our system can achieve fault tolerance comparable to conventional two-dimensional grids while saving substantial hardware. Furthermore, our architecture can produce magic states at 1.22x the baseline rate given a fixed number of transmon qubits. Here, this is a critical benchmark for future fault-tolerant quantum computers as magic states are essential and machines will spend the majority of their resources continuously producing them. This architecture substantially reduces the hardware requirements for fault-tolerant quantum computing and puts within reach a proof-of-concept experimental demonstration of around 10 logical qubits, requiring only 11 transmons and 9 attached cavities in total.

quantum computing↗

Fast and Adaptive Lossless Onboard Hyperspectral Data Compression System

Modern hyperspectral imaging systems are able to acquire far more data than can be downlinked from a spacecraft. Onboard data compression helps to alleviate this problem, but requires a system capable of power efficiency and high throughput. Software solutions have limited throughput performance and are power-hungry. Dedicated hardware solutions can provide both high throughput and power efficiency, while taking the load off of the main processor. Thus a hardware compression system was developed. The implementation uses a field-programmable gate array (FPGA). The implementation is based on the fast lossless (FL) compression algorithm reported in Fast Lossless Compression of Multispectral-Image Data (NPO-42517), NASA Tech Briefs, Vol. 30, No. 8 (August 2006), page 26, which achieves excellent compression performance and has low complexity. This algorithm performs predictive compression using an adaptive filtering method, and uses adaptive Golomb coding. The implementation also packetizes the coded data. The FL algorithm is well suited for implementation in hardware. In the FPGA implementation, one sample is compressed every clock cycle, which makes for a fast and practical realtime solution for space applications. Benefits of this implementation are: 1) The underlying algorithm achieves a combination of low complexity and compression effectiveness that exceeds that of techniques currently in use. 2) The algorithm requires no training data or other specific information about the nature of the spectral bands for a fixed instrument dynamic range. 3) Hardware acceleration provides a throughput improvement of 10 to 100 times vs. the software implementation. A prototype of the compressor is available in software, but it runs at a speed that does not meet spacecraft requirements. The hardware implementation targets the Xilinx Virtex IV FPGAs, and makes the use of this compressor practical for Earth satellites as well as beyond-Earth missions with hyperspectral instruments.

Aranki, Nazeeh I.↗

Methodology for Assessing Reusability of Spaceflight Hardware

In 2011 the Space Shuttle, the only Reusable Launch Vehicle (RLV) in the world, returned to earth for the final time. Upon retirement of the Space Shuttle, the United States (U.S.) no longer possessed a reusable vehicle or the capability to send American astronauts to space. With the National Aeronautics and Space Administration (NASA) out of the RLV business and now only pursuing Expendable Launch Vehicles (ELV), not only did companies within the U.S. start to actively pursue the development of either RLVs or reusable components, but entities around the world began to venture into the reusable market. For example, SpaceX and Blue Origin are developing reusable vehicles and engines. The Indian Space Research Organization is developing a reusable space plane and Airbus is exploring the possibility of reusing its first stage engines and avionics housed in the flyback propulsion unit referred to as the Advanced Expendable Launcher with Innovative engine Economy (Adeline). Even United Launch Alliance (ULA) has announced plans for eventually replacing the Atlas and Delta expendable rockets with a family of RLVs called Vulcan. Reuse can be categorized as either fully reusable, the situation in which the entire vehicle is recovered, or partially reusable such as the National Space Transportation System (NSTS) where only the Space Shuttle, Space Shuttle Main Engines (SSME), and Solid Rocket Boosters (SRB) are reused. With this influx of renewed interest in reusability for space applications, it is imperative that a systematic approach be developed for assessing the reusability of spaceflight hardware. The partially reusable NSTS offered many opportunities to glean lessons learned; however, when it came to efficient operability for reuse the Space Shuttle and its associated hardware fell short primarily because of its two to four-month turnaround time. Although there have been several attempts at designing RLVs in the past with the X-33, Venture Star and Delta Clipper Experimental (DC-X), reusability within the spaceflight arena is still in its infancy. With unlimited resources (namely, time and money), almost any launch vehicle and its associated hardware can be made reusable. However, an endless supply of funds for space exploration is not the case in today's economy for neither government agencies nor their commercial counterparts. Therefore, any organization wanting to be a leader in space exploration and remain competitive in this unforgiving space faring industry must confront shrinking budgets with more cost conscious and efficient designs. Therefore, standards for developing reusable spaceflight hardware need to be established. By having standards available to existing and emerging companies, some of the potential roadblocks and limitations that plagued previous attempts at reuse may be minimized or completely avoided.

Childress-Thompson, Rhonda↗

Avoiding pitfalls in simulating real-time computer systems

The software simulation of a computer target system on a computer host system, known as an interpretive computer simulator (ICS), functionally models and implements the action of the target hardware. For an ICS to function as efficiently as possible and to avoid certain pitfalls in designing an ICS, it is important that the details of the hardware architectural design of both the target and the host computers be known. This paper discusses both host selection considerations and ICS design features that, without proper consideration, could make the resulting ICS too slow to use or too costly to maintain and expand.

Smith, R. S.↗

Leveraging dendritic complexity for neuromorphic computing

Abstract Beyond-von Neumann computing approaches are necessary to sustain the growth of microelectronics and the increasing appetite for artificial intelligence/machine learning algorithms. Neuromorphic computing is an emerging paradigm that takes inspiration from the brain to provide a path forward to improve the computational efficiency and computational density of next-generation computing architectures. In nature, we observe brains performing complex computations with a much smaller energy footprint than conventional computing approaches. Current neuromorphic systems are focused primarily on scalability, namely, increasing the number of computational units (neurons) and connections between units (synapses). However, for brain-like cognition and efficiency in next-generation computing hardware, we need increased complexity in function, as well as improved connection density for scalability. Here, we present our work that aims to incorporate dendrites for ‘compute-on-wire’ in neuromorphic architectures to increase the computational complexity (e.g. number of programmable parameters, nonlinear dynamics) as well as computational efficiency (energy/compute) of artificial neural networks (ANNs). We do this by showcasing neuromorphic dendrite elements that can be leveraged for various applications. We will present examples of neuroscience-inspired direction-selective circuits and an ANN with active dendrites leveraging shunting inhibition. We also demonstrate the benefits of using dendrites in deep neural networks. To conclude, we discuss how we can utilize emerging hardware devices in these systems and design next-generation neuromorphic architectures with dendrites.

Cardwell, Suma G. (ORCID:0000000226575545)↗

Hamilton: Flexible, Open Source $10 Wireless Sensor System for Energy Efficient Building Operation

Sensors for improving building performance are rapidly populating the market, driven in part by the drive to reduce greenhouse gas emissions resulting from energy production as well as improve the interior environment for healthy and more productive spaces. UC Berkeley has led wireless sensor development over the past 25 years (e.g., Telos mote), with the Hamilton (named after Alexander Hamilton on the US $10 bill) as the most recent. The Hamilton sensor was designed as a low-cost high-performance sensor that is modular and interoperable. The objective of the Hamilton project was to create, evaluate and establish the technological foundations for secure and easy to deploy building energy efficiency applications utilizing pervasive, low-cost wireless sensors integrated with traditional Building Management Systems (BMS), consumer-sector building components, and powerful data analytics. The project included iterative hardware design, incorporating a high-performance database (BTrDb, http://btrdb.io/), creating and iterating the development of secure data middleware (BOSSwave, WAVE/WAVEMQ), working with and pushing the development of an open-source tiny operating system RiotOS, and implementing and improving protocols such as Thread/OpenThread and TCP/IP. The hardware benefited from careful design to drive down the cost; the design included a System-on-a-Chip (SoC), chip antenna, single crystal and five passive components. Careful design of the operating system created a low-power design to enable a long life with small batteries. The hardware included several sensors: temperature, radiant temperature, relative humidity, magnetometer, accelerometer, and light, with an optional occupancy (Passive InfraRed) sensor. The project was the basis of several applications, both internal to the research team and other researchers and professionals at other institutions. Several applications used the sensor hardware as the basis for other complex devices. Other applications used the sensors to improve building performance through interoperating with the building Heating Ventilation and Air-Conditioning (HVAC) system, such as using occupancy and/or distributed temperature sensing to reduce HVAC zone energy while still providing thermal comfort and to reduce peak loads in small commercial buildings. We demonstrated cloud-based energy analytics, implemented a schedule and a Model Predictive Controller in a small commercial building to optimize HVAC energy, occupancy and electricity price. Initial integration of these technological innovations was performed through the creation of execution containers containing the WAVE agent and various driver, proxy, or building system function logic. The research added to the understanding of efficient sensor hardware, secure middleware, time-series data management (high performance database), efficient communication protocols, and interoperating with applications and building systems. The project showed the technical effectiveness and economic feasibility of creating a low-cost, modular, and easy-to-deploy sensor. Through conversations with multiple end users, the research team discovered that many customers wanted data management and services in addition to the sensors. HamiltonIOT developed packages of sensors, border router, and data services to provide a seamless “plug-and-play” sensor deployment. Some customers were willing to pay for higher quality sensors (such as light); some customers wanted a robust enclosure (waterproof).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗