Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Geometry Optimization of Cable-Based Actuation for Small-Scale Model Testing of a Floating Marine Turbine: Preprint

This research aims to apply combined wave and tidal current loads to a small-scale floating marine current turbine in a wave tank, where an actuation system applies hydrodynamic and mooring forces on the hardware based on results from a simulation. We use a real-time hybrid test setup with physical wave forcing from the wave tank and simulated current and mooring forces implemented through a tensioned cable array. For this paper, we developed an optimization algorithm that adjusts the hardware geometry of the cable array to achieve more efficient tension allocation across the cables for the loads that need to be actuated. By adjusting the points where the cables are attached to the floating platform and the angles between the platform and the winches, an optimal cable geometry can be found to minimize tension variations in the lines, maintain the desired pretension, and prevent excessive tensions. We present the optimization problem formulation, the actuation system evaluation approach, and optimization results that show effective cable actuation setups that are being considered for implementation in the wave tank tests.

cable-based actuation↗

Measurement of the Neutron Electromagnetic Form Factor Ratio at High Momentum Transfer

The inner structure of the nucleon (proton and neutron) remains a topic of great interest in nuclear and particle physics, after many decades of study. For example, understanding the quark-gluon dynamics inside the nucleon would shed light on how 99% of the nucleon mass is created. The neutron electromagnetic form factors, Gn E and Gn M , give important insights into the neutron structure. The Super BigBite Spectrometer (SBS) program at Jefferson Lab (JLab) seeks to extend the form factor measurements for both the proton and the neutron. The neutron electric form actor, Gn E , has been historically difficult to measure due to the short lifetime of the free neutron and the small value of Gn E . The GEn-II experiment is part of the SBS program and seeks to measure Gn E , significantly increasing the high momentum transfer coverage. A newly designed polarized 3He target increased the figure of merit by three times compared to previous measurements. The analysis of this data is especially challenging due to the unprecedented high-rate environment caused by the open nature of the spectrometer with a direct line of sight to the target. This required developing new Gas Electron Multiplier (GEM) particle trackers which can cover large areas demanded by this setup and handle particle rates up to 500 kHz/cm2. Rates this high over a large area is unprecedented in particle tracking systems and came with a number of challenges. Data taken in the SBS program was critical to understanding hardware and software solutions that improved the track reconstruction efficiency to be >97% with a position resolution of 70 ?m. In previous experiments the proton electromagnetic form factors, Gp E and Gp M were measured up to Q2 = 8.5 GeV2 and Q2 = 30 GeV2, respectively, while Gn E has only been measured up to Q2 = 3.4 GeV2. The GEn-II experiment has measured the neutron form factor ratio, Gn E/Gn M, at Q2 values of 2.90, 6.50, and 9.47 GeV2 by scattering a polarized electron beam with a polarized 3He target, used here as an effective polarized neutron target, and measuring the double spin asymmetry of the cross section. Previous Gn E measurements do not extend above Q2 = 3.4 GeV2, and therefore this analysis has extended the world data by almost three times. The background correction is especially difficult at the higher Q2 settings leading to large systematic errors. As very exploratory results from this early analysis of the data, we find for Q2 = 2.90 GeV2, Gn E = 0.0157 ±stat 0.0016 ±sys 0.0011, for Q2 = 6.50 GeV2, Gn E = 0.0067 ±stat 0.0019 ±sys 0.0005, and for Q2 = 9.46 GeV2, Gn E = 0.0046 ±stat 0.0023 ±sys 0.0005. These results are compared to predictions from the Dyson-Schwinger Equations (DSE) model and a Relativistic Constituent Quark Model (RCQM).

Jeffas, Sean↗

Machine Learning-Assisted Stability Boundary Determination of Multiport Autonomous Reconfigurable Solar Power Plants

The multiport autonomous reconfigurable solar (MARS) power plant is a promising solution to integrate renewable resources and energy storage systems into the alternating current (ac) power grid and an high-voltage direct current (HVdc) link. In the MARS system, various input power sources are connected to the individual submodules (SMs) through direct current (dc)–dc converters. However, the presence of external power sources can result in unbalanced capacitor voltages of SMs, thereby violating stability constraints under multiple/diverse operating conditions. This article aims to address the gap by accurately determining the stability boundary of the MARS system. As such, a novel machine learning (ML)-assisted energy balancing control (EBC) criterion is proposed. Further, in conjunction with a refined EBC, this approach ensures balanced capacitor voltages across various types of SMs, significantly enhancing the overall system efficiency. The proposed EBC criterion effectively controls EBC activation and deactivation, achieving remarkable accuracy. Both power systems computer aided design (PSCAD)/electromagnetic transients including direct current (EMTDC) simulations and control hardware-in-the-loop (cHIL) tests are conducted to validate the feasibility and efficiency of the proposed method. By combining the EBC and ML-assisted EBC criterion, efficient energy management is achieved for systems featuring multiple input power sources, such as MARS. This approach enables the system to fully exploit its potential across an expanded operational range while upholding high-efficiency standards.

14 SOLAR ENERGY↗

Flexible AI Models for Grid Resilience

The rapid growth in size and complexity of artificial intelligence (AI) and machine learning (ML) models has led to increased energy demands, posing a threat to the reliability of the existing power grid. This project addresses the challenge of highly intermittent and energy-intensive inference workloads by (1) developing fidelity-adaptive neural networks capable of dynamic response to grid conditions and (2) integrating these networks with power flow simulations to assess their impact on power grid reliability. We will explore both top-down and bottom-up approaches to create hierarchies of submodels that provide a controlled trade-off between power draw and prediction accuracy. The top-down method utilizes NN pruning to reduce a flagship model into progressively smaller, energy-efficient variants. The bottom-up approach employs geometrically principled weight setting strategies to construct depth-efficient models from the ground up. A real-time hardware-in-the-loop (HIL) platform will be developed to simulate a scaled AC power grid, integrating live AI workload power draw and enabling dynamic model switching in response to grid feedback. This work will provide a novel framework for evaluating the impact of flexible AI/ML workloads on grid performance and establish new methodologies for energy-aware computing in data centers. The outcomes will demonstrate that adaptive AI/ML can play a critical role in improving grid stability while advancing NREL's leadership in energy-efficient computing research.

24 POWER TRANSMISSION AND DISTRIBUTION↗

SODA: a New Synthesis Infrastructure for Agile Hardware Design of Machine Learning Accelerators

Next generation systems, such as edge devices, will have to provide efficient processing of machine learning (ML) algorithms along several metrics, including energy, performance, area, and latency. However, the quickly evolving field of ML makes it extremely difficult to generate accelerators able to support a wide variety of algorithms. At the same time, designing accelerators in hardware description languages (HDLs) by hand is hard and time consuming, and does not allow quick exploration of the design space. This paper discusses the SODA synthesizer, an automated open source high-level ML framework-to-Verilog compiler targeting ML Application-Specific Integrated Circuits (ASICs) chiplets based on the LLVM infrastructure. The SODA synthesizers will allow implementing optimal designs by combining templated and fully tunable IPs and macros, and fully custom components generated through high-level synthesis. All these components will be provided through an extendable resource library, characterized with both commercial and open source logic design flows. Through a closed loop design space exploration engine, developers will be able to quickly explore their hardware designs along different dimension

Minutoli, Marco↗

Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities

Although deep learning has made great progress in recent years, the exploding economic and environmental costs of training neural networks are becoming unsustainable. To address this problem, there has been a great deal of research on algorithmically-efficient deep learning, which seeks to reduce training costs not at the hardware or implementation level, but through changes in the semantics of the training program. In this paper, we present a structured and comprehensive overview of the research in this field. First, we formalize the algorithmic speedup problem, then we use fundamental building blocks of algorithmically efficient training to develop a taxonomy. Our taxonomy highlights commonalities of seemingly disparate methods and reveals current research gaps. Next, we present evaluation best practices to enable comprehensive, fair, and reliable comparisons of speedup techniques. To further aid research and applications, we discuss common bottlenecks in the training pipeline (illustrated via experiments) and offer taxonomic mitigation strategies for them. Finally, we highlight some unsolved research challenges and present promising future directions.

97 MATHEMATICS AND COMPUTING↗

Blade Designs For Improved Multi-Phase Performance In sCO 2 Compressors; Part II - Optical Diagnostics In sCO 2 And Experimental Evaluation With Particle Image Velocimetry

This paper presents the second part of a study in which the leading-edge and suction surface of a compressor blade was modified to delay onset of phase change for sCO 2 compressors operating near the critical point. Using a first-of-its-kind apparatus for the measurement of sCO 2 flow fields, Particle Image Velocimetry (PIV) is used for local flow field measurements of two compressor blade geometries: the modified “biased-wedge,” and a conventional constant thickness blade. Utilizing the developed hardware, the feasibility of a simple, laser-based diagnostic for qualitatively measuring liquid phase regions, is also presented. The design of the optical diagnostics rig, a discussion of numerous challenges, and necessary considerations involved in performing optical-based measurements like PIV, in sCO 2 , are discussed. Velocity field measurements for the modified compressor profile show a much lower suction peak compared to a conventional blade. Furthermore, these results validate numerical results at the tested conditions, where the suction side profile of the biased wedge works to minimize the local pressure gradient.

14 SOLAR ENERGY↗

Coupling Noah-Multiparameterization land-surface Model with Energy Research and Forecasting Model

The Energy Research and Forecasting (ERF) model is a high-performance atmospheric model built on the AMReX adaptive mesh refinement (AMR) framework, enabling efficient simulations on heterogeneous computing platforms that combine multicore processors with hardware accelerators. To support land–atmosphere interactions within ERF’s AMR-based environment, a land-surface model must be capable of operating directly on hierarchically refined meshes. In this work, we present a methodology for coupling the Fortran-based Noah-Multiparameterization (Noah-MP) land-surface model with ERF’s C++ codebase. Rather than rewriting Noah-MP, we construct a Fortran–C interoperability layer using CodeScribe, a tool that leverages large language models (LLMs) to automate the generation of interface code. CodeScribe applies structured prompting techniques to generate bindings that support efficient data exchange and function calls between ERF and Noah-MP. The coupling framework also incorporates AMR-aware data handling strategies, allowing NoahMP to operate seamlessly within ERF’s hierarchical mesh structure. This work provides a structured approach for integrating legacy Fortran models into modern C++-based modeling systems using LLM-assisted code generation.

54 ENVIRONMENTAL SCIENCES↗

Inferring Quantum Network Topology Using Local Measurements

Statistical correlations that can be generated across the nodes in a quantum network depend crucially on its topology. However, this topological information might not be known a priori, or it may need to be verified. In this paper, we propose an efficient protocol for distinguishing and inferring the topology of a quantum network. We leverage entropic quantities-namely, the von Neumann entropy and the measured mutual information-as well as measurement covariance to uniquely characterize the topology. We show that the entropic quantities are sufficient to distinguish two networks that prepare GHZ states. Moreover, if qubit measurements are available, both entropic quantities and covariance can be used to infer the network topology without state-preparation assumptions. We show that the protocol can be entirely robust to noise and can be implemented via quantum variational optimization. Numerical experiments on both classical simulators and quantum hardware show that covariance is generally more reliable for accurately and efficiently inferring the topology, whereas entropy-based methods are often better at identifying the absence of entanglement in the low-shot regime.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An artificial spiking quantum neuron

Abstract Artificial spiking neural networks have found applications in areas where the temporal nature of activation offers an advantage, such as time series prediction and signal processing. To improve their efficiency, spiking architectures often run on custom-designed neuromorphic hardware, but, despite their attractive properties, these implementations have been limited to digital systems. We describe an artificial quantum spiking neuron that relies on the dynamical evolution of two easy to implement Hamiltonians and subsequent local measurements. The architecture allows exploiting complex amplitudes and back-action from measurements to influence the input. This approach to learning protocols is advantageous in the case where the input and output of the system are both quantum states. We demonstrate this through the classification of Bell pairs which can be seen as a certification protocol. Stacking the introduced elementary building blocks into larger networks combines the spatiotemporal features of a spiking neural network with the non-local quantum correlations across the graph.

Kristensen, Lasse Bjørn (ORCID:0000000239398170)↗

Adaptive hyperparameter updating for training restricted Boltzmann machines on quantum annealers

Restricted Boltzmann Machines (RBMs) have been proposed for developing neural networks for a variety of unsupervised machine learning applications such as image recognition, drug discovery, and materials design. The Boltzmann probability distribution is used as a model to identify network parameters by optimizing the likelihood of predicting an output given hidden states trained on available data. Training such networks often requires sampling over a large probability space that must be approximated during gradient based optimization. Quantum annealing has been proposed as a means to search this space more efficiently which has been experimentally investigated on D-Wave hardware. D-Wave implementation requires selection of an effective inverse temperature or hyperparameter (β) within the Boltzmann distribution which can strongly influence optimization. Here, we show how this parameter can be estimated as a hyperparameter applied to D-Wave hardware during neural network training by maximizing the likelihood or minimizing the Shannon entropy. We find both methods improve training RBMs based upon D-Wave hardware experimental validation on an image recognition problem. Neural network image reconstruction errors are evaluated using Bayesian uncertainty analysis which illustrate more than an order magnitude lower image reconstruction error using the maximum likelihood over manually optimizing the hyperparameter. The maximum likelihood method is also shown to out-perform minimizing the Shannon entropy for image reconstruction.

97 MATHEMATICS AND COMPUTING↗

A Photovoltaic MPPT Charge Controller Real-Time Testbed for Cybersecurity Applications

The increasing deployment of distributed energy resources (DER) over the last decade is a great ally to combat climate change and strengthen the grid during increasingly common extreme weather events. However, DER systems, combined with the ongoing transition to a digital power grid, also pose substantial cybersecurity threats. One of the most common communication protocols used in DER integration is the Distributed Network Protocol 3 (DNP3), which is known to have many security vulnerabilities. Thus, it is essential to investigate cyberattack behaviors and mitigation on power systems using DNP3. In this paper, we designed and implemented a cybersecurity testbed for a simulated photovoltaic (PV) maximum power point tracking (MPPT) charge controller. Our testbed uses an MPPT charge controller simulated on a Typhoon HIL602+ real-time simulator with a real DNP3 communication connection over TCP/IP, allowing for safe and efficient monitoring and manipulation of data traffic between the simulated hardware and supervisory control and data acquisition (SCADA) systems.

14 SOLAR ENERGY↗

Experimental Validation of Subsystem Models for a Novel Variable Displacement Hydraulic Motor

A novel, variable displacement, low-speed high-torque hydraulic motor is being developed that is expected to be highly efficient across a broad operating range. To ensure the final hardware achieves the expected performance, the models used in the development of the motor must be experimentally validated and revised, as necessary. The specific focus in this work is on mechanical energy loss models that were used to guide the design of a single-cylinder motor prototype and on experimental tests used for model validation. Ideally each model, whether friction loss in a piston/cylinder interface or energy loss due to leakage in a valve , would be individually validated by an independent test. This granular approach would remove any question of where the error lies and result in highly accurate models. However, many motor components have multiple forms of energy loss, creating difficulty in validating individual losses. Additionally, it is not physically realizable to divide many components into an individual model equivalent test. Testing the motor as a full assembly is possible, but pinpointing the source of discrepancies between the model and the hardware becomes difficult with dozens of models potentially being partially responsible. A compromise was found by separating the motor into functional component groups that are characterized by the type of loss and ability to test each subcomponent independently. By checking for correlation between test observations and model predictions, revisions could be implemented into the models. This allows future solutions to be more accurately predicted in the design phase to drive the design of better machines.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toward Large-Scale Image Segmentation on Summit

Semantic segmentation of images is an important computer vision task that emerges in a variety of application domains such as medical imaging, robotic vision and autonomous vehicles to name a few. While these domain-specific image analysis tasks involve relatively small image sizes (~ 10 2 × 10 2 ), there are many applications that need to train machine learning models on image data with extents that are orders of magnitude larger (~10 4 × 10 4 ). Training deep neural network (DNN) models on large extent images is extremely memory-intensive and often exceeds the memory limitations of a single graphical processing unit, a hardware accelerator of choice for computer vision workloads. Here, an efficient, sample parallel approach to train U-Net models on large extent image data sets is presented. Its advantages and limitations are analyzed and near-linear strong-scaling speedup demonstrated on 256 nodes (1536 GPUs) of the Summit supercomputer. Using a single node of the Summit supercomputer, an early evaluation of a recently released model parallel framework called GPipe is demonstrated to deliver ~ 2X speedup in executing a U-Net model with an order of magnitude larger number of trainable parameters than reported before. Performance bottlenecks for pipelined training of U-Net models are identified and mitigation strategies to improve the speedups are discussed. Together, these results open up the possibility of combining both approaches into a unified scalable pipelined and data parallel algorithm to efficiently train U-Net models with very large receptive fields on data sets of ultra-large extent images.

Seal, Sudip↗

Real-time evaluation of cybersecurity threats to DER inverter grid-support functions

In this project we aim to contribute to the understanding of the type and severity of potential cybersecurity attacks to the grid-support functionalities of DER systems interconnected to the AC distribution grid via inverters. Our preliminary work focused on developing a small-scale testbed allowing to study cybersecurity threats to an isolated photovoltaic-battery system using a real-time simulator (Typhoon HIL602+) with a real DNP3 communication connection over TCP/IP, allowing for safe and efficient monitoring and manipulation of data traffic between the simulated hardware and supervisory control and data acquisition (SCADA) system. In this project we propose to expand upon this development by utilizing a) a recently acquired NovaCor RTDS (Real Rime digital Simulator) to emulate the DER-inverter-grid topology including main grid-support functions as defined by IEEE Std. 1547-2018, and b) an industrial control and automation device to enable realistic evaluation of control functions and utilization of communication protocols for real-time data transmission.

24 POWER TRANSMISSION AND DISTRIBUTION↗

EJFAT Scientific Perspective

Presented new computing model to the test by deploying the EJFAT system alongside a data-stream processing framework running the production-level CLAS12 event reconstruction application. In this experiment, a continuous stream of CLAS12 Level-1 identified events was processed in real-time using the EJFAT load balancer, distributing the workload across 90 computing nodes located across the U.S. This marks the first-ever large-scale, real-time distributed data stream processing experiment, demonstrating that scientific data-streaming pipelines can efficiently scale across four dimensions, thanks to EJFAT’s advanced hardware and software capabilities.

Gyurjyan, Vardan [Thomas Jefferson National Accele↗

Will Stochastic Devices Play Nice With Others in Neuromorphic Hardware?: There’s More to a Probabilistic System Than Noisy Devices

Achieving brain-like efficiency in computing requires a co-design between the development of neural algorithms, brain-inspired circuit design, and careful consideration of how to use emerging devices. The recognition that leveraging device-level noise as a source of controlled stochasticity represents an exciting prospect of achieving brain-like capabilities in probabilistic neural algorithms, but the reality of integrating stochastic devices with deterministic devices in an already-challenging neuromorphic circuit design process is formidable. Here, we explore how the brain combines different signaling modalities into its neural circuits as well as consider the implications of more tightly integrated stochastic, analog, and digital circuits. Further, by acknowledging that a fully CMOS implementation is the appropriate baseline, we conclude that if mixing modalities is going to be successful for neuromorphic computing, it will be critical that device choices consider strengths and limitations at the overall circuit level.

42 ENGINEERING↗

Predicting Biological Cleanliness: An Empirical Bayes Approach for Spacecraft Bioburden Accounting

To comply with the international planetary protection policy set forth by the Committee on Space Research and NASA Agency level requirements, spacecraft destined to biologically sensitive planetary bodies have to minimize terrestrial biological contamination. Analysis, testing and inspection are the standard forward verification activities that are used to demonstrate compliance with the biological contamination requirements. For testing of spacecraft surface areas, a swab or wipe sample is collected from surfaces prior to last access and subsequently processed in the lab using NASA Approved Planetary Protection Methods for Culture Based Assays. Raw data resulting from this assay is then statistically treated employing a mathematical paradigm stemming from the 1970’s Viking Lander Project to generate the bioburden density and total microbial bioburden present. This standard approach arbitrarily accounts for error and provides an upper conservative bound as it reports the maximum number of spores estimated to be present on flight hardware surfaces. A bioburden density estimate factors in the following variables: the observed bioburden count, representative volume processed, sampling efficiencies. Notably, to account for error in the approach, a 0 observed count is arbitrarily changed to a count of 1 for each hardware grouping. The data generated by spacecraft bioburden verification campaigns in the past have resulted in <80% of wipes and <90% of swabs containing a bioburden count of 0. As such, having a robust and well documented statistical approach for dealing with the probability of low incident rates is necessary to be able to estimate spacecraft bioburden. Being able to statistically describe the bioburden distribution and associated confidence level is a gamechanger for the development of bioburden allocations during mission design and will allow for tighter management of risk throughout spacecraft build. Thus, Empirical Bayes statistical approach was evaluated to estimate the microbial bioburden on spacecraft to mitigate the aforementioned mathematical concerns and provide a probabilistic bioburden distribution of the flight hardware surface. For application of this approach to performing bioburden calculations, a range of non-informative prior assumptions on hardware surfaces are explored for Bayesian analyses while informative priors using posterior distributions from prior assays are utilized for Empirical Bayes analyses. Several non-informative priors are currently under investigation to assess fitness including use of these priors to serve as a foundation to build off of NASA specification values or a basis of risk to account for unknowns during the integration and testing process. Informative priors under consideration are generated using sampled bioburden values from hardware originating within like processing environments (e.g. vendor cleaning process or similar assembly process), temporal spacecraft status events as a prediction for hardware cleanliness of future samples, and heritage system bioburden actuals to predict allocation for subsequent missions. Informative priors and probabilistic bioburden distributions are then validated using data sets from the Mars Exploration Rover, Mars Science Laboratory, and InSight missions. Using Empirical Bayes approach to generate a probabilistic bioburden distribution as demonstrated through mission use cases provides a valid approach for use in the end-to-end requirements verification process.

97 MATHEMATICS AND COMPUTING↗