Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Computational Efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Efficient Computation of Doppler-Broadened Elastic Scattering Kernel Moments Using Ladder-Operator Formulation

Anefficient routine for computing Legendre moments of the Doppler-broadened elastic scattering kernel, including resonance scattering effects, has been implemented in the ISOXML module of Griffin. Isotropic scattering in the center-of-mass system and the ideal gas model for target motion are assumed. A ladder-operator formulation is introduced to compute all Legendre moments from order 0 to N simultaneously, enabling near-linear scaling of computational cost with respect to the maximum Legendre order. A physics-based strategy for constructing outgoing energy grids has also been developed, in which a tailored base grid is combined with adaptive refinement to maintain accuracy while limiting the number of outgoing energy points. For energies between resonances, a constant cross-section model is employed to further reduce computational cost. In addition, a quantitative criterion is derived to determine isotope-wise cut-off incident energies based on a prescribed up-scattering probability coverage. For 238U, up to incident energies of approximately 75, 230, and 661 eV at 294, 900, and 2500 K (corresponding to a 2% up-scattering probability threshold), computation of P0 kernels requires 1–8 s and computation of P0–P5 kernels requires 0.4–4 min using a single thread, while maintaining 1–3% relative error in up-scattering probability. These results demonstrate that the proposed formulation enables accurate and computationally practical Doppler-broadened kernel generation for online multigroup cross-section production in Griffin.

Doppler-broadening

Decode the Workload: Training Deep Learning Models for Efficient Compute Cluster Representation

Monitoring the status of a high throughput computing cluster running computationally intensive production jobs is a crucial yet challenging system administration task due to the complexity of such systems. To this end, we train autoencoders using the Linux kernel CPU metrics of the cluster. Additionally, we explore assisting these models with graph neural networks to share information across threads within a compute node. The models are compared in terms of their ability to: 1) Produce a compressed latent representation that captures the salient features of the input, 2) Detect anomalous activity, and 3) Make distinction between different kinds of jobs run at Jefferson Lab. The goal is to have a robust encoder whose compressed embeddings are used for several downstream tasks. We extend this study further by deploying these models in a human-in-the-loop production-based setting for the anomaly detection task and discuss the associated implementation aspects such as continual learning and the criterion to generate alarms. This study represents a first step in the endeavor towards building self-supervised large-scale foundation models for computing centers.

Mohammed, Ahmed

Efficiently Computable Limits on EPR Pair Generation in Quantum Broadcast Channels

We investigate the generation of EPR pairs between three observers in a general causally structured setting, where communication occurs via a noisy quantum broadcast channel. The most general quantum codes for this setup take the form of tripartite quantum channels. Since the receivers are constrained by causal ordering, additional temporal relationships naturally emerge between the parties. These causal constraints enforce intrinsic no-signalling conditions on any tripartite operation, ensuring that it constitutes a physically realizable quantum code for a quantum broadcast channel. We analyze these constraints and, more broadly, characterize the most general quantum codes for communication over such channels. We examine the capabilities of codes that are fully no-signalling among the three parties, positive partial transpose (PPT)-preserving, or both, and derive simple semidefinite programs to compute the achievable entanglement fidelity. We then establish a hierarchy of semidefinite programming converse bounds -- both weak and strong -- for the capacity of quantum broadcast channels for EPR pair generation, in both one-shot and asymptotic regimes. Notably, in the special case of a point-to-point channel, our strong converse bound recovers and strengthens existing results. Finally, we demonstrate how the PPT-preserving codes we develop can be leveraged to construct PPT-preserving entanglement combing schemes, and vice versa.

FOS: Physical sciences

SPARTAN (Scalable Probabilistic Application Reconfigurable Tensor Autonomous Network)

The technical founder of Ludwig Computing Inc has been competitively selected for support by Cyclotron Road, a U.S. Department of Energy (DOE) Advanced Manufacturing Office (AMO) Lab-Embedded Entrepreneurship Program (LEEP) through an approved merit review process. Ludwig Computing Inc, supported by the U.S. Department of Energy's Advanced Manufacturing Office through the Cyclotron Road program, has investigated the advantages of probabilistic computing for real-world compute-intensive applications. This research adds to the understanding of alternative computing paradigms by exploring a unique hardware-software co-design that integrates quantum computing methods with nature-inspired problem-solving techniques. The project's focus on areas such as combinatorial optimization, graph analytics, and machine learning demonstrates the potential for significant advancements in computational efficiency and performance. By harnessing natural randomness to streamline large circuits into fewer devices, Ludwig's approach enables massive parallelism, potentially offering higher throughput, speed, and energy efficiency compared to conventional hardware solutions. This work benefits the public by paving the way for more efficient computing solutions that could address complex real-world problems while potentially reducing energy consumption in data-intensive industries.

97 MATHEMATICS AND COMPUTING

Energy Efficiency Scaling for 2 Decades (EES2) Roadmap for Computing

In response to the looming crisis in global energy consumption required for advanced computing applications, the United States Department of Energy (DOE) Advanced Materials and Manufacturing Technology Office (AMMTO) is leading a multi-organizational effort to define a roadmap for energy efficiency scaling for two decades (EES2) with the aim to reduce energy use in all aspects of computation by more than a factor of 1000 in two decades. By July of 2024, over 60 organizations representing industry, academia, and the national laboratories have pledged to work in various aspects of research and development to enable energy efficiency in computing including in the development of the EES2 roadmap, with an initial public release in 2024 as the first phase of an ongoing commitment to energy-efficient and sustainable computation.

Kaarsberg, Tina [U.S. Department of Energy (DOE)]

Leveraging dendritic complexity for neuromorphic computing

Abstract Beyond-von Neumann computing approaches are necessary to sustain the growth of microelectronics and the increasing appetite for artificial intelligence/machine learning algorithms. Neuromorphic computing is an emerging paradigm that takes inspiration from the brain to provide a path forward to improve the computational efficiency and computational density of next-generation computing architectures. In nature, we observe brains performing complex computations with a much smaller energy footprint than conventional computing approaches. Current neuromorphic systems are focused primarily on scalability, namely, increasing the number of computational units (neurons) and connections between units (synapses). However, for brain-like cognition and efficiency in next-generation computing hardware, we need increased complexity in function, as well as improved connection density for scalability. Here, we present our work that aims to incorporate dendrites for ‘compute-on-wire’ in neuromorphic architectures to increase the computational complexity (e.g. number of programmable parameters, nonlinear dynamics) as well as computational efficiency (energy/compute) of artificial neural networks (ANNs). We do this by showcasing neuromorphic dendrite elements that can be leveraged for various applications. We will present examples of neuroscience-inspired direction-selective circuits and an ANN with active dendrites leveraging shunting inhibition. We also demonstrate the benefits of using dendrites in deep neural networks. To conclude, we discuss how we can utilize emerging hardware devices in these systems and design next-generation neuromorphic architectures with dendrites.

Cardwell, Suma G. (ORCID:0000000226575545)

Efficient derivative computation for unsteady fatigue-constrained nonlinear aero-structural wind turbine blade optimization

Gradient-based optimization offers significant efficiency advantages for wind turbine blade design, but its application has often been limited by the cost and accuracy of finite-difference derivative calculations, especially when fatigue constraints are considered. In this work, we systematically compare and evaluate four differentiation techniques, namely algorithmic differentiation, implicit differentiation, sparsity exploitation, and parallelization, to determine their effectiveness in computing accurate gradients through time-domain aero-structural simulations. By integrating these techniques with unsteady nonlinear aerodynamic and structural models, we develop software designed for accurate gradient computation. We show that combining these techniques addresses memory and runtime challenges associated with long simulations required by design load cases. Specifically, the most effective combination reduces derivative computation wall time by over an order of magnitude compared to finite differencing while maintaining superior accuracy. We demonstrate this approach in a proof-of-concept aero-structural optimization of a wind turbine blade that improves the cost of energy by 12.78 %. This comparative study establishes a viable approach for fatigue-aware blade design that balances computational efficiency with modeling accuracy.

17 WIND ENERGY

Evaluation of fluxon synapse device based on superconducting loops for energy efficient neuromorphic computing

With Moore’s law nearing its end due to the physical scaling limitations of CMOS technology, alternative computing approaches have gained considerable attention as ways to improve computing performance. Here, we evaluate performance prospects of a new approach based on disordered superconducting loops with Josephson-junctions for energy efficient neuromorphic computing. Synaptic weights can be stored as internal trapped fluxon states of three superconducting loops connected with multiple Josephson-junctions (JJ) and modulated by input signals applied in the form of discrete fluxons (quantized flux) in a controlled manner. The stable trapped fluxon state directs the incoming flux through different pathways with the flow statistics representing different synaptic weights. We explore implementation of matrix–vector-multiplication (MVM) operations using arrays of these fluxon synapse devices. We investigate the energy efficiency of online-learning of MNIST dataset. Our results suggest that the fluxon synapse array can provide ~100× reduction in energy consumption compared to other state-of-the-art synaptic devices. This work presents a proof-of-concept that will pave the way for development of high-speed and highly energy efficient neuromorphic computing systems based on superconducting materials.

42 ENGINEERING

The Extended Embedded Self-Shielding Method in SCALE 6.3/Polaris

The SCALE transport lattice code, Polaris, has been previously developed to generate few-group homogenized cross sections for whole-core nodal diffusion simulators in which the embedded self-shielding method (ESSM) is used for resonance self-shielding calculations to process cross sections. Although the ESSM capability has been very successful in light-water reactor analysis, it may require enhancements in computational efficiency; treatment of spatially dependent resonance self-shielding effects; and handling of interrelated resonance effects among fuel, cladding, and control rod materials. Therefore, this study focuses on improving computational efficiency by using a Dancoff-based Wigner–Seitz approximation combined with a material-based resonance categorization, through which a spatially dependent ESSM capability is developed to accurately estimate self-shielded cross sections inside the fuel. Benchmark results show that the new capability significantly enhances computational efficiency and accuracy for spatially dependent local zones within the fuel and through depletion.

ESSM

Energy-efficient scientific computing using chemical reservoirs

The rapid growth of computing demands driven by scientific computing, data analytics, and artificial intelligence (AI) advancements has exposed the limitations of traditional digital processing systems. These systems are nearing physical energy barriers, making significant gains in energy efficiency increasingly unattainable. As we advance toward post-exascale computing, disruptive approaches are critical to overcoming these limitations. Among emerging analog solutions, biochemical computing offers a transformative path for achieving orders-of-magnitude improvements in energy efficiency. By leveraging the natural optimization capabilities of chemical reaction networks (CRNs), biochemical systems have the potential to meet high-performance computing needs through natural scalability. However, numerous challenges remain, including theoretical limitations in mapping computational problems to CRNs and practical barriers in implementing biochemical computing devices. In this paper, we present a framework for chemical computation using biochemical systems and introduce key components of our approach for energy-efficient scientific computing. We showcase the feasibility of this framework by solving a system of ordinary differential equations by emulating a chemical reservoir device, demonstrating its potential for addressing modern computing challenges. This work lays a foundational step toward harnessing the computational power of chemistry to design energy-efficient, scalable, high-performance next-generation computing systems.

Johnson, Connah G. M. [Pacific Northwest National

Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications

Abstract Robust quantification of predictive uncertainty is a critical addition needed for machine learning applied to weather and climate problems to improve the understanding of what is driving prediction sensitivity. Ensembles of machine learning models provide predictive uncertainty estimates in a conceptually simple way but require multiple models for training and prediction, increasing computational cost and latency. Parametric deep learning can estimate uncertainty with one model by predicting the parameters of a probability distribution but does not account for epistemic uncertainty. Evidential deep learning, a technique that extends parametric deep learning to higher-order distributions, can account for both aleatoric and epistemic uncertainties with one model. This study compares the uncertainty derived from evidential neural networks to that obtained from ensembles. Through applications of the classification of winter precipitation type and regression of surface-layer fluxes, we show evidential deep learning models attaining predictive accuracy rivaling standard methods while robustly quantifying both sources of uncertainty. We evaluate the uncertainty in terms of how well the predictions are calibrated and how well the uncertainty correlates with prediction error. Analyses of uncertainty in the context of the inputs reveal sensitivities to underlying meteorological processes, facilitating interpretation of the models. The conceptual simplicity, interpretability, and computational efficiency of evidential neural networks make them highly extensible, offering a promising approach for reliable and practical uncertainty quantification in Earth system science modeling. To encourage broader adoption of evidential deep learning, we have developed a new Python package, Machine Integration and Learning for Earth Systems (MILES) group Generalized Uncertainty for Earth System Science (GUESS) (MILES-GUESS) ( https://github.com/ai2es/miles-guess ), that enables users to train and evaluate both evidential and ensemble deep learning. Significance Statement This study demonstrates a new technique, evidential deep learning, for robust and computationally efficient uncertainty quantification in modeling the Earth system. The method integrates probabilistic principles into deep neural networks, enabling the estimation of both aleatoric uncertainty from noisy data and epistemic uncertainty from model limitations using a single model. Our analyses reveal how decomposing these uncertainties provides valuable insights into reliability, accuracy, and model shortcomings. We show that the approach can rival standard methods in classification and regression tasks within atmospheric science while offering practical advantages such as computational efficiency. With further advances, evidential networks have the potential to enhance risk assessment and decision-making across meteorology by improving uncertainty quantification, a longstanding challenge. This work establishes a strong foundation and motivation for the broader adoption of evidential learning, where properly quantifying uncertainties is critical yet lacking.

Schreck, John S.

Latent heat thermal energy storage performance maps enabling fast & accurate building energy simulations

Thermal energy storage (TES) using phase change materials (PCMs) has gained attention as an effective approach to manage energy demand fluctuations and shift peak building loads. PCM embedded heat exchangers (PCM-HXs) offer high energy storage density and low temperature variation during phase change, being suitable for load-shifting applications. However, this component is typically evaluated using computationally expensive methods, which present significant challenges when the ultimate goal is to assess the performance of PCM-HX integrated thermal energy storage systems in the full building context. In this paper, we present a methodology to generate highly accurate and computationally efficient PCM-HX performance maps which can be easily integrated into building energy simulation tools to analyze the feasibility of space conditioning systems with latent heat PCM-based TES. The performance maps are generated using a computationally efficient PCM-HX simulation tool based on a Generalized Resistance-Capacitance Model (GRCM) which can simulate arbitrary PCM-HXs with high accuracy and significantly less computational effort compared to full CFD simulations. The methodology was verified for a case study considering a 5-ton (~17.5 kW) air-to-water heat pump-thermal energy storage system (HP-TES), which was co-simulated in Modelica for a DOE prototype small-office building in Vienna, Austria, using Spawn of EnergyPlus™. The TES performance maps provided accurate predictions of PCM-HX behavior when used as Modelica component, with deviations within 2-4% while also achieving at least 103 computational time reduction. Leveraging this faster prediction capability, four PCMs with different melting temperatures for cooling (12°C, 16°C) and heating (31°C, 36°C) were assessed to investigate their impact on system performance. This work highlights the importance of robust PCM-HX models for efficient and high-fidelity building-level simulations, presenting new opportunities for advanced control strategy development and parametric analysis of TES configurations in a computationally efficient manner

Modelica Building Simulations

A Scale‐Adaptive Urban Hydrologic Framework: Incorporating Network‐Level Storm Drainage Pipes Representation

Abstract Below‐ground urban stormwater networks (BUSNs) significantly influence urban flood dynamics, yet their representation at the watershed or larger scales remains challenging. We introduce a scalable urban hydrologic framework that centers on a novel network‐level BUSN representation, balancing the needs for physical basis, parameter parsimony, and computational efficiency. Our framework conceptualizes an urban watershed into four interacting zones: hillslopes (natural), storm‐sewersheds (urban), a sub‐network channel (tributaries), and a main channel. We develop an innovative Graph Theory‐based algorithm to derive network‐level BUSN parameters from publicly available datasets, enabling efficient, scalable parameterization. We demonstrate this framework's applicability at nine representative watersheds in the Houston metropolitan region, USA, with urban imperviousness ranging from 0% to 64% and drainage areas ranging from 24 to 302 . Our model achieves satisfying computational efficiency, completing hourly time step simulations for 18 years in less than 5 sec per watershed on a standard PC. Validation against observed daily streamflow confirms that the model can capture small‐to‐large flood peaks and seasonal and annual water balance over these watersheds. Comparisons with the National Water Model show better performance in predicting flood peaks and overall water balance, underscoring the promises of our new framework for urban hydrologic modeling at large scales. Furthermore, analysis reveals nonlinear relationships between BUSNs' designed capacities and flood reduction effects. Our approach bridges the gap between detailed hydraulic and large‐scale hydrologic models, providing a valuable tool for urban flood prediction and management across broader spatial and temporal scales.

54 ENVIRONMENTAL SCIENCES

Non -degenerate marginal-likelihood calibration with application to quantum characterization

Here, we propose a marginal likelihood strategy within the Kennedy-O’Hagan (KOH) Bayesian framework, where a Gaussian process (GP) models the discrepancy between a physical system and its simulator. Our approach introduces a novel marginalized likelihood by integrating out the degenerate eigenspace of the covariance matrix, rather than approximating the original likelihood. Unlike approximation methods that compromise accuracy for computational efficiency, our method defines an exact likelihood—distinct from the original but preserving all relevant information. This formulation achieves computational efficiency and stability, even for large datasets where the covariance matrix nears degeneracy. Applied to the characterization of a superconducting quantum device at Lawrence Livermore National Laboratory, the approach enhances the predictive accuracy of the Lindblad master equations for modeling Ramsey measurement data by effectively quantifying uncertainties consistent with the quantum data.

general physics

Coupled Induction Machine and HVAC Models for Simulating HVAC Performance Considering Grid Dynamics in Buildings

This paper presents the development of novel models that integrate induction machines with HVAC equipment, such as pumps, heat pumps, and chillers, to analyze the impact of electrical parameters on the operational performance of thermo-fluid systems. The proposed model employs a coupling technique that captures the dynamic interactions between induction machines and HVAC systems. By integrating electrical, thermal, and mechanical dynamics, the models provide a comprehensive framework for simulating real-world scenarios, including interactions with the electrical grid. This achievement was made possible through the development of a Computationally Efficient and Accurate Induction Machine (CEAIM) model. Implemented using the equation-based Modelica language, the CEAIM model has been validated against experimental results, manufacturer data sheets, and various operating conditions. Its performance has been compared with existing induction machine models in the Modelica Standard Library (MSL), demonstrating superior accuracy and computational efficiency. The CEAIM model predicts torque, speed, and power consumption with a coefficient of determination (R 2 ) ranging from 0.98 to 1 and a coefficient of variation of root mean square error (CVRMSE) between 0.27% and 6.67%. Additionally, CEAIM scales more efficiently than conventional MSL models, with a slower computational growth rate in large-scale simulations. After thorough validation of the CEAIM model, it was coupled with HVAC equipment as this approach provides a detailed multi-dimensional view of capturing electrical transients and mechanical performance. To support this, a case study was conducted to showcase its capabilities.

24 POWER TRANSMISSION AND DISTRIBUTION