Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Incremental Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Then and Now: Improving Software Portability, Productivity, and 100× Performance

The US Exascale Computing Project (ECP) has succeeded in preparing applications to run efficiently on the first reported Exascale supercomputers in the world. To achieve this, it modernized the whole leadership software stack, from libraries to simulation codes. In this article, we contrast selected leadership software before and after ECP. We discuss how sustainable research software development for leadership computing can embrace the conversation with the hardware vendors, the leadership computing facilities, the software community, and the domain scientists who are the application developers and integrators of software products. We elaborate on how software needs to take portability as a central design principle and to benefit from interdependent teams; we also demonstrate how moving to programming languages with high momentum, like modern C++, can help improve the sustainability, interoperability, and performance of research software. Finally, we showcase how cross-institutional efforts can enable algorithm advances that are beyond incremental performance optimization.

97 MATHEMATICS AND COMPUTING↗

Experimental and theoretical analysis of carbon driven detonation waves in a heterogeneously premixed Rotating Detonation Engine

Coal dust explosions can be hazardous; however, they can also generate a significant rise in stagnation pressure if adequately harnessed. Rotating detonation combustors seek to take advantage of the stagnation pressure rise phenomenon in a more sustained and controlled manner via confinement to a physical annulus, leading to increased overall thermodynamic efficiency. Here this investigation presents an analysis of detonations fueled by Carbon Black, a solid particulate consisting of virtually pure carbon molecules and lean Hydrogen-Air mixtures. It is realized that with the addition of Carbon Black, an increase of lean mixture detonability and detonation velocities extending the operating limit over that of a pure hydrogen-air mixture is experienced. For all testing conditions, the total equivalence ratio is held at φ = 1, while the fuel mixture's carbon mass fraction is increased from 0 to 0.7 while the hydrogen is decreased. Detonation wave velocities are extracted from high-speed imaging through applying a Discrete Fourier Transform algorithm to determine changes to the wave speed as Carbon Black particles are introduced. As a result, due to the addition of Carbon Black as an auxiliary fuel source, detonations were formed instead of deflagrations in operating conditions where one would expect deflagrations at the same hydrogen-air equivalence ratios without Carbon Black addition. The detonation formation provides evidence that the coal particles are reacting within the detonation wave in a large enough capacity to support a detonation wave within the annulus. Furthermore, the wave speed is shown to increase with the additional of carbon particles. At a constant global equivalence ratio, the detonation wave velocities were found to decrease with hydrogen's incremental replacement with coal particles. Whereby, through a theoretical comparison of the heat of combustion as computed from the experimentally derived detonation wave velocities, a linear relationship of the two was shown to exist. Therefore, the heat of combustion has the potential to describe an operational limit to sustaining a detonation wave.

42 ENGINEERING↗

Disentangling the physics of the attractive Hubbard model as a fully interacting model of fermions via the accessible and symmetry-resolved entanglement entropies

The complicated ways in which electrons interact in many-body systems such as molecules and materials have long been viewed through the lens of local electron correlation and associated correlation functions. However, quantum information science has demonstrated that more global diagnostics of quantum states like the entanglement entropy can provide a complementary and clarifying lens on electronic behavior. One particularly useful measure that can be used to distinguish between quantum and classical sources of entanglement is the accessible entanglement, the entanglement available as a quantum resource for systems subject to conservation laws, such as fixed particle number, due to superselection rules. In this work, we introduce an algorithm and demonstrate how to compute accessible and symmetry-resolved entanglements for interacting fermion systems. This is accomplished by combining an incremental version of the swap algorithm with a recursive auxiliary field quantum Monte Carlo algorithm recently developed by the authors. We apply these tools to study the pairing and charge density waves exhibited in the paradigmatic attractive Hubbard model via entanglement. We find that the particle and spin symmetry-resolved entanglements and their related full probability distribution functions show very clear—and unique—signatures of the underlying electronic behavior even when those features are less pronounced in conventional correlation functions. Altogether, this work provides a systematic means of characterizing the entanglement within quantum systems that can grant a deeper understanding of the complicated electronic behavior that underlies quantum phase transitions and crossovers in many-body systems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Iteration-based Linearized Distribution-level Locational Marginal Price for Three-phase Unbalanced Distribution Systems

Distributed energy resources (DERs) are rocking the utilities’ business landscape. It calls for competitive market environments that incentivize DERs to form maximum operating efficiency. Among proposed pricing schemes, distribution-level locational marginal price (DLMP) is effective in signaling the marginal generation cost differences driven by energy losses and network constraints. It can be derived from a distribution-level optimal power flow (OPF) framework, as it essentially presents the sensitivity of optimized generation cost towards incremental loads. However, due to the high resistance-to-inductance ratio and unbalanced characteristics of distribution networks, computational affordable DLMPs are highly challenged. This article provides a linear-approximated DLMP that can be solved efficiently and generalized to account for reactive power flow, three-phase unbalanced loads and meshed network structure. The successive linear programming technique is introduced to enhance the model accuracy. Case studies on an IEEE 123-Bus system validate its accuracy against a nonlinear benchmark and capability in offering proper incentives.

24 POWER TRANSMISSION AND DISTRIBUTION↗

NREL Infrastructure Perception and Control Workshop

A lack of highly reliable, full state-space awareness of roadway situations is the current bottleneck for the incremental introduction of smart infrastructure control. NREL's Infrastructure Perception and Control (IPC) lab applies advanced sensing and computation controls to the coordinated movement of vehicles on the road as well as people in large facilities and has produced field test results from a Colorado Springs intersection. In this presentation, NREL discusses the state of smart infrastructure control and opportunities for partnership.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Fragmentation analysis of a bar with the Lip-field approach

The Lip-field approach was introduced in Moës and Chevaugeon (2021) as a new way to regularize softening material models. It was tested in 1D quasistatic in Moës and Chevaugeon (2021) and 2D quasistatic in Chevaugeon and Moës (2021): this paper extends it to 1D dynamics, on the challenging problem of dynamic fragmentation. The Lip-field approach formulates the mechanical problem to be solved as an optimization problem, where the incremental potential to be minimized is the non-regularized one. Spurious localization is prevented by imposing a Lipschitz constraint on the damage field. Here, the displacement and damage field at each time step are obtained by a staggered algorithm, that is the displacement field is computed for a fixed damage field, then the damage field is computed for a fixed displacement field. Indeed, these two problems are convex, which is not the case of the global problem where the displacement and damage fields are sought at the same time. The incremental potential is obtained by equivalence with a cohesive zone model, which makes material parameters calibration simple. A non-regularized local damage equivalent to a cohesive zone model is also proposed. It is used as a reference for the Lip-field approach, without the need to implement displacement jumps. These approaches are applied to the brittle fragmentation of a 1D bar with randomly perturbed material properties to accelerate spatial convergence. Both explicit and implicit dynamic implementations are compared. Favorable comparison to several analytical, numerical and experimental references serves to validate the modeling approach.

36 MATERIALS SCIENCE↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

Addressing Load Imbalance in Bioinformatics and Biomedical Applications: Efficient Scheduling across Multiple GPUs

Computational bioinformatics and biomedical applications frequently contain heterogeneously sized units of work or tasks, for instance due to variability in the sizes of biological sequences and molecules. Variable-sized workloads lead to load imbalances in parallel implementations which detract from efficiency and performance. Many modern computing resources now have multiple graphics processing units(GPUs) per computer for acceleration. These multiple GPU resources need to be used efficiently through balancing of workloads across the GPUs. OpenMP is a portable directive-based parallel programming API used ubiquitously in bioscience applications to program CPUs; recently, the use of OpenMP directives for GPU acceleration has become possible. Here, motivated by experiences with imbalanced loads in GPU-accelerated bioinformatics applications, we address the load balancing problem using OpenMP task-to-GPU scheduling combined with OpenMP GPU offloading for multiply heterogeneous workloads – loads with both variable input sizes, and simultaneously, variable convergence rates for algorithms with a stochastic component – scheduled across multiple GPUs. We aim to develop strategies which are both easy to use and have lower overheads, and may be incorporated incrementally in existing programs which already make use of OpenMP for CPU-based threading in order to make use of multi-GPU computers. We test different combinations of input size variability and convergence rate variability, and characterize the effects of these different scenarios on the performance of scheduling strategies across multiple GPUs with OpenMP. We present several dynamic scheduling solutions for different parallel patterns, explore optimizations, and provide publicly available example computational kernels to make these strategies easy to use in programs. This work will enable application developers to efficiently and easily use multiple GPUs for imbalanced workloads found in bioinformatics and biomedical applications.

Thavappiragasam, Mathialakan↗

Watermarks in stream processing systems: semantics and comparative analysis of Apache Flink and Google cloud dataflow

Streaming data processing is an exercise in taming disorder: from oftentimes huge torrents of information, we hope to extract powerful and timely analyses. But when dealing with streaming data, the unbounded and temporally disordered nature of real-world streams introduces a critical challenge: how does one reason about the completeness of a stream that never ends? In this paper, we present a comprehensive definition and analysis of watermarks, a key tool for reasoning about temporal completeness in infinite streams.First, we describe what watermarks are and why they are important, highlighting how they address a suite of stream processing needs that are poorly served by eventually-consistent approaches:• Computing a single correct answer, as in notifications.• Reasoning about a lack of data, as in dip detection.• Performing non-incremental processing over temporal subsets of an infinite stream, as in statistical anomaly detection with cubic spline models.• Safely and punctually garbage collecting obsolete inputs and intermediate state.• Surfacing a reliable signal of overall pipeline health.Second, we describe, evaluate, and compare the semantically equivalent, but starkly different, watermark implementations in two modern stream processing engines: Apache Flink and Google Cloud Dataflow.

Akidau, Tyler↗

Solving the $k$-Sparse Eigenvalue Problem with Reinforcement Learning

We examine the possibility of using a reinforcement learning (RL) algorithm to solve large-scale eigenvalue problems in which the desired the eigenvector can be approximated by a sparse vector with at most k nonzero elements, where k is relatively small compare to the dimension of the matrix to be partially diagonalized. Here, this type of problem arises in applications in which the desired eigenvector exhibits localization properties and in large-scale eigenvalue computations in which the amount of computational resource is limited. When the positions of these nonzero elements can be determined, we can obtain the k-sparse approximation to the original problem by computing eigenvalues of a k × k submatrix extracted from k rows and columns of the original matrix. We review a previously developed greedy algorithm for incrementally probing the positions of the nonzero elements in a k-sparse approximate eigenvector and show that the greedy algorithm can be improved by using an RL method to refine the selection of k rows and columns of the original matrix. We describe how to represent states, actions, rewards and policies in an RL algorithm designed to solve the k-sparse eigenvalue problem and demonstrate the effectiveness of the RL algorithm on two examples originating from quantum many-body physics.

97 MATHEMATICS AND COMPUTING↗

Ab-initio Cu alloy design for high-gradient accelerating structures

Operation of normal conducting accelerator structures at high accelerating gradients is beneficial for many accelerator applications in basic science, industry, medicine, and National Security. RF breakdown is the major factor that limits the achievable accelerating gradients. Previous experiments on copper (Cu) have demonstrated that RF breakdown probability can be significantly decreased by hardening the material and alloying Cu with solutes such as silver (Ag). In this paper, we propose a figure-of-merit (FOM) that characterizes the ability of Cu alloys to withstand high-gradients. The FOM represents a trade-off between hardening through solid solution strengthening and the additional thermal stress induced by incremental RF pulse heating resulting from changes in electronic properties induced by alloying. We performed high-throughput ab initio calculations and computed the FOM for a large number of binary Cu alloys. Several promising candidate alloys for high-gradient accelerating structures were identified, such as CuAg, CuCd, CuHg, CuAu, CuIn, and CuMg. CuAg alloys have previously exhibited low RF breakdown rates in experiments. The results provide guidance for selecting alloys for the future high-gradient normal conducting accelerating structures operating at very high gradients.

36 MATERIALS SCIENCE↗

CICE on a C-grid: new momentum, stress, and transport schemes for CICEv6.5

Abstract. This article presents the C-grid implementation of the CICE sea ice model, including the C-grid discretization of the momentum equation, the boundary conditions (BCs), and the modifications to the code required to use the incremental remapping transport scheme. To validate the new C-grid implementation, many numerical experiments were conducted and compared to the B-grid solutions. In idealized experiments, the standard advection method (incremental remapping with C-grid velocities interpolated to the cell corners) leads to a checkerboard pattern. A modal analysis demonstrates that this computational noise originates from the spatial averaging of C-grid velocities at corners. The checkerboard pattern can be eliminated by adjusting the departure regions to match the divergence obtained from the solution of the momentum equation. We refer to this novel approach as the edge flux adjustment (EFA) method. The C-grid discretization with edge flux adjustment allows for transport in channels that are one grid cell wide – a capability that is not possible with the B-grid discretization nor with the C-grid and standard remapping advection. Simulation results match the predicted values of a novel analytical solution for one-grid-cell-wide channels.

Lemieux, Jean-François (ORCID:0000000320845759)↗

Computational Study of Additively Manufactured Internally Cooled Airfoils for Industrial Gas Turbine Applications

Internal cooling features such as pin-fins, impingement jets, and rib-turbulators are necessary to keep turbine components cool, but if sufficiently advanced can potentially also eliminate the need for film cooling on turbine blades particularly in industrial gas turbines where temperatures are not extreme. Furthermore, by leveraging additive manufacturing, other advanced designs such as lattice and incremental impingement configurations are possible and have recently been experimentally tested. While the performance of such configurations has been quantified through means of overall cooling effectiveness, it is not as clear why certain designs were better than others, or what the mechanisms were behind the observed external cooling patterns. The purpose of this study was to computationally analyze different advanced turbine blade internal cooling designs previously tested by the National Energy Technology Laboratory.

CFD↗

Toward Full Configuration Interaction for Transition-Metal Complexes

In this work, an efficacious approximation to full configuration interaction (FCI) is adapted to calculate singlet-triplet gaps for transition-metal complexes. This strategy, incremental FCI (iFCI), uses a many-body expansion to systematically add correlation to a simple reference wave function and therefore achieves greatly reduced computational costs compared to FCI. iFCI through the 3-body expansion is demonstrated on four model transition-metal complexes involving the metals Zn, V, and Cu. Screening techniques to increase the computational efficiency of iFCI are proposed and tested, showing reduction in the number of 3-body terms by more than 90% with controlled errors. The largest complex treated by iFCI has 142 valence electrons, all of which are correlated among the full set of 444 active orbitals. Computed spin gaps approach experimental results for the four complexes, though room for improvement remains.

74 ATOMIC AND MOLECULAR PHYSICS↗

Decomposing causality into its synergistic, unique, and redundant components

Causality lies at the heart of scientific inquiry, serving as the fundamental basis for understanding interactions among variables in physical systems. Despite its central role, current methods for causal inference face significant challenges due to nonlinear dependencies, stochastic interactions, self-causation, collider effects, and influences from exogenous factors, among others. While existing methods can effectively address some of these challenges, no single approach has successfully integrated all these aspects. Here, we address these challenges with SURD: Synergistic-Unique-Redundant Decomposition of causality. SURD quantifies causality as the increments of redundant, unique, and synergistic information gained about future events from past observations. The formulation is non-intrusive and applicable to both computational and experimental investigations, even when samples are scarce. We benchmark SURD in scenarios that pose significant challenges for causal inference and demonstrate that it offers a more reliable quantification of causality compared to previous methods.

applied mathematics↗

Packet router with virtual channel hop buffer control

An integrated circuit includes a network on chip (NOC) that includes a plurality of processing elements and a plurality of NOC nodes, interconnected to the plurality of processing elements. The integrated circuit includes logic that is configured to: increment by one, a virtual channel identifier to produce an incremented destination VC identifier, the virtual channel (VC) identifier associated with at least portion of a packet stored in at least one virtual channel buffer; determine that a destination virtual channel buffer corresponding to the incremented destination VC identifier in a destination NOC node in the NOC is available to store the portion of the packet; and in response to the determination, send the portion of the packet and the incremented destination VC identifier to the destination NOC node.

97 MATHEMATICS AND COMPUTING↗

iBLAST: Incremental BLAST of new sequences via automated e-value correction

Search results from local alignment search tools use statistical scores that are sensitive to the size of the database to report the quality of the result. For example, NCBI BLAST reports the best matches using similarity scores and expect values (i.e., e-values) calculated against the database size. Given the astronomical growth in genomics data throughout a genomic research investigation, sequence databases grow as new sequences are continuously being added to these databases. As a consequence, the results (e.g., best hits) and associated statistics (e.g., e-values) for a specific set of queries may change over the course of a genomic investigation. Thus, to update the results of a previously conducted BLAST search to find the best matches on an updated database, scientists must currently rerun the BLAST search against the entire updated database, which translates into irrecoverable and, in turn, wasted execution time, money, and computational resources. To address this issue, we devise a novel and efficient method to redeem past BLAST searches by introducing iBLAST. iBLAST leverages previous BLAST search results to conduct the same query search but only on the incremental (i.e., newly added) part of the database, recomputes the associated critical statistics such as e-values, and combines these results to produce updated search results. Our experimental results and fidelity analyses show that iBLAST delivers search results that are identical to NCBI BLAST at a substantially reduced computational cost, i.e., iBLAST performs (1 + δ )/ δ times faster than NCBI BLAST, where δ represents the fraction of database growth. We then present three different use cases to demonstrate that iBLAST can enable efficient biological discovery at a much faster speed with a substantially reduced computational cost.

59 BASIC BIOLOGICAL SCIENCES↗

Entropy–Preserving and Entropy–Stable Relaxation IMEX and Multirate Time–Stepping Methods

In this work, we propose entropy-preserving and entropy-stable partitioned Runge–Kutta (RK) methods. In particular, we extend the explicit relaxation Runge–Kutta methods to IMEX–RK methods and a class of explicit second-order multirate methods for stiff problems arising from scale-separable or grid-induced stiffness in a system. The proposed approaches not only mitigate system stiffness but also fully support entropy-preserving and entropy-stability properties at a discrete level. The key idea of the relaxation approach is to adjust the step completion with a relaxation parameter so that the time-adjusted solution satisfies the entropy condition at a discrete level. The relaxation parameter is computed by solving a scalar nonlinear equation at each timestep in general; however, as for a quadratic entropy function, we theoretically derive the explicit form of the relaxation parameter and numerically confirm that the relaxation parameter works the Burgers equation. Several numerical results for ordinary differential equations and the Burgers equation are presented to demonstrate the entropy-conserving/stable behavior of these methods. We also compare the relaxation approach and the incremental direction technique for the Burgers equation with and without a limiter in the presence of shocks.

97 MATHEMATICS AND COMPUTING↗