Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computers in other domains”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Synthetic-domain computing and neural networks using lithium niobate integrated nonlinear phononics

Analogue computing uses the physical behaviours of devices to provide energy-efficient arithmetic operations. However, scaling up analogue computing platforms by simply increasing the number of devices leads to challenges such as device-to-device variation. Here, in this study, we report scalable analogue computing and neural networks in the synthetic frequency domain using an integrated nonlinear phononic platform on lithium niobate. This synthetic-domain computing is robust to device variations, as vectors and matrices are concurrently encoded at different frequencies within a single device, achieving a high throughput per area. Leveraging inherent nonlinearities, our device-aware neural network can perform a four-class classification task with an accuracy of 98.2%. The nonlinear phononic computing hardware also maintains consistent performance over a wide operational temperature range (characterized up to 192 °C). Our synthetic-domain computing combines single-device parallelism, inherent nonlinearity and environmental stability, and could be of use in edge computing applications in which power efficiency and environmental resilience are crucial.

Ji, Jun [Virginia Polytechnic Inst. and State Univ↗

Frequency-domain computing using nonlinear acoustic-wave device on lithium niobate

Abstract Multiply-accumulation are crucial computing operations in signal processing, numerical simulations, and machine learning. In recent years, optical analog approaches have demonstrated higher computing performance and better power efficiency than their digital counterparts. However, analog computing chips usually need large areas and complex structures for parallel computing, as a single device element only executes one computing operation at a single time. Here, we demonstrate frequency-domain computing using the nonlinear acoustic-wave devices on lithium niobate, featuring a normalized external second-harmonic generation conversion efficiency of ~ 5.7 × 10-4 W-1. The second-order sum-frequency nonlinear process of lithium niobate enables multiplication of inputs encoded in the frequency domain. Compared to the analog schemes, our device features a notably simpler design, and nanofabrication requires only one lift-off. Using a single acoustic-wave device within an area of 0.03 mm2, we can simultaneously conduct over 130,000 multiply-accumulation operations. Our acoustic-wave device shows applications in real and complex vector convolutions and image processing. This demonstration sets the stage for experimental realizations into frequency-domain integrated nonlinear acoustic computing systems, potentially shaping future developments in acoustic neural networks and quantum computing.

chai, mingzhao (ORCID:0009000466226341)↗

Accelerating Multivariate Functional Approximation Computation with Domain Decomposition Techniques⋆

Modeling large datasets through Multivariate Functional Approximations (MFA) provide an elegant way to handle many visualization and scientific analysis workflows. The process necessitates scalable data partitioning methods to compute MFA representations efficiently without compromising the accuracy or continuity of the reconstructed solution. We propose a domain -decomposed method for computing the MFA with B -spline bases, which reduces the total work per task and uses a restricted Additive Schwarz (RAS) method to converge the control point data degrees -of -freedom along subdomain boundaries. We provide an in-depth analysis of the parallel approach with domain decomposition solvers, aiming to minimize local subdomain error residuals and recover high -order continuity at subdomain interfaces with appropriate choices of knot overlaps. The communication cost, determined by the overlap regions in the RAS implementation, is optimized to recover the numerical error profile of the single subdomain case. Our proposed method stands in contrast to previous methods, which typically only recover either C 0 or at best C 1 continuity for arbitrary B -spline degree expansions, or those that require post -processing to blend discontinuities in the reconstructed data. We demonstrate the effectiveness of our approach using analytical and real -world datasets in 1D, 2D, and 3D through both strong and weak scaling studies. The performance results indicate that the overall cost of computing the approximation is directly proportional to the underlying nearest -neighbor communication implementation, and is only weakly dependent on the overlap region size that determines the size of the messages. This finding underscores the efficiency and scalability of our proposed method, making it a promising solution for handling large datasets in scientific workflows.

additive Schwarz solvers↗

Perfectly Matched Layers and Characteristic Boundaries in Lattice Boltzmann: Accuracy vs Cost

Artificial boundary conditions (BCs) play a ubiquitous role in numerical simulations of transport phenomena in several diverse fields, such as fluid dynamics, electromagnetism, acoustics, geophysics, and many more. They are essential for accurately capturing the behavior of physical systems whenever the simulation domain is truncated for computational efficiency purposes. Ideally, an artificial BC would allow relevant information to enter or leave the computational domain without introducing artifacts or unphysical effects. Boundary conditions designed to control spurious wave reflections are referred to as nonreflective boundary conditions (NRBCs). Another approach is given by the perfectly matched layers (PMLs), in which the computational domain is extended with multiple dampening layers, where outgoing waves are absorbed exponentially in time. Here, in this work, the definition of PML is revised in the context of the lattice Boltzmann method. The impact of adopting different types of BCs at the edge of the dampening zone is evaluated and compared, in terms of both accuracy and computational costs. It is shown that for sufficiently large buffer zones, PMLs allow stable and accurate simulations even when using a simple zeroth-order extrapolation BC. Moreover, employing PMLs in combination with NRBCs potentially offers significant gains in accuracy at a modest computational overhead, provided the parameters of the BC are properly tuned to match the properties of the underlying fluid flow.

97 MATHEMATICS AND COMPUTING↗

Improving the secretion of designed protein assemblies through negative design of cryptic transmembrane domains

Computationally designed protein nanoparticles have recently emerged as a promising platform for the development of new vaccines and biologics. For many applications, secretion of designed nanoparticles from eukaryotic cells would be advantageous, but in practice, they often secrete poorly. Here we show that designed hydrophobic interfaces that drive nanoparticle assembly are often predicted to form cryptic transmembrane domains, suggesting that interaction with the membrane insertion machinery could limit efficient secretion. We develop a general computational protocol, the Degreaser, to design away cryptic transmembrane domains without sacrificing protein stability. The retroactive application of the Degreaser to previously designed nanoparticle components and nanoparticles considerably improves secretion, and modular integration of the Degreaser into design pipelines results in new nanoparticles that secrete as robustly as naturally occurring protein assemblies. Both the Degreaser protocol and the nanoparticles we describe may be broadly useful in biotechnological applications.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

A Minkowski difference-based advancing front packing technique for generating convex noncircular particles in complex domains

In this work, a Minkowski difference-based advancing front approach is proposed to generate convex and non-circular particles in a predefined computational domain. Two specific algorithms are developed to handle the contact conformity of generated particles with the boundaries of the computational domain. The first, called the open form, is used to handle the smooth contact of generated particles with (external) boundaries, while the other, called the closed form, is proposed to handle the internal boundaries of a computational domain with a complex cavity. The Gilbert-Johnson-Keerthi (GJK) method is used to efficiently solve the contact detection between the newly generated particle at the front and existing particles. Furthermore, the problem of one-sided particle lifting, which can cause some defects in the packing structure in existing advancing front methods during packing generation, is highlighted and an effective solution is developed. Several examples of increasing complexity are used to demonstrate the efficiency and applicability of the proposed packing generation approach. The numerical results show that the generated packing is not only more uniform, but also achieves a higher packing density than existing advancing front methods.

42 ENGINEERING↗

Variation Tolerant and Energy-Efficient Charge Domain Compute-in-Memory Array with Binary and Multi-Level Cell Ferroelectric FET

Here, in this work, we present a variation-tolerant and energy-efficient charge-domain Ferroelectric FET (FeFET) based Compute-in-Memory (CiM) array design that is compatible with both binary and multi-level cell memory sensing. We demonstrate that: 1) by exploiting FeFET as a nonvolatile switch, its high ON/OFF ratio in the subthreshold region can suppress the error introduced by the inaccurate ON state conductance, thus realizing robust CiM operations, unlike the current-domain CiM design where the computation results is highly sensitive to the device conductance variation; 2) by leveraging a dense dynamic random access memory (DRAM)-like 1FeFET1C cell structure, the proposed design benefits from the existing high density DRAM establishment while also significantly relaxing the capacitor retention and transistor leakage requirement; 3) the charge-domain CiM supports both binary FeFET with minimum overhead and MLC FeFET with tolerable latency for MLC state sensing, whose efficacy is validated experimentally on both cell-level and array-level; 4) the proposed CiM shows much better device variation resilience than conventional current-domain CiM, and also improves inference accuracy. Macro-level evaluation results demonstrate significantly higher energy efficiency and area efficiency compared to prior CiM works.

Duan, Jiahui [University of Notre Dame, IN (United↗

Embedded symmetric positive semi-definite machine-learned elements for reduced-order modeling in finite-element simulations with application to threaded fasteners

Here, we present a machine-learning strategy for finite element analysis of solid mechanics wherein we replace complex portions of a computational domain with a data-driven surrogate. In the proposed strategy, we decompose a computational domain into an “outer” coarse-scale domain that we resolve using a finite element method (FEM) and an “inner” fine-scale domain. We then develop a machine-learned (ML) model for the impact of the inner domain on the outer domain. In essence, for solid mechanics, our machine-learned surrogate performs static condensation of the inner domain degrees of freedom. This is achieved by learning the map from displacements on the inner-outer domain interface boundary to forces contributed by the inner domain to the outer domain on the same interface boundary. We consider two such mappings, one that directly maps from displacements to forces without constraints, and one that maps from displacements to forces by virtue of learning a symmetric positive semi-definite (SPSD) stiffness matrix. We demonstrate, in a simplified setting, that learning an SPSD stiffness matrix results in a coarse-scale problem that is well-posed with a unique solution. We present numerical experiments on several exemplars, ranging from finite deformations of a cube to finite deformations with contact of a fastener-bushing geometry. We demonstrate that enforcing an SPSD stiffness matrix drastically improves the robustness and accuracy of FEM–ML coupled simulations, and that the resulting methods can accurately characterize out-of-sample loading configurations with significant speedups over the standard FEM simulations.

97 MATHEMATICS AND COMPUTING↗

Absorbing boundary conditions in material point method adopting perfectly matched layer theory

This study focuses on solving the numerical challenges of imposing absorbing boundary conditions for dynamic simulations in the material point method (MPM). To attenuate elastic waves leaving the computational domain, the current work integrates the Perfectly Matched Layer (PML) theory into the implicit MPM framework. The proposed approach introduces absorbing particles surrounding the computational domain that efficiently absorb outgoing waves and reduce reflections, allowing for accurate modeling of wave propagation and its further impact on geotechnical slope stability analysis. The study also includes several benchmark tests to validate the effectiveness of the proposed method, such as several types of impulse loading and symmetric and asymmetric base shaking. The conducted numerical tests also demonstrate the ability to handle large deformation problems, including the failure of elasto-plastic soils under gravity and dynamic excitations. The findings extend the capability of MPM in simulating continuous analysis of earthquake-induced landslides, from shaking to failure.

58 GEOSCIENCES↗

Flexible User-Defined Domain Decomposition in Kilometer-Scale E3SM Land Model Simulation

The Energy Exascale Earth System Model (E3SM) Land Model (ELM) has been extended to kilometer-scale (km-ELM) resolutions, enabling high-fidelity simulations of terrestrial processes at 1 km x 1 km grid spacing. In ELM, domain decomposition partitions the computational domain across processors, ensuring efficient parallel execution. Currently, round-robin decomposition is applied, providing a straightforward way to distribute computational workload. As ELM continues evolving at the kilometer-scale (km-scale), particularly with integrating lateral flow modeling, decomposition strategies must also account for the increased workload and data movement. This paper introduces a flexible user-defined domain decomposition framework, allowing users to customize domain partitioning based on application requirements. The impact of different decomposition strategies is evaluated across various applications concerning computation, communication, and I/O. Results demonstrate that while 1D partitioning yields superior I/O performance, k-nearest neighbors (KNN) clustering effectively reduces inter-process communication overhead. This study lays the groundwork for scalable partitioning in large-scale land surface simulations, enhancing next-generation Earth system modeling.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

Impact of Cloud-Base Turbulence on CCN Activation: Single-Size CCN

Abstract This paper examines the impact of cloud-base turbulence on activation of cloud condensation nuclei (CCN). Following our previous studies, we contrast activation within a nonturbulent adiabatic parcel and an adiabatic parcel filled with turbulence. The latter is simulated by applying a forced implicit large-eddy simulation within a triply periodic computational domain of 64 3 m 3 . We consider two monodisperse CCN. Small CCN have a dry radius of 0.01 μ m and a corresponding activation (critical) radius and critical supersaturation of 0.6 μ m and 1.3%, respectively. Large CCN have a dry radius of 0.2 μ m and feature activation radius of 5.4 μ m and critical supersaturation 0.15%. CCN are assumed in 200-cm −3 concentration in all cases. Mean cloud-base updraft velocities of 0.33, 1, and 3 m s −1 are considered. In the nonturbulent parcel, all CCN are activated and lead to a monodisperse droplet size distribution above the cloud base, with practically the same droplet size in all simulations. In contrast, turbulence can lead to activation of only a fraction of all CCN with a nonzero spectral width above the cloud base, of the order of 1 μ m, especially in the case of small CCN and weak mean cloud-base ascent. We compare our results to studies of the turbulent single-size CCN activation in the Pi chamber. Sensitivity simulations that apply a smaller turbulence intensity, smaller computational domain, and modified initial conditions document the impact of specific modeling assumptions. The simulations call for a more realistic high-resolution modeling of turbulent cloud-base activation.

54 ENVIRONMENTAL SCIENCES↗

PyJMAK: An Open-Source Python Toolkit for Modeling Solid-State Metallurgical Phase Transformations

Accurate prediction of metallurgical phase transformations is an essential basis for autonomous optimization and rapid part qualification. Several methods can be used to estimate the evolution of phase fractions such as JMAK kinetics-based models, phase-field models, thermodynamic models, and data-driven machine learning models. Thermodynamic and phase-field-based methodologies solve multiphysics equations requiring numerous calibration parameters and significant computational resources. As a result, the computation domain is limited to a point or on order of micron-meters. The data-driven models rely on large datasets from experiments and simulations. While the JMAK model only provides information about phase fraction evolution, it can predict this evolution in near real-time using thermal history and thermodynamic data without restriction on the domain. JMAK models have been popularly used by researchers to model phase transformations occuring during additive manufacturing or over arbitrary temperature profiles. Commercial proprietary software such as Abaqus and Ansys or closed-source in-house implementations offer the ability to model JMAK based kinetics to predict phase transformation. However, these software packages are not open-source or freely available for use and development in conjunction with manufacturing machines, sensors, and machine learning algorithms. In addition, the use of the model is restricted by a license token. In contrast, given temperature profiles at multiple points in the domain, this Python-based PyJMAK model can compute phase evolution in parallel due to its stand-alone modular, voxel-based structure, and it can be executed on high-performance computing resources without any license restrictions.

Prabhune, Bhagya [Oak Ridge National Laboratory (O↗

Scalable Simulation of Pressure Gradient-Driven Transport of Rarefied Gases in Complex Permeable Media Using Lattice Boltzmann Method

Accurate representations of slip and transitional flow regimes present a challenge in the simulation of rarefied gas flow in confined systems with complex geometries. In these regimes, continuum-based formulations may not capture the physics correctly. This work considers a regularized multi-relaxation time lattice Boltzmann (LB) method with mixed Maxwellian diffusive and halfway bounce-back wall boundary treatments to capture flow at high Kn. The simulation results are validated against atomistic simulation results from the literature. We examine the convergence behavior of LB for confined systems as a function of inlet and outlet treatments, complexity of the geometry, and magnitude of pressure gradient and show that convergence is sensitive to all three. The inlet and outlet boundary treatments considered in this work include periodic, pressure, and a generalized periodic boundary condition. Compared to periodic and pressure treatments, simulations of complex domains using a generalized boundary treatment conserve mass but require more iterations to converge. Convergence behavior in complex domains improves at higher magnitudes of pressure gradient across the computational domain, and lowering the porosity deteriorates the convergence behavior for complex domains.

42 ENGINEERING↗

A BOUT++ extension for full annular tokamak edge MHD and turbulence simulations

For tokamak edge plasma simulation, a plasma simulation framework BOUT++ employs a dual coordinate system to simulate moderate-n and high-n plasma instability with reasonable computational cost, where n is the toroidal mode number. This coordinate system however limits the computational domain to the toroidal wedge (full torus divided into N parts in the toroidal direction) for computational efficiency and the use of flute-ordering approximation in the field solver calculating the flow potential from the vorticity which may not be valid for low-n modes. Improving numerical treatment of low-n modes is however indispensable to address simulations of low-n current-driven edge localized mode (ELM), ELM control by resonant magnetic perturbations (RMPs), edge turbulence with RMPs and so on. In this work, BOUT++ is extended to simulate the interplay between $n=0$, low-n and high-n plasma components in a full annular tokamak edge domain through hybrid modeling of the flow potential and the vorticity. Low-n modes of flow potential are calculated in an orthogonal flux surface coordinate and high-n modes in the dual coordinate system separately in Fourier space. Finally, the proposed scheme can capture an interplay between $n=1$ global modes and high-n turbulence during pedestal collapse in a full annular torus domain with a circular cross section.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Low-Rank QTT-based Finite Element Method for Elasticity Problems

We present an efficient and robust numerical algorithm for solving the linear elasticity problem that combines the Quantized Tensor Train format and a domain partitioning strategy. This approach makes it possible to solve the linear elasticity problem on a computational domain that is more general than a square. By integrating Z-ordering and subdomain concatenation, our method substantially decreases memory usage and achieves a notable reduction in rank compared to established Finite Element implementations like the FEniCS platform. This efficiency is maintained while still guaranteeing exponential convergence with respect to the number of degrees of freedom. This performance gain, however, requires a fundamental rethinking of how core finite element operations are implemented. This includes changes to mesh discretization, node and degree of freedom ordering, stiffness matrix and internal nodal force assembly, and the execution of algebraic matrix-vector operations. In this work, we discuss all these aspects in detail and assess the method’s performance in the numerical approximation of three representative test cases.

97 MATHEMATICS AND COMPUTING↗

RANS Simulation of Variable Density Turbulent Round Jets with Coflow using xRAGE Hydrodynamic Code and BHR Turbulence Models

This work for the fiscal year 2023 (FY23) is a continuation of previous efforts to evaluate the BHR turbulence models for their ability to accurately simulate variable density turbulent round jets with coflow. As before, RANS simulations are carried out using the xRAGE hydrodynamic code. The following are some of the previous findings. Israel showed that i) the symmetry boundary conditions for the BHR models in axisymmetric simulations were in error, ii) three grids of different resolutions did not lead to converging solutions, and iii) BHR 3.1 simulation exhibited instabilities and did not reach a steady state. Saenz and Rauenzahn derived and implemented into xRAGE the correct BHR boundary conditions at the symmetry axis. Cline conducted sensitivity studies with various parameters including the BHR models (versions 2, 3.1 and 4), gravity, material pressure, specific heat, initial turbulent kinetic energy and initial turbulent length scale, and found that the largest factor impacting on simulation results was the BHR model version, followed by the initial turbulent length scale. In addition, freeze boundary conditions at the exit and the side wall of the computational domain were used to remove anomalous flow behavior. Cline adjusted the inlet jet velocity, initial turbulent kinetic energy and initial turbulent length scale to obtain the best reasonable match with the experimental data of Charonko and Prestridge. It was found that the BHR 2 and 3.1 models performed in a similar manner, but the BHR 2 model produced much lower levels of density-specific-volume covariance and turbulent kinetic energy. The main focus for the FY23 is to study the effects of computational parameters related to the boundary conditions, mesh, domain size and timestep size. The reasoning behind this is that, unless simulation results can be shown to be reasonably independent from the aforementioned computational parameters, it would be difficult to attribute any discrepancies between simulation and experimental results to turbulence models. This important aspect has largely been overlooked in the previous years, and therefore will be studied comprehensively here. Additionally, effects of varying the initial turbulent length scale will be examined because it was previously identified as a major factor affecting the flow fields.

42 ENGINEERING↗

A high accuracy/resolution spectral element/Fourier–Galerkin method for the simulation of shoaling non-linear internal waves and turbulence in long domains with variable bathymetry

A high-order hybrid continuous-Galerkin numerical method, designed for the simulation of non-linear, non -hydrostatic internal waves and turbulence in long computational domains with complex bathymetry, is presented. The spatial discretization in the non-periodic wave-propagating directions, utilizes the nodal spectral element method. Such a high-order element-based discretization allows the highly accurate representation of complex domain geometry along with the flexibility of concentrating resolution in areas of interest. Under the assumption of the normal-to-isobath propagation of non-linear internal waves, a third periodic direction is incorporated via a Fourier-Galerkin discretization. The distinct non-hydrostatic nature of non-linear internal waves and, any instabilities and turbulence therein, necessitates the numerically challenging solution of the pressure Poisson problem. A defining feature of this work is the application of a domain decomposition approach, combined with block-Jacobi/deflation-based preconditioning to the pressure Poisson problem. Such a combined approach is particularly suitable for the long high aspect-ratio complex domains of interest and enables the efficient high-accuracy reproduction of the non-hydrostatic dynamics of non-linear internal waves. Implementation details are also described in the context of the stability of the solver and its parallelization strategy. A series of benchmarks of increasing complexity demonstrate the robustness of the flow solver. The benchmarks culminate with the three-dimensional simulation of a convectively breaking mode-one non-linear internal wave over a realistic South-China-Sea bathymetric transect and background current/stratification profiles.

Deflation↗

A Survey on the Expanding Scope and Interdisciplinary Opportunities for Processing-in-Memory Techniques

Processing-in-Memory (PIM) is emerging as a practical path to overcome the limitations of traditional von Neumann architectures. At its core, PIM systems implement computing primitives such as logic operations and multiply-accumulate acceleration through compute-in-memory, near-memory processing, or hybrid designs. The role of memory cells varies widely across technologies, acting as inputs, outputs, or analog accumulators through bit-lines and sense amplifiers. This diversity creates trade-offs in precision, bandwidth, latency, and programmability, making it difficult to build a unified understanding on the progress of the field. In this survey, we organize recent advances of PIM into three areas. First, we discuss the progress on the architectural optimizations of PIM and its integration with both DRAM and emerging non-volatile memories. Second, we examine how PIM is being used to accelerate key computing domains, including generative AI workloads and high-performance kernels, along with new approaches. Third, we highlight the growing adoption of PIM in computational sciences, where it is being applied to solve interdisciplinary problems such as genome analysis, mRNA quantification, mass spectrometry, quantum circuit simulation, wave modeling, and secure computation. Finally, we synthesize the major challenges that continue to slow PIM adoption, including manufacturing constraints, power delivery, thermal reliability, data consistency, runtime and memory-management coordination, and the difficulty of building portable software abstractions without sacrificing commercial viability. This work provides an updated, structured perspective on PIM’s potential across computing and computational sciences and the barriers that must be solved for it to reach its full impact.

Asifuzzaman, Kazi [Oak Ridge National Laboratory (↗