Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computational framework”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing

This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.

Miniskar, Narasinga Rao↗

Progress Towards a Predictive Eagle Behavior and Risk Modeling Framework: Overview and Recent Validation Efforts

This presentation summarizes progress to date of the U.S. Department of Energy project, "Development of a computational framework for modeling golden eagles (Aquila chrysaetos) near wind farms," which focuses on stochastic behavioral modeling of soaring raptors across landscape, facility, and turbine spatiotemporal scales. This publicly available, open-source modeling framework includes behavioral models based on three different underlying principles: energy minimization at landscape scale, behavioral heuristics at landscape-facility scale, and data-driven behaviors at the facility-micro-scale. We will briefly overview the key advancements in the behavioral modeling state of the art, which leverages multiple high-resolution telemetry data sources combined with high-fidelity atmospheric flow modeling insights. We then present preliminary results from a validation study in Altamont, California. This new study involves a novel application of the Stochastic Soaring Raptor Simulator (SSRS), in a new geographic locale, to understand facility scale eagle movement patterns over time scales representative of a wind project's lifetime. For this desktop analysis (that does not depend on any high-performance computing resources), SSRS simultaneously considers a variety of wind conditions and eagle approach vectors toward a project site of interest. This work demonstrates the integration of publicly available landscape-scale atmospheric datasets, our recently improved engineering updraft models (see presentation from Thedin et al.), and our energy minimization behavioral models within the SSRS framework. While we only present results from a single behavioral model, the integration of these three modeling components forms the foundation for our more sophisticated behavioral models (see presentations from Brandes et al., Sandhu et al.) that are under active development. Results are presented in the form of presence maps, which may be applied to estimate risk to wildlife, augment ground survey data, inform wind-plant operations, or incorporated into wind-plant designs.

agent-based modeling↗

White Box Access to Quantum Testbeds for Co-Design

At Lawrence Livermore National Laboratory (LLNL), we operate and maintain the Quantum Device and Integration Testbed (QuDIT) facility, a small quantum testbed that supports about 10 active research teams (including our own) and over 50 internal and external collaborators. This testbed is designed to give remote white box access to users for research, training, and outreach. A guiding principle behind the development of our testbed infrastructure, software and user interfaces is to empower users to perform experiments at the cutting edge of quantum information science at any level of abstraction, from materials studies, device physics and control and characterization techniques to algorithm development and quantum operating system design. Our testbed targets a multilevel quantum system (qudit) to expand the accessible Hilbert space of a simple-to-manufacture quantum device and focuses on quantum simulation, typically implemented through custom gates designed with quantum optimal control methods, rather than on a universal computing framework with a fixed gate set. We leverage the Lab’s high-performance computing (HPC) program and related expertise to simulate quantum systems, develop hybrid algorithms, and generate gates optimized for given simulations. Additionally, we have adopted a co-design philosophy from the HPC community in designing new hardware, so that the systems we develop are optimized for the specific physics simulations we plan to use them for.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Privacy-Aware Federated Learning Framework for Distributed Energy Resource Analytics in Constrained Environments

To be resilient against extreme weather events, the rural communities in Puerto Rico are leveraging distributed energy resources (DER). However, computing frameworks sup-porting the grid in critical decision-making are still largely centralized. Sensitive consumer data are transmitted over the Internet or cellular networks to a secondary or tertiary node. It guarantees better situational awareness at the cost of a wider attack surface, jeopardizing user privacy, as more DER come online. Cloud, Edge, and Fog computing all require data aggregation at some level. This paper introduces a privacy-aware federated learning framework that leverages the Fog model by pushing analytics all the way to the DER and load assets. These local models train on individual asset data and transmit only learned parameters (such as weights) over secure communications to a global decision-maker. By abstracting personally identifiable consumer data without impacting decision optimality, this framework better aligns with distributed power generation paradigm.

Sundararajan, Aditya↗

Quantum embedding theories to simulate condensed systems on quantum computers.

Quantum computers hold promise to improve the efficiency of quantum simulations of materials and to enable the investigation of systems and properties that are more complex than tractable at present on classical architectures. Here, we discuss computational frameworks to carry out electronic structure calculations of solids on noisy intermediate-scale quantum computers using embedding theories, and we give examples for a specific class of materials, that is, solid materials hosting spin defects. These are promising systems to build future quantum technologies, such as quantum computers, quantum sensors and quantum communication devices. Although quantum simulations on quantum architectures are in their infancy, promising results for realistic systems appear to be within reach.

Vorwerk, Christian↗

AdditiveFOAM: A Continuum Multiphysics Code for Additive Manufacturing

AdditiveFOAM is a computational framework that simulates transport phenomena in Additive Manufacturing (AM) processes. It is built on OpenFOAM (Weller et al., 1998), the leading free, open-source software package for computational fluid dynamics (CFD). OpenFOAM offers an extensible platform for solving complex multiphysics problems using state-of-the-art finite volume methods. AdditiveFOAM leverages these capabilities to develop specialized tools aimed at addressing challenges in AM processing. Metal additive manufacturing, also known as metal 3D printing, is an advanced manufacturing technique that creates physical parts from a three-dimensional (3D) digital model by melting metal powder or wire feedstock. A significant area of research in metal AM focuses on process planning to mitigate anomalous features during printing that are deleterious to part performance (e.g., porosity and cracking), as well as controlling localized microstructure and material properties. Given the high costs and substantial time requirements associated with experimental methods for qualifying new materials and processes, there is a compelling incentive for researchers to utilize advanced computational simulations. In this context, AdditiveFOAM offers a simulation framework to better understand undesirable features in printing, thereby enhancing process planning and reducing the reliance on labor-intensive experimental campaigns.

Coleman, John [Oak Ridge National Laboratory (ORNL↗

Enhanced read resolution in reconfigurable memristive synapses for Spiking Neural Networks

Abstract The synapse is a key element circuit in any memristor-based neuromorphic computing system. A memristor is a two-terminal analog memory device. Memristive synapses suffer from various challenges including high voltage, SET or RESET failure, and READ margin issues that can degrade the distinguishability of stored weights. Enhancing READ resolution is very important to improving the reliability of memristive synapses. Usually, the READ resolution is very small for a memristive synapse with a 4-bit data precision. This work considers a step-by-step analysis to enhance the READ current resolution or the read current difference between two resistance levels for a current-controlled memristor-based synapse. An empirical model is used to characterize the $${\hbox {HfO}}_{2}$$ HfO 2 based memristive device. $$1\textrm{st}$$ 1 st and $$2\textrm{nd}$$ 2 nd stage device of our proposed synapse design can be scaled to enhance the READ current margin up to $$\sim$$ ∼ 4.3 $$\times$$ × and $$\sim$$ ∼ 21%, respectively. Moreover, READ current resolution can be enhanced with run-time adaptation techniques such as READ voltage scaling and body biasing. The READ voltage scaling and body biasing can improve the READ current resolution by about 46% and 15%, respectively. TENNLab’s neuromorphic computing framework is leveraged to evaluate the effect of READ current resolution on classification, control, and reservoir computing applications. Higher READ current resolution shows better accuracy than lower resolution even when facing different levels of read noise.

97 MATHEMATICS AND COMPUTING↗

Autonomous nondestructive evaluation of resistance spot welded joints

The application of non-destructive evaluation approaches has attracted strong interests in modern automotive industries. Here, we present an autonomous deep-computing framework to analyze raw videos from infrared systems and to predict weld nugget shape and size with unprecedented accuracy and speed. In a comprehensive training and testing experiment with 90 videos (seven sets of welding material stack-ups), a new method was developed to assemble sufficient datasets for neural network training. Our framework successfully predicts all the nugget shapes with F1 scores that range from 0.84 to 0.92. The total training time on Nvidia DGX station takes less than 10 min for each set of welding material stack-up. The real inference time of an individual dataset (with 30 video frames) takes about 0.005 s. The procedure and methods developed in the study can be applied to other image-based weld property prediction, as well as other manufacturing processes. Furthermore, our well-trained neural networks take limited memory resources (2.3 MB) and are suitable for embedded microprocessors for in-situ welding quality control as edge computing within an intelligent welding framework.

42 ENGINEERING↗

Nanoindentation mapping defects filtration for heterogeneous materials using generative adversarial networks

Advanced composite materials with multiple phases and heterogeneous microstructure necessitate spatial mapping characterization of elastic modulus to develop constitutive relations and overall mechanical response. Such modulus mapping can be obtained using the nanoindentation technique, where the indenter tip raster over the selected microstructure region. Typically, a surface preparation procedure is done in the specimens to ensure proper contact between the indenter tip and sample surface. However, a near-perfect surface finish is unachievable in heterogeneous materials, primarily with ceramic reinforcements, due to the differential material removal rate during polishing. Thus, the nanoindenter records localized erroneous measurements due to differences in surface roughness and corresponding force response. This study establishes a novel deep learning-based strategy to rectify incorrect experimental spatial measurements acquire during nanoindentation modulus mapping. Here, the integrated bicubic interpolation and generative adversarial networks (GANs) model was trained using 14 ceramic and 18 metallic data sets, each comprising 65,536 measurements. The developed algorithm was validated against experimental measurements on four unknown specimens. The standard deviation in measured elastic modulus reduces by ~50% in ceramics and ~72% in metallic samples. This computational framework proposes a novel approach to reducing uncertainty in materials’ properties using state-of-the-art computer vision techniques.

36 MATERIALS SCIENCE↗

A Framework for Neural Network Inference on FPGA-Centric SmartNICs

FPGA-based SmartNICs offer great potential to significantly improve the performance of high-performance computing and warehouse data processing by tightly coupling support for reconfigurable data-intensive computation with cross-node communication, thereby mitigating the von Neumann bottleneck. Existing work, however, has been generally been limited in that it assumes an accelerator model where kernels are offloaded to SmartNICs, but most control tasks are left to the CPUs. This leads to frequent waiting, inferior performance, and scaling challenges. In this work, we propose a new distributive data-centric computing framework, named FCsN, for reconfigurable SmartNIC-based systems. Through a lightweight task circulation execution model and its implementation architecture, FCsN allows the complete detaching of kernel execution, control logic, system scheduling, and network communication to the SmarNICs. This boosts performance by: (i) avoiding the control dependency with CPUs and (ii) supporting streaming kernel execution and network communication at line rate and in a very fine-grained manner. We demonstrate the efficiency and flexibility of FCsN using various types of neural network applications including graph neural networks; as these last are both irregular and data intensive they offer an especially robust demonstration. Evaluations using commonly-used neural network models and graph datasets show that a system with the support of FCsN can achieve, on average, 144 speedups over the MPI-based standard CPU baselines.

Guo, Anqi↗

A Machine Learning Framework for Modeling Ensemble Properties of Atomically Disordered Materials

Atomic disorder can strongly influence material properties such as charge transport, optical response, and catalytic activity. However, efficiently modeling these disorder effects remains challenging for first-principles methods due to the cost of sampling large configurational spaces and computing complex physical quantities. Recent advances of machine learning techniques, particularly graph neural networks (GNNs), has enabled the efficient and accurate predictions of complex material properties, offering promising tools for studying disordered systems. In this work, we present a general machine-learning-assisted computational framework that integrates equivariant GNNs with Monte Carlo simulations to compute the thermodynamic and ensemble-averaged functional properties of disordered materials. Using the surface-termination-disordered MXene monolayer Ti 3 C 2 T 2–x as a representative system, we find that electrical conductivity exhibits an emergent peak near the order–disorder phase transition temperature due to the interplay between electron scattering and doping. In contrast, optical conductivity remains largely insensitive to local atomic disorder and reflects the global surface chemical composition. These results highlight the role of atomic disorder in affecting material properties and demonstrate the potential of our approach for statistically modeling disorder effects in a wide range of materials such as high-entropy alloys and spin liquids.

MXene↗

An Atomistic Study of Reactivity in Solid-State Electrolyte Interphase Formation for Li/Li7P3S11

Lithium metal batteries offer superior volumetric and gravimetric specific capacities compared to those based on traditional graphite anodes. Although advancements in solid-state electrolytes address safety concerns, challenges remain, particularly regarding interphase formation in lithium metal anodes. This work presents a computational framework based on high-throughput first-principles density functional theory and machine-learning interatomic potentials (MLIPs) including automated iterative, active learning to enable robust computational exploration of interphase formation between lithium metal anodes and an inorganic solid-state electrolyte. As a demonstration, we apply the framework to a Li/Li7P3S11 interface and find that it accurately identifies the experimentally observed, thermodynamically stable interphase products as well as their overall spatial arrangement within a heterogeneous, amorphous layered structure, with Li2S domains of nanocrystallinity. Our simulations show two stages, a fast and slow diffusion reaction regime, that corroborate the relative phase formation rate of Li x P, Li2S, and Li3P. Using the Onsager transport theory, we capture time-dependent ionic diffusion within the reacting interface, including cross-correlation effects. We found that cross-correlation effects between Li-P and P-S ionic motion significantly influence P-ion diffusion, making it highly sensitive to the local environment and potentially leading to "kinetic trapping" of Li-P phases. The passivation of the interface is shown as the ionic fluxes all approach zero, effectively halting interphase growth.

Diffusion↗

Adaptive simulations enable computational design of electron beam processing of nanomaterials with supersonic micro-jet precursor

Focused Electron Beam Induced Processing (FEBIP) is a powerful tool for the “direct-write” of nanomaterials with the possibility of atomistic control on suspended 2D material substrates. FEBIP capabilities have been significantly expanded by using a localized jet-based delivery of precursors, and especially when a thermally energized supersonic micro-jet enhances the delivery of mass flux along with controlling the far-from-equilibrium thermodynamic state of adsorbed adatoms. The possibilities of growing nanomaterials with “dialed-in” composition with ultra-high growth rates and aspect ratios up to 100:1 have been demonstrated using the supersonic micro-jet-FEBIP. Bringing this scientific discovery to the level of maturity required for practical applications in additive nanomanufacturing requires simulation tools that are capable of capturing the complex flow physics of micro-jet-substrate interactions that bridge a wide range of flow regimes from the high-density gas micro-jet expanding into the vacuum environment of FEBIP. To address this significant computational challenge, a new approach has been developed and described in this work for a multiscale adaptive DSMC (Direct Simulation Monte Carlo) algorithm to predict the micro-jet gas dynamics in an FEBIP environment, spanning the full range of flow regimes from low Knudsen (Kn) number O(0.01) continuum flow to high Kn of O(10) for the molecular flow within a unified computational framework. Here, the fundamental principles of the adaptive DSMC algorithm are described, its viability as a DSMC technique is demonstrated, and computational improvements for the benchmark cases are discussed in the context of advancing 3D nanofabrication with FEBIP. Ultimately, combining the first principle simulations via adaptive DSMC with the complementary experimental data will enable the creation of the powerful CAD tools for in silico design and optimal operation of micro-jet-FEBIP.

36 MATERIALS SCIENCE↗

Polarization consistent dielectric screening in polarizable continuum model calculations of solvation energies

A polarization consistent framework, where dielectric screening is affected consistently in polarizable continuum model (PCM) calculations, is employed for the study of solvation energies. The computational framework combines a screened range-separated-hybrid functional (SRSH) with PCM calculations, SRSH-PCM, where dielectric screening is imposed in both PCM self-consistent reaction field (SCRF) iterations and the electronic structure Hamiltonian. We begin by demonstrating the impact of modifying the Hamiltonian to include such dielectric screening in SCRF iterations by considering the solutions of electrostatically embedded Hartree–Fock (HF) exact exchange equations. Long-range screened HF-PCM calculations are shown to capture properly the linear dependence of gap energy of frontier orbitals on the inverse of the dielectric constant, whereas unscreened HF-PCM orbital energies are fallaciously semi-constant with respect to the dielectric constant and, therefore, inconsistent with the ionization energy gaps. Similar trends affect density functional theory (DFT) calculations that aim to achieve predictive quality. Importantly, the dielectric screened calculations are shown to significantly affect DFT- and HF PCM-based solvation energies, where screened solvation energies are smaller compared to the unscreened values. Importantly, SRSH-PCM, therefore, appears to reduce the tendency of DFT-PCM to overestimate solvation energies, where we find the effect to increase with the dielectric constant and the polarity of the molecular solute, trends that enhance the quality of DFT-PCM calculations of solvation energy. Understanding the relationship of dielectric screening in the Hamiltonian and DFT-PCM calculations can ultimately benefit on-going efforts for the design of predictive and parameter free descriptions of solvation energies.

Chemistry↗

Relativistic Exact Two-Component Theory in the Generalized Pseudospectral Representation

We present a formulation and implementation of exact two-component (X2C) relativistic theory in the generalized pseudospectral representation. When combined with the Hartree-Fock-Slater framework, this approach enables efficient and accurate treatments of scalar-relativistic and spin-orbit effects in atomic electronic structures without the computational overhead of four-component methods. Benchmark calculations across light and heavy elements demonstrate that the our X2C scheme yields substantially more accurate relativistic corrections to core-electron binding energies than perturbational Breit-Pauli treatments while converging more rapidly with respect to basis size. The method provides an improved computational framework for modeling ultrafast x-ray-induced processes in heavy-element systems.

Wang, Xubo↗

An ultrahigh-resolution E3SM land model simulation framework and its first application to the Seward Peninsula in Alaska

The availability of supercomputers and state-of-science datasets has made it possible to conduct large-scale land simulations at an ultrahigh-resolution. This study reported a computational framework for land surface simulation using the E3SM land model (ELM) at an unprecedented resolution (1 km x 1 km gridcell). The ultrahigh-resolution ELM (uELM) simulation framework includes three parts: (1) high-resolution atmospheric forcing and surface properties dataset generation, (2) massive gridcell-based simulation, and (3) large-scale simulation results analysis. Additionally, we implemented the uELM simulation framework and completed the first 1 km x 1 km terrestrial ecosystem simulation (from 1850 to 2014) over the Seward Peninsula in Alaska (78,000 km 2 ). The experiment contained two phases: a spin-up simulation and a transient simulation, and required five weeks of calculations using 320 cores in a 44-node Linux HPC computer. It created approximately 1.3 TB of data from the transient simulation alone (1850 - present). We selected sample results (monthly and daily simulation outputs) to illustrate the temporal and spatial variations of several variables in high-latitude Arctic ecosystems’ water, energy, and carbon cycles. At last, we summarized the lessons learned and proposed new developments for full-scale uELM simulations over the entire North American continent (approximately 22,000,000 km 2 ).

54 ENVIRONMENTAL SCIENCES↗

SPARC-X: Quantum simulations at extreme scale - reactive dynamics from first principles

We have developed the massively parallel electronic structure code SPARC-X: a computational framework for performing Kohn-Sham Density Functional Theory (DFT) calculations that can scale linearly with the number of atoms in the system, while being able to leverage petascale and emerging exascale parallel computers to study chemical phenomena at unprecedented length and time scales. SPARC-X exploits a recent breakthrough in electronic structure methodologies: systematically improvable, strictly local, orthonormal, discontinuous real-space bases that efficiently and systematically capture the local chemistry of the system. With further adaptation using new machine-learning techniques and the use of the massively parallel Spectral Quadrature (SQ) electronic structure method, the algorithmic complexity and prefactor associated with DFT calculations involving semilocal as well as hybrid functionals are dramatically reduced. Using petascale computational resources, SPARC-X enables quantum mechanical simulations at length and time scales previously accessible only by empirical approaches, e.g., 1,000,000 atoms for a few picoseconds using semilocal functionals or 1,000 atoms for a few picoseconds using hybrid functionals. Using exascale resources, the sizes and times targeted are two orders of magnitude larger. Such a capability has applications in a wide variety of chemical sciences, including reactive interfaces where large length- and/or long time-scales are needed and traditional force fields fail. This is particularly important in dynamic catalysis, where bond breaking and formation must be understood in detail. We developed, tested, and employed the SPARC-X framework to understand the photocatalytic properties of TiO 2 nanoparticles, revealing finite size effects that cannot be captured with standard model systems or functionals. This integrated development and application strategy ensures that SPARC-X remains a robust, efficient, and scalable software package for quantum simulations on current petascale and emerging exascale computing resources.

97 MATHEMATICS AND COMPUTING↗

Unified, Geometric Framework for Nonequilibrium Protocol Optimization

Controlling thermodynamic cycles to minimize the dissipated heat is a long-standing goal in thermodynamics, and more recently, a central challenge in stochastic thermodynamics for nanoscale systems. Here, we introduce a theoretical and computational framework for optimizing nonequilibrium control protocols that can transform a system between two distributions in a minimally dissipative fashion. These protocols optimally transport a system along paths through the space of probability distributions that minimize the dissipative cost of a transformation. Furthermore, we show that the thermodynamic metric—determined via a linear response approach—can be directly derived from the same objective function that is optimized in the optimal transport problem, thus providing a unified perspective on thermodynamic geometries. As a result, we investigate this unified geometric framework in two model systems and observe that our procedure for optimizing control protocols is robust beyond linear response.

36 MATERIALS SCIENCE↗