Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Domain Distributed Framework”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Quantum Simulators and Applications on Quantum Framework

Simulating quantum circuits is essential for validating quantum algorithms. However, no single simulator consistently performs best - efficiency depends on circuit structure, entanglement, and depth. In this work, we integrate Qiskit-Aer (state-vector and matrix product state) and QTensor, a tree-tensor-network based simulator, into the Quantum Framework (QFw), a modular platform that supports multiple quantum backends via a unified interface. We also enable distributed quantum approximate optimization algorithm (DQAOA) application compatibility with QFw, allowing sub-problems to be solved in parallel at scale. We then benchmark DQAOA and TFIM (transverse field Ising model) circuits across supported simulators, showing how performance varies significantly with problem type. All simulations are deployed on the Frontier supercomputer using QFw's MPI-based orchestration for distributed, multinode execution. These results underscore the need for simulatoragnostic infrastructure to enable systematic evaluation and highperformance scaling of quantum workloads. QFw provides a practical and extensible path toward reproducible quantum algorithm development across diverse application domains.

Chundury, Srikar [ORNL] (ORCID:0009000183359259)↗

Determining Levels of Detail for Simulators of Parallel and Distributed Computing Systems via Automated Calibration

There are two sources of inaccuracy when simulating parallel and distributed computing systems: (i) a simulator implemented at an insufficient level of detail; and (ii) incorrectly calibrated simulation parameter values. Increasing the simulator’s level of detail can improve accuracy, but at the cost of higher space, time, and/or software complexity. Furthermore, evaluating the intrinsic accuracy of a simulator requires that its parameters be well-calibrated. Making decisions regarding the level of detail is thus challenging. We propose a methodology for instantiating the simulation calibration process and a framework for automating this process, which makes it possible to pick appropriate levels of detail for any simulator. We demonstrate the usefulness of our approach via two case studies for two different domains.

McDonald, Jessie [University of Hawaii at Manoa, H↗

TRIM: AI Guided Random Number Generation for Resource-Constrained IoT Systems

Random numbers often serve as the backbone for many security solutions in diverse domains such as cryptography, side channel leakage prevention, and moving target defense. However, generating true random numbers requires a physical source of entropy (e.g. hardware, quantum, environmental phenomenon) making it difficult to realize at a large scale and at a low cost. On the flip side, pseudorandom number generators (easy to implement) following a specific distribution (e.g. Gaussian) can be easily compromised given a sufficient amount of traces. In this work, we have developed a machine learning-guided generative approach that can be used to create portable, resource-efficient, and cost-effective random number generators with high throughput and true randomness characteristics. We implement the proposed approach as a highly parameterized framework and perform extensive evaluation for different settings. The framework was able to learn from true random sources such as irrational numbers and environmental audio noise and imitate those sources towards generating new good quality random numbers on demand. We have generated more than 1 billion bits and observed robust performance in terms of true randomness metrics obtained from NIST SP 800-22 and FIPS 140-1 randomness test suites achieving a throughput of up to 142.85 Mbps. Compared to the state-of-the-art (SOTA) technique, the iso-cost setup of our framework can achieve more than 500 Mbps in a distributed setting. We have evaluated the efficacy of running the true randomness imitation AI models on target edge devices such as Raspberry Pi 4 (Model B), Nvidia Jetson Nano, Nvidia Jetson Orin Nano and Nvidia Jetson Xavier. We have also looked at the security of the TRIM framework itself against different adversarial threat models.

Cybersecurity↗

Preparing For Advanced Air Mobility: A System-Level Framework For Infrastructure, Investment, and Deployment

Advanced air mobility (AAM) is moving from demonstration toward early deployment, supported by a growing national strategy that outlines how these systems may evolve across airspace, infrastructure, and operations. As this transition takes shape, a more practical question comes into focus: what does it mean to be ready? This paper introduces a Capability Maturity Model (CMM) as a structured way to think about that challenge. Rather than treating readiness as a fixed condition, it frames it as a progression - one that develops across multiple, interdependent domains over time.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Assessing the Sensitivity of Pourbaix Diagrams to Computational Protocols: Electrochemical Stability of Ni Oxides as a Case Study

Pourbaix diagrams stand as a useful tool in assessing and visualizing materials’ electrochemical stability and are widely used for electrocatalyst design. However, their reliability hinges on the accuracy of the chemical potentials of involved phases, which may bear uncertainties and can be significantly impacted by decision-making steps in the computational protocol. Here, this study introduces a robust sensitivity analysis framework, exemplified through a detailed examination of the computational Pourbaix diagram of Ni, the oxides of which are used as high-activity and cost-friendly catalysts for many electrochemical reactions. Quantities of interest derived from the Pourbaix diagram include the appearance and stability domain of the catalytically active Ni oxide phases along with the onset electrochemical potentials of phase transitions. These metrics can guide the design of operational conditions for Ni oxide electrocatalysts. We find that the employed DFT exchange-correlation functional has the most significant influence on the computed Pourbaix diagram. Uncertainties on crystal structures, along with their related ab initio energetics, are also found to affect the size of the phase stability domain. Higher-order coupling among input parameters is found to play a crucial role in influencing the appearance and distribution of Ni phases in the diagram. Our findings suggest a need to consider variations and uncertainties associated with the computational procedures on predicted Pourbaix diagrams for materials design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multi-head physics-informed neural networks for learning functional priors and uncertainty quantification

In numerous applications, the integration of prior knowledge and historical information is essential, particularly for tasks requiring the solution of ordinary or partial differential equations (ODEs/PDEs) in data-sparse or noisy environments. For instance, achieving accurate solutions to time-dependent PDEs with limited initial condition measurements necessitates an effective strategy for embedding prior knowledge. Hard-parameter sharing architectures in neural networks (NNs) have demonstrated success in both traditional and scientific machine learning domains, facilitating the learning of informative representations. Here, in this study, we introduce a novel, yet efficient, method to enhance physics-informed neural networks (PINNs) by incorporating a multi-head structure that enables the learning of functional priors from both empirical data and governing physical laws. This prior information can then be used to address data sparsity and high-level noise in solving ODE/PDE problems with uncertainty quantification (UQ). The approach, termed Multi-Head PINN (MH-PINN), consists of a shared body NN and multiple head NNs, each corresponding to an individual PINN instance. Our framework for functional prior learning is carried out in two stages: (1) training the MH-PINNs to develop a shared body NN alongside multiple head NNs, and (2) employing these trained head NNs to estimate a prior distribution through a normalizing flow-based density estimator. The learned functional prior can then be applied as a regularization mechanism in deterministic contexts or as an informative prior within a Bayesian inference framework, aiding in the resolution of subsequent ODE/PDE tasks. We evaluate the efficacy of MH-PINNs across five benchmark problems, including a high-dimensional parametric PDE, all characterized by data sparsity or substantial noise levels. Our findings reveal that MH-PINNs deliver accurate solutions and robust UQ, demonstrating adaptability across a range of complex and challenging scenarios.

Bayesian inference↗

Parallel quantum computing simulations via quantum accelerator platform virtualization

Quantum circuit execution is a central task in quantum computation. Due to inherent quantum-mechanical constraints, quantum computing workflows often involve a considerable number of independent measurements over a large set of slightly different quantum circuits. Here we discuss a simple model for parallelizing such quantum circuit executions that is based on introducing a large array of virtual quantum processing units (mapped to HPC nodes in our case) as a parallel quantum computing platform. Implemented within the XACC framework, the model can readily take advantage of its backend-agnostic features, enabling parallel quantum computing/simulation over any target backend supported by XACC. We illustrate the performance of this approach by demonstrating strong scaling in two pertinent domain science problems, namely in computing the gradients for the multi-contracted variational quantum eigensolver and in data-driven quantum circuit learning, where we vary the number of qubits and the number of circuit layers. Here, the latter simulation leverages the cuQuantum library to run efficiently on GPU-accelerated HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Enabling probabilistic learning on manifolds through double diffusion maps

Here, we present a generative learning framework for probabilistic sampling that extends Probabilistic Learning on Manifolds (PLoM), which is designed to generate statistically consistent realizations of a random vector in a finite-dimensional Euclidean space, informed by a (representative) set of observations. In its original form, PLoM constructs a reduced-order probabilistic model by combining three main components: (a) kernel density estimation to approximate the underlying probability measure, (b) Diffusion Maps to characterize the manifold of the data, and (c) a reduced-order Itô Stochastic Differential Equation (ISDE) to sample from the learned distribution. However, its sampling dynamics are posed in the ambient space and the retained number of reduced coordinates is chosen by projection-reconstruction error. In practice, this often (i) requires more coordinates than the data’s intrinsic dimension to achieve stable sampling and (ii) lacks a smooth, basis-independent lifting back to the data domain; moreover, standard Diffusion Maps emphasize harmonic eigenfunctions and can miss non-harmonic latent structure. We address these limitations by decoupling geometry learning from sampling: a first Diffusion Maps pass identifies non-harmonic coordinates on which we formulate a full-order ISDE directly in the latent space, while Double Diffusion Maps captures multiscale geometric features and Geometric Harmonics (GH) learns a smooth lifting map to the ambient variables that is independent of the particular diffusion basis. This hybrid design preserves the system’s dynamical richness with a compact geometric representation and enables principled out-of-sample inference. The effectiveness and robustness of the proposed method are illustrated through two numerical studies: one based on data generated from two-dimensional Hermite polynomial functions and another based on high-fidelity simulations of a detonation wave in a reactive flow.

Double diffusion maps↗

HARD: A performance portable radiation hydrodynamics code based on FleCSI framework

Hydrodynamics And Radiation Diffusion (HARD) is an open-source application for high-performance simulations of compressible hydrodynamics with radiation-diffusion coupling. Built on the FleCSI (Bergen et al., 2021 [1]) (Flexible Computational Science Infrastructure) framework, HARD expresses its computational units as tasks whose execution can be orchestrated by multiple back-end runtimes, including Legion (Bauer et al., 2012 [2]), MPI (Forum, 1994 [3]), and HPX (Kaiser et al., 2020 [4]). Node-level parallelism is handled through Kokkos (Edwards et al., 2014 [5]), providing a single-source, portable code base that runs efficiently on laptops, small homogeneous clusters, and the largest heterogeneous supercomputers currently available. To ensure scientific reliability, HARD includes a regression test suite that automatically reproduces canonical verification problems such as the Sod and LeBlanc shock tubes, and the Sedov blast wave, comparing numerical solutions against known analytical results. The project is distributed under an OSI-approved license, hosted on GitHub, and accompanied by reproducible build scripts and continuous integration workflows. This combination of performance portability, verification infrastructure, and community-focused development makes HARD a sustainable platform for advancing radiation hydrodynamics research across multiple domains.

97 MATHEMATICS AND COMPUTING↗

Poisson-response Tensor-on-Tensor Regression and Applications

We introduce Poisson-response tensor-on-tensor regression (PToTR), a novel regression framework designed to handle tensor responses composed element-wise of random Poisson-distributed counts. Tensors, or multi-dimensional arrays, composed of counts are common data in fields such as inter national relations, social networks, epidemiology, and medical imaging, where events occur across multiple dimensions like time, location, and dyads. PToTR accommodates such tensor responses alongside tensor covariates, providing a versatile tool for multi dimensional data analysis. We propose algorithms for maximum likelihood estimation under a canonical polyadic (CP) structure on the regression coefficient tensor that satisfy the positivity of Poisson parameters and then provide an initial theoretical error analysis for PToTR estimators. We also demonstrate the utility of PToTR through three concrete applications: longitudinal data analysis of the Integrated Crisis Early Warning System database, positron emission tomography (PET) image reconstruction, and change-point detection of communication patterns in longitudinal dyadic data. These applications highlight the versatility of PToTR in addressing complex, structured count data across various domains.

97 MATHEMATICS AND COMPUTING↗

Real-Time Lifetime Prediction of Semiconductor Devices Using Hardware-in-the-Loop

This paper presents a unique approach to enable real-time lifespan prediction of semiconductor power modules using a Hardware-in-the-Loop (HIL) system. By integrating the module's overall loss characteristics-specifically switching and conduction losses-with a thermoelectric model of the thermal management system, this research demonstrates that the model can dynamically estimates the junction temperature profile of the semiconductor devices in response to a changing torque demand profile for the motor drive system. This capability enables continuous monitoring of the module's operational time and cumulative stress induced on the devices to compute accumulated remaining lifetime or time-to-failure (TTF). This study provides an architectural framework for the HIL system with high-fidelity component models of multiple physical domains, allowing simulation of dynamic behaviors of a closely-coupled motor drive system. The advanced real-time computation and measurement functionalities of the HIL system allow for both dynamic lifetime calculations based on simulated data and aggregate lifetime predictions utilizing historical data. Moreover, this paper details an algorithm that not only computes cumulative damage but also synthesizes these data into a comprehensive aggregated lifetime metric. This methodology can enhance the maintenance scheduling strategies and operational reliability of semiconductor devices in critical applications, ultimately extending their service life while optimizing performance.

hardware-in-the-loop (HIL)↗

Tailoring composition and deformation modes at the microstructural level for next generation low-cost high-strength austenitic stainless steels

The objective of this project is to enable deliberate development of cost-effective, hydrogen resistant alloys by establishing detailed relationships specific to the effects of alloy composition, short-range order (SRO), and microsegregation in the presence of hydrogen on the transition between homogeneous deformation and localized plasticity in shear bands. In collaboration with the International Institute for Carbon-Neutral Energy Research, I2CNER, at Kyushu University in Japan, we conceptualized, designed, and manufactured four austenitic alloys that maintain corrosion resistance and ensure lower cost relative to baseline commercial alloys. The mechanical properties and deformation modes of the novel alloys (KU alloys) were assessed in the presence of hydrogen (H). Correlations between composition and performance revealed that two of the KU alloys are suitable replacements for 316 steel, while another is a viable replacement for 304 steel at room temperature. We found that, in the presence of other austenite stabilizing elements namely Mn and N, replacing Ni with Cu does not lead to martensite formation as has been previously reported.1–3 Furthermore, we found that the addition of Cu leads to an earlier onset of multiple slip resulting in an relative earlier onset of a higher work hardening rate (WHR). Greater understanding of the relationships between alloy composition and SRO required the development of a novel advanced electron diffraction methodology to characterize SRO in complex FCC alloys. This innovative approach, which combines fluctuation and correlation analyses of diffuse-scattering signals, successfully differentiated between SRO and long-range ordering (LRO). Further investigations into annealed austenitic stainless steels could provide insights into manipulating SRO and its effects on material properties. Atomistic simulations provided understanding of SRO behavior that was difficult to capture experimentally. This project created the first spin cluster expansion model that is able to capture and describe SRO effects in Fe-Ni-Cr FCC alloys, accounting for the non-negligible effects of magnetism. An automated computational workflow was established to provide reliable predictions of SRO in Fe-Ni-Cr austenitic alloys, both with and without the presence of H atoms. Analysis of the propensity for SRO in Fe-Ni-Cr alloys revealed that H tends to cluster with specific, well-defined SRO domains. The computational framework is general purpose and can be extended to realistic stainless steels across diverse composition ranges. With confidence that SRO is possible in austenitic stainless steels, we developed a discrete dislocation finite element code to understand the interaction of dislocations with SRO in the presence of H. By incorporating H effects on the dislocation emission and SRO stress field we show that the critical stress for the dislocation pileup to breakthrough the SRO domain decreases in the presence of H, which directly contributes localized deformation at the macroscale. Through the simulation of a uniaxial tension test, we demonstrated that H-induced weakening of SRO stress field and H-enhanced dislocation emission can lead to the onset of shear localization at lower macroscopic strains. As a whole, this project identified three novel alloys that show improvements in performance and cost efficiency for H-facing applications by studying correlations between alloy chemistry and deformation behavior. We also made significant advancements to experimental and computational methodologies necessary to study the chemistry and distribution of SRO across a range of alloys, which in turn allowed us to demonstrate how deformation mechanisms change due to the contributions of SRO in austenitic alloys in the presence of H. The combined advancements in fundamental understanding with novel alloy development in this project has increased the viability of next generation H-technologies for the broader public through accessible low-cost alloys and accelerated development towards future H-infrastructure.

08 HYDROGEN↗

A segmented approach to modeling building height: Delineating high-rise and low-rise buildings for enhanced height estimation

Understanding building height is imperative to the overall study of energy efficiency, population distribution, urban morphologies, emergency response, among others. Currently, existing approaches for modeling building height at scale are hindered by two pervasive issues. First, there is no consistent approach to quantify what a high-rise building is at a macro scale, leaving researchers unable to accurately compare results across geographies and domains. Second, high-rise buildings represent a small fraction of the built environment, implying data imbalance challenges that negatively affect current approaches. This is a problem of practical relevance since information on high-rise buildings is important for studies on urban heat islands, population dynamics, and pollution dispersion. Here, we introduce a novel approach to map building height which first identifies two distinct distributions within the built environment, with one being composed of low-rise buildings and one composed of high-rise buildings. We then develop an ensemble scheme where discrete specialist models are trained for each subset of low-rise buildings and high-rise buildings to infer building height from morphology features. For experiments mapping heights of 4.85 million buildings in Japan, we show an increase of 34 % in accuracy within 3m error when compared to the current state-of-the-art when modeling high-rise buildings, which based on KNN experimentation we define as any building > 12m . Our findings show that such an ensemble framework outperforms the current state-of-the-art approaches, which is especially relevant in relation to inferring height for high-rise buildings, a prominent issue of existing approaches for mapping the built environment.

97 MATHEMATICS AND COMPUTING↗

Correspondence between Color Glass Condensate and High-Twist Formalism

The color glass condensate (CGC) effective theory and the collinear factorization at high twist (HT) are two well-known frameworks describing perturbative QCD multiple scatterings in nuclear media. It has long been recognized that these two formalisms have their own domain of validity in different kinematic regions. Taking direct photon production in proton-nucleus collisions as an example, we clarify for the first time the relation between CGC and HT at the level of a physical observable. We show that the CGC formalism beyond shock-wave approximation, and with the Landau-Pomeranchuk-Migdal interference effect is consistent with the HT formalism in the transition region where they overlap. Such a unified picture paves the way for mapping out the phase diagram of parton density in nuclear medium from dilute to dense region.

Parton distribution functions↗

Grid Architecture Mapping to Understand Transformation (GAMUT): Methods and Framework Architecture

Grid architecture (GA) is a concept that was developed to address the need for a comprehensive view of power grid challenges. GA can be viewed as a relatively consistent and fixed high-level approach; however, for any instantiation of grid structures, a combinatorial explosion results from each lower-layer expansion. This constitutes the main challenge with GA—it is a grid architect’s view of the system, which might not be very informative at the implementation level. Grid Architecture Mapping to Understand Transformation (GAMUT project) seeks to bridge that gap by integrating subject matter expertise across GA structures, providing users who lack expertise in GA approaches with valuable insights and informational materials. GAMUT seeks to answer feasibility questions for the approach. System-level expectations are that a GA baseline needs to be established in order for GA to be the common framework to which any lower layer approach is tied. This report explores a potential information ingestion and documentation framework to support GAMUT. The main concepts that enable the solution domain of GAMUT are discussed, and examples are provided. The solution domain leverages already-existing technology and concepts related to GA, knowledge management, and other relevant areas. To assess GAMUT building blocks and the overall approach, a feasibility assessment is proposed, rooted in systems engineering and GA architecture evaluation concepts.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Summary of GFL Inverter Capability and Performance: Malawi [Slides]

This presentation provides a comprehensive comparison of inverter-based resource (IBR) technical requirements across multiple international grid codes - including IEEE 2800, ESO, EirGrid, and Malawi's evolving regulatory framework - to support capacity building for electricity-sector institutions in Malawi. It highlights key performance expectations for grid-following (GFL) inverters across critical domains such as primary and fast frequency response, ramping behavior, frequency and voltage ride-through capabilities, reactive power and voltage control, ROCOF withstand, phase jump tolerance, system strength considerations, protection schemes, and dynamic modeling requirements. Emphasizing a "Do no harm, do something, do more" philosophy, the document synthesizes best practices and operational benchmarks to guide the integration and stability of inverter-based resources within Malawi's power system. The presentation forms part of the SABESS Center of Excellence initiative supported by the Global Energy Alliance for People and the Planet.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Joint scheduling of energy, fast and primary frequency response reserves in integrated transmission–distribution networks

Inverter-based distributed energy resources (DERs) connected to distribution networks (DNs) can provide fast frequency support, but their reserve deliverability depends on feeder constraints and differs from synchronous primary frequency response (PFR). Existing transmission–distribution coordination studies usually treat reserve generically or neglect feeder-level feasibility, while frequency-security scheduling studies rarely represent distribution feeders explicitly. This paper develops a bi-level day-ahead scheduling framework for integrated transmission–distribution networks that jointly clears energy, transmission-side PFR, and distribution-side fast frequency response (FFR) under exogenous hourly inertia and largest-loss inputs from an external unit commitment (UC) schedule. The transmission problem is modeled with DC-optimal power flow (OPF) and closed-form second-order cone (SOC) frequency-security constraints, whereas each DN is represented by a reserve-aware branch-flow AC-OPF so that scheduled fast reserves remain deliverable during activation. The bi-level problem is reformulated through Karush–Kuhn–Tucker (KKT) conditions into a mixed-integer SOC program, and a penalty term is used to tighten the distribution-network relaxation. In the reduced test system, lower exogenous inertia increased the required primary response from 179.64 MW to 191.08 MW, distribution-side fast response reduced total frequency-response procurement by up to 4.9%, and neglecting distribution constraints overstated the combined distribution-side energy and reserve award by up to 18%. In the expanded study, the largest case was solved in 2.02 s with a 0.00% optimality gap. Time-domain simulations kept the frequency nadir above 59.0 Hz in all tested hours. These results demonstrate the value of fast-response modeling and distribution-feasible reserve delivery in coordinated market clearing.

Noh, Seung-Gil↗

Direct numerical simulations for hybrid rocket boundary layers: Performance modeling and scaling

This paper presents a comprehensive performance and scaling analysis of direct numerical simulations for reacting boundary layers, focusing on slab burner configurations. Using a PETSc-based finite volume CFD framework, the study evaluates the scalability and computational cost of flow, chemistry, and radiation evaluations across 2D and 3D simulations. Polymethyl methacrylate (PMMA) is the fuel with pure O 2 as the oxidizer, modeled using a detailed chemical kinetics mechanism with 113 species and 660 reactions. A ray-tracing-based radiation solver, designed for distributed memory applications, is implemented to model radiation heat transfer. Parallel scalability is analyzed for the coupled flow, chemistry, and radiation heat transfer processes. Weak and strong scaling studies are conducted on up to 15,000 computational ranks, revealing robust performance when flow cells exceed 200 per rank. Chemistry evaluations dominate the computational cost in large 3D simulations, accounting for approximately 40% of the total runtime, while flow processes contribute around 35%, and radiation solver contributions remain below 10% due to reduced evaluation frequencies. GPU accelerated chemistry evaluation, implemented with Zero-RK, demonstrates significant promise, achieving up to a 4x speedup for workloads exceeding 30,000 cells per GPU. However, diminishing returns are observed for smaller workloads due to CPU-GPU communication overhead. This study identifies key challenges, including memory bottlenecks and the effects of domain partitioning on flow scalability, while highlighting the potential of GPU-accelerated chemistry to reduce computational costs. In conclusion, these findings provide realizable run configurations for 2D, 3D, and GPU-accelerated cases, offering insights for optimizing reactive flow solvers.

CFD Scalability↗