Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

Asynchronous domain decomposition methods for nonlinear PDEs

One- and two-level parallel asynchronous methods for the numerical solution of nonlinear systems of equations, especially those arising from (nonlinear) partial differential equations, are studied. The proposed methods are based on domain decomposition techniques. Local convergence theorems are presented in several cases, with appropriate hypotheses. Computational results on a shared memory multiprocessor machine for various problems exhibiting nonlinearities are reported, illustrating the potential of these asynchronous methods, especially for heterogeneous clusters.

97 MATHEMATICS AND COMPUTING↗

Roles of Carburization and Oxidation on Transition-Metal Catalyzed Surface Reactions

Many industrial chemical transformation processes are based on reactions catalyzed by transition metals. In some cases the nano-size particles used as catalysts are synthesized simultaneously or right before the initiation of the main catalytic reaction. Synthesis treatments involve complex processes such as reductions followed by calcinations in oxygenated environments coupled to the presence of gases used for the main reaction. Synthesis details are crucial because they have important consequences on the catalyst shape, size, and chemical composition, and may also affect the structure and composition of the catalyst support. All these effects may alter the main catalyzed reaction in unexpected ways. This work focuses on the synthesis of the overall catalytic system that includes catalyst and support taking place in parallel with the initial stages of the main catalyzed reaction. Two important processes are characterized: carburization and oxidation of the transition metal catalyst during synthesis or initial stages of the main reaction. The objective is to understand how the coking mechanism coupled to surface oxidation and changes in the substrate influence the dry reforming of methane taken as a model reaction. The approach uses first principles computational methods to investigate the catalyst synthesis and characterization and catalytic reaction mechanisms. These studies are complemented with experimental characterization done by the group of Dr. Renu Sharma at the National Institute of Standards and Technology (unfunded collaborator) using high resolution transmission electron microscopy. Specific goals are to identify the role of the synthesis process on the initial nanocatalyst structure and composition, and to determine how that structure and composition may evolve after being exposed to the reactants atmosphere. These studies facilitate a systematic analysis of conditions where the nanoparticle composition can be optimized regarding activity and stability. Moreover, the fundamental study of catalyst/support interactions yields valuable insights for other catalytic processes and is a first step toward controlled nanocatalyst synthesis.

36 MATERIALS SCIENCE↗

Methods and systems for fabricating high quality superconducting tapes

An MOCVD system fabricates high quality superconductor tapes with variable thicknesses. The MOCVD system can include a gas flow chamber between two parallel channels in a housing. A substrate tape is heated and then passed through the MOCVD housing such that the gas flow is perpendicular to the tape's surface. Precursors are injected into the gas flow for deposition on the substrate tape. In this way, superconductor tapes can be fabricated with variable thicknesses, uniform precursor deposition, and high critical current densities.

36 MATERIALS SCIENCE↗

Methods and systems for fabricating high quality superconducting tapes

An MOCVD system fabricates high quality superconductor tapes with variable thicknesses. The MOCVD system can include a gas flow chamber between two parallel channels in a housing. A substrate tape is heated and then passed through the MOCVD housing such that the gas flow is perpendicular to the tape's surface. Precursors are injected into the gas flow for deposition on the substrate tape. In this way, superconductor tapes can be fabricated with variable thicknesses, uniform precursor deposition, and high critical current densities.

36 MATERIALS SCIENCE↗

Methods and systems for fabricating high quality superconducting tapes

An MOCVD system fabricates high quality superconductor tapes with variable thicknesses. The MOCVD system can include a gas flow chamber between two parallel channels in a housing. A substrate tape is heated and then passed through the MOCVD housing such that the gas flow is perpendicular to the tape's surface. Precursors are injected into the gas flow for deposition on the substrate tape. In this way, superconductor tapes can be fabricated with variable thicknesses, uniform precursor deposition, and high critical current densities.

Majkic, Goran↗

Fixing Amdahl's Law within the Limits of Accelerated Systems: FALLACY

Closeout report for FALLACY project. The performance of Data Model Convergence Initiative (DMC) applications on parallel machines is far below the limit set by Amdahl’s law. Whether the machine is based on many-core, GPUs, FPGAs, or a heterogeneous combination, usually the most significant bottleneck is accessing data from the memory system. Aligning with DMC’s HW/architecture thrust, this project developed a set of memory-centric tools called ‘MemGaze’ that inform the HW/SW stack about an application’s memory behavior, including data access latency and diagnosing poor data layout and data composition. Our approach uses architectural modeling and analysis of workload data accesses.

97 MATHEMATICS AND COMPUTING↗

Credit based flow control mechanism for use in multiple link width interconnect systems

Flow control credit management is provided when converting traffic from a first parallel link width on a first link to a second parallel link width on a second link A current value is calculated for a variable flow control credit exchange rate (R) associated with the first and second links. A first flow control credit indicator is received on the second link, and a credit amount calculated based on the first flow control credit indicator and R. A second flow control credit indicator for the credit amount is then transmitted on the first link.

42 ENGINEERING↗

Quantum Overlapping Tomography

It is now experimentally possible to entangle thousands of qubits, and efficiently measure each qubit in parallel in a distinct basis. To fully characterize an unknown entangled state of $\textit{n}$ qubits, one requires an exponential number of measurements in $\textit{n}$, which is experimentally unfeasible even for modest system sizes. By leveraging (i) that single-qubit measurements can be made in parallel, and (ii) the theory of perfect hash families, we show that all $\textit{k}$-qubit reduced density matrices of an $\textit{n}$ qubit state can be determined with at most $e^{\mathcal{O}}(k) \text{log}^2(n)$ rounds of parallel measurements. In this work, we provide concrete measurement protocols which realize this bound. As an example, we argue that with near-term experiments, every two-point correlator in a system of 1024 qubits could be measured and completely characterized in a few days. This corresponds to determining nearly 4.5 million correlators.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Case Study: BESS Capacity Building for Malawi

To strengthen Malawi's national electricity grid and to enhance reliability, the government of Malawi and the Electricity Supply Corporation of Malawi are partnering with the Global Alliance for People and Planet to develop a 20-megawatt/30-megawatt-hour battery energy storage system at Kanengo, Malawi, expected to come online in 2026. In parallel, the Global Energy Alliance for People and Planet and the National Laboratory of the Rockies implemented a comprehensive technical capacity-building program to equip the Electricity Supply Corporation of Malawi, the Malawi Energy Regulatory Authority, and the Malawi University of Business and Applied Sciences with the expertise needed to plan, regulate, and operate advanced storage systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Systems and methods for application performance profiling and improvement

Methods for analyzing and improving a target computer application and corresponding systems and computer-readable mediums. A method includes receiving the target application. The method includes generating a parallel control flow graph (ParCFG) corresponding to the target application. The method includes analyzing the ParCFG by the computer system. The method includes generating and storing the modified ParCFG for the target application.

Leidel, John D.↗

An exploration of online-simulation-driven portfolio scheduling in Workflow Management Systems

Workflow Management Systems used to automate the execution of scientific workflow applications on parallel and distributed computing platforms must make scheduling decisions at runtime. A large number of workflow scheduling algorithms have been proposed in the literature, but often these algorithms are evaluated based on simplifying assumptions that may not hold in practice. Furthermore, published algorithm evaluation and/or comparison results are necessarily only for a subset of all possible scenarios, and thus may not include scenarios relevant to particular use-cases. Consequently, it is difficult for Workflow Management Systems (WMSs) developers to decide which scheduling algorithm should be implemented. To obviate this difficulty, one possible approach is to implement a portfolio of scheduling algorithms and select the most effective algorithm at runtime. One method for performing this selection is to run an online simulation for each algorithm in the portfolio. The algorithm that leads to the best performance, in simulation, is selected for future use. The above simulation-driven portfolio scheduling (SDPS) approach has been proposed in a few parallel and distributed computing contexts. The main objective of this work is to evaluate the feasibility and potential merit of SDPS if implemented in WMSs. Here we perform this evaluation using simulated WMS executions, where the simulations are instantiated from real-world platform and workflow configurations. Our main finding is that SDPS is on par with or outperforms an approach in which a single algorithm is used, where this algorithm is the one that performs best on average across all our experimental scenarios. Furthermore, we find that SDPS remains an attractive proposition even in the presence of high levels of simulation error and for simulators with relatively low levels of sophistication. In many of our experimental scenarios we find that mitigating simulation error at runtime can further improve performance. Finally, we show that simulation overhead can be made sufficiently low for SDPS to be feasible in practice.

97 MATHEMATICS AND COMPUTING↗

Standard model physics and the digital quantum revolution: thoughts about the interface

Advances in isolating, controlling and entangling quantum systems are transforming what was once a curious feature of quantum mechanics into a vehicle for disruptive scientific and technological progress. Pursuing the vision articulated by Feynman, a concerted effort across many areas of research and development is introducing prototypical digital quantum devices into the computing ecosystem available to domain scientists. Through interactions with these early quantum devices, the abstract vision of exploring classically-intractable quantum systems is evolving toward becoming a tangible reality. Beyond catalyzing these technological advances, entanglement is enabling parallel progress as a diagnostic for quantum correlations and as an organizational tool, both guiding improved understanding of quantum many-body systems and quantum field theories defining and emerging from the standard model. Here, from the perspective of three domain science theorists, this article compiles thoughts about the interface on entanglement, complexity, and quantum simulation in an effort to contextualize recent NISQ-era progress with the scientific objectives of nuclear and high-energy physics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Hydrogen Component Leak Rate Quantification for System Risk and Reliability Assessment through QRA and PHM Frameworks: Preprint

The National Renewable Energy Laboratory's (NREL) Hydrogen Safety Research and Development (HSR&D) program in collaboration with the University of Maryland's Systems Risk and Reliability Analysis Laboratory (SyRRA) are working to improve reliability and reduce risk in hydrogen systems. This approach strives to use quantitative data on component leaks and failures, together with Prognosis and Health Management (PHM), and Quantitative Risk Assessment (QRA) to identify at-risk components, reduce component failures and downtime, and predict when components require maintenance. Hydrogen component failures increase facility maintenance cost, facility downtime, and reduce public acceptance of hydrogen technologies, ultimately increasing facility size and cost because of potentially overly conservative requirements. Leaks are a predominant failure mode for hydrogen components. However, uncertainties in the amount of hydrogen emitted from leaking components and the frequency of those failure events limit the understanding of the risks that they present under real-world operational conditions. NREL has deployed a test fixture, the Leak Rate Quantification Apparatus (LRQA), to quantify the mass flow rate of leaking gases from medium and high-pressure components that have failed while in service. Quantitative hydrogen leak rate data from this system could ultimately be used to better inform risk assessment and Regulation Codes and Standards (RCS). Parallel activity explores the use of PHM and QRA techniques to assess and reduce risk, thereby improving safety and reliability of hydrogen systems. The results of QRAs could further provide a systematic and science-based foundation for the design and implementation of RCS, as in the latest versions of the NFPA 2 code for gaseous hydrogen stations. Alternatively, data-driven techniques of PHM could provide new damage diagnosis and health-state prognosis tools. This research will help end users, station owners and operators, and regulatory bodies move towards risk-informed preventative maintenance versus emergency corrective maintenance, reducing cost and improving reliability. Predictive modelling of failures could improve safety and affect RCS requirements such as setback distances at liquid refuelling sites. The combination of leak rate quantification research, PHM, and QRA can lead to better informed models enabling data-based decision to be made for hydrogen system safety improvements.

codes and standards↗

Interface microstructure effects on dynamic failure behavior of layered Cu/Ta microstructures

Abstract Structural metallic materials with interfaces of immiscible materials provide opportunities to design and tailor the microstructures for desired mechanical behavior. Metallic microstructures with plasticity contributors of the FCC and BCC phases show significant promise for damage-tolerant applications due to their enhanced strengths and thermal stability. A fundamental understanding of the dynamic failure behavior is needed to design and tailor these microstructures with desired mechanical responses under extreme environments. This study uses molecular dynamics (MD) simulations to characterize plasticity contributors for various interface microstructures and the damage evolution behavior of FCC/BCC laminate microstructures. This study uses six model Cu/Ta interface systems with different orientation relationships that are as- created, and pre-deformed to understand the modifications in the plasticity contributions and the void nucleation/evolution behavior. The results suggest that pre-existing misfit dislocations and loading orientations (perpendicular to and parallel to the interface) affect the activation of primary and secondary slip systems. The dynamic strengths are observed to correlate with the energy of the interfaces, with the strengths being highest for low-energy interfaces and lowest for high-energy interfaces. However, the presence of pre-deformation of these interface microstructures affects not only the dynamic strength of the microstructures but also the correlation with interface energy.

42 ENGINEERING↗

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations. To make advancements in all the aforementioned scientific areas, many technical aspects need to be more thoroughly studied and deeply understood in parallel discrete event simulation (PDES). The unique dynamics inherent in a discrete event modeling approach, by their very nature, intersect and influence the entire stack of the computing system, including (a) the unique nature of the instruction sets exercised in PDES workloads without a predominance of high-precision floating point operations, (b) virtual time-constrained multi-threaded execution of many logical processes per processor, (c) extremely variable and difficult to predict network traffic characteristics, (d) interfaces and inter-dependencies with machine learning and artificial intelligence codes at higher software layers, and (e) highly challenging load balancing needs, especially in effectively accounting for accelerated/extremely heterogeneous computing in current and future high-performance computing systems. Efficient and accurate parallel execution of PDES workloads is also dominated by challenges in dealing with their asynchronous concurrency fundamentally present at the model level. Conservative synchronization, optimistic/speculative synchronization, and their hybrid schemes open new questions in fundamental computer science with respect to reversibility of computation and prediction (lookahead) of behaviors inherent within model codes. On the implementation front, there are relatively few scalable, general-purpose parallel discrete event simulators in the world, and even fewer have been studied on emerging hardware platforms. To enable scientific advances using PDES, the research needs in computer science must also be pursued and met in the intersection of the algorithmic and hardware-aware aspects of scalable PDES engines. This report is aimed at capturing a computer science-oriented view of this important area of research in PDES, presenting a sample of important applications with their inherent discrete event technology elements. Needs are outlined in core areas of parallel discrete event research as well as cross-cutting directions in computer science research that positively impact scientific advancements across several important application areas. A selection of priority research opportunities in advanced computing for PDES is identified to serve as reference for key research topics and their order of importance for scientific advancements.

97 MATHEMATICS AND COMPUTING↗

Metaplastic and energy-efficient biocompatible graphene artificial synaptic transistors for enhanced accuracy neuromorphic computing

CMOS-based computing systems that employ the von Neumann architecture are relatively limited when it comes to parallel data storage and processing. In contrast, the human brain is a living computational signal processing unit that operates with extreme parallelism and energy efficiency. Although numerous neuromorphic electronic devices have emerged in the last decade, most of them are rigid or contain materials that are toxic to biological systems. In this work, we report on biocompatible bilayer graphene-based artificial synaptic transistors (BLAST) capable of mimicking synaptic behavior. The BLAST devices leverage a dry ion-selective membrane, enabling long-term potentiation, with ~50 aJ/µm 2 switching energy efficiency, at least an order of magnitude lower than previous reports on two-dimensional material-based artificial synapses. The devices show unique metaplasticity, a useful feature for generalizable deep neural networks, and we demonstrate that metaplastic BLASTs outperform ideal linear synapses in classic image classification tasks. With switching energy well below the 1 fJ energy estimated per biological synapse, the proposed devices are powerful candidates for bio-interfaced online learning, bridging the gap between artificial and biological neural networks.

97 MATHEMATICS AND COMPUTING↗

Chip-scale frequency combs for data communications in computing systems

Recent developments in chip-based frequency-comb technology demonstrate that comb devices can be implemented in applications where photonic integration and power efficiency are required. The large number of equally spaced comb lines that are generated make combs ideal for use in communication systems, where each line can serve as an optical carrier to allow for massively parallel wavelength-division multiplexing (WDM) transmission. In this review, we summarize the developments in integrated frequency-comb technology for use as a WDM source for communication systems in data centers and high-performance computing systems. We highlight the following three approaches for chip-scale comb generation: semiconductor modelocked lasers, electro-optic combs, and Kerr frequency combs.

Okawachi, Yoshitomo (ORCID:0000000196390549)↗