Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

The Thermal Plumbing System of Stromboli Volcano, Aeolian Islands (Italy) Inferred From Electrical Conductivity and Induced Polarization Tomography

Abstract We performed the first 3D island‐scale tomography of the electrical conductivity of Stromboli volcano (Aeolian Islands, Italy) using 2D acquisition lines (37.2 km) and a total of 18,880 measurements and 2,402 unique electrode locations. This 3D data set was inverted using a Gauss‐Newton algorithm, parallel‐processing on an unstructured tetrahedral mesh containing 678,420 finite‐element nodes and 3,580,145 elements to account for the topography of the volcanic island. The tomogram exhibits a conductive body (10 −2 –1.0 S m −1 ) consistent with the location of CO 2 and temperature anomalies observed at the ground surface. It corresponds to the hydrothermal system with high electrical conductivity associated with alteration. In order to confirm this interpretation, a 2.5D large‐scale induced polarization tomography was performed crossing the volcano. The joint interpretation of the conductivity and normalized chargeability is done with a petrophysical model previously tested and verified at both shield‐ and strato‐volcanoes. This model implies that alteration (through the effect of the cation exchange capacity associated with clay minerals and zeolites) plays a strong role in both controlling the electrical conductivity and normalized chargeability at Stromboli volcano. A temperature tomogram, derived from the geoelectrical measurements, is consistent with surface temperature anomalies and the Very Long Period (VLP) seismicity related to the mild‐explosive activity. This survey displays at 600 m a.s.l. a lateral shift in the highest temperature location, also corresponding to the source of VLP seismicity. Structural boundaries have a major role in the hottest hydrothermal fluids rising below the active crater terrace of Stromboli volcano.

58 GEOSCIENCES↗

Broadband unidirectional visible imaging using wafer-scale nano-fabrication of multi-layer diffractive optical processors

We present a broadband and polarization-insensitive unidirectional imager that operates at the visible part of the spectrum, where image formation occurs in one direction, while in the opposite direction, it is blocked. This approach is enabled by deep learning-driven diffractive optical design with wafer-scale nano-fabrication using high-purity fused silica to ensure optical transparency and thermal stability. Our design achieves unidirectional imaging across three visible wavelengths (covering red, green, and blue parts of the spectrum), and we experimentally validated this broadband unidirectional imager by creating high-fidelity images in the forward direction and generating weak, distorted output patterns in the backward direction, in alignment with our numerical simulations. This work demonstrates wafer-scale production of diffractive optical processors, featuring 16 levels of nanoscale phase features distributed across two axially aligned diffractive layers for visible unidirectional imaging. This approach facilitates mass-scale production of ~0.5 billion nanoscale phase features per wafer, supporting high-throughput manufacturing of hundreds to thousands of multi-layer diffractive processors suitable for large apertures and parallel processing of multiple tasks. Beyond broadband unidirectional imaging in the visible spectrum, this study establishes a pathway for artificial-intelligence-enabled diffractive optics with versatile applications, signaling a new era in optical device functionality with industrial-level, massively scalable fabrication.

36 MATERIALS SCIENCE↗

Reconstruction framework advancements to support streaming for the ePIC detector at the EIC

The ePIC collaboration adopted the JANA2 framework to manage its reconstruction algorithms. This framework has since evolved substantially in response to ePIC’s needs. There have been three main design drivers: integrating cleanly with the Podio-based data models and other layers of the key4hep stack, enabling external configuration of existing components, and supporting timeframe splitting for streaming readout. The result is a unified component model featuring a new declarative interface for specifying inputs, outputs, parameters, services, and resources. This interface enables the user to instantiate, configure, and wire components via an external file. One critical new addition to the component model is a hierarchical decomposition of data boundaries into levels such as Run, Timeframe, PhysicsEvent, and Subevent. Two new component abstractions, Folder and Unfolder, are introduced in order to traverse this hierarchy, e.g. by splitting or merging. The pre-existing components can now operate at different event levels, and JANA2 will automatically construct the corresponding parallel processing topology. This means that a user may write an algorithm once, and configure it at runtime to operate on timeframes or on physics events. Overall, these changes mean that the user requires less knowledge about the framework internals, obtains greater flexibility with configuration, and gains the ability to reuse the existing abstractions in new streaming contexts.

Brei, Nathan [Thomas Jefferson National Accelerato↗

A New Integrated Analysis Suite for Fast-Ion Study in KSTAR

Here, an integrated workflow for fast-ion analysis was developed by adapting the One Modeling Framework for Integrated Task (OMFIT) workflow manager to support a standard and unified analysis platform for KSTAR users. The newly established analysis suite offers a graphical user interface–based workflow to enable users to readily access and handle experimental data archived in various data formats and servers. Further, users can analyze the data by importing modules designed for conducting certain tasks, such as profile fitting, equilibrium reconstruction, and postprocessing of tokamak data. The procedures for preparing the inputs for fast-ion simulations are streamlined by a common workflow manager, which enables the parallel processing of various tasks to efficiently analyze large fast-ion datasets. The OMFIT platform comprises a flexible Python-based application that enables users to freely manipulate the Python scripts for applications that are unavailable in the standard workflow. The framework also offers mapping tools to translate the output data into the Integrated Modeling and Analysis Suite format to maintain application compatibility for future ITER burning plasma experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Scaling Ultrahigh-Resolution E3SM Land Model for Leadership-Class Supercomputers

This paper presents advancements in scaling the ultrahigh-resolution E3SM Land Model (uELM) for deployment on leadership-class supercomputers, addressing the increased demand for km-scale Earth system modeling. By focusing on km-scale ELM simulations, we enhance predictive capabilities for climate interactions, facilitating improved responses to climate change impacts on energy systems, agriculture, and water resources. Our approach leverages innovative software architecture optimizations, sophisticated data handling techniques, and advanced parallel processing, achieving strong scalability on two leadership supercomputers (2400 nodes (105,600 cores) on Summit, and 1200 nodes (76,800 cores) on Frontier). Results from extensive scalability assessments on the Summit and Frontier also demonstrate outstanding I/O performance (close to 400 GB/s write throughput) and the model's ability to efficiently handle increasing computational demands. This study not only establishes uELM's capability for high-resolution simulations over vast geographical domains, but also sets a foundation for future Earth system modeling breakthroughs.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

Really Embedding Domain-Specific Languages into C++

The following topics are dealt with: program compilers; optimising compilers; parallel processing; software engineering; learning (artificial intelligence); multiprocessing systems; shared memory systems; optimisation; computational complexity; specification languages.

Finkel, Hal J.↗

AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload Rebalancing

The recent development of deep learning has been mostly focusing on Euclidean data, such as images, videos, audios, etc. However, most real-world information and relation are often expressed as graphs. To efficiently learn from graph data, graph convolutional networks (GCNs) emerge as a promising approach, showing advantages in several practical applications such as social network analysis, knowledge discovery, 3D modeling, motion capturing, etc. Real-world graphs are usually extremely large and imbalanced, posting significant performance demand and design challenges on the hardware dedicated for GCN inference. In this paper, we propose an architecture design called UW-GCN to accelerate graph convolutional network inference. To tackle the major performance bottleneck from workload imbalance, we propose dynamic neighborhood stealing and remote chunk shuffling techniques, relying on hardware flexibility to achieve hardware auto-tuning under negligible area or delay overhead. Specifically, UW-GCN is able to smartly profile the sparse graph pattern while continuously adjusting the workload distribution via routing reconfiguration among parallel processing elements (PEs). The ideal configuration is then reused in the remaining iterations. To the best of our knowledge, this is the first accelerator design particularly for GCN and the first work relying on hardware auto-tuning, which is normally based on software, to achieve near-optimal workload balance in processing sparse structures.

Geng, Tong↗

Lessons from α Dragon Fly's Brain: Evolution Built a Small, Fast, Efficient Neural Network in a Dragonfly. Why Not Copy It for Missile Defense?

In each of our brains, 86 billion neurons work in parallel, processing inputs from senses and memories to produce the many feats of human cognition. The brains of other creatures are less broadly capable, but those animals often exhibit innate aptitudes for particular tasks, abilities honed by millions of years of evolution.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Safe Reinforcement Learning-Based Transient Stability Control for Islanded Microgrids With Topology Reconfiguration

This paper proposes a safe reinforcement learning (RL)-based transient stability emergency control (TSEC) method for islanded microgrids. RL requires extensive interaction with the environment to learn control strategies, hence, a data-driven approach is used as a substitute for time-consuming time-domain simulation calculations. Deep sigma point processes (DSPP), which is a Gaussian process model, is utilized to predict the normal distribution of transient stability of microgrids and to construct a transient stability chance constraint. Reward-constrained policy optimization (RCPO) can simultaneously achieve objective prediction, policy learning, and constraint cost coefficient update across multiple timescales. RCPO interacts with the DSPP-based microgrid environment through a multi-process parallel manner, greatly increasing the training speed. Case studies on a real islanded microgrid demonstrate that the proposed method can efficiently and quickly obtain the optimal emergency control strategy while adhering to all hard constraints.

14 SOLAR ENERGY↗

Library for Evolutionary Algorithms in Python (LEAP)

There are generally three types of scientific software users: users that solve problems using existing science software tools, researchers that explore new approaches by extending existing code, and educators that teach students scientific concepts. Python is a general-purpose programming language that is accessible to beginners, such as students, but also as a language that has a rich scientific programming ecosystem that facilitates writing research software. Additionally, as high-performance computing (HPC) resources become more readily available, software support for parallel processing becomes more relevant to scientific software.There currently are no Python-based evolutionary computation frameworks that support all three types of scientific software users. Moreover, some support synchronous concurrent fitness evaluation that do not efficiently use HPC resources. We pose here a new Python-based EC framework that uses an established generalized unified approach to EA concepts to provide an easy to use toolkit for users wishing to use an EA to solve a problem, for researchers to implement novel approaches, and for providing a low-bar to entry to EA concepts for students. Additionally, this toolkit provides a scalable asynchronous fitness evaluation implementation friendly to HPC that has been vetted on hardware ranging from laptops to the world’s fastest supercomputer, Summit.

Coletti, Mark↗

Axolotl: a scalable genomics library based on Apache Spark (Axolotl) v1.0.0

Axolotl is a Python library for scalable distributed genome and metagenome data analysis. Existing tools and systems that we rely on are struggling to keep up with the rapid explosion of genomic data. Compounding this issue, developing scalable solutions require a steep learning curve in parallel programming, which presents a barrier to academic researchers. While we do have scalable solutions for specific tasks, we lack comprehensive, end-to-end solutions. It's this gap in our toolkit that we aim to address with Axolotl. The Axolotl library is built for easy parallel processing, efficiently handling multiple tasks or large datasets simultaneously, and scaling up to meet the demands of extensive genomic data analysis.

Wang, Zhong↗

Control And Optimization Modular Modeling Application For Nuclear Deployment

The purpose of the COMMAND code is to provide a flexible, scalable tool for use in developing, integrating, and testing the technologies necessary for achieving autonomous operations of advanced nuclear reactors. The code enables users to efficiently implement custom simulations and experiments by combining key methods from different software modules. These modules are focused on: modeling and simulation tools, such as nuclear simulation tools used for high-fidelity modeling (e.g., Reactor Excursion and Leak Analysis Program [RELAP5-3D] and Monte Carlo N-Particle [MCNP]); machine learning and optimization tools (e.g., anomaly detection and data-driven modeling techniques); advanced control in its digital, high-performance, and supervisory control forms (e.g., proportional integral derivative (PID) control and model predictive control (MPC); and integration with hardware through industrial communication protocols. To ensure flexibility and scalability, COMMAND was designed to be both modular—the software “pieces” all inherit from generic building blocks and can be combined and connected to create complicated simulations—and high performing—designed for parallel processing, enabling simulations and experiments to take advantage of multi-core computers, servers, and nodes. The code is written in the Python programming language due to the language's popularity, active community, and open-source and cross-platform nature. Maintaining consistency with other simulation tools used within the nuclear energy community, users implement simulations and experiments through text input files, which define components, parameters, connections, etc., through lines of text. Given that COMMAND is written in Python, these input files are native Python scripts, and so use the standard Python structure and formatting. This also enables users to take advantage of Python's extensive package library to develop custom capabilities for their specific use cases.

Faber, Jacob [Idaho National Laboratory (INL), Ida↗

Adorym: a multi-platform generic X-ray image reconstruction framework based on automatic differentiation

We describe and demonstrate an optimization-based X-ray image reconstruction framework called Adorym. Our framework provides a generic forward model, allowing one code framework to be used for a wide range of imaging methods ranging from near-field holography to fly-scan ptychographic tomography. By using automatic differentiation for optimization, Adorym has the flexibility to refine experimental parameters including probe positions, multiple hologram alignment, and object tilts. It is written with strong support for parallel processing, allowing large datasets to be processed on high-performance computing systems. We demonstrate its use on several experimental datasets to show improved image quality through parameter refinement.

36 MATERIALS SCIENCE↗

Massively Parallel Capability in Sierra/SD for Simulation Vibration with Piezoelectrics

Sierra/SD is an engineering structural dynamics code that provides Sandia and other customers a tool to model structural and acoustic physics on large complex physical systems using massively parallel processing. This report provides a detailed overview on Sierra/SD’s most recent physics package: coupled electro-mechanical physics. This capability uses the finite element method to model coupled electro-mechanical physics exhibited by piezoelectric materials. This report provides an applications overview, theory overview, and verification examples demonstrating the electro-mechanical physics modeling capabilities of Sierra/SD.

97 MATHEMATICS AND COMPUTING↗

Rapid QSTS Simulations for High-Resolution Comprehensive Assessment of Distributed PV

The rapid increase in penetration of distributed energy resources on the electric power distribution system has created a need for more comprehensive interconnection modeling and impact analysis. Unlike conventional scenario-based studies, quasi-static time-series (QSTS) simulations can realistically model time-dependent voltage controllers and the diversity of potential impacts that can occur at different times of year. However, to accurately model a distribution system with all its controllable devices, a yearlong simulation at 1-second resolution is often required, which could take conventional computers a computational time of 10 to 120 hours when an actual unbalanced distribution feeder is modeled. This computational burden is a clear limitation to the adoption of QSTS simulations in interconnection studies and for determining optimal control solutions for utility operations. The solutions we developed include accurate and computationally efficient QSTS methods that could be implemented in existing open-source and commercial software used by utilities and the development of methods to create high-resolution proxy data sets. This project demonstrated multiple pathways for speeding up the QSTS computation using new and innovative methods for advanced time-series analysis, faster power flow solvers, parallel processing of power flow solutions and circuit reduction. The target performance level for this project was achieved with year-long high-resolution time series solutions run in less than 5 minutes within an acceptable error.

24 POWER TRANSMISSION AND DISTRIBUTION↗

QA4, a language for artificial intelligence.

Introduction of a language for problem solving and specifically robot planning, program verification, and synthesis and theorem proving. This language, called question-answerer 4 (QA4), embodies many features that have been found useful for constructing problem solvers but have to be programmed explicitly by the user of a conventional language. The most important features of QA4 are described, and examples are provided for most of the material introduced. Language features include backtracking, parallel processing, pattern matching, set manipulation, and pattern-triggered function activation. The language is most convenient for use in an interactive way and has extensive trace and edit facilities.

Derksen, J. A. C.↗