Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “synchronized sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

62 records · Page 4

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations. To make advancements in all the aforementioned scientific areas, many technical aspects need to be more thoroughly studied and deeply understood in parallel discrete event simulation (PDES). The unique dynamics inherent in a discrete event modeling approach, by their very nature, intersect and influence the entire stack of the computing system, including (a) the unique nature of the instruction sets exercised in PDES workloads without a predominance of high-precision floating point operations, (b) virtual time-constrained multi-threaded execution of many logical processes per processor, (c) extremely variable and difficult to predict network traffic characteristics, (d) interfaces and inter-dependencies with machine learning and artificial intelligence codes at higher software layers, and (e) highly challenging load balancing needs, especially in effectively accounting for accelerated/extremely heterogeneous computing in current and future high-performance computing systems. Efficient and accurate parallel execution of PDES workloads is also dominated by challenges in dealing with their asynchronous concurrency fundamentally present at the model level. Conservative synchronization, optimistic/speculative synchronization, and their hybrid schemes open new questions in fundamental computer science with respect to reversibility of computation and prediction (lookahead) of behaviors inherent within model codes. On the implementation front, there are relatively few scalable, general-purpose parallel discrete event simulators in the world, and even fewer have been studied on emerging hardware platforms. To enable scientific advances using PDES, the research needs in computer science must also be pursued and met in the intersection of the algorithmic and hardware-aware aspects of scalable PDES engines. This report is aimed at capturing a computer science-oriented view of this important area of research in PDES, presenting a sample of important applications with their inherent discrete event technology elements. Needs are outlined in core areas of parallel discrete event research as well as cross-cutting directions in computer science research that positively impact scientific advancements across several important application areas. A selection of priority research opportunities in advanced computing for PDES is identified to serve as reference for key research topics and their order of importance for scientific advancements.

97 MATHEMATICS AND COMPUTING↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Using scalable computer vision to automate high-throughput semiconductor characterization

Abstract High-throughput materials synthesis methods, crucial for discovering novel functional materials, face a bottleneck in property characterization. These high-throughput synthesis tools produce 10 4 samples per hour using ink-based deposition while most characterization methods are either slow (conventional rates of 10 1 samples per hour) or rigid (e.g., designed for standard thin films), resulting in a bottleneck. To address this, we propose automated characterization (autocharacterization) tools that leverage adaptive computer vision for an 85x faster throughput compared to non-automated workflows. Our tools include a generalizable composition mapping tool and two scalable autocharacterization algorithms that: (1) autonomously compute the band gaps of 200 compositions in 6 minutes, and (2) autonomously compute the environmental stability of 200 compositions in 20 minutes, achieving 98.5% and 96.9% accuracy, respectively, when benchmarked against domain expert manual evaluation. These tools, demonstrated on the formamidinium (FA) and methylammonium (MA) mixed-cation perovskite system FA 1−x MA x PbI 3 , 0 ≤ x ≤ 1, significantly accelerate the characterization process, synchronizing it closer to the rate of high-throughput synthesis.

Science & Technology - Other Topics↗

From BEYONDPLANCK to COSMOGLOBE: Preliminary WMAP Q -band analysis

We present the first application of the COSMOGLOBE analysis framework by analyzing nine-year WMAP time-ordered observations that uses similar machinery to that of BEYONDPLANCK for the Planck Low Frequency Instrument (LFI). We analyzed only the Q-band (41 GHz) data and report on the low-level analysis process based on uncalibrated time-ordered data to calibrated maps. Most of the existing BEYONDPLANCK pipeline may be reused for WMAP analysis with minimal changes to the existing codebase. The main modification is the implementation of the same preconditioned biconjugate gradient mapmaker used by the WMAP team. Producing a single WMAP Q1-band sample requires 22 CPU-hrs, which is slightly more than the cost of a Planck 44 GHz sample of 17 CPU-hrs; this demonstrates that a full end-to-end Bayesian processing of the WMAP data is computationally feasible. In general, our recovered maps are very similar to the maps released by the WMAP team, although with two notable differences. In terms of temperature, we find a ~2 μK quadrupole difference that most likely is caused by different gain modeling, while in polarization we find a distinct 2.5 μK signal that has been previously referred to as poorly measured modes by the WMAP team. In the COSMOGLOBE processing, this pattern arises from temperature-to-polarization leakage from the coupling between the CMB Solar dipole, transmission imbalance, and sidelobes. No traces of this pattern are found in either the frequency map or TOD residual map, suggesting that the current processing has succeeded in modeling these poorly measured modes within the assumed parametric model by using Planck information to break the sky-synchronous degeneracies inherent in the WMAP scanning strategy.

79 ASTRONOMY AND ASTROPHYSICS↗

Covariation of Slab Tracers, Volatiles, and Oxidation During Subduction Initiation

Abstract Subduction‐related lavas have higher Fe 3+ /∑Fe than midocean ridge basalts (MORB). Hypotheses for this offset include imprint from subducting slabs and differentiation in thickened crust. These ideas are readily tested through examination of the time‐dependent evolution of slab‐derived signatures, thickening crust of the overriding plate, and evolving redox during subduction initiation. Here, we present Fe 3+ /ΣFe and volatile element abundances of volcanic glasses recovered from International Ocean Discovery Program (IODP) Expedition 352 to the Izu‐Bonin‐Mariana (IBM) forearc. The samples include forearc basalts (FAB) that are stratigraphically overlain by low‐ and high‐silica boninite lavas. The FAB glasses have 0.18–0.85 wt% H 2 O, 75–233 ppm CO 2 , S contents controlled by saturation with a sulfide phase (602–1,386 ppm), Ba/La from 3.9‐10, and Fe 3+ /ΣFe ratios from 0.136 to 0.177. These compositions are similar to MORB and suggest that decompression melting of dry and reduced mantle dominates the earliest stages of subduction initiation. Low‐ and high‐silica boninite glasses have 1.51–3.19 wt% H 2 O, CO 2 below detection, S contents below those required for sulfide saturation (5–235 ppm), Ba/La from 11 to 29, and Fe 3+ /∑Fe from 0.181 to 0.225. The compositions are broadly similar to modern arc lavas in the IBM arc. These data demonstrate that the establishment of fluid‐fluxed melting of the mantle, which occurs in just 0.6–1.2 my after subduction initiation, is synchronous with the production of oxidized, mantle‐derived magmas. The coherence of high Fe 3+ /∑Fe and Ba/La ratios with high H 2 O contents in Expedition 352 glasses and the modern IBM arc rocks strongly links the production of oxidized arc magmas to signatures of slab dehydration.

Brounce, Maryjo↗

Toward an AI-Powered Software Pipeline for Real-Time Tracking and Analysis of Wildfire and Smoke

Real-time tracking of wildfires and smoke is crucial for effective response, minimizing damage, protecting lives, and efficiently managing resources during fire emergencies. We develop a web-based AI-powered pipeline that detects wildfires in aerial video and estimates deployment-relevant behavior metrics, including cumulative burned area, burned-area growth rate, fire spread direction, and smoke dispersion. The system combines a YOLO-based detector with YCbCr-based fire segmentation, HSV-based smoke segmentation, Farneback optical flow, and centroid-based spatiotemporal tracking. Using ground sampling distance (GSD), pixel-level fire masks are converted to physical burned-area measurements by correlating fire pixel counts with camera altitude and tilt angle. We benchmark YOLO variants and non-YOLO baselines (GoogLeNet, CNN, DBN, Autoencoder, U-Net, and AlexNet) on the IEEE FLAME dataset and a newly created aerial frame dataset, Wildfire-DB. Cross-dataset evaluation uses a strict threshold-transfer protocol: decision thresholds are selected on FLAME validation and transferred unchanged to Wildfire-DB to quantify generalization under domain shift. YOLOv6 achieves the strongest cross-dataset frame-level fire detection on Wildfire-DB (ROC-AUC 0.8200, PR-AUC 0.8044, and transferred-threshold F1 0.7596). For tracking-oriented deployment requiring oriented localization, YOLO11-OBB provides the most reliable cross-dataset behavior among OBB-capable models while remaining computationally feasible. To analyze the feasibility of UAV deployment, we further measure inference efficiency using synchronized GPU and CPU power logs on a fixed workload of 1569 frames. YOLO-family models process the video in 5.73–12.47 seconds with net energy of 1247.28–1775.39 J, substantially lower latency and energy than heavier classification and reconstruction baselines. Overall, model optimality depends on operational objectives: YOLOv6 is best for cross-dataset detection robustness, whereas YOL...

Color segmentation↗

Detection of surface water temperature variations of Mongolian lakes benefiting from the spatially and temporally gap-filled MODIS data

Lakes provide critical water resources for human activities and ecosystems, particularly in the Mongolian Plateau (MP), which is characterized by a dry climate and a harsh environment. As a region that is sensitive to anthropogenic warming, tracking lake surface water temperature (LSWT) changes in Mongolian lakes is crucial for understanding the consequences of a warming climate on lake ecosystems. However, the long-term monitoring of LSWT is restricted by the spatiotemporal gaps in the raw imagery of remote sensing-based land surface temperature (LST), e.g., the commonly used Moderate Resolution Imaging Spectroradiometer (MODIS) LST products. This study applied an improved gap-filling method by utilizing the discrete cosine transform-based penalized least squares (DCT-PLS) strategy in the spatial domain combined with the linear interpolation (LI) algorithm in the temporal domain. The method was applied to fill gaps in the LSWT imagery of 12 representative lakes across MP. The randomly sampled high-quality MODIS LSWT values in the spatial and temporal domains were excavated as false data gaps and considered “virtual true” validation datasets. The spatial validation results showed that the estimated LSWT for all the lake cases were comparable with the “virtual true” LSWT values, with the average values of the coefficient of determination, mean absolute error, mean square error, and root mean square error being 0.98, 0.38 °C, 0.45 °C, and 0.59 °C, respectively. Meanwhile, the error of nighttime LSWT results was relatively lower than that of daytime LSWT. For temporal interpolation validation, the LI algorithm exhibited relatively better performance and could more objectively indicate the variation in LSWT. Benefiting from the spatially and temporally well-constrained data, we analyzed the interannual and intra-annual change characteristics of the LSWTs of the 12 lakes. The long-term variations of annual and seasonal mean LSWTs in the 12 selected lakes exhibited no evident trends in 2000–2020, while presented apparent interannual fluctuations. The slight changes in the average LSWTs of the 12 selected lakes were in excellent synchronization with the surrounding LST derived from the reanalysis datasets, confirming the widely reported phenomenon of “global warming hiatus” that occurred in the early 21st century. This study improves the understanding of the LSWT variations in Mongolian lakes in response to global climate change. It has the potential to provide an effective approach for monitoring LSWT changes in other large-scale studies.

54 ENVIRONMENTAL SCIENCES↗

Digital Grid Twin–Direct Communication Scheme Test Bed for Assessing Relay-to-Relay Radio Antenna and Optical Fiber Performance and Misoperations

This study introduces a novel “Digital Grid Twin–Direct Communication Scheme” test bed. This advanced platform evaluates point-to-point communication between transmitter and receiver relays with optical fiber and radio omnidirectional antenna systems, implemented at the Advanced Protection lab in the Grid Research Innovation and Development Center at Oak Ridge National Laboratory. The increased diversity of energy sources has led to more protective relay misoperations. In North America, microgrid protection schemes now use point-to-point communication along distribution lines between relays to implement advanced logic in nonradial grids that include both high- and low-inertia generators. This trend challenges utilities to minimize misoperations while ensuring rapid fault clearance and accurate selectivity coordination between primary and backup relays. This study assesses relay-to-relay communication schemes by introducing an advanced testing platform based on a digital grid twin protection test bed using a synchronized time source system. The platform evaluates the communication system using radio antennas or optical fiber links by integrating protective relays that operate breakers within the digital twin and record relay events and communication signals. In the experiments, transmitter and receiver relays were configured with inverse time overcurrent and breaker trip detection logic to assess the total time of the communication protection schemes based on the sum of the relay protection element operating time, radio latency, propagation delay, baud rate delay, and relay processing time. These delays were derived from recorded relay events and communication signals from the interface of a real-time simulator set as a digital grid twin. The test bed successfully simulated various electrical faults along a distribution line while ensuring effective and reliable point-to-point communication between transmitter and receiver relays. The radio antenna communication system exhibited latency because of the radio. This latency depends on the baud rate setting and type of radio application; in general, the higher the baud rate, the lower the radio latency. The measured radio latency (for Mirrored Bits with an encryption card at 9,600 bps) was about 9–10 ms. Additionally, calculated propagation delay per mile for radio antennas and optical fiber was 5.36 µs/mi and 8.04 µs/mi, respectively. Optical fiber communication did not demonstrate radio latency. Instead, the protection element operating time depends mainly on the protection logic function set in the relay, and the relay processing time depends on the processing rate of the relay in samples per power system cycle.

24 POWER TRANSMISSION AND DISTRIBUTION↗