Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data transfer pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

100 records · Page 6

The Simons Observatory: validation of reconstructed power spectra from simulated filtered maps for the small aperture telescope survey

We present a transfer function-based method to estimate angular power spectra from filtered maps for cosmic microwave background (CMB) surveys. This is especially relevant for experiments targeting the faint primordial gravitational wave signatures in CMB polarisation at large scales, such as the Simons Observatory (SO) small aperture telescopes. While timestreams can be filtered to mitigate the contamination from low-frequency noise, usual methods that calculate the mode coupling at individual multipoles can be challenging for experiments covering large sky areas or reaching few-arcminute resolution. The method we present here, although approximate, is more practical and faster for larger data volumes. We validate it through the use of simulated observations approximating the first year of SO data, going from half-wave plate-modulated timestreams to maps, and using simulations to estimate the mixing of polarisation modes induced by an example of time-domain filtering. We show its performance through an example null test and with an end-to-end pipeline that performs inference on cosmological parameters, including the tensor-to-scalar ratio r. The performance demonstration uses simulated observations at multiple frequency bands. We find that the method can recover unbiased parameters for our simulated noise levels.

CMBR experiments↗

PATOKA: Simulating Electromagnetic Observables of Black Hole Accretion

The Event Horizon Telescope (EHT) has released analyses of reconstructed images of horizon-scale millimeter emission near the supermassive black hole at the center of the M87 galaxy. Parts of the analyses made use of a large library of synthetic black hole images and spectra, which were produced using numerical general relativistic magnetohydrodynamics fluid simulations and polarized ray tracing. In this article, we describe the PATOKA pipeline, which was used to generate the Illinois contribution to the EHT simulation library. We begin by describing the relevant accretion systems and radiative processes. We then describe the details of the three numerical codes we use, iharm, ipole, and igrmonty, paying particular attention to differences between the current generation of the codes and the originally published versions. Finally, we provide a brief overview of simulated data as produced by PATOKA and conclude with a discussion of limitations and future directions.

supermassive black holes↗

Baryon fraction from the BAO amplitude: a consistent approach to parameterizing perturbation growth

Galaxy clustering constrains the baryon fraction Omega_b/Omega_m through the amplitude of baryon acoustic oscillations and the suppression of perturbations entering the horizon before recombination. This produces a different pre-recombination distribution of baryons and dark matter. After recombination, the gravitational potential responds to both components in proportion to their mass, allowing robust measurement of the baryon fraction. This is independent of new-physics scenarios altering the recombination background (e.g. Early Dark Energy). The accuracy of such measurements does, however, depend on how baryons and CDM are modeled in the power spectrum. Previous template-based splitting relied on approximate transfer functions that neglected part of information. We present a new method that embeds an extra parameter controlling the balance between baryons and dark matter in the growth terms of the perturbation equations in the CAMB Boltzmann solver. This approach captures the baryonic suppression of CDM prior to recombination, avoids inconsistencies, and yields a clean parametrization of the baryon fraction in the linear power spectrum, separating out the simple physics of growth due to the combined matter potential. We implement this framework in an analysis pipeline using Effective Field Theory of Large-Scale Structure with HOD-informed priors and validate it against noiseless LCDM and EDE cosmologies with DESI-like errors. The new scheme achieves comparable precision to previous splitting while reducing systematic biases, providing a more robust way to baryon-fraction measurements. In combination with BBN constraints on the baryon density and Alcock-Paczynski estimates of the matter density, these results strengthen the use of baryon fraction measurements to derive a Hubble constant from energy densities, with future DESI and Euclid data expected to deliver competitive constraints.

Crespi, Andrea [U. Waterloo (main); Waterloo U., I↗

Leveraging generative artificial intelligence to bridge domain gaps in wind turbine research

A central challenge in wind turbine health monitoring is the scarcity of real-world data due to limited instrumentation, leading researchers to rely on simulation models that often suffer from reduced fidelity. However, even within simulation environments, discrepancies arise because of modeling assumptions, and configuration fidelities, creating domain gaps that limit the transferability of learned representations. Here, to investigate domain translation under controlled conditions, this project explores the use of generative artificial intelligence, specifically cycle-consistent generative adversarial networks (CGANs), to bridge the gap between OpenFAST simulation models representing 1.5 MW and 5 MW wind turbines. A physics-informed CGAN architecture is introduced, where a simplified turbine tower dynamics model is incorporated into the training loss to ensure physically consistent outputs. Quantitative results showed moderate to high agreement in frequency-domain features. Incorporating the physics-informed loss function improved the R 2 values by 30%, reduced the RMSE from 1.39 to 1.1 m/s 2 , and reduced training time by 82%. Furthermore, under increased turbulence intensity (IEC Category A), the RMSE remained stable at approximately 1.1 m/s 2 . While the present study is entirely simulation-based, it establishes a pipeline for evaluating physics-informed generative domain translation, which may serve as a foundation for future simulation-to-reality validation studies.

17 WIND ENERGY↗

A machine learning pipeline for identifying infiltration managed aquifer recharge locations from satellite imagery in the San Joaquin Valley, California

This study focuses on an agricultural region in California’s Central Valley, USA, where Managed Aquifer Recharge (MAR) is widely implemented to mitigate groundwater depletion under increasing water demand and climate variability. A deep learning and machine learning framework was developed to identify infiltration-MAR locations using satellite imagery and environmental data. The framework integrates surface water detection from Sentinel-2 imagery, geospatial delineation of water bodies, spatiotemporal tracking of water body dynamics, and supervised classification using meteorological, environmental, and topographic variables. The framework was applied to a 2379 km² study area southwest of Fresno, where 765 water bodies were detected, including 139 identified MAR sites based on publicly available datasets and expert knowledge. The classification model achieved an accuracy of 0.94 and an F1 score of 0.85. Feature importance analysis indicates that cropland, normalized difference vegetation index (NDVI), and evaporation are among the most influential predictors for infiltration-MAR. Notably, the framework suggests that engineered water management in infiltration-MAR systems can disrupt or even reverse the expected positive correlation between surface water extent and precipitation. These findings provide physically interpretable insights into the characteristics of existing infiltration-MAR facilities and demonstrate the potential of the proposed framework as a reproducible, interpretable, and potentially transferable tool for data-driven infiltration-MAR identification and inventory development under growing climatic and hydrological uncertainty.

Classification↗

Deep Cyber-Physical Situational Awareness for Energy Systems: A Secure Foundation for Next-Generation Energy Management

This document provides the final report for the CYPRES project. The purpose is (1) to highlight and summarize its major accomplishments and (2) to provide guidance on how its outcomes have informed and can inform important additional research and technology transfer. The goal of CYPRES was the research, development, and demonstration of a security-oriented next generation cyber-physical EMS for electric power systems that detects malicious and abnormal events through the fusion of cyber and physical data. To achieve this, the CYPRES project team researched, developed, and built a prototype of the solution, referred to as the CYPRES EMS. The CYPRES EMS is a proof-of-concept cyber-physical platform that demonstrates the management of the energy system, communications, security, and cyber-physical grid modeling and analytics. As part of the capabilities of the CYPRES EMS, the team designed and developed a suite of power system applications for monitoring, risk analyses, detection, and control that are inherently cyberaware. At its core, the project aimed to research, develop, and demonstrate a security-oriented next-generation cyber-physical Energy Management System (EMS) capable of detecting malicious and abnormal events through the innovative fusion of cyber and physical data. This approach represents a fundamental shift from traditional EMS, reimagining how critical infrastructure can be protected through unified cyber-aware and physics-aware secure data flow pipelines. The project’s cornerstone deliverable, the CYPRES EMS, serves as a proof-of-concept cyber-physical platform that revolutionizes the management of energy systems, communications, security, and cyber-physical grid modeling and analytics. This prototype implements a comprehensive suite of power system applications for monitoring, risk analyses, detection, and control, all designed with inherent cyber awareness. The system’s architecture extends from end-devices in the field through to control center applications, establishing a secure and resilient control framework that addresses the challenges posed by diverse devices of unknown trustworthiness connecting to modern power systems. Through this innovative approach to deep cyber-physical situational awareness, the CYPRES project not only advances the state-of-the-art in energy infrastructure protection but also establishes a new paradigm for how EMS can be designed, deployed, and operated in an increasingly complex threat landscape. The findings and developments from this project provide crucial insights for stakeholders across the energy sector, offering a blueprint for enhancing the reliability and resilience of our nation’s critical energy infrastructure in the face of evolving cyber threats.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

Automated defect identification in electroluminescence images of solar modules

Solar photovoltaic (PV) modules are susceptible to manufacturing defects, mishandling problems or extreme weather events that can limit energy production or cause early device failure. Trained professionals use electroluminescence (EL) images to identify defects in modules, however, field surveys or inline image acquisition can generate millions of EL images, which are infeasible to analyze by rote inspection. Here, we develop a rapid automatic computer vision pipeline (~0.5 seconds/module) to analyze EL images and identify defects including cracks, intra-cell defects, oxygen-induced defects, and solder disconnections. Defect identification is achieved with a machine learning model (Random Forest, ResNet models and YOLO) trained on 762 manually-labeled EL images of PV modules. We compare model performance on an imbalanced real-world validation set containing 134 EL images and determine that ResNet18 and YOLO are the optimal models; we next evaluated these models on a dedicated testing set (129 module images) with resulting macro F1 scores of 0.83 (ResNet18) and 0.78 (YOLO). Using a field EL survey of a PV power plant damaged in a vegetation fire, we analyze 18,954 EL images (2.4 million cells) and inspect the spatial distribution of defects on the solar modules. The results find increased frequency of ‘crack’, ‘solder’ and ‘intra-cell’ defects on the edges of the solar module closest to the ground after fire. We also find an abnormal increase of striation rings on cells which were assumed to be caused mainly in fabrication process. Our methods are published as open-source software. It can also be used to identify other kinds of defects or process different types of solar cells with minor modification on models by transfer learning.

14 SOLAR ENERGY↗

MODELING AND PARAMETRIC STUDY OF END-GAS AUTOIGNITION TO ALLOW THE REALIZATION OF ULTRA-LOW EMISSIONS, HIGH-EFFICIENCY HEAVY-DUTY SPARK-IGNITED NATURAL GAS ENGINES

Engine knock and misfire are barriers to pathways leading to high-efficiency Spark-Ignited (SI) Natural Gas (NG) engines. The general tendency to knock is highly dependent on engine operating conditions and the fuel reactivity. The problem is further complicated by the low emission limits and the wide range of chemical reactivity in pipeline-quality natural gas. Depending on the region and the source of the natural gas, its reactivity, described by its Methane Number (MN), which is analogous to the Octane Number for liquid SI fuels, can span from 65 to 95. In order to realize diesel-like efficiencies, SI NG engines must be designed to operate at high Brake Mean Effective Pressures (BMEP), near or beyond knock limits, over a wide range of fuel reactivity. This requires a deep understanding of the combustion-engine interactions pertaining to flame propagation and End-Gas Autoignition (EGAI), i.e., the autoignition of the unburned gas (end gas) ahead of the flame front. However, EGAI, if controlled, provides an opportunity to increase SI NG engine efficiency by increasing the combustion rate and the total fraction of burned fuel, mitigating the effects of the slow flame speeds characteristic of natural gas fuels, which generally reduce BMEP and increase unburned hydrocarbon emissions. For this reason, to realize diesel-like efficiencies and ultra-low emissions on SI NG engines, this work proposes the study of the main parameters influencing the modeling and prediction of NG EGAI to allow for its control. In this work, a novel EGAI detection and onset determination method was developed to reliably quantify EGAI for data analysis and engine control. The new method allowed the prediction of EGAI on SI NG engines without the need to use engine- and operating-condition-dependent thresholds and reduced the error in quantifying the fraction of the total energy released by the EGAI event by up to 40%pts. One- and three-dimensional engine models were then developed to study the engine/fuel interactions that lead to NG EGAI and its performance benefits. These models, although having decent agreement with experimental data, showed the need to account for NOx chemistry when predicting NG EGAI due to a consistently later prediction of the EGAI onset (~1.65 crank-angle degrees) and thus, a new reduced chemical mechanism for real NG fuels was developed containing NOx chemistry. The new reduced mechanism improved the EGAI onset prediction agreement to within ±0.5 crank-angle degrees and decreased simulation time during combustion by nearly 50% when using the further reduced AREIS50NOx chemical mechanism. These models were then used to study the role of NG composition on EGAI, evaluate the engine/fuel interactions leading to NG EGAI, and perform engine optimization while leveraging EGAI to increase thermal efficiency. Piston design optimization combined with a Controlled EGAI (C-EGAI) combustion mode allowed a Heavy-Duty (HD) SI NG engine to operate at diesel-like efficiencies, i.e., Brake Thermal Efficiency (BTE) ≥44%. Experimental and modeling data analysis revealed that earlier and faster heat release increases combustion efficiency by an average of 1%pts, increases work transferred to the piston resulting in a decrease in exhaust losses by 50% depending on the engine operating condition while slightly increasing heat losses. Finally, the simulation results revealed an opportunity to further enhance the BTE (up to 50%) by enabling C-EGAI combustion at leaner conditions, λ=1.4-1.6.

Bestel, Diego Bernardi↗