Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A Data-Driven Approach to Real-World Degradation of Backsheets

The objectives of this project are as follows: • The population behavior of fielded modules in various conditions of use • Predictions of materials in specific climatic zones • Understanding of a module’s local environment in the field on its degradation It aims to understand how and what backsheet materials of photovoltaic modules degrade in the real world field. • Field Survey Protocol This project is started from the protocol, because all the data, information, domain knowledge are from experience of the real world field surveys, which is based on the protocol. The Protocol is explored from the experience of the field survey observations. With the increasing of the field surveys, it is refined for three versions, which are Task 1.0 (Section 3.1), Task 6.0 (Section 3.6), Task 10.0 (Section 3.10) respectively. It includes a document and a training video, which is able to direct other teams to follow the same procedure with the sites surveyed during this project. The documents include the detailed information but is not limited to the terminology definition, instruments SOP, preparation items for the surveys, form for the data collections, the order of the information collection. The final version of the protocol can be found at Appendix A, see Section 5. Additional, It can also be found at Open Science Framework (OSF), see Section 3.18 for detail. • Written Waiver Request Before staring the field surveys, the request of the waiver for international surveys is completed, because the limitation of climate zone in the United States, see Appendix B in Section 6 for the request documents. Unfortunately, only 1 international site from Taiwan, China can be finished, due to the COVID-19. • Field Survey According to the protocol we built in Section 3.1, 3.6 and 3.10, 41 sites have been surveyed across seven different climate zones (Cfa, Csa, Csb, BSk, Dfa, Dfb, Am). A variety of materials, including Polyethylene Naphthalate (PEN), Polyethylene Terephthalate (PET), Polyvinyl Fluoride (PVF), Polyvinylidene Fluoride (PVDF), Acrylic PVDF, Fluoroethylene Vinyl Ether (FEVE), and Glass, were identified. These sites are located in various states including California, South Carolina, New Mexico, Maryland, Ohio, Tennessee, Florida, Massachusetts, Illinois, Minnesota, Oregon, Colorado, and Taiwan, Republic of China. The ages of the sites ranged from 2 - 38 years in service and the field size varied from 1 MW - 25 MW. All requirements for the modeling have been satisfied. Some observations like ’Edge Effect’ for the rows and Junction box heating will also be a useful knowledge to build the model. Section 3.7 provides detailed information on the sites visited during this reporting period.

14 SOLAR ENERGY↗

Investigation and Development of Data-Driven D-Region Model for HF Systems Impacts

Space Environment Corporation (SEC) and RP Consultants (RPC) are to develop and validate a weather-capable D region model for making High Frequency (HF) absorption predictions in support of the HF communications and radar communities. The weather-capable model will assimilate solar and earth space observations from NASA satellites. The model will account for solar-induced impacts on HF absorption, including X-rays, Solar Proton Events (SPE's), and auroral precipitation. The work plan includes: I . Optimize D-region model to quickly obtain ion and electron densities for proper HF absorption calculations. 2. Develop indices-driven modules for D-region ionization sources for low, mid, & high latitudes including X-rays, cosmic rays, auroral precipitation, & solar protons. (Note: solar spectrum & auroral modules already exist). 3. Setup low-cost monitors of existing HF beacons and add one single-frequency beacon. 4. Use PENEX HF-link database with HF monitor data to validate D-region/HF absorption model using climatological ionization drivers. 5. Develop algorithms to assimilate NASA satellite data of solar, interplanetary, and auroral observations into ionization source modules. 6. Use PENEX HF-link & HF-beacon data for skill score comparison of assimilation versus climatological D-region/HF absorption model. Only some satellites are available for the PENEX time period, thus, HF-beacon data is necessary. 7. Use HF beacon monitors to develop HF-link data assimilation algorithms for regional improvement to the D-region/HF absorption model.

Eccles, J. V.↗

Towards a data-driven model of hadronization using normalizing flows

We introduce a model of hadronization based on invertible neural networks that faithfully reproduces a simplified version of the Lund string model for meson hadronization. Additionally, we introduce a new training method for normalizing flows, termed MAGIC, that improves the agreement between simulated and experimental distributions of high-level (macroscopic) observables by adjusting single-emission (microscopic) dynamics. Our results constitute an important step toward realizing a machine-learning based model of hadronization that utilizes experimental data during training. Finally, we demonstrate how a Bayesian extension to this normalizing-flow architecture can be used to provide analysis of statistical and modeling uncertainties on the generated observable distributions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Data-Driven Method for Synthetic Extreme Weather Generation and Solar Impact Assessment: Preprint

High-resolution, high-fidelity weather datasets are essential for testing and evaluating the resilience of power systems, particularly under extreme weather conditions. However, existing extreme weather datasets are typically derived from historical events that are localized and may lack the spatial and temporal resolution or scenario diversity needed to test largescale power systems. In this work, we propose a synthetic extreme weather simulation approach capable of generating targeted extreme events, such as hurricanes, using publicly available data sources. Preliminary results demonstrate the impact of a simulated Category 1 hurricane on renewable generation and critical infrastructure in California. The work aims to provide a flexible approach for creating multiple types of extreme weather scenarios across different regions, enabling comprehensive system stress testing, training, and resilience assessment.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Robust Data-Driven Predictive Run-to-Run Control for Automated Serial Sectioning

This letter presents a one-step predictive run-to-run controller (R2R-MPC) for the automation of mechanical serial sectioning (MSS), a destructive material analysis process. To address the inherent uncertainty and disturbances in the MSS process, a robust closed-loop approach is presented. Here, the robust R2R-MPC models the uncertainty of the MSS process using a linear differential inclusion. As an analytical model of the MSS process is unavailable, the differential inclusion is identified from historical data. The R2R-MPC is posed as an optimization problem that computes incremental changes to the control input which minimize the worst-case material removal errors. This optimization-based controller is combined with a run-to-run controller to provide integral action that rejects constant disturbances and tracks constant reference removal rates. To demonstrate the efficacy of our robust R2R-MPC, we present simulation results which compare the presented controller with a conventional non-robust R2R.

42 ENGINEERING↗

Improving Trustworthiness of Data-Driven Power Grid Contingency Analysis With Bayesian Residual Graph Neural Networks

The evolving energy landscape requires novel tools to efficiently perform contingency analysis and reliability assessment of power grids, potentially in real-time. The high computational cost of traditional power flow solvers limits their applicability in practice. Machine learning (ML) surrogates such as deep neural networks (NNs) accelerate power flow solvers computations, enabling high-order contingency analysis and real-time decision-making by learning highly nonlinear functions and integrating grid topology via graph architectures. However, (graph) NNs lack predictive power away from training data and do not provide predictive confidence estimates. Here, we present a Bayesian residual graph NN that integrates knowledge from low-fidelity data via residual training and embeds granular quantification of uncertainties, improving trustworthiness critical for high-consequence decision-making. Applying Bayesian concepts to NNs is challenging due to the high-dimensionality of both the parameter space, complicating derivation of a meaningful prior, and the output space in large grid systems, requiring enhanced techniques to assess the predicted high-dimensional uncertainties. Our contributions include: (1) Deriving a prior for fully connected and graph NNs that leverages low-fidelity data to guide mean predictions and appropriately control prior predictive uncertainty. (2) Integrating this prior within an ensembling with anchoring scheme for efficient approximate posterior inference. (3) Deriving enhanced metrics to assess accuracy of both the mean and uncertainty predictions in high dimensions, appropriately accounting for correlations propagated through graph layers. The resulting Bayesian residual graph NN is tested on a contingency analysis task for 14-bus and 118-bus grids.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Data-driven Expectations for Electromagnetic Counterpart Searches Based on LIGO/Virgo Public Alerts

Searches for electromagnetic counterparts of gravitational-wave signals have redoubled since the first detection in2017 of a binary neutron star merger with a gamma-ray burst, optical/infrared kilonova, and panchromatic after glow. Yet, one LIGO/Virgo observing run later, there has not yet been a second, secure identification of an electromagnetic counterpart. This is not surprising given that the localization uncertainties of events in LIGO and Virgo’s third observing run, O3, were much larger than predicted. We explain this by showing that improvements in data analysis that now allow LIGO/Virgo to detect weaker and hence more poorly localized events have increased the overall number of detections, of which well-localized, gold-plated events make up a smaller proportion overall. We present simulations of the next two LIGO/Virgo/KAGRA observing runs, O4 and O5, that are grounded in the statistics ofO3 public alerts. To illustrate the significant impact that the updated predictions can have, we study the follow-up strategy for the Zwicky Transient Facility. Realistic and timely forecasting of gravitational-wave localization accuracy is paramount given the large commitments of telescope time and the need to prioritize which events are followed up. We include a data release of our simulated localizations as a public proposal planning resource for astronomers

Polina Petrov↗

Mechanical separations of corn stover anatomical fractions in an integrated feedstock preprocessing system: An experimental and data-driven modeling study

High variabilities of material attributes in lignocellulosic biomass present risks for biofuel and biochemical productions and must be mitigated via preprocessing. Since almost no mechanical device is originally designed for processing biomass, how to operate existing apparatuses with efficient performance has not been investigated extensively. This work presents a study on an integrated screening and air classification to separate cobs and stalks from husks and leaves in corn stover. Prototype machine learning models were developed to assess the feasibility of predicting the process outcome based on the measurable parameters. The models trained upon limited experimental data rendered decent predictive accuracy of yield and purity. The experimental data and modeling results collectively suggest decreasing throughput leads to a higher purity. To the contrary, if throughput increases, a lower purity is likely. A possible trade-off between yield and purity of the separated streams indicates the need for optimal combinations of feedstock size, moisture, and throughput to achieve optimized separations. The results of this study also suggest the need to further improve model predictability by developing more accurate formulations for physics governing the integrated unit operations. To accomplish this, additional experimental data needs to be generated for model training.

09 - BIOMASS FUELS↗

A high-throughput experimentation platform for data-driven discovery in electrochemistry

Automating electrochemical analyses combined with artificial intelligence is poised to accelerate discoveries in renewable energy sciences and technologies. This study presents an automated high-throughput electrochemical characterization (AHTech) platform as a cost-effective and versatile tool for rapidly assessing liquid analytes. The Python-controlled platform combines a liquid handling robot, potentiostat, and customizable microelectrode bundles for diverse, reproducible electrochemical measurements in microtiter plates, minimizing chemical consumption and manual effort. To showcase the capability of AHTech, we screened a library of 180 small molecules as electrolyte additives for aqueous zinc metal batteries, generating data for training machine learning models to predict Coulombic efficiencies. Key molecular features governing additive performance were elucidated using Shapley Additive exPlanations and Spearman’s correlation, pinpointing high-performance candidates like cis-4-hydroxy-d-proline, which achieved an average Coulombic efficiency of 99.52% over 200 cycles. The workflow established herein is highly adaptable, offering a powerful framework for accelerating the exploration and optimization of extensive chemical spaces across diverse energy storage and conversion fields.

Lin, Dian-Zhao [Johns Hopkins University, Baltimor↗

Estimation and Bias Correction of Aerosol Abundance using Data-driven Machine Learning and Remote Sensing

Air quality information is increasingly becoming a public health concern, since some of the aerosol particles pose harmful effects to peoples health. One widely available metric of aerosol abundance is the aerosol optical depth (AOD). The AOD is the integrated light extinction coefficient over a vertical atmospheric column of unit cross section, which represents the extent to which the aerosols in that vertical profile prevent the transmission of light by absorption or scattering. The comparison between the AOD measured from the ground-based Aerosol Robotic Network (AERONET) system and the satellite MODIS instruments at 550 nm shows that there is a bias between the two data products. We performed a comprehensive analysis exploring possible factors which may be contributing to the inter-instrumental bias between MODIS and AERONET. The analysis used several measured variables, including the MODIS AOD, as input in order to train a neural network in regression mode to predict the AERONET AOD values. This not only allowed us to obtain an estimate, but also allowed us to infer the optimal sets of variables that played an important role in the prediction. In addition, we applied machine learning to infer the global abundance of ground level PM2.5 from the AOD data and other ancillary satellite and meteorology products. This research is part of our goal to provide air quality information, which can also be useful for global epidemiology studies.

Malakar, Nabin K.↗

Data Driven Commercial Building Energy Code Compliance and Technology Inventory for New York City

Building Performance Standards (BPS) are gaining national traction. A BPS will require new processes in the design, construction, and operation of buildings that take the occupants into account and enable predictive analysis to ensure compliance with current and future GHG emissions caps. In New York City, most buildings over 25,000 square feet will be regulated by a BPS starting in 2024, regardless of whether it is new construction permitted under current energy codes or an existing building. This research is one of the first to begin the evaluation of a long-term series of building policies in the context of an open data ecosystem, in cooperation with city agencies. Existing building policies enacted in NYC have ranged from building energy benchmarking and labeling to energy audits to the regulation of GHG emission in buildings. Through the development of a dataset related to building technologies and energy consumption, this project can help to evaluate if meaningful conclusions can be drawn for the data that has been largely self-reported in compliance with city regulations. This project will also provide lessons learned from a deep dive into these types of datasets to provide best practices for municipalities or states seeking to embark on policies like those enacted in NYC. In addition, a Building Automation System (BAS) Stretch Standard of Care (SSOC) for owners, designers, and building operators will enable the measurement and predictive analysis of energy consumption and GHG emissions at the plant, system, or component level, in anticipation of regulated GHG limits on buildings based on energy use. The SSOC is expected to be suitable for use on a national level. The primary feature of an SSOC is a standardized format for a set of BAS points that can be used to control and to gather data from individual plants, systems, or components that are related to building energy consumption. This project examined how measurements compare to prescriptive or simulation-based energy code targets, finding little correlation between predictive 8760-hour energy modeling and actual energy consumption for a small sample (n=27) of buildings constructed after 2015. Other analysis found that, while large multifamily housing (MFH) buildings showed a general trend similar to predicted reductions in energy use from the implementation of model commercial energy codes, this trend was not evident in the office, K-12 school, and hotel use groups in NYC. No upward or downward trends in energy consumption were found when buildings were grouped by size. Energy audit data were analyzed and it appears that there is bias by audit company on measures recommended to clients. Further research should be performed to cross-analyze this with other attributes, such as building size, vintage, and number of stories. Analysis found that for 281 buildings that were permitted and completed after 2015 and had submitted benchmarking data in 2022, between 81% and 96% (by use group) were found to be in compliance with the 2024 to 2029 NYC BPS emission caps, and between 55% and 89% were in compliance with the 2030-2034 caps. This work is beneficial to the public in helping policymakers and building stakeholders better understand the wide-ranging implications of a BPS.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Data-driven Modeling for Grid Edge IBRs: A Digital Twin Perspective of User-Defined Models

Recent events in Odessa have brought attention to the challenges associated with the interaction between Inverter- Based Resources (IBRs) and the transmission and distribution system. The NERC event diagnosis report has highlighted sev- eral issues, emphasizing the need for continuous performance monitoring of these IBRs by system operators. Key areas of concern include the mismatch of control and protection perfor- mance of IBRs between the original equipment manufacturer (OEM)-provided models and field measurements. The inability to replicate the realistic response can result in incorrect reliability and resilience studies. In this paper, we developed an approach on how to emulate the behavior of an IBR using measurement data obtained for system operators to utilize in real-time and long- term planning. Two experiments are conducted in the phasor domain and electromagnetic transients (EMT) domain to emulate the behavior for grid forming and grid following inverters under various operating conditions and the effectiveness of the proposed model is demonstrated in terms of accuracy and ease of utilizing user-defined models (UDMs)

Mahapatra, Kaveri [BATTELLE (PACIFIC NW LAB)]↗

Data-driven emulation of modal aerosol microphysics via neural operator-based modeling

The complexity and the small characteristic scales of aerosol microphysical processes pose a big challenge for accurate and efficient Earth system simulations at regional and global scales. In this work, we construct and evaluate a surrogate model: the aerosol deep operator network (ADON), a physics-inspired dual-net architecture for emulating the aerosol microphysics parameterization suite in the version 2 of the Energy Earth System Model (E3SMv2). The current version of the surrogate model is trained on a dataset comprising 9.8 million samples obtained from a global E3SMv2 simulation with the horizontal resolution of about one degree under cloud-free conditions. Incorporating domain spatial and temporal coordinates, as well as principle components extracted from training data, the dual-net surrogate model effectively captures the intricate representations of aerosol and the relationship with atmospheric state variables, achieving an R-squared score over $$95.7\%$$ for all the lognormal aerosol modes in the extrapolated regime. The validated model provides feature importance of input variables and their impact on the predictive capacity of the surrogate model in relation to the E3SM. The computational cost of online inference time deployed on CPUs and GPUs with lower precisions highlights ADON’s efficiency and potential in robust predictive modeling for large-scale Earth system computations.

Bai, Zhe↗

Data-driven organic solubility prediction at the limit of aleatoric uncertainty

Abstract Small molecule solubility is a critically important property which affects the efficiency, environmental impact, and phase behavior of synthetic processes. Experimental determination of solubility is a time- and resource-intensive process and existing methods for in silico estimation of solubility are limited by their generality, speed, and accuracy. This work presents two models derived from the FASTPROP and CHEMPROP architectures and trained on BigSolDB which are capable of predicting solubility at arbitrary temperatures for a wide range of small molecules in organic solvent. Both extrapolate to unseen solutes 2–3 times more accurately than the current state-of-the-art model and we demonstrate that they are approaching the aleatoric limit (0.5–1$$\log S$$ log S ) of available test data, suggesting that further improvements in prediction accuracy require more accurate datasets. The FASTPROP-derived model (called FASTSOLV) and the CHEMPROP-based model are open source, freely accessible via a Python package and web interface, highly reproducible, and up to 2 orders of magnitude faster than current alternatives.

Science & Technology - Other Topics↗

Continuous integration data-driven platform of industrial-scale subsurface storage for real-time analytics

This project helped address the growing need for efficient and scalable models to support geological carbon and energy storage, which are crucial for achieving net-zero emissions. Traditionally accurate high-fidelity numerical models have been used to simulate relevant storage processes under a handful of processes, however such models are computationally demanding, making uncertainty quantification impractical. Consequently, we first developed a machine learning framework, based on Graph Neural Operators (GNOs), to improving the accuracy of model predictions for a fixed computational budget. We then developed an Ensemble of Improved Neural Operators (ENO), which uses bagging and Monte Carlo dropout techniques, to further improve prediction accuracy. Lastly, we developed the way to explain progressive transfer learning methods to reduce the amount of training data and computational cost of training (i.e., reduce trainable parameters) when using our models for multiple storage sites. Our numerical investigation, which used real-world case studies, demonstrated that our framework can significantly improve the safety and efficiency of geological storage operations, with potential applications in other domains such as geothermal reservoirs and climate modeling.

54 ENVIRONMENTAL SCIENCES↗

EMPDF : inferring the Milky Way mass with data-driven distribution function in phase space

We introduce the emPDF (empirical distribution function), a novel dynamical modelling method that infers the gravitational potential from kinematic tracers with optimal statistical efficiency under the minimal assumption of steady state. emPDF determines the best-fitting potential by maximizing the similarity between instantaneous kinematics and the time-averaged phase-space distribution function (DF), which is empirically constructed from observation upon the theoretical foundation of oPDF (Han et al. 2016). This approach eliminates the need for presumed functional forms of DFs or orbit libraries required by conventional DF- or orbit-based methods. emPDF stands out for its flexibility, efficiency, and capability in handling observational effects, making it preferable to the popular Jeans equation or other minimal assumption methods, especially for the Milky Way (MW) outer halo where tracers often have limited sample size and poor data quality. We apply emPDF to infer the MW mass profile using Gaia DR3 data of satellite galaxies and globular clusters, obtaining enclosed masses of M (,r) = 26±8, 46±8, 90±13⁠, and 149±40 x 10 10 M ⊙ at r = 30, 50, 100⁠, and 200 kpc, respectively. These are consistent with the updated constraints from simulation-informed DF fitting (Li et al. 2020). While the simulation-informed DF offers superior precision owing to the additional information extracted from simulations, emPDF is independent of such supplementary knowledge and applicable to general tracer populations. emPDF is currently implemented for tracers with complete 6D kinematics within spherical potentials, but it can potentially be extended to address more general problems.

Astrophysics of Galaxies (astro-ph.GA)↗

A data-driven approach to real-time vertical position estimation for NSTX-U vertical stability control

In this paper, a database of 77 996 plasma equilibrium reconstructions from 727 discharges during the initial operation of the NSTX-U spherical tokamak is analyzed to develop a statistically robust model of the plasma vertical position for real-time control. A variety of regression models are developed and tested, ranging in complexity from linear models to deep neural networks, and including input signals ranging from the four pairs of flux loops used historically on NSTX-U up to the full set of 389 real-time signals available to the plasma control system. A linear model based on 140 real-time magnetics signals is found to offer excellent accuracy, with a coefficient of determination R 2 = 0.906. The robustness of this model to limited training data, new operating scenarios, and signal errors is tested, and a procedure is demonstrated to tune the model parameters to optimize its robustness. A time-dependent plasma equilibrium solver, TokaMaker, is used to simulate vertical stability control in NSTX-U, demonstrating that it should be possible to iteratively tune the parameters of a linear vertical position model to stabilize both positive and negative triangularity plasmas in future experiments.

magnetic diagnostics↗

A data-driven method to estimate the antiproton background in the Mu2e experiment

The Mu2e experiment at Fermilab will search for the Charged Lepton Flavour Violating (CLFV) process of coherent, neutrinoless µ− → e − conversion in the field of an aluminum nucleus. The expected signal is a monochromatic electron with the energy of 104.97 MeV, slightly below the muon rest mass. Observation of a CLFV process would provide unambiguous evidence for Beyond the Standard Model (BSM) physics. Mu2e is sensitive to a wide range of BSM models and has the capability to distinguish between them, guiding us towards the most accurate models. The key features of the Mu2e experiment are: (1) a high intensity pulsed negative muon beam with about 1010 stopped µ −/s, and (2) a sophisticated superconducting solenoid system with a gradient magnetic field to form and guide the intense muon beam to the target. The Mu2e physics data taking is expected to begin in 2027. For Run I, the expected 5σ discovery sensitivity is Rµe = 1.2 × 10−15, with a total expected background of 0.11 ± 0.03 events. In the absence of a signal, the expected upper limit is Rµe < 6.2 × 10−16 at 90% CL. The success of this experiment hinges on the accurate estimation of the background from various SM processes that could provide signal-like electrons. One of the background processes is antiprotons annihilating in the stopping target to produce signal like electrons through π0 → γγ decays followed by γ conversions, and π− → µ−ν¯ decays followed by µ− decay. It is a relatively small background with large uncertainty (100%) due to the lack of antiproton production cross section information for the Mu2e proton beam energy of 8 GeV. We have developed a novel methodology to estimate the antiproton background in-situ. This forms the main theme of the thesis. We observed that at Mu2e energies, antiproton annihilation in the stopping target is the only source of events with multiple, simultaneous particle trajectories. From Geant4 simulations, only about 0.2% of the simulated antiproton annihilation events have a signal-like electron. Meanwhile, ∼ 5% of events have multiple reconstructible particle tracks per event. Therefore, we have devised a methodology to reconstruct the multi-track events and estimate the antiproton background by exploiting the large ratio of the production rates of the two final states.

Chithirasreemadam, Namitha [Pisa U.] (ORCID:000000↗