Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data consistent”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Probabilistic Deliverability Assessment of Distributed Energy Resources via Scenario-Based AC Optimal Power Flow

As electric grids decarbonize and distributed energy resources (DERs) become increasingly prevalent, interconnection assessments must evolve to reflect operational variability and control flexibility. This paper highlights key modeling limitations observed in practice and reviews approaches for modeling uncertainty. It then introduces a Probabilistic Deliverability Assessment (PDA) framework designed to complement and extend existing procedures. The framework integrates scenario-based AC optimal power flow (AC OPF), corrective dispatch, and optional multi-temporal constraints. Together, these form a structured methodology for quantifying DER utilization, deliverability, and reliability under uncertainty in load, generation, and topology. Outputs include interpretable metrics with confidence intervals that inform siting decisions and evaluate compliance with reliability thresholds across sampled operating conditions. A case study on Puerto Rico’s publicly available bulk power system model demonstrates the framework’s application using minimal input data, consistent with current interconnection practice. Across staged fossil generation retirements, the PDA identifies high-value DER sites and regions requiring additional reactive power support. Results are presented through mean dispatch signals, reliability metrics, and geospatial visualizations, demonstrating how the framework provides transparent, data-driven siting recommendations. The framework’s modular design supports incremental adoption within existing workflows, encouraging broader use of AC OPF in interconnection and planning contexts.

14 SOLAR ENERGY↗

Characterization of Partially Observed Epidemics - Application to COVID-19

This report documents a statistical method for the "real-time" characterization of partially observed epidemics. Observations consist of daily counts of symptomatic patients, diagnosed with the disease. Characterization, in this context, refers to estimation of epidemiological parameters that can be used to provide short-term forecasts of the ongoing epidemic, as well as to provide gross information for the time-dependent infection rate. The characterization problem is formulated as a Bayesian inverse problem, and is predicated on a model for the distribution of the incubation period. The model parameters are estimated as distributions using a Markov Chain Monte Carlo (MCMC) method, thus quantifying the uncertainty in the estimates. The method is applied to the COVID-19 pandemic of 2020, using data at the country, provincial (e.g., states) and regional (e.g. county) levels. The epidemiological model includes a stochastic component due to uncertainties in the incubation period. This model-form uncertainty is accommodated by a pseudo-marginal Metropolis-Hastings MCMC sampler, which produces posterior distributions that reflect this uncertainty. We approximate the discrepancy between the data and the epidemiological model using Gaussian and negative binomial error models; the latter was motivated by the over-dispersed count data. For small daily counts we find the performance of the calibrated models to be similar for the two error models. For large daily counts the negative-binomial approximation is numerically unstable unlike the Gaussian error model. Application of the model at the country level (for the United States, Germany, Italy, etc.) generally provided accurate forecasts, as the data consisted of large counts which suppressed the day-to-day variations in the observations. Further, the bulk of the data is sourced over the duration before the relaxation of the curbs on population mixing, and is not confounded by any discernible country-wide second wave of infections. At the state-level, where reporting was poor or which evinced few infections (e.g., New Mexico), the variance in the data posed some, though not insurmountable, difficulties, and forecasts were able to capture the data with large uncertainty bounds. The method was found to be sufficiently sensitive to discern the flattening of the infection and epidemic curve due to shelter-in-place orders after around 90% quantile for the incubation distribution (about 10 days for COVID-19). The proposed model was also used at a regional level to compare the forecasts for the central and north-west regions of New Mexico. Modeling the data for these regions illustrated different disease spread dynamics captured by the model. While in the central region the daily counts peaked in the late April, in the north-west region the ramp-up continued for approximately three more weeks.

59 BASIC BIOLOGICAL SCIENCES↗

Emergent magnetic properties of biphase iron oxide nanorods

We report on the magnetic properties of biphase iron oxide nanorods (NRs) consisting of ferrimagnetic Fe 3 O 4 and antiferromagnetic α-Fe 2 O 3 phases. Annealing as-prepared NRs at 250 °C for 5h, significantly improved the crystallinity of the Fe 3 O 4 phase and enhanced the volume fraction of the α-Fe 2 O 3 phase. Magnetometry data consistently reveal these two magnetically distinct phases, which are not in proximity to each other but separated by a region of disordered spins giving rise to enhanced magnetization at low temperatures when the sample was cooled down from 300 K in the presence of a 1T field to 10 K. This phenomenon which is also known as the pinning effect is much more pronounced in the annealed sample, resulting from the increased volume fraction of the α-Fe 2 O 3 phase which could strengthen the interfacial spin frustration between these two phases and enhance the density of disordered spins at the interface.

36 MATERIALS SCIENCE↗

An assessment of ocean alkalinity enhancement using aqueous hydroxides: kinetics, efficiency, and precipitation thresholds

Abstract. Ocean alkalinity enhancement (OAE) is a promising approach to marine carbon dioxide removal (mCDR) that leverages the large surface area and carbon storage capacity of the oceans to sequester atmospheric CO2 as dissolved bicarbonate (HCO3-). One OAE method involves the conversion of salt in seawater into aqueous alkalinity (NaOH), which is returned to the ocean. The resulting increase in seawater pH and alkalinity causes a shift in dissolved inorganic carbon (DIC) speciation toward carbonate and a decrease in the surface ocean pCO2. The shift in the pCO2 results in enhanced uptake of atmospheric CO2 by the seawater due to gas exchange. In this study, we systematically test the efficiency of CO2 uptake in seawater treated with NaOH at aquarium (15 L) and tank (6000 L) scales to establish operational boundaries for safety and efficiency in advance of scaling up to field experiments. CO2 equilibration occurred on the order of weeks to months, depending on circulation, air forcing, and air bubbling conditions within the test tanks. An increase of ∼0.7–0.9 mol DIC per mol added alkalinity (in the form of NaOH) was observed through analysis of seawater bottle samples and pH sensor data, consistent with the value expected given the values of the carbonate system equilibrium calculations for the range of salinities and temperatures tested. Mineral precipitation occurred when the bulk seawater pH exceeded 10.0 and Ωaragonite exceeded 30.0. This precipitation was dominated by Mg(OH)2 over hours to 1 d before shifting to CaCO3,aragonite precipitation. These data, combined with models of the dilution and advection of alkaline plumes, will allow the estimation of the amount of carbon dioxide removal expected from OAE pilot studies. Future experiments should better approximate field conditions including sediment interactions, biological activity, ocean circulation, air–sea gas exchange rates, and mixing zone dynamics.

Ringham, Mallory C. (ORCID:0000000348020185)↗

Construction of Women’s All-Around Speed Skating Event Performance Prediction Model and Competition Strategy Analysis Based on Machine Learning Algorithms

Introduction Accurately predicting the competitive performance of elite athletes is an essential prerequisite for formulating competitive strategies. Women’s all-around speed skating event consists of four individual subevents, and the competition system is complex and challenging to make accurate predictions on their performance. Objective The present study aims to explore the feasibility and effectiveness of machine learning algorithms for predicting the performance of women’s all-around speed skating event and provide effective training and competition strategies. Methods The data, consisting of 16 seasons of world-class women’s all-around speed skating competition results, used in the present study came from the International Skating Union (ISU). According to the competition rules, distinct features are filtered using lasso regression, and a 5,000 m race model and a medal model are built using a fivefold cross-validation method. Results The results showed that the support vector machine model was the most stable among the 5,000 m race and the medal models, with the highest AUC (0.86, 0.81, respectively). Furthermore, 3,000 m points are the main characteristic factors that decide whether an athlete can qualify for the final. The 11th lap of the 5,000 m, the second lap of the 500 m, and the fourth lap of the 1,500 m are the main characteristic factors that affect the athlete’s ability to win medals. Conclusion Compared with logistic regression, random forest, K-nearest neighbor, naive Bayes, neural network, support vector machine is a more viable algorithm to establish the performance prediction model of women’s all-around speed skating event; excellent performance in the 3,000 m event can facilitate athletes to advance to the final, and athletes with outstanding performance in the 500 m event are more likely competitive for medals.

Liu, Meng↗

Welch Method and Bootstrapping Applied to Subcritical Gamma Noise

We measured the prompt neutron decay constant 𝛼 of the CROCUS zero-power reactor at the Swiss Federal Institute of Technology Lausanne using cross-power spectral density (CPSD) analysis of gamma-gamma correlations from two trans-stilbene organic scintillators positioned near the reactor core. We measured critical and subcritical states, with water levels ranging from 960 mm (critical) to 800 mm (𝜌=−1.4 $ subcritical). Our analysis used the Welch method, dividing signal segments for fast Fourier transform (FFT) frequency analysis and applying bootstrapping uncertainty quantification that uses Welch-defined segments. Results demonstrated a clear increase in the measured 𝛼 as reactor reactivity decreased, distinguishing critical from subcritical conditions. At the 960-mm critical level, 𝛼 was estimated at 155.9 ± 0.7 s −1 , and for the 800-mm subcritical level, 𝛼 increased significantly to 367.3 ± 6.9 s –1 . A linear regression of subcritical states yielded a critical estimate of 154.0 ± 3.1 s –1 , aligning with the static 𝛼 estimate at critical. The bootstrapping method produced normally distributed 𝛼 estimates, confirming data consistency. The gamma CPSD 𝛼 estimates clearly distinguish reactor states and improve monitoring of zero-power reactors. The future deployment of modular and microreactors as potential candidates for noise analysis is demonstrated in CROCUS, particularly zero-power mock-ups of new designs. The improvement of noise analysis in the subcritical domain from this work will support experimental data for reactor deployment and procedure.

CROCUS↗

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Soil Greenhouse Gas Fluxes 04/2020-06/2024

Both raw and processed data measured using a LI-COR 7810 Greenhouse Gas Analyzer in the TEMPEST experiment, part of COMPASS-FME (https://compass.pnnl.gov/FME/COMPASSFME), at Smithsonian Environmental Research Center. This ecosystem-scale experiment probes the effects of saltwater versus freshwater flooding in a coastal deciduous forest.The data consist of both concentrations and fluxes of carbon dioxide and methane measured using static chambers on the soil surface, approximately every two weeks from early 2020 to mid 2024. Some of the measurement points are controls, and some subject to root-exclusion techniques; all measurements are embedded in the TEMPEST control, freshwater, and saltwater plots (see Hopple et al. 2023). These data were generated to understand changing soil greenhouse gas (CO2 and CH4) production and consumption. File types are comma-separated value (.csv) for data, and markdown (.md) for supplementary information.

54 ENVIRONMENTAL SCIENCES↗

Machine learning to discover mineral trapping signatures due to CO 2 injection

Mineral trapping is pursued as a geological CO 2 sequestration (GCS) mechanism because it permanently stores CO 2 in solid phases or minerals. However, CO 2 mineral-trapping mechanisms are poorly understood due to (1) lack of sufficient field and laboratory data characterizing these complex processes, and (2) challenges to develop site-specific reactive-transport models coupling fluid flow and geochemical reactions occurring at various temporal (from milliseconds to years) and spatial (from pore (millimeters) to field (kilometers)) scales. Reactive transport with additional complexities such as heterogeneity can make the simulation outputs even more difficult to interpret because of complex nonlinearity and multi-scale interdependencies. Furthermore, the values of model outputs such as concentrations can vary by several orders of magnitude, making it harder to correlate and characterize the impact of the variables via traditional data interpretation techniques such as exploratory data analyses. Recently, machine learning (ML) has shown promise in feature discovery and in highlighting hidden mechanisms that cannot be obtained by existing data-analytics and statistical methods. In this study, we applied an unsupervised ML approach, non-negative matrix factorization with custom -means clustering (NMF) to the data generated by reactive-transport simulations of GCS. The reactive-transport data consisted of 19 attributes, including four physio-chemical variables (pH, porosity, aqueous CO 2 , and sequestered CO 2 ), six chemical species (K + , Na + , HCO, Ca 2+ , Mg 2+ , Fe 2+ ), and four carbonate minerals (calcite, dolomite, siderite, and ankerite), a feldspar mineral (albite), and four clay minerals (illite, clinochlore, kaolinite, and smectite) over a period of 200 years of simulation time. Furthermore, the simulation data used was for Morrow B sandstone at the Farnsworth hydrocarbon unit in Texas. Data are sampled at two locations within the model domain: (1) at the injection well and (2) 200 m west of the injection well. The injection was performed for a period of 10 years. Using NMF, we estimated the temporal interdependencies among the 19 attributes over a span of 200 years. We found that NMF was able to identify four reaction stages and their dominant attributes; these cannot be directly discerned through traditional visualization (e.g., line plots, Pareto analysis, Glyph-based visualization methods) or exploratory data analysis tools of the simulation data. The four stages were: reactions in the injection phase followed by short-, mid-, and long-term reactions. The NMF analysis also revealed that 10 among the 19 attributes are dominant. These dominant attributes for mineral trapping include calcite, dolomite at injection well, siderite at 200 m away from the injection well, clinochlore, kaolinite, Na + , K + , Ca 2+ , Mg 2+ , pH, and aqeuous CO 2 . Finally, at late times (65–200 years), our results showed that calcite plays a major role in mineral trapping with insignificant contribution from siderite, ankerite, and clay minerals. These findings make the proposed unsupervised ML-model attractive for reactive-transport sensing towards real-time GCS monitoring.

54 ENVIRONMENTAL SCIENCES↗

Extending Component Lifetime And Improving Inverter Reliability (ECLAIIR)

Inverter reliability remains one of the most persistent challenges limiting the performance, availability, and economic viability of utility‑scale photovoltaic (PV) plants. Industry data consistently show that inverters account for the highest share of corrective maintenance events and unplanned outages across PV fleets. These failures result in energy losses, increased O&M costs, and reduced confidence in long‑term solar asset performance. Motivated by these challenges, this project—Extending Component Lifetime and Improving Inverter Reliability (ECLAIIR)—was undertaken to systematically investigate inverter degradation and failure mechanisms, develop predictive maintenance capabilities, and establish data‑driven pathways to improve service life and reduce the Levelized Cost of Energy (LCOE) for large‑scale PV systems. The primary goal of the project was to identify pre‑failure signatures in string inverters using both lab‑based accelerated lifetime testing and field‑based data and to develop predictive maintenance algorithms that can anticipate inverter faults before they occur. Through collaboration with inverter testing laboratory, solar PV plant owner, and failure‑analysis experts, the project advanced the technical understanding of inverter reliability. By instrumenting inverters with thermistors, humidity sensors, power‑quality meters, and acoustic sensors, the research established how multiple sensing modalities can reliably detect deviations from normal behavior hours to days before failure. These findings substantially enhance scientific understanding of inverter failure kinetics and provide the PV industry with the most comprehensive cross‑OEM characterization of early‑stage failure indicators reported to date. Technically, the project demonstrated the effectiveness of predictive maintenance by developing and validating the PreDICT (Predictive Diagnostics of PV Inverters Using Condition Monitoring and Trend Analysis) framework—a multi‑layer diagnostic architecture combining peer‑to‑peer analytics, historical trend modeling, and advanced machine‑learning techniques such as the Sequential Conditional Variational Autoencoder (SCVAE). This predictive model achieved more than 90% accuracy in detecting pre‑failure conditions and provided up to four days of lead time before inverter failure in field scenarios. Economically, the project’s LCOE analysis showed that predictive maintenance can reduce lifetime energy losses and minimize corrective maintenance interventions. Modeling indicated that, depending on inverter failure rates and replacement timelines, predictive maintenance can significantly reduce LCOE impacts associated with inverter downtime: from as high as 19.4% under conventional maintenance strategies to 0.1%–10.17% when predictive analytics are adopted. These results confirm that predictive maintenance is both technically feasible and economically advantageous for utilities and plant operators. The project’s findings also have broad public benefit. By improving inverter reliability and reducing downtime, predictive maintenance directly increases electricity generation from existing PV assets. Enhanced reliability lowers operational costs for utilities, which can translate over time into lower energy costs for consumers. Furthermore, the project’s technical publications, conference presentations, and industry workshops ensure that knowledge gained is shared broadly across the solar industry, supporting workforce development and enabling utilities of all sizes to adopt modern asset‑health monitoring practices. The retrofitting case study and service‑life prediction framework further support informed decision‑making for aging PV fleets, helping operators extend system life and reduce electronic waste. In summary, the ECLAIIR project significantly advanced the state of knowledge on inverter degradation, demonstrated the technical and economic value of predictive maintenance, and delivered actionable tools and insights that support more reliable, cost‑effective, and sustainable PV plant operation. The outcomes of this project will continue to inform utility practices, guide inverter design improvements, and strengthen the long‑term performance of solar assets nationwide.

14 SOLAR ENERGY↗

Results and lessons learned from accelerating radio frequency modeling using machine learning [slides]

The “advanced tokamak” reactor concept is a leading candidate for a steady state fusion pilot plant. An advanced tokamak (AT) sustains a majority of the required plasma current with effects resulting from maintenance of the peaked pressure at the device center. This current is augmented by auxiliary current drive sources. These auxiliary actuators may consist of neutral particle beams and/or radio frequency (RF) systems such as lower hybrid current drive (LHCD) and high harmonic fast wave (HHFW) current drive using radio and microwaves from antennas. The primary focus of this work is to develop models of RF current profile control suitable for use in integrated modeling frameworks and for real-time control in experiments. Direct physics models of RF current drive can be computationally intensive. In order to achieve predictive times appropriate for the thousands of calls needed in real-time control of experiments and for use in integrated models, we will apply modern machine learning (ML) techniques to accelerate these models and interpolate their results. To generate the fast and accurate models for use in control level algorithms and integrated modeling we need to replace present models with high dimensional interpolation of their results. We will perform additional simulations across a broader parameter range for EAST and other tokamaks in different physics regimes (Alcator C-Mod, DIII-D, WEST, CFETR, ARC, ITER) and combine them into a larger database for training and testing of the ML models. Further testing of the control level models with experimental current profile data from EAST and C-Mod tokamaks will provide additional confirmation of the control level model before integration in a tokamak control system or integrated modeling suite. ML will be used to optimize the selection of training data consisting of RF current driven at different values of density profile, temperature profile, plasma current, and wavenumber. ML will also be used to facilitate classification of current drive from these input data. The output of this effort will be a validated classifier capable of determining the current drive profiles for HHFW CD and LHCD on a mille-second timescale. This will provide a breakthrough capability enabling real-time control of RF driven current profiles in experiments including ITER ICRF and use integrated modeling frameworks requiring thousands of current profile calculations in discharge simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Temporal Subsampling Diminishes Small Spatial Scales in Recurrent Neural Network Emulators of Geophysical Turbulence

The immense computational cost of traditional numerical weather and climate models has sparked the development of machine learning (ML) based emulators. Because ML methods benefit from long records of training data, it is common to use data sets that are temporally subsampled relative to the time steps required for the numerical integration of differential equations. Here, we investigate how this often overlooked processing step affects the quality of an emulator's predictions. We implement two ML architectures from a class of methods called reservoir computing: (a) a form of Nonlinear Vector Autoregression (NVAR), and (b) an Echo State Network (ESN). Despite their simplicity, it is well documented that these architectures excel at predicting low dimensional chaotic dynamics. We are therefore motivated to test these architectures in an idealized setting of predicting high dimensional geophysical turbulence as represented by Surface Quasi-Geostrophic dynamics. In all cases, subsampling the training data consistently leads to an increased bias at small spatial scales that resembles numerical diffusion. Interestingly, the NVAR architecture becomes unstable when the temporal resolution is increased, indicating that the polynomial based interactions are insufficient at capturing the detailed nonlinearities of the turbulent flow. The ESN architecture is found to be more robust, suggesting a benefit to the more expensive but more general structure. Spectral errors are reduced by including a penalty on the kinetic energy density spectrum during training, although the subsampling related errors persist. Future work is warranted to understand how the temporal resolution of training data affects other ML architectures.

58 GEOSCIENCES↗

Cosmological parameter estimation with a joint-likelihood analysis of the cosmic microwave background and big bang nucleosynthesis

Here, we present a joint-likelihood analysis of big bang nucleosynthesis (BBN) and cosmic microwave background (CMB) data, consistently combining likelihoods and taking into account uncertainties in nuclear reaction rates for the first time. Bayesian inference is performed on the baryon abundance and the effective number of neutrino species, 𝑁 eff , using a CMB Boltzmann solver in combination with LINX , a new flexible and efficient BBN code. We marginalize over Planck nuisance parameters and nuclear rates to find 𝑁 eff =3.0⁢8$^{+0.15}_{−0.14}$, 2.9⁢4$^{+0.16}_{−0.15}$, or 2.96$^{+0.13}_{−0.14}$, for three separate reaction networks. This framework enables robust testing of the lambda cold dark matter paradigm and its variants with CMB and BBN data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Expansion-history preferences of DESI DR2 and external data

We explore the origin of the preference of Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) baryon acoustic oscillation measurements and external data from cosmic microwave background (CMB) and type Ia supernovae (SNIa) that dark energy behavior departs from that expected in the standard cosmological model with vacuum energy (Λ ⁢CDM). In our analysis, we allow a flexible scaling of the expansion rate with redshift that nevertheless allows reasonably tight constraints on the quantities of interest, and adopt and validate a simple yet accurate compression of the CMB data that allows us to constrain our phenomenological model of the expansion history. We find that data consistently show a preference for a 3%–4% increase in the expansion rate at 𝑧 ≃ 0.7 relative to that predicted by the standard Λ⁢ CDM model, in excellent agreement with results from the less flexible (𝑤 0 ,𝑤 𝑎 ) parametrization which was used in previous analyses. Even though our model allows a departure from the best-fit Λ⁢ CDM model at zero redshift, we find no evidence for such a signal. We also find no evidence (at greater than 1⁢𝜎 significance) for a departure of the expansion rate from the Λ ⁢CDM predictions at higher redshifts for any of the data combinations that we consider. Altogether, our results strengthen the robustness of the findings using the combination of DESI, CMB, and SNIa data to dark-energy modeling assumptions.

Cosmological parameters↗

Seismic Signal Detection on International Monitoring System 3-Component Stations using PhaseNet

In this report we discuss training a deep learning seismic signal detection model on 3-component stations from the International Monitoring System (IMS) using the PhaseNet architecture. Using 14 years of associated signals from the International Data Centre’s (IDC) Late Event Bulletin (LEB), we auto-curated training data consisting of signal windows containing associated arrivals, and noise windows that contain no LEB-associated signals. We trained several models using different waveform window durations (30 seconds and 100 seconds), with and without bandpass filtering. We evaluated the effectiveness of our models using associated signals from the Unconstrained Global Event Bulletin (UGEB) and found that several of our models outperformed the signal detections from the IDC’s Selected Event List 3 (SEL3) arrival table. The SEL3 bulletin evaluated on the UGEB dataset with 100-second waveform windows registered a precision and recall of .15 and .48, respectively, versus .19 and .59 for our filtered-data model. For the 30-second waveform window dataset, the SEL3 bulletin achieved a precision and recall of .31 and .47, respectively, versus .32 and .60 for our filtered-data model. Finally, our models detected signals from all source-to-receiver distances, suggesting it is feasible to use a single PhaseNet model for the IMS network.

58 GEOSCIENCES↗

ATOMIC Simulations and Experimental Data for CaCO3 Mixtures

This data consists of simulations and experimental measurements of laser-induced breakdown spectroscopy (LIBS). The simulations are produced by ATOMIC, a general purpose plasma modeling and kinetics code that has been designed to compute emission (or absorption) spectra from plasmas [1] and are used to develop a statistical characterization of matrix effects. Our overall suite of simulations includes contains several sets of simulations: training and validation sets of simulations for three and four element mixtures of calcium, carbon, oxygen, and nitrogen (included to account for atmosphere) along with simulations of the individual elements. The 4-element simulations include the mixture of all four elements mentioned and for each of the four individual elements. The 3-element simulations include output for the mixture of calcium, carbon, oxygen and for these three individual elements. The training data were produced using a 600-run design, shown in Figure 1, that varies input parameters temperature (T), electron density (Ne), and proportion of the elements calcium, carbon, oxygen, and nitrogen (Ca; C; O; N) for the 4 element output. The 3-element output includes all parameters except for the proportion of nitrogen. The element proportions (all the variables but T and Ne) sum to one and are unused in the single-element simulations. The validation data was produced with a 80-run design shown in Figure 2. The training and validation simulation outputs for the 4-element simulations for the mixture and for the single element calcium are shown as sample simulations in Figures 3 and 4 respectively. The simulations produce spectra over a range of 190nm - 950nm that roughly mimics the range collected by the SciAps Z-300 LIBS instrument that was used for the experimental data. The measured spectra for a CaCO3 (which may include contribution from Earth's atmosphere) in the experiment is shown in in Figure 5. All files are kept in directories whose names indicate the elemental composition (CaCO3, Ca, C, O, or N), number of elements (3 or 4), and purpose (training, which is not labeled in the file name, or validation) with file names numbered to indicate the line in the design files used to produce the simulation. The designs are provided as text files with names indicating their purpose. The experimental data is provided as a CSV file. [1] J Colgan, EJ Judge, DP Kilcrease, and JE Barefield II. Ab-initio modeling of an iron laser-induced plasma: Comparison between theoretical and experimental atomic emission spectra. Spectrochimica Acta Part B: Atomic Spectroscopy, 97:65{73}, 2014.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Underwater Target Detection Software Demonstration on the RivGen Turbine

This repository contains data and processing scripts necessary to train the object detection models utilized in the underwater target detection software demonstration on the RivGen turbine project and to produce performance metrics (precision, recall, mAP50, mAP50-95). - Contents - Data consist of "images" and "labels". Each image has an associated label, both share the same time string in its file name (e.g., 2024_05_25_09_01_57.98.jpg and 2024_05_25_09_01_57.98.txt). Time strings have the format %yyyy_%mm_%dd_%HH_%MM_%SS.%3f. Images and labels were curated from 2021 and 2024 smolt outmigration periods at the project site in Igiugig, AK. Images are monochrome 8-bit images of objects (smolt, debris, and other) passing through the field of view of the deployed cameras during various operational stages of the RivGen turbine. Labels are text files indicating the class and bounding polygon of each object in an image. The provided labels use the "YOLO" label format. - Requirements - Python3.8+ is required to install and run the train and validation script. The README.md provides instruction for installing the requirements from the requirements.py file. - Instructions - The "example_train.py" file ingests the provided data, trains a model, and produces model performance metrics at completion. NOTE: model performance metrics will vary from run to run as a consequence of the random selection of training and validation data.

16 TIDAL AND WAVE POWER↗

Generative models on phase space

Deep generative models such as diffusion and flow matching are powerful machine learning tools capable of learning and sampling from high-dimensional distributions. They are particularly useful when the training data appears to be concentrated on a submanifold of the data embedding space. For high-energy physics data, consisting of collections of relativistic energy-momentum 4-vectors, this submanifold can enforce extremely strong physically-motivated priors, such as energy and momentum conservation. If these constraints are learned only approximately, rather than exactly, this can inhibit the interpretability and reliability of such generative models. To remedy this deficiency, we introduce generative models which are, by construction, confined at every step of their sampling trajectory to the manifold of massless N-particle Lorentz-invariant phase space in the center-of-momentum frame. In the case of diffusion models, the "pure noise" forward process endpoint corresponds to the uniform distribution on phase space, which provides a clear starting point from which to identify how correlations among the particles emerge during the reverse (de-noising) process. We demonstrate that our models are able to learn both few-particle and many-particle distributions with various singularity structures, paving the way for future interpretability studies using generative models trained on simulated jet data.

Bogorad, Zachary [Fermilab]↗

Plug-and-Play Methods for Integrating Physical and Learned Models in Computational Imaging: Theory, algorithms, and applications

Plug-and-play (PnP) priors constitute one of the most widely used frameworks for solving computational imaging problems through the integration of physical models and learned models. PnP leverages high-fidelity physical sensor models and powerful machine learning methods for prior modeling of data to provide state-of-the-art reconstruction algorithms. PnP algorithms alternate between minimizing a data fidelity term to promote data consistency and imposing a learned regularizer in the form of an image denoiser. Recent highly successful applications of PnP algorithms include biomicroscopy, computerized tomography (CT), magnetic resonance imaging (MRI), and joint ptychotomography. This article presents a unified and principled review of PnP by tracing its roots, describing its major variations, summarizing main results, and discussing applications in computational imaging. Additionally, we also point the way toward further developments by discussing recent results on equilibrium equations that formulate the problem associated with PnP algorithms.

97 MATHEMATICS AND COMPUTING↗