Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Constrained regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

34 records · Page 2

Discrete generative diffusion models without stochastic differential equations: A tensor network approach

Diffusion models (DMs) are a class of generative machine learning methods that sample a target distribution by transforming samples of a trivial (often Gaussian) distribution using a learned stochastic differential equation. In standard DMs, this is done by learning a “score function” that reverses the effect of adding diffusive noise to the distribution of interest. Here we consider the generalisation of DMs to lattice systems with discrete degrees of freedom, and where noise is added via Markov chain jump dynamics. We show how to use tensor networks (TNs) to efficiently define and sample such “discrete diffusion models” (DDMs) without explicitly having to solve a stochastic differential equation. We show the following: (i) by parametrising the data and evolution operators as TNs, the denoising dynamics can be represented exactly; (ii) the auto-regressive nature of TNs allows to generate samples efficiently and without bias; (iii) for sampling Boltzmann-like distributions, TNs allow to construct an efficient learning scheme that integrates well with Monte Carlo. We illustrate this approach to study the equilibrium of two models with non-trivial thermodynamics, the d = 1 constrained Fredkin chain and the d = 2 Ising model. Published by the American Physical Society 2025

Causer, Luke (ORCID:0000000194243473)↗

Extratropical Cloud Feedback Constrained by Cloud Sources and Sinks in Cyclones

Constraining cloud feedback in global climate models (GCMs) using observations is important for establishing accurate predictions of future climate. Uncertainty in shortwave cloud feedback (SW FB ) dominates uncertainty in total cloud feedback. Recent studies show a shift toward more positive extratropical SW FB in the latest generations of GCMs leading to the emergence of very high equilibrium climate sensitivity (ECS). In this study, we use precipitation efficiency and albedo susceptibility to constrain liquid water path (LWP) response to warming and SW FB in the Southern Ocean (SO; 50°–80°S). We analyze precipitation in extratropical cyclones (ECs) to learn about extratropical condensed water sink processes, combined with observations of clouds and moisture convergence, and use the analysis to better understand and constrain SW FB . We utilize a perturbed parameter ensemble (PPE) hosted in the Community Atmosphere Model, version 6 (CAM6), to provide a constraint on SW FB based on observations from Clouds and the Earth’s Radiant Energy System (CERES) and Multisensor Advanced Climatology of LWP (MAC-LWP). We apply Gaussian process regression to emulate the model response to all parameters perturbed in the PPE. Confronting the emulator output with observations provides a new estimated response of Earth to global warming. Furthermore, our new estimates of SO LWP reduce the PPE range by 66%–72%, which results in a shortwave cloud radiative effect estimated range that is 27%–34% less than the PPE range. Observations suggest a more positive SO SW FB than the Community Earth System Model, version 2 (CESM2), and consequently do not reject the high climate sensitivity GCMs emerging from the Coupled Model Intercomparison Project phase 6 (CMIP6).

Atmosphere↗

Gaussian-process generative model for the QCD equation of state

We develop a generative model for the nuclear matter equation of state at zero net baryon density using the Gaussian process regression method. We impose first-principles theoretical constraints from lattice quantum chromodynamics and hadron resonance gas at high- and low-temperature regions, respectively. By allowing the trained Gaussian process regression model to vary freely near the phase transition region, we generate random smooth crossover equations of state with different speeds of sound that do not rely on specific parametrizations. Here, we explore a collection of experimental observable dependencies on the generated equations of state, which paves the groundwork for future Bayesian inference studies to use experimental measurements from relativistic heavy-ion collisions to constrain the nuclear matter equation of state.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Regularization via f -Divergence: An Application to Multi-Oxide Spectroscopic Analysis

In this paper, we explore the application of convolutional neural networks (CNNs) for predicting the chemical composition of complex geologic samples in a simulated Martian atmospheric environment. Specifically, we aim to characterize oxide weight percentages (wt.%) of rock samples analyzed by remote Laser-Induced Breakdown Spectroscopy (LIBS), framing the problem as a multi-target regression task . Neural networks trained on LIBS spectra are prone to overfitting due to high spectral complexity, limited labeled data, and measurement noise. While regularization is critical for improving generalization, common methods (e.g., ℓ 2 regularization) impose constraints not directly tied to data distribution properties. We propose a novel regularization method based on a specific ƒ-divergence induced by a graph-based estimator, designed to constrain the distributional discrepancy between predictions and targets. This regularizer serves a dual purpose: (a) mitigating overfitting by enforcing a constraint on the distributional difference between predictions and noisy targets, and (b) acting as an auxiliary loss that penalizes large divergences. To enable backpropagation, we develop a differentiable approximation of this particular ƒ-divergence, making the method feasible for neural networks. Experiments on ChemCam and SuperCam LIBS calibration spectra show that mathematical equation-divergence regularization outperforms or matches standard regularization methods (ℓ 1 , ℓ 2 , dropout) and the classical baseline, partial least squares (PLS). Combining ƒ-divergence regularization with standard regularization yields further performance gains, indicating that distributional regularization is useful in this context giving a promising direction for robust model training in planetary science applications. Source code is publicly available at Klein and Li (2025), https://doi.org/10.11578/dc.20250530.7.

58 GEOSCIENCES↗

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS↗

Arm and shoulder muscle segmentation in axial MRI with UNet deep learning model

Quantifying individual upper-limb muscle volumes from MRI provides key insight into muscle-specific strength, deficits, and adaptations. Manual delineation is the gold standard but time‑intensive, and the performance of current deep learning approaches, particularly for small or anatomically complex muscles, remains incompletely characterized. We evaluated a state‑of‑the‑art deep learning framework across the entire upper limb and analyzed factors governing segmentation performance, with attention to the forearm. Three previously published MRI datasets (1.5 T, 3D GRE T1‑weighted; total n = 39) spanning young, middle‑aged, and older adults were curated and quality‑checked, including expert manual segmentations for 31 muscles. Following multiclass mask reconstruction, we trained three 3D nnU‑Net multiclass models matched to the muscle subsets present across datasets, using five‑fold cross‑validation and a composite Dice Similarity Coefficient (DSC) + cross entropy loss. Segmentation accuracy was assessed with DSC. Performance varied across muscles (mean DSC = 0.806 ± 0.098), ranging from 0.920 (Deltoid) to 0.461 (Extensor pollicis brevis). In uncertainty‑weighted regressions, muscle volume was positively associated with DSC (R2 = 0.36, p < 0.001), whereas training segmentation count and muscle orientation showed negligible associations (R2 ≤ 0.06). A weighted mixed‑effects model identified volume as the strongest evaluated predictor, explaining 23.9% of variance in DSC; orientation and training count each contributed <1%, leaving 61.5% unexplained. These results indicate that deep learning–based segmentation can accurately quantify muscle volume for many upper‑limb muscles but remains constrained for small, low‑contrast forearm muscles.

Gillespie, Samuel↗

Frequency-Nadir-Constrained Unit Commitment for Low-Inertia, High-IBR Island Power Systems [Slides]

The process of energy decarbonization in island power systems is accelerated due to the swift integration of inverter-based renewable energy resources (IBRs). The unique features of such systems, including rapid frequency changes resulting from potential generation outages or imbalances due to the unpredictability of renewable power, pose a significant challenge in maintaining the frequency nadir without external support. This paper presents a unit commitment (UC) model with data-driven frequency nadir constraints, including either frequency nadir or minimum inertia requirements, helping to limit frequency deviations after significant generator outages. The constraints are formulated using a linear regression model that takes advantage of real-world, year-long generation scheduling and dynamic simulation data. The efficacy of the proposed UC model is verified through a year-long simulation in an actual island power system using historical weather data. The alternative minimum inertia constraint, derived from actual system operation assumptions, is also evaluated. Findings demonstrate that the proposed frequency nadir constraint notably improves the system's frequency nadir under high photovoltaic (PV) penetration levels, albeit with a slight increase in generation costs, when compared to the alternative minimum inertia constraint.

14 SOLAR ENERGY↗

Event-Based Energy Impact Tracking and Forecasting with Limited Measurements for Rooftop Units

Packaged air conditioning units and heat pumps, also known as rooftop units (RTUs), are responsible for almost 133 billion kWh of electricity usage annually on site for space cooling U.S. commercial buildings. In addition, the use of heat pumps is a trend we expect to accelerate as buildings transition from fossil fuel-based heating to electricity as a key step for decarbonizing the U.S. commercial buildings sector. However, the operation conditions and energy use of RTUs and heat pumps are usually not well monitored as they are not commonly integrated with building automation systems and lack exposed sensing and control points. To fill this gap, this paper proposes a framework for tracking and forecasting energy impacts resulting from degradation of performance and improved performance for unit servicing using limited data. The proposed framework makes use of a constrained dataset, specifically measurements of the outdoor air temperature and the power demand of individual RTUs, to track and forecast changes in energy use associated with changes in performance over various temporal horizons ranging from days to weeks. Following the detection of an RTU fault, performance degradation, or performance improvement, the framework employs a prediction model to assess the cumulative energy impact. We demonstrate the effectiveness of the method with field-collected data for servicing and degradation examples and compare the predicting accuracy of Gradient Boosting Decision Tree (GBDT) Regression models to Support Vector Regression and Linear Regression models. The results show that GBDT achieved the best accuracy for time-series validation datasets for the servicing and degradation cases, and the prediction model was able to track the cumulative energy impacts of events. The proposed framework can inform building owners of the cumulative change in energy usage of RTUs associated with performance degradation, performance improvement, or a fault.

packaged air conditioners, packaged heat pumps, ro↗

Machine Learning for Mapping Multipactor Susceptibility in RF Systems: Capabilities and Generalization Constraints

Multipactor is a surface-driven electron avalanche phenomenon that degrades the performance and reliability of radio-frequency (RF) systems in particle accelerator and vacuum electronics applications. Multipactor behavior in a given device structure is conventionally assessed through susceptibility charts, which provide a parameter-space characterization of the instability. In this work, we assess the capabilities of machine-learning (ML) models to learn and predict such susceptibility charts and analyze the constraints governing their generalization across materials. Using a simulation-derived dataset spanning six distinct secondary-electron-yield material profiles in a canonical two-surface planar geometry, we train supervised regression models and artificial neural networks to predict the time-averaged electron growth rate, δavg, across the relevant parameter space. Model performance is evaluated using metrics that explicitly probe the structure of susceptibility charts, including Intersection over Union, Structural Similarity Index, and correlation analysis. Tree-based ensemble models outperform neural-network models in reconstructing susceptibility regions and in generalizing across material domains. Principal-component analysis reveals disjoint material feature distributions, indicating that the piecewise mode structure of multipactor susceptibility is difficult to represent with a single global model and that generalization is constrained by data coverage rather than by model complexity. An exhaustive reduced-coverage study further shows that sparse material-space coverage can yield mean performance in the same general range but producing large variability in the susceptibility-region overlap. These results clarify the capabilities of ML-based surrogate models for parameter-space characterization of multipactor discharge. They also provide guidance for their appropriate use in RF system design.

43 PARTICLE ACCELERATORS↗

Refining localtype primordial non-Gaussianity: Sharpened bϕ constraints through bias expansion

Local-type primordial non-Gaussianity (PNG), predicted by many nonminimal models of inflation, creates a scale-dependent contribution to the power spectrum of large-scale structure tracers. Its amplitude is characterized by the product bϕfNLloc, where bϕ is an astrophysical parameter dependent on the properties of the tracer. However, bϕ exhibits significant secondary dependence on halo concentration and other astrophysical properties, which may bias and weaken the constraints on fNLloc. In this work, we demonstrate that incorporating knowledge of the relation between Lagrangian bias parameters and bϕ can significantly enhance PNG constraints. We employ the hybrid effective field theory approach at the field level and a linear regression model to seek a connection between the bias parameters and bϕ for halo and galaxy samples, constructed using the abacussummit simulation suite and mimicking the luminous red galaxies and quasistellar objects of the Dark Energy Spectroscopic Instrument survey. For the fixed-mass halo samples, our full bias model reduces the uncertainty by more than 70%, with most of that improvement coming from b∇, which we find to be an excellent proxy for concentration. For the galaxy samples, our model reduces the uncertainty on bϕ by 80% for all tracers. By adopting Lagrangian-bias informed priors on the parameter bϕ, future analyses can thus constrain fNLloc with less bias and smaller errors.

Hadzhiyska, Boryana↗

Contrasting Carbon–Water–Energy Dynamics in Perennial and Annual Bioenergy Agroecosystems Using Eddy Covariance and Interpretable Machine Learning

Understanding how agroecosystems respond to environmental variability is fundamental to predicting productivity and sustainability under a changing climate. We analyzed 55 site-years of high-frequency eddy covariance observations from five agroecosystems—two perennial grasses (miscanthus and switchgrass), two annual rotation systems (maize–soybean and sorghum–soybean), and a restored native prairie—to examine ecosystem-scale carbon, water, and energy fluxes. Using an interpretable machine-learning framework with regression tree ensembles, Shapley Additive Explanations, and Accumulated Local Effects, we quantified how environmental and temporal factors regulate gross primary productivity (GPP), evapotranspiration (ET), water-use efficiency, and the Bowen ratio. Perennials exhibited stronger physiological buffering and maintained fluxes across a broader range of temperature and moisture conditions, reflecting deeper rooting and persistent canopy cover. Annuals, in contrast, showed greater short-term variability and stronger coupling to atmospheric demand, with GPP and ET declining rapidly under low humidity or soil moisture. Differences in temperature sensitivity of Bowen ratio further revealed that perennials sustained proportionally greater sensible heat flux under cool conditions, whereas annuals exhibited constrained energy exchange when evaporative demand was low. Together, these results demonstrate that crop life cycle and canopy structure are fundamental determinants of ecosystem-scale carbon–water–energy coupling. By integrating long-term flux observations with interpretable machine learning, this study identifies the environmental drivers that shape agroecosystem function and highlights how conversion from annual to perennial feedstocks can enhance climatic resilience and alter land–atmosphere energy feedbacks. These findings provide a data-driven basis for improving crop and Earth-system models and for guiding bioenergy landscape design under future climate scenarios.

Accumulated Local Effects↗

Inequalities in global residential cooling energy use to 2050

Intersecting socio-demographic transformations and warming climates portend increasing worldwide heat exposures and health sequelae. Cooling adaptation via air conditioning (AC) is effective, but energy-intensive and constrained by household-level differences in income and adaptive capacity. Using statistical models trained on a large multi-country household survey dataset (n = 673,215), we project AC adoption and energy use to mid-century at fine spatial resolution worldwide. Globally, the share of households with residential AC could grow from 27% to 41% (range of scenarios assessed: 33-48%), implying up to a doubling of residential cooling electricity consumption, from 1220 to 1940 (scenarios range: 1590-2377) terawatt-hours yr. –1 , emitting between 590 and 1,365 million tons of carbon dioxide equivalent (MtCO 2 e). AC access and utilization will remain highly unequal within and across countries and income groups, with significant regressive impacts. Up to 4 billion people may lack air-conditioning in 2050. Our global gridded projections facilitate incorporation of AC’s vulnerability, health, and decarbonization effects into integrated assessments of climate change.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Suppressing the sample variance of DESI-like galaxy clustering with fast simulations

Ongoing and upcoming galaxy redshift surveys, such as the Dark Energy Spectroscopic Instrument (DESI) survey, will observe vast regions of sky and a wide range of redshifts. In order to model the observations and address various systematic uncertainties, N-body simulations are routinely adopted, however, the number of large simulations with sufficiently high mass resolution is usually limited by available computing time. Therefore, achieving a simulation volume with the effective statistical errors significantly smaller than those of the observations becomes prohibitively expensive. In this study, we apply the Convergence Acceleration by Regression and Pooling (CARPool) method to mitigate the sample variance of the DESI-like galaxy clustering in the AbacusSummit simulations, with the assistance of the quasi-N-body simulations FastPM. Based on the halo occupation distribution (HOD) models, we construct different FastPM galaxy catalogs, including the luminous red galaxies (LRGs), emission line galaxies (ELGs), and quasars, with their number densities and two-point clustering statistics well matched to those of AbacusSummit. We also employ the same initial conditions between AbacusSummit and FastPM to achieve high cross-correlation, as it is useful in effectively suppressing the variance. Our method of reducing noise in clustering is equivalent to performing a simulation with volume larger by a factor of 5 and 4 for LRGs and ELGs, respectively. We also mitigate the standard deviation of the LRG bispectrum with the triangular configurations k 2 = 2k 1 = 0.2 h Mpc -1 by a factor of 1.6. With smaller sample variance on galaxy clustering, we are able to constrain the baryon acoustic oscillations (BAO) scale parameters to higher precision. The CARPool method will be beneficial to better constrain the theoretical systematics of BAO, redshift space distortions (RSD) and primordial non-Gaussianity (NG).

79 ASTRONOMY AND ASTROPHYSICS↗

EFIT-Prime: Probabilistic and physics-constrained reduced-order neural network model for equilibrium reconstruction in DIII-D

We introduce EFIT-Prime, a novel machine learning surrogate model for EFIT (Equilibrium FIT) that integrates probabilistic and physics-informed methodologies to overcome typical limitations associated with deterministic and ad hoc neural network architectures. EFIT-Prime utilizes a neural architecture search-based deep ensemble for robust uncertainty quantification, providing scalable and efficient neural architectures that comprehensively quantify both data and model uncertainties. Physically informed by the Grad–Shafranov equation, EFIT-Prime applies a constraint on the current density J tor and a smoothness constraint on the first derivative of the poloidal flux, ensuring physically plausible solutions. Furthermore, the spatial location of the diagnostics is explicitly incorporated in the inputs to account for their spatial correlation. Extensive evaluations demonstrate EFIT-Prime's accuracy and robustness across diverse scenarios, most notably showing good generalization on negative-triangularity discharges that were excluded from training. Timing studies indicate an ensemble inference time of 15 ms for predicting a new equilibrium, offering the possibility of plasma control in real-time, if the model is optimized for speed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Nonlinearity of the post-spinel transition and its expression in slabs and plumes worldwide

Phase transitions in the mantle control its internal dynamics and structure. The post-spinel transition marks the upper–lower mantle boundary, where ringwoodite dissociates into bridgmanite plus ferropericlase, and its Clapeyron slope regulates mantle flow across it. This interaction has previously been assumed to have no lateral spatial variations, based on the assumption of a linear post-spinel boundary in pressure and temperature. Here we present laser-heated diamond anvil cell experiments with synchrotron X-ray diffraction to better constrain this boundary, especially at higher temperatures. Combining our data with results from the literature, and using a global analysis based on machine learning, we find a pronounced nonlinearity in the post-spinel boundary, with its slope ranging from –4 MPa/K at 2100 K, to –2 MPa/K at 1950 K, and to 0 MPa/K at 1600 K. Changes in temperature over time and space can therefore cause the post-spinel transition to have variable effects on mantle convection and the movement of subducting slabs and upwelling plumes.

58 GEOSCIENCES↗