Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model fitting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Strongly temperature dependent ferroelectric switching in AlN, Al 1-x Sc x N, and Al 1-x B x N thin films

This manuscript reports the temperature dependence of ferroelectric switching in Al 0.84 Sc 0.16 N, Al 0.93 B 0.07 N, and AlN thin films. Polarization reversal is demonstrated in all compositions and is strongly temperature dependent. Between room temperature and 300 °C, the coercive field drops by almost 50% in all samples, while there was very small temperature dependence of the remanent polarization value. Furthermore, over this same temperature range, the relative permittivity increased between 5% and 10%. Polarization reversal was confirmed by piezoelectric coefficient analysis and chemical etching. Applying intrinsic/homogeneous switching models produces nonphysical fits, while models based on thermal activation suggest that switching is regulated by a distribution of pinning sites or nucleation barriers with an average activation energy near 28 meV.

36 MATERIALS SCIENCE↗

Case study: Accounting for response measurement error in fitting a regression model

This article presents a case study motivated by a plot of data that suggested an emerging trend that the authors were faced with explaining. When the measurement error of the data is accounted for, it turns out there was no real trend. Furthermore, this article shows how to use a Bayesian modeling approach to account for the measurement error.

42 ENGINEERING↗

Characterizing the gene–environment interaction underlying natural morphological variation in Neurospora crassa conidiophores using high-throughput phenomics and transcriptomics

Abstract Neurospora crassa propagates through dissemination of conidia, which develop through specialized structures called conidiophores. Recent work has identified striking variation in conidiophore morphology, using a wild population collection from Louisiana, United States of America to classify 3 distinct phenotypes: Wild-Type, Wrap, and Bulky. Little is known about the impact of these phenotypes on sporulation or germination later in the N. crassa life cycle, or about the genetic variation that underlies them. In this study, we show that conidiophore morphology likely affects colonization capacity of wild N. crassa isolates through both sporulation distance and germination on different carbon sources. We generated and crossed homokaryotic strains belonging to each phenotypic group to more robustly fit a model for and estimate heritability of the complex trait, conidiophore architecture. Our fitted model suggests at least 3 genes and 2 epistatic interactions contribute to conidiophore phenotype, which has an estimated heritability of 0.47. To uncover genes contributing to these phenotypes, we performed RNA-sequencing on mycelia and conidiophores of strains representing each of the 3 phenotypes. Our results show that the Bulky strain had a distinct transcriptional profile from that of Wild-Type and Wrap, exhibiting differential expression patterns in clock-controlled genes (ccgs), the conidiation-specific gene con-6, and genes implicated in metabolism and communication. Combined, these results present novel ecological impacts of and differential gene expression underlying natural conidiophore morphological variation, a complex trait that has not yet been thoroughly explored.

59 BASIC BIOLOGICAL SCIENCES↗

Model Exploration of an Information-Based Healthcare Intervention Using Parallelization and Active Learning

This paper describes the application of a large-scale active learning method to characterize the parameter space of a computational agent-based model developed to investigate the impact of CommunityRx, a clinical information-based health intervention that provides patients with personalized information about local community resources to meet basic and self-care needs. Additionally, the diffusion of information about community resources and their use is modeled via networked interactions and their subsequent effect on agents' use of community resources across an urban population. A random forest model is iteratively fitted to model evaluations to characterize the model parameter space with respect to observed empirical data. We demonstrate the feasibility of using high-performance computing and active learning model exploration techniques to characterize large parameter spaces; by partitioning the parameter space into potentially viable and non-viable regions, we rule out regions of space where simulation output is implausible to observed empirical data. We argue that such methods are necessary to enable model exploration in complex computational models that incorporate increasingly available micro-level behavior data. We provide public access to the model and high-performance computing experimentation code.

97 MATHEMATICS AND COMPUTING↗

Stochastic Models, Indices & Optimization Algorithms for Pricing & Hedging Reliability Risks in Modern Power Grids: Data Plan - Princeton

We collected and cleaned the synthetic grid data produced by NREL for the Texas and New York synthetic grids. We developed a high dimensional joint stochastic model for load at the zone level, and solar and wind power productions at the asset level, capturing the spatial and temporal dependencies between all the variables, and demonstrated how such a model could be fitted to historical data. We designed and implemented a simulation engine which can produce Monte Carlo scenarios for the hourly day-ahead values of load, and solar and wind power productions at the spatial and temporal resolutions of the historical data used to fit the model. Finally we developed an open-source Python package which can, from an input grid model, efficiently use forecasts and large numbers of Monte Carlo scenarios to provide unit commitment and economic dispatch for each of these scenarios. The high dimensional stochastic model and the subsequent Monte Carlo simulation engine were implemented in the package PGscen and the corresponding UC and ED optimization programs in the package Vatic.

14 SOLAR ENERGY↗

Empirical estimation of weather-driven yield shocks using biophysical characteristics for U.S. rainfed and irrigated maize, soybeans, and winter wheat

Agricultural yields are highly susceptible to changes in weather system patterns, including both annual and sub-annual changes in temperature and precipitation. Understanding the impacts of future changes in meteorological variables on crop yields have been widely studied in both empirical and process-based models. This work presents a structural econometric approach that combines historical weather data with information about the biophysical growth cycles of maize, winter wheat, and soy to predict the year-to-year yield responses for rainfed crops. Modeled soil moisture and temperature are taken as key predictors, which allows testing the fitted rainfed model’s ability to predict irrigated yields by assuming irrigation produces the numerically optimal level of soil moisture. This approach is grounded in known biophysical processes, which improves confidence in the model predictions. It produces estimates of interannual variability in yields, which enable study of the effects of extreme events on agricultural productivity. Finally, it enables prediction of the potential impacts of changing weather patterns on irrigated crops in areas that are currently primarily rainfed. We present the results of the empirical model, fitted with rainfed data; out-of-sample validation on irrigated crops; and projections of yield shocks under multiple future climate scenarios. Under a bias-corrected GFDL RCP8.5 scenario, this approach predicts average yield change, relative to 2006-2020 average yields, across U.S. counties in the 2040-2060 period of -16.9% and -13.4% for rainfed and irrigated maize, of -19.1% and -18.4% for rainfed and irrigated soy, and of -4.0% and -2.0% for rainfed and irrigated winter wheat. And in the 2070-2090 period of -30.3% and -28.8 % for rainfed and irrigated maize, of -33.1% and -32.8% for rainfed and irrigated soy, and of -6.5% and -5.2%for rainfed and irrigated winter wheat.

54 ENVIRONMENTAL SCIENCES↗

The Atacama Cosmology Telescope: A Measurement of the DR6 CMB Lensing Power Spectrum and Its Implications for Structure Growth

We present new measurements of cosmic microwave background (CMB) lensing over 9400 deg of the sky. These lensing measurements are derived from the Atacama Cosmology Telescope (ACT) Data Release 6 (DR6) CMB data set, which consists of five seasons of ACT CMB temperature and polarization observations. We determine the amplitude of the CMB lensing power spectrum at 2.3% precision (43σ significance) using a novel pipeline that minimizes sensitivity to foregrounds and to noise properties. To ensure that our results are robust, we analyze an extensive set of null tests, consistency tests, and systematic error estimates and employ a blinded analysis framework. Our CMB lensing power spectrum measurement provides constraints on the amplitude of cosmic structure that do not depend on Planck or galaxy survey data, thus giving independent information about large-scale structure growth and potential tensions in structure measurements. The baseline spectrum is well fit by a lensing amplitude of A lens = 1.013 ± 0.023 relative to the Planck 2018 CMB power spectra best-fit ΛCDM model and A lens = 1.005 ± 0.023 relative to the ACT DR4 + WMAP best-fit model. From our lensing power spectrum measurement, we derive constraints on the parameter combination $S^{CMBL}_{8}$ ≡ σ 8 (Ω m /0.3) 0.25 of $S^{CMBL}_{8}$ = 0.818 ± 0.022 from ACT DR6 CMB lensing alone and $S^{CMBL}_{8}$ = 0.813 ± 0.018 when combining ACT DR6 and Planck NPIPE CMB lensing power spectra. These results are in excellent agreement with ΛCDM model constraints from Planck or ACT DR4 + WMAP CMB power spectrum measurements. Our lensing measurements from redshifts z ~ 0.5–5 are thus fully consistent with ΛCDM structure growth predictions based on CMB anisotropies probing primarily z ~ 1100. We find no evidence for a suppression of the amplitude of cosmic structure at low redshifts.

79 ASTRONOMY AND ASTROPHYSICS↗

Developing Practical Models of Complex Salts for Molten Salt Reactors

Molten salt reactors (MSRs) utilize salts as coolant or as the fuel and coolant together with fissile isotopes dissolved in the salt. It is necessary to therefore understand the behavior of the salts to effectively design, operate, and regulate such reactors, and thus there is a need for thermodynamic models for the salt systems. Molten salts, however, are difficult to represent as they exhibit short-range order that is dependent on both composition and temperature. A widely useful approach is the modified quasichemical model in the quadruplet approximation that provides for consideration of first- and second-nearest-neighbor coordination and interactions. Its use in the CALPHAD approach to system modeling requires fitting parameters using standard thermodynamic data such as phase equilibria, heat capacity, and others. A shortcoming of the model is its inability to directly vary coordination numbers with composition or temperature. Another issue is the difficulty in fitting model parameters using regression methods without already having very good initial values. The proposed paper will discuss these issues and note some practical methods for the effective generation of useful models.

Besmann, Theodore M. (ORCID:0000000155980550)↗

Encoding nonlinear and unsteady aerodynamics of limit cycle oscillations using nonlinear sparse Bayesian learning

This article investigates the applicability of a recently proposed, nonlinear sparse Bayesian learning (NSBL) algorithm to identify and estimate the complex aerodynamics of limit cycle oscillations. NSBL provides a semi-analytical framework for determining the data-optimal sparse model nested within a (potentially) over-parameterized model. This is particularly relevant to nonlinear dynamical systems where modelling approaches involve the use of physics-based and data-driven components. In such cases, the data-driven components, where analytical descriptions of the physical processes are not readily available, are often prone to overfitting, meaning that the empirical aspects of these models will often involve the calibration of an unnecessarily large number of parameters. While an overparameterized model may fit the observed data well, such models may be inadequate for making predictions in regimes that are different from those wherein the data were recorded. In view of this, it is desirable to not only calibrate the model parameters, but also identify the optimal compromise between data fit and model complexity. In this article, we exhibit the optimal model discovery for an aeroelastic system wherein the structural dynamics are well-known and described by a differential equation model, coupled with a semi-empirical aerodynamic model for laminar separation flutter, resulting in low-amplitude limit cycle oscillations (LCO). To illustrate the performance of the algorithm, in this article, we use synthetic data and demonstrate the ability of the algorithm to correctly rediscover the optimal model and model parameters, given a known data-generating model. The synthetic data are generated from a forward simulation of a known differential equation model with parameters selected so as to mimic the dynamics observed in wind-tunnel experiments. Subsequently, we demonstrate the performance of the algorithm for model selection using noisy LCO data from wind tunnel experiments. As there is no ground truth available for the experimental data case, we provide a comparison between NSBL and Bayesian model selection to validate the results, and demonstrate the use of NSBL as an efficient alternative to traditional methods.

97 MATHEMATICS AND COMPUTING↗

Aboveground biomass density models for NASA’s Global Ecosystem Dynamics Investigation (GEDI) lidar mission

NASA's Global Ecosystem Dynamics Investigation (GEDI) is collecting spaceborne full waveform lidar data with a primary science goal of producing accurate estimates of forest aboveground biomass density (AGBD). This paper presents the development of the models used to create GEDI's footprint-level (~25 m) AGBD (GEDI04_A) product, including a description of the datasets used and the procedure for final model selection. The data used to fit our models are from a compilation of globally distributed spatially and temporally coincident field and airborne lidar datasets, whereby we simulated GEDI-like waveforms from airborne lidar to build a calibration database. We used this database to expand the geographic extent of past waveform lidar studies, and divided the globe into four broad strata by Plant Functional Type (PFT) and six geographic regions. GEDI's waveform-to-biomass models take the form of parametric Ordinary Least Squares (OLS) models with simulated Relative Height (RH) metrics as predictor variables. From an exhaustive set of candidate models, we selected the best input predictor variables, and data transformations for each geographic stratum in the GEDI domain to produce a set of comprehensive predictive footprint-level models. We found that model selection frequently favored combinations of RH metrics at the 98th, 90th, 50th, and 10th height above ground-level percentiles (RH98, RH90, RH50, and RH10, respectively), but that inclusion of lower RH metrics (e.g. RH10) did not markedly improve model performance. Second, forced inclusion of RH98 in all models was important and did not degrade model performance, and the best performing models were parsimonious, typically having only 1-3 predictors. Third, stratification by geographic domain (PFT, geographic region) improved model performance in comparison to global models without stratification. Fourth, for the vast majority of strata, the best performing models were fit using square root transformation of field AGBD and/or height metrics. There was considerable variability in model performance across geographic strata, and areas with sparse training data and/or high AGBD values had the poorest performance. These models are used to produce global predictions of AGBD, but will be improved in the future as more and better training data become available.

54 ENVIRONMENTAL SCIENCES↗

In Situ Inference for Earth System Predictability

An understanding of future evolution in precipitation extremes is critical to numerous DOE mission questions. Extreme events are by nature short time-scale events that are difficult to diagnose in available model data. Accurate modeling of extreme events necessarily requires high spatial resolution at the storm scale locally. However, the environment in which storms grow is dependent on global, remote, processes. These complex spatiotemporal relationships are impossible to diagnose at resolutions required to accurately model storms responsible for extreme precipitation. At exascale, climate simulations will produce results at fine enough resolution to investigate these relationships. However, the resulting data from these simulations will be far too large to save for post-simulation analysis. We advocate for fitting statistical models inside the simulations as they run, a context known as in situ, which will facilitate scientific investigations using the full fine-scale data stream. Figure 1 shows an example of the type of model we could consider, a Bayesian hierarchical spatial regression model. Precipitation extremes at each grid cell are modeled using extreme value distributions. Since extremes are rare, fitting models to individual grid cells can result in high variance and poor estimates. Instead, the model can be made more robust by smoothing the parameters of the extreme value model across space. Additionally, the parameters themselves can be functionally linked to other variables elsewhere in the simulation. Thus, we can use the fine-scale data to build more robust models for extremes that link extreme behavior to other climate patterns.

54 ENVIRONMENTAL SCIENCES↗

In Situ Inference for Earth System Predictability

Focal Area: Focal Area 3: Insight gleaned from complex simulated data using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI. Science Challenge: An understanding of future evolution in precipitation extremes is critical to numerous DOE mission questions. Extreme events are by nature short time-scale events that are difficult to diagnose in available model data. Accurate modeling of extreme events necessarily requires high spatial resolution at the storm scale locally. However, the environment in which storms grow is dependent on global, remote, processes. These complex spatiotemporal relationships are impossible to diagnose at resolutions required to accurately model storms responsible for extreme precipitation. At exascale, climate simulations will produce results at fine enough resolution to investigate these relationships. However, the resulting data from these simulations will be far too large to save for post-simulation analysis. We advocate for fitting statistical models inside the simulations as they run, a context known as in situ, which will facilitate scientific investigations using the full fine-scale data stream. Figure 1 shows an example of the type of model we could consider, a Bayesian hierarchical spatial regression model. Precipitation extremes at each grid cell are modeled using extreme value distributions. Since extremes are rare, fitting models to individual grid cells can result in high variance and poor estimates. Instead, the model can be made more robust by smoothing the parameters of the extreme value model across space. Additionally, the parameters themselves can be functionally linked to other variables elsewhere in the simulation. Thus, we can use the fine-scale data to build more robust models for extremes that link extreme behavior to other climate patterns.

54 ENVIRONMENTAL SCIENCES↗

Precision measurements of EFT parameters and BAO peak shifts for the Lyman- α forest

We present precision measurements of the bias parameters of the one-loop power spectrum model of the Lyman- α (Ly- α ) forest, derived within the effective field theory (EFT) of large-scale structure. We fit our model to the three-dimensional flux power spectrum measured from the ACCEL 2 hydrodynamic simulations. The EFT model fits the data with an accuracy of below 2% up to k = 2 h Mpc − 1 . Further, we analytically derive how nonlinearities in the three-dimensional clustering of the Ly- α forest introduce biases in measurements of the baryon acoustic oscillations (BAOs) scaling parameters in radial and transverse directions. From our EFT parameter measurements, we obtain a theoretical error budget of Δ α ∥ = − 0.2 % ( Δ α ⊥ = − 0.3 % ) for the radial (transverse) parameters at redshift z = 2.0 . This corresponds to a shift of − 0.3 % (0.1%) for the isotropic (anisotropic) distance measurements. We provide an estimate for the shift of the BAO peak for Ly- α -quasar cross-correlation measurements assuming analytical and simulation-based scaling relations for the nonlinear quasar bias parameters resulting in a shift of − 0.2 % ( − 0.1 % ) for the radial (transverse) dilation parameters, respectively. This analysis emphasizes the robustness of Ly- α forest BAO measurements to the theory modeling. We provide informative priors and an error budget for measuring the BAO feature—a key science driver of the currently observing Dark Energy Spectroscopic Instrument (DESI). Our work paves the way for full-shape cosmological analyses of Ly- α forest data from DESI and upcoming surveys such as the Prime Focus Spectrograph, WEAVE-QSO, and 4MOST. Published by the American Physical Society 2025

de Belsunce, Roger (ORCID:0000000336604028)↗

Probing the Diffuse Ly α Emission on Cosmological Scales: Ly α Emission Intensity Mapping Using the Complete SDSS-IV eBOSS

Based on Sloan Digital Sky Survey Data Release 16, we have detected the large-scale structure of Lyα emission in the universe at redshifts z = 2–3.5 by cross-correlating quasar positions and Lyα emission imprinted in the residual spectra of luminous red galaxies. We apply an analytical model to fit the corresponding Lyα surface brightness profile and multipoles of the redshift-space quasar–Lyα emission cross-correlation function. The model suggests an average cosmic Lyα luminosity density of $6.6^{+3.3}_{-3.1}$ x 10 40 erg s -1 cMpc -3 , a ~2σ detection with a median value about 8–9 times those estimated from deep narrowband surveys of Lyα emitters at similar redshifts. Although the low signal-to-noise ratio prevents us from a significant detection of the Lyα forest–Lyα emission cross-correlation, the measurement is consistent with the prediction of our best-fit model from quasar–Lyα emission cross-correlation within current uncertainties. We rule out the scenario where the Lyα photons mainly originate from quasars. We find that Lyα emission from star-forming galaxies, including contributions from that concentrated around the galaxy centers and that in diffuse Lyα-emitting halos, is able to explain the bulk of the Lyα luminosity density inferred from our measurements. Ongoing and future surveys can further improve the measurements and advance our understanding of the cosmic Lyα emission field.

79 ASTRONOMY AND ASTROPHYSICS↗

HighDimMixedModels.jl: Robust high-dimensional mixed-effects models across omics data

High-dimensional mixed-effects models are an increasingly important form of regression in which the number of covariates rivals or exceeds the number of samples, which are collected in groups or clusters. The penalized likelihood approach to fitting these models relies on a coordinate descent algorithm that lacks guarantees of convergence to a global optimum. Here, we empirically study the behavior of this algorithm on simulated and real examples of three types of data that are common in modern biology: transcriptome, genome-wide association, and microbiome data. Our simulations provide new insights into the algorithm’s behavior in these settings, and, comparing the performance of two popular penalties, we demonstrate that the smoothly clipped absolute deviation (SCAD) penalty consistently outperforms the least absolute shrinkage and selection operator (LASSO) penalty in terms of both variable selection and estimation accuracy across omics data. To empower researchers in biology and other fields to fit models with the SCAD penalty, we implement the algorithm in a Julia package, HighDimMixedModels.jl .

Gorstein, Evan↗

Improving neutrino-nuclei interaction models: Recommendations and case studies on Peelle’s Pertinent Puzzle

Improving the modeling of neutrino-nuclei interactions using data-driven methods is crucial for high-precision neutrino oscillation experiments. This paper investigates Peelle’s Pertinent Puzzle (PPP) in the context of neutrino measurements, a longstanding challenge to fitting theoretical models to experimental data. Inconsistencies in data-model comparisons hinder efforts to enhance the accuracy and reliability of model predictions. We analyze various sources contributing to these inconsistencies and propose strategies to address them, supported by practical case studies. We advocate for incorporating model fitting exercises as a standard practice in cross section publications to enhance the robustness of results. We use a common analysis framework to explore PPP-related challenges with MicroBooNE and T2K data in an unified manner. Our findings offer valuable insights for improving the accuracy and reliability of neutrino-nuclei interaction models, particularly by systematically tuning models using data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗