Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “linear regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Simulation and Regression Modeling of Nasa'S X-59 Low-Boom Carpets Across America

NASA’s X-59 aircraft is predicted to produce a significantly quieter cruise sonic boom than traditional N-wave-producing aircraft. A propagation simulation study was undertaken to quantify loudness levels, exposure size, and variability of the X-59’s low-boom carpet using realistic atmospheric profiles across the contiguous United States of America (CONUS). Near-field pressure data of the X-59 in supersonic cruise from NASA’s fully unstructured Navier–Stokes three-dimensional (known as FUN3D) computational fluid dynamics code were propagated using NASA’s PCBoom code, which solves an enhanced Burgers equation along acoustic rays. Atmospheric profiles from the National Oceanic and Atmospheric Administration’s Climate Forecast System Version 2 database were used for propagation at 138 locations across the CONUS. Carpets at each location were generated for aircraft headings in the four cardinal directions. Over one million X-59 carpets were generated in total. The effects of the heading, season, geography, and climate zone on boom levels and exposure size are presented. Multiple linear regression models were developed to estimate carpet width and loudness metrics across the CONUS. These results inform regulators and mission planners on expected variations in boom levels and carpet extent from atmospheric variations. Understanding potential carpet variability is important when planning community noise surveys using the X-59.

X-59↗

Spacebased Estimation of Moisture Transport in Marine Atmosphere Using Support Vector Regression

An improved algorithm is developed based on support vector regression (SVR) to estimate horizonal water vapor transport integrated through the depth of the atmosphere ((Theta)) over the global ocean from observations of surface wind-stress vector by QuikSCAT, cloud drift wind vector derived from the Multi-angle Imaging SpectroRadiometer (MISR) and geostationary satellites, and precipitable water from the Special Sensor Microwave/Imager (SSM/I). The statistical relation is established between the input parameters (the surface wind stress, the 850 mb wind, the precipitable water, time and location) and the target data ((Theta) calculated from rawinsondes and reanalysis of numerical weather prediction model). The results are validated with independent daily rawinsonde observations, monthly mean reanalysis data, and through regional water balance. This study clearly demonstrates the improvement of (Theta) derived from satellite data using SVR over previous data sets based on linear regression and neural network. The SVR methodology reduces both mean bias and standard deviation comparedwith rawinsonde observations. It agrees better with observations from synoptic to seasonal time scales, and compare more favorably with the reanalysis data on seasonal variations. Only the SVR result can achieve the water balance over South America. The rationale of the advantage by SVR method and the impact of adding the upper level wind will also be discussed.

Support vector regression↗

Observed Seasonal to Decadal-Scale Responses in Mesospheric Water Vapor

The 14-yr (1991-2005) time series of mesospheric water vapor from the Halogen Occultation Experiment (HALOE) are analyzed using multiple linear regression (MLR) techniques for their6 seasonal and longer-period terms from 45S to 45N. The distribution of annual average water vapor shows a decrease from a maximum of 6.5 ppmv at 0.2 hPa to about 3.2 ppmv at 0.01 hPa, in accord with the effects of the photolysis of water vapor due to the Lyman-flux. The distribution of the semi-annual cycle amplitudes is nearly hemispherically symmetric at the low latitudes, while that of the annual cycles show larger amplitudes in the northern hemisphere. The diagnosed 11-yr, or solar cycle, max minus min, water vapor values are of the order of several percent at 0.2 hPa to about 23% at 0.01 hPa. The solar cycle terms have larger values in the northern than in the southern hemisphere, particularly in the middle mesosphere, and the associated linear trend terms are anomalously large in the same region. Those anomalies are due, at least in part, to the fact that the amplitudes of the seasonal cycles were varying at northern mid latitudes during 1991-2005, while the corresponding seasonal terms of the MLR model do not allow for that possibility. Although the 11-yr variation in water vapor is essentially hemispherically-symmetric and anti-phased with the solar cycle flux near 0.01 hPa, the concurrent temperature variations produce slightly colder conditions at the northern high latitudes at solar minimum. It is concluded that this temperature difference is most likely the reason for the greater occurrence of polar mesospheric clouds at the northern versus the southern high latitudes at solar minimum during the HALOE time period.

Remsberg, Ellis↗

Selected contribution: redistribution of pulmonary perfusion during weightlessness and increased gravity

To compare the relative contributions of gravity and vascular structure to the distribution of pulmonary blood flow, we flew with pigs on the National Aeronautics and Space Administration KC-135 aircraft. A series of parabolas created alternating weightlessness and 1.8-G conditions. Fluorescent microspheres of varying colors were injected into the pulmonary circulation to mark regional blood flow during different postural and gravitational conditions. The lungs were subsequently removed, air dried, and sectioned into approximately 2 cm(3) pieces. Flow to each piece was determined for the different conditions. Perfusion heterogeneity did not change significantly during weightlessness compared with normal and increased gravitational forces. Regional blood flow to each lung piece changed little despite alterations in posture and gravitational forces. With the use of multiple stepwise linear regression, the contributions of gravity and vascular structure to regional perfusion were separated. We conclude that both gravity and the geometry of the pulmonary vascular tree influence regional pulmonary blood flow. However, the structure of the vascular tree is the primary determinant of regional perfusion in these animals.

short duration↗

Neural network-based classification and regression of magnetohydrodynamic modes in tokamaks

We present a machine learning-based magnetohydrodynamic (MHD) classifier and regressor that utilizes real or complex-valued 3D magnetic sensor array data to determine neoclassical tearing mode (NTM) onset times in tokamaks with millisecond accuracy. The input dataset consists of poloidal profiles of complex Fourier amplitudes with an n = 1 toroidal mode number from 144 human-labeled ITER Baseline Scenario discharges in the DIII-D tokamak, spanning both tearing-dominated and sawtooth-dominated regimes. Since m, n = 2,1 NTMs frequently emerge alongside sawteeth at the same frequency in this scenario, the focus is on isolating the m = 1 and m = 2 components of the n = 1 MHD mode near the tearing onset. To improve model regularization and prediction stability, singular value decomposition was applied to balance the sawtooth and tearing datasets. The enriched datasets facilitated training neural networks that learn the key distinguishing features of sawtooth and tearing modes in the poloidal profiles of their magnetic amplitude and phase. When the modes occur independently, the networks achieve perfect classification due to the modes’ distinct characteristics and low measurement noise. In the more experimentally relevant case where both modes coexist, the networks maintain exceptional performance across key metrics. Tests on synthetic data with known ground truth demonstrate the superior accuracy of the neural network trained on complex-valued input compared to models using real amplitude, phase, or pseudo-complex data, achieving both a mean time delay and standard deviation below 1 ms. Notably, standard linear regression methods fitting the dominant singular modes to the data closely match the neural network’s performance. Applying these methods across a broad range of H-mode scenarios will enable future studies to systematically identify dominant NTM triggers as scenario-specific variables, paving the way for more effective tearing mode avoidance strategies in future fusion reactor designs.

machine learning↗

Prediction of muscle performance during dynamic repetitive movement

BACKGROUND: During long-duration spaceflight, astronauts experience progressive muscle atrophy and often perform strenuous extravehicular activities. Post-flight, there is a lengthy recovery period with an increased risk for injury. Currently, there is a critical need for an enabling tool to optimize muscle performance and to minimize the risk of injury to astronauts while on-orbit and during post-flight recovery. Consequently, these studies were performed to develop a method to address this need. METHODS: Eight test subjects performed a repetitive dynamic exercise to failure at 65% of their upper torso weight using a Lordex spinal machine. Surface electromyography (SEMG) data was collected from the erector spinae back muscle. The SEMG data was evaluated using a 5th order autoregressive (AR) model and linear regression analysis. RESULTS: The best predictor found was an AR parameter, the mean average magnitude of AR poles, with r = 0.75 and p = 0.03. This parameter can predict performance to failure as early as the second repetition of the exercise. CONCLUSION: A method for predicting human muscle performance early during dynamic repetitive exercise was developed. The capability to predict performance to failure has many potential applications to the space program including evaluating countermeasure effectiveness on-orbit, optimizing post-flight recovery, and potential future real-time monitoring capability during extravehicular activity.

Physical Endurance/physiology↗

Decadal-Scale Responses in Middle and Upper Stratospheric Ozone From SAGE II Version 7 Data

Stratospheric Aerosol and Gas Experiment (SAGE II) version 7 (v7) ozone profiles are analyzed for their decadal-scale responses in the middle and upper stratosphere for 1991 and 1992-2005 and compared with those from its previous version 6.2 (v6.2). Multiple linear regression (MLR) analysis is applied to time series of its ozone number density vs. altitude data for a range of latitudes and altitudes. The MLR models that are fit to the time series data include a periodic 11 yr term, and it is in-phase with that of the 11 yr, solar UV (Ultraviolet)-flux throughout most of the latitude/ altitude domain of the middle and upper stratosphere. Several regions that have a response that is not quite in-phase are interpreted as being affected by decadal-scale, dynamical forcings. The maximum minus minimum, solar cycle (SClike) responses for the ozone at the low latitudes are similar from the two SAGE II data versions and vary from about 5 to 2.5% from 35 to 50 km, although they are resolved better with v7. SAGE II v7 ozone is also analyzed for 1984-1998, in order to mitigate effects of end-point anomalies that bias its ozone in 1991 and the analyzed results for 1991-2005 or following the Pinatubo eruption. Its SC-like ozone response in the upper stratosphere is of the order of 4%for 1984-1998 vs. 2.5 to 3%for 1991-2005. The SAGE II v7 results are also recompared with the responses in ozone from the Halogen Occultation Experiment (HALOE) that are in terms of mixing ratio vs. pressure for 1991-2005 and then for late 1992- 2005 to avoid any effects following Pinatubo. Shapes of their respective response profiles agree very well for 1992-2005. The associated linear trends of the ozone are not as negative in 1992-2005 as in 1984-1998, in accord with a leveling off of the effects of reactive chlorine on ozone. It is concluded that the SAGE II v7 ozone yields SC-like ozone responses and trends that are of better quality than those from v6.2.

Remsberg, E. E.↗

The tonotopic map in the embryonic chicken cochlea

The purpose of the present study was to determine the tonotopic map in the chicken cochlea at 19 days of incubation (E19) by obtaining characteristic frequencies (CFs) for primary afferents, labeling the characterized neurons, and documenting their projections to the papilla. The lowest and highest CFs recorded were 188 and 1623 Hz respectively. The embryonic tonotopic map coincided with maps reported for post-hatch chicks. There were no evidence that neurons selective to low frequencies project inappropriately to more basal locations of the embryonic papilla. Linear regression was used to estimate the frequency gradient (b = 0.037 +/- 0.012 In Hz/% [b +/- SEb]) and intercept (In C, where C = 111 Hz) of the semilog plot of frequency versus cochlear position (in % distance from apex). From these estimates the octave distribution was calculated to be 18.7%/octave or 0.58 mm/octave. These quantities were not significantly different from those found in post hatch chickens. We conclude that the tonotopic map of the avian cochlea for CFs between 100 and 1700 Hz is stable and relatively mature from age E19 to post-hatch day 21 (P21). The most striking sign of immaturity in the E19 embryo is the limited range of high CFs. We offer the hypothesis that, between the ages of E19 and P21, improvements in middle ear admittance alone or in combination with functional maturation of the cochlear base may be the principal factors responsible for the appearance of adult-like high CF limits and not an apically shifting tonotopic map.

NASA Discipline Neuroscience↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

The Evolution of Randomized Clinical Trial Designs to Assess Therapeutics in Alzheimer Disease

Importance The success of recent randomized clinical trials (RCTs) for Alzheimer disease (AD), particularly those focusing on anti-amyloid therapies, has been discussed at length. However, the evolution of RCT design features for AD that preceded this success remain underexplored. Objective To describe temporal changes in the features of RCT design for interventions in AD. Evidence Review PubMed, Scopus, and Web of Science databases were searched in January 2025 for phase 2 and 3 AD RCTs published between January 1992 and December 2024. RCTs that investigated an intervention for AD, with a placebo or standard-of-care control group, were included. Four assessors independently reviewed full-text articles to capture study characteristics. Main Outcomes and Measures The number of participants and the duration of RCTs as well as the target population, outcomes, and funding were extracted from published reports. These features were analyzed with respect to time using linear regression and χ 2 analyses. Results The study included 203 RCTs with 79 589 participants testing interventions in AD. From 1992 to 2024, the mean sample size increased by 464% for phase 2 RCTs (from 42 to 237), and 50% for phase 3 RCTs (from 632 to 951), while the mean trial duration increased by 188% (from 16 to 46 weeks) for phase 2, and 256% (from 20 to 71 weeks) for phase 3 RCTs. This longer duration of RCTs may be partially attributed by a greater share of disease-modifying rather than symptomatic treatments. Similarly, more recent trials required AD biomarker evidence for enrollment (from 1 of 36 [2.7%] before 2006 to 40 of 76 [52.6%] since 2019). A substantial difference in the type of therapeutics researched was observed, with anti-amyloid and anti-tau RCTs being more likely to be funded by the pharmaceutical industry compared with neurotransmitter or other RCTs (anti-amyloid or anti-tau, 68 of 71 [95.8%]; neurotransmitter, 52 of 69 [77.6%]; other, 33 of 52 [63.5%]). RCT transparency improved, with more frequent data accessibility statements, registered reports, and better reporting on race and ethnicity. Conclusions and Relevance This methodology research of AD RCTs highlights substantial changes in key features of AD clinical trials from 1992 to 2024. AD RCTs have become larger and longer, such that they are powered to detect smaller clinical differences. The increased sample sizes and duration should enable the detection of smaller and more slowly occurring outcomes, which may lead to successful RCTs of therapies with slower and more subtle efficacy.

General & Internal Medicine↗

Factorization Machine‐Based Active Learning for Functional Materials Design with Optimal Initial Data

The optimization of functional materials is important to enhance their properties, but their complex geometries pose great challenges to optimization. Data-driven algorithms efficiently navigate such complex design spaces by learning relationships between material structures and performance metrics to discover high-performance functional materials. Surrogate-based active learning, continually improving its surrogate model by iteratively including high-quality data points, has emerged as a cost-effective data-driven approach. Furthermore, it can be coupled with quantum computing to enhance optimization processes, especially when paired with a special form of surrogate model (i.e., quadratic unconstrained binary optimization), formulated by factorization machine (FM). However, current practices often overlook the variability in design space sizes when determining the initial data size for optimization. In this work, we investigate the optimal initial data sizes required for efficient convergence across various design space sizes. By employing averaged piecewise linear regression, we identify initiation points where convergence begins, highlighting the crucial role of employing adequate initial data in achieving efficient optimization. These results contribute to the efficient optimization of functional materials by ensuring faster convergence and reducing computational costs in FM-based active learning.

active learning↗

Optical image analysis for graphene layer detection: Enhanced green channel methodology

Graphene, a material of increasing research interest, requires accurate layer identification due to its sensitivity to layer count. Existing methods for graphene layer number identification are either time-consuming or of low accuracy, with high-accuracy methods often requiring expensive processes. This paper aims to address this challenge by proposing a cost-effective and efficient approach. Specifically, the current work highlights only the green channel—one of the three primary color channels (red, green, blue) that make up an optical image—from images of exfoliated graphene flakes for layer count identification. A linear regression is performed between pixel position and substrate green channel value, and this effect is subtracted from the entire optical image to mitigate background effects. By storing the range of green channel values for each type of flake (monolayer, bilayer, or tri-layer) based on a few images, we establish thresholds for identifying different types of layers in a particular setup. Additionally, our methodology allows for flexible threshold tuning using a single reference image, enabling adjustment to changes in detection setup such as illumination level, magnification, or microscope used. Finally, demonstrating high accuracy and flexibility, this methodology presents a suitable technique for graphene layer number identification without the need for large datasets or expensive instruments.

2D materials↗

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han↗

Surrogate model evaluation and building energy benchmarking for commercial buildings

Building energy consumption benchmarking involves challenges associated with various energy patterns for different building types; heating, ventilating, and air-conditioning (HVAC) system types; and climates. Given significant variation in energy use patterns, accurate prediction of long-term energy use using surrogate models remains challenging. Multiple linear regression (MLR) is commonly used for building energy benchmarking because of its simple structure; however, it lacks accuracy compared to other black-box models. Although many studies have compared surrogate models and offer guidance on model selection based on metrics, they do not provide detailed analysis on improving the surrogate model accuracy. In this paper, we implement a surrogate model using polynomial ridge regression (i.e., MLR with interaction terms combined with ridge regularization) for small office and retail strip mall buildings across six HVAC system types and all climate zones, for electricity and natural gas in baseline and proposed scenarios. A simulation workflow is developed using OpenStudio TM /EnergyPlus TM to generate simulation data using measures over a wide range of efficiency inputs. Enhancements based on statistical insights are used for improving the model accuracy using filters, input transformations, and change points. Surrogate models achieved average coefficient of variation of the root mean squared error (CVRMSE) values of 2.17, 1.06, 2.05, and 3.26 for proposed electricity, proposed natural gas, baseline electricity, and baseline natural gas, respectively, with enhancements reducing CVRMSE by an average of 14.9% across all combinations. We provide model interpretation via Shapley additive explanations to determine which input variables most influence energy consumption and provide supportive arguments for enhancements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Construction of 3D MHD pressure drop correlation and flow characterization in the contraction region of a fusion blanket manifold

Inlet and outlet manifolds are typical components of liquid metal (LM) blanket designs of a fusion power reactor to be used to distribute the LM flow into breeding channels and collect it at the exit of the blanket. High pressure loss in the magnetohydrodynamic (MHD) flows featuring abrupt geometrical changes is one of the main feasibility issues of such designs. Recently, optimization studies were conducted to construct 3D MHD pressure drop correlations for a LM flow in an electrically insulating manifold with gradual expansion. Here, the 3D computational approach developed in that study is applied to the outlet manifold featuring gradual contraction. A systematic analysis was performed with a total number of 135 flow cases computed with COMSOL Multiphysics for Hartmann numbers 1000 < Ha < 10,000, Reynolds numbers 100 < Re < 12,000, and contraction angles 45° < θ < 75° for a fixed contraction ratio of 4. The effects of Ha, Re and θ on the flow recirculation, development length and the total pressure drop were carefully examined. A linear regression analysis was used to determine the power rule of pressure drop coefficient k related to Ha and Re, demonstrating a good match with the Ludford layer theory. Eventually, a correlation for the 3D MHD pressure drop coefficient was constructed as a function of Ha, Re and θ. Further, the results were compared against the inlet manifold. It was found that the flow in the inlet manifold exhibits larger recirculation zones. In the investigated range of Ha, Re and θ, the pressure drop coefficient k of the LM MHD flow in the gradual contraction is only slightly lower (< 8 %) than that in the gradual expansion.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Propagation method and planting density influence canopy developmental transition and biomass productivity in Miscanthus × giganteus

Understanding how establishment practices influence the mechanisms underlying Miscanthus × giganteus (miscanthus) productivity and canopy development is critical for optimizing management. Data was collected during the juvenile (2011–2013) and mature (2024) phases of a long-term field experiment established in Urbana, Illinois, to evaluate the effects of propagation method (plug propagation [PP] and rhizome propagation [RP]), planting density (1.0, 0.75, and 0.25 plants m⁻²), and nitrogen application (0 and 67 kg N ha⁻¹) on end-of-season biomass yield, tiller mass, tiller density, and tiller height. Linear regression models identified the dominant predictors of yield across stand ages and management regimes. Planting density, nitrogen (N) application, and propagation method significantly influenced early yield and canopy development. During the juvenile phase, biomass yield was driven by tiller density due to canopy expansion; in the mature phase, yield became driven by tiller mass. The PP plots produced higher tiller density than the RP plots, resulting in faster canopy closure and higher juvenile-phase yields. Rhizome-propagated (RP) plots produced lower tiller density, but individual tillers were 3.3–6.4 g tiller −1 heavier than PP tillers. After the canopy reached equilibrium, the PP and RP yields were similar because greater RP tiller mass compensated for its lower tiller density. Higher planting density resulted in greater yield and tiller density during the second year (2012), but this effect was absent from the third year (2013) onward. In the juvenile phase, N fertilization enhanced yield by 1.6–3.4 Mg ha −1 . Initiating fertilization in 2013 on unfertilized plots produced biomass similar to that in fertilized plots, suggesting yield recovery in the mature phase. These findings revealed that establishment strategies, including propagation method and planting density, influence juvenile miscanthus canopy development and productivity, transitioning from tiller-density- to mass-dominated yields, but not mature phase productivity.

09 BIOMASS FUELS↗

Latent space dynamics identification for interface tracking with application to shock-induced pore collapse

Capturing sharp, evolving interfaces remains a central challenge in reduced-order modeling, especially when data is limited and the system exhibits localized nonlinearities or discontinuities. Here, we propose LaSDI-IT (Latent Space Dynamics Identification for Interface Tracking), a data-driven framework that combines low-dimensional latent dynamics learning with explicit interface-aware encoding to enable accurate and efficient modeling of physical systems involving moving material boundaries. At the core of LaSDI-IT is a revised autoencoder architecture that jointly reconstructs the physical field and an indicator function representing material regions or phases, allowing the model to track complex interface evolution without requiring detailed physical models or mesh adaptation. The latent dynamics are learned through linear regression in the encoded space and generalized across parameter regimes using Gaussian process interpolation with greedy sampling. We demonstrate LaSDI-IT on the problem of shock-induced pore collapse in high explosives, a process characterized by sharp temperature gradients and dynamically deforming pore geometries. The method achieves relative prediction errors below 9% across the parameter space, accurately recovers key quantities of interest such as pore area and hot spot formation, and matches the performance of dense training with only half the data. This latent dynamics prediction was 10 6 times faster than the conventional high-fidelity simulation, proving its utility for multi-query applications. These results highlight LaSDI-IT as a general, data-efficient framework for modeling discontinuity-rich systems in computational physics, with potential applications in multiphase flows, fracture mechanics, and phase change problems.

Gaussian process↗

Performance of a dynamic single bubbler in single and two-phase immiscible liquids

Ensuring nonproliferation and safeguards of special nuclear materials (SNM) is a critical aspect of advancing the nuclear fuel cycle. Traditional bubbler systems used to estimate liquid levels and densities in nuclear recycling processes have limitations, particularly in harsh environments where dip-tube corrosion and buildup necessitate frequent maintenance and recalibration. This study explores the Dynamic Single Bubbler (DSB) method, which utilizes a single dip-tube attached to a linear actuator to estimate liquid properties dynamically. This approach is extended to estimate liquid-liquid interfaces in immiscible liquids and employs a linear regression method to reduce uncertainties and improve accuracy. The DSB method achieved density estimate uncertainties of less than 0.5% and surface level estimate uncertainties typically under 0.5%, across various fluids including water, acetone, methanol, mineral oil, glycerol, and aqueous salt solutions. Results indicate that the DSB method provides accurate and robust estimates of liquid density and surface levels with minimal maintenance and without the need for calibration. Additionally, the method's applicability to immiscible liquids and various dip-tube geometries was demonstrated, showing promise for widespread use in nuclear and other industrial applications.

98 - NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL↗