Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Process Variability Effects on Tensile Response in Injection Molded, Fluorinated Thermoplastics

The mechanical properties of fluorinated thermoplastics (i.e., tensile strength and elongation) can vary with changes in injection molding processing parameters. Four fluoropolymers are examined: poly(vinylidene fluoride) (PVDF) and random poly(vinylidene fluoride-co-chlorotrifluoroethylene) (PVDF-CTFE) with three CTFE concentrations. Dog bones were manufactured with various cylinder dwell times and mold cooling times to assess the manufacturing sensitivity to the tensile response. Dwell and cooling times increasingly impact mechanical performance as CTFE concentration increases. Specimens exhibit higher tensile strength as a function of injection order. The first injected specimen exhibits the lowest tensile strength and highest elongation in all copolymers. This trend becomes more pronounced among fluoropolymers with higher CTFE concentration and lower weight-averaged molecular weight. Parallel plate rheology was used to obtain the zero-shear viscosity as a function of material type, process, and injection order. We found that in the copolymers, the first injected sample exhibited a lower zero-shear viscosity than the next, which indicates a lower molecular weight in the first injected specimen. This phenomenon was not present for the PVDF homopolymer. Copolymer mechanical uncertainties are hypothesized to result from the shorter molecular weight chains extruding out of the specimens’ sides as a flash due to higher mobility with CTFE segments.

36 MATERIALS SCIENCE↗

Deep Learning Classification of Cheatgrass Invasion in the Western United States Using Biophysical and Remote Sensing Data

Cheatgrass (Bromus tectorum) invasion is driving an emerging cycle of increased fire frequency and irreversible loss of wildlife habitat in the western US. Yet, detailed spatial information about its occurrence is still lacking for much of its presumably invaded range. Deep learning (DL) has demonstrated success for remote sensing applications but is less tested on more challenging tasks like identifying biological invasions using sub-pixel phenomena. We compare two DL architectures and the more conventional Random Forest and Logistic Regression methods to improve upon a previous effort to map cheatgrass occurrence at >2% canopy cover. High-dimensional sets of biophysical, MODIS, and Landsat-7 ETM+ predictor variables are also compared to evaluate different multi-modal data strategies. All model configurations improved results relative to the case study and accuracy generally improved by combining data from both sensors with biophysical data. Cheatgrass occurrence is mapped at 30 m ground sample distance (GSD) with an estimated 78.1% accuracy, compared to 250-m GSD and 71% map accuracy in the case study. Furthermore, DL is shown to be competitive with well-established machine learning methods in a limited data regime, suggesting it can be an effective tool for mapping biological invasions and more broadly for multi-modal remote sensing applications.

54 ENVIRONMENTAL SCIENCES↗

Structure-aware Initialization via Numerical Continuation and Informed Priors

Scientific machine learning (SciML) often operates in ill-conditioned, weakly identifiable regimes due to limited data or indirect observations. In such settings, optimization and inference are highly sensitive to the starting point, making initialization--often under-reported--a consequential degree of freedom. Random initialization is not a neutral default as it induces an implicit prior over candidate solutions and can systematically bias the result, producing large run-to-run variability. Here, we formalize this view by treating initialization as a hidden confounder in SciML and develop a unifying theory for structure-aware initialization via numerical continuation, constructing warm starts from related problem instances. Across representative tasks, including physics-informed neural networks, maximum likelihood estimation, and variational inference, warm starts have been shown to consistently reduce optimization effort and improve reliability.

Data integrity↗

2001 KIPDA Regional Household Travel Survey

The Kentuckiana Planning and Development Agency (KIPDA) conducted this four-month study from October 2000 through January 2001 to update regional travel demand models. The universe for the survey consisted of households in the five-county Louisville transportation planning study area, including Clark and Floyd counties in Indiana and Bullitt, Jefferson, and Oldham counties in Kentucky. Demographic variables and travel behavior characteristics were collected for 4,433 households in a 24-hour period. Of these, 4,113 were completed with households that were selected at random, and the remaining 320 were completed with representatives. The data from the survey only reflect travel by the selected residents of the five counties listed above and do not account for travel in the region by persons who were not residents of those five counties.

1Hz data↗

Urban morphology from a landscape perspective: How building morphology distribution land models (BMDLM) emulate pattern and process

Urban form (e.g., building morphology such as height or footprint) can be used to predict environmental footprints, such as energy/water consumption and carbon emissions. Although progress has been made in predicting building characteristics to fill gaps in observation or derive 3-D representations, the relationships between morphology and other variables such as land use and population are poorly understood. Understanding these relationships may enable projections for how cities will evolve with landscapes in the future. A suite of random forest models, the Building Morphology Distribution Land Models (BMDLM), was developed to determine how well building morphology for two distinct statistical measures (central tendency and frequency) can be predicted using land use (e.g., zoning) and population at different resolutions. Clark County, Nevada and Los Angeles County, California are explored as case studies. Generally, 1-km models outperformed 30-m models. Frequency distribution models had the best performance, especially in LA County. Frequency models significantly outperformed spatial autocorrelative models using inverse distance weighting (IDW). BMDLM offers a new take on modeling urban form in which generalized landscape patterns are characterized to understand the influence of population and zoning on urban development, as described by urban scaling theory.

Sturtevant, Jillian [Baylor Univ., Waco, TX (Unite↗

Reducing Southern Ocean Shortwave Radiation Errors in the ERA5 Reanalysis with Machine Learning and 25 Years of Surface Observations

Earth system models struggle to simulate clouds and their radiative effects over the Southern Ocean, partly due to a lack of measurements and targeted cloud microphysics knowledge. We have evaluated biases of downwelling shortwave radiation in the ERA5 climate reanalysis using 25 years (1995–2019) of summertime surface measurements, collected on the Research and Supply Vessel (RSV) Aurora Australis, the Research Vessel (R/V) Investigator, and at Macquarie Island. During October–March daylight hours, the ERA5 simulation of SW down exhibited large errors (mean bias = 54 W m -2 , mean absolute error = 82 W m -2 , root-mean-square error = 132 W m -2 , and R 2 = 0.71). To determine whether we could improve these statistics, we bypassed ERA5’s radiative transfer model for SW down with machine learning–based models using a number of ERA5’s gridscale meteorological variables as predictors. These models were trained and tested with the surface measurements of SW down using a 10-fold shuffle split. An extreme gradient boosting (XGBoost) and a random forest–based model setup had the best performance relative to ERA5, both with a near complete reduction of the mean bias error, a decrease in the mean absolute error and root-mean-square error by 25% ± 3%, and an increase in the R 2 value of 5% ± 1% over the 10 splits. Large improvements occurred at higher latitudes and cyclone cold sectors, where ERA5 performed most poorly. We further interpret our methods using Shapley additive explanations. Our results indicate that data-driven techniques could have an important role in simulating surface radiation fluxes and in improving reanalysis products.

54 ENVIRONMENTAL SCIENCES↗

Estimating Subhourly Inverter Clipping Loss From Satellite-Derived Irradiance Data

Photovoltaic system production simulations are conventionally run using hourly weather datasets. Hourly simulations are sufficiently accurate to predict the majority of long-term system behavior but cannot resolve high-frequency effects like inverter clipping caused by short-duration irradiance variability. Direct modeling of this subhourly clipping error is only possible for the few locations with high-resolution irradiance datasets. This paper describes a method of predicting the magnitude of this error using a machine learning regressor ensemble model, comprised of a random forest and an XGBoost model, and 30-minute satellite irradiance data. The method predicts a correction for each 30-minute interval with the potential to roll up into 60-minute corrections to match an hourly energy model. The model is trained and validated at locations where the error can be directly simulated from 1-minute ground data. The validation shows low bias at most ground station locations. The model is also applied to gridded satellite irradiance to produce a heatmap of the estimated clipping error across the United States. Finally, the relative importance of each predictor satellite variable is retrieved from the model and discussed.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Say where you sample: Increasing site selection transparency in urban ecology

Urban ecological studies have the potential to extend our understanding of socio-ecological systems beyond that of an individual city or region. Cross-comparative empirical work and synthesis are imperative to develop a general urban ecological theory. This can be achieved only if studies are replicable and generalizable. Transparency in methods reporting facilitates generalizability and replicability by documenting the decisions scientists make during the various steps of research design; this is particularly true for sampling design and selection because of their impact on both internal and external validity and the potential to unintentionally introduce bias. Three interdependent aspects of sample design are study sample selection (e.g., specific organisms, soils, or water), sample specification (measurement of specific variable of interest), and site selection (locations sampled). Of these, documentation of site selection—the where component of sample design—is underrepresented in the urban ecology literature. Using a stratified random sample of 158 papers from 12 major urban ecology journals, we investigated how researchers selected study sites in urban ecosystems and evaluated whether their site selection methods were transparent. We extracted data from these papers using a 50-question, theory-based questionnaire and a multiple-reviewer approach. Our sample represented almost 45 years of urban ecology research across 40 different countries. We found that more than 80% of the papers read were not transparent in their site selection methodology. We do not believe site selection methods are replicable for 70% of the papers read. Key weaknesses include incomplete descriptions of populations and sampling frames, urban gradients, sample selection methods, and property access. Low transparency in reporting the where methodology limits urban ecologists’ ability to assess the internal and external validity of studies’ findings and to replicate published studies; it also limits the generalizability of existing studies. The challenges of low transparency are particularly relevant in urban ecology, a field where standard protocols for site selection and delineation are still being developed. These limitations interfere with the fields’ ability to build theory and inform policy. We conclude by offering a set of recommendations to increase transparency, replicability, and generalizability.

54 ENVIRONMENTAL SCIENCES↗

Hardware-in-the-Loop Testing of Wide-Area Damping Controller for Field Implementation in Large-scale Power Grid

In our previous work, an adaptive measurement-driven wide-area damping controller (WADC) for suppressing inter-area oscillations has been proposed and a hardware prototype was developed and validated through hardware-in-the-loop tests. As a continuation of the work, this paper introduces a WADC software prototype to handle the realistic challenges for field implementation in the control room of the power grid. The WADC software is developed and operated as an openPDC adapter with a graphical user interface (GUI) to monitor the WADC inputs and output, the communication delays and other variables. The software prototype has been fully tested through an enhanced hardware-in-the-loop (HIL) test setup. Its performance is verified under various realistic communication uncertainties, such as random time delays and data losses, with different communication protocols. The experiment results have proven the WADC software can deliver sufficient damping to suppress the targeted oscillation mode in handling various communication uncertainties for future field deployment.

Jia, Xinlan↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

Application of the AI2 Climate Emulator to E3SMv2's Global Atmosphere Model, With a Focus on Precipitation Fidelity

Abstract Can the current successes of global machine learning‐based weather simulators be generalized beyond 2‐week forecasts to stable and accurate multiyear runs? The recently developed AI2 Climate Emulator (ACE) suggests this is feasible, based upon 10‐year simulations with a network trained on output from a physics‐based global atmosphere model using a grid spacing of approximately 110 km and forced by a repeating annual cycle of sea‐surface temperature. Here we show that ACE, without modification, can be trained to emulate another major atmospheric model, EAMv2, run at a comparable grid spacing for at least 10 years with similarly small climate biases—a prerequisite to wider applicability. With an analysis that combines multiple temporal, spatial, and frequency domain perspectives, we show that ACE faithfully represents the spatiotemporal structure of EAMv2 precipitation and related variables. Finally, we show that a pretrained ACE network is able to adapt to a new global climate model simulation data set with 10 fewer training steps than when starting from random initialization, all while still maintaining low levels of climate bias. Further analysis of these fine‐tuning experiments reveal ACE's intriguing ability to interpolate between distinct global climate models.

Duncan, James P. C.↗

Using automated machine learning for the upscaling of gross primary productivity

Estimating gross primary productivity (GPP) over space and time is fundamental for understanding the response of the terrestrial biosphere to climate change. Eddy covariance flux towers provide in situ estimates of GPP at the ecosystem scale, but their sparse geographical distribution limits larger-scale inference. Machine learning (ML) techniques have been used to address this problem by extrapolating local GPP measurements over space using satellite remote sensing data. However, the accuracy of the regression model can be affected by uncertainties introduced by model selection, parameterization, and choice of explanatory features, among others. Recent advances in automated ML (AutoML) provide a novel automated way to select and synthesize different ML models. In this work, we explore the potential of AutoML by training three major AutoML frameworks on eddy covariance measurements of GPP at 243 globally distributed sites. We compared their ability to predict GPP and its spatial and temporal variability based on different sets of remote sensing explanatory variables. Explanatory variables from only Moderate Resolution Imaging Spectroradiometer (MODIS) surface reflectance data and photosynthetically active radiation explained over 70 % of the monthly variability in GPP, while satellite-derived proxies for canopy structure, photosynthetic activity, environmental stressors, and meteorological variables from reanalysis (ERA5-Land) further improved the frameworks' predictive ability. We found that the AutoML framework Auto-sklearn consistently outperformed other AutoML frameworks as well as a classical random forest regressor in predicting GPP but with small performance differences, reaching an r 2 of up to 0.75. We deployed the best-performing framework to generate global wall-to-wall maps highlighting GPP patterns in good agreement with satellite-derived reference data. This research benchmarks the application of AutoML in GPP estimation and assesses its potential and limitations in quantifying global photosynthetic activity.

54 ENVIRONMENTAL SCIENCES↗

Stochastic parametric skeletal dosimetry model for humans: General approach and application to active marrow exposure from bone-seeking beta-particle emitters

The objective of this study is to develop a skeleton model for assessing active marrow dose from bone-seeking beta-emitting radionuclides. This article explains the modeling methodology which accounts for individual variability of the macro- and microstructure of bone tissue. Bone sites with active hematopoiesis are assessed by dividing them into small segments described by simple geometric shapes. Spongiosa, which fills the segments, is modeled as an isotropic three-dimensional grid (framework) of rod-like trabeculae that “run through” the bone marrow. Randomized multiple framework deformations are simulated by changing the positions of the grid nodes and the thickness of the rods. Model grid parameters are selected in accordance with the parameters of spongiosa microstructures taken from the published papers. Stochastic modeling of radiation transport in heterogeneous media simulating the distribution of bone tissue and marrow in each of the segments is performed by Monte Carlo methods. Model output for the human femur at different ages is provided as an example. The uncertainty of dosimetric characteristics associated with individual variability of bone structure was evaluated. An advantage of this methodology for the calculation of doses absorbed in the marrow from bone-seeking radionuclides is that it does not require additional studies of autopsy material. The biokinetic model results will be used in the future to calculate individual doses to members of a cohort exposed to 89,90 Sr from liquid radioactive waste discharged to the Techa River by the Mayak Production Association in 1949–1956. Further study of these unique cohorts provides an opportunity to gain more in-depth knowledge about the effects of chronic radiation on the hematopoietic system. In addition, the proposed model can be used to assess the doses to active marrow under any other scenarios of 90 Sr and 89 Sr intake to humans.

59 BASIC BIOLOGICAL SCIENCES↗

Resilient information and inference networks under mixed-trust sensing

With ubiquitous digitization, sensing, and computational intelligence deployed in increasingly more and broader domains, including critical infrastructure, potentially misleading and destabilizing effects of multimodal anomalies and adversarial behavior are growing in importance. Here, we develop randomized and reinforcement learning-based strategies for strategically recruiting and utilizing deployed (and, thus, vulnerable and potentially faulty and/or compromised) nodes from information and inference networks, while defending against adversaries that attempt to misguide assessments of inferred variables. Recognizing that, besides communication and other costs, sampling from any observable node can either provide true data or dangerously expose our inference to misinformation (without being easily distinguishable what actually happens), the proposed strategies proceed by progressively recruiting nodes and cautiously scaling their information contribution based on assumed, or, in our reinforcement learning approach, intelligently weighed trustworthiness, with the learning approach also considering network-wide, threat-inclusive risk/value tradeoffs. While avoiding the hardware, communication, analytical and computational burden of explicit redundancy, the proposed defensive schemes enable on-the-fly assessments of underlying processes, and system-wide situational awareness with demonstrable resilience against adversarial activities.

97 - MATHEMATICS AND COMPUTING↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

Danovo Energy Solution's presented its paper named: Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events at the 2026 Georgia Tech Fault & Disturbance Analysis Conference. The full paper can be found at OSTI ID# 3169150 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danova Energy Solutions]↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]↗

A unifying modeling abstraction for infinite-dimensional optimization

Infinite-dimensional optimization (InfiniteOpt) problems involve modeling components (variables, objectives, and constraints) that are functions defined over infinite-dimensional domains. Examples include continuous-time dynamic optimization (time is an infinite domain and components are functions of time), PDE optimization problems (space and time are infinite domains and components are functions of space-time), as well as stochastic and semi-infinite optimization (random space is an infinite domain and components are a function of such random space). InfiniteOpt problems also arise from combinations of these problem classes (e.g., stochastic PDE optimization). Given the infinite-dimensional nature of objectives and constraints, one often needs to define appropriate quantities (measures) to properly pose the problem. Moreover, InfiniteOpt problems often need to be transformed into a finite dimensional representation so that they can be solved numerically. Here in this work, we present a unifying abstraction that facilitates the modeling, analysis, and solution of InfiniteOpt problems. The proposed abstraction enables a general treatment of infinite-dimensional domains and provides a measure-centric paradigm to handle associated variables, objectives, and constraints. This abstraction allows us to transfer techniques across disciplines and with this identify new, interesting, and useful modeling paradigms (e.g., event constraints and risk measures defined over time domains). Our abstraction serves as the backbone of an intuitive Julia-based modeling package that we call InfiniteOpt.jl. We demonstrate the developments using diverse case studies arising in engineering.

42 ENGINEERING↗

Organic molecules are deterministically assembled in variably inundated river sediments, but drivers remain unclear

Dissolved organic matter (DOM) is central to ecosystem function. A fundamental challenge is understanding the processes leading to variation in the chemistry of organic molecules that comprise DOM. Here we study these processes in variably inundated riverbed sediments, as an understudied, yet ubiquitous component of rivers. Using null-model approaches adopted from community ecology, we found that within-site variation in environmental conditions caused non-random (i.e., deterministic) shifts in DOM chemistry. Deterministic shifts were observed across diverse biomes, though the strength of determinism varied substantially. We found that the strength of determinism decreased with increasing sediment moisture, but in the form of a constraint space. Many systems fell below the upper constraint boundary, however. We propose a conceptual model based on our results and other publications in which DOM assemblages are hypothesized to be increasingly deterministic across the continuum from the river water column to saturated sediment pore spaces to unsaturated and dry soils/sediments.

59 BASIC BIOLOGICAL SCIENCES↗