Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Increasing sensitivity of dryland vegetation greenness to precipitation due to rising atmospheric CO 2

Water availability plays a critical role in shaping terrestrial ecosystems, particularly in low- and mid-latitude regions. The sensitivity of vegetation growth to precipitation strongly regulates global vegetation dynamics and their responses to drought, yet sensitivity changes in response to climate change remain poorly understood. Here we use long-term satellite observations combined with a dynamic statistical learning approach to examine changes in the sensitivity of vegetation greenness to precipitation over the past four decades. We observe a robust increase in precipitation sensitivity (0.624% yr –1 ) for drylands, and a decrease (–0.618% yr –1 ) for wet regions. Using model simulations, we show that the contrasting trends between dry and wet regions are caused by elevated atmospheric CO 2 (eCO 2 ). eCO 2 universally decreases the precipitation sensitivity by reducing leaf-level transpiration, particularly in wet regions. However, in drylands, this leaf-level transpiration reduction is overridden at the canopy scale by a large proportional increase in leaf area. The increased sensitivity for global drylands implies a potential decrease in ecosystem stability and greater impacts of droughts in these vulnerable ecosystems under continued global change.

54 ENVIRONMENTAL SCIENCES↗

Exacerbated drought impacts on global ecosystems due to structural overshoot

Vegetation dynamics are affected not only by the concurrent climate but also by memory-induced lagged responses. For example, favourable climate in the past could stimulate vegetation growth to surpass the ecosystem carrying capacity, leaving an ecosystem vulnerable to climate stresses. This phenomenon, known as structural overshoot, could potentially contribute to worldwide drought stress and forest mortality but the magnitude of the impact is poorly known due to the dynamic nature of overshoot and complex influencing timescales. Here, we use a dynamic statistical learning approach to identify and characterize ecosystem structural overshoot globally and quantify the associated drought impacts. We find that structural overshoot contributed to around 11% of drought events during 1981-2015 and is often associated with compound extreme drought and heat, causing faster vegetation declines and greater drought impacts compared to non-overshoot related droughts. The fraction of droughts related to overshoot is strongly related to mean annual temperature, with biodiversity, aridity and land cover as secondary factors. These results highlight the large role vegetation dynamics play in drought development and suggest that soil water depletion due to warming-induced future increases in vegetation could cause more frequent and stronger overshoot droughts.

54 ENVIRONMENTAL SCIENCES↗

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We find that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no significant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption.

Oles, Vlad↗

Automated CT registration, segmentation, and quantification (AutoCT) v1.0

Processing and analyzing brain imaging is crucial in both scientific development and clinical field. In this software package, we build a pipeline that integrates automatic registration, segmentation, and quantitative analysis for subjects' CT scans. Leveraging diffeomorphic transformations, we enable optimized forward and inverse mappings between an image and the reference. Furthermore, we extract localized features from deformation field based on an online template process, which advances statistical learning downstream. The created templates, atlas as well as our methods provide the brain imaging community tools for AI implementations.

Essiari, Abdelilah↗

Automated CT registration, segmentation, and quantification (AutoCT) v1.1

Processing and analyzing brain imaging is crucial in both scientific development and clinical field. In this software package, we build a pipeline that integrates automatic registration, segmentation, and quantitative analysis for subjects' CT scans. Leveraging diffeomorphic transofrmations, we enable optimized forward and inverse mappings between an image and the reference. Furthermore, we extract localized features from deformation field based on an online template process, which advances statistical learning downstream. The created templates, atlas as well as our methods provide the brain imaging community tools for AI implementations

Bai, Zhe↗

ForceFinder

SAND2025-11750O ForceFinder extends the Structural Dynamics Python Libraries (SDynPy) with comprehensive tools for inverse source estimation (ISE) tasks via frequency response function (FRF) matrix inversion. The software is designed for transfer path analysis and multiple-input/multiple-output (MIMO) vibration control problems. It allows users to estimate sources through various algorithms, from the basic Moore-Penrose pseudo-inverse to statistical learning methods such as Tikhonov regularization via an L-curve and elastic net regularization via an information criterion. ForceFinder uses an object-oriented framework, where all components of the ISE problem—such as FRFs, responses, and transformations—are stored in a "SourcePathReceiver" object. This software can be applied to any noise and vibration problem. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Carter, Steven [Sandia National Lab. (SNL-CA), Liv↗

Investigating Aerosol and Meteorological Influences on Convective Clouds in Houston, Texas, during the TRACER/ESCAPE Field Campaigns

Aerosols serve as cloud condensation nuclei, shaping the microphysical properties of cloud droplets. Aerosol effects on convective clouds are complex and remain controversial. The debate centers around the process of aerosol-induced invigoration of deep convection, a phenomenon that could significantly affect convective cloud properties but lacks robust evidence due to methodological limitations in observational approaches and questions about the robustness of modeling studies. Resolving these discrepancies is crucial for understanding how aerosols affect the atmosphere. Here, this study examines the effects of meteorological and aerosol parameters in a weakly synoptic-driven convective environment, where the influence of aerosols may be more pronounced and observable. Daily atmospheric soundings and aerosol concentrations from several ground instruments collected during the summer of 2022 in Houston, Texas, as part of the Tracking Aerosol Convection interactions Experiment (TRACER) and Experiment of Sea Breeze Convection, Aerosols, Precipitation, and Environment (ESCAPE) field campaigns are analyzed. Statistical learning methods are applied to uncover the complex relationships between aerosols, meteorology, and convective cloud characteristics, such as cell area and echo-top height. The findings reveal that higher aerosol concentrations are associated with narrower convective cells, which we argue contradicts the idea of stronger convection with increased aerosol loading. However, once the data are clustered by the synoptic environment, the relationship between aerosol loading and convective cell area diminishes, indicating that the covariablity between synoptic-scale weather patterns, local thermodynamics, and aerosol loading makes it challenging to draw definitive conclusions about the specific impacts of aerosols on convective cloud properties.

54 ENVIRONMENTAL SCIENCES↗

Large-Scale Trajectory Analysis via Feature Vectors

The explosion of both sensors and GPS-enabled devices has resulted in position/time data being the next big frontier for data analytics. However, many of the problems associated with large numbers of trajectories do not necessarily have an analog with many of the historic big-data applications such as text and image analysis. Modern trajectory analytics exploits much of the cutting-edge research in machine-learning, statistics, computational geometry and other disciplines. We will show that for doing trajectory analytics at scale, it is necessary to fundamentally change the way the information is represented through a feature-vector approach. We then demonstrate the ability to solve large trajectory analytics problems using this representation.

58 GEOSCIENCES↗

Advanced Computing, Data Science, and Artificial Intelligence Research Opportunities for Energy-Focused Transportation Science

The Energy Efficient Mobility Systems (EEMS) technology landscape is complex and rapidly evolving, which provides both tremendous opportunities and formidable challenges. Significant alterations to the mobility landscape are underway due to the advent of vehicle and infrastructure connectivity, autonomous driving, and rapid passenger- and freight-vehicle electrification. Advanced computing will play an increasingly important role in enabling the EEMS program to understand and identify the most important levers to improve the energy productivity of future integrated mobility systems. It is also driving new approaches to mobility and the research to unlock an affordable, efficient, safe, and accessible transportation future. Driving much of this change is the collection, analysis, and strategic use of massive amounts of diverse, complex data from infrastructure and vehicles with on-board sensors and data storage and transmission capabilities. Diverse and representative data are key to implementing approaches to maximize mobility energy productivity. While high-fidelity modeling of integrated transportation networks has strengthened our understanding of dynamic movement and behavior patterns, existing tools must be expanded beyond their current focus. This work necessitates data infrastructure investments (e.g., secure-streaming data platforms driven by ubiquitous sensors and video analytics) as well as investments in critical capabilities for large-scale automated analysis and organization using modern machine learning, statistics, and artificial intelligence. Other chief needs include agile, large-scale storage that can be quickly searched and queried for relevant data to support validation and model development, data-sharing agreements, and formatting standards for key data types. The future of public transit must be explored in greater detail, research must inform design, and opportunities must be identified for improving the mobility productivity of public transit in both urban and rural America.

33 ADVANCED PROPULSION SYSTEMS↗

3P Program: Phenotyping X Prediction = Productivity (Final Scientific/Technical Report)

The goal of the 3P Program was to establish integrated, real-time phenotyping and to analyze above- and below-ground plant architecture and total carbon partitioning and allocation to predict heterosis and develop superior crop hybrids by fully leveraging the Sorghum gene pool. There were two overarching themes: 1) the development of a new crop improvement approach utilizing advances in high-throughput phenotyping (HTP), computing, and genomics for public dissemination and 2) leveraging this platform for sorghum crop improvement and commercialization. The Clemson team worked on creating genomic resources and using both statistical learning and high-throughput phenotyping in genomics-assisted breeding. Research was broadly interested in the genetics of carbon partitioning, with the aim of improving crop performance and achieving sustainability. The technology and resources created can be readily found in the public domain and serve to advance scientific understanding of crop genomics and breeding. Genomic prediction was able to identify top crosses to be made, and a hybrid prediction pipeline is in place to drive year-over-year genetic gain. Roots have long been ignored by plant breeders and agronomists, not because they are unimportant but because they are hard to measure. This is an untapped white space of potential insight and innovation. To address this, Hi Fidelity Genetics developed the RootTracker to measure roots in the field on a continuous basis. A database system called RootTracker Tracker was developed to handle data coming from the RootTrackers. In using this device, valuable data was observed for plant breeding, hydrochemical development, and other agricultural biology applications. Carnegie Mellon’s goal was developing new techniques to generate high-resolution 3D models of plants from data collected in the field. The idea was that more useful and more informative phenotypes could be extracted by resolving small features, such as seeds and flowers, and that by modeling in 3D, the spatial structure of plants could be examined. To achieve this, multiple images collected by a new small format structured light stereo imager were fused together. A sorghum panicle modeling pipeline was developed to allow the collection and processing of data. Carolina Seed Systems is an agricultural technology company focused on decarbonizing the agricultural system. Their technology pipeline serves to drive fundamental progress towards creation and distribution of carbon negative crops. The genomic and the engineering technology developed through the 3P Program was leveraged to deliver both value and sustainability from the grower to the consumer. Promising sorghum hybrids were scaled up and commercialized. The overall goal of our research was to integrate, create, and deploy genetic and engineering concepts and technologies to enhance crop productivity in a sustainable fashion. The combination of public and private partners allowed the basic research and hypothesis testing to be quickly accelerated for commercial application by the companies yet maintained that the core framework and academic insights remain in the public domain for continued market disruption, competition, and innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Missing data in multi-omics integration: Recent advances through artificial intelligence

Biological systems function through complex interactions between various ‘omics (biomolecules), and a more complete understanding of these systems is only possible through an integrated, multi-omic perspective. This has presented the need for the development of integration approaches that are able to capture the complex, often non-linear, interactions that define these biological systems and are adapted to the challenges of combining the heterogenous data across ‘omic views. A principal challenge to multi-omic integration is missing data because all biomolecules are not measured in all samples. Due to either cost, instrument sensitivity, or other experimental factors, data for a biological sample may be missing for one or more ‘omic techologies. Recent methodological developments in artificial intelligence and statistical learning have greatly facilitated the analyses of multi-omics data, however many of these techniques assume access to completely observed data. A subset of these methods incorporate mechanisms for handling partially observed samples, and these methods are the focus of this review. We describe recently developed approaches, noting their primary use cases and highlighting each method's approach to handling missing data. We additionally provide an overview of the more traditional missing data workflows and their limitations; and we discuss potential avenues for further developments as well as how the missing data issue and its current solutions may generalize beyond the multi-omics context.

97 MATHEMATICS AND COMPUTING↗

Race-Specific Risk Factors for Homeownership Disparity in the Continental United States

The United States has a racial homeownership gap due to a legacy of historic inequality and discriminatory policies, but factors that contribute to the racial disparity in homeownership rates between White Americans and people of color have not been fully characterized. In order to alleviate this issue, policymakers need a better understanding of how risk factors affect the homeownership rates of racial and ethnic groups differently. In this study, data from several publicly available surveys, including the American Community Survey and United States Census, were leveraged in combination with statistical learning models to investigate potential factors related to homeownership rates across racial and ethnic categories, with a focus on how risk factors vary by race or ethnicity. Our models indicated that job availability for specific demographics, and specific regions of the United States were factors that affect homeownership rates in Black, Hispanic, and Asian populations in different ways. Based on the results of this study, it is recommended policymakers promote strategies to increase access to jobs for people of color (POC), such as vocational training and programs to reduce implicit bias in hiring practices. These interventions could ultimately increase homeownership rates for POC and be a step toward reducing the racial wealth gap.

99 GENERAL AND MISCELLANEOUS↗

Exploring for Superhot Geothermal Targets in Magmatic Settings: Developing a Methodology

This paper presents preliminary results from a subset of work carried out as part of a multinational research project entitled DErisking Exploration for multiple geothermal Plays in magmatic ENvironments (DEEPEN). One objective of DEEPEN is to develop a customized approach to exploration for superhot geothermal plays in magmatic systems. This paper summarizes key geologic components, risk factors, and exploration methods for geothermal plays in magmatic settings based on a review and comparative analysis of international training sites. As part of a Play Fairway Analysis (PFA) approach to exploring for multiple play types in a single magmatic system, training data were compiled and weights assigned to various evidence layers. Two different approaches for weighting exploration datasets are described in this paper - one based on expert opinions and the other using statistical learning. Weights produced by both approaches will be input into a 3D PFA workflow that combines multiple exploration datasets to generate 3D geothermal favorability models, which will be applied to two international demonstration sites.

GEOTHERMAL ENERGY↗

AI and ML Applications for PV Reliability and System Performance

This poster discusses AI and ML topics in PV reliability and system performance. In particular, automated metadata extraction and QA for fielded solar installations is covered for the PV Fleets Project. Additionally, statistical learning topics for the PVInsight Project are addressed, as well as development of the PV Validation Hub.

algorithm↗

1st Computational Physics School for Fusion Research (2019 CPS-FR)

The rising number of applications of machine learning and computational statistics in fusion energy research requires flexibility in adopting a growing variety of tools. The Computational Physics School for Fusion Research (CPS-FR) aims at providing young researchers with critical skill sets to deal with modern fusion energy research challenges. The School aims at covering essentials of: Computational Statistics, Machine Learning, Deep Learning and optimization methods, Parallel Programming and HPC. As the first edition of the CPS-FR just concluded, this report highlights its main results and summarizes its contents.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Statistical Complexity of Quantum Learning

Abstract Learning problems involve settings in which an algorithm has to make decisions based on data, and possibly side information such as expert knowledge. This study has two main goals. First, it reviews and generalizes different results on the data and model complexity of quantum learning, where the data and/or the algorithm can be quantum, focusing on information‐theoretic techniques. Second, it introduces the notion of copy complexity, which quantifies the number of copies of a quantum state required to achieve a target accuracy level. Copy complexity arises from the destructive nature of quantum measurements, which irreversibly alter the state to be processed, limiting the information that can be extracted about quantum data. As a result, empirical risk minimization is generally inapplicable. The paper presents novel results on the copy complexity for both training and testing. To make the paper self‐contained and approachable by different research communities, an extensive background material is provided on classical results from statistical learning theory, as well as on the distinguishability of quantum states. Throughout, the differences between quantum and classical learning are highlighted by addressing both supervised and unsupervised learning, and extensive pointers are provided to the literature.

97 MATHEMATICS AND COMPUTING↗

Statistically-informed deep learning for gravitational wave parameter estimation

We introduce deep learning models to estimate the masses of the binary components of black hole mergers, $(m_1,m_2)$, and three astrophysical properties of the post-merger compact remnant, namely, the final spin, $a_\mathrm f$, and the frequency and damping time of the ringdown oscillations of the fundamental $\ell = m = 2$ bar mode, $(\omega_\mathrm R, \omega_\mathrm I)$. Our neural networks combine a modified WaveNet architecture with contrastive learning and normalizing flow. We validate these models against a Gaussian conjugate prior family whose posterior distribution is described by a closed analytical expression. Upon confirming that our models produce statistically consistent results, we used them to estimate the astrophysical parameters $(m_1,m_2, a_\mathrm f, \omega_\mathrm R, \omega_\mathrm I)$ of five binary black holes: GW150914, GW170104, GW170814, GW190521 and GW190630. We use PyCBC Inference to directly compare traditional Bayesian methodologies for parameter estimation with our deep learning based posterior distributions. Our results show that our neural network models predict posterior distributions that encode physical correlations, and that our data-driven median results and 90% confidence intervals are similar to those produced with gravitational wave Bayesian analyses. This methodology requires a single V100 NVIDIA GPU to produce median values and posterior distributions within two milliseconds for each event. Furthermore, this neural network, and a tutorial for its use, are available at the Data and Learning Hub for Science.

79 ASTRONOMY AND ASTROPHYSICS↗