Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Optimal Transport as a Tool for Scientific Discovery in Radiation Biology

This report summarizes findings from research conducted for the “Exploration of the Poten tial for Artificial Intelligence and Machine Learning to Advance Low-Dose Radiation Biology Re search” (RadBio-AI) program, supported by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research, under Awards KP1601011/FWP CC121 and KP1601017/FWP CC121. The research reported here was undertaken in an effort to assess the potential of optimal measure transport methods as components within the larger scope of a com putational framework envisioned to support research in the radiation biology domain. Within this effort, our interest centered on enabling a unified generic framework where probabilistic modeling, inference, and statistical learning can be carried out for a wide range of data distributions. As described next in Section 1 (and in more detail in our original publication), optimal measure transport offers the possibility of such unified approach.

97 MATHEMATICS AND COMPUTING↗

Correlative piezoresponse and micro-Raman imaging of CuInP 2 S 6 –In 4/3 P 2 S 6 flakes unravels phase-specific phononic fingerprint via unsupervised learning

Characterizing the novel properties of layered van der Waals materials is key for their application in functional devices. A better understanding of this type of material requires correlative imaging of diverse nanoscale material properties. Within this class of materials, CuInP 2 S 6 (CIPS) has received a significant degree of interest due to its ionically mediated room temperature ferroelectricity. Moreover, it is possible to form stable self-assembled heterostructures of ferroelectric CuInP 2 S 6 (CIPS) and non-ferroelectric (i.e., lacking Cu) In 4/3 P 2 S 6 (IPS) phases, by controlling the targeted composition and kinetics of synthesis. In this work, we present a correlative nanometric imaging study of the phononic modes and piezoelectricity of the phase-separated thin heteroepitaxial CIPS/IPS flakes. Here, we show that it is possible to isolate the different phononic modes of the two phases by spatially correlating them with their distinct ferroelectric behavior. The coupling of our experimental data with unsupervised learning statistical methods enables unraveling specific Raman peaks that are characteristic of each chemical phase (CIPS and IPS) present in the composite sample, discarding the less significant ones.

correlative microscopy↗

Report on the AAPM grand challenge on deep generative modeling for learning medical image statistics

Abstract Background The findings of the 2023 AAPM Grand Challenge on Deep Generative Modeling for Learning Medical Image Statistics are reported in this Special Report. Purpose The goal of this challenge was to promote the development of deep generative models for medical imaging and to emphasize the need for their domain‐relevant assessments via the analysis of relevant image statistics. Methods As part of this Grand Challenge, a common training dataset and an evaluation procedure was developed for benchmarking deep generative models for medical image synthesis. To create the training dataset, an established 3D virtual breast phantom was adapted. The resulting dataset comprised about 108 000 images of size 512 512. For the evaluation of submissions to the Challenge, an ensemble of 10 000 DGM‐generated images from each submission was employed. The evaluation procedure consisted of two stages. In the first stage, a preliminary check for memorization and image quality (via the Fréchet Inception Distance [FID]) was performed. Submissions that passed the first stage were then evaluated for the reproducibility of image statistics corresponding to several feature families including texture, morphology, image moments, fractal statistics, and skeleton statistics. A summary measure in this feature space was employed to rank the submissions. Additional analyses of submissions was performed to assess DGM performance specific to individual feature families, the four classes in the training data, and also to identify various artifacts. Results Fifty‐eight submissions from 12 unique users were received for this Challenge. Out of these 12 submissions, 9 submissions passed the first stage of evaluation and were eligible for ranking. The top‐ranked submission employed a conditional latent diffusion model, whereas the joint runners‐up employed a generative adversarial network, followed by another network for image superresolution. In general, we observed that the overall ranking of the top 9 submissions according to our evaluation method (i) did not match the FID‐based ranking, and (ii) differed with respect to individual feature families. Another important finding from our additional analyses was that different DGMs demonstrated similar kinds of artifacts. Conclusions This Grand Challenge highlighted the need for domain‐specific evaluation to further DGM design as well as deployment. It also demonstrated that the specification of a DGM may differ depending on its intended use.

Radiology, Nuclear Medicine & Medical Imaging↗

DEEPEN 3D PFA Weights for Exploration Datasets in Magmatic Environments

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This GDR submission includes those weights. The weighting was done using two different approaches: one based on expert opinions, and one based on statistical learning. The weights are intended to describe how useful a particular exploration method is for imaging each component of each play type. They may be adjusted based on the characteristics of the resource under investigation, knowledge of the quality of the dataset, or simply to reduce the impact a single dataset has on the resulting outputs. Within the DEEPEN PFA, separate sets of weights are produced for each component of each play type, since exploration methods hold different levels of importance for detecting each play component, within each play type. The weights for conventional hydrothermal systems were based on the average of the normalized weights used in the DOE-funded PFA projects that were focused on magmatic plays. This decision was made because conventional hydrothermal plays are already well-studied and understood, and therefore it is logical to use existing weights where possible. In contrast, a true PFA has never been applied to superhot EGS or supercritical plays, meaning that exploration methods have never been weighted in terms of their utility in imaging the components of these plays. To produce weights for superhot EGS and supercritical plays, two different approaches were used: one based on expert opinion and the analytical hierarchy process (AHP), and another using a statistical approach based on principal component analysis (PCA). The weights are intended to provide standardized sets of weights for each play type in all magmatic geothermal systems. Two different approaches were used to investigate whether a more data-centric approach might allow new insights into the datasets, and also to analyze how different weighting approaches impact the outcomes. The expert/AHP approach involved using an online tool (https://bpmsg.com/ahp/) with built-in forms to make pairwise comparisons which are used to rank exploration methods against one-another. The inputs are then combined in a quantitative way, ultimately producing a set of consensus-based weights. To minimize the burden on each individual participant, the forms were completed in group discussions. While the group setting means that there is potential for some opinions to outweigh others, it also provides a venue for conversation to take place, in theory leading the group to a more robust consensus then what can be achieved on an individual basis. This exercise was done with two separate groups: one consisting of U.S.-based experts, and one consisting of Iceland-based experts in magmatic geothermal systems. The two sets of weights were then averaged to produce what we will from here on refer to as the "expert opinion-based weights," or "expert weights" for short. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. More information on this approach along with the dataset used to produce the statistical weights may be found in the linked dataset below.

15 GEOTHERMAL ENERGY↗

Artificial Intelligence/Machine Learning Technologies for Advanced Reactors (Workshop Summary Report)

A workshop on artificial intelligence and machine learning (AI/ML) for advanced reactors (AR) was held October 5-6, 2021. The workshop was to be attended in-person at ANL but COVID restrictions forced the workshop to go virtual. The objectives of the workshop were to identify the most promising AI/ML opportunities for improving advanced reactor design, optimizing plant performance, and enhancing economic competitiveness and to develop an understanding of the scientific, engineering and licensing challenges facing their application. The workshop planning committee included GAIN, EPRI and NEI and members of three national laboratories (ANL, INL, and ORNL). The workshop was attended by more than 200 individuals representing academic and scientific institutions and the nuclear power industry. The definition put forth for an AI/ML system was one that perceives its environment and takes actions that maximize its chance of achieving its goals. In this report AI/ML refers to next generation algorithms that include deep learning, statistical analysis and data analytics and associated scientific computing and their potential application to the design, licensing, operation and maintenance of ARs. These methods typically incorporate models built from process data and may also include data generated by simulations that represent the behavior of a system. The workshop was organized in response to the growing interest in application of AI/ML for improving the economic competitiveness of nuclear energy. Increasingly more resources are being allocated to investigating the benefits of AI/ML methods. The DOE created the Artificial Intelligence & Technology Office to promote their development. And within the Office of Nuclear Energy, resources have been allocated to explore and understand the potential benefits of AI/ML. Additionally, the national laboratories are strategically positioned with DOE computing facilities such as Summit, Perlmutter, Aurora and Frontier that support large-scale simulations, hybrid HPC models with AI surrogates, and the exploration of new types of generative models emerging from multi-model data streams and sources. The workshop was organized with members of the AR community to understand the effort and to identify the level of interest and progress in this emerging technology. The workshop discussions focused on identifying opportunities for AI/ML across diverse areas of the nuclear industry and identifying current scientific and engineering challenges for advanced reactors that might be addressed through transformational uses of AI/ML. Discussion panels focused on four high-interest technical domains for advanced reactors: design, maintenance and operations, energy storage, and materials. The results of those discussions are summarized in this report. This includes opportunities that were identified for exploiting AI techniques and methods to improve the efficacy and efficiency of reactor analysis and to improve the operation and optimization of advanced reactors. Advanced reactor developers expressed an interest in learning more about AI/ML methods and their application. This included understanding whether ML methods can provide an advantage over existing nonlinear data regression methods for collapsing high-fidelity simulation results into faster running models. A consensus emerged that AR advances planned for the next decade will benefit from the use of AI/ML tools. The need exists to understand and model complex systems across length scales and modalities. AI/ML is a tool for discovery that can yield a set of engineering principles for use by nuclear engineers, licensing bodies, and operators to solve problems in plant design, safety analyses, autonomous operation, and predictive maintenance. While AI/ML represents a new set of tools, an awareness by the nuclear community of the full potential is still in the early stages so there is a need to increase awareness. It appears that the wide-spread adoption of AI/ML tools for ARs would be facilitated by future educational workshops that describe foundational methods and capabilities and describe successful applications.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Benchmarking FFTF LOFWOS Test# 13 using SAM code: Baseline model development and uncertainty quantification

The development and deployment of advanced reactors, such as the sodium-cooled fast reactor (SFR), relies on sophisticated modeling tools to ensure the safety of the design under various transients. The predictive capability of these advanced modeling tools requires validation to garner trust in supporting the licensing of the advanced reactors. For this reason, the International Atomic Energy Agency (IAEA) initiated a coordinated research project (CRP) in 2018 for the analysis of the Fast Flux Test Facility (FFTF) Loss of Flow Without Scram (LOFWOS) Test #13.In this study, we present and discuss the benchmarking efforts of the modern system code SAM on the FFTF LOFWOS Test #13. Further, the SAM baseline model was developed according to the benchmark specification, which included a detailed core model with reactivity feedback. Generally, good agreement was observed between the baseline results and benchmark measurements; however, discrepancies persisted, particularly in predicted fuel assembly coolant outlet temperatures. Utilizing the baseline model, uncertainty quantification (UQ) and sensitivity analysis (SA) were conducted with the assistance of various statistical learning and machine learning methods, including kernel density estimation, Gaussian processes, and Sobol indices. Following the baseline model prediction and UQ and SA results, we discuss the reasons for the simulation discrepancies and propose further improvements to the model. This benchmarking effort adheres to the best-estimate plus uncertainty approach and can serve as a valuable example for supporting risk-informed licensing of advanced reactors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine learning in materials science: From explainable predictions to autonomous design

The advent of big data and algorithmic developments in the field of machine learning (and artificial intelligence, in general) have greatly impacted the entire spectrum of physical sciences, including materials science. Materials data, measured or computed, combined with various techniques of machine learning have been employed to address a myriad of challenging problems, such as, development of efficient and predictive surrogate models for a range of materials properties, screening and down-selection of novel candidate materials for targeted applications, new methodologies to improve and further expedite molecular and atomistic simulations, with likely many more important developments to come in the foreseeable future. While the applications thus far have provided a glimpse of the true potential data-enabled routes have to offer, it has also become clear that further progress in this direction hinges on our ability to understand, explain and rationalize findings of a machine learning model in light of the domain-knowledge. This focused review provides an overview of the main areas where machine learning has been widely and successfully used in materials science. Subsequently, a brief discussion of several techniques that have been helpful in extracting physically-meaningful insights, causal relationships and design-centric knowledge from materials data is provided. Finally, we identify some of the imminent opportunities and challenges that materials community faces in this exciting and rapidly growing field.

36 MATERIALS SCIENCE↗

Database-wide hazard modelling of the onset of DIII-D tearing modes with field features

The rate of onset (hazard) of tearing modes is modelled probabilistically using statistical learning algorithms. Axisymmetric energy-density equilibrium fields are taken as raw high-dimensional input features which are reduced with principal component analysis. Signal processing of non-axisymmetric magnetics fluctuation array data provides the target information from which to learn. Model selection, visualization and calibration assessment procedures are detailed. Here, the analysis is deployed at large scale across the DIII-D tokamak database. Standard model selection criteria suggest that the energy-density post-processed feature is a better choice for modelling the onset rate compared to the non-processed equilibrium reconstruction solution. Two example applications of the learned rate function are demonstrated: (i) proximity-to-onset discharge monitoring and (ii) database analysis showing an (expected) observational global trend that the general hazard increases as a plasma performance metric increases. An important connection between the hazard function and its use as a conditional probability generator is reviewed in the Appendix.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Increasing sensitivity of dryland vegetation greenness to precipitation due to rising atmospheric CO 2

Water availability plays a critical role in shaping terrestrial ecosystems, particularly in low- and mid-latitude regions. The sensitivity of vegetation growth to precipitation strongly regulates global vegetation dynamics and their responses to drought, yet sensitivity changes in response to climate change remain poorly understood. Here we use long-term satellite observations combined with a dynamic statistical learning approach to examine changes in the sensitivity of vegetation greenness to precipitation over the past four decades. We observe a robust increase in precipitation sensitivity (0.624% yr –1 ) for drylands, and a decrease (–0.618% yr –1 ) for wet regions. Using model simulations, we show that the contrasting trends between dry and wet regions are caused by elevated atmospheric CO 2 (eCO 2 ). eCO 2 universally decreases the precipitation sensitivity by reducing leaf-level transpiration, particularly in wet regions. However, in drylands, this leaf-level transpiration reduction is overridden at the canopy scale by a large proportional increase in leaf area. The increased sensitivity for global drylands implies a potential decrease in ecosystem stability and greater impacts of droughts in these vulnerable ecosystems under continued global change.

54 ENVIRONMENTAL SCIENCES↗

Exacerbated drought impacts on global ecosystems due to structural overshoot

Vegetation dynamics are affected not only by the concurrent climate but also by memory-induced lagged responses. For example, favourable climate in the past could stimulate vegetation growth to surpass the ecosystem carrying capacity, leaving an ecosystem vulnerable to climate stresses. This phenomenon, known as structural overshoot, could potentially contribute to worldwide drought stress and forest mortality but the magnitude of the impact is poorly known due to the dynamic nature of overshoot and complex influencing timescales. Here, we use a dynamic statistical learning approach to identify and characterize ecosystem structural overshoot globally and quantify the associated drought impacts. We find that structural overshoot contributed to around 11% of drought events during 1981-2015 and is often associated with compound extreme drought and heat, causing faster vegetation declines and greater drought impacts compared to non-overshoot related droughts. The fraction of droughts related to overshoot is strongly related to mean annual temperature, with biodiversity, aridity and land cover as secondary factors. These results highlight the large role vegetation dynamics play in drought development and suggest that soil water depletion due to warming-induced future increases in vegetation could cause more frequent and stronger overshoot droughts.

54 ENVIRONMENTAL SCIENCES↗

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We find that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no significant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption.

Oles, Vlad↗

Automated CT registration, segmentation, and quantification (AutoCT) v1.0

Processing and analyzing brain imaging is crucial in both scientific development and clinical field. In this software package, we build a pipeline that integrates automatic registration, segmentation, and quantitative analysis for subjects' CT scans. Leveraging diffeomorphic transformations, we enable optimized forward and inverse mappings between an image and the reference. Furthermore, we extract localized features from deformation field based on an online template process, which advances statistical learning downstream. The created templates, atlas as well as our methods provide the brain imaging community tools for AI implementations.

Essiari, Abdelilah↗

Automated CT registration, segmentation, and quantification (AutoCT) v1.1

Processing and analyzing brain imaging is crucial in both scientific development and clinical field. In this software package, we build a pipeline that integrates automatic registration, segmentation, and quantitative analysis for subjects' CT scans. Leveraging diffeomorphic transofrmations, we enable optimized forward and inverse mappings between an image and the reference. Furthermore, we extract localized features from deformation field based on an online template process, which advances statistical learning downstream. The created templates, atlas as well as our methods provide the brain imaging community tools for AI implementations

Bai, Zhe↗

ForceFinder

SAND2025-11750O ForceFinder extends the Structural Dynamics Python Libraries (SDynPy) with comprehensive tools for inverse source estimation (ISE) tasks via frequency response function (FRF) matrix inversion. The software is designed for transfer path analysis and multiple-input/multiple-output (MIMO) vibration control problems. It allows users to estimate sources through various algorithms, from the basic Moore-Penrose pseudo-inverse to statistical learning methods such as Tikhonov regularization via an L-curve and elastic net regularization via an information criterion. ForceFinder uses an object-oriented framework, where all components of the ISE problem—such as FRFs, responses, and transformations—are stored in a "SourcePathReceiver" object. This software can be applied to any noise and vibration problem. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Carter, Steven [Sandia National Lab. (SNL-CA), Liv↗

Investigating Aerosol and Meteorological Influences on Convective Clouds in Houston, Texas, during the TRACER/ESCAPE Field Campaigns

Aerosols serve as cloud condensation nuclei, shaping the microphysical properties of cloud droplets. Aerosol effects on convective clouds are complex and remain controversial. The debate centers around the process of aerosol-induced invigoration of deep convection, a phenomenon that could significantly affect convective cloud properties but lacks robust evidence due to methodological limitations in observational approaches and questions about the robustness of modeling studies. Resolving these discrepancies is crucial for understanding how aerosols affect the atmosphere. Here, this study examines the effects of meteorological and aerosol parameters in a weakly synoptic-driven convective environment, where the influence of aerosols may be more pronounced and observable. Daily atmospheric soundings and aerosol concentrations from several ground instruments collected during the summer of 2022 in Houston, Texas, as part of the Tracking Aerosol Convection interactions Experiment (TRACER) and Experiment of Sea Breeze Convection, Aerosols, Precipitation, and Environment (ESCAPE) field campaigns are analyzed. Statistical learning methods are applied to uncover the complex relationships between aerosols, meteorology, and convective cloud characteristics, such as cell area and echo-top height. The findings reveal that higher aerosol concentrations are associated with narrower convective cells, which we argue contradicts the idea of stronger convection with increased aerosol loading. However, once the data are clustered by the synoptic environment, the relationship between aerosol loading and convective cell area diminishes, indicating that the covariablity between synoptic-scale weather patterns, local thermodynamics, and aerosol loading makes it challenging to draw definitive conclusions about the specific impacts of aerosols on convective cloud properties.

54 ENVIRONMENTAL SCIENCES↗

Large-Scale Trajectory Analysis via Feature Vectors

The explosion of both sensors and GPS-enabled devices has resulted in position/time data being the next big frontier for data analytics. However, many of the problems associated with large numbers of trajectories do not necessarily have an analog with many of the historic big-data applications such as text and image analysis. Modern trajectory analytics exploits much of the cutting-edge research in machine-learning, statistics, computational geometry and other disciplines. We will show that for doing trajectory analytics at scale, it is necessary to fundamentally change the way the information is represented through a feature-vector approach. We then demonstrate the ability to solve large trajectory analytics problems using this representation.

58 GEOSCIENCES↗

Advanced Computing, Data Science, and Artificial Intelligence Research Opportunities for Energy-Focused Transportation Science

The Energy Efficient Mobility Systems (EEMS) technology landscape is complex and rapidly evolving, which provides both tremendous opportunities and formidable challenges. Significant alterations to the mobility landscape are underway due to the advent of vehicle and infrastructure connectivity, autonomous driving, and rapid passenger- and freight-vehicle electrification. Advanced computing will play an increasingly important role in enabling the EEMS program to understand and identify the most important levers to improve the energy productivity of future integrated mobility systems. It is also driving new approaches to mobility and the research to unlock an affordable, efficient, safe, and accessible transportation future. Driving much of this change is the collection, analysis, and strategic use of massive amounts of diverse, complex data from infrastructure and vehicles with on-board sensors and data storage and transmission capabilities. Diverse and representative data are key to implementing approaches to maximize mobility energy productivity. While high-fidelity modeling of integrated transportation networks has strengthened our understanding of dynamic movement and behavior patterns, existing tools must be expanded beyond their current focus. This work necessitates data infrastructure investments (e.g., secure-streaming data platforms driven by ubiquitous sensors and video analytics) as well as investments in critical capabilities for large-scale automated analysis and organization using modern machine learning, statistics, and artificial intelligence. Other chief needs include agile, large-scale storage that can be quickly searched and queried for relevant data to support validation and model development, data-sharing agreements, and formatting standards for key data types. The future of public transit must be explored in greater detail, research must inform design, and opportunities must be identified for improving the mobility productivity of public transit in both urban and rural America.

33 ADVANCED PROPULSION SYSTEMS↗