Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP)

Anthropogenic climate change is unfolding rapidly, yet its regional manifestation can be obscured by internal variability. A primary goal of climate science is to identify the externally forced climate response from among the noise of internal variability. Separating the forced response from internal variability can be addressed in climate models by using a large ensemble to average over different possible realizations of internal variability. However, with only one realization of the real world, it is a major challenge to isolate the forced response directly in observations. In the Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP), contributors used existing and newly developed statistical and machine learning methods to estimate the forced response over 1950–2022 within individual realizations of the climate system. Participants used neural networks, linear inverse models, fingerprinting methods, and low-frequency component analysis, among other approaches. These methods were trained using large ensembles from multiple climate models and then applied to observations. Here, we evaluate method performance within large ensembles and investigate the estimates of the forced response in observations. Our results show that many different types of methods are skillful for estimating the forced response in climate models, though the relative skill of individual methods varies depending on the variable and evaluation metric. Methods with comparable skill in models can give a wide range of estimates of the forced response pattern in observations, illustrating the epistemic uncertainty in forced response estimates. ForceSMIP gives new insights into the forced response in observations, its uncertainty, and methods for its estimation.

Climate attribution↗

Classification of events from α -induced reactions in the MUSIC detector via statistical and ML methods

The Multi-Sampling Ionization Chamber (MUSIC) detector is typically used to measure nuclear reaction cross sections relevant for nuclear astrophysics, fusion studies, and other applications. From the MUSIC data produced in one experiment scientists carefully extract an order of 10 3 events of interest from about 10 9 total events, where each event can be represented by an 18-dimensional vector. However, the standard data classification process is based on expert driven, manually intensive data analysis techniques that require several months to identify patterns and classify the relevant events from the collected data. Here, to address this issue, we present a method for the classification of events originating from specific α-induced reactions by combining statistical and machine learning methods that require significantly less input from the domain scientist, relative to the standard technique. Here, we applied the new method to two experimental data sets and compared our results with those obtained using traditional methods. With few exceptions, the number of events classified by our method agrees within ±20% with the results obtained using traditional methods. With the present method, which is the first of its kind for the MUSIC data, we have established the foundation for the automated extraction of physical events of interest from experiments using the MUSIC detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dynamical coupled-channel models for hadron dynamics

Dynamical coupled-channel (DCC) approaches parametrize the interactions and dynamics of two and more hadrons and their response to different electroweak probes. The inclusion of unitarity, three-body channels, and other properties from scattering theory allows for a reliable extraction of resonance spectra and their properties from data. Here, we review the formalism and application of the ANL-Osaka, the Juelich-Bonn-Washington, and other DCC approaches in the context of light baryon resonances from meson, (virtual) photon, and neutrino-induced reactions, as well as production reactions, strange baryons, light mesons, heavy meson systems, exotics, and baryon-baryon interactions. Finally, we also provide a connection of the formalism to study finite-volume spectra obtained in Lattice QCD, and review applications involving modern statistical and machine learning tools.

Amplitude analysis↗

Navigating the Expansive Landscapes of Soft Materials: A User Guide for High-Throughput Workflows

Synthetic polymers are highly customizable with tailored structures and functionality, yet this versatility generates challenges in the design of advanced materials due to the size and complexity of the design space. Thus, exploration and optimization of polymer properties using combinatorial libraries has become increasingly common, which requires careful selection of synthetic strategies, characterization techniques, and rapid processing workflows to obtain fundamental principles from these large data sets. Herein, we provide guidelines for strategic design of macromolecule libraries and workflows to efficiently navigate these high-dimensional design spaces. We describe synthetic methods for multiple library sizes and structures as well as characterization methods to rapidly generate data sets, including tools that can be adapted from biological workflows. We further highlight relevant insights from statistics and machine learning to aid in data featurization, representation, and analysis. This Perspective acts as a “user guide” for researchers interested in leveraging high-throughput screening toward the design of multifunctional polymers and predictive modeling of structure–property relationships in soft materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Distributed Acoustic Sensing to Estimate the Permeability

Optical fiber in a borehole can be interrogated with distributed acoustic sensors (DAS) to capture fracture displacements with the potential to map surrounding fracture networks. We designed a laboratory experiment to test the capability of DAS to determine borehole flow characteristics, and we show that for the first time DAS can be used to remotely estimate permeability. Optical fiber was wrapped around a bead filled pipe and the pressure drop and flow velocity were measured to directly calculate permeability. A machine learning model using statistical features from continuous DAS estimated the bulk permeability. Fluid interactions with the permeable material demonstrate insufficient resolution using DAS amplitude-based measurements for estimating pressure drop to infer permeability. Variations in the spectral domain relate DAS measurements to the pressure drop and provide consistent permeability estimates. Resolution with DAS is sufficient to estimate permeability and provides a reliable method to monitor at depth in borehole conditions.

58 GEOSCIENCES↗

Improving the Quasi‐Biennial Oscillation via a Surrogate‐Accelerated Multi‐Objective Optimization

Accurate simulation of the quasi-biennial oscillation (QBO) is challenging due to uncertainties in representing convectively generated gravity waves. We develop an end-to-end uncertainty quantification workflow that calibrates these gravity wave processes in E3SM for a realistic QBO. Central to our approach is a domain knowledge-informed, compressed representation of high-dimensional spatio-temporal wind fields. By employing a parsimonious statistical model that learns the fundamental frequency from complex observations, we extract interpretable and physically meaningful quantities capturing key attributes. Building on this, we train a probabilistic surrogate model that approximates the fundamental characteristics of the QBO as functions of critical physics parameters governing gravity wave generation. Leveraging the Karhunen–Loève decomposition, our surrogate efficiently represents these characteristics as a set of orthogonal features, capturing cross-correlations among multiple physics quantities evaluated at different pressure levels and enabling rapid surrogate-based inference at a fraction of the computational cost of full-scale simulations. Finally, we analyze the inverse problem using a multi-objective approach. Our study reveals a tension between amplitude and period that constrains the QBO representation, precluding a single optimal solution. To navigate this, we quantify the bi-criteria trade-off and generate a set of Pareto optimal parameter values that balance the conflicting objectives. This integrated workflow improves the fidelity of QBO simulations and offers a versatile template for uncertainty quantification in complex geophysical models.

54 ENVIRONMENTAL SCIENCES↗

Multi-fidelity information fusion with concatenated neural networks

Recently, computational modeling has shifted towards the use of statistical inference, deep learning, and other data-driven modeling frameworks. Although this shift in modeling holds promise in many applications like design optimization and real-time control by lowering the computational burden, training deep learning models needs a huge amount of data. This big data is not always available for scientific problems and leads to poorly generalizable data-driven models. This gap can be furnished by leveraging information from physics-based models. Exploiting prior knowledge about the problem at hand, this study puts forth a physics-guided machine learning (PGML) approach to build more tailored, effective, and efficient surrogate models. For our analysis, without losing its generalizability and modularity, we focus on the development of predictive models for laminar and turbulent boundary layer flows. In particular, we combine the self-similarity solution and power-law velocity profile (low-fidelity models) with the noisy data obtained either from experiments or computational fluid dynamics simulations (high-fidelity models) through a concatenated neural network. We illustrate how the knowledge from these simplified models results in reducing uncertainties associated with deep learning models applied to boundary layer flow prediction problems. The proposed multi-fidelity information fusion framework produces physically consistent models that attempt to achieve better generalization than data-driven models obtained purely based on data. While we demonstrate our framework for a problem relevant to fluid mechanics, its workflow and principles can be adopted for many scientific problems where empirical, analytical, or simplified models are prevalent. In line with grand demands in novel PGML principles, this work builds a bridge between extensive physics-based theories and data-driven modeling paradigms and paves the way for using hybrid physics and machine learning modeling approaches for next-generation digital twin technologies.

42 ENGINEERING↗

Bayesian stability and force modeling for uncertain machining processes

Accurately simulating machining operations requires knowledge of the cutting force model and system frequency response. However, this data is collected using specialized instruments in an ex-situ manner. Bayesian statistical methods instead learn the system parameters using cutting test data, but to date, these approaches have only considered milling stability. This paper presents a physics-based Bayesian framework which incorporates both spindle power and milling stability. Initial probabilistic descriptions of the system parameters are propagated through a set of physics functions to form probabilistic predictions about the milling process. The system parameters are then updated using automatically selected cutting tests to reduce parameter uncertainty and identify more productive cutting conditions, where spindle power measurements are used to learn the cutting force model. The framework is demonstrated through both numerical and experimental case studies. Results show that the approach accurately identifies both the system natural frequency and cutting force model.

42 ENGINEERING↗

Quantum machine learning for chemistry and physics

Machine learning (ML) has emerged as a formidable force for identifying hidden but pertinent patterns within a given data set with the objective of subsequent generation of automated predictive behavior. In recent years, it is safe to conclude that ML and its close cousin, deep learning (DL), have ushered in unprecedented developments in all areas of physical sciences, especially chemistry. Not only classical variants of ML, even those trainable on near-term quantum hardwares have been developed with promising outcomes. Such algorithms have revolutionized materials design and performance of photovoltaics, electronic structure calculations of ground and excited states of correlated matter, computation of force-fields and potential energy surfaces informing chemical reaction dynamics, reactivity inspired rational strategies of drug designing and even classification of phases of matter with accurate identification of emergent criticality. In this review we shall explicate a subset of such topics and delineate the contributions made by both classical and quantum computing enhanced machine learning algorithms over the past few years. We shall not only present a brief overview of the well-known techniques but also highlight their learning strategies using statistical physical insight. The objective of the review is not only to foster exposition of the aforesaid techniques but also to empower and promote cross-pollination among future research in all areas of chemistry which can benefit from ML and in turn can potentially accelerate the growth of such algorithms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An implementation of neural simulation-based inference for parameter estimation in ATLAS

Neural simulation-based inference (NSBI) is a powerful class of machine-learning-based methods for statistical inference that naturally handles high-dimensional parameter estimation without the need to bin data into low-dimensional summary histograms. Such methods are promising for a range of measurements, including at the Large Hadron Collider, where no single observable may be optimal to scan over the entire theoretical phase space under consideration, or where binning data into histograms could result in a loss of sensitivity. This work develops a NSBI framework for statistical inference, using neural networks to estimate probability density ratios, which enables the application to a full-scale analysis. It incorporates a large number of systematic uncertainties, quantifies the uncertainty due to the finite number of events in training samples, develops a method to construct confidence intervals, and demonstrates a series of intermediate diagnostic checks that can be performed to validate the robustness of the method. As an example, the power and feasibility of the method are assessed on simulated data for a simplified version of an off-shell Higgs boson couplings measurement in the four-lepton final states. This approach represents an extension to the standard statistical methodology used by the experiments at the Large Hadron Collider, and can benefit many physics analyses.

frequentist statistics↗

TransPlatformer

We propose TransPlatformer for translating toxicogenomics from one platform to another. Transcriptomic profiling has evolved through multiple generations of technology, from microarrays (e.g., Affymetrix, CodeLink) to more recent high-throughput sequencing and targeted panels such as S1500+. Microarrays, which dominated gene expression studies in the early 2000s, provided affordable and high-throughput transcript quantification but suffered from cross-hybridization issues and limited dynamic range . RNA-Seq, introduced in the late 2000s, revolutionized transcriptomics by enabling unbiased and comprehensive gene expression analysis, albeit at higher costs and computational demands . Despite advances, many studies rely on historical microarray data, necessitating the translation of legacy data into modern platforms to ensure continuity and comparability. This translation is complicated by factors such as platform-specific probe design, differences in transcript coverage, and batch effects . Existing methods for cross-platform mapping include statistical normalization, machine learning models, and biological anchoring approaches. The ability to translate transcriptomic data between platforms has broad implications, including enhanced meta-analyses, improved toxicological modeling, and better integration of historical datasets with contemporary research. TransPlatformer seeks to contribute to this effort by evaluating translation methodologies and proposing novel strategies to improve cross-platform gene expression harmonization. In this repository there are code examples for TransPlatformer implementation

Cong, Guojing↗

Data for Yield from Iowa’s first commercial miscanthus fields: implications of spatial variability for productivity and sustainability beyond research plots

This dataset contains biomass yield measurements and associated vegetation index data collected from commercial Miscanthus × giganteus fields in eastern Iowa during the 2022–2023 growing seasons. The data support the analyses presented in the article: “Yield From Iowa's First Commercial Miscanthus Fields: Implications of Spatial Variability for Productivity and Sustainability Beyond Research Plots.” We collected 105 ground-truth biomass samples from four mature commercial fields (>4 years old) covering 92.81 ha. Samples were taken from 3 m² quadrats that were hand-harvested in alignment with commercial harvest timing. Stem biomass (excluding leaves) was weighed, moisture-corrected, and converted to dry-matter yield expressed in Mg DM ha⁻¹. Sampling locations were selected to capture spatial variability visible in aerial imagery and were recorded using RTK GPS. Each biomass observation was paired with vegetation indices derived from high-resolution PlanetScope satellite imagery (3 m resolution). Images were acquired throughout the growing season, and indices were calculated to evaluate their ability to predict end-of-season biomass yield. Statistical and machine learning approaches were used to identify key predictors, and a linear regression model based on end-of-July Green Normalized Difference Vegetation Index (GNDVI) was developed and evaluated. This repository includes the data used in that modeling workflow. Management practices, economic data, full imagery time series, and additional methodological details are described in the associated publication and are not included here. The dataset consists of three comma-separated value (CSV) files: 1. Combine_Groundtruth_Yield_VI_22_23.csv This file contains ground-truth biomass yield measurements and associated key vegetation index values collected during the 2022 and 2023 growing seasons. Rows: 105 observations Columns: Year — Year of observation (2022 or 2023) Field — Field location identifier Sample_number — Unique sample identifier GNDVI_End_Jul — Green Normalized Difference Vegetation Index calculated at end of July GNDVI_End_Aug — Green Normalized Difference Vegetation Index calculated at end of August NDRE_End_Aug — Normalized Difference Red Edge index calculated at end of August Biomass_Stem_Yield_MgDM/ha — Measured stem biomass yield (megagrams dry matter per hectare) 2. trainData_GNDVI.csv This file contains the subset of observations used to train the predictive relationship between July GNDVI and biomass yield. Rows: 76 observations Columns: Unnamed: 0 — Row index retained from the original data processing workflow GNDVI_End_Jul — GNDVI at end of July Stem_Yield_MgDM/ha — Observed stem biomass yield (Mg DM ha⁻¹) 3. testData_GNDVI.csv This file contains the test dataset used to evaluate model performance. Rows: 29 observations Columns: Unnamed: 0 — Row index retained from the original data processing workflow GNDVI_End_Jul — GNDVI at end of July Predicted_Yield_MgDM/ha — Model-predicted stem biomass yield (Mg DM ha⁻¹) Observed_Yield_MgDM/ha — Measured stem biomass yield (Mg DM ha⁻¹)

Potential yield, yield gap, in-field management, y↗

A Provably Accurate Randomized Sampling Algorithm for Logistic Regression

In statistics and machine learning, logistic regression is a widely-used supervised learning technique primarily employed for binary classification tasks. When the number of observations greatly exceeds the number of predictor variables, we present a simple, randomized sampling-based algorithm for logistic regression problem that guarantees high-quality approximations to both the estimated probabilities and the overall discrepancy of the model. Our analysis builds upon two simple structural conditions that boil down to randomized matrix multiplication, a fundamental and well-understood primitive of randomized numerical linear algebra. We analyze the properties of estimated probabilities of logistic regression when leverage scores are used to sample observations, and prove that accurate approximations can be achieved with a sample whose size is much smaller than the total number of observations. To further validate our theoretical findings, we conduct comprehensive empirical evaluations. Overall, our work sheds light on the potential of using randomized sampling approaches to efficiently approximate the estimated probabilities in logistic regression, offering a practical and computationally efficient solution for large-scale datasets.

Chowdhury, Agniva↗

Ocean & Geohazard Analysis Tool

The Ocean & Geohazard Analysis (OGA) software tool is designed to summarize insights into key offshore hazards drawing from a diverse set of approaches, including artificial intelligence, machine learning, probabilistic and statistical, and offshore data sources. The offshore hazards that can be analyzed include submarine landslides, extreme wind/wave/current event probabilities, earthquakes, and metocean pathways (CIAM Climatological Isolation and Attraction Model–Climatological Lagrangian Coherent Structures - Submissions - EDX (doe.gov)). Currently, the tool is developed for use in the Gulf of Mexico. The data underlying the offshore hazard analyses can be found here: https://edx.netl.doe.gov/dataset/gulf-of-mexico-risk-analysis-database-gomrad This work was conducted under the Advanced Offshore Research Portfolio, FWP Number 1022409 at National Energy Technology Laboratory, U.S. Dept. of Energy. Disclaimer This project was funded by the United States Department of Energy, National Energy Technology Laboratory, in part, through a site support contract. Neither the United States Government nor any agency thereof, nor any of their employees, nor the support contractor, nor any of their employees, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof.

AIML↗

LCSL-II Cryomodule Testing at Fermilab

Cold powered testing of all LCLS-II production cryomodules at Fermilab is complete as of February 2021. A total of twenty-five tests on both 1.3 GHz and 3.9 GHz cryomodules were conducted over a nearly five year time span beginning in the summer of 2016. During the course of this campaign cutting-edge results for cavity Q₀ and gradient in continuous wave operation were achieved. A summary of all test results will be presented, with a comparison to established acceptance criteria, as well as overall test stand statistics and lessons learned.

43 PARTICLE ACCELERATORS↗

DeepBench: A simulation package for physical benchmarking data

We introduce **DeepBench**, a python library that generates simple simulated image data from first principles, such as basic geometric shapes and astronomical objects. These data are highly valuable for developing (calibration, testing, and benchmarking) statistical and machine learning models because they make it possible to connect the final data product to physically interpretable inputs. This software includes tools to curate and store the datasets to maximize reproducibility.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Coordinated Ramping Product and Regulation Reserve Procurements in CAISO and MISO using Multi-Scale Probabilistic Solar Power Forecasts (Pro2R)

How can probabilistic solar forecasts lower costs and improve reliability for independent system operator (ISO) markets? We tackle this question in three steps. First, we enhance an existing solar forecasting system to provide well-calibrated hours-ahead probabilistic forecasts. We then relate the degree of uncertainty in those forecasts to error distributions for net load ramps for the California ISO (CAISO) using statistical and machine learning methods. Projected net load errors conditioned on solar uncertainty are translated into flexible ramp requirements that therefore reflect real-time meteorological and solar conditions, improving on typical ISO procedures. Finally, a multi-period look-ahead production cost model quantifies how conditional ramp requirements can a) decrease operating costs by lowering requirements compared to often conservative unconditional methods, and b) reduce generation scarcity events and consequently improve reliability by increasing flexibility requirements at times when unconditional forecast-based requirements understate actual ramp uncertainty. In addition to the products just described (quantification of solar uncertainty, its translation into requirements for ramp capability product, and quantification of the benefits of more accurate ramp requirements), this project also developed a visualization system that alerts system operators of ramp and uncertainty conditions within the network based on solar forecasts. The system is called Resource Forecast and Ramp Visualization for Situational Awareness (RaVIS). These four products represent significant advances in the state-of-the-art of probabilistic solar forecasting, development of weather-informed reserve requirements, production costing methods for estimating the benefits of more accurate reserve requirements, and visualization of system status, respectively. Yet the products are also practical and can be immediately implemented, potentially enabling system operators to save millions of dollars in ramp product procurement costs per year.

14 SOLAR ENERGY↗

Technical Report on the Belle II Summer Workshop and Explorer Workshop 2023

The 2023 Belle II Summer workshop took place July 24-28, 2023 at Duke University in Durham. The meeting webpage can be found at https://indico.belle2.org/event/8841/. The meeting had 58 registered participants with the overwhelming majority attending in person. The event at giving beginning graduate students and postdocs an overview over the Belle II physics program and detector as well as an in-depth exploration of the Belle II software. For the latter, several hands-on sessions were organized to introduce participants to the Belle II software as well as more specialized topics in statistics and Machine Learning. The workshop also featured an ML/AI competition. In addition to the DOE support, the workshop was also supported by the Duke Physics department. We were able to host almost all students that so wished in the Duke Dorms and provide a meal plan. The support enabled us to waive the registration fee for all participants and cover also part of the dorm costs. Furthermore, we covered travel costs for external speakers on ML/AI topics. Figures 1 and 2 show the group picture and a scene from the hands-on sessions, respectively.

99 GENERAL AND MISCELLANEOUS↗