Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “correspondence learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Machine learning inversion from scattering for mechanically driven polymers

A machine learning inversion method is developed for analyzing scattering functions of mechanically driven polymers and extracting the corresponding feature parameters, which include energy parameters and conformation variables. The polymer is modeled as a chain of fixed-length bonds constrained by bending energy, and it is subject to external forces such as stretching and shear. We generate a data set consisting of random combinations of energy parameters, including bending modulus, stretching and shear force, along with Monte Carlo-calculated scattering functions and conformation variables such as end-to-end distance, radius of gyration and off-diagonal component of the gyration tensor. The effects of the energy parameters on the polymer are captured by the scattering function, and principal component analysis ensures the feasibility of the machine learning inversion. Finally, we train a Gaussian process regressor using part of the data set as a training set and validate the trained regressor for inversion using the rest of the data. The regressor successfully extracts the feature parameters.

Gaussian process regressors↗

Machine learning BPS spectra and the gap conjecture

We explore statistical properties of Bogomol’nyi-Prasad-Sommerfield q-series for strongly coupled supersymmetric theories that correspond to a particular family of three-manifolds. We discover that gaps between exponents in the -series are statistically more significant at the beginning of the -series compared to gaps that appear in higher powers of. Our observations are obtained by calculating saliencies of -series features used as input data for principal component analysis, which is a standard example of an explainable machine learning technique that allows for a direct calculation and a better analysis of feature saliencies.

97 MATHEMATICS AND COMPUTING↗

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection Description This dataset contains input and output data for the manuscript Mongird, K. et al. (under review) titled "Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection". Input data corresponds to gridded spatial siting attributes that are necessary to conduct a random forest machine learning analysis of siting feature importance. Output data includes SHAP feature analysis outputs, and classification report values. For data on power plant siting results referred to in the manuscript, please refer to the CERF: IM3 Projected Western US Power Plant Locations data download page. The downloadable data includes values for eight different future scenarios for the Western US. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 Technical Information The dataset includes two sets of data files: (1) CERF gridded siting parameters and (2) Feature analysis outputs and classification reports. All downloadable data is in csv file format. Files with x/y coordinate information use the Albers Equal Area Conic projection (ESRI:102003). 1. CERF Gridded Siting Parameters This directory provides a balanced sample of gridded CERF siting parameters data for eight different scenarios for the Western US through 2055, seven different technologies, and eight timesteps. This data serves as input to the feature analysis. It contains the following parameters. region_name - name of region (i.e., state) sited - binary value representing whether the grid cell received a siting of that technology type (1=True) rcp - binary value representing scenario resource concentration pathway (0 = RCP4.5, 1 = RCP8.5) ssp - binary value representing scenario shared socioeconomic pathway (0 = SSP3, 1 = SSP5) climate - binary value representing cooler (0) or hotter (1) GCM forcing tech_name - generation technology name sited_year - year that values correspond to transmission_cost - cost of transmission interconnection pipeline_cost - cost of natural gas pipeline interconnection interconnection_cost - total interconnection cost (sum of transmission cost and gas pipeline cost) lmp - associated locational marginal value ($/MWh) associated with the grid cell, timestep, scenario, and technology xcoord - x-coordinate of location ycoord - y-coordinate of location 2a. Feature Analysis Output The dataset includes the feature analysis shap output for locational marginal price and interconnection cost. It contains the following parameters. technology - generator technology name scenario - name of scenario feature - name of feature, either locational_marginal_price or interconnection_cost value - the mean of absolute value of SHAP values for given feature 2b. Feature Analysis Classification Report This download includes the classification report associated with each random forest model. The dataset contains the following parameters. technology - generation technology name scenario - name of scenario test - one of precision (the proportion of predicted positives that are actually correct), recall (the proportion of actual positives that were correctly identified), f1-score (the harmonic mean of precision and recall) 0.0 - value of test for classification of 0 (grid cell not chosen for siting) 1.0 - value of test for classification of 1 (grid cell chosen for siting) accuracy - accuracy of model (i.e., fraction of all predictions that were right) macro avg - Simple average of test values for all classes weighted avg - Weighted average of test values for all classes, weighted based on Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

Simulated wildfire burned area over the CONUS during 2001-2020

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM). A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

Liu, Ye↗

Improving North American Wildfire Prediction by Integrating a Machine-Learning Fire Model in a Land Surface Model

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM) version 2.1. A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

54 ENVIRONMENTAL SCIENCES↗

Uncertainty Quantification and Sensitivity Analysis of Low-Dimensional Manifold via Co-Kurtosis PCA in Combustion Modeling

For multi-scale multi-physics applications e.g., the turbulent combustion code Pele, robust and accurate dimensionality reduction is crucial to solving problems at exascale and beyond. A recently developed technique, Co-Kurtosis based Principal Component Analysis (CoK-PCA) which leverages principal vectors of co-kurtosis, is a promising alternative to traditional PCA for complex chemical systems. To improve the effectiveness of this approach, we employ Artificial Neural Networks for reconstructing thermo-chemical scalars, species production rates, and overall heat release rates corresponding to the full state space. Our focus is on bolstering confidence in this deep learning based non-linear reconstruction through Uncertainty Quantification (UQ) and Sensitivity Analysis (SA). UQ involves quantifying uncertainties in inputs and outputs, while SA identifies influential inputs. One of the noteworthy challenges is the computational expense inherent in both endeavors. To address this, we employ the Monte Carlo methods to effectively quantify and propagate uncertainties in our reduced spaces while managing computational demands. Our research carries profound implications not only for the realm of combustion modeling but also for a broader audience in UQ. By showcasing the reliability and robustness of CoK-PCA in dimensionality reduction and deep learning predictions, we empower researchers and decision-makers to navigate complex combustion systems with greater confidence.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ionic Interdiffusion at Cathode|Solid-Electrolyte Interface: A Machine Learning–Assisted Multiscale Investigation and Mitigation Strategies

Future lithium batteries are expected to use solid electrolytes to achieve higher energy density and fast charge capabilities. However, most solid electrolytes are thermodynamically unstable against layered oxide cathodes. In this study, the stability of LiCoO2 (LCO) cathode with Li10GeP2S12 (LGPS) solid electrolyte is investigated using ab initio molecular dynamics (AIMD) and machine learning molecular dynamics (MLMD). The propensity of ionic interdiffusion, formation of a passivating interphase layer, and corresponding decay in cell performance is addressed using a continuum model. Large-scale MLMD simulations confirm that the LCO|LGPS interface permits interdiffusion of cobalt (Co) and other ionic species, leading to the formation and growth of a resistive interphase and to dramatic capacity fade even in the first cycle. We examine the literature evidence that incorporating a thin layer of LiNb0.5Ta0.5O3 (LNTO) between LCO and LGPS prevents the interdiffusion of ions. Atomistic simulations suggest that substituting lithium (Li) in LNTO with Co is thermodynamically unfavorable, thereby inhibiting ionic interdiffusion. The stable Nb5+/Ta5+ states form a rigid metal-oxide framework, which consequently also prevents the substitution of niobium (Nb) or tantalum (Ta). However, continuum-level analysis suggests that the higher mechanical stiffness of LNTO can lead to interfacial delamination between the LCO and LNTO. This phenomenon reduces the effectiveness of the protective layer. This paper, therefore, highlights the need to develop novel interlayers that balance low ionic interdiffusion with low mechanical stiffness.

Ncube, Musawenkosi K.↗

Utilizing Experimental-based Line Positions for Semi-empirical IR Line Lists

Molecular IR line lists computed from semi-empirically refined ab initio potential energy surface and high quality ab initio dipole moment surface may have line position accuracy typically within 0.01 - 0.05 cm−1. For high resolution spectroscopy databases, it is necessary to integrate the computed theoretical IR intensities with the line positions accurately determined from experiments and (or) reliably derived from Effective Hamiltonian models. However, it is not a trivial task, and the choices in reality are heavily contingent upon the coverage, consistency, and accuracy of the data available for a specific molecule or isotopologue. We will present several approaches applied to recent Ames IR line lists of CO 2 1, N 2 O 2 , and OCS to demonstrate how challenging such integration may become, and what we have learned From these projects. Discussions will focus on the relation between each different scenario and the corresponding choice or solution, including their advantages and limits, how the semi-empirical IR line lists can help, and what else may be needed for future improvements. We will emphasize that the “Theory+Experiment” synergy may still play significant role in the determination of the best line positions.

Effective Hamiltonian models↗

Clustering of tethered satellite system simulation data by an adaptive neuro-fuzzy algorithm

Recent developments in neuro-fuzzy systems indicate that the concepts of adaptive pattern recognition, when used to identify appropriate control actions corresponding to clusters of patterns representing system states in dynamic nonlinear control systems, may result in innovative designs. A modular, unsupervised neural network architecture, in which fuzzy learning rules have been embedded is used for on-line identification of similar states. The architecture and control rules involved in Adaptive Fuzzy Leader Clustering (AFLC) allow this system to be incorporated in control systems for identification of system states corresponding to specific control actions. We have used this algorithm to cluster the simulation data of Tethered Satellite System (TSS) to estimate the range of delta voltages necessary to maintain the desired length rate of the tether. The AFLC algorithm is capable of on-line estimation of the appropriate control voltages from the corresponding length error and length rate error without a priori knowledge of their membership functions and familarity with the behavior of the Tethered Satellite System.

Mitra, Sunanda↗

Human factors aspects of control room design

A plan for the design and analysis of a multistation control room is reviewed. It is found that acceptance of the computer based information system by the uses in the control room is mandatory for mission and system success. Criteria to improve computer/user interface include: match of system input/output with user; reliability, compatibility and maintainability; easy to learn and little training needed; self descriptive system; system under user control; transparent language, format and organization; corresponds to user expectations; adaptable to user experience level; fault tolerant; dialog capability user communications needs reflected in flexibility, complexity, power and information load; integrated system; and documentation.

Jenkins, J. P.↗

A variational framework for residual-based adaptivity in neural PDE solvers and operator learning

Residual-based adaptive strategies are widely used in scientific machine learning yet remain largely heuristic. We introduce a variational framework that formalizes these methods through convex transformations of the residual, where different transformations correspond to distinct objective functionals. For instance, exponential weights target uniform error minimization, while linear weights recover quadratic error minimization. This perspective reveals adaptive weighting as a means of selecting sampling distributions that optimize a primal objective, directly linking discretization choices to error metrics. This principled approach yields three key benefits: it enables systematic design of adaptive schemes, reduces discretization error by lowering estimator variance, and enhances learning dynamics by improving gradient signal-to-noise ratio. Extending the framework to operator learning, we demonstrate substantial performance gains across diverse optimizers and architectures. Our results provide a theoretical perspective for residual-based adaptivity and establish a foundation for principled discretization and training.

97 MATHEMATICS AND COMPUTING↗

Residual symmetries and scalar multiplet vacuum alignment in non-Abelian flavour models

We demonstrate that, upon minimizing a renormalizable, single-scalar potential invariant under a non-Abelian symmetry, special orientations in the associated vacuum alignment of the scalar multiplet correspond to the preservation of a discrete residual flavour symmetry in the broken phase of the theory. Conversely, we show that these special scalar alignments are perturbed when additional Lagrangian operators (e.g. renormalizable, multi-flavon operators and/or effective, higher-dimensional operators) are present that break said residual symmetry, leading to a vacuum reorientation and phenomenological consequences. We therefore construct a one-to-one correspondence principle between broken residual symmetries and vacuum alignment corrections, providing a mechanism to identify (and correct) a subtle but persistent form of phenomenologically relevant fine-tuning embedded in — but often ignored by — most successful non-Abelian flavour models. We first establish this correspondence in a set of toy models based on the S4 permutation symmetry, and then apply the lessons learned to the more realistic A4 Altarelli-Feruglio and ∆(27) Universal Texture Zero models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Neural Networks to Find the Optimal Forcing for Offsetting the Anthropogenic Climate Change Effects

Abstract Of great relevance to climate engineering is the systematic relationship between the radiative forcing to the climate system and the response of the system, a relationship often represented by the linear response function (LRF) of the system. However, estimating the LRF often becomes an ill-posed inverse problem due to high-dimensionality and nonunique relationships between the forcing and response. Recent advances in machine learning make it possible to address the ill-posed inverse problem through regularization and sparse system fitting. Here, we develop a convolutional neural network (CNN) for regularized inversion. The CNN is trained using the surface temperature responses from a set of Green’s function perturbation experiments as imagery input data together with data sample densification. The resulting CNN model can infer the forcing pattern responsible for the temperature response from out-of-sample forcing scenarios. This promising proof of concept suggests a possible strategy for estimating the optimal forcing to negate certain undesirable effects of climate change. The limited success of this effort underscores the challenges of solving an inverse problem for a climate system with inherent nonlinearity. Significance Statement Predicting the climate response for a given climate forcing is a direct problem, while inferring the forcing for a given desired climate response is often an inverse, ill-posed, problem, posing a new challenge to the climate community. This study makes the first attempt to infer the radiative forcing for a given target pattern of global surface temperature response using a deep learning approach. The resulting deeply trained convolutional neural network inversion model shows promise in capturing the forcing pattern corresponding to a given surface temperature response, with a significant implication on the design of an optimal solar radiation management strategy for curbing global warming. This study also highlights the technical challenges that future research should prioritize in seeking feasible solutions to the inverse climate problem.

Ren, Huiying↗

A Census of Young Stellar Objects in Two Line-of-Sight Star-Forming Regions Toward IRAS 22147+5948 in the Outer Galaxy

Context. Star formation in the outer Galaxy, namely, outside of the Solar circle, has not been extensively studied in part due to the low CO brightness of the molecular clouds linked with the negative metallicity gradient. Recent infrared surveys provide an overview of dust emission in large sections of the Galaxy, but they suffer from cloud confusion and poor spatial resolution at far-infrared wavelengths. Aims. We aim to develop a methodology to identify and classify young stellar objects (YSOs) in star-forming regions in the outer Galaxy and use it to resolve a long-standing disparity in terms of the distance and evolutionary status of IRAS 22147+5948. Methods. We used a support vector machine learning algorithm to complement standard color–color and color–magnitude diagrams in our search for YSOs in the IRAS 22147 region, based on publicly available data from the Spitzer Mapping of the Outer Galaxy survey. The agglomerative hierarchical clustering algorithm was used to identify clusters. Then the physical properties of individual YSOs were calculated. The distances were determined using CO 1–0 from the Five College Radio Astronomy Observatory survey. Results. We identified 13 Class I and 13 Class II YSO candidates using the color–color diagrams, along with an additional 2 and 21 sources, respectively, using the applied machine learning techniques. The spectral energy distributions of 23 sources were modeled with a star and a passive disk, corresponding to Class II objects. The models of three sources include envelopes that are typical for Class I objects. The objects were grouped into two clusters located at a distance of 2:2 kpc and 5 clusters at 5:6 kpc. The spatial extent of CO, radio continuum, and dust emission confirms the origin of YSOs in two distinct star-forming regions along a similar line of sight. Conclusions. The outer Galaxy may serve as a unique laboratory for exploring star formation across environments, on the condition that complementary methods and ancillary data are used to properly account for cloud confusion and distance uncertainties.

Agata Karska↗

Cascade Back-Propagation Learning in Neural Networks

The cascade back-propagation (CBP) algorithm is the basis of a conceptual design for accelerating learning in artificial neural networks. The neural networks would be implemented as analog very-large-scale integrated (VLSI) circuits, and circuits to implement the CBP algorithm would be fabricated on the same VLSI circuit chips with the neural networks. Heretofore, artificial neural networks have learned slowly because it has been necessary to train them via software, for lack of a good on-chip learning technique. The CBP algorithm is an on-chip technique that provides for continuous learning in real time. Artificial neural networks are trained by example: A network is presented with training inputs for which the correct outputs are known, and the algorithm strives to adjust the weights of synaptic connections in the network to make the actual outputs approach the correct outputs. The input data are generally divided into three parts. Two of the parts, called the "training" and "cross-validation" sets, respectively, must be such that the corresponding input/output pairs are known. During training, the cross-validation set enables verification of the status of the input-to-output transformation learned by the network to avoid over-learning. The third part of the data, termed the "test" set, consists of the inputs that are required to be transformed into outputs; this set may or may not include the training set and/or the cross-validation set. Proposed neural-network circuitry for on-chip learning would be divided into two distinct networks; one for training and one for validation. Both networks would share the same synaptic weights.

Duong, Tuan A.↗

A deep learning and finite element approach for exploration of inverse structure–property designs of lightweight hybrid composites

Hybrid composites have important applications, such as high-performance and lightweight materials in aerospace and automotive industries. Hybrid composites utilize the synergy of diverse fillers to achieve desired material properties, but usually have more complicated microstructures. While topology optimization can optimize a particular property, designing hybrid composites for customized mechanical performances, e.g. full-range stress–strain curve, remains challenging. Here, a computational framework that integrated finite element analysis (FEA) and artificial intelligence (AI) methods of Conditional Generative Adversarial Networks (cGAN) deep learning and transfer learning was developed to establish inverse structure–property relationships and design tailor-made hybrid composites. Based on FEA-generated datasets of hybrid fiber-particle–matrix microstructures and their corresponding full-range stress–strain curves, a cGAN architecture was trained to generate tailored microstructures and establish structure–property relationships. Similarity in microstructural features and well-matched stress–strain curves based on the AI-generated composites were achieved. In conclusion, transfer learning was used to expand the pre-trained model for designing different materials systems.

Hybrid composites↗

Expanding the Operational Use of Total Lightning Ahead of GOES-R

NASA's Short‐term Prediction Research and Transition Center (SPoRT) has been transitioning real‐time total lightning observations from ground‐based lightning mapping arrays since 2003. This initial effort was with the local Weather Forecast Offices (WFO) that could use the North Alabama Lightning Mapping Array (NALMA). These early collaborations established a strong interest in the use of total lightning for WFO operations. In particular the focus started with warning decision support, but has since expanded to include impact‐based decision support and lightning safety. SPoRT has used its experience to establish connections with new lightning mapping arrays as they become available. The GOES‐R / JPSS Visiting Scientist Program has enabled SPoRT to conduct visits to new partners and expand the number of operational users with access to total lightning observations. In early 2014, SPoRT conducted the most recent visiting scientist trips to meet with forecast offices that will used the Colorado, Houston, and Langmuir Lab (New Mexico) lightning mapping arrays. In addition, SPoRT met with the corresponding Center Weather Service Units (CWSUs) to expand collaborations with the aviation community. These visits were an opportunity to learn about the forecast needs of each office visited as well as to provide on‐site training for the use of total lightning, setting the stage for a real‐time assessment during May‐July 2014. With five lightning mapping arrays covering multiple geographic locations, the 2014 assessment has demonstrated numerous uses of total lightning in varying situations. Several highlights include a much broader use of total lightning for impact‐based decision support ranging from airport weather warnings, supporting fire crews, and protecting large outdoor events. The inclusion of the CWSUs has broadened the operational scope of total lightning, demonstrating how these data can support air traffic management, particularly in the Terminal Radar Approach Control Facilities (TRACON) region around an airport. These collaborations continue to demonstrate, from the operational perspective, the utility of total lightning and the importance of continued training and preparation in advance of the Geostationary Lightning Mapper.

GLM↗

Observable optimization for precision theory: machine learning energy correlators

The practice of collider physics typically involves the marginalization of multi-dimensional collider data to uni-dimensional observables relevant for some physics task. In many cases, such as classification or anomaly detection, the observable can be arbitrarily complicated, such as the output of a neural network. However, for precision measurements, the observable must correspond to something computable systematically beyond the level of current simulation tools. In this work, we demonstrate that precision-theory-compatible observable space exploration can be systematized by using neural simulation-based inference techniques from machine learning. We illustrate this approach by exploring the space of marginalizations of the energy 3-point correlator to optimize sensitivity to the top quark mass. We first learn the energy-weighted probability density from simulation, then search in the space of marginalizations for an optimal triangle shape. Although simulations and machine learning are used in the process of observable optimization, the output is an observable definition which can be then computed to high precision and compared directly to data without any memory of the computations which produced it. We find that the optimal marginalization is isosceles triangles on the sphere with a side ratio approximately $1 : 1 : \sqrt{2}$ (i.e. right triangles) within the set of marginalizations we consider.

Jets and Jet Substructure↗