Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Confidence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Confidence Assessment for Automatic Target Recognition

Confidence assessment is critical for effective automatic target recognition (ATR). Productive use and interpretation of ATR results by analysts or downstream algorithms requires not only algorithmic declarations of target presence and identity, but also algorithmic assessment of the certainty of those declarations in comparison to the certainties of alternative target-identity possibilities. Unfortunately, despite its importance, confidence assessment is an understudied, underdeveloped, and often-neglected function of ATR systems. This lack of regard stems not only from the difficulty of accurate algorithmic determination of target-identity certainty, but also from a general lack of understanding and careful consideration about what confidence should actually represent. We present a framework for confidence assessment that establishes a clear definition of confidence and provides a straightforward theoretical basis for its calculation. This framework is grounded in a hypothesis-theoretic consideration of ATR and it springs from from a handful of axiomatic principles concerning the nature and meaning of confidence in this context. This framework establishes a rigorous mathematical definition of confidence and it provides equations relating confidence to other information that is almost always provided by ATRs. We present an approach for computing confidence within this framework, using an advance process of ATR characterization followed by a simple computation at the time of ATR execution. We discuss practical difficulties with our approach, and we suggest methods for effective mitigation of these difficulties in implemented systems.

97 MATHEMATICS AND COMPUTING↗

Method for Generating Expert Derived Confidence Scores

We executed a pilot demonstration of a methodology for developing a new confidence metric to help operators calibrate their trust in ML event classifiers. This confidence metric was derived from domain expert judgment and was accompanied with a qualitative description describing the reason for each confidence rating. After learning the boundaries of an ML’s performance by studying a subset of events an SME rated his confidence in the ML’s ability to classify similar events and provided an explanation for his ratings. To demonstrate this methodology, we developed our expert driven confidence scores for the ML event classifier within the ESAMS. Next, we assessed the accuracy of the human expert confidence scores relative to the ML’s uncertainty quantification scores. This report includes a description of our methodology, summary of our findings and future directions.

97 MATHEMATICS AND COMPUTING↗

Confidence Calibration Metrics

This technical report serves to summarize a literature search conducted that covered confidence calibration. This report is meant to serve as a solid starting reference for individuals interested in learning more about the confidence calibration domain as well as for individuals more familiar with this work – as a summarizing document for calibration metrics is notably lacking in the literature. This report is not meant to serve as a comprehensive review of everything that has been done in this field – in fact, the reader is encouraged to look further into this domain. We describe confidence and calibration and discuss properties of good calibration metrics. We detail various calibration and calibration-tangential metrics, presenting equations, algorithms, parameters, and an analysis of strengths and weaknesses. We apply a subset of these metrics to eight proxy confidence assessment datasets. We examine the various metrics in the context of model confidence. Finally, we discuss promising future directions and outstanding questions.

99 GENERAL AND MISCELLANEOUS↗

The DECOVALEX international collaboration on modeling of coupled subsurface processes and its contribution to confidence building in radioactive waste disposal

Abstract The long-lived radiotoxicity of the high-level radioactive waste generated by nuclear power plants requires safe isolation from the biosphere for many hundreds of thousands of years. An international consensus has emerged that such isolation can best be provided by disposal in mined geologic repositories, a strategy that today is pursued by most countries dealing with radioactive waste. However, the need to predict the performance of such repositories over very long time periods generates large uncertainties that have to be accounted for in safety assessments. The findings from such safety assessments need to be conveyed to all stakeholders in a clear way, such that public confidence in geologic disposal solutions can be achieved. It is suggested here that close international collaboration on the technical aspects of geologic waste disposal has helped, and will continue to help, building trust and increasing confidence. This paper discusses a particular international collaboration initiative referred to as DECOVALEX, which brings together multiple teams and disciplines to collectively tackle complex experimental and modeling challenges related to geologic disposal. By describing how DECOVALEX works and by providing joint research examples, a case is made that such international collaboration contributes to knowledge transfer and confidence building in radioactive waste disposal science.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Dense Image Matching Uncertainty Estimation and Confidence Metrics

Dense stereo matching takes overlapping image pairs as input and outputs a disparity map which encodes pixel-by-pixel matches between the images. Recently, there has been an interest in ranking the quality, or even quantifying the accuracy, of disparity estimates. The proposed methods can be described as either uncertainty estimators or confidence metrics. Uncertainty estimators are a small minority of the research. However, they have the potential to be the most useful because they estimate disparity accuracy (in pixel units) that can be used to threshold matches or carried forward using error propagation. The majority of the research deals with confidence metrics which give an ordinal (or binary) ranking of a match’s quality relative to other matches. Confidence metrics do not have units and thus are useful primarily for thresholding matches from mismatches. The methods could also be described as handcrafted or deep-learning based. The majority of the research focused on outdoor driving scenes. Hence, our interest–application to a satellite semi-global matching pipeline–is a domain shift that may challenge deep-learning based methods. We conclude by recommending five handcrafted and two deep-learning based methods for evaluation in our pipeline.

97 MATHEMATICS AND COMPUTING↗

Confidence-weighted integration of human and machine judgments for superior decision-making

Large language models (LLMs) can surpass humans in certain forecasting tasks. What role does this leave for humans in the overall decision process? One possibility is that humans, despite performing worse than LLMs, can still add value when teamed with them. A human and machine team can surpass each individual teammate when team members’ confidence is well calibrated and team members diverge in which tasks they find difficult (i.e., calibration and diversity are needed). We simplified and extended a Bayesian approach to combining judgments using a logistic regression framework that integrates confidence-weighted judgments for any number of team members. Using this straightforward method, we demonstrated its effectiveness in both image classification and neuroscience forecasting tasks. Combining human judgments with one or more machines consistently improved overall team performance. Our hope is that this simple and effective strategy for integrating the judgments of humans and machines will lead to productive collaborations.

97 MATHEMATICS AND COMPUTING↗

Comparative Analysis of Confidence Metrics for Nuclear Criticality Safety

Nuclear criticality safety standards provide guidance on the requirements and recommendations to establish confidence in computerized model results used to support operation with fissionable materials. By design, the guidance is not prescriptive, leaving the analysts free to determine how various sources of uncertainties are to be statistically aggregated. This report compares the analyses and key assumptions behind four notable methodologies documented in the nuclear criticality safety literature: the parametric, nonparametric, Whisper, and TSURFER methodologies. Because of the involved use of statistics entangled with heuristic recipes, the results of these methodologies are often difficult to interpret. Also, they are augmented by additional large administrative margins, eliminating the incentive to understand their differences. With the new resurgent wave of advanced nuclear systems focused on economizing operation—including advanced reactors, fuel cycles, and fuel concepts—there is a strong need to develop a clear understanding of uncertainties and their fusion methodologies to reduce uncertainties in a scientifically defensible manner. This report offers a deep dive into the various assumptions of the four noted methodologies, their adequacy, and their limitations, to provide guidance on developing confidence for the emergent nuclear systems. These systems are expected to be challenged by the scarcity of experimental data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Developing Confidence in Machine Learning Results

As the field of deep learning has emerged in recent years, the amount of knowledge and expertise that data scientists are expected to absorb and maintain has correspondingly increased. One of the challenges experienced by data scientists working with deep learning models is developing confidence in the accuracy of their approach and the resulting findings. In this study, we conducted semi-structured interviews with data scientists at a National Laboratory to understand the processes that data scientists use when attempting to develop their models and the ways that they gain confidence that the results they obtained were accurate. These interviews were analysed to provide an overview of the techniques currently used when working with machine learning (ML) models and opportunities for collaboration with human factors researchers to develop new tools are identified.

Baweja, Jessica A.↗

Bounded-Confidence Models of Multidimensional Opinions with Topic-Weighted Discordance

People’s opinions on a wide range of topics often evolve over time through their interactions with others. Models of opinion dynamics primarily focus on one-dimensional opinions, which represent opinions on one topic. However, opinions on various topics are rarely isolated; instead, they can be interdependent and correlated. In a bounded-confidence model (BCM) of opinion dynamics, agents are receptive to each other only if their opinions are sufficiently similar. Here, we extend classical agent-based BCMs—namely, the Hegselmann–Krause BCM, which has synchronous interactions, and the Deffuant–Weisbuch BCM, which has asynchronous interactions—to a multidimensional setting, in which the opinions are multidimensional vectors representing opinions of different topics and opinions on different topics are interdependent. To measure opinion differences between agents, we introduce topic-weighted discordance functions that account for opinion differences in all topics. We define regions of receptiveness for our models, and we use them to characterize the steady-state opinion clusters and provide an analytical approach to compute these regions. In addition, we numerically simulate our models on various networks with initial opinions drawn from a variety of distributions. When initial opinions are correlated across different topics, our topic-weighted BCMs yield significantly different results in both transient and steady states compared to baseline models, where the dynamics of each opinion topic are independent.

Mathematics and Computing↗

Monte Carlo method for constructing confidence intervals with unconstrained and constrained nuisance parameters in the NOvA experiment

Measuring observables to constrain models using maximum-likelihood estimation is fundamental to many physics experiments. Wilks' theorem provides a simple way to construct confidence intervals on model parameters, but it only applies under certain conditions. These conditions, such as nested hypotheses and unbounded parameters, are often violated in neutrino oscillation measurements and other experimental scenarios. Monte Carlo methods can address these issues, albeit at increased computational cost. In the presence of nuisance parameters, however, the best way to implement a Monte Carlo method is ambiguous. Furthermore, this paper documents the method selected by the NOvA experiment, the profile construction. It presents the toy studies that informed the choice of method, details of its implementation, and tests performed to validate it. It also includes some practical considerations which may be of use to others choosing to use the profile construction.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Combining Eddy Covariance Towers, Field Measurements, and the MEMS 2 Ecosystem Model Improves Confidence in the Climate Impacts of Bioenergy With Carbon Capture and Storage

ABSTRACT Carbon dioxide removal technologies such as bioenergy with carbon capture and storage (BECCS) are required if the effects of climate change are to be reversed over the next century. However, BECCS demands extensive land use change that may create positive or negative radiative forcing impacts upstream of the BECCS facility through changes to in situ greenhouse gas fluxes and land surface albedo. When quantifying these upstream climate impacts, even at a single site, different methods can give different estimates. Here we show how three common methods for estimating the net ecosystem carbon balance of bioenergy crops established on former grassland or former cropland can differ in their central estimates and uncertainty. We place these net ecosystem carbon balance forcings in the context of associated radiative forcings from changes to soil N 2 O and CH 4 fluxes, land surface albedo, embedded fossil fuel use, and geologically stored carbon. Results from long term eddy covariance measurements, a soil and plant carbon inventory, and the MEMS 2 process‐based ecosystem model all agree that establishing perennials such as switchgrass or mixed prairie on former cropland resulted in net negative radiative forcing (i.e., global cooling) of −26.5 to −39.6 fW m −2 over 100 years. Establishing these perennials on former grassland sites had similar climate mitigation impacts of −19.3 to −42.5 fW m −2 . However, the largest climate mitigation came from establishing corn for BECCS on former cropland or grassland, with radiative forcings from −38.4 to −50.5 fW m −2 , due to its higher plant productivity and therefore more geologically stored carbon. Our results highlight the strengths and limitations of each method for quantifying the field scale climate impacts of BECCS and show that utilizing multiple methods can increase confidence in the final radiative forcing estimates.

Falvo, Grant [Department of Plant, Soil and Microb↗

Variance-Reduced Accelerated First-Order Methods: Central Limit Theorems and Confidence Statements

In this paper, we consider a strongly convex stochastic optimization problem and propose three classes of variable sample-size stochastic first-order methods: (i) the standard stochastic gradient descent method, (ii) its accelerated variant, and (iii) the stochastic heavy-ball method. In each scheme, the exact gradients are approximated by averaging across an increasing batch size of sampled gradients. We prove that when the sample size increases at a geometric rate, the generated estimates converge in mean to the optimal solution at an analogous geometric rate for schemes (i)–(iii). Based on this result, we provide central limit statements, whereby it is shown that the rescaled estimation errors converge in distribution to a normal distribution with the associated covariance matrix dependent on the Hessian matrix, the covariance of the gradient noise, and the step length. If the sample size increases at a polynomial rate, we show that the estimation errors decay at a corresponding polynomial rate and establish the associated central limit theorems (CLTs). Under certain conditions, we discuss how both the algorithms and the associated limit theorems may be extended to constrained and nonsmooth regimes. As a result, we provide an avenue to construct confidence regions for the optimal solution based on the established CLTs and test the theoretical findings on a stochastic parameter estimation problem.

Lei, Jinlong↗

Robust Explanations using Diverse Adversarially Trained Ensembles, Multi-Modal Contrastive Learning, and Attribution-based Confidence Metrics

The primary objective of this project is to strengthen the trustworthiness of AI systems by designing algorithms that make their internal decision-making processes more understandable to human users. This involves creating clear, interpretable explanations for AI decisions and developing metrics to assess these explanations' validity and reliability. Significant progress has been achieved through (i) developing symbolic explanations, (ii) generating meaningful interpretive insights, (iii) establishing accuracy and confidence metrics, and (iv) devising methods to evaluate the knowledge boundaries of AI models. To date, the research findings have been shared in peer-reviewed publications, with accompanying scientific and technical information (STI) detailed below.

97 MATHEMATICS AND COMPUTING↗

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip↗