Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “confidence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Confidence Assessment for Automatic Target Recognition

Confidence assessment is critical for effective automatic target recognition (ATR). Productive use and interpretation of ATR results by analysts or downstream algorithms requires not only algorithmic declarations of target presence and identity, but also algorithmic assessment of the certainty of those declarations in comparison to the certainties of alternative target-identity possibilities. Unfortunately, despite its importance, confidence assessment is an understudied, underdeveloped, and often-neglected function of ATR systems. This lack of regard stems not only from the difficulty of accurate algorithmic determination of target-identity certainty, but also from a general lack of understanding and careful consideration about what confidence should actually represent. We present a framework for confidence assessment that establishes a clear definition of confidence and provides a straightforward theoretical basis for its calculation. This framework is grounded in a hypothesis-theoretic consideration of ATR and it springs from from a handful of axiomatic principles concerning the nature and meaning of confidence in this context. This framework establishes a rigorous mathematical definition of confidence and it provides equations relating confidence to other information that is almost always provided by ATRs. We present an approach for computing confidence within this framework, using an advance process of ATR characterization followed by a simple computation at the time of ATR execution. We discuss practical difficulties with our approach, and we suggest methods for effective mitigation of these difficulties in implemented systems.

97 MATHEMATICS AND COMPUTING↗

Rats use memory confidence to guide decisions

Memory enables access to past experiences to guide future behavior. Humans can determine which memories to trust (high confidence) and which to doubt (low confidence). How memory retrieval, memory confidence, and memory-guided decisions are related, however, is not understood. In particular, how confidence in memories is used in decision making is unknown. We developed a spatial memory task in which rats were incentivized to gamble their time: betting more following a correct choice yielded greater reward. Rat behavior reflected memory confidence, with higher temporal bets following correct choices. We applied machine learning to identify a memory decision variable and built a generative model of memories evolving over time that accurately predicted both choices and confidence reports. Our results reveal in rats an ability thought to exist exclusively in primates and introduce a unified model of memory dynamics, retrieval, choice, and confidence.

60 APPLIED LIFE SCIENCES↗

Method for Generating Expert Derived Confidence Scores

We executed a pilot demonstration of a methodology for developing a new confidence metric to help operators calibrate their trust in ML event classifiers. This confidence metric was derived from domain expert judgment and was accompanied with a qualitative description describing the reason for each confidence rating. After learning the boundaries of an ML’s performance by studying a subset of events an SME rated his confidence in the ML’s ability to classify similar events and provided an explanation for his ratings. To demonstrate this methodology, we developed our expert driven confidence scores for the ML event classifier within the ESAMS. Next, we assessed the accuracy of the human expert confidence scores relative to the ML’s uncertainty quantification scores. This report includes a description of our methodology, summary of our findings and future directions.

97 MATHEMATICS AND COMPUTING↗

Confidence Calibration Metrics

This technical report serves to summarize a literature search conducted that covered confidence calibration. This report is meant to serve as a solid starting reference for individuals interested in learning more about the confidence calibration domain as well as for individuals more familiar with this work – as a summarizing document for calibration metrics is notably lacking in the literature. This report is not meant to serve as a comprehensive review of everything that has been done in this field – in fact, the reader is encouraged to look further into this domain. We describe confidence and calibration and discuss properties of good calibration metrics. We detail various calibration and calibration-tangential metrics, presenting equations, algorithms, parameters, and an analysis of strengths and weaknesses. We apply a subset of these metrics to eight proxy confidence assessment datasets. We examine the various metrics in the context of model confidence. Finally, we discuss promising future directions and outstanding questions.

99 GENERAL AND MISCELLANEOUS↗

Do LIGO/Virgo Black Hole Mergers Produce AGN Flares? The Case of GW190521 and Prospects for Reaching a Confident Association

The recent report of an association of the gravitational-wave (GW) binary black hole (BBH) merger GW190521 with a flare in the active galactic nuclei (AGNs) J124942.3 + 344929 has generated tremendous excitement. However, GW190521 has one of the largest localization volumes among all of the GW events detected so far. The 90% localization volume likely contains 7400 unobscured AGNs brighter than g ≤ 20.5 AB mag, and it results in a ≳70% probability of chance coincidence for an AGN flare consistent with the GW event. We present a Bayesian formalism to estimate the confidence of an AGN association by analyzing a population of BBH events with dedicated follow-up observations. Depending on the fraction of BBHs arising from AGNs, counterpart searches of $\mathcal{O}(1) - \mathcal{O}(100)$ GW events are needed to establish a confident association, and more than an order of magnitude more for searches without follow-up (i.e., using only the locations of AGN and GW events). Follow-up campaigns of the top ~5% (based on volume localization and binary mass) of BBH events with total rest-frame mass ≥50 M ⊙ are expected to establish a confident association during the next LIGO/Virgo/KAGRA observing run (O4), as long as the true value of the fraction of BBHs giving rise to AGN flares is >0.1. Our formalism allows us to jointly infer cosmological parameters from a sample of BBH events that include chance coincidence flares. Until the confidence of AGN associations is established, the probability of chance coincidence must be taken into account to avoid biasing astrophysical and cosmological constraints.

79 ASTRONOMY AND ASTROPHYSICS↗

The DECOVALEX international collaboration on modeling of coupled subsurface processes and its contribution to confidence building in radioactive waste disposal

Abstract The long-lived radiotoxicity of the high-level radioactive waste generated by nuclear power plants requires safe isolation from the biosphere for many hundreds of thousands of years. An international consensus has emerged that such isolation can best be provided by disposal in mined geologic repositories, a strategy that today is pursued by most countries dealing with radioactive waste. However, the need to predict the performance of such repositories over very long time periods generates large uncertainties that have to be accounted for in safety assessments. The findings from such safety assessments need to be conveyed to all stakeholders in a clear way, such that public confidence in geologic disposal solutions can be achieved. It is suggested here that close international collaboration on the technical aspects of geologic waste disposal has helped, and will continue to help, building trust and increasing confidence. This paper discusses a particular international collaboration initiative referred to as DECOVALEX, which brings together multiple teams and disciplines to collectively tackle complex experimental and modeling challenges related to geologic disposal. By describing how DECOVALEX works and by providing joint research examples, a case is made that such international collaboration contributes to knowledge transfer and confidence building in radioactive waste disposal science.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Dense Image Matching Uncertainty Estimation and Confidence Metrics

Dense stereo matching takes overlapping image pairs as input and outputs a disparity map which encodes pixel-by-pixel matches between the images. Recently, there has been an interest in ranking the quality, or even quantifying the accuracy, of disparity estimates. The proposed methods can be described as either uncertainty estimators or confidence metrics. Uncertainty estimators are a small minority of the research. However, they have the potential to be the most useful because they estimate disparity accuracy (in pixel units) that can be used to threshold matches or carried forward using error propagation. The majority of the research deals with confidence metrics which give an ordinal (or binary) ranking of a match’s quality relative to other matches. Confidence metrics do not have units and thus are useful primarily for thresholding matches from mismatches. The methods could also be described as handcrafted or deep-learning based. The majority of the research focused on outdoor driving scenes. Hence, our interest–application to a satellite semi-global matching pipeline–is a domain shift that may challenge deep-learning based methods. We conclude by recommending five handcrafted and two deep-learning based methods for evaluation in our pipeline.

97 MATHEMATICS AND COMPUTING↗

Confidence-weighted integration of human and machine judgments for superior decision-making

Large language models (LLMs) can surpass humans in certain forecasting tasks. What role does this leave for humans in the overall decision process? One possibility is that humans, despite performing worse than LLMs, can still add value when teamed with them. A human and machine team can surpass each individual teammate when team members’ confidence is well calibrated and team members diverge in which tasks they find difficult (i.e., calibration and diversity are needed). We simplified and extended a Bayesian approach to combining judgments using a logistic regression framework that integrates confidence-weighted judgments for any number of team members. Using this straightforward method, we demonstrated its effectiveness in both image classification and neuroscience forecasting tasks. Combining human judgments with one or more machines consistently improved overall team performance. Our hope is that this simple and effective strategy for integrating the judgments of humans and machines will lead to productive collaborations.

97 MATHEMATICS AND COMPUTING↗

Key Constructs for Reasonable Confidence in System Validation Programs

This paper presents four key constructs developed in preparing IEEE Std P2411TM, Human Factors Engineering Guide for the Validation of System Designs and Integrated Systems Operations at Nuclear Facilities for submission. These constructs reflect judgments that responsible parties make in order to bring validation to closure. Such judgments weigh competing concerns and cost-benefits for stakeholders. Improved processes and methods may support such judgments but cannot alone eliminate uncertainty. Stating and clarifying these constructs aims to reduce uncertainty in the validation process, to better prepare the implementers and reviewers of future validation programs, both in the nuclear industry and beyond. Reasonable Confidence – A proof standard, as used to reach conclusions in a legal case. Legal proof standards are bounded between “preponderance of evidence” and “beyond reasonable doubt”. Reasonable confidence is comparable to an intermediate legal proof standard of “clear and convincing evidence.” Representative Test Set – A collection of scenarios of sufficient variety to represent the anticipated range of operational conditions, events, evolutions, and activities for validating the system(s) under test. Dispositive vs. Diagnostic Criteria – Relevant performance criteria are placed in one of two categories. Dispositive criteria are assessed, without exception, to determine whether a validation test passes or fails overall. Diagnostic criteria are assessed, subject to justified exception, to evaluate the quality or degree of some particular aspect of performance. Repeatability – Consistency in achieving passing results is to be demonstrated for each scenario in the representative test set; thus, the minimum number of formal repetitions for each scenario is two.

99 GENERAL AND MISCELLANEOUS↗

Accurate and confident prediction of electron beam longitudinal properties using spectral virtual diagnostics

Abstract Longitudinal phase space (LPS) provides a critical information about electron beam dynamics for various scientific applications. For example, it can give insight into the high-brightness X-ray radiation from a free electron laser. Existing diagnostics are invasive, and often times cannot operate at the required resolution. In this work we present a machine learning-based Virtual Diagnostic (VD) tool to accurately predict the LPS for every shot using spectral information collected non-destructively from the radiation of relativistic electron beam. We demonstrate the tool’s accuracy for three different case studies with experimental or simulated data. For each case, we introduce a method to increase the confidence in the VD tool. We anticipate that spectral VD would improve the setup and understanding of experimental configurations at DOE’s user facilities as well as data sorting and analysis. The spectral VD can provide confident knowledge of the longitudinal bunch properties at the next generation of high-repetition rate linear accelerators while reducing the load on data storage, readout and streaming requirements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Confidence-Guided Technique for Tracking Time-Varying Features

Application scientists often employ feature tracking algorithms to capture the temporal evolution of various features in their simulation data. However, as the complexity of the scientific features is increasing with the advanced simulation modeling techniques, quantification of reliability of the feature tracking algorithms is becoming important. One of the desired requirements for any robust feature tracking algorithm is to estimate its confidence during each tracking step so that the results obtained can be interpreted without any ambiguity. To address this, we develop a confidence-guided feature tracking algorithm that allows reliable tracking of user-selected features and presents the tracking dynamics using a graph-based visualization along with the spatial visualization of the tracked feature. Here, the efficacy of the proposed method is demonstrated by applying it to two scientific datasets containing different types of time-varying features.

97 MATHEMATICS AND COMPUTING↗

Comparative Analysis of Confidence Metrics for Nuclear Criticality Safety

Nuclear criticality safety standards provide guidance on the requirements and recommendations to establish confidence in computerized model results used to support operation with fissionable materials. By design, the guidance is not prescriptive, leaving the analysts free to determine how various sources of uncertainties are to be statistically aggregated. This report compares the analyses and key assumptions behind four notable methodologies documented in the nuclear criticality safety literature: the parametric, nonparametric, Whisper, and TSURFER methodologies. Because of the involved use of statistics entangled with heuristic recipes, the results of these methodologies are often difficult to interpret. Also, they are augmented by additional large administrative margins, eliminating the incentive to understand their differences. With the new resurgent wave of advanced nuclear systems focused on economizing operation—including advanced reactors, fuel cycles, and fuel concepts—there is a strong need to develop a clear understanding of uncertainties and their fusion methodologies to reduce uncertainties in a scientifically defensible manner. This report offers a deep dive into the various assumptions of the four noted methodologies, their adequacy, and their limitations, to provide guidance on developing confidence for the emergent nuclear systems. These systems are expected to be challenged by the scarcity of experimental data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Optical Microscopy-Guided Laser Ablation Electrospray Ionization Ion Mobility Mass Spectrometry: Ambient Single Cell Metabolomics with Increased Confidence in Molecular Identification

Single cell analysis is a field of increasing interest as new tools are continually being developed to understand intercellular differences within large cell populations. Laser-ablation electrospray ionization mass spectrometry (LAESI-MS) is an emerging technique for single cell metabolomics. Over the years, it has been validated that this ionization technique is advantageous for probing the molecular content of individual cells in situ. Here, we report the integration of a microscope into the optical train of the LAESI source to allow for visually informed ambient in situ single cell analysis. Additionally, we have coupled this ‘LAESI microscope’ to a drift-tube ion mobility mass spectrometer to enable separation of isobaric species and allow for the determination of ion collision cross sections in conjunction with accurate mass measurements. This combined information helps provide higher confidence for structural assignment of molecules ablated from single cells. Here, we show that this system enables the analysis of the metabolite content of Allium cepa epidermal cells with high confidence structural identification together with their spatial locations within a tissue.

59 BASIC BIOLOGICAL SCIENCES↗

Developing Confidence in Machine Learning Results

As the field of deep learning has emerged in recent years, the amount of knowledge and expertise that data scientists are expected to absorb and maintain has correspondingly increased. One of the challenges experienced by data scientists working with deep learning models is developing confidence in the accuracy of their approach and the resulting findings. In this study, we conducted semi-structured interviews with data scientists at a National Laboratory to understand the processes that data scientists use when attempting to develop their models and the ways that they gain confidence that the results they obtained were accurate. These interviews were analysed to provide an overview of the techniques currently used when working with machine learning (ML) models and opportunities for collaboration with human factors researchers to develop new tools are identified.

Baweja, Jessica A.↗

Bounded-Confidence Models of Multidimensional Opinions with Topic-Weighted Discordance

People’s opinions on a wide range of topics often evolve over time through their interactions with others. Models of opinion dynamics primarily focus on one-dimensional opinions, which represent opinions on one topic. However, opinions on various topics are rarely isolated; instead, they can be interdependent and correlated. In a bounded-confidence model (BCM) of opinion dynamics, agents are receptive to each other only if their opinions are sufficiently similar. Here, we extend classical agent-based BCMs—namely, the Hegselmann–Krause BCM, which has synchronous interactions, and the Deffuant–Weisbuch BCM, which has asynchronous interactions—to a multidimensional setting, in which the opinions are multidimensional vectors representing opinions of different topics and opinions on different topics are interdependent. To measure opinion differences between agents, we introduce topic-weighted discordance functions that account for opinion differences in all topics. We define regions of receptiveness for our models, and we use them to characterize the steady-state opinion clusters and provide an analytical approach to compute these regions. In addition, we numerically simulate our models on various networks with initial opinions drawn from a variety of distributions. When initial opinions are correlated across different topics, our topic-weighted BCMs yield significantly different results in both transient and steady states compared to baseline models, where the dynamics of each opinion topic are independent.

Mathematics and Computing↗

Multi-year incubation experiments boost confidence in model projections of long-term soil carbon dynamics

Abstract Global soil organic carbon (SOC) stocks may decline with a warmer climate. However, model projections of changes in SOC due to climate warming depend on microbially-driven processes that are usually parameterized based on laboratory incubations. To assess how lab-scale incubation datasets inform model projections over decades, we optimized five microbially-relevant parameters in the Microbial-ENzyme Decomposition (MEND) model using 16 short-term glucose (6-day), 16 short-term cellulose (30-day) and 16 long-term cellulose (729-day) incubation datasets with soils from forests and grasslands across contrasting soil types. Our analysis identified consistently higher parameter estimates given the short-term versus long-term datasets. Implementing the short-term and long-term parameters, respectively, resulted in SOC loss (–8.2 ± 5.1% or –3.9 ± 2.8%), and minor SOC gain (1.8 ± 1.0%) in response to 5 °C warming, while only the latter is consistent with a meta-analysis of 149 field warming observations (1.6 ± 4.0%). Comparing multiple subsets of cellulose incubations (i.e., 6, 30, 90, 180, 360, 480 and 729-day) revealed comparable projections to the observed long-term SOC changes under warming only on 480- and 729-day. Integrating multi-year datasets of soil incubations (e.g., > 1.5 years) with microbial models can thus achieve more reasonable parameterization of key microbial processes and subsequently boost the accuracy and confidence of long-term SOC projections.

54 ENVIRONMENTAL SCIENCES↗

Monte Carlo method for constructing confidence intervals with unconstrained and constrained nuisance parameters in the NOvA experiment

Measuring observables to constrain models using maximum-likelihood estimation is fundamental to many physics experiments. Wilks' theorem provides a simple way to construct confidence intervals on model parameters, but it only applies under certain conditions. These conditions, such as nested hypotheses and unbounded parameters, are often violated in neutrino oscillation measurements and other experimental scenarios. Monte Carlo methods can address these issues, albeit at increased computational cost. In the presence of nuisance parameters, however, the best way to implement a Monte Carlo method is ambiguous. Furthermore, this paper documents the method selected by the NOvA experiment, the profile construction. It presents the toy studies that informed the choice of method, details of its implementation, and tests performed to validate it. It also includes some practical considerations which may be of use to others choosing to use the profile construction.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Combining Eddy Covariance Towers, Field Measurements, and the MEMS 2 Ecosystem Model Improves Confidence in the Climate Impacts of Bioenergy With Carbon Capture and Storage

ABSTRACT Carbon dioxide removal technologies such as bioenergy with carbon capture and storage (BECCS) are required if the effects of climate change are to be reversed over the next century. However, BECCS demands extensive land use change that may create positive or negative radiative forcing impacts upstream of the BECCS facility through changes to in situ greenhouse gas fluxes and land surface albedo. When quantifying these upstream climate impacts, even at a single site, different methods can give different estimates. Here we show how three common methods for estimating the net ecosystem carbon balance of bioenergy crops established on former grassland or former cropland can differ in their central estimates and uncertainty. We place these net ecosystem carbon balance forcings in the context of associated radiative forcings from changes to soil N 2 O and CH 4 fluxes, land surface albedo, embedded fossil fuel use, and geologically stored carbon. Results from long term eddy covariance measurements, a soil and plant carbon inventory, and the MEMS 2 process‐based ecosystem model all agree that establishing perennials such as switchgrass or mixed prairie on former cropland resulted in net negative radiative forcing (i.e., global cooling) of −26.5 to −39.6 fW m −2 over 100 years. Establishing these perennials on former grassland sites had similar climate mitigation impacts of −19.3 to −42.5 fW m −2 . However, the largest climate mitigation came from establishing corn for BECCS on former cropland or grassland, with radiative forcings from −38.4 to −50.5 fW m −2 , due to its higher plant productivity and therefore more geologically stored carbon. Our results highlight the strengths and limitations of each method for quantifying the field scale climate impacts of BECCS and show that utilizing multiple methods can increase confidence in the final radiative forcing estimates.

Falvo, Grant [Department of Plant, Soil and Microb↗