Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Judgment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Confidence-weighted integration of human and machine judgments for superior decision-making

Large language models (LLMs) can surpass humans in certain forecasting tasks. What role does this leave for humans in the overall decision process? One possibility is that humans, despite performing worse than LLMs, can still add value when teamed with them. A human and machine team can surpass each individual teammate when team members’ confidence is well calibrated and team members diverge in which tasks they find difficult (i.e., calibration and diversity are needed). We simplified and extended a Bayesian approach to combining judgments using a logistic regression framework that integrates confidence-weighted judgments for any number of team members. Using this straightforward method, we demonstrated its effectiveness in both image classification and neuroscience forecasting tasks. Combining human judgments with one or more machines consistently improved overall team performance. Our hope is that this simple and effective strategy for integrating the judgments of humans and machines will lead to productive collaborations.

97 MATHEMATICS AND COMPUTING

Neural correspondence to spectrum of environmental uncertainty in multiple-cue probability judgment system with time delay

Despite state-of-the-art technologies like artificial intelligence, human judgment is critically essential in cooperative systems, such as the multi-agent system (MAS), which collect information among agents based on multiple-cue judgment. Human agents can prevent impaired situational awareness of automated agents by confirming situations under environmental uncertainty. System error caused by uncertainty can result in an unreliable system environment, and this environment affects the human agent, resulting in non-optimal decision-making in MAS. Thus, it is necessary to know how human behavior is changed to capture system reliability under uncertainty. Another issue affecting MAS is time delay, which can delay agent information transfer, resulting in low performance and instability. However, it is difficult to find studies on the influence of time delay on human agents. This study is about understanding the human decision-making process under a specific system reliability environment by uncertainty with time delay. We used concepts of expected and unexpected uncertainty to implement reliability of the system usage environment with three types of time delay conditions: no time delay, regular time delay, and irregular time delay conditions. We used electroencephalogram (EEG) for human cognitive neural mechanisms in multiple-cue judgment systems to understand human decision-making. In the reliability of system usage environment, the unreliable system environment significantly creates less memory load by less utilization of system rules for decision-making. In terms of time delay, delayed information delivery does not significantly affect memory load for decision-making.

cognitive process

Criteria for Retention of 3013 S1 Containers Based on Relative Risk and Expert Judgment

An evaluation was performed to assess the suitability of thirty-three 3013 containers proposed for retention. These containers have moisture levels greater than 0.08 wt.% – the S1 population. The remainder of the S1 population stored at SRS will be down blended and disposed of by the end of 2028. Based on field surveillance and shelf-life data available to date as well as informed technical judgment, no container is currently expected to fail in its 50-year storage period. However, corrosion risk varies across the S1 population. Relative risks were evaluated using predicted Consensus Scores and their 95% Upper Prediction Limits (UPLs). The predicted values are based on a statistical model of Consensus Score as a function of moisture, chloride, and whether the packaged material was electrorefining scrap packaged at Hanford. Consensus Score has been shown to be a useful indicator of corrosion potential, and the UPL captures uncertainty in the model predictions, providing a conservative indicator of corrosion potential. Using UPLs to determine relative risks, together with expert review, three containers were identified as not suitable for retention, and the remainder were determined to be suitable.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Identifying preferential flow from soil moisture time series: Review of methodologies

Abstract Identifying and quantifying preferential flow (PF) through soil—the rapid movement of water through spatially distinct pathways in the subsurface—is vital to understanding how the hydrologic cycle responds to climate, land cover, and anthropogenic changes. In recent decades, methods have been developed that use measured soil moisture time series to identify PF. Because they allow for continuous monitoring and are relatively easy to implement, these methods have become an important tool for recognizing when, where, and under what conditions PF occurs. The methods seek to identify a pattern or quantification that indicates the occurrence of PF. Most commonly, the chosen signature is either (1) a nonsequential response to infiltrated water, in which soil moisture responses do not occur in order of shallowest to deepest, or (2) a velocity criterion, in which newly infiltrated water is detected at depth earlier than is possible by nonpreferential flow processes. Alternative signatures have also been developed that have certain advantages but are less commonly utilized. Choosing among these possible signatures requires attention to their pertinent characteristics, including susceptibility to errors, possible bias toward false negatives or false positives, reliance on subjective judgments, and possible requirements for additional types of data. We review 77 studies that have applied such methods to highlight important information for readers who want to identify PF from soil moisture data and to inform those who aim to develop new methods or improve existing ones. Core Ideas Soil moisture data can be used to identify the occurrence of preferential flow (PF) and its initiating conditions. Various data‐analysis methods to identify PF differ in susceptibility to error, bias, and subjectivity. These methods can utilize vast amounts of data from soil moisture monitoring networks to develop understanding of when, where, and under what conditions PF occurs. Newly developed methods may lead to better accuracy and reliability, and reduce the need for subjective judgments. Plain Language Summary Preferential flow through soil occurs when a large amount of water is suddenly available, as during an intense storm. This type of flow moves rapidly through the soil in distinct narrow pathways rather than moving evenly throughout the body of soil, with major consequences for groundwater resources, ecosystems, spreading of contaminants, and other vital concerns. Methods of detecting preferential flow have been developed that utilize measurements of soil water content made by sensors installed at various depths. This measurement technology has been widely implemented, many locations now having datasets years in length, and various methods have been developed for using these to identify preferential flow. The various methods are based on different features in the soil moisture records and vary in their advantages and shortcomings. In this review, we explain and evaluate these methods, highlighting important information for their implementation to identify preferential flow from soil moisture data and for efforts to develop new methods or improve existing ones.

Nimmo, John R

A mathematical approach to using the forgetting curve to evaluate experience and training factors in human reliability analysis

Traditional human reliability analysis (HRA) methods have difficulty dealing with the dynamic nature of factors such as time and rely on static and expert-judgment-based assessments of performance-shaping factors (PSFs) across limited levels. In this study, we introduce a mathematical approach for dynamically evaluating the experience and training PSF. Our proposed method integrates the psychological concept of the “forgetting curve” to evaluate how PSFs are impacted by the number of trainings and the time elapsed since training. To confirm the validity of the model, we provide experimental data fitted by identifying the quantitative relationship between training and human performance. This research enables dynamic and objective assessments, thus reducing reliance on subjective expert judgment and improving the accuracy of HRA.

99 - GENERAL AND MISCELLANEOUS

Virtual refrigerant charge sensor for variable-speed heat pumps based on feature selection

The refrigerant charge level in heat pump systems significantly impacts their energy efficiency. Virtual refrigerant charge (VRC) sensing technology has been comprehensively investigated and well-established due to its lower cost compared to physical sensors. However, the previous VRC research often relied on expert judgment and physical reasoning for their variable selection, which can potentially select redundant (or highly correlated) or insignificant features, and it is also primarily focused on single-speed systems. To address these challenges, this study proposes a VRC algorithm for variable-speed heat pumps that selects features through a rigorous feature selection method in combination with physical insights. We also propose a piecewise linear model structure segmented by subcooling temperature to accurately predict charge levels, particularly when subcooling temperatures are substantially low. The proposed algorithm was evaluated using experimental data of a residential R410A heat pump, and the performance was compared with two baseline VRC algorithms. The results are: (1) The proposed algorithm outperforms for the case with subcooling temperature less than 1 °C. (2) The proposed algorithm achieves a tested mean absolute percentage error (MAPE) of 4.23%, and improves the overall accuracy for cooling conditions by approximately 60%, compared with the two baseline algorithms. (3) The proposed algorithm uses two fewer features and improves the accuracy for undercharge cooling conditions by 68.0%, compared with baseline algorithm 2. These improvements enhance prediction accuracy and prevent overfitting, providing a more reliable refrigerant charge level prediction and helping improve the heat pump energy efficiency.

Liang, Chenjiyu

Governing in Time: Temporal Capacity and the Feasibility of Energy Transitions

Energy systems function as both technological systems and temporal institutions that shape how societies coordinate, justify, and support collective choices over time. This paper introduces the concept of governance horizons to explain why energy transitions can remain morally supported yet become institutionally weak under increasing pressure. We argue that governability depends on institutions' capacity to synchronize across multiple timeframes - aligning short-term decisions with intermediate coordination and long-term commitments. When this synchronization fails, transitions struggle not because their goals are dismissed, but because governance lacks sufficient time to justify, coordinate, and uphold decisions. Comparative analysis of San Antonio, Texas, and Interior Alaska reveals how energy system pressures generate distinct temporal configurations: San Antonio exhibits governance horizon stretching, where institutions must simultaneously meet near-term reliability demands and long-term transformation goals, while Interior Alaska exhibits horizon compression, where extreme environmental constraints force decision-making into short stabilization cycles. In both contexts, public support for sustainability goals coexists with institutional strain because evaluative judgments are unevenly distributed over time. A temporal configuration analysis is introduced as a diagnostic analytic stance for identifying these patterns. By treating temporal alignment as an explanatory variable rather than a background condition, this approach clarifies how feasibility, sequencing, and legitimacy are shaped by constraints on institutional time. The analysis demonstrates that successful energy transitions depend not only on technological innovation or institutional support, but on governance systems’ ability to sustain credible coordination across multiple time horizons.

Comparative case study

Impact of representative ground motion level on seismic PSA with the boundary between overestimation and underestimation

One commonly used approach in seismic probabilistic safety assessment (PSA) is the discrete method. This method follows the standard PSA framework and can be applied to various models, such as multi-unit models, while reducing computational costs using standard software. However, due to the inability to subdivide intervals infinitely, the discrete method approximates with a finite number of subintervals. In practice, different numbers of subintervals are applied, and the representative ground motion level is selected based on expert judgment. When employing a smaller number of subintervals, it is important to take caution to prevent underestimation. This study analyzes the impact of the representative ground motion level on seismic risk. It confirms that underestimation can occur with a small number of subintervals depending on the representative ground motion level. This study also proposes a method for determining the boundary of underestimation and overestimation. The method is demonstrated through examples, providing a mathematical foundation for selecting appropriate representative ground motion levels. By avoiding underestimation, this research helps prevent the oversight of significant risk contributors and enhances the understanding of seismic risk.

99 - GENERAL AND MISCELLANEOUS

Sensitivity analysis, surrogate modeling, and optimization of pebble-bed reactors considering normal and accident conditions

This research provides a valuable tool that streamlines the optimization process while significantly increasing its accuracy. This study creates a robust framework for reactor design optimization by incorporating comprehensive modeling using the Comprehensive Reactor Analysis Bundle, or BlueCRAB, within the Multiphysics Object-Oriented Simulation Environment (MOOSE). BlueCRAB is the United States Nuclear Regulatory Commission's code suite for non-light water reactor analysis and includes the Griffin, Pronghorn, and Bison applications. This not only improves the efficiency of the optimization process but also enhances the reliability of the results. Such a tool is essential for advancing the state-of-the-art in pebble-bed reactor technology and is critical for achieving the goals of Generation IV reactors, which aim for safe, sustainable, and economically viable nuclear energy solutions. This work presents and applies this workflow on pebble-bed reactors while considering both normal and off-normal conditions. A representative gas-cooled pebble-bed reactor at equilibrium core conditions serves as the nominal design specification for normal operation and is based on previous research. The depressurized loss-of-forced-cooling accident is deployed for off-normal conditions in this work. After defining design-related parameters and quantities of interest regarding reactor safety and performance, this multiphysics model is sampled using the MOOSE stochastic tools module. The result is a comprehensive dataset of configurations, enabling sensitivity analysis and the generation of surrogate models. Subsequently, the dataset and surrogate models are employed in two optimization studies aimed at maximizing fuel utilization and economic profit while adhering to safety and operational constraints. Performing the optimization process with fuel utilization as the metric leads to an improvement of approximately 10%, compared to engineering-judgment-based nominal conditions. The optimization on economic profit leads to an estimated increase of ~300 million USD over the lifetime of the reactor.

97 MATHEMATICS AND COMPUTING

Technoeconomic Design Optimization for Fast Reactors. Part I: Workflow Development and Case Study for Small LFR District Energy Application

The nuclear industry is developing small reactor designs that can target a variety of deployment locations and energy products. Smaller nuclear designs have traditionally struggled to handle the steep trade-offs between size and cost that have historically incentivized large reactors. This motivates computational optimization of small reactors to minimize costs and quantify the trade-off between size and cost. In this paper, the cost/size trade-off for a small fast reactor is derived using a multi-objective genetic algorithm optimization, with steady-state, transient, and cost analysis of the fast reactor being performed. Specifically, the method is demonstrated on a small 10- to 120-MW(thermal) U-Pu-Zr–fueled lead-cooled fast reactor with a 10-year core life for district energy applications, which can have a thermal load compatible with this range. The results reinforced that fast reactor cores at the lower end of this power range suffer cost penalties due to critical mass considerations. It was found that high power density cores with strong reactivity swings and many control rods were favored over designing to minimize reactivity swing. Furthermore, this contrasts with some traditional configurations designed using engineering judgment and demonstrates that optimizers can find nontraditional but realistic solutions, along with demonstrating the value of incorporating cost functions into whole-reactor design optimization.

Fast reactor

The ballad of LLM agents: philosophical reasoning for chemistry

Large language models (LLMs) show remarkable potential for scientific reasoning but often produce unreliable or scientifically unactionable outputs when faced with multi-step logic, domain grounding, and interpretability challenges, especially in complex fields like chemistry and materials science. Here, we introduce a framework of philosophical reasoning agents, inspired by canonical thinkers such as Socrates, Descartes, Kant, and Hume, to guide LLM behavior via structured prompt engineering. These agents embody distinct reasoning paradigms (dialectical inquiry, deductive logic, rule-based judgment, and empirical validation) and are evaluated across multiple chemistry subdomains, physical, analytical, general, inorganic, and organic chemistry, using the ChemBench benchmark. Our agentic prompting approach yields substantial accuracy gains on open-ended numerical chemistry questions, with gains of +11.5 percentage points for GPT-4o with Hume, +4.5 percentage points for GPT-5 with Kant, and +21.8 percentage points for GPT-5.1 with Socrates at the strict 1% error threshold, relative to the corresponding base models. Beyond accuracy, we observe benchmark-level model–agent performance patterns, suggesting that different prompting styles interact differently with each base model. These findings demonstrate that embedding philosophy-of-science principles into multi-agent frameworks can improve and produce interpretable, adaptive, and domain-aligned scientific LLMs.

Harb, Hassan [Argonne National Laboratory (ANL), A

Feedback and Oscillations: Constructing Feedback Systems for Root Cause Analysis of Oscillations in Power Grids

Dynamic phenomena linked to inverter-based resources (IBRs) have gained global attention. Several IBR-induced dynamics have caused bulk power system-connected wind or solar power plants to trip, and some have even led to widespread outages. In addition, many oscillations have been observed involving IBR power plants. In 2023, the IEEE Power & Energy Society (PES) IBR Subsynchronous Oscillations (SSO) task force published a journal article, “Real-World Subsynchronous Oscillation Events in Power Grids With High Penetrations of Inverter-Based Resources,” in which 19 IBR oscillation events were examined for their causation. Earlier in 2020, another PES task force article, “Definition and Classification of Power System Stability-Revisited & Extended,” authored by prominent academics, introduced converter-driven stability as a new category of stability. The international power grid industry community also took action by publishing the CIGRE Green Book, Power System Dynamic Modelling and Analysis in Evolving Networks (led by Babak Badrzadeh and Zia Emin) in 2024. In August 2024, the Energy Systems Integration Group (ESIG) released a practical guide led by Nick Miller, “Diagnosis and Mitigation of Observed Oscillations in IBR-Dominant Power System: A Practical Guide.” The goal of the guide is to assist practicing engineers in making initial judgments and conducting detailed analyses about oscillations. Finally, when addressing the classification of stability and oscillations, the guide emphasizes a causality-based taxonomy for grouping, such as voltage control-induced oscillations, synchronization-induced oscillations, and frequency or active power control-induced oscillations.

Fan, Lingling [Univ. of South Florida, Tampa, FL (

C‐Cracking of Brittle‐Material Spheres From Eccentric Hertzian Contact

Contact of brittle-material spheres can occur while they are manufactured, handled, or used in operation as roller elements or in deformably processed composite material. Statistically, the directional contact between two spheres will nearly always be oblique or eccentric in an unconfined and mobile sphere population. Although analyses of oblique contact exist, the portrayal of its maximum (tensile) first principal stress ( S 1 ) field is lacking. This information is needed for improved judgment of the prospect of crack initiation. Given these, the S 1 due to eccentric contact and subsequent crack initiation was analyzed using finite element analysis (FEA) and corroborative demonstration of c-crack creation. The FEA shows that an asymmetric S 1 field is created about the contact patch and is caused by superimposed shear intrinsic to eccentric contact. A crack will initiate if the maximum S 1 exceeds the material's tensile strength, and its field will cause the crack to propagate and arrest to form a C-shaped crack, or c-crack. C-crack initiation trends with lower applied forces, with greater amount of contact eccentricity, higher coefficient of friction, and smaller spherical radius. Finally, these metrics deserve recognition, especially when c-cracked brittle-material spheres are used in a system whose functionality or reliability is potentially compromised by the c-cracks.

Hertzian contact

Examining Cloud Feedback Components in the Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM)

Cloud feedback remains the main source of uncertainty in climate sensitivity estimated by global climate models (GCMs), largely because subgrid cloud responses are parameterized in GCMs due to their coarse resolution. Here, this study examines cloud feedback in the global 3.25-km Simple Cloud-Resolving Energy Exascale Earth System Model (E3SM) Atmosphere Model (SCREAM 3 km) through a pair of 1-yr atmosphere-only simulations with control and +4-K sea surface temperature perturbations. SCREAM 3 km produces a positive cloud feedback that falls within but at the upper end of the range of Coupled Model Intercomparison Project phase 5 (CMIP5) and CMIP phase 6 (CMIP6) models and expert judgment. The positive cloud feedback arises from positive contributions from both high- and low-level clouds, with increases in high-cloud altitude and decreases in low-cloud amount and optical depth playing key roles. The stronger-than-CMIP-average feedback is mainly attributable to the high-cloud altitude feedback, owing to cloud tops rising nearly isothermally in SCREAM 3 km. The positive low-cloud amount feedback is weaker in SCREAM than in GCMs because estimated inversion strength (EIS) increases more dramatically with warming. A coarser 12-km resolution version of SCREAM exhibits a weaker positive cloud feedback than SCREAM 3 km, mainly because its low-cloud-radiative flux is more sensitive to EIS, leading to a stronger negative low-cloud amount feedback. With this process-level assessment of cloud feedback, this study reveals where SCREAM aligns with and diverges from conventional GCMs and expert assessment, providing insights to inform further model improvement and future expert assessment.

Cloud radiative effects

Understanding Decision-Relevant Regional Data Products: Workshop Report

A broad community of climate adaptation practitioners, stakeholders and policymakers rely on historical reconstructions and future projections of local to regional climate. To be of value to these users, climate data must be credible, salient, and authoritative (Cash et al. 2002). Namely, data must be consistent with our physical understanding of the global Earth system, must be relevant for informing the decision-making process, and must be backed by expert judgment. As more and more data products have become available, multiple challenges have emerged around the production, evaluation, selection, and use of these data products. Consequently, to ensure crucial decisions leverage the best possible historical and future physical climate data, there is a pressing need to develop a coordinated national climate data strategy that is inclusive of all relevant communities of practice.

54 ENVIRONMENTAL SCIENCES

Virtual Reality for Shoot/No-Shoot Decision Training in Law Enforcement: A Literature Review and Research Agenda

Virtual reality (VR) can materially improve “shoot / no-shoot” (SNS) training by giving officers realistic, repeatable practice making high-stakes decisions under pressure. Traditional tools—live-fire ranges and video simulators—build basics, but they cannot adapt to each officer in real time or fully mirror the complexity of the field. VR closes that gap by creating immersive scenarios that are safer, more flexible, easier to scale across units, and able to capture objective performance data. SNS decisions are not just about marksmanship; they rely on perception, judgment, memory, and the ability to hold fire when a threat is uncertain. Effective training therefore needs realism, decision complexity, and branching outcomes that reflect the true consequences of choices. These elements strengthen recognition of hostile intent while reducing false positives and building the self-control required in ambiguous situations. VR brings specific advantages: dynamic environments, full-body interaction, and the ability to measure performance with precision—enabling targeted feedback and better transfer of learning to the street. At the same time, responsible deployment must address scenario quality (credible environments and behaviors), lawful decision models, and user wellbeing (appropriate stress levels, comfort, and safety). Sandia’s VIPER Lab is positioned to lead this work. The team combines human-performance science, AI/ML, and VR/AR development with a deep equipment bench (e.g., omnidirectional treadmill, eye-tracking, haptics, multiple HMDs). This ecosystem supports building and validating next-generation SNS training that is immersive, measurable, and trustworthy. Bottom line: Investment in VR-enabled SNS training that blends evidence-based design with careful validation and legal safeguards is expected to pay off in safer, more consistent decision-making and improved community trust, delivered through training that is practical to deploy at scale.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF

Assessing and Enabling Trustworthy Predictions for High-Consequence Decisions

Predictions from physics-based computational models provide critical information to inform high consequence decisions, e.g., engineering design decisions. The ability to assess the reliability of such predictions is therefore critical. However, to date, reliability assessment rely heavily on expert judgment and qualitative arguments. This report details the efforts of LDRD 233072 to develop quantitative methods to assess reliability of model predictions, especially in the context of simplifying assumptions that can impact their reliability.

42 ENGINEERING