Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model based definitions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

From Well Log to Formation Model: A Novel Laboratory Calibrated Methodology with Demonstration

This work demonstrates how the characterization and modeling of both elastic and creep properties are essential to describe zones in layered rock formations as either low-stress targets for stimulation or high-stress barriers to fracture growth. Prediction of fracture height is critical for designing stimulation operations in oil and gas wells. Ideally, fractures are placed in target zones which will produce hydrocarbons and should not propagate into zones expected to be unproductive or to produce unwanted fluids such as water which in turn must be treated and/or disposed. The essential task in designing stimulation plans is predicting which zones have low horizontal stresses and which will be high-stress barriers to fracture growth. Despite this importance, there are gaps in current knowledge and a complete workflow from laboratory characterization to a finite element model which includes time dependent rock deformation is required. While the research and methodology presented here also have application to CO 2or hydrogen storage, wastewater injection, and geothermal applications, the focus will be on hydrocarbon extraction. This thesis presents the results of a characterization-to-prediction workflow for the Caney shale, which is an emerging hydrocarbon resource in Oklahoma, USA. It begins with an investigation to enable critical evaluation of the Caney zonation into nominally “brittle” and “ductile” zones based on properties observed from well logs. It shows none of the zones are consistently “brittle” or “ductile” mechanical behavior based on the variety of definitions of these terms. However, the nominally ductile zones are weaker and more prone to creep. A laboratory investigation of samples including strength, elastic, and creep properties, is then used in a finite element model of stress evolution. The model includes both elastic deformation and viscoplastic creep. Results predict the least creep-prone layers to have the lowest horizontal stresses, therefore comprising hydraulic fracturing targets. The most creep-prone layers attain a horizontal stress similar to the vertical stress and therefore are predicted to be high stress barriers to hydraulic fracture stimulation. In addition to defining stimulation target intervals, the model shows how as tectonic strain rate increases, there is a transition from creep-dominated stresses to stresses dominated by elasticity.

Benge, Margaret↗

Modeling, Performance Assessment, and Nodal Data Analysis of TRISO-Fueled Systems with Shift

This technical report documents several enhancements to the Shift Monte Carlo (MC) code under the US Department of Energy (DOE) Nuclear Energy Advanced Modeling and Simulation (NEAMS) program in fiscal year (FY) 2022. Performance enhancements were added to Shift specifically for tristructural isotropic (TRISO)–fueled reactor systems and guided based on performance analysis in FY 2021. For the pebble performance model developed in previous studies, the runtime improved by ~ 91× compared to the original model and ~ 2× compared to the user-optimized model. Compared to Serpent, Shift is ~ 3× slower if Serpent delta-tracking is enabled but ~ 2× faster when delta-tracking is disabled. The multigroup cross section generation was improved through simplifying tally input definitions, porting several post-processing tally operations from Python scripts into the Shift code base, and accounting for production reactions in the scattering multiplicity. Progress was also made on two emerging capabilities: (1) the development of Titan (a Shift reactor physics user interface) and (2) initial investigation into path-length tallies for computing multigroup scattering matrices.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Assessment of integral models for non-Boussinesq lazy plumes using numerical simulations

Integral modelling of turbulent buoyant plumes is crucial for rapid predictions of plume characteristics. While the governing equations are typically derived using self-similarity and a Boussinesq approximation, these assumptions may not hold for plumes originating from finite-area sources with large density ratios. Here, this work evaluates the accuracy of integral-scale models for non-Boussinesq lazy plumes using high-fidelity numerical simulations of turbulent helium plumes. We analyse the plume kinematics by computing vertical fluxes, plume radius and radial profiles, establishing some disparities between common practice and physical accuracy. We identify how the definition of the plume radius changes the perception of the plume structure when the flow is not self-similar and derive a relationship between the flux-based and threshold-based definitions without requiring self-similarity. We then examine the plume dynamics by evaluating the source terms from the governing plume equations. Our results support neglecting diffusive and viscous effects but emphasise the importance of the mean pressure gradient, even in the self-similar regime. Two coefficients need to be modelled: the well-known entrainment coefficient and the lesser-known momentum correction coefficient, which is a correction required for the momentum equation to account for self-similar and slender approximations. The momentum correction coefficient is found to be approximately constant and slightly greater than the assumed value of 1. The standard entrainment coefficient models perform well up to a local Richardson number three times the asymptotic value but overpredict entrainment for larger Richardson numbers. We propose a correction using the known finite limit of entrainment at infinite Richardson number.

Meehan, Michael Alexander [Sandia National Laborat↗

Collective dynamics of polarized spin-half fermions in relativistic heavy-ion collisions

Standard relativistic hydrodynamics has been successful in describing the properties of the strongly interacting matter produced in the heavy-ion collision experiments. Recently, there has been a significant theoretical advancement in this field to explain spin polarization of hadrons emitted in these processes. Although current models have successfully explained some of the experimental data based on the coupling between spin polarization and vorticity of the medium, they still lack a clear understanding of the differential measurements. This is commonly interpreted as an indication that the spin needs to be treated as an independent degree of freedom whose dynamics is not entirely bound to flow circulation. In particular, if the spin is a macroscopic property of the system, in equilibrium its dynamics should follow hydrodynamic laws. Here, we develop a framework of relativistic hydrodynamics which includes spin degrees of freedom from the quantum kinetic theory for Dirac fermions and use it for modeling the dynamics of matter. Following experimental observations, we assume that the polarization effects are small and derive conservation laws for the net baryon current, the energy–momentum tensor and the spin tensor based on the de Groot–van Leeuwen–van Weert definitions of these currents. We present various properties of the spin polarization tensor and its components, analyze the propagation properties of the spin polarization components, and derive the spin-wave velocity for arbitrary statistics. We find that only the transverse spin components propagate, analogously to the electromagnetic waves. Finally, using our framework, we study the space–time evolution of the spin polarization for the systems respecting certain space–time symmetries and calculate the mean spin polarization per particle, which can be compared to the experimental data. We find that, for some observables, our spin polarization results agree qualitatively with the experimental findings and other model calculations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Agent-based modeling and simulation for the circular economy: Lessons learned and path forward

Circular economy aims at decoupling human activities from resource use and creating wealth. However, many have questioned the link between increased circularity and sustainability, resulting in several methodological approaches being developed to answer that question. This article analyzes and discusses the insights gained from applying agent-based modeling and simulation to study the techno-economic and social conditions promoting circularity and sustainability. This article analyzes the benefits and limitations of this technology and discusses future methodology developments within the circular economy context. Moreover, six limits of the circular economy concept are used to interpret insights from the literature: thermodynamic limits, system boundary limits, limits posed by the physical scale of the economy, limits posed by path dependencies and lock-in, limits of governance and management, and limits of social and cultural definitions. Promising research avenues are to use this methodology with machine learning, industrial ecology methods, and detailed geographic information.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Traceable Black-Box Watermarks For Federated Learning

Due to the distributed nature of Federated Learning (FL) systems, each local client has access to the global model, which poses a critical risk of model leakage. Existing works have explored injecting watermarks into local models to enable intellectual property protection. However, these methods either focus on non-traceable watermarks or traceable but white-box watermarks. We identify a gap in the literature regarding the formal definition of traceable black-box watermarking and the formulation of the problem of injecting such watermarks into FL systems. In this work, we first formalize the problem of injecting traceable black-box watermarks into FL. Based on the problem, we propose a novel server-side watermarking method, TraMark, which creates a traceable watermarked model for each client, enabling verification of model leakage in black-box settings. To achieve this, TraMark partitions the model parameter space into two distinct regions: the main task region and the watermarking region. Subsequently, a personalized global model is constructed for each client by aggregating only the main task region while preserving the watermarking region. Each model then learns a unique watermark exclusively within the watermarking region using a distinct watermark dataset before being sent back to the local client. Extensive results across various FL systems demonstrate that TraMark ensures the traceability of all watermarked models while preserving their main task performance.

Xu, Jiahao [University of Nevada, Reno]↗

The Baseline Performance Reference for Irradiance in PV System Applications

This report proposes the definition of a new baseline performance reference (BPR). The definition goes beyond existing standards pertaining to photovoltaic (PV) reference cells and devices to define the response under all possible operating conditions in the field. Field evaluations using BPR devices will be more sensitive to performance anomalies than pyranometers because they track PV system power output more closely. At the same time, they will be able to detect a broader range of performance anomalies than traditional matched reference devices, which might have matching defects. The BPR definition also opens the door to new practices in resource assessment and yield prediction. Solar resource data can be collected or modeled and validated directly as BPR irradiance, and PV system simulations based on BPR irradiance need fewer assumptions and less processing to obtain the effective irradiance on modules. As a result, lower uncertainty in yield assessments can be expected.

14 SOLAR ENERGY↗

A deep learning approach to identify missing is-a relations in SNOMED CT

Abstract Objective SNOMED CT is the largest clinical terminology worldwide. Quality assurance of SNOMED CT is of utmost importance to ensure that it provides accurate domain knowledge to various SNOMED CT-based applications. In this work, we introduce a deep learning-based approach to uncover missing is-a relations in SNOMED CT. Materials and Methods Our focus is to identify missing is-a relations between concept-pairs exhibiting a containment pattern (ie, the set of words of one concept being a proper subset of that of the other concept). We use hierarchically related containment concept-pairs as positive instances and hierarchically unrelated containment concept-pairs as negative instances to train a model predicting whether an is-a relation exists between 2 concepts with containment pattern. The model is a binary classifier leveraging concept name features, hierarchical features, enriched lexical attribute features, and logical definition features. We introduce a cross-validation inspired approach to identify missing is-a relations among all hierarchically unrelated containment concept-pairs. Results We trained and applied our model on the Clinical finding subhierarchy of SNOMED CT (September 2019 US edition). Our model (based on the validation sets) achieved a precision of 0.8164, recall of 0.8397, and F1 score of 0.8279. Applying the model to predict actual missing is-a relations, we obtained a total of 1661 potential candidates. Domain experts performed evaluation on randomly selected 230 samples and verified that 192 (83.48%) are valid. Conclusions The results showed that our deep learning approach is effective in uncovering missing is-a relations between containment concept-pairs in SNOMED CT.

97 MATHEMATICS AND COMPUTING↗

Advancing representations of equity and justice in climate mitigation futures

THIS PAPER WAS PRIMARILY COMPLETED PRIOR TO THE AUTHOR JOINING PNNL AND NO DOE FUNDING WAS USED FOR THIS PAPER. In this work, we review how equity and justice issues in global climate mitigation scenarios are addressed within Integrated Assessment Models (IAMs) and propose a new research agenda to strengthen their integration in model development and application. We begin by examining prominent concerns at the science-policy interface. We introduce a typology of equity and justice limitations in climate mitigation scenarios, distinguishing among structural, methodological, and epistemological biases that shape what integrated assessment models can reveal at policy-relevant scales. Reflecting on these concerns, we propose a research agenda that describes new avenues of work and draws together distinct emerging initiatives. This agenda is based on the feasibility and depth of required interventions, from incremental improvements to structural reforms and alternative participatory approaches. Drawing on reflexive insights from integrated assessment practitioners, it addresses the operational challenges of translating justice concepts into metrics, including risks of reductionism, tokenism, and narrow definitions. Underlying this research agenda is a recognition that modeling communities must engage more critically with implicit assumptions in model design and use that have equity and justice implications. Achieving equitable climate futures will require transformative actions that integrate diverse justice concerns, advance sustainable development goals, and confront systemic inequities across both human and ecological dimensions. Although models will never capture all these aspects, they can be significantly enhanced to support more informed discussion and practical application. Our contribution proposes a way forward to achieving this goal.

Pachauri, Shonali↗

A taxonomy of constraints in black-box simulation-based optimization

The types of constraints encountered in black-box simulation-based optimization problems differ significantly from those addressed in nonlinear programming. Here, we introduce a characterization of constraints to address this situation. We provide formal definitions for several constraint classes and present illustrative examples in the context of the resulting taxonomy. This taxonomy, denoted KARQ, is useful for modeling and problem formulation, as well as optimization software development and deployment. It can also be used as the basis for a dialog with practitioners in moving problems to increasingly solvable branches of optimization.

42 ENGINEERING↗

Co-optimized machine-learned manifold models for large eddy simulation of turbulent combustion

Many modeling approaches in large eddy simulation (LES) of turbulent combustion employ a projection of the thermochemical state onto a low-dimensional manifold within state space to reduce the number of transported variables and hence computational cost. Flamelet-generated manifolds (FGM) is an example of a well-established, physics-based approach, but increasingly, principal component analysis (PCA) is being used as a data-driven method for generating manifold models. For both approaches, the nonlinear relationship between the location on the predefined manifold and the outputs of interest, such as reaction rates, can be tabulated or encoded in a neural network. This work proposes a new approach for manifold modeling that extends these existing approaches. A modified neural network structure simultaneously encodes the definition of the manifold variables, the nonlinear mapping, and the subfilter closure for LES. This allows all three of these aspects of the model to be co-optimized, generating a model from any source of combustion thermochemical state data. The manifold parameterizing variables are constrained to be linear combinations of species, as in FGM and PCA-based models, to aid in interpretability and implementation. For LES, subfilter variances of the manifold variables are also included as inputs. Two types of a priori analysis are performed to evaluate the new approach. In the first, the model is trained on data from one-dimensional premixed flames. In this case, the approach recovers the behavior of flamelet-based manifold approaches, and in fact slightly improves performance by identifying an optimized progress variable. The approach is also applied to data from direct numerical simulations of spherical ignition kernels in isotropic turbulence. For any specified manifold dimensionality, the new approach provides substantially lower prediction errors than a PCA-based model developed from the same data set. Additionally, the LES formulation of the new approach can provide accurate predictions for filtered reaction rates across a variety of filter widths.

97 MATHEMATICS AND COMPUTING↗

Towards philosophical reasoning with agentic LLMs: Socratic method for scientific assistance

As large language models (LLMs) become central tools in science, improving their reasoning capabilities is critical for meaningful and trustworthy applications. We introduce a Socratic agent for scientific reasoning, implemented through a structured system prompt that guides LLMs via classical principles of inquiry. Unlike typical prompt engineering or retrieval-based methods, our approach leverages definition, analogy, hypothesis elimination, and other Socratic techniques to generate more coherent, critical, and domain-aware responses. We evaluate the agent across diverse scientific domains and benchmark it on the abstraction and reasoning corpus challenge dataset, achieving 97.15% under a fixed prompting protocol and without fine-tuning or external tools. Expert evaluation shows improved reasoning depth, clarity, and adaptability over conventional LLM outputs, suggesting that structured prompting rooted in philosophical reasoning can improve the scientific utility of language models.

LLM reasoning↗

Selective mass scaling for single-layer thick shell elements in DYNA3D

Hexahedral elements can be adapted to model thin and moderately thick structures by neglecting the coupling of through-thickness stress, resulting in a fully three-dimensional, but simplified, state of stress. These specialized elements, often referred to as “thick” or “solid” shells, are generally employed to model thin-walled structures using continuum mechanics-based material models. In explicit dynamics simulations, where computational speed is important, these elements are integrated with a single quadrature point and a set of anti-hourglassing (stabilizing) forces. Thick shells, by definition, have a thickness dimension smaller than their in-plane dimensions, and this small thickness often determines the stable time step size in simulations, despite the mechanics being approximated. To alleviate this limitation while retaining the relevant dynamics of thin-walled structures, selective mass scaling (SMS), or selective mass “augmentation,” has been proposed in the literature. In this technical report, we explore the application of SMS to single-layer thick shells in the simulation software DYNA3D.

42 ENGINEERING↗

Estimation of pipe failure frequencies in the absence of operational experience data: A pilot study

Probabilistic failure metrics such as leak frequency and rupture frequency are commonly used to characterize piping reliability. The methodologies for calculating the failure metrics rely on a complex set of input parameters. Operating experience data and experimental data play an important role in informing the different input parameters. The paper describes results and conclusions of a coordinated research project to benchmark three different reliability models using a four-step procedure: reference case definition of relevance to advanced reactor designs, input parameter calibration, validation of results, and application of different methodologies upon completion of the calibration and validation steps. The reference case is a weld consisting of nickel-base alloy 152/52 and located within a primary pressure boundary of an advanced reactor. This alloy is a class of structural materials known to be highly resistant to stress corrosion cracking. Synergies between the different methods are noted and the importance of a multi-disciplinary approach to input parameter development is underscored. A key conclusion is that the three methods are equally suitable for estimating failure frequencies. In any specific application, a selection of the most practical or effective computational tool can be considered. The comparison of alternative models confirms and helps to gain confidence in the computed failure frequency estimates. The study was part of a coordinated research project organized by the International Atomic Energy Agency.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Particle configurations in the $NN\bar K$ system

Three-body AAB model for the $NN\bar K$(s NN = 0) kaonic cluster is considered based on the configuration space Faddeev equations. Within a single channel approach, the difference between masses of nucleons and kaons and the charge independence breaking of nucleon-nucleon interaction are taken into consideration. We definite the particle configurations in the system according to the particle masses and pair potentials. There are two sets of the particle configurations, ppK¯, npK¯ 0 and nnK¯ 0 , npK¯, charged and neutral. The three-body calculations are performed by applying NN and NK¯ phenomenological isospin-dependent potentials. The mass and energy spectra related to the particle configurations are presented. We evaluate the mass and energy uncertainties for the NNK¯ model. As a result, an analogy to NNN model for the 3 H and 3 He nuclei is proposed.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Physics-informed machine learning for building performance simulation-A review of a nascent field

Building performance simulation (BPS) is critical for understanding building dynamics and behavior, analyzing the performance of the built environment, optimizing energy efficiency, improving demand flexibility, and enhancing building resilience. However, conducting BPS is not trivial. Traditional BPS relies on accurate building energy models, which are primarily physics-based and heavily dependent on detailed building information, expert knowledge, and case-by-case model calibrations, significantly limiting their scalability. With the development of sensing technology and the increased availability of data, there is growing attention and interest in data-driven BPS. However, purely data-driven models often suffer from limited generalization ability and a lack of physical consistency, resulting in poor performance in real-world applications. To address these limitations, recent studies have begun integrating physics priors into data-driven models, a methodology known as physics-informed machine learning (PIML). PIML is an emerging field where its definitions, methodologies, evaluation criteria, application scenarios, and future directions remain open. To bridge those gaps, this study systematically reviews the state-of-the-art PIML for BPS, offering a comprehensive definition of PIML and comparing it to traditional BPS approaches regarding data requirements, modeling effort, performance, and computational cost. We also summarize the commonly used methodologies, validation approaches, application domains, available data sources, open-source packages, and testbeds. In addition, this study provides a general guideline for selecting appropriate PIML models based on BPS applications. Finally, this study identifies key challenges and outlines future research directions, providing a solid foundation and valuable insights to advance R&D of PIML in BPS.

Jiang, Zixin↗

AOI 3 Life Modelling of Critical Steam Cycle Components in Coal-Fueled Power Plants

Microstructural damage accumulation models have been used to produce calibrated life estimation models for a DR22/P22 steel wye-block welds, and a Jethete stainless steel turbine blade (bucket). The calibrated life estimation models will aid the power plant operator in determining optimal maintenance and operation schedules based upon historical operational data as well as current, or future operation schemes. The impact of this project will enable existing coal-fueled power plants to operate safely for longer periods of time and at higher efficiencies, thereby reducing the economic and environmental impact of the existing coal power plant fleet. Testing, characterization, and modelling indicates that the operating life of P22 pipelines and their welds are dominated by fatigue damage mechanisms. Specifically, fatigue is of no concern in these materials when operating under realistic conditions manifesting in the main steam piping of coal-fueled power plants. However, if a low-temperature overload ever occurs during operation, fatigue will manifest as a damage mechanism of interest. Primary impact provided by the completion of this work is in the manifestation of a detailed ABAQUS solid model providing accurate boundary conditions to enable the prediction of operational stresses and strains. The completed plant solid model, in conjunction with the calibrated fatigue and creep life models provide definitive confirmation that creep is the dominant damage mechanism during operational conditions. Jethete life modelling has been completed by the manifestation of a material-specific, and temperature-specific, Kitagowa diagram. The Kitagowa diagram provides maintenance and operation decision making with scientifically-based go/no-go support based upon the crack-like features that have been identified by use of inspection.

01 COAL, LIGNITE, AND PEAT↗