Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model based definitions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Computational fluid dynamics analysis of Space Shuttle main engine multiple plume flows at high-altitude flight conditions

Computational fluid dynamics (CFD) analysis is providing verification of Space Shuttle flight performance details and is being applied to Space Shuttle Main Engine Multiple plume interaction flow field definition. Advancements in real-gas CFD methodology that are described have allowed definition of exhaust plume flow details at Mach 3.5 and 107,000 ft. The specific objective includes the estimate of flow properties at oblique shocks between plumes and plume recirculation into the Space Shuttle Orbiter base so that base heating and base pressure can be modeled accurately. The approach utilizes the Rockwell USA Real Gas 3-D Navier-Stokes (USARG3D) Code for the analysis. The code has multi-zonal capability to detail the geometry of the plumes based region and utilizes finite-rate chemistry to compute the plume expansion angle and relevant flow properties at altitude correctly. Through an improved definition of the base recirculation flow properties, heating, and aerodynamic design environments of the Space Shuttle Vehicle can be further updated.

Dougherty, N. S.↗

Design and Optimization of a 3-Stage Axial Supercritical CO2 Compressor

This paper presents the detailed design and optimization of a three-stage axial supercritical carbon dioxide (sCO2) compressor to support the advancement of sCO 2 power cycles, recognized for their efficiency and compactness in energy conversion. The three-stage design is a scaled model of the original nine-stage 100 MW design. The aerodynamic design and structural analysis were performed at the University of Cincinnati, and the results are presented. The first stage of the design has been meticulously manufactured and experimentally tested at University of Notre Dame Turbomachinery Laboratory. The design process starts with the preliminary design process, considering the required operating boundary conditions and geometrical constraints for full stage design. The design is then scaled down to 3-stage based on the power limitation. The definition for design parameters has been inspired from the EEE HPC design and tweaked to fit the sCO 2 operation. Axisymmetric analysis is performed next and the parametric geometry modeler -Tblade3 has been used to create 3D blade geometry from the results which can then be exported for the 3D CFD analysis and optimization. An efficiency of 89.85% is predicted for the final design at the design point.

Compressor↗

Accelerating full-waveform inversion using source stacking: synthetic experiments at the global scale in a realistic 3-D earth model

SUMMARY The spectral element method is currently the method of choice for computing accurate synthetic seismic wavefields in realistic 3-D earth models at the global scale. However, it requires significantly more computational time, compared to normal mode-based approximate methods. Source stacking, whereby multiple earthquake sources are aligned on their origin time and simultaneously triggered, can reduce the computational costs by several orders of magnitude. We present the results of synthetic tests performed on a realistic radially anisotropic 3-D model, slightly modified from model SEMUCB-WM1 with three component synthetic waveform ‘data’ for a duration of 10 000 s, and filtered at periods longer than 60 s, for a set of 273 events and 515 stations. We consider two definitions of the misfit function, one based on the stacked records at individual stations and another based on station-pair cross-correlations of the stacked records. The inverse step is performed using a Gauss–Newton approach where the gradient and Hessian are computed using normal mode perturbation theory. We investigate the retrieval of radially anisotropic long wavelength structure in the upper mantle in the depth range 100–800 km, after fixing the crust and uppermost mantle structure constrained by fundamental mode Love and Rayleigh wave dispersion data. The results show good performance using both definitions of the misfit function, even in the presence of realistic noise, with degraded amplitudes of lateral variations in the anisotropic parameter ξ. Interestingly, we show that we can retrieve the long wavelength structure in the upper mantle, when considering one or the other of three portions of the cross-correlation time series, corresponding to where we expect the energy from surface wave overtone, fundamental mode or a mixture of the two to be dominant, respectively. We also considered the issue of missing data, by randomly removing a successively larger proportion of the available synthetic data. We replace the missing data by synthetics computed in the current 3-D model using normal mode perturbation theory. The inversion results degrade with the proportion of missing data, especially for ξ, and we find that a data availability of 45 per cent or more leads to acceptable results. We also present a strategy for grouping events and stations to minimize the number of missing data in each group. This leads to an increased number of computations but can be significantly more efficient than conventional single-event-at-a-time inversion. We apply the grouping strategy to a real picking scenario, and show promising resolution capability despite the use of fewer waveforms and uneven ray path distribution. Source stacking approach can be used to rapidly obtain a starting 3-D model for more conventional full-waveform inversion at higher resolution, and to investigate assumptions made in the inversion, such as trade-offs between isotropic, anisotropic or anelastic structure, different model parametrizations or how crustal structure is accounted for.

Geochemistry & Geophysics↗

A data management system for engineering and scientific computing

Data elements and relationship definition capabilities for this data management system are explicitly tailored to the needs of engineering and scientific computing. System design was based upon studies of data management problems currently being handled through explicit programming. The system-defined data element types include real scalar numbers, vectors, arrays and special classes of arrays such as sparse arrays and triangular arrays. The data model is hierarchical (tree structured). Multiple views of data are provided at two levels. Subschemas provide multiple structural views of the total data base and multiple mappings for individual record types are supported through the use of a REDEFINES capability. The data definition language and the data manipulation language are designed as extensions to FORTRAN. Examples of the coding of real problems taken from existing practice in the data definition language and the data manipulation language are given.

Elliot, L.↗

User modeling techniques for enhanced usability of OPSMODEL operations simulation software

The PC based OPSMODEL operations software for modeling and simulation of space station crew activities supports engineering and cost analyses and operations planning. Using top-down modeling, the level of detail required in the data base can be limited to being commensurate with the results required of any particular analysis. To perform a simulation, a resource environment consisting of locations, crew definition, equipment, and consumables is first defined. Activities to be simulated are then defined as operations and scheduled as desired. These operations are defined within a 1000 level priority structure. The simulation on OPSMODEL, then, consists of the following: user defined, user scheduled operations executing within an environment of user defined resource and priority constraints. Techniques for prioritizing operations to realistically model a representative daily scenario of on-orbit space station crew activities are discussed. The large number of priority levels allows priorities to be assigned commensurate with the detail necessary for a given simulation. Several techniques for realistic modeling of day-to-day work carryover are also addressed.

Davis, William T.↗

Fault diagnosis based on continuous simulation models

The results are described of an investigation of techniques for using continuous simulation models as basis for reasoning about physical systems, with emphasis on the diagnosis of system faults. It is assumed that a continuous simulation model of the properly operating system is available. Malfunctions are diagnosed by posing the question: how can we make the model behave like that. The adjustments that must be made to the model to produce the observed behavior usually provide definitive clues to the nature of the malfunction. A novel application of Dijkstra's weakest precondition predicate transformer is used to derive the preconditions for producing the required model behavior. To minimize the size of the search space, an envisionment generator based on interval mathematics was developed. In addition to its intended application, the ability to generate qualitative state spaces automatically from quantitative simulations proved to be a fruitful avenue of investigation in its own right. Implementations of the Dijkstra transform and the envisionment generator are reproduced in the Appendix.

Feyock, Stefan↗

GMI-IPS: Python Processing Software for Aircraft Campaigns

NASA's Atmospheric Tomography Mission (ATom) seeks to understand the impact of anthropogenic air pollution on gases in the Earth's atmosphere. Four flight campaigns are being deployed on a seasonal basis to establish a continuous global-scale data set intended to improve the representation of chemically reactive gases in global atmospheric chemistry models. The Global Modeling Initiative (GMI), is creating chemical transport simulations on a global scale for each of the ATom flight campaigns. To meet the computational demands required to translate the GMI simulation data to grids associated with the flights from the ATom campaigns, the GMI ICARTT Processing Software (GMI-IPS) has been developed and is providing key functionality for data processing and analysis in this ongoing effort. The GMI-IPS is written in Python and provides computational kernels for data interpolation and visualization tasks on GMI simulation data. A key feature of the GMI-IPS, is its ability to read ICARTT files, a text-based file format for airborne instrument data, and extract the required flight information that defines regional and temporal grid parameters associated with an ATom flight. Perhaps most importantly, the GMI-IPS creates ICARTT files containing GMI simulated data, which are used in collaboration with ATom instrument teams and other modeling groups. The initial main task of the GMI-IPS is to interpolate GMI model data to the finer temporal resolution (1-10 seconds) of a given flight. The model data includes basic fields such as temperature and pressure, but the main focus of this effort is to provide species concentrations of chemical gases for ATom flights. The software, which uses parallel computation techniques for data intensive tasks, linearly interpolates each of the model fields to the time resolution of the flight. The temporally interpolated data is then saved to disk, and is used to create additional derived quantities. In order to translate the GMI model data to the spatial grid of the flight path as defined by the pressure, latitude, and longitude points at each flight time record, a weighted average is then calculated from the nearest neighbors in two dimensions (latitude, longitude). Using SciPya's Regular Grid Interpolator, interpolation functions are generated for the GMI model grid and the calculated weighted averages. The flight path points are then extracted from the ATom ICARTT instrument file, and are sent to the multi-dimensional interpolating functions to generate GMI field quantities along the spatial path of the flight. The interpolated field quantities are then written to a ICARTT data file, which is stored for further manipulation. The GMI-IPS is aware of a generic ATom ICARTT header format, containing basic information for all flight campaigns. The GMI-IPS includes logic to edit metadata for the derived field quantities, as well as modify the generic header data such as processing dates and associated instrument files. The ICARTT interpolated data is then appended to the modified header data, and the ICARTT processing is complete for the given flight and ready for collaboration. The output ICARTT data adheres to the ICARTT file format standards V1.1. The visualization component of the GMI-IPS uses Matplotlib extensively and has several functions ranging in complexity. First, it creates a model background curtain for the flight (time versus model eta levels) with the interpolated flight data superimposed on the curtain. Secondly, it creates a time-series plot of the interpolated flight data. Lastly, the visualization component creates averaged 2D model slices (longitude versus latitude) with overlaid flight track circles at key pressure levels. The GMI-IPS consists of a handful of classes and supporting functionality that have been generalized to be compatible with any ICARTT file that adheres to the base class definition. The base class represents a generic ICARTT entry, only defining a single time entry and 3D spatial positioning parameters. Other classes inherit from this base class; several classes for input ICARTT instrument files, which contain the necessary flight positioning information as a basis for data processing, as well as other classes for output ICARTT files, which contain the interpolated model data. Utility classes provide functionality for routine procedures such as: comparing field names among ICARTT files, reading ICARTT entries from a data file and storing them in data structures, and returning a reduced spatial grid based on a collection of ICARTT entries. Although the GMI-IPS is compatible with GMI model data, it can be adapted with reasonable effort for any simulation that creates Hierarchical Data Format (HDF) files. The same can be said of its adaptability to ICARTT files outside of the context of the ATom mission. The GMI-IPS contains just under 30,000 lines of code, eight classes, and a dozen drivers and utility programs. It is maintained with GIT source code management and has been used to deliver processed GMI model data for the ATom campaigns that have taken place to date.

Damon, M. R.↗

Experimental and Theoretical Studies of Pulsating Turbulent Flow

The objective of this investigation was to study the effects of small amplitude sinusoidal pulsations on fully developed turbulent flow in a tube from both experimental and theoretical viewpoints. Theoretical models for the macroscopic behavior of pulsating turbulent tube flow were developed for the two cases of very low and very high pulsation frequencies. The models are based on assumptions of quasi-steady and frozen eddy viscosity flow behavior, respectively. The models successfully predict unsteady velocity profiles, thereby supporting the currently proposed definitions of frequency regimes in pulsating turbulent flow. Experimental measurements were made of the time-dependent pressure drop and velocity profiles over the range of frequency-to-Reynolds number ratios from 0.0095 to 0.24. The two macroscopic models developed in this study predict unsteady velocity profiles which are in moderately good agreement with the experiments in their respective frequency regimes, and a previously developed quasi-steady model is found to predict experimental velocity profiles well in both the quasisteady and the frozen eddy viscosity frequency regimes. The effect of flow pulsations on the dissipation of turbulence energy in the vicinity of the wall was measured in the lower transition frequency regime. The long-time averaged dissipation was observed to be unchanged from the steady flow dissipation, within the accuracy of the experiment. A theoretical model of the periodic viscous sublayer was also developed and applied to pulsating flow in a tube, in order to investigate the effects of flow pulsations on the rate of production of turbulence in the region of the wall. The periodic viscous sublayer model predicts sublayer growth periods in steady flow which agree with the published experimental data. When the model is applied to pulsating flow, the response of the sublayer growth period falls into three frequency regimes, the parameters of which are in approximate agreement with the frequency regimes which are defined on the basis of macroscopic flow behavior. The sublayer renewal cycle exhibits quasi-steady flow behavior when the sublayer growth period is much less than the pulsation period, transition behavior when these two periods are approximately equal, and frozen eddy viscosity behavior when the sublayer period is much longer than the pulsation period. The effect of the sublayer growth and renewal cycle on the level of turbulence was investigated by two methods. The velocity fluctuations seen by a point velocity probe located close to the wall were predicted from the model in one method and the rate of turbulence production was estimated from the frequency of sublayer renewal events in the other.

Kingston, G. C.↗

Systems biology markup language (SBML) level 3 package: multistate, multicomponent and multicompartment species, version 1, release 2

Rule-based modeling is an approach that permits constructing reaction networks based on the specification of rules for molecular interactions and transformations. These rules can encompass details such as the interacting sub-molecular domains and the states and binding status of the involved components. Conceptually, fine-grained spatial information such as locations can also be provided. Through “wildcards” representing component states, entire families of molecule complexes sharing certain properties can be specified as patterns. This can significantly simplify the definition of models involving species with multiple components, multiple states, and multiple compartments. The systems biology markup language (SBML) Level 3 Multi Package Version 1 extends the SBML Level 3 Version 1 core with the “type” concept in the Species and Compartment classes. Therefore, reaction rules may contain species that can be patterns and exist in multiple locations. Multiple software tools such as Simmune and BioNetGen support this standard that thus also becomes a medium for exchanging rule-based models. This document provides the specification for Release 2 of Version 1 of the SBML Level 3 Multi package. No design changes have been made to the description of models between Release 1 and Release 2; changes are restricted to the correction of errata and the addition of clarifications.

59 BASIC BIOLOGICAL SCIENCES↗

Thermal Modeling Method Improvements for SAGE III on ISS

The Stratospheric Aerosol and Gas Experiment III (SAGE III) instrument is the fifth in a series of instruments developed for monitoring aerosols and gaseous constituents in the stratosphere and troposphere. SAGE III will be delivered to the International Space Station (ISS) via the SpaceX Dragon vehicle. A detailed thermal model of the SAGE III payload, which consists of multiple subsystems, has been developed in Thermal Desktop (TD). Many innovative analysis methods have been used in developing this model; these will be described in the paper. This paper builds on a paper presented at TFAWS 2013, which described some of the initial developments of efficient methods for SAGE III. The current paper describes additional improvements that have been made since that time. To expedite the correlation of the model to thermal vacuum (TVAC) testing, the chambers and GSE for both TVAC chambers at Langley used to test the payload were incorporated within the thermal model. This allowed the runs of TVAC predictions and correlations to be run within the flight model, thus eliminating the need for separate models for TVAC. In one TVAC test, radiant lamps were used which necessitated shooting rays from the lamps, and running in both solar and IR wavebands. A new Dragon model was incorporated which entailed a change in orientation; that change was made using an assembly, so that any potential additional new Dragon orbits could be added in the future without modification of the model. The Earth orbit parameters such as albedo and Earth infrared flux were incorporated as time-varying values that change over the course of the orbit; despite being required in one of the ISS documents, this had not been done before by any previous payload. All parameters such as initial temperature, heater voltage, and location of the payload are defined based on the case definition. For one component, testing was performed in both air and vacuum; incorporating the air convection in a submodel that was only built for the in-air cases allowed correlation of all testing to be done in a single model. These modeling improvements and more will be described and illustrated in the paper.

Liles, Kaitlin↗

Space Shuttle noise suppression concepts for the Eastern Test Range

The basic objectives of the Space Shuttle noise suppression program for the Eastern Test Range were the definition of the acoustic environment of the Shuttle and the adjacent ground plane for both on-pad and liftoff conditions and the definition of realistic noise suppression techniques and modification of the launch facility that could reduce engine noise associated with supersonic flow. Scaling considerations for acoustic model testing based on the principle of dynamic similarity are detailed. Suppression approaches are described with emphasis placed on barriers and shields: exhaust flow trench covers, solid dividers between the SSME and SRB exhaust flows and a crossed pipe water injection system over the SSME exhaust flow trench. Graphs are presented summarizing noise data gathered for various noise sources and using different suppression approaches.

Guest, S. H.↗

Repair Concepts as Design Constraints of a Stiffened Composite PRSEUS Panel

A design and analysis of a repair concept applicable to a stiffened thin-skin composite panel based on the Pultruded Rod Stitched Efficient Unitized Structure is presented. The concept is a bolted repair using metal components, so that it can easily be applied in the operational environment. The damage scenario considered is a midbay-to-midbay saw-cut with a severed stiffener, flange and skin. In a previous study several repair configurations were explored and their feasibility confirmed but refinement was needed. The present study revisits the problem under recently revised design requirements and broadens the suite of loading conditions considered. The repair assembly design is based on the critical tension loading condition and subsequently its robustness is verified for a pressure loading case. High fidelity modeling techniques such as mesh-independent definition of compliant fasteners, elastic-plastic material properties for metal parts and geometrically nonlinear solutions are utilized in the finite element analysis. The best repair design is introduced, its analysis results are presented and factors influencing the design are assessed and discussed.

Przekop, Adam↗

STARTR: An Open-Source MARVEL model for the NRIC Virtual Test Bed [Poster]

The National Reactor Innovation Center (NRIC) seeks to improve the understanding of microreactor physics in industry and academia through the development of a Microreactor Applications Research Validation and Evaluation (MARVEL) reactor-based model, published on the Virtual Test Bed (VTB). To achieve this goal, the Sodium-cooled Thermal-spectrum Advanced Research Test Reactor (STARTR) model was built using publicly available MARVEL specifications where possible and approximations where applicable, and was optimized for fast runtimes for researchers to receive rapid simulation feedback. STARTR will fill a gap between stakeholder interest and available models, as the first Sodium-cooled Thermal Reactor (STR) hosted on the VTB with baseline performance sanctioned by INL. This project involved the definition of all materials used in the reactor, geometry and all reactor subcomponents, and assertion of tallies and simulation settings within OpenMC 0.13.3. This poster details a small subset of the overall reactor physics testing: the two-dimensional power peaking factors and the flux energy spectrum, as well as plots of the created geometry. Future work includes code-to-code verification between the OpenMC-based model and a separately designed MCNP 6.2-based model.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Development of the Artemis Distributed Simulation FOMs

The National Aeronautics and Space Administration (NASA) is formulating and developing the Artemis Program, a collaboration with domestic commercial and international partners that will establish a long term human presence on the Moon and extend human exploration beyond the Earth-Moon system ahead of exploring Mars. These Artemis partners are developing a portfolio of space and surface systems to support human missions to the lunar surface and beyond. The Artemis systems will provide the mobility, habitation, and logistics infrastructure that will support human exploration and foster robust scientific investigations. Each partner will contribute one or more elements to the Artemis Program with NASA having the overarching responsibility for defining the Artemis architecture and guiding the integration of this complex system of space systems. To successfully accomplish this audacious task, NASA will rely on the development and execution of many complex models and simulations. Many of these simulations will be provided by the Artemis partners. While each of these simulations will provide important insight into the characteristics and performance of an associated system, individually they will not provide insight into the integrated performance of the architecture and the system of systems working in concert to execute a given Artemis mission. To address this need, NASA is developing a distributed simulation capability called the Artemis Distributed Simulation (ADS). ADS’s distributed nature supports the complex aggregation of constituent Artemis element simulations. Artemis partner simulations will be able to join into an ADS-based distributed simulation and interact with other Artemis element simulations while limiting the exposure of proprietary designs and data. ADS is defining a distributed simulation capability built on international simulation interoperability standards, specifically the High Level Architecture (HLA) and the Space Reference Federation Object Model (SpaceFOM). While HLA and SpaceFOM provide the substantive necessary technology basis for ADS, additional common datatypes, message definitions, and execution protocols are required. These extensions constitute the ADS Federation Object Model (FOM). This paper describes the fundamental architectural elements of ADS and the FOM extensions needed to support the complex nature of the Artemis Program. This includes the examination of the ADS FOM modules, ADS base datatypes, ADS SpaceFOM Object Class extensions, new ADS Object Classes, and new ADS Interaction Classes.

HLA↗

Development of the Artemis Distributed Simulation FOMs

The National Aeronautics and Space Administration (NASA) is formulating and developing the Artemis Program, a collaboration with domestic commercial and international partners that will establish a long term human presence on the Moon and extend human exploration beyond the Earth-Moon system ahead of exploring Mars. These Artemis partners are developing a portfolio of space and surface systems to support human missions to the lunar surface and beyond. The Artemis systems will provide the mobility, habitation, and logistics infrastructure that will support human exploration and foster robust scientific investigations. Each partner will contribute one or more elements to the Artemis Program with NASA having the overarching responsibility for defining the Artemis architecture and guiding the integration of this complex system of space systems. To successfully accomplish this audacious task, NASA will rely on the development and execution of many complex models and simulations. Many of these simulations will be provided by the Artemis partners. While each of these simulations will provide important insight into the characteristics and performance of an associated system, individually they will not provide insight into the integrated performance of the architecture and the system of systems working in concert to execute a given Artemis mission. To address this need, NASA is developing a distributed simulation capability called the Artemis Distributed Simulation (ADS). ADS’s distributed nature supports the complex aggregation of constituent Artemis element simulations. Artemis partner simulations will be able to join into an ADS-based distributed simulation and interact with other Artemis element simulations while limiting the exposure of proprietary designs and data. ADS is defining a distributed simulation capability built on international simulation interoperability standards, specifically the High Level Architecture (HLA) and the Space Reference Federation Object Model (SpaceFOM). While HLA and SpaceFOM provide the substantive necessary technology basis for ADS, additional common datatypes, message definitions, and execution protocols are required. These extensions constitute the ADS Federation Object Model (FOM). This paper describes the fundamental architectural elements of ADS and the FOM extensions needed to support the complex nature of the Artemis Program. This includes the examination of the ADS FOM modules, ADS base datatypes, ADS SpaceFOM Object Class extensions, new ADS Object Classes, and new ADS Interaction Classes.

HLA↗

An encoder–decoder LSTM-based EMPC framework applied to a building HVAC system

Numerous studies have demonstrated the benefit of economic model predictive control (EMPC) applied to building heating, ventilation, and air conditioning (HVAC) systems. However, the construction and training of predictive models for building HVAC systems are widely recognized as a key technological barrier preventing large-scale adoption of EMPC for buildings. In this work, an encoder–decoder long short-term memory-based EMPC framework is developed. The key advantage of the approach is that a model may be automatically generated from a list of inputs and outputs. From the definition of inputs and outputs, the constructed model may be trained and automatically embedded into the EMPC framework for real-time estimation and control. The overall end-to-end EMPC framework from model training to on-line estimation and control are described. To this end, the encoder–decoder model provides a natural framework for state estimation (encoder), which is required to provide an initial condition for the predictive model of EMPC (decoder). Closed-loop simulations using EnergyPlus are performed to demonstrate the approach. The simulated closed-loop system consists of a building zone from a multi-zone building, which is served by an air handling unit-variable air volume HVAC system. For the HVAC example considered, the trained encoder–decoder model can predict the indoor air temperature and HVAC sensible cooling rate of a building zone over a two-day horizon with high accuracy. Overall, we find that considering a time-of-use electric rate structure, the EMPC, which manipulates the zone temperature setpoint, can reduce the HVAC power consumption cost relative to keeping the zone temperature setpoint at its maximum value (i.e., minimum energy approach).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

Magnetograph response to canopy-type fields

The response of longitudinal-field magnetographs to magnetic fields which are semi-infinite or confined to a horizontal layer is discussed with respect to the interpretation of solar diffuse fields, observed towards the limb, in terms of magnetic canopy models. Numerical results are presented for several reference solar models and typical 'calibration' curves are shown for the CI 9111 A, Fe I 8688 A, and Ca II 8542 A lines in magnetostatic atmospheres derived from a mean model. A procedure is developed for determining the base heights of magnetic canopies from observations with an uncertainty not exceeding the order of a pressure scale height. Until definitive information regarding atmospheric structure inside flux tubes can be developed from theory or observation, reliable field strengths cannot be derived from the data.

Jones, H. P.↗