Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “virtual training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Real Time-Optimal Power Flow-Based Distributed Energy Resource Management System (DERMS)

This project aims to promote lab-proven clean energy technology to commercially scalable versions of the technology, integrate the technology with broader systems, provide extended performance data, and validate the manufacturability and reliability of the technology. The lab-proven technology, RT-OPF DERMS, was developed and validated through previous U.S. Department of Energy-funded efforts, including Advanced Research Projects Agency-Energy funding under the Network Optimized Distributed Energy Systems program and Holy-Cross Energy High Impact Project. In the Advanced Research Projects Agency-Energy Network Optimized Distributed Energy Systems project, the RT-OPF DERMS was developed and implemented in multiple hardware platforms, demonstrating its performance and capabilities in the lab and field environments. The technology was also evaluated and matured via a participation in the U.S. Department of Energy I-Corps program, whose goal is to pair teams of researchers with industry mentors for an intensive 2-month training in which the researchers define technology value propositions, conduct customer discovery interviews, and develop viable market pathways for their technologies. These activities indicate the high technology maturity and Technology Readiness Level of the RT-OPF DERMS.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Predicting beam transmission using 2-dimensional phase space projections of hadron accelerators

We present a method to compress the 2D transverse phase space projections from a hadron accelerator and use that information to predict the beam transmission. This method assumes that obtaining at least three projections of the 4D transverse phase space is possible and that an accurate simulation model is available for the beamline. Using a simulated model, we show that—a computer can train a convolutional autoencoder to reduce phase-space information which can later be used to predict the beam transmission. Finally, we argue that although using projections from a realistic nonlinear distribution produces less accurate results, the method still generalizes well.

43 PARTICLE ACCELERATORS↗

Neural Network Analysis of Nuclear Magnetic Resonance and Infrared Spectra

Nuclear magnetic resonance (NMR) spectroscopy and infrared (IR) spectroscopy are powerful chemical characterization techniques with broad general usage. However, the manual evaluation of the resulting spectra is time-consuming and requires significant expertise, preventing insights from being used in real-time applications. With recent advances in computation and artificial intelligence (AI), new tools are available for automating spectral interpretation. In this work, machine learning (ML) algorithms using 1-dimensional convolutional neural networks (CNNs) were applied to identify common functional groups from spectral information. Raw spectra were collected virtually from the Human Metabolome Database (HMDB) and National Institute of Standards and Technology (NIST) Chemistry WebBook and processed into a suitable standard. Algorithm design was tailored to best fit the nature of the problem, with built-in flexibility to accommodate relevant parameters beyond the raw spectral input, specifically solvent identity and magnetic frequency for NMR. The predictive capability of the algorithm in identifying functional groups is displayed in several examples. This methodology has been compiled into a code repository and could easily be modified to adapt alternative data sources, including other spectrum types. To mitigate overfitting, a common problem in mathematical modeling where overfamiliarity with training data produces trends that are not representative of the general data, a novel metric was developed, referred to as Accufit. Accufit includes a parameter that penalizes substantial differences in the training accuracy and the accuracy of an independent validation set. Examples are presented showing the effectiveness of Accufit in maintaining the model’s predictive capability while controlling the overfitting when used as a custom metric for hyperparameter tuning.

Sturgill, James↗

An Autonomous Critical Data Extrapolator for the AGN-201m

Nuclear nonproliferation serves as a key goal, being undertaken by the International Atomic Energy Agency (IAEA). To recognize proliferation there are two pathways that states, who intend to use nuclear material for malicious purposes can take, diversion can misuse. Diversion is when fissile nuclear material is declared to the IAEA for non-weapon purposes, but then covertly removed. If the source of nuclear material, that is not declared and not fissionable, is placed inside the reactor core to create fissile material used to create weapons then the state is using the second pathway of proliferation, misuse. With the emerging development in areas of simulation and machine learning the creation of virtual models of reactor systems, digital twins, serve as a potential method to identify proliferation through detecting anomalous behavior in the reactor. A digital twin for a physical nuclear reactor has never been developed, as digital twins serve as an emerging technology. To investigate the process for development and use of a digital twin for a nuclear reactor Idaho State University’s AGN-201m serves as the nuclear reactor used for development of this digital twin. A data acquisition system has been installed to the reactor system allowing for the transfer of collected data from a reactor operation to Idaho National Laboratory’s Deeplynx data warehouse. When utilizing data to train reactor physics and machine learning models, a significant challenge encountered is the initial state of the data. Nuclear proliferation will have the capacity to be detected when the reactor immediately starts up, nor will it occur after the reactor shuts down. Generally, it will be detected when the reactor is operating at some desired power over a sufficient period for that specific reactor design. For the AGN-201m this will be when the reactor is critical (generally 1 mW or above) for a timespan that is within or less than the range of a regular business day. Datasets sent to Deeplynx have had to be manually cut to when the reactor is critical based on plots of power levels. This method is inefficient and laborious, especially when using multiple datasets at once to train a model. To provide a more streamlined approach an automated critical data extrapolator is developed, with capabilities of recognizing when the reactor operation first reaches criticality, and when the reactor undergoes a SCRAM and is shutdown.

99 GENERAL AND MISCELLANEOUS↗

SARS-CoV2 billion-compound docking

Abstract This dataset contains ligand conformations and docking scores for 1.4 billion molecules docked against 6 structural targets from SARS-CoV2, representing 5 unique proteins: MPro, NSP15, PLPro, RDRP, and the Spike protein. Docking was carried out using the AutoDock-GPU platform on the Summit supercomputer and Google Cloud. The docking procedure employed the Solis Wets search method to generate 20 independent ligand binding poses per compound. Each compound geometry was scored using the AutoDock free energy estimate, and rescored using RFScore v3 and DUD-E machine-learned rescoring models. Input protein structures are included, suitable for use by AutoDock-GPU and other docking programs. As the result of an exceptionally large docking campaign, this dataset represents a valuable resource for discovering trends across small molecule and protein binding sites, training AI models, and comparing to inhibitor compounds targeting SARS-CoV-2. The work also gives an example of how to organize and process data from ultra-large docking screens.

60 APPLIED LIFE SCIENCES↗

Reinforcement Learning Approach to Cybersecurity in Space (RELACSS)

Securing satellite groundstations against cyber-attacks is vital to national security missions. However, these cyber threats are constantly evolving. As vulnerabilities are discovered and patched, new vulnerabilities are discovered and exploited. In order to automate the process of discovering existing vulnerabilities and the means to exploit them, a reinforcement learning framework is presented in this report. We demonstrate that this framework can learn to successfully navigate an unknown network and detect nodes of interest despite the presence of a moving target defense. The agent then exfiltrates a file of interest from the node as quickly as possible. This framework also incorporates a defensive software agent that learns to impede the attacking agents progress. This setup allows for the agents to work against each other and improve their abilities. We anticipate that this capability will help uncover unforeseen vulnerabilities and the means to mitigate them. The modular nature of the framework enables users to swap out learning algorithms and modify the reward functions in order to adapt the learning tasks to various use cases and environments. Several algorithms, viz., tabular Q learning, deep Q networks, proximal policy optimization, advantage actor-critic, generative adversarial imitation learning, are explored for the agents and the results highlighted. The agent learns to solve the tasks in a light-weight abstract environment. Once the agent learns to perform sufficiently well, it can be deployed in a minimega virtual machine environment (or a real network) with wrappers that map abstract actions to software commands. The agent also uses a local representation of the actions called a ‘slot-mechanism’. This allows the agent to learn in a certain network and generalize it to different networks. The defensive agent learns to predict the actions taken by an offensive agent and uses that information to anticipate the threat. This information can then either be used to raise an alarm or to take actions to thwart the attack. We believe that with the appropriate reward design, a representative environment, and action set, this framework can be generalized to tackle other cybersecurity tasks. By sufficiently training these agents, we can anticipate vulnerabilities leading to robust future designs. We can also deploy automated defensive agents that can help secure satellite groundstation and their vital national security missions.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Network Anomaly Detection Using Federated Learning

The internet is turning out to be an integral part of every-one's lives as more and more devices are being connected to serve societal needs. Our work is motivated by two ma-jor observations. Firstly, one drawback of connecting to the network is the threat of network attacks that can compromise users' private information, leading to data loss and adversely affecting productivity. There are several traditional security mechanisms to defend against these attacks, such as firewalls, virtual private networks (VPNs), demilitarized zones (DMZs), and vulnerability scanners. One way to prevent these attacks is early detection and prevention. However, these kinds of architecture do not scale very well because of their centralized nature. Secondly, we observe from heuristics and data set distributions that the majority of the requests made to a server are innocuous. Therefore, almost all server request data sets are highly imbalanced, weighted highly towards the harmless requests.

Marfo, William↗

A hybrid surrogate modeling framework for the Digital Twin of a Fluoride-salt-cooled High-temperature Reactor (FHR)

While nuclear energy is a non-greenhouse-gas emitting energy source, expensive operational costs due to the high-level of safety requirements decreases their competitiveness in the sustainable energy market. Advanced reactor concepts paired with Digital Twins aim to increase the commercialization gains of nuclear energy by reducing operational costs, increasing reactor reliability and enhancing power generation. To support Digital Twin tasks such as real-time autonomous control, proactive maintenance monitoring or optimizing power demand operations, a fast and accurate virtual representation of the Nuclear Power Plant (NPP) is required. The computational cost of high-fidelity, physics-based models are unsuitable for real-time analysis or scalability. Here, in this work, a hybrid surrogate modeling framework is developed fora Fluoride-salt-cooled High-temperature Reactor (FHR) that leverages physics-inspired models for key reactor components and uses data-driven methods for rapid system state space prediction. The Xenon reactivity feedback model is integrated to inform the surrogate model about the reactor core and the homologous pump theory model is the basis for representing pump degradation. Using a detailed, two dimensional thermal hydraulics model to generate data on the FHR, we train a network of Vectorized Autoregressive Moving-Average with eXogenous input (VARMAX) models to predict the remaining state values. The result is a surrogate model that provides a detailed reactor state representation of 41 system states and a pump degradation analysis. The framework is applied to Load Follows profiles, yielding high accuracy and a speedup that is more than 4000x faster compared to the higher- fidelity thermal hydraulics model, enabling real-time operational intelligence and applications in long horizon predictions. While the surrogate model framework is demonstrated for the particular case of FHR, the hybrid physical/data-driven modeling approach including the network of surrogates and the underlying modularity has the potential to be applied to other physical asset systems.

Digital Twins↗

Student Support for the “Frontiers in Attosecond & Ultrafast X-ray Science” School

The new millennium witnessed two revolutionary breakthroughs in ultrafast x-ray science: table-top XUV sources based on high harmonic generation in gases ushered in the attosecond era while facility-based x-ray free-electron lasers opened the path for intense, femtosecond hard x-rays. This award requested scholarship funds for a 2019 and 2023 School whose prime objective was the training of young scientists in these emerging complementary areas both relevant to DOE BES mission. The school entitled the “Frontiers of Attosecond and Ultrafast X-ray Science (FAXS)” (http://www.erice-attosecond.it/) was the second and fourth in a series, which began in 2017. The two Schools were held during March 10-16, 2019 and March 26-31, 2023 at the Ettore Majorana Foundation and Centre for Scientific Culture (http://www.ccsem.infn.it/) in Erice, Sicily. The DOE funds supported the registration fee for young scientists (graduate students and postdocs) from US institutions. The registration fees included School participation, lodging and meals over the duration of the School. Note, the third addition of the School was held in 2022 as a virtual event due to the pandemic, DOE funds were not necessary for this event. The FAXS School is a course of the 62th and 63rd International School of Quantum Electronics under the directorship of Prof. Diederik Wiersma (University of Florence). The Directors for the FAXS School are Louis DiMauro (The Ohio State University, USA) and Mauro Nisoli (Politecnico di Milano, Italy). The FAXS School program consisted of approximately 10 lectures by leading experts in attosecond and x-ray science (see attached list). Most lecturers delivered a series of three 1-hour lectures. The lecturers were required to spend the full 5 days at the school so to promote interaction with the students. The Erice Majorana Center venue accommodated ~75 young scientists. The registration fee covered the cost of participating in the school, lodging and meals for the entire duration of the school. The schedule consisted of lectures every morning and afternoon except for one afternoon that was reserved for an archaeological excursion. Every evening had a student/postdoc poster session and social gatherings at the Majorana Center to encourage further interaction of all participants and lecturers. The two FAXS Schools attracted an international group of young scientists. The DOE funds supported the registration fee for 11 students/postdocs from US institutions (5 supported in 2019 and 6 supported in 2023). The management of the DOE fellowships were administered through the Research Foundation of The Ohio State University. FAXS scholarships for European students/postdocs were provided by European funding sources administered by co-Director, Prof. Nisoli. All students/postdocs were encouraged to present a poster. Travel expenses to the FAXS School were the responsibility of the student/postdoc home institution.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Emerging materials intelligence ecosystems propelled by machine learning

We report that the age of cognitive computing and artificial intelligence (AI) is just dawning. Inspired by its successes and promises, several AI ecosystems are blossoming, many of them within the domain of materials science and engineering. These materials intelligence ecosystems are being shaped by several independent developments. Machine learning (ML) algorithms and extant materials data are utilized to create surrogate models of materials properties and performance predictions. Materials data repositories, which fuel such surrogate model development, are mushrooming. Automated data and knowledge capture from the literature (to populate data repositories) using natural language processing approaches is being explored. The design of materials that meet target property requirements and of synthesis steps to create target materials appear to be within reach, either by closed-loop active-learning strategies or by inverting the prediction pipeline using advanced generative algorithms. AI and ML concepts are also transforming the computational and physical laboratory infrastructural landscapes used to create materials data in the first place. Surrogate models that can outstrip physics-based simulations (on which they are trained) by several orders of magnitude in speed while preserving accuracy are being actively developed. Automation, autonomy and guided high-throughput techniques are imparting enormous efficiencies and eliminating redundancies in materials synthesis and characterization. The integration of the various parts of the burgeoning ML landscape may lead to materials-savvy digital assistants and to a human-machine partnership that could enable dramatic efficiencies, accelerated discoveries and increased productivity. Here, we review these emergent materials intelligence ecosystems and discuss the imminent challenges and opportunities. The materials research landscape is being transformed by the infusion of approaches based on machine learning. This Review discusses the emerging materials intelligence ecosystems and the potential of human-machine partnerships for fast and efficient virtual materials screening, development and discovery.

36 MATERIALS SCIENCE↗

Advancing Cross-Disciplinary Understanding of Land-Atmosphere Interactions

The evolution of disciplinary silos and increasingly narrow disciplinary boundaries have together resulted in one-sided approaches to the study of land-atmosphere interactions—a field that requires a bi-directional approach to understand the complex feedbacks and interactions that occur. The integration of surface flux and atmospheric boundary layer measurements is therefore essential to advancing our understanding. The Land-Atmosphere 2021 workshop (held virtually, June 10-11, 2021) involved almost 300 participants from around the world and promoted cross-discipline collaboration by way of talks from invited speakers, moderated discussions, breakout sessions, and a virtual poster session. The workshop focused on five main theme areas: “big picture” overview, instrumentation and remote sensing, modeling, water, and aerosols and clouds. In talks and breakout groups, there were frequent calls for more AmeriFlux sites to be instrumented for boundary layer height measurements, and for the development of some “super sites” where profiling instruments would be deployed. There was further agreement on the need for the standardization of various datasets. There was also a consensus that funding agencies need to be willing to support the sorts of large projects (including associated instrumentation) which can drive interdisciplinary work. Early-career scientists, in particular, expressed enthusiasm for working across disciplinary boundaries but noted that there need to be more financial support and training opportunities so they would be better prepared for interdisciplinary work. Investment in these career development opportunities would enable today's cohort of early-career scientists to advance the frontiers of interdisciplinary work over the next couple of decades.

54 ENVIRONMENTAL SCIENCES↗

Virtual refrigerant charge sensing algorithm for residential CO₂ heat pumps

Natural refrigerants are increasingly adopted in next-generation heat pump systems, among which CO₂ heat pumps have attracted significant attention. However, due to their high operating pressures, the leakage risk is higher, resulting in undercharge conditions and degraded heat pump performance. Thus, developing an accurate refrigerant charge level detection technique is necessary to guarantee safe and efficient operation. Although virtual refrigerant charge (VRC) level calculation algorithms for CO₂ heat pumps exist, they typically rely on empirically selected features without a systematic selection framework, leading to multicollinearity and potential overfitting, which limit their prediction accuracy and generalizability. To address these issues, this study proposes a VRC algorithm framework with a systematic feature selection method that identifies physically meaningful and statistically significant features, and is applied using a residential CO₂ heat pump as a case study. The method is extended from previous work on conventional refrigerants to account for charge behavior in CO₂ gas coolers. The selected features include gas cooler outlet density, evaporator pressure, and superheat temperature. The results demonstrate that the proposed feature selection method significantly improves prediction accuracy compared to existing VRC approaches. A relatively small training dataset (∼30 samples) is sufficient for feature identification and model development. The developed algorithm achieves less than 3% prediction error under both undercharge and overcharge conditions, representing reductions of 46.7% and 35.3% compared to two recent reference VRC algorithms for transcritical CO₂ heat pumps reported in the literature. The proposed algorithm and feature selection method enhance leakage detection capability, facilitate the deployment of CO₂ heat pump systems, and contribute to reduced energy waste and maintenance costs.

Guo, Fangzhou [Lawrence Berkeley National Laborato↗

Robust Design Under Uncertainty in Quantum Error Mitigation

Error mitigation techniques are crucial to achieving near-term quantum advantage. Classical postprocessing of quantum computation outcomes is a popular approach for error mitigation, which includes methods, such as zero noise extrapolation, virtual distillation, and learning-based error mitigation. However, these techniques have limitations due to the propagation of uncertainty resulting from the finite shot number of a quantum measurement. In this work, we introduce general and unbiased methods for quantifying the uncertainty and error of error-mitigated observables based on the strategic sampling of error mitigation outcomes. We then extend our approach to demonstrate the optimization of performance and robustness of error mitigation under uncertainty. To illustrate our methods, we apply them to zero noise extrapolation and Clifford date regression in the ground state of the XY model simulated using depolarizing and International Business Machines Corporation (IBM) Toronto noise models, respectively. In particular, we optimize the choice of noise levels and the allocation of shots for zero noise extrapolation and the distribution of the training circuits for Clifford data regression. While our methods are readily applicable to any postprocessing-based error mitigation approach, in practice they must not be prohibitively expensive—even though they perform optimizations of the error mitigation hyperparameters requiring sampling of a statistical distribution of error mitigation outcomes. By leveraging surrogate-based optimization, we show that our methods can efficiently perform optimal design for a zero noise extrapolation implementation. We then further demonstrate the transferability of learned zero noise extrapolation hyperparameters to other similar circuits.

97 MATHEMATICS AND COMPUTING↗

Deep Learning Approaches to Surrogates for Solving the Diffusion Equation for Mechanistic Real-World Simulations

In many mechanistic medical, biological, physical, and engineered spatiotemporal dynamic models the numerical solution of partial differential equations (PDEs), especially for diffusion, fluid flow and mechanical relaxation, can make simulations impractically slow. Biological models of tissues and organs often require the simultaneous calculation of the spatial variation of concentration of dozens of diffusing chemical species. One clinical example where rapid calculation of a diffusing field is of use is the estimation of oxygen gradients in the retina, based on imaging of the retinal vasculature, to guide surgical interventions in diabetic retinopathy. Furthermore, the ability to predict blood perfusion and oxygenation may one day guide clinical interventions in diverse settings, i.e., from stent placement in treating heart disease to BOLD fMRI interpretation in evaluating cognitive function (Xie et al., 2019; Lee et al., 2020). Since the quasi-steady-state solutions required for fast-diffusing chemical species like oxygen are particularly computationally costly, we consider the use of a neural network to provide an approximate solution to the steady-state diffusion equation. Machine learning surrogates, neural networks trained to provide approximate solutions to such complicated numerical problems, can often provide speed-ups of several orders of magnitude compared to direct calculation. Surrogates of PDEs could enable use of larger and more detailed models than are possible with direct calculation and can make including such simulations in real-time or near-real time workflows practical. Creating a surrogate requires running the direct calculation tens of thousands of times to generate training data and then training the neural network, both of which are computationally expensive. Often the practical applications of such models require thousands to millions of replica simulations, for example for parameter identification and uncertainty quantification, each of which gains speed from surrogate use and rapidly recovers the up-front costs of surrogate generation. We use a Convolutional Neural Network to approximate the stationary solution to the diffusion equation in the case of two equal-diameter, circular, constant-value sources located at random positions in a two-dimensional square domain with absorbing boundary conditions. Such a configuration caricatures the chemical concentration field of a fast-diffusing species like oxygen in a tissue with two parallel blood vessels in a cross section perpendicular to the two blood vessels. To improve convergence during training, we apply a training approach that uses roll-back to reject stochastic changes to the network that increase the loss function. The trained neural network approximation is about 1000 times faster than the direct calculation for individual replicas. Because different applications will have different criteria for acceptable approximation accuracy, we discuss a variety of loss functions and accuracy estimators that can help select the best network for a particular application. We briefly discuss some of the issues we encountered with overfitting, mismapping of the field values and the geometrical conditions that lead to large absolute and relative errors in the approximate solution.

60 APPLIED LIFE SCIENCES↗

Leveraging generative artificial intelligence to bridge domain gaps in wind turbine research

A central challenge in wind turbine health monitoring is the scarcity of real-world data due to limited instrumentation, leading researchers to rely on simulation models that often suffer from reduced fidelity. However, even within simulation environments, discrepancies arise because of modeling assumptions, and configuration fidelities, creating domain gaps that limit the transferability of learned representations. Here, to investigate domain translation under controlled conditions, this project explores the use of generative artificial intelligence, specifically cycle-consistent generative adversarial networks (CGANs), to bridge the gap between OpenFAST simulation models representing 1.5 MW and 5 MW wind turbines. A physics-informed CGAN architecture is introduced, where a simplified turbine tower dynamics model is incorporated into the training loss to ensure physically consistent outputs. Quantitative results showed moderate to high agreement in frequency-domain features. Incorporating the physics-informed loss function improved the R 2 values by 30%, reduced the RMSE from 1.39 to 1.1 m/s 2 , and reduced training time by 82%. Furthermore, under increased turbulence intensity (IEC Category A), the RMSE remained stable at approximately 1.1 m/s 2 . While the present study is entirely simulation-based, it establishes a pipeline for evaluating physics-informed generative domain translation, which may serve as a foundation for future simulation-to-reality validation studies.

17 WIND ENERGY↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

A Hybrid Energy System Workflow for Energy Portfolio Optimization

This manuscript develops a workflow, driven by data analytics algorithms, to support the optimization of the economic performance of an Integrated Energy System. The goal is to determine the optimum mix of capacities from a set of different energy producers (e.g., nuclear, gas, wind and solar). A stochastic-based optimizer is employed, based on Gaussian Process Modeling, which requires numerous samples for its training. Each sample represents a time series describing the demand, load, or other operational and economic profiles for various types of energy producers. These samples are synthetically generated using a reduced order modeling algorithm that reads a limited set of historical data, such as demand and load data from past years. Numerous data analysis methods are employed to construct the reduced order models, including, for example, the Auto Regressive Moving Average, Fourier series decomposition, and the peak detection algorithm. All these algorithms are designed to detrend the data and extract features that can be employed to generate synthetic time histories that preserve the statistical properties of the original limited historical data. The optimization cost function is based on an economic model that assesses the effective cost of energy based on two figures of merit: the specific cash flow stream for each energy producer and the total Net Present Value. An initial guess for the optimal capacities is obtained using the screening curve method. The results of the Gaussian Process model-based optimization are assessed using an exhaustive Monte Carlo search, with the results indicating reasonable optimization results. The workflow has been implemented inside the Idaho National Laboratory’s Risk Analysis and Virtual Environment (RAVEN) framework. The main contribution of this study addresses several challenges in the current optimization methods of the energy portfolios in IES: First, the feasibility of generating the synthetic time series of the periodic peak data; Second, the computational burden of the conventional stochastic optimization of the energy portfolio, associated with the need for repeated executions of system models; Third, the inadequacies of previous studies in terms of the comparisons of the impact of the economic parameters. The proposed workflow can provide a scientifically defendable strategy to support decision-making in the electricity market and to help energy distributors develop a better understanding of the performance of integrated energy systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

A comparison of histopathology imaging comprehension algorithms based on multiple instance learning

Whole slide imaging (WSI), also called digital virtual microscopy, is a new imaging modality. It allows for the application of AI and machine learning methods to cancer pathology to help establish a means for the automatic diagnosis of cancer cases. However, designing machine-learning models for WSI is computationally challenging due to its required ultra-high resolution. The current state-of-the-art models use multiple instance learning (MIL). MIL is a weakly-supervised learning method in which the model uses an array of inferences from many smaller instances to make a final classification about the entire set. In the context of WSI, researchers divide the ultra-high-resolution image into many patches. The model then classifies the slide based on an array of inferences from the patches. Among several ways of making the final classification, attention-based mechanisms have resulted in superb accuracy scores. The Transformer, one attention-based algorithm, has reported substantial improvements for WSI comprehension tasks. In this project, we studied and compared several WSI comprehension algorithms. We used the following three datasets: CAMELYON16+17, TCGALung, and TCGA-Kidney. We found that attention-based MIL algorithms performed better than standard MIL algorithms for classifying WSI images, achieving a higher mean accuracy and AUC. However, none of the attention-based algorithms performed significantly better than the others, reporting accuracy scores that varied widely. Presumably, it is due to the limited availability of training samples in the data corpus. Since it is not easy to increase the samples from human subjects, some machine learning techniques like transfer learning could help mitigate this issue.

Saunders, Adam↗