Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Steepest-descent algorithm for simulating plasma-wave caustics via metaplectic geometrical optics

The design and optimization of radiofrequency-wave systems for fusion applications is often performed using ray-tracing codes, which rely on the geometrical-optics (GO) approximation. However, GO fails at wave cutoffs and caustics. To accurately model the wave behavior in these regions, more advanced and computationally expensive “full-wave” simulations are typically used, but this is not strictly necessary. A new generalized formulation called metaplectic geometrical optics (MGO) has been proposed that reinstates GO near caustics. The MGO framework yields an integral representation of the wavefield that must be evaluated numerically in general. We present an algorithm for computing these integrals using Gauss-Freud quadrature along the steepest-descent contours. Benchmarking is performed on the standard Airy problem, for which the exact solution is known analytically. Furthermore, the numerical MGO solution provided by the new algorithm agrees remarkably well with the exact solution and significantly improves on previously derived analytical approximations of the MGO integral.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Online transfer learning strategy for enhancing the scalability and deployment of deep reinforcement learning control in smart buildings

In recent years, advanced control strategies based on Deep Reinforcement Learning (DRL) proved to be effective in optimizing the management of integrated energy systems in buildings, reducing energy costs and improving indoor comfort conditions when compared to traditional reactive controllers. However, the scalability and implementation of DRL controllers are still limited since they require a considerable amount of time before converging to a near-optimal solution. This issue is currently addressed in literature through the offline pre-training of the DRL agent. However this solution results in two main critical issues: (1) the need to develop a building surrogate model to perform the training task, and (2) the need to perform a fine-tuning process over several training episodes to obtain a near-optimal control policy. In this context, this paper introduces an Online Transfer Learning (OTL) strategy that exploits two knowledge-sharing techniques, weight-initialization and imitation learning, to transfer a DRL control policy from a source office building to various target buildings in a simulation environment coupling EnergyPlus and Python. A DRL controller based on discrete Soft Actor–Critic (SAC) is trained on the source building to manage the operation of a cooling system consisting of a chiller and a thermal storage. Several target buildings are defined to benchmark the performance of the OTL strategy with that of a Rule-Based Controller (RBC) and two DRL-based control strategies, deployed in offline and online fashion. The strategy adopted for OTL emulates the real world implementation with a simulation process by implementing the transferred DRL agent for a single episode in the target buildings. Target buildings have the same geometrical features and are served by the same energy system as the source building, but differ in terms of weather conditions, electricity price schedules, occupancy patterns, and building envelope efficiency levels. The results show that the OTL strategy can reduce the cumulated sum of temperature violations on average by 50% and 80% respectively when compared to RBC and online DRL while enhancing the energy system operation with electricity cost savings ranging between 20% and 40%. Furthermore, the OTL agent performs slightly worse than the offline DRL controller but it does not require any modeling effort and can be implemented directly on target buildings emulating a real-world implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Development of an Exploration-Class Cascade Distillation Subsystem: Performance Testing of the Generation 1.0 Prototype

The ability to recover and purify water is crucial for realizing long-term human space missions. The National Aeronautics and Space Admininstration and Honeywell co-developed a five-stage vacuum rotary distillation water recovery system referred to as the Cascade Distillation Subsystem (CDS). Over the past three years, NASA's Advanced Exploration Systems (AES) Water Recovery Project (WRP) has been working toward the development of a flight-forward CDS design. In 2012 the original CDS prototype underwent a series of incremental upgrades and tests intened to both demonstrate the feasibility of a on-orbit demonstration of the system and to collect operational and performance data to be used to inform a second generation design. The latest testing of the CDS Generation 1.0 prototype was conducted May 29 through July 2, 2014. Initial system performance was benchmarked by processing deionized water and sodium chloride. Following, the system was challenged with analogue urine waste stream solutions stabilized with an Oxone-based and the two International Space Station baseline and alternative pretreatment solutions. During testing, the system processed more than 160 kilograms of wastewater with targeted water recoveries between 75 and 85% depending on the specific waste stream tested. For all wastewater streams, contaminant removals from wastewater feed to product water distillate, were estimated at greater than 99%. The average specific energy of the system was less than 120 Watt-hours/kilogram. The following paper provides detailed information and data on the performance of the CDS as challenged per the WRP test objectives.

Callahan, Michael R.↗

Architecture and evolution of Goddard Space Flight Center Distributed Active Archive Center

The Goddard Space Flight Center (GSFC) Distributed Active Archive Center (DAAC) has been developed to enhance Earth Science research by improved access to remote sensor earth science data. Building and operating an archive, even one of a moderate size (a few Terabytes), is a challenging task. One of the critical components of this system is Unitree, the Hierarchical File Storage Management System. Unitree, selected two years ago as the best available solution, requires constant system administrative support. It is not always suitable as an archive and distribution data center, and has moderate performance. The Data Archive and Distribution System (DADS) software developed to monitor, manage, and automate the ingestion, archive, and distribution functions turned out to be more challenging than anticipated. Having the software and tools is not sufficient to succeed. Human interaction within the system must be fully understood to improve efficiency to improve efficiency and ensure that the right tools are developed. One of the lessons learned is that the operability, reliability, and performance aspects should be thoroughly addressed in the initial design. However, the GSFC DAAC has demonstrated that it is capable of distributing over 40 GB per day. A backup system to archive a second copy of all data ingested is under development. This backup system will be used not only for disaster recovery but will also replace the main archive when it is unavailable during maintenance or hardware replacement. The GSFC DAAC has put a strong emphasis on quality at all level of its organization. A Quality team has also been formed to identify quality issues and to propose improvements. The DAAC has conducted numerous tests to benchmark the performance of the system. These tests proved to be extremely useful in identifying bottlenecks and deficiencies in operational procedures.

Bedet, Jean-Jacques↗

Degradation in Photovoltaic Encapsulant Transmittance: Results of the Second PVQAT TG5 Artificial Weathering Study

The optical degradation of encapsulants from ultraviolet (UV) radiation has historically resulted in a significant loss in performance throughout the life of a photovoltaic (PV) module. IEC test methods have recently been developed to screen for PV encapsulants prone to loss in optical performance. The present study was performed to benchmark polymeric packaging materials relative to IEC 62788-1-4 (covering the measurement of optical transmittance) and IEC 62788-1-7 (on the durability of transmittance), provide feedback toward improvement of the methods, and develop insight regarding optical degradation. Contemporary materials were examined, including: poly(ethylene-co-vinyl acetate) (EVA), thermoplastic polyolefin (TPO), polyolefin elastomer (POE), and polyvinyl butyral (PVB) encapsulants; a poly(ethene-co-tetrafluoroethene)/poly(ethylene terephthalate) (ETFE/PET) transparent backsheet; and a polystyrene (PS) working reference material. The use of silica-, specialty-, and rolled-glass was also compared in laminated coupons. Specimen size was separately examined from 2.5 cm to 12.5 cm. Weathering was performed with a xenon source, using IEC TS 62788-7-2 methods A2, A3, A4, and A5 (chamber temperature of 55 Degrees C, 65 Degrees C, 75 Degrees C, or 85 Degrees C), respectively. Characterizations were made using a UV-VIS-NIR spectrophotometer (transmittance and reflectance, with and without an integrating sphere), a UV-VIS fluorescence spectrophotometer, a camera, and an optical microscope. Performance was analyzed, including solar weighted transmittance, yellowness index, UV cut-off wavelength, and haze (scattering). Separate Arrhenius analyses were performed to assess retention of transmittance and changes in yellowness index. The activation energy for both characteristics was found to range from 15 kJ?mol-1-80 kJ/mol-1, with an average of 48 kJ/mol-1, similar to the average of 45 kJ/mol-1 identified in the previous international PV Quality Assurance Task Force (PVQAT) Task Group 5 (TG5) study of more traditional encapsulants. The separate degradation modes of discoloration and scattering were distinguished in the encapsulants using a comprehensive spectral characterization. Based on these results, the IEC 62788-1-7 pass/fail criteria of 5% change in transmittance was confirmed to identify a known bad encapsulant.

durability↗

Impact of battery cell imbalance on electric vehicle range

Due to manufacturing variation, battery cells often possess heterogeneous characteristics, leading to battery state-of-charge variation in real-time. Since the lowest cell state-of-charge determines the useful life of battery pack, such variation can negatively impact the battery performance and electric vehicles range. Existing research has been focused on control design to mitigate cell imbalance. However, it is yet unclear how much impacts the cell imbalance can have on electric vehicle range. This paper closes this knowledge gap by using a simulation environment consisting of real-world driving speed data, vehicle longitudinal control, propulsion and vehicle dynamics, and cell level battery modeling. In particular, each battery cell is modeled as an equivalent circuit model, and variations among cell parameters are introduced to assess their impact on electric vehicles range and to identify the most influential parameter variations. Simulation results and analysis can be used to assist balancing control design and to benchmark control performance.

25 ENERGY STORAGE↗

Risk-Aware Framework Development for Disruption Prediction: Alcator C-Mod and DIII-D Survival Analysis

Abstract Survival regression models can achieve longer warning times at similar receiver operating characteristic performance than previously investigated models. Survival regression models are also shown to predict the time until a disruption will occur with lower error than other predictors. Time-to-event predictions from time-series data can be obtained with a survival analysis statistical framework, and there have been many tools developed for this task which we aim to apply to disruption prediction. Using the open-source Auton-Survival package we have implemented disruption predictors with the survival regression models Cox Proportional Hazards, Deep Cox Proportional Hazards, and Deep Survival Machines. To compare with previous work, we also include predictors using a Random Forest binary classifier, and a conditional Kaplan-Meier formalism. We benchmarked the performance of these five predictors using experimental data from the Alcator C-Mod and DIII-D tokamaks by simulating alarms on each individual shot. We find that developing machine-relevant metrics to evaluate models is an important area for future work. While this study finds cases where disruptive conditions are not predicted, there are instances where the desired outcome is produced. Giving the plasma control system the expected time-to-disruption will allow it to determine the optimal actuator response in real time to minimize risk of damage to the device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Discrete fracture network model benchmarks developed and applied in a DECOVALEX-2023 repository performance assessment study

This study presents newly developed benchmarks for modeling flow and transport within discrete fracture networks (DFNs) and useful methods for analyzing the results. The new benchmarks are designed to test modeling approaches for use in probabilistic performance assessment models of deep geologic repositories in fractured rock. The benchmarks simulate flow and transport through a 1 km 3 block of fractured rock. The first simulates migration of a short pulse of tracer through a simple network of four intersecting fractures. The second adds 1089 stochastically generated fractures. The third changes the pulse to a continuous point source. Evaluation of model performance relies on moment analysis and comparison of the results of different models. The expected nondimensional first moment of the conservative tracer for each benchmark is 1. The benchmarks were simulated by teams from Canada, Czechia, Germany, Korea, Sweden, Taiwan, and the United States as part of a DECOVALEX-2023 study (decovalex.org). The teams used various approaches, including explicit DFN modeling, DFN upscaling to an equivalent continuous porous medium (ECPM), and a combination of both methods. Transport mechanisms are modeled using either the advection-dispersion equation or particle tracking. Results demonstrate strong agreement among the models in breakthrough behavior up to the 75th percentile. Significant deviations in first moments and well-clustered outputs led to the identification of inaccuracies in several models. Such findings exemplify the benefit of exercising these benchmarks and using the presented methods to test DFN flow and transport models.

Benchmark↗

Visualization of shocked material instabilities using a fast-framing camera and XFEL four-pulse train

Many questions regarding dynamic materials could be answered by using time-resolved ultra-fast imaging techniques to characterize the physical and chemical behavior of materials in extreme conditions and their evolution on the nanosecond scale. In this work, we perform multi-frame phase-contrast imaging (PCI) of micro-voids in low density polymers under laser-driven shock compression. At the Matter in Extreme Conditions (MEC) Instrument at the Linac Coherent Light Source (LCLS), we used a train of four x-ray free electron laser (XFEL) pulses to probe the evolution of the samples. To visualize the void and shock wave interaction, here, we deployed the Icarus V2 detector to record up to four XFEL pulses, separated by 1-3 nanoseconds. In this work, we image elastic waves interacting with the micro-voids at a pressure of several GPa. Monitoring how the material’s heterogeneities, like micro-voids, dictate its response to a compressive wave is important for benchmarking the performances of inertial confinement fusion energy materials. For the first time in a single sample, we have combined an ultrafast x-ray framing camera and four XFEL pulse train to create an ultrafast movie of micro-void evolution under laser-driven shock compression. Eventually, we hope this technique will resolve the material density as it evolves dynamically under laser shock compression.

fusion↗

LBNL Fault Detection and Diagnostics Datasets

These datasets can be used to evaluate and benchmark the performance accuracy of Fault Detection and Diagnostics (FDD) algorithms or tools. It contains operational data from simulation, laboratory experiments, and field measurements from real buildings for seven HVAC systems/equipment (rooftop unit, single-duct air handler unit, dual-duct air handler unit, variable air volume box, fan coil unit, chiller plant, and boiler plant). Each dataset includes a .pdf file to document key information necessary to understand the content and scope, multiple csv files containing all the time-series data for faults at different severity levels and one fault-free case, and a ttl file to visualize the data according to BRICK schema. The dataset was created by LBNL, PNNL, NREL, ORNL and Drexel University.

AC↗

Calibrating the Classical Hardness of the Quantum Approximate Optimization Algorithm

The trading of fidelity for scale enables approximate classical simulators such as matrix product states (MPSs) to run quantum circuits beyond exact methods. A control parameter, the so-called bond dimension $\mathcal{χ}$ for MPSs, governs the allocated computational resources and the output fidelity. Here, we characterize the fidelity for the quantum approximate optimization algorithm by the expectation value of the cost function that it seeks to minimize and find that it follows a scaling law $\mathscr{F}$(ln $\mathcal{χ}$/N), where N is the number of qubits. With ln $\mathcal{χ}$ amounting to the entanglement that a MPS can encode, we show that the relevant variable for investigating the fidelity is the entanglement per qubit. Importantly, our results calibrate the classical computational power required to achieve the desired fidelity and benchmark the performance of quantum hardware in a realistic setup. For instance, we quantify the hardness of performing better classically than a noisy superconducting quantum processor by readily matching its output to the scaling function. Moreover, we relate the global fidelity to that of individual operations and establish its relationship with $\mathcal{χ}$ and N. We sharpen the requirements for noisy quantum computers to outperform classical techniques at running a quantum optimization algorithm in speed, size, and fidelity.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Observations on Cost Modeling and Performance Measurement of Long Term Archives

This paper describes a prototype suite of Excel-based tools that could be used for estimating lifecycle costs for newly planned or modified long-term archival facilities. These tools may also prove valuable for monitoring the long-term performance of such facilities once operational. Cost estimation is by analogy, using statistical curve-fitting techniques across a database of comparable data activities. The database currently includes 29 operational data centers ranging from small (2 FTEs) to large (66 FTEs), and is readily expandable to include additional activities specifically involving data preservation and added value. Each comparable data center is described in terms of its staffing, throughput workload, archival and distribution requirements, levels of user service, overall complexity, degree of automation, and other data, comprising 94 distinct descriptors in all. The descriptors were developed by normalizing heterogeneous data from the various centers and mapping them into an Excel framework consistent with the OAIS reference model. A user-friendly tool is provided for generating input to and updating the comparables database. This tool can also be used to benchmark the performance (in terms of cost versus throughput) of an operational data center, and to update the staffing, cost and workload data on a periodic basis. The comparables database could thus provide a history of staffing and throughput over time, as a means of performance monitoring and providing feedback for continuous improvement. Ancillary tools are also provided for performing "what-if' cost exercises for planning purposes, and for graphical display of data and results. We provide a high-level description of the tools; present our experiences and observations on gathering the information and maintaining the database; and discuss how this tool set might be applied to long term archives.

Fontaine, Kathy↗

TRANSLATE - a Monte Carlo simulation of electron transport in liquid argon

Here, the microphysics of electron and photon propagation in liquid argon is a key component of detector design and calibrations needed to construct and perform measurements within a wide range of particle physics experiments. As experiments grow in scale and complexity, and as the precision of their intended measurements increases, the development of tools to investigate important microphysics effects impacting such detectors becomes necessary. In this paper we present a new time-domain Monte Carlo simulation of electron transport in liquid argon. The simulation models the TRANSport in Liquid Argon of near-Thermal Electrons (TRANSLATE) with the aim of providing a multi-purpose software package for the study and optimization of detector environments, with a particular focus on ongoing and next generation liquid argon neutrino experiments utilizing the time projection chamber technology. TRANSLATE builds on previous work of Wojcik and Tachiya, amongst others, introducing additional processes, including ionization, thus modeling the full range of drift electron scattering interactions. The simulation is validated by benchmarking its performance with swarm parameters from data collected in experimental setups operating in gas and liquid.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

PandExo: A Community Tool for Transiting Exoplanet Science with JWST and HST

As we approach the James Webb Space Telescope (JWST) era, several studies have emerged that aim to (1) characterize how the instruments will perform and (2) determine what atmospheric spectral features could theoretically be detected using transmission and emission spectroscopy. To some degree, all these studies have relied on modeling of JWST's theoretical instrument noise. With under two years left until launch, it is imperative that the exoplanet community begins to digest and integrate these studies into their observing plans, as well as think about how to leverage the Hubble Space Telescope (HST) to optimize JWST observations. To encourage this and to allow all members of the community access to JWST & HST noise simulations, we present here an open-source Python package and online interface for creating observation simulations of all observatory-supported timeseries spectroscopy modes. This noise simulator, called PandExo, relies on some aspects of Space Telescope Science Institute's Exposure Time Calculator, Pandeia. We describe PandExo and the formalism for computing noise sources for JWST. Then we benchmark PandExoʼs performance against each instrument team's independently written noise simulator for JWST, and previous observations for HST. We find that PandExo is within 10% agreement for HST/WFC3 and for all JWST instruments.

Batalha, Natasha E.↗

The physics potential of a reactor neutrino experiment with Skipper CCDs: Measuring the weak mixing angle

We analyze in detail the physics potential of an experiment like the one recently proposed by the vIOLETA collaboration: a kilogram-scale Skipper CCD detector deployed 12 meters away from a commercial nuclear reactor core. This experiment would be able to detect coherent elastic neutrino nucleus scattering from reactor neutrinos, capitalizing on the exceptionally low ionization energy threshold of Skipper CCDs. To estimate the physics reach, we elect the measurement of the weak mixing angle as a case study. We choose a realistic benchmark experimental setup and perform variations on this benchmark to understand the role of quenching factor and its systematic uncertainties, background rate and spectral shape, total exposure, and reactor antineutrino flux uncertainty. We take full advantage of the reactor flux measurement of the Daya Bay collaboration to perform a data driven analysis which is, up to a certain extent, independent of the theoretical un- certainties on the reactor antineutrino flux. We show that, under reasonable assumptions, this experimental setup may provide a competitive measurement of the weak mixing angle at few MeV scale with neutrino-nucleus scattering.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Random projection using random quantum circuits

The random sampling task performed by Google's Sycamore processor gave us a glimpse of the “quantum supremacy era.” This has definitely shed some light on the power of random quantum circuits in this abstract task of sampling outputs from the (pseudo)random circuits. In this paper, we explore a practical near-term use of local random quantum circuits in dimensional reduction of large low-rank data sets. We make use of the well-studied dimensionality reduction technique called the random projection method. This method has been extensively used in various applications such as image processing, logistic regression, entropy computation of low-rank matrices, etc. We prove that the matrix representations of local random quantum circuits with sufficiently shorter depths [ ∼ O ( n ) ] serve as good candidates for random projection. We demonstrate numerically that their projection abilities are not far off from the computationally expensive classical principal components analysis on MNIST and CIFAR-100 image datasets. We also benchmark the performance of quantum random projection against the commonly used classical random projection in the tasks of dimensionality reduction of image data sets and computing von Neumann entropies of large low-rank density matrices. And finally, using variational quantum singular value decomposition, we demonstrate a near-term implementation of extracting the singular vectors with dominant singular values after quantum random projecting a large low-rank matrix to lower dimensions. All such numerical experiments unequivocally demonstrate the ability of local random circuits to randomize a large Hilbert space at sufficiently shorter depths with robust retention of properties of large data sets in reduced dimensions. Published by the American Physical Society 2024

Kumaran, Keerthi (ORCID:0009000949125721)↗

Objective Structured Clinical Evaluation (OSCE) of an Artificial Intelligence (AI) Clinical Decision Support System (CDSS) Tool

BACKGROUND Objective Structured Clinical Evaluations (OSCEs) have long been established as a robust methodology for summative assessment of clinical skills and decision-making during medical education. The recent integration of Artificial Intelligence (AI) into clinical decision-making processes has prompted the need for novel evaluation frameworks to assess the efficacy and reliability of AI clinical decision support system (CDSS) tools. This abstract outlines the process of quantitatively evaluating a novel CDSS (“Doc in a Box” Google 2024) trained on curated medical spaceflight data in the psychomotor domain as it interfaces with a human volunteer acting as the crew medical officer (CMO). PURPOSE The AI CDSS under review was developed as part of the Lunar Command and Control Interoperability (LuCCI) project, which is intended to address a gap in how Lunar Surface Systems (LSS) would interoperate across multiple programs, commercial partners, and international partners. The project objective is to define, prototype, integrate, and evaluate an interoperable lunar command, control, data, and software reference architecture to enable autonomy and informatics capability through common standards across LSS. A multi-modal AI-based CDSS compatible with Federated LSS will assist clinicians in diagnosing and managing complex medical conditions by providing evidence-based recommendations through predictive analytics. Given the critical role of decision-support as NASA continues to evolve its Earth-independent medical operations (EIMO), it is imperative to ensure that such AI tools perform reliably and align with clinical standards during progressive lunar and Martian exploration class missions. METHODS The OSCE framework, traditionally used for evaluating human clinicians, was adapted to assess the AI tool's decision-making capabilities in simulated clinical scenarios. In this adapted OSCE, the AI CDSS was tested across a series of structured clinical scenarios designed to mimic real-life spaceflight patient cases. These scenarios included a range of conditions and complexities, allowing for comprehensive assessment of the tool's performance. Key evaluation metrics included accuracy of diagnosis, timeliness of decision-making, and appropriate recommendations for therapies. The OSCE was scored by human physician evaluators who assessed the AI's recommendations in comparison with expert clinicians' medical decision making to ensure alignment with best practices and the standard of care. RESULTS Preliminary results indicate that the AI CDSS demonstrated high accuracy in diagnostic recommendations and decision support across various scenarios. However, certain limitations were noted, such as occasional discrepancies in handling complex or nuanced cases that required a more contextual understanding. Additionally, the tool scored higher on the diagnostic portion of the rubric, with lower scores in the therapeutic recommendations. These findings highlight the importance of continuous refinement and validation of AI tools through rigorous evaluation frameworks like the OSCE. The adaptation of OSCEs for AI tools presents several advantages, including a structured and reproducible approach to evaluation, the ability to test AI systems in diverse clinical scenarios, and the opportunity to benchmark AI performance against established clinical standards to permit charting of future progress as aerospace medicine evolves as a discipline. Remaining challenges include ensuring that these evaluations capture the full spectrum of clinical decision-making scenarios that will be confronted by CMOs during missions and adequately reflecting real-world variability of the austere spaceflight environment. CONCLUSION Employing OSCEs to evaluate AI clinical decision support tools offers a promising approach to validating their clinical utility and efficacy. This methodology not only provides insights into the tool's performance but also fosters ongoing improvement and alignment with standard of care practices. Future research should focus on refining these evaluation processes and addressing limitations to enhance the integration of AI tools in clinical spaceflight settings. REFERENCES Scott S, Hearns V, Barker MA. Testing Clinical Skills: A Look at the OSCE and USMLE Clinical Skills Exams. S D Med. 2019 Oct;72(10):451-453. Majumder MAA, Kumar A, Krishnamurthy K, Ojeh N, Adams OP, Sa B. An evaluative study of objective structured clinical examination (OSCE): students and examiners perspectives. Adv Med Educ Pract. 2019 Jun 5;10:387-397. Karam VY, Park YS, Tekian A, Youssef N. Evaluating the validity evidence of an OSCE: results from a new medical school. BMC Med Educ. 2018 Dec 20;18(1):313.

Ariana M Nelson↗

The Importance of Being Adaptable: An Exploration of the Power and Limitations of Domain Adaptation for Simulation-Based Inference with Galaxy Clusters

The application of deep machine learning methods in astronomy has exploded in the last decade, with new models showing remarkably improved performance on benchmark tasks. Not nearly enough attention is given to understanding the models' robustness, especially when the test data are systematically different from the training data, or "out of domain." Domain shift poses a significant challenge for simulation-based inference, where models are trained on simulated data but applied to real observational data. In this paper, we explore domain shift and test domain adaptation methods for a specific scientific case: simulation-based inference for estimating galaxy cluster masses from X-ray profiles. We build datasets to mimic simulation-based inference: a training set from the Magneticum simulation, a scatter-augmented training set to capture uncertainties in scaling relations, and a test set derived from the IllustrisTNG simulation. We demonstrate that the Test Set is out of domain in subtle ways that would be difficult to detect without careful analysis. We apply three deep learning methods: a standard neural network (NN), a neural network trained on the scatter-augmented input catalogs, and a Deep Reconstruction-Regression Network (DRRN), a semi-supervised deep model engineered to address domain shift. Although the NN improves results by 17% in the Training Data, it performs 40% worse on the out-of-domain Test Set. Surprisingly, the Scatter-Augmented Neural Network (SANN) performs similarly. While the DRRN is successful in mapping the training and Test Data onto the same latent space, it consistently underperforms compared to a straightforward Yx scaling relation. These results serve as a warning that simulation-based inference must be handled with extreme care, as subtle differences between training simulations and observational data can lead to unforeseen biases creeping into the results.

Ntampaka, Michelle [Baltimore, Space Telescope Sci↗