Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “EVALUATION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Commercial Smallsat Data Acquisition Program On-ramp #2 Airbus U.S. Synthetic Aperture Radar (SAR) Evaluation Report

In 2017, NASA’s Earth Science Division (ESD) launched the Private-Sector Small Constellation Satellite Data Product Pilot, now referred to as the Commercial Smallsat Data Acquisition (CSDA) program. The objective of CSDA is to identify, evaluate, and acquire commercial remote sensing data that support NASA’s Earth science research and application activities. The Pilot successfully concluded in early 2020, when CSDA transitioned into a sustained program with on-ramping opportunities for new vendors as the industry emerges with new candidates and capabilities. In October 2019, a Request for Information (RFI) seeking capability statements from parties interested in providing data from spaceborne platforms was released for the CSDA on-ramp #2 evaluations. To be responsive to the RFI, the commercial satellite constellations had to consist of three or more operating spacecraft actively collecting data in a non-geostationary orbit with full latitudinal coverage and be U.S. companies. Two vendors responded to the RFI and were evaluated by a committee composed of NASA ESD leadership, program managers, and scientists. Both vendors satisfied the RFI requirements and were asked to respond to a Request for Proposal (RFP). After review of the proposals, NASA entered into a Blanket Purchase Agreement (BPA) with Airbus Defense and Space GEO, Inc. (Airbus) U.S. in September 2021 and with BlackSky Geospatial Solutions, Inc. (BlackSky) in November 2021. In this report, CSDA provides an evaluation of the usefulness of data provided by the Airbus U.S. Synthetic Aperture Radar (SAR) satellite constellation, consisting of TerraSAR-X (launched in 2007), TanDEM-X (launched in 2010), and PAZ (launched in 2018), for advancing NASA’s Earth system science research and applications. The evaluation of the BlackSky commercial data will be provided in a separate report. To conduct the Airbus evaluation, NASA’s ESD augmented 13 existing research projects that could potentially benefit from, and had the expertise to evaluate, the commercial data being considered for longer-term purchase. Investigators from NASA’s Research and Analysis Program science focus areas and from NASA’s Applied Sciences Program elements participated in the evaluation. A summary of the research areas evaluated by the Principal Investigator (PI) teams is presented in Figure 3. CSDA also funded a dedicated activity to evaluate the satellite data quality (calibration and geolocation) independently by assessing the accuracy of data from Airbus. Evaluation activities were carried out by the selected PIs from December 7, 2022, to December 7, 2023. Delivery of datasets requested by the researchers began in January 2023. The vendors were evaluated on the accessibility of data, accuracy and completeness of metadata, and promptness and quality of user support services. Datasets purchased during the evaluation have been archived by NASA and will be made available to current and future government-funded researchers in accordance with the End User License Agreement (EULA). This synthesis report distills and integrates the findings of research reports commissioned by NASA for the Airbus evaluation. This report also includes recommendations that inform the way ahead for the program. The scientific results from the evaluations demonstrated that the commercial data from Airbus were able to advance NASA research and applications. However, the PIs encountered limitations that diminished the usefulness of the data due to the amount of effort that was required to access, preprocess, and analyze these data. One significant issue encountered was the limited spatial and temporal coverage of the data in the Airbus archive that could be used to conduct time series analyses or assessments over large spatial scales. Overall, however, the utility and the quality of the evaluated data outweighed the difficulties encountered, and NASA has concluded that the Airbus SAR data would complement NASA’s existing Earth observation capabilities and Airbus U.S. would qualify to participate in the sustained phase of the program.

Batuhan Osmanoglu↗

Objective Structured Clinical Evaluation (OSCE) of an Artificial Intelligence (AI) Clinical Decision Support System (CDSS) Tool

BACKGROUND Objective Structured Clinical Evaluations (OSCEs) have long been established as a robust methodology for summative assessment of clinical skills and decision-making during medical education. The recent integration of Artificial Intelligence (AI) into clinical decision-making processes has prompted the need for novel evaluation frameworks to assess the efficacy and reliability of AI clinical decision support system (CDSS) tools. This abstract outlines the process of quantitatively evaluating a novel CDSS (“Doc in a Box” Google 2024) trained on curated medical spaceflight data in the psychomotor domain as it interfaces with a human volunteer acting as the crew medical officer (CMO). PURPOSE The AI CDSS under review was developed as part of the Lunar Command and Control Interoperability (LuCCI) project, which is intended to address a gap in how Lunar Surface Systems (LSS) would interoperate across multiple programs, commercial partners, and international partners. The project objective is to define, prototype, integrate, and evaluate an interoperable lunar command, control, data, and software reference architecture to enable autonomy and informatics capability through common standards across LSS. A multi-modal AI-based CDSS compatible with Federated LSS will assist clinicians in diagnosing and managing complex medical conditions by providing evidence-based recommendations through predictive analytics. Given the critical role of decision-support as NASA continues to evolve its Earth-independent medical operations (EIMO), it is imperative to ensure that such AI tools perform reliably and align with clinical standards during progressive lunar and Martian exploration class missions. METHODS The OSCE framework, traditionally used for evaluating human clinicians, was adapted to assess the AI tool's decision-making capabilities in simulated clinical scenarios. In this adapted OSCE, the AI CDSS was tested across a series of structured clinical scenarios designed to mimic real-life spaceflight patient cases. These scenarios included a range of conditions and complexities, allowing for comprehensive assessment of the tool's performance. Key evaluation metrics included accuracy of diagnosis, timeliness of decision-making, and appropriate recommendations for therapies. The OSCE was scored by human physician evaluators who assessed the AI's recommendations in comparison with expert clinicians' medical decision making to ensure alignment with best practices and the standard of care. RESULTS Preliminary results indicate that the AI CDSS demonstrated high accuracy in diagnostic recommendations and decision support across various scenarios. However, certain limitations were noted, such as occasional discrepancies in handling complex or nuanced cases that required a more contextual understanding. Additionally, the tool scored higher on the diagnostic portion of the rubric, with lower scores in the therapeutic recommendations. These findings highlight the importance of continuous refinement and validation of AI tools through rigorous evaluation frameworks like the OSCE. The adaptation of OSCEs for AI tools presents several advantages, including a structured and reproducible approach to evaluation, the ability to test AI systems in diverse clinical scenarios, and the opportunity to benchmark AI performance against established clinical standards to permit charting of future progress as aerospace medicine evolves as a discipline. Remaining challenges include ensuring that these evaluations capture the full spectrum of clinical decision-making scenarios that will be confronted by CMOs during missions and adequately reflecting real-world variability of the austere spaceflight environment. CONCLUSION Employing OSCEs to evaluate AI clinical decision support tools offers a promising approach to validating their clinical utility and efficacy. This methodology not only provides insights into the tool's performance but also fosters ongoing improvement and alignment with standard of care practices. Future research should focus on refining these evaluation processes and addressing limitations to enhance the integration of AI tools in clinical spaceflight settings. REFERENCES Scott S, Hearns V, Barker MA. Testing Clinical Skills: A Look at the OSCE and USMLE Clinical Skills Exams. S D Med. 2019 Oct;72(10):451-453. Majumder MAA, Kumar A, Krishnamurthy K, Ojeh N, Adams OP, Sa B. An evaluative study of objective structured clinical examination (OSCE): students and examiners perspectives. Adv Med Educ Pract. 2019 Jun 5;10:387-397. Karam VY, Park YS, Tekian A, Youssef N. Evaluating the validity evidence of an OSCE: results from a new medical school. BMC Med Educ. 2018 Dec 20;18(1):313.

Ariana M Nelson↗

The Uniform Methods Project: Smart Thermostat Evaluation Protocol

A smart thermostat is an internet-connected device that controls home heating, ventilation, and air-conditioning (HVAC) equipment and can automatically adjust temperature set points to optimize performance and achieve energy savings. Smart thermostat features often include two way communication, occupancy detection (such as geofencing and occupancy sensors), schedule learning, and seasonal optimization algorithms. Smart thermostats can control most conventional HVAC systems, including central air conditioners, heat pumps, and forced air furnaces. Several types of residential utility programs offer smart thermostats as replacements measures. Working with smart thermostat vendors, utilities can offer separate optimization programs to produce energy savings beyond those achieved by installing a smart thermostat. From an evaluation perspective, smart thermostat programs have several noteworthy features. First, the energy savings from a smart thermostat may change over the life of the device. As a smart thermostat is connected to the internet, original equipment manufacturers can update the thermostat software to improve the thermostat's energy efficiency. Likewise, users can adjust the thermostat settings and schedules over time in response to changes in weather, thermal comfort, energy prices, or preferences for energy efficiency. Additionally, many thermostat manufacturers offer seasonal optimization programs that recommend changes or make minor, automated adjustments to the thermostat settings to improve energy efficiency. These opt-in programs are now standard offerings for many smart thermostat manufacturers and provided at no additional cost to users. The potential for software updates and continuous optimization and the evolving nature of user interactions mean future energy savings may differ from first-year savings and the energy savings of smart thermostats may need to be evaluated more than once. Second, smart thermostats often have small unit energy savings relative to a home's total energy consumption, especially in comparison to whole- home retrofit programs. This can make it difficult to detect the smart thermostat savings in billing or advanced metering infrastructure (AMI) meter consumption data. For example, as cooling loads in many regions average about 20% of annual electricity consumption, smart thermostat savings of 10% of cooling energy use would equate to a 2% reduction in home electricity consumption. Evaluators should use regression analysis of whole-home billing consumption or advanced metering infrastructure (AMI) meter consumption data to evaluate smart thermostat savings because, as explained at greater length below , these data are usually available to evaluators and regression can control for the impacts of weather and other potentially confounding factors on a home's energy consumption. Finally, as with other energy efficiency programs, participation in smart thermostat programs is self-selective. As discussed at greater length below , smart thermostat participants tend to be, among other things, younger, higher-income, and more likely to adopt electric vehicles (EVs) and internet connected devices than nonparticipants. These differences are often unobservable to the evaluator and correlated with a home's energy consumption, creating the potential for bias in estimating savings. Due to the small unit savings of thermostats, errors and biases from self-selection that may not be very consequential when evaluating a whole- home retrofits (e.g., ±2% of home electricity consumption) can have a major impact when evaluating the savings and cost-effectiveness of smart thermostat programs. A percentage point change in the estimated savings could affect the cost-effectiveness of a program. This means it is important for evaluators to assess and to minimize the potential for error from selection bias in estimating smart thermostat program savings. The Uniform Methods Project provides model protocols for determining energy savings and demand reductions that result from specific energy efficiency measures implemented through state and utility programs. In most cases, the measure protocols are based on a particular option identified by the International Performance Verification and Measurement Protocol ; however, this work provides a more detailed approach to implementing that option. Each chapter is written by technical experts in collaboration with their peers, reviewed by industry experts, and subject to public review and comment. The UMP protocols can be used by utilities, program administrators, public utility commissions, evaluators, and other stakeholders for both program planning and evaluation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Development of a Technical, Economic, and Risk Assessment Tool for the Evaluation of Work Reduction Opportunities

Efficient and cost-effective operation of a nuclear power plant (NPP) is essential to ensuring long-term economical and safe operation. Multiple cost saving opportunities exist, referred to here as work reduction opportunities (WRO). These WROs reduce plant operating costs by employing various cost-effective strategies (e.g., implementation of modern technologies). Identifying and objectively screening WROs is an essential task to help reduce overall costs. However, there is no comprehensive framework for assessing WROs in the nuclear industry and evaluating their impact on plant operations. This report presents a novel framework for systematically evaluating WROs from a technical, economic, and risk perspective. As NPPs continue to add new technology and implement modernization strategies into their current processes, potential WROs are commonly identified. Although most WROs have the potential to reduce costs, not all opportunities will result in significant cost savings due to unforeseen risks, large implementation costs, or benefits that fall short of expectations. Examples of this can be the result of a technology that is not fully developed, uncertainty in the amount of cost reduction, or difficulties introducing a new process into an organization. These uncertainties can manifest several ways and can result in a WRO with limited cost savings or even a loss of investment. The framework developed emphasizes the importance of effectively screening the WROs from a holistic perspective to objectively identify inefficiencies and ensure a positive impact to the organization. This report presents the Technical, Economic, and Risk Assessment (TERA) as a key methodology for the screening and evaluation of potential WROs. The TERA framework begins with a screening phase where the process is examined through a hybrid combination of Lean Six Sigma and Integrated Operations for Nuclear (ION) guiding principles. This framework examines the current processes using the Lean Six Sigma SIPOC (Suppliers, Inputs, Process, Outputs, Consumers) methodology but retains the ION key elements of People, Technology, Process, and Governance as important factors to the nuclear decision-making process. By combining the principles of Lean Six Sigma and ION, the developed screening process is specific to the nuclear industry in that it systematically evaluates WROs in order to implement new technology that is comprehensively evaluated. The TERA begins by mapping current processes as they relate to WROs and examining the inefficiencies. Furthermore, the created process map can be used to identify and evaluate potential solutions. Using key performance indicators (KPIs), the TERA evaluates each area—technology, economics, and risk—for uncertainties and to perform cost-benefit analysis. The results of the TERA are important KPIs that allow for an evaluation of different processes and technology implementations. This assessment enables decision-makers to compare various WROs based on metrics and then make informed decisions for which opportunity to implement first. This research includes not only the creation of the TERA framework, but also the evaluation of its performance. A case study for screening potential WROs at Southern Nuclear Company is presented that utilizes the TERA methodology. Through the use of TERA, various WROs were screened, and the solutions evaluated for cost-benefit expectations. The report concludes by summarizing the overall effort and implications for utility modernization. The performance of the screening and TERA are discussed as well as the impact on the nuclear industry. The TERA process enables utilities to evaluate and inform investment decisions for WROs and mitigate any potential risks. Through this research, we provide utilities with a valuable framework to optimize operations, reduce costs, and drive continuous process improvement.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A Review of the Evaluation of Building Energy Code Compliance in the United States

Building energy codes are essential tools for achieving energy efficiency in buildings. However, the full energy savings potential of these codes can only be realized if buildings are constructed in compliance with them. Therefore, evaluating building energy code compliance is crucial in bridging the gap between the energy efficiency requirements set by energy codes and the actualized energy savings achieved. An energy code compliance evaluation serves as a mechanism to assess construction practices, evaluate the effectiveness of code enforcement, identify gaps in compliance, and guide strategies for improvement through training and education. Conducting code compliance evaluation activities involves field studies that require careful design and significant resources. Historically, more emphasis has been placed on developing and adopting building energy codes, while efforts to evaluate compliance have been relatively limited and lacking consistent approaches. The passage of the 2009 American Recovery and Reinvestment Act (ARRA), which mandated that states create plans for achieving 90% compliance within eight years, stimulated the need for an energy code compliance evaluation. As a result, federal, state, and local governments, and utilities have invested in the development of methodologies and tools for code compliance evaluation studies. This paper reviews the code compliance evaluation studies conducted in the United States over the past three decades. It describes and compares the methodologies and metrics used to assess building energy code compliance, summarizes the general elements and steps involved in the evaluation process, and discusses common issues in these studies. Over time, code compliance evaluation methodologies have evolved from isolated development within individual states, regions, and utilities, to widely accepted protocols applicable across different states and local jurisdictions. There has been a transition in compliance metrics, shifting from historical compliance rates to energy-consumption-oriented approaches.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Producing Evaluation-quality 239 Pu Average Prompt Fission Neutron Multiplicities using a Correlated Fission Model

An evaluation of the average prompt fission neutron multiplicity, $\bar{ν}_p$, of 239 Pu(n,f) is shown. This evaluation includes (a) the correlated fission model CGMF, and (b) a detailed analysis of past and recently published experimental data. Using CGMF-calculated $\bar{ν}_p$ as prior enables to link, through the use of evaluated model input parameters, $\bar{ν}_p$ to other fission observables such as the prompt fission neutron spectrum (PFNS), preneutron emission fission yields as a function of mass, and the average total kinetic energy of the fragments. These evaluated parameters produce realistic predictions of many fission observables, while the evaluated $\bar{ν}_p$ agrees well (χ 2 ≈ 1) with data. Moreover, with the new evaluated $\bar{ν}_p$, the effective neutron multiplication factor of fast Pu ICSBEP critical assemblies are predicted with a mean bias of 58 pcm compared to 18 pcm with ENDF/B-VIII.0, when paired with a new 239 Pu PFNS and fission cross section. Due to these encouraging validation results, the evaluated $\bar{ν}_p$ is currently part of a release candidate for the 239 Pu ENDF/B-VIII.1 file. Hence, a correlated fission model was used for the first time for evaluating $\bar{ν}_p$ that is of evaluation quality. This is an important step towards consistent evaluations of prompt fission observables.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Evaluated 238 U(n,f) Average Prompt Fission Neutron Multiplicities Including the CGMF Model

This report documents an evaluation of the average prompt fission neutron multiplicity, $\overline{v}_p$, of 238 U from 800 keV to MeV. This evaluation had to be re-done from “scratch” as the input to previous $\overline{v}_p$ evaluations, specifically ENDF/B-VIII.0, was not found. That means that all available experimental data were re-analyzed and uncertainties were re-estimated. The new evaluated 238 U $\overline{v}_p$ based on only experimental data differs distinctly from ENDF/B-VIII.0 $\overline{v}_p$ from 2 to 4.5 MeV, and from 6 to 7 MeV, and is otherwise similar. The difference from 2 to 4.5 MeV stems from the fact that ENDF/B-VIII.0 was tweaked in this energy range to data of Frehaut, while two other, equally trustworthy, data sets would indicate an evaluated 238 U $\overline{v}_p$ that is up to 2% higher. Also, second chance fission in ENDF/B-VIII.0 was smoothed over from 6–7 MeV. Another major difference to ENDF/B-VIII.0 is that one of the evaluations presented here includes model information from the Hauser-Feshbach fission fragment decay code CGMF, while ENDF/B-VIII.0 is based purely on experimental data. CGMF links several fission quantities with each other; $\overline{v}_p$ is predicted by assumptions made on, e.g., pre-neutron emission yields as a function of mass, the total kinetic energy, or spin and parity of fission fragments. This allows to validate the new 238 U $\overline{v}_p$ by using CGMF parameters obtained from fitting to experimental 238 U $\overline{v}_p$ to predict yields as a function of mass, the average total kinetic energy, or the mean energy of the prompt fission neutron spectrum. These model-predicted values can then be compared to experimental and evaluated data. The model-predicted fission-observable values using evaluated parameters obtained here are reasonably close to experimental data for some observables, but are farther away from experimental data related to TKE observables. In addition to that, the evaluated 238 U(n,f) $\overline{v}_p$ shows similar deviations from ENDF/B-VIII.0 as for the evaluation with only experimental data. This difference is expected to lead to changes in simulated effective neutron multiplication factor, $k_{eff}$ of ICSBEP critical assemblies that are sensitive to 238 U in the fast range (BigTen, Flattop, Flattop-Pu). These changes in $k_{eff}$ need to be counter-balanced. Chi-Nu PFNS experimental data are expected to be released in the next few months that might lead to the needed changes in the PFNS. Until then, we hold off in benchmarking the new 238 U(n,f) $\overline{v}_p$ as well as submitting it to ENDF/B-VIII.1. Also, new high-precision 238 U $\overline{v}_p$ are expected to be measured by the CEA in the next two years that will shed further light on question on 238 U $\overline{v}_p$ from 2–4.5 and 6–7 MeV.

238U↗

Consistent $\overline{ν}$ evaluation for minor U isotopes with $\tt{CGMF}$

Following several successful prompt $\overline{ν}$ evaluations using $\tt{CGMF}$, including consistent evaluations for minor Pu isotopes, we detail in this report our efforts to perform a consistent $\overline{ν}$ evaluation for minor U isotopes during FY25. Although we have not yet produced a finalized evaluation, we present the progress that we have made towards such an evaluation for 232,233,234,236,237,239 U prompt $\overline{ν}$. Our milestone explicitly calls out evaluations for 233 U, 234 U, and 236 U, however, to better constrain the model with reliable experimental $\overline{ν}$ data, we also include 235 U and 238 U in the evaluation procedure. Then, we additionally produce evaluations for 232 U, 237 U and 239 U $\overline{ν}$ as a byproduct. Elsewhere, we will report our efforts on a stand-alone 233 U $\overline{ν}$ evaluation. This report is organized in the following manner. In Sec. 2, we briefly outline the updates to CGMF that were needed to be able to calculate all of these minor U fission reactions. The experimental data overview is given in Sec. 3. The evaluation methodology and results are presented in Secs. 4 and 5, respectively. Finally, we conclude and outline work for FY26 in Sec. 6.

07 ISOTOPE AND RADIATION SOURCES↗

Evaluation of image quality

This presentation outlines in viewgraph format a general approach to the evaluation of display system quality for aviation applications. This approach is based on the assumption that it is possible to develop a model of the display which captures most of the significant properties of the display. The display characteristics should include spatial and temporal resolution, intensity quantizing effects, spatial sampling, delays, etc. The model must be sufficiently well specified to permit generation of stimuli that simulate the output of the display system. The first step in the evaluation of display quality is an analysis of the tasks to be performed using the display. Thus, for example, if a display is used by a pilot during a final approach, the aesthetic aspects of the display may be less relevant than its dynamic characteristics. The opposite task requirements may apply to imaging systems used for displaying navigation charts. Thus, display quality is defined with regard to one or more tasks. Given a set of relevant tasks, there are many ways to approach display evaluation. The range of evaluation approaches includes visual inspection, rapid evaluation, part-task simulation, and full mission simulation. The work described is focused on two complementary approaches to rapid evaluation. The first approach is based on a model of the human visual system. A model of the human visual system is used to predict the performance of the selected tasks. The model-based evaluation approach permits very rapid and inexpensive evaluation of various design decisions. The second rapid evaluation approach employs specifically designed critical tests that embody many important characteristics of actual tasks. These are used in situations where a validated model is not available. These rapid evaluation tests are being implemented in a workstation environment.

Pavel, M.↗

A Framework for Evaluating Distributed Electric Propulsion on the SUSAN Electrofan Aircraft

This work presents a framework for evaluating models and algorithms for Distributed Electric Propulsion (DEP) on the SUSAN Electrofan Aircraft. Throughout the development of the SUSAN aircraft, the performance of various configurations of the aircraft will need to be analyzed. However, the static behavior alone is not sufficient to describe the performance of these configurations. Therefore, simulation with fully integrated subsystem models is required. The proposed framework considers the vehicle aerodynamic, propulsion, and control subsystems. The presented framework automatically generates control laws for any vehicle configuration in response to changes in these subsystems. To compare these different vehicle configurations, various time and frequency domain performance metrics are compared. Three different system modifications are used as cases to evaluate this framework. The first modification integrates the propulsion control system with the flight controller to enable differential thrust without stalling the main engine. This evaluation case is used to validate the framework for aircraft configurations with coupled subsystems. The second modification compares the effect of the vertical tail size on open and closed loop performance. This evaluation case is used to validate the framework for controlling different configurations and tuning towards comparable closed loop performance despite changes to the aircraft's aerodynamic model. The third modification implements two different control allocation schemes. This evaluation case demonstrates the framework's ability to evaluate allocation modifications needed to take advantage of DEP. The first evaluation case is used to show that controller integration enables differential thrust, improving realized wingfan bandwidth by up to 40\% in simulation. The second evaluation case demonstrates that the framework can stabilize the reduced tail size aircraft with closed loop control. The third evaluation case demonstrates that a pseudoinverse control allocation scheme improves lateral velocity settling time by approximately 17~seconds over a symmetric-thrust allocation. These cases show that the framework is useful for evaluating the performance of integrated system designs, enabling analyses of new models and algorithms for the SUSAN distributed electric propulsion vehicle.

Nicholas C Ogden↗

Status of the International Criticality Safety Benchmark Evaluation Project

The International Criticality Safety Benchmark Evaluation Project (ICSBEP) has continued its work generating evaluations of new and historical criticality benchmark experiments since the last update to the nuclear criticality safety community at the 11th International Conference on Nuclear Criticality Safety (ICNC 2019) in Paris, France. Three additional versions of the ICSBEP Handbook have been published since that update, and the Technical Review Group (TRG) held two in-person (in 2019 and 2023) and three virtual (2020 and 2021) meetings to review and approve additional benchmarks. The 2019 edition of the ICSBEP Handbook included five new evaluations with 79 new configurations, the 2020 version of the ICSBEP Handbook contained five new evaluations totaling 76 new configurations, and the 2021 version of the ICSBEP handbook contained five new evaluations with a total of 57 different configurations. The ICSBEP TRG met in October and December 2021, to review benchmarks for potential inclusion in the 2022 ICSBEP Handbook, with seven evaluations receiving provisional approval pending resolution of review group comments. Final comment resolution for some of these evaluations is currently underway and handbook publication should be completed soon. The ICSBEP TRG met again in person in April 2023 to review benchmarks for the 2023 ICSBEP Handbook, provisionally approving 7 new evaluations. The ICSBEP continues to deliver high-quality, peer reviewed evaluations of experiments relevant to the nuclear criticality safety community.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Criticality Safety Evaluation Project Development for University of California Berkely Nuclear Criticality Safety Pipeline Course

The Nuclear Criticality Safety Division at Lawrence Livermore National Laboratory (LLNL) has taken a unique approach to developing criticality safety evaluation topics in support of the University of California Berkeley criticality safety pipeline course. The evaluation topics are designed to go beyond the typical evaluation examples used for many training courses including vault storage and variations on storage arrays. These types of evaluations provide in-depth analysis into the fundamentals of criticality safety and are complex but may be far off from what a new criticality safety engineer may actually be evaluating. To provide more practical examples of criticality safety evaluation topics that are better fit for the knowledge level of a criticality safety engineer in-training, variations of current and future operations and research operations performed at LLNL are used as evaluation topics. Additionally, an emphasis on research is included in all evaluation topics as it allows students to take advantage of the concepts learned in class to apply them for process improvement, engineering equipment that is favorable for criticality safety, and negotiation tactics to work with operations personnel. The process used by LLNL to develop project topics for the pipeline course is provided in this paper. The intent is to provide an alternative technique for training students and potentially younger staff members in criticality safety on developing criticality safety evaluations.

42 ENGINEERING↗

Evaluation of proposed Section III Division 5 Class B rules: Piping example problem

Evaluation of proposed Section III Division 5 Class B rules: Piping example problem to be presented include Z pipe geometry and model problem statement, Z pipe loads and design inputs problem statement, Primary Stress Limit overview, Primary stress limit pseudo yield stress calculation, Primary stress limited factored load procedure, Strain limit evaluation composite load cycle, Strain limit evaluation pseudo yield stress, Strain limit evaluation strain limit criteria and ratcheting check, Creep fatigue damage evaluation overview, Creep fatigue damage evaluation alternating stress calculations, Creep fatigue damage evaluation lower bound stress, Creep fatigue damage evaluation stress relaxation history, Creep fatigue damage evaluation damage fractions, and Recommendations to consider for Z-pipe problem.

42 ENGINEERING↗

Release of Evaluated 235 U(n,f) Average Prompt Fission Neutron Multiplicities Including the CGMF Model

This report documents an evaluation of the average prompt fission neutron multiplicity, $\overline{v}_p$, of 235 U from 200 keV to 15 MeV that is a potential release candidate for the upcoming U.S. nuclear data library, ENDF/B-VIII.1. This evaluation had to be re-done from "scratch", as the input to the $\overline{v}_p$ evaluation of the previous library, ENDF/B-VIII.0, was lost. That means that all available experimental data were re-analyzed and uncertainties were re-estimated. Another major difference to ENDF/B-VIII.0 is that this evaluation includes model information from the Hauser-Feshbach fission fragment decay code CGMF, while ENDF/B-VIII.0 is based purely on experimental data. CGMF links several fission quantities with each other; $\overline{v}_p$ is predicted by assumptions made on, e.g., pre-neutron emission yields as a function of mass, the total kinetic energy, or spin and parity of fission fragments. This allows to perform two types of validation for the new 235 U $\overline{v}_p$: On the one hand, one can employ evaluated CGMF parameters obtained from fitting to experimental 235 U $\overline{v}_p$ to predict yields as a function of mass, the average total kinetic energy, or the mean energy of the prompt fission neutron spectrum. These model-predicted values can then be compared to experimental and evaluated data. The model-predicted fission-observable values using evaluated parameters obtained here are reasonably close to experimental data indicating the evaluated 235 U(n,f) $\overline{v}_p$ are physical. On the other hand, one can validate 235 U $\overline{v}_p$ with respect to integral responses such as fast ICSBEP critical assemblies or LLNL pulsed spheres. LLNL pulsed-sphere neutron-leakage spectra are minimally impacted by the new 235 U $\overline{v}_p$ as these experimental data are shape data and the $\overline{v}_p$ would mostly lead to a change in normalization of the data as the spheres are relatively thin (0.7 and 1.5 mean-free path) and, thus, mostly depend on 235 U $\overline{v}_p$ from 12-15 MeV. The change in the predicted effective neutron multiplication factor, k eff , of selected ICSBEP critical assemblies, however, is large compared to values using ENDF/B-VIII.0 and experimental k eff : The average bias is 108 pcm across all studied k eff values versus 12 pcm for ENDF/B-VIII.0. A reasonable performance in simulating keff (mean bias of 14 pcm) can be retained by tweaking 235 U $\overline{v}_p$ from 3-5 MeV, and combining it with a recent 235 U PFNS evaluation that is also a ENDF/B-VIII.1 release candidate.

235U↗

Engineering flight and guest pilot evaluation report, phase 2

Prior to the flight evaluation, the two-segment profile capabilities of the DC-8-61 were evaluated and flight procedures were developed in a flight simulator at the UA Flight Training Center in Denver, Colorado. The flight evaluation reported was conducted to determine the validity of the simulation results, further develop the procedures and use of the area navigation system in the terminal area, certify the system for line operation, and obtain evaluations of the system and procedures by a number of pilots from the industry. The full area navigation capabilities of the special equipment installed were developed to provide terminal area guidance for two-segment approaches. The objectives of this evaluation were: (1) perform an engineering flight evaluation sufficient to certify the two-segment system for the six-month in-service evaluation; (2) evaluate the suitability of a modified RNAV system for flying two-segment approaches; and (3) provide evaluation of the two-segment approach by management and line pilots.

Morrison, J. A.↗

NASA PC software evaluation project

The USL NASA PC software evaluation project is intended to provide a structured framework for facilitating the development of quality NASA PC software products. The project will assist NASA PC development staff to understand the characteristics and functions of NASA PC software products. Based on the results of the project teams' evaluations and recommendations, users can judge the reliability, usability, acceptability, maintainability and customizability of all the PC software products. The objective here is to provide initial, high-level specifications and guidelines for NASA PC software evaluation. The primary tasks to be addressed in this project are as follows: to gain a strong understanding of what software evaluation entails and how to organize a structured software evaluation process; to define a structured methodology for conducting the software evaluation process; to develop a set of PC software evaluation criteria and evaluation rating scales; and to conduct PC software evaluations in accordance with the identified methodology. Communication Packages, Network System Software, Graphics Support Software, Environment Management Software, General Utilities. This report represents one of the 72 attachment reports to the University of Southwestern Louisiana's Final Report on NASA Grant NGT-19-010-900. Accordingly, appropriate care should be taken in using this report out of context of the full Final Report.

Dominick, Wayne D.↗

Towards Reliable Evaluation of Anomaly-Based Intrusion Detection Performance

This report describes the results of research into the effects of environment-induced noise on the evaluation process for anomaly detectors in the cyber security domain. This research was conducted during a 10-week summer internship program from the 19th of August, 2012 to the 23rd of August, 2012 at the Jet Propulsion Laboratory in Pasadena, California. The research performed lies within the larger context of the Los Angeles Department of Water and Power (LADWP) Smart Grid cyber security project, a Department of Energy (DoE) funded effort involving the Jet Propulsion Laboratory, California Institute of Technology and the University of Southern California/ Information Sciences Institute. The results of the present effort constitute an important contribution towards building more rigorous evaluation paradigms for anomaly-based intrusion detectors in complex cyber physical systems such as the Smart Grid. Anomaly detection is a key strategy for cyber intrusion detection and operates by identifying deviations from profiles of nominal behavior and are thus conceptually appealing for detecting "novel" attacks. Evaluating the performance of such a detector requires assessing: (a) how well it captures the model of nominal behavior, and (b) how well it detects attacks (deviations from normality). Current evaluation methods produce results that give insufficient insight into the operation of a detector, inevitably resulting in a significantly poor characterization of a detectors performance. In this work, we first describe a preliminary taxonomy of key evaluation constructs that are necessary for establishing rigor in the evaluation regime of an anomaly detector. We then focus on clarifying the impact of the operational environment on the manifestation of attacks in monitored data. We show how dynamic and evolving environments can introduce high variability into the data stream perturbing detector performance. Prior research has focused on understanding the impact of this variability in training data for anomaly detectors, but has ignored variability in the attack signal that will necessarily affect the evaluation results for such detectors. We posit that current evaluation strategies implicitly assume that attacks always manifest in a stable manner; we show that this assumption is wrong. We describe a simple experiment to demonstrate the effects of environmental noise on the manifestation of attacks in data and introduce the notion of attack manifestation stability. Finally, we argue that conclusions about detector performance will be unreliable and incomplete if the stability of attack manifestation is not accounted for in the evaluation strategy.

cyber defense↗