Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Code of Record”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Hindsight logging for model training

In modern Machine Learning, model training is an iterative, experimental process that can consume enormous computation resources and developer time. To aid in that process, experienced model developers log and visualize program variables during training runs. Exhaustive logging of all variables is infeasible, so developers are left to choose between slowing down training via extensive conservative logging, or letting training run fast via minimalist optimistic logging that may omit key information. As a compromise, optimistic logging can be accompanied by program checkpoints; this allows developers to add log statements post-hoc, and "replay" desired log statements from checkpoint---a process we refer to as hindsight logging. Unfortunately, hindsight logging raises tricky problems in data management and software engineering. Done poorly, hindsight logging can waste resources and generate technical debt embodied in multiple variants of training code. In this paper, we present methodologies for efficient and effective logging practices for model training, with a focus on techniques for hindsight logging. Our goal is for experienced model developers to learn and adopt these practices. To make this easier, we provide an open-source suite of tools for Fast Low-Overhead Recovery (flor) that embodies our design across three tasks: (i) efficient background logging in Python, (ii) adaptive periodic checkpointing, and (iii) an instrumentation library that codifies hindsight logging for efficient and automatic record-replay of model-training. Model developers can use each flor tool separately as they see fit, or they can use flor in hands-free mode, entrusting it to instrument their code end-to-end for efficient record-replay. Our solutions leverage techniques from physiological transaction logs and recovery in database systems. Evaluations on modern ML benchmarks demonstrate that flor can produce fast checkpointing with small user-specifiable overheads (e.g. 7%), and still provide hindsight log replay times orders of magnitude faster than restarting training from scratch.

Computer Science↗

Software Quality Assurance for EBR-II Fuels Irradiation and Physics Database (FIPD)

The Fuels Irradiation and Physics Database (FIPD) is an ongoing DOE project on archival of the EBR-II metal-alloy fuel irradiation experiments. As part of its use in support of license applications, the Quality Assurance Program Plan (QAPP) was drafted and endorsed by NRC in an effort to demonstrate its compliance with regulatory expectations. Software Quality Assurance (SQA) for the physics portion of FIPD is intended to qualify the calculated quantities such as fuel and cladding temperatures, neutron fluence and axially varying burnup estimates for irradiated fuel elements. This report covers the initial evaluation of SQA status of three neutron physics and thermo-fluid codes (REBUS, RCT and SE2RCT) that form the basis of calculated quantities for as-irradiated characteristics of the tested metallic fuel elements. The report also introduces an SQA plan to address the identified deficiencies. The REBUS, RCT, and SE2RCT codes are all part of the Argonne Reactor Code (ARC) code system. There is considerable knowledge and experience on REBUS and RCT but relatively less on SE2RCT. During FY2021, efforts focused on an assessment of how the data in the EBR-II Physics and Analysis DataBase (PADB) is generated with SE2RCT and used in FIPD. Additional tasks included considerations of uncertainties for power estimates in REBUS and RCT calculations and their impact on the combined RCT methodology. The RCT software usage in FIPD was assessed this year and the input/output details studied. A “requirements” document was created that identifies the key features of the RCT software being used in FIPD that need to have SQA documentation. A brief discussion on the history of RCT and its input is included in this report along with the basic SQA roadmap laid out in the requirements document. The SE2RCT software usage in FIPD is still being studied noting that there is no current manual. As part of the work done this year, two bugs were identified in the SE2RCT software which have a minor impact on the accuracy of the results it produces. No requirements document has been created, but one identified feature of SE2RCT being used that needs verification was its fuel pin temperature calculation. The work completed this year confirms that the approximations which will be included in the software verification report for SE2RCT are accurate. In addition to software quality assurance work for RCT and SE2RCT, an automated verification framework is proposed to simplify the software quality assurance process. The purpose of this framework is to streamline code verification and documentation while minimizing repetitive tasks for code developers and reviewers. The reduction of repeated input (between reference solution, software, and documentation input) throughout the SQA process reduces potential for human errors during the preparation of the supporting software quality records. The automation of the verification and documentation process proposed for this project leverages the existing verification structure already in place for the SAS4A/SASSYS-1 code.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

r -process nucleosynthesis and kilonovae from hypermassive neutron star post-merger remnants

ABSTRACT We investigate r-process nucleosynthesis and kilonova emission resulting from binary neutron star (BNS) mergers based on a three-dimensional (3D) general-relativistic magnetohydrodynamic (GRMHD) simulation of a hypermassive neutron star (HMNS) remnant. The simulation includes a microphysical finite-temperature equation of state (EOS) and neutrino emission and absorption effects via a leakage scheme. We track the thermodynamic properties of the ejecta using Lagrangian tracer particles and determine its composition using the nuclear reaction network SkyNet. We investigate the impact of neutrinos on the nucleosynthetic yields by varying the neutrino luminosities during post-processing. The ejecta show a broad distribution with respect to their electron fraction Ye, peaking between ∼0.25–0.4 depending on the neutrino luminosity employed. We find that the resulting r-process abundance patterns differ from solar, with no significant production of material beyond the second r-process peak when using luminosities recorded by the tracer particles. We also map the HMNS outflows to the radiation hydrodynamics code SNEC and predict the evolution of the bolometric luminosity as well as broadband light curves of the kilonova. The bolometric light curve peaks on the timescale of a day and the brightest emission is seen in the infrared bands. This is the first direct calculation of the r-process yields and kilonova signal expected from HMNS winds based on 3D GRMHD simulations. For longer-lived remnants, these winds may be the dominant ejecta component producing the kilonova emission.

79 ASTRONOMY AND ASTROPHYSICS↗

Probabilistic Seismic Hazard Analysis for Iraq Based on the Updated Earthquake Catalog (1900-2021) and Ground Motion Characteristics

Onur et al. (2017) compiled the first comprehensive earthquake catalog for Iraq, covering 1900 to 2009 within 26°–40°N latitude and 36°–51°E longitude. This catalog was utilized in a probabilistic seismic hazard assessment (PSHA) by Abdulnaby et al. (2020) to aid in updating Iraq’s building code seismic provisions. Recently, we have updated the earthquake catalog for Iraq by adding earthquakes recorded from 2010 to 2021 and directly calculating moment magnitude (Mw) for about 2,800 earthquakes using the coda envelope methodology and waveform data from the Mesopotamian Seismological Network (MPSN) in Iraq.

58 GEOSCIENCES↗

HAPPA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded

High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method -- HAppA-LSTM -- achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of keywords representing the source code. A comprehensive importance analysis of these keywords further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.

Jiang, Hailong [Kent State University]↗

Qualification of Digitized Legacy Fast Reactor Data

The Integral Fast Reactor (IFR) fuel compatibility test program (1984-1994) included a variety of fuel pin examinations conducted at the Hot Fuel Examination Facility (HFEF) and the Alpha-Gamma Hot Cell Facility (AGHCF). Hard copy data records of these examinations have been recovered, scanned, and preserved in PDF format. Many hard copy records are now qualified in accordance with an NRC-approved Quality Assurance Program Plan (QAPP), and there is an ongoing effort to qualify additional legacy records. This legacy fuel performance data is vital to support design and licensing of fast reactors with validation of state-of-the-art codes and advanced methods for design and analysis. Stakeholders can most easily utilize this data when the PDF scans have been converted into digital data tables. However, qualification of the scanned hard copy data does not qualify the digital data file resulting from the digitization of the data contained in the record; the subject matter expert (SME) must make a review of the digitized data table as well before it can be designated as qualified. This report outlines a peer review process to qualify the digital data file(s), typically in CSV format, corresponding to hard copy records in accordance with the existing QAPP.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Mneme

A simple tool allowing recording the execution of a GPU (CUDA) kernel and replaying that kernel as an independent executable. The tool operates in 3 phases. During compile time the user needs to apply a provided LLVM pass to instrument the code. The pass detects all device global variables and device functions and stores this information with the respective LLVM-IR in the global device memory. The compilation generates a record-able executable. The second phase involves running the application executable with a desired input and using LD_PRELOAD to enable recording. When recording before invoking a device kernel the pre-loaded library stores device memory in persistent storage and associates the memory with the device kernel and an LLVM IR file. At the end of the recorded execution the pre-load library generates a database in the form of a JSON file containing information regarding the LLVM-IR files and the snapshots of device memory. During the third and last phase the user can replay the execution of an kernel as a separate independent executable. Besides executing it the user can modify the LLVM IR file and auto-tune parameters such as kernel launch-bounds or kernel runtime execution parameters (e.g. Kernel Block and Grid Dimensions). Is

Parasyris, Konstantinos↗

Women in Data Science Livermore Datathon 2024

The WiDS Datathon 2024 focuses on a prediction task using a roughly 39k record dataset (split into training and test sets) representing patients and their characteristics (age, race, BMI, zip code), their diagnosis and treatment information (breast cancer diagnosis code, metastatic cancer diagnosis code, metastatic cancer treatments, … etc.), their geo (zip-code level) demographic data (income, education, rent, race, poverty, …etc), as well as toxic air quality data (Ozone, PM25 and NO2) that tie health outcomes to environmental conditions. Each row in the data corresponds a single patient and her Diagnosis Period.

59 BASIC BIOLOGICAL SCIENCES↗

Controlling accesses to a branch prediction unit for sequences of fetch groups

An electronic device handles accesses of a branch prediction functional block when executing instructions in program code. The electronic device includes a processor having the branch prediction functional block that provides branch prediction information for control transfer instructions (CTIs) in the program code and a minimum predictor use (MPU) functional block. The MPU functional block determines, based on a record associated with a given fetch group of instructions, that a specified number of subsequent fetch groups of instructions that were previously determined to include no CTIs or conditional CTIs that were not taken are to be fetched for execution in sequence following the given fetch group. The MPU functional block then, when each of the specified number of the subsequent fetch groups is fetched and prepared for execution, prevents corresponding accesses of the branch prediction functional block for acquiring branch prediction information for instructions in that subsequent fetch group.

Agrawal, Varun↗

Wind and Temperature Consensus at Horn Point, HU-Beltsville, Piney Run (Maryland) in support of CoURAGE

The Maryland Department of the Environment (MDE) operates a ground-based atmospheric profiling network consisting of collocated radar wind profilers (RWP) and radio acoustic sounding systems (RASS) as part of its Ambient Air Monitoring Program. This network provides continuous observations of wind and temperature structure in the lower troposphere to support air quality forecasting, regulatory analysis, and atmospheric research. The network currently includes three fixed sites across Maryland: Horn Point (HP, lower eastern shore) [38.587525°,-76.141006°], Howard University-Beltsville (HUB, central Maryland) [39.055277°, -76.878632°], and Piney Run (PR, western Maryland) [39.705950°, -79.012000°] The network is designed to capture regional variability in atmospheric transport and boundary-layer processes. These systems measure vertical profiles of horizontal wind speed and direction using Doppler radar techniques, with observations typically spanning from ~100 m above ground level up to approximately 2.5–4 km. Measurements are derived from the Doppler shift of backscattered electromagnetic signals, enabling retrieval of wind vectors at multiple altitudes with high temporal resolution (e.g., 30-minute averages reported every 6 minutes). Each radar wind profiler is paired with a Radio Acoustic Sounding System (RASS) to provide profiles of virtual temperature in the lower atmosphere (~100–200 m AGL) by measuring the propagation speed of acoustic waves. Together, the RWP/RASS system yields a coupled data set of thermodynamic and kinematic atmospheric structure, including additional parameters such as vertical velocity, radial velocity, signal-to-noise ratio, and spectral width for advanced analysis. There are two types of files for each station: wind data (files with a "w" prefix) and virtual temperature RASS data (files with a "t" prefix). The wind data files are in the format wYYDDD.cns, where YY is the 2-digit year and DDD is the day of the year. The RASS virtual temperature data files are in the format tYYDDD.cns. Each record has the following header structure: Line 1 : Station Name RASS files Line 2 : RASS rev DeTect_2.0, WINDS files Line 2 : WINDS rev ATI 5.1 Line 3 : N latitude, W longitude, and site elevation (m) Line 4 : Date and begin time of consensus: yy mm dd hh mn ss plus # minutes to add to get UTC Line 5 : Consensus averaging time (minutes); number of beams; number of range gates Line 6 : Number of records required to make consensus (num) total number of records (tot) and the consensus window size (m/s) in the format: num:tot (window) RASS files Line 7 : no. of coded cells, no. of spec, pulse width (ns), and inter-pulse period (µs), WINDS files Line 7 : No. of coded cells, no. of spectra, pulse width (ns), and inter-pulse period (µs), each with a pair of values: first value is for oblique beams, second for vertical RASS files Line 8 : Full scale Doppler value (m/s) Delay to first gate (ns) Number of gates Spacing of gates (ns), WINDS files Line 8 : Full scale Doppler velocity (m/s), oblique and vertical Vertical correction applied to oblique beams? (0 = no, 1 = yes) Delay to first gate (ns), oblique and vertical Number of gates, oblique and vertical Spacing of gates (ns), oblique and vertical Line 9 : Azimuth and elevation (9s indicate vertical beam not used) RASS files Line 10, values : HT = Height above ground (km), T = Uncorrected virtual temperature consensus (deg C), Tc = Corrected virtual temperature consensus (deg C), W = Vertical wind consensus (9s indicate vertical beam not used, w-component, positive upward, m/s), CNT = Number of records that made consensus (for the 3 values in same order), SNR = Average signal to noise ratio (dB) of records in consensus (same order) WINDS files Line 10, values : HT = Height above ground (km), SPD = Wind speed (m/s), DIR = Wind direction (deg E of N from N), RAD = Radial velocities for each beam (m/s) in order given in azimuth and elevation line (positive toward radar; 9s indicate vertical beam not used, CNT = Number of records that made consensus, SNR = Average signal to noise ratio (dB) of records in consensus

{"wind speed and direction",temperature}↗

Clinical knowledge extraction via sparse embedding regression (KESER) with multi-center large scale electronic health record data

The increasing availability of electronic health record (EHR) systems has created enormous potential for translational research. However, it is difficult to know all the relevant codes related to a phenotype due to the large number of codes available. Traditional data mining approaches often require the use of patient-level data, which hinders the ability to share data across institutions. In this project, we demonstrate that multi-center large-scale code embeddings can be used to efficiently identify relevant features related to a disease of interest. We constructed large-scale code embeddings for a wide range of codified concepts from EHRs from two large medical centers. We developed knowledge extraction via sparse embedding regression (KESER) for feature selection and integrative network analysis. We evaluated the quality of the code embeddings and assessed the performance of KESER in feature selection for eight diseases. Besides, we developed an integrated clinical knowledge map combining embedding data from both institutions. The features selected by KESER were comprehensive compared to lists of codified data generated by domain experts. Features identified via KESER resulted in comparable performance to those built upon features selected manually or with patient-level data. The knowledge map created using an integrative analysis identified disease-disease and disease-drug pairs more accurately compared to those identified using single institution data. Analysis of code embeddings via KESER can effectively reveal clinical knowledge and infer relatedness among codified concepts. KESER bypasses the need for patient-level data in individual analyses providing a significant advance in enabling multi-center studies using EHR data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Question-answering system extracts information on injection drug use from clinical notes

Background. Injection drug use (IDU) can increase mortality and morbidity. Therefore, identifying IDU early and initiating harm reduction interventions can benefit individuals at risk. However, extracting IDU behaviors from patients’ electronic health records (EHR) is difficult because there is no other structured data available, such as International Classification of Disease (ICD) codes, and IDU is most often documented in unstructured free-text clinical notes. Although natural language processing can efficiently extract this information from unstructured data, there are no validated tools. Methods. Here, to address this gap in clinical information, we design a question-answering (QA) framework to extract information on IDU from clinical notes for use in clinical operations. Our framework involves two main steps: (1) generating a gold-standard QA dataset and (2) developing and testing the QA model. We use 2323 clinical notes of 1145 patients curated from the US Department of Veterans Affairs (VA) Corporate Data Warehouse to construct the gold-standard dataset for developing and evaluating the QA model. We also demonstrate the QA model’s ability to extract IDU-related information from temporally out-of-distribution data. Results. Here, we show that for a strict match between gold-standard and predicted answers, the QA model achieves a 51.65% F1 score. For a relaxed match between the gold-standard and predicted answers, the QA model obtains a 78.03% F1 score, along with 85.38% Precision and 79.02% Recall scores. Moreover, the QA model demonstrates consistent performance when subjected to temporally out-of-distribution data. Conclusions. Our study introduces a QA framework designed to extract IDU information from clinical notes, aiming to enhance the accurate and efficient detection of people who inject drugs, extract relevant information, and ultimately facilitate informed patient care.

60 APPLIED LIFE SCIENCES↗

2003 Interstate 595 Vehicle Trip-Length Study

# 2003 Interstate 595 Vehicle Trip-Length Study The 2003 Vehicle Trip-Length Study focused on Interstate 595 between Davie Road and University Drive in Florida. Survey participants answered questions about their trip's origin and destination—including the type of location such as work, home, store, etc.—the on- and off-ramps used, how many people were in the car, and the type of vehicle. The survey also collected household demographic data such as annual household income, available vehicles, the number of people living in their household, and the number of workers in their household above the age of 16. It also asked if they would use proposed bus-only lanes or train service along the corridor, if available. ## Data Collection Agency The survey was conducted by and for the Florida Department of Transportation. ## Survey Methodology The survey was conducted via mail and online in March 2003. ## Survey Records, Data, and Documentation Survey records include 7,917 participants. Origin and destination locations include street addresses, nearest intersections or landmarks, city, state, zip code, and latitude/longitude.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2003 Interstate 595 Vehicle Trip-Length Study

# 2003 Interstate 595 Vehicle Trip-Length Study The 2003 Vehicle Trip-Length Study focused on Interstate 595 between Davie Road and University Drive in Florida. Survey participants answered questions about their trip's origin and destination—including the type of location such as work, home, store, etc.—the on- and off-ramps used, how many people were in the car, and the type of vehicle. The survey also collected household demographic data such as annual household income, available vehicles, the number of people living in their household, and the number of workers in their household above the age of 16. It also asked if they would use proposed bus-only lanes or train service along the corridor, if available. ## Data Collection Agency The survey was conducted by and for the Florida Department of Transportation. ## Survey Methodology The survey was conducted via mail and online in March 2003. ## Survey Records, Data, and Documentation Survey records include 7,917 participants. Origin and destination locations include street addresses, nearest intersections or landmarks, city, state, zip code, and latitude/longitude.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2003 Interstate 595 Vehicle Trip-Length Study

# 2003 Interstate 595 Vehicle Trip-Length Study The 2003 Vehicle Trip-Length Study focused on Interstate 595 between Davie Road and University Drive in Florida. Survey participants answered questions about their trip's origin and destination—including the type of location such as work, home, store, etc.—the on- and off-ramps used, how many people were in the car, and the type of vehicle. The survey also collected household demographic data such as annual household income, available vehicles, the number of people living in their household, and the number of workers in their household above the age of 16. It also asked if they would use proposed bus-only lanes or train service along the corridor, if available. ## Data Collection Agency The survey was conducted by and for the Florida Department of Transportation. ## Survey Methodology The survey was conducted via mail and online in March 2003. ## Survey Records, Data, and Documentation Survey records include 7,917 participants. Origin and destination locations include street addresses, nearest intersections or landmarks, city, state, zip code, and latitude/longitude.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2003 Interstate 595 Vehicle Trip-Length Study

# 2003 Interstate 595 Vehicle Trip-Length Study The 2003 Vehicle Trip-Length Study focused on Interstate 595 between Davie Road and University Drive in Florida. Survey participants answered questions about their trip's origin and destination—including the type of location such as work, home, store, etc.—the on- and off-ramps used, how many people were in the car, and the type of vehicle. The survey also collected household demographic data such as annual household income, available vehicles, the number of people living in their household, and the number of workers in their household above the age of 16. It also asked if they would use proposed bus-only lanes or train service along the corridor, if available. ## Data Collection Agency The survey was conducted by and for the Florida Department of Transportation. ## Survey Methodology The survey was conducted via mail and online in March 2003. ## Survey Records, Data, and Documentation Survey records include 7,917 participants. Origin and destination locations include street addresses, nearest intersections or landmarks, city, state, zip code, and latitude/longitude.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2003 Interstate 595 Vehicle Trip-Length Study

# 2003 Interstate 595 Vehicle Trip-Length Study The 2003 Vehicle Trip-Length Study focused on Interstate 595 between Davie Road and University Drive in Florida. Survey participants answered questions about their trip's origin and destination—including the type of location such as work, home, store, etc.—the on- and off-ramps used, how many people were in the car, and the type of vehicle. The survey also collected household demographic data such as annual household income, available vehicles, the number of people living in their household, and the number of workers in their household above the age of 16. It also asked if they would use proposed bus-only lanes or train service along the corridor, if available. ## Data Collection Agency The survey was conducted by and for the Florida Department of Transportation. ## Survey Methodology The survey was conducted via mail and online in March 2003. ## Survey Records, Data, and Documentation Survey records include 7,917 participants. Origin and destination locations include street addresses, nearest intersections or landmarks, city, state, zip code, and latitude/longitude.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2003 Interstate 595 Vehicle Trip-Length Study

# 2003 Interstate 595 Vehicle Trip-Length Study The 2003 Vehicle Trip-Length Study focused on Interstate 595 between Davie Road and University Drive in Florida. Survey participants answered questions about their trip's origin and destination—including the type of location such as work, home, store, etc.—the on- and off-ramps used, how many people were in the car, and the type of vehicle. The survey also collected household demographic data such as annual household income, available vehicles, the number of people living in their household, and the number of workers in their household above the age of 16. It also asked if they would use proposed bus-only lanes or train service along the corridor, if available. ## Data Collection Agency The survey was conducted by and for the Florida Department of Transportation. ## Survey Methodology The survey was conducted via mail and online in March 2003. ## Survey Records, Data, and Documentation Survey records include 7,917 participants. Origin and destination locations include street addresses, nearest intersections or landmarks, city, state, zip code, and latitude/longitude.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗