Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data requirements”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

When physics-informed data analytics outperforms black-box machine learning: A case study in thickness control for additive manufacturing

Aerosol jet printing (AJP) has emerged as a promising noncontact additive manufacturing method for high-resolution printing for a wide range of material systems. A key challenge limiting the broader adoption of AJP in the material science community is the lack of methods to precisely control thickness. Herein, we develop a model-based design of experiment (MBDoE) framework that integrates physics-informed models, nonlinear regression, and information criteria to postulate, select and calibrate the best model to describe and optimize the AJP manufacturing process. Starting with already available data from system commissioning (e.g., prior single variable sensitivity analysis), four candidate physics-informed models are postulated and trained. MBDoE identifies a single additional optimal experiment to validate these predictive models with quantified uncertainties, which are then used to determine the best experimental conditions to control printed film thickness. As a comparative benchmark, the analysis is repeated using the same dataset with nonparametric Gaussian process regression (GPR) model that does not incorporate physical information. Using MBDoE principles, we find that only five experiments are necessary to calibrate the nonlinear physics-informed parametric model, and with said limited data, this model outperforms the black-box machine learning GPR model. This key result underscores an emerging trend in the data science community: incorporating physical information into predictive models often drastically reduces the data requirements. Leveraging MBDoE further increased the data efficiency. By design, the proposed data science framework is general in nature and can be easily extended to other experimental and additive manufacturing systems beyond AJP.

Aerosol jet printing↗

Requirements for Cataloging Hanford Geophysical Datasets

Environmental management activities at the Hanford Site produce extensive data about site conditions, contaminants, cleanup, and more. Managing and archiving that data requires a high degree of collaboration among site contractors and a high level of awareness by project managers and staff. Part of that effort is developing a Hanford Environmental Information and Data Index (HEIDI) to organize the data and maximize its value by making it findable and available for reuse. The objective is to catalog the disparate data sets collected to address the evolving needs of planning, executing, and documenting cleanup over several decades up to the present day, including links to active data sources when available. A properly implemented data catalog makes finding environmental datasets related to an area or theme a routine, reliable process, without requiring the searcher to have special knowledge that a data set exists and where it may be stored. In this project, a working group, including the U.S. Department of Energy, the Hanford Site contractors, and Pacific Northwest National Laboratory staff, identified needs and requirements for handling complex site data. Geophysical data was chosen as a test case because it can be large and complex and often involves multiple processing steps to extract the information incorporated into deliverables. The ability to document those steps was one of the requirements identified for the catalog. In addition to developing requirements, other activities included selecting a metadata schema and initial testing with the objective of determining whether the workflow and capabilities of selected data catalog software platforms were sufficient to implement and impose the identified requirements. This initial testing involved running the default catalog instance using the software platform of interest and altering the configuration to achieve each requirement, if possible. Where configuration alone was insufficient, the possibility of modifying the software by changing the code was examined, but not implemented. A follow-on task is planned to reprogram the code as necessary to implement requirements in a prototype catalog.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Data Science Infrastructure SOFTWARE

LANL science workflows generate complex data sets and ensembles of data requiring significant compute and storage resources. The Data Science Infrastructure (DSI) project focuses on data-driven approaches to make data more readily available to LANL projects. DSI workflows leverage metadata stored in data-agnostic databases, supported by an abstraction layer API to simplify searching and accessing data across simulation runs, experimental runs, filesystems and environments. The abstraction layer API allows the user to query a range of data types: raw output, processed data, configuration data, machine learning models, performance data, etc. In addition to the abstraction backend API, the DSI project is developing client-driven query APIs and UIs to support specific user workflows.

Turton, Terece↗

National Ignition Facility Opacity Time Resolved Spectrometer Systems Engineering Final Project

The National Ignition Facility (NIF) is the world’s largest and most energetic laser facility. The NIF system is designed to produce high energy density (temperature and pressure) conditions through the application of its 192 laser beams. One of the users of NIF is the opacity platform developed to study the opacities at temperatures and densities relevant to the solar interior and stellar evolution. The platform was developed to study iron (Fe) opacity at temperatures relevant to the solar interior. The opacity campaign uses spectrometers to gather data. Spectrometers utilize crystals to produce x-ray spectra that are recorded on time-integrated and time-resolved detectors. The opacity spectrometer (OpSpec) currently fielded and in use at NIF uses a time integrated film channel to collect data. The opacity spectrometer time resolved (OpSpecTR) will utilize novel hCMOS detectors to capture time resolved images of spectra of interest. The key stakeholders identified for OpSpecTR included the physicists responsible for OpSpec and OpSpecTR, the Target Area Science and Engineering (TASE) department at NIF, the NIF and Photon Science (NIF & PS) Opacity program, the Nevada National Security Site (NNSS) Physics and Engineering program, the Sandia hCMOS manufacturing and testing program, and the Los Alamos National Laboratory (LANL) program sponsor. The Target and Experimental Operations (TEXOPS) was identified as a key stakeholder because the group includes the individuals that will physically interact with the OpSpecTR system as it participates in NIF experiments. The opacity platform collects data in a unique orientation relative to the existing diagnostics fielded at NIF. The existing infrastructure at NIF uses a diagnostic manipulator (DIM) to insert the diagnostic near the target chamber center to collect data during a NIF shot. Existing diagnostics collect data through the center line of the DIM axis and collect relevant data perpendicular to this axis. The opacity platform requires crystals mounted in a specific orientation which requires data collection parallel to the DIM axis. This deviation from standard NIF practices was a key factor in developing requirements.

42 ENGINEERING↗

Decay Curve Correction Analysis Report

The decay curve analysis that is done on the short-lived radionuclide gas samples is used to differentiate between gaseous radionuclides that have the same characteristic gamma decay energy, 511 kiloelectron-volts (keV). A sample of stack gas is isolated and the total counts in the 511 keV peak are counted repeatedly in 10-second intervals to evaluate the decay rate of the sample over time. Analysis of this decay data required a series of steps. First, a raw data report is generated by the gamma acquisition system, based on an analysis template within the acquisition software. The data report file was then loaded into Microsoft Word, and a macro was used to perform minor formatting (remove colons and insert tabs between data columns) to allow analysis within Excel. The file is then saved as a text file at this point. The text file is then uploaded into Excel and a series of macros are used to add labels, calculate radioactive decay constants, and analyze the gamma decay data using linear regression techniques. The analysis template has been used since 1998 for stack 53000303 (TA-53, building 0003, exhaust stack 03) and 2000 for stack 53000702. The overall process, including the gamma report format and the macros used in Word and Excel for processing the report, had remained unchanged until 2015. In October of 2015, staff made a change to the report template in the gamma acquisition software which resulted in an error in the calculations later performed by the Excel macro. This error was not caught until a more in-depth review of the analysis took place regarding 2020 data. This report covers a much more complete review of the issue that occurred regarding the decay curve analysis, a review of the calculations completed to correct the issue, a review of the updated decay curve analysis process, and recommendations for moving forward.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Advancing research on compound weather and climate events via large ensemble model simulations

Societally relevant weather impacts typically result from compound events, which are rare combinations of weather and climate drivers. Focussing on four event types arising from different combinations of climate variables across space and time, here we illustrate that robust analyses of compound events — such as frequency and uncertainty analysis under present-day and future conditions, event attribution to climate change, and exploration of low-probability-high-impact events — require data with very large sample size. In particular, the required sample is much larger than that needed for analyses of univariate extremes. We demonstrate that Single Model Initial-condition Large Ensemble (SMILE) simulations from multiple climate models, which provide hundreds to thousands of years of weather conditions, are crucial for advancing our assessments of compound events and constructing robust model projections. Combining SMILEs with an improved physical understanding of compound events will ultimately provide practitioners and stakeholders with the best available information on climate risks.

54 ENVIRONMENTAL SCIENCES↗

Pulse: An Outlier Sensitive Downsampling Algorithm For Timeseries Data

Pulse is a downsampling algorithm for timeseries data. Frequently datasets become so large that visualization tools and web browsers cannot effectively render graphics due to memory constraints. Downsampling algorithms are commonly applied to minimize the quantity of data required to visualize important features or trends in the data, but some datasets are composed by distinct enough features and trends that most existing downsampling algorithms fail to preserve them. Pule was developed to downsample timeseries data for galvanostatic stack test data at the Idaho National Laboratory. These datasets were composed by approximately 4 million records, most of them being extremely uniform. However, during relatively brief time periods when the stack test changes state, for example when the test article is powered on, or a load is added, the data produce sparse asymptotes. No existing downsampling algorithm was capable of preserving the sparse asymptotes in electrolysis stack test data. Instead, we develop a downsampling algorithm that preserves important outliers in data, and otherwise aggressively downsamples uniform data. The algorithm has applications in other domains like seismology, in the measurement of earthquakes, or astronomy, in the measurement of quasars or transit photometry.

Woodruff, Nathan [Idaho National Laboratory (INL),↗

Lighting System Control Data to Improve Design and Operation: Tunable Lighting System Data from NICU Patient Rooms

The advancement of LED and controls technology, computing capacity, and software provides new opportunities for researchers and designers to work together to further optimize spaces for occupant benefit. Here, lighting system control data from five neonatal intensive care unit patient rooms was collected over a 25-week monitoring period and analyzed to better understand occupant response to a tunable lighting system with automatic transitions throughout the day. Lighting systems are very rarely refined after installation based on actual use. Objective data detailing how the lighting system is used by the actual occupants highlights the opportunities for optimization after installation and provides insight for improving the next design. As use of the data becomes more commonplace, it can be leveraged for design recommendations. The collection of the data required no additional cost beyond the time for examining the data. The analysis revealed several clear opportunities for improvement, including adjustments to the default control setting at night, re-labeling of the control stations, and adjustments to the nighttime fade rate. The patient room occupants were active users of the different zones, dimming options, and manual overrides made available by the lighting system.

60 APPLIED LIFE SCIENCES↗

Marine Boundary Layer Decoupling and the Stable Isotopic Composition of Water Vapor

Decoupling of the subcloud layer in stratocumulus-topped marine boundary layers (STMBL) influences low-cloud cover by limiting the supply of water vapor from the surface. However, the relative importance of mixing between surface fluxes and other reservoirs of water vapor as a function of the degree of decoupling is poorly understood. Water vapor transport within the STMBL and its response to decoupling is explored using surface measurements of water vapor isotopic composition that were obtained from Graciosa Island, Azores, during the summer and fall of 2018. The data show an inverse relationship between the decoupling metric, Δq, and the degree of drying from an isotopically depleted water vapor source. The isotopic data require some degree of drying with an isotopically depleted source for coupled conditions, and likely require a small amount of mixing even for strongly decoupled conditions. The data are consistent with mixing between purely local reservoirs of water vapor, subjected to a small amount of condensation and fractionation, and do not require the invocation of large-scale transport of water vapor aloft or additional cloud-formation effects aloft.

54 ENVIRONMENTAL SCIENCES↗

OTERR Theory Manual

OTERR is a python code designed to couple an external transport/depletion capability with an internal genetic algorithm for fuel reloading optimization. OTERR stores the state information required for creating neutronics code input, output from ARC codes, and data required for running optimization in HDF5 files.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Late Breaking Results: COPPER: Computation Obfuscation by Producing Permutations for Encoding Randomly

Deployed embedded devices face security risks due to increased ease of physical access to the devices by unauthorized users. Capable adversaries can intercept a device to recover the data in memory, including results of performed sensitive computations. Device owners require data confidentiality on their physically insecure devices. To satisfy this goal we implement a novel method, COPPER (Computation Obfuscation by Producing Permutations for Encoding Randomly), to create data which never exists on the device digitally in plaintext format and which is subsequently used for computation. In this paper we utilize COPPER to calculate a moving average computation on encoded data.

embedded systems↗

Out-of-distribution generalization for learning quantum dynamics

Abstract Generalization bounds are a critical tool to assess the training data requirements of Quantum Machine Learning (QML). Recent work has established guarantees for in-distribution generalization of quantum neural networks (QNNs), where training and testing data are drawn from the same data distribution. However, there are currently no results on out-of-distribution generalization in QML, where we require a trained model to perform well even on data drawn from a different distribution to the training distribution. Here, we prove out-of-distribution generalization for the task of learning an unknown unitary. In particular, we show that one can learn the action of a unitary on entangled states having trained only product states. Since product states can be prepared using only single-qubit gates, this advances the prospects of learning quantum dynamics on near term quantum hardware, and further opens up new methods for both the classical and quantum compilation of quantum circuits.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Out-of-distribution generalization for learning quantum dynamics

Generalization bounds are a critical tool to assess the training data requirements of Quantum Machine Learning (QML). Recent work has established guarantees for in-distribution generalization of quantum neural networks (QNNs), where training and testing data are drawn from the same data distribution. However, there are currently no results on out-of-distribution generalization in QML, where we require a trained model to perform well even on data drawn from a different distribution to the training distribution. Here, we prove out-of-distribution generalization for the task of learning an unknown unitary. In particular, we show that one can learn the action of a unitary on entangled states having trained only product states. Since product states can be prepared using only single-qubit gates, this advances the prospects of learning quantum dynamics on near term quantum hardware, and further opens up new methods for both the classical and quantum compilation of quantum circuits.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A portable application framework for energy management and information systems (EMIS) solutions using Brick semantic schema

This paper introduces a portable framework for developing, scaling and maintaining energy management and information systems (EMIS) applications using an ontology-based approach. Key contributions include an interoperable layer based on Brick schema, the formalization of application constraints pertaining metadata and data requirements, and a field demonstration. The framework allows for querying metadata models, fetching data, preprocessing, and analyzing data, thereby offering a modular and flexible workflow for application development. Its effectiveness is demonstrated through a case study involving the development and implementation of a data-driven anomaly detection tool for the photovoltaic systems installed at the Politecnico di Torino, Italy. During eight months of testing, the framework was used to tackle practical challenges including: (i) developing a machine learning-based anomaly detection pipeline, (ii) replacing data-driven models during operation, (iii) optimizing model deployment and retraining, (iv) handling critical changes in variable naming conventions and sensor availability (v) extending the pipeline from one system to additional ones.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Electrical Load Forecasting Over Multihop Smart Metering Networks With Federated Learning

Electric load forecasting is essential for power management and stability in smart grids. This is mainly achieved via advanced metering infrastructure, where smart meters (SMs) record household energy data. Traditional machine learning (ML) methods are often employed for load forecasting, but require data sharing, which raises data privacy concerns. Federated learning (FL) can address this issue by running distributed ML models at local SMs without data exchange. However, current FL-based approaches struggle to achieve efficient load forecasting due to imbalanced data distribution across heterogeneous SMs. Here, this article presents a novel personalized FL (PFL) method for high-quality load forecasting in metering networks. A meta-learning-based strategy is developed to address data heterogeneity at local SMs in the collaborative training of local load forecasting models. Moreover, to minimize the load forecasting delays in our PFL model, we study a new latency optimization problem based on optimal resource allocation at SMs. A theoretical convergence analysis is also conducted to provide insights into FL design for federated load forecasting. Extensive simulations from real-world datasets show that our method outperforms existing approaches regarding better load forecasting and reduced operational latency costs.

Rahman, Ratun [Univ. of Alabama, Huntsville, AL (U↗

MRCI Subtask 2.3: Developing Industrial Partnerships and Regional Technical Collaboration Final Technical Summary Report

Under the objective of regional data collection and helping accelerate deployment, MRCI collaborated with industrial stakeholders in their project planning, characterization, and analysis. Some examples of these collaborations are given below. The data and information shared by the industrial collaborations added to the regional CCS framework development and were incorporated into the overall datasets, while addressing any proprietary data requirements. Three examples of collaborative partnerships with industry that have provided geologic characterization data relevant and beneficial to the MRCI program are discussed below, including: the UIC Class II Injection Facility in Eastern Ohio, the Core Energy CO2-EOR (enhanced oil recovery) operation in Otsego County Michigan, and the Marquis ethanol plant in Hennepin Illinois.

CCS,CCUS,MRCI,Midwest USA,Technical Challenges,inj↗

FY24 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

Algorithms for Machine Learning (ML) and data analysis for the 3013 Surveillance Program have been developed in an ongoing collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). The objective of the algorithms is to automate the identification of corrosion and crack formation in the Inner Container Closure Weld Region (ICCWR) of the canister system used to store Pu-bearing material. Data for corrosion and cracking is collected from large binary files generated by a Laser Confocal Microscope (LCM), the Wide Area 3D Measurement System (WAMS), or,in a recent proposal, by a Scanning Electron Microscope (SEM). The ML software uses the physical attributes in the data files (e.g., one or more of: height, color, and 16-bit grayscale values as functions of position in a plane projection) to detect signs of surface corrosion and cracking after being trained on similar data, with the features to be detected. Although the initial scope included screening for broader indicators of corrosion, e.g., pitting, identification of potential cracks was prioritized for the past several years at the request of program leadership. Labeled training data is essential to developing the ML algorithm, and enhancements to data labeling capability have been developed to address this essential precursor to application of ML routines. Efficient labeling is particularly important in view of the large volume of data required to train ML algorithms and the relative rarity of cracks in the ICCWR data set. The updated program will read binary data from either LCM, WAMS or SEM files, interrogate data attributes, facilitate user labeling of data for training ML algorithms, execute ML algorithms, output parameters from trained ML algorithms, report ML model accuracy with respect to labeled data, and generate graphical representations for various analyses. In FY24, hourglass neural networks (HNNs) that were initiated in FY22 were further developed and tested using available LCM data, and their performance was tested against that of the alternative U-Net Neural Network algorithm structure. HNNs along with previously developed Convolutional Neural Networks (CNNs) and Deep Neural Networks (DNNs) comprise a suite of ML tools for identification of cracks in the ICCWR

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Carbon Storage Technical Viability Approach and National Data Assessment

Carbon Storage Technical Viability Approach is a comprehensive evaluation of potential geologic carbon storage sites that includes classic subsurface resource assessment components such as reservoir suitability in addition to hazard, environmental, social, regulatory, and jurisdiction components that may enhance or hinder successful carbon storage. Data for assessments of technically viable carbon storage span many categories, types; are disparate and numerous. Variables and types of data required to support CS assessments were previously not clearly defined and documented. Contextualizing available data indicates their utility, potential uncertainty, and gaps. This presentation details the Carbon Storage Technical Viability Approach, the associated matrix and database, and initial results from a portion of the national data availability assessment that contextualizes the data available to support carbon storage projects. Presented at the American Association of Petroleum Geologists' Carbon Capture Utilization and Storage (SPE-AAPG-SEG CCUS 2024) conference in Houston, TX, March 11-13, 2024.

Mark-Moser, Mackenzie K.↗