Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Utah FORGE InSAR Data from 2021

Interferometric Synthetic Aperture Radar data from the TerraSAR-X and the TanDEM-X satellite missions operated by the German Space Agency (DLR). Interferometric pairs (interferograms) were created using generic mapping tool GMT-SAR processing software (see link in Resources). Data from January through November 2021.

15 GEOTHERMAL ENERGY↗

Utah FORGE InSAR Data from 2022

Interferometric Synthetic Aperture Radar data from the TerraSAR-X and the TanDEM-X satellite missions operated by the German Space Agency (DLR). Interferometric pairs (interferograms) were created using generic mapping tool GMT-SAR processing software (see link in Resources). Data from January through June 2022.

15 GEOTHERMAL ENERGY↗

Automating Traffic Microsimulation from SYNCHRO UTDF to SUMO

Modern transportation research relies on seamlessly integrating traffic signal data with robust network representation and simulation tools. This study presents utdf2gmns, an open-source Python tool that automates conversion of the Universal Traffic Data Format, including network representation, signalized intersections, and turning volumes into the General Modeling Network Specification (GMNS) Standard. The resulting GMNS-compliant network can be converted for microsimulation in SUMO. By automatically extracting intersection control parameters and aligning them with GMNS conventions, utdf2gmns minimizes manual preprocessing and data loss. utdf2gmns also integrates with the Sigma-X engine to extract and visualize key traffic control metrics, such as phasing diagrams, turning volumes, volume-tocapacity ratios, and control delays. This streamlined workflow enables efficient scenario testing, accurate model building, and consistent data management. Validated through case studies, utdf2gmns reliably models complex urban corridors, promoting reproducibility and standardization. Documentation is available on GitHub and PyPI, supporting easy integration and community engagement.

Luo, Roy [ORNL] (ORCID:0009000312909983)↗

Data and scripts associated with the manuscript "Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning"

This package contains the data and scripts used in "Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning" (Jiang et al., 2022). The data.zip file contains the flux tower and automated chamber observations used for developing the deep learning model for modeling soil respiration. The scripts.zip file contains the Jupyter notebooks and python scripts for preprocessing the data, training the deep learning models, and postprocessing the results. The src.zip contains the source code for training the deep learning model, performing mutual information analysis, and plotting functions. The trained_models.zip contains multiple folders used for hosting the trained deep-learning models and the associated soil respiration predictions. The whole process is performed using python. We include the REAMD.md to document the python package requirements.Soil respiration in dryland ecosystems is challenging to model due to its complex interactions with environmental drivers. Knowledge-guided deep learning provides a much more effective means of accurately representing these complex interactions than traditional Q10-based models. Mutual information analysis revealed that future soil temperature shares more information with soil respiration than past soil temperature, consistent with their clockwise diel hysteresis. We explicitly encoded diel hysteresis, soil drying, and soil rewetting effects on soil respiration dynamics in a newly designed Long Short Term Memory (LSTM) model. The model takes both past and future environmental drivers as inputs to predict soil respiration. The new LSTM model substantially outperformed three Q10-based models and the Community Land Model when reproducing the observed soil respiration dynamics in a semi-arid ecosystem. The new LSTM model clearly demonstrated its superiority for temporally extrapolating soil respiration dynamics, such that the resulting correlation with observational data is up to 0.7 while the correlations of both Q10-based models and the Community Land Model (CLM) are less than 0.4. Our results underscore the high potential for knowledge-guided deep learning to replace Q10-based soil respiration modules in Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Model Data Archive for Manuscript Titled "Evaluation of a Coupled Surface–Subsurface Hydrologic Model Using Dense Water‑Level Sensors in a Mixed Urban–Rural Watershed"

This archive provides scripts, input files, and datasets used for the implementation and evaluation of a fully coupled surface–subsurface hydrologic model in the Neches River Basin, southeast Texas. The study uses the Advanced Terrestrial Simulator (ATS) to simulate coupled surface–subsurface hydrologic processes over a mixed urban–rural watershed and evaluates model performance using a dense network of 136 in situ water-level sensors, nine U.S. Geological Survey (USGS) stream gauges, and SSEBop-derived evapotranspiration estimates during the period October 2014–June 2024. The workflow is implemented primarily in Python 3 using the Watershed Workflow package. The Jupyter notebooks can be executed using open-source software such as Anaconda JupyterLab or Visual Studio Code. Other data files include TXT, CSV, XML, SHP, TIF, NetCDF, HDF5, and ExodusII files, which can be processed using the provided Python scripts. ATS input files are provided in XML format and can be edited using any commonly used text editor. This archive contains: *Scripts and input files used to generate the ATS model setup, including watershed discretization, mesh generation, parameter mapping, and model configuration. *Jupyter notebooks used for preprocessing observational data, evaluating streamflow, water levels, and evapotranspiration, computing performance metrics, and generating the figures presented in the manuscript. *ATS simulation outputs and processed observational datasets, including OneRain and DD6 water-level sensors, USGS streamflow observations, GIS data, and supporting spatial datasets used throughout the study.

Dense water-level sensor network↗

Forecasting of in situ electron energy loss spectroscopy

Abstract Forecasting models are a central part of many control systems, where high-consequence decisions must be made on long latency control variables. These models are particularly relevant for emerging artificial intelligence (AI)-guided instrumentation, in which prescriptive knowledge is needed to guide autonomous decision-making. Here we describe the implementation of a long short-term memory model (LSTM) for forecasting in situ electron energy loss spectroscopy (EELS) data, one of the richest analytical probes of materials and chemical systems. We describe key considerations for data collection, preprocessing, training, validation, and benchmarking, showing how this approach can yield powerful predictive insight into order-disorder phase transitions. Finally, we comment on how such a model may integrate with emerging AI-guided instrumentation for powerful high-speed experimentation.

36 MATERIALS SCIENCE↗

Why is the winner the best?

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multi- center study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), im- age preprocessing (97%), data curation (79%), and post- processing (66%). The “typical” lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work.

Eisenmann, Matthias↗

Offshore Wind ENergy Simulation Toolkit (OWENS)

SAND2021-2751 O The Offshore Wind ENergy Simulation Toolkit (OWENS) is a collection of aerodynamic, structural, hydrodynamic, drivetrain, controls, composite structure and mesh preprocessing, and data postprocessing. OWENS is primarily an ontology, or glue code, pulling together many open-source and Sandia-developed libraries to model the aero-servo-hydro-elastic physics of wind and marine energy turbines. The toolkit’s intended use is for arbitrary aeroelastic rotor configurations analysis including vertical-axis wind turbines, horizontal-axis wind turbines, and analogous marine energy applications for fixed-bottom and floating configurations. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525

Owens, Brian↗

HAIMOS Ensemble Forecasts for Intra-day and Day- Ahead GHI, DNI and Ramps

The objective of this research is to develop a hybrid physics-based/data-driven forecast model to improve direct normal and global horizontal irradiance (DNI and GHI) prediction for horizons ranging from 1 to 72 hours. Project objectives also address key gaps in state-of-the-art solar forecasting: accurate probabilistic solar forecasts and the forecasting of large irradiance ramps (ramp onset and magnitude). The proposed model ensembles Numerical Weather Prediction (NWP) forecasts, determinist physics-based algorithms, and new-generation cloud cover products (high-resolution rapid refresh satellite images and Large Eddy Simulations). The result is the Hybrid Adaptive Input Model Objective Selection (HAIMOS) ensemble model. HAIMOS blends state of the art machine learning methodologies with physics-based models for cloud cover and cloud optical depth forecasts. The technical activities followed a two-pronged strategy. First, the preprocessing of data, the selection of inputs to the nonlinear approximators, the type of approximator and objective functions, and post-processing ensembling techniques included in HAIMOS were all optimized adaptively to find the best model for a specific goal (reduce DNI/GHI forecast error, improve the prediction of ramp onset, etc.). Second, a large effort was put in improving cloud identification and the forecast of cloud cover and cloud optical depth. To this end, new-generation cloud parametrization products were developed in this work. These include improved algorithms to assist in cloud identification, cloud classification and cloud parametrization from satellite images – three key factors in the accuracy of 1 to 6-hours irradiance forecasts and prediction of ramp onset. Furthermore, we also included cloud information extracted from high resolution rapid refresh satellite images (GOES-16) and Large Eddy Simulations (LES). LES was used to model the atmosphere in detail over locations of interest and produce cloud optical depth forecasts. Once these data streams were validated, they were used as input data to the HAIMOS forecast. The model was developed using data from several climatologically distinct locations with potential for high solar penetration. In the last year of the project, we conducted a validation campaign according to the guidelines stipulated by the Topic Area 1 project as described in the FOA. This effort brings, for the first time, proven machine-learning methodologies for generating state-of-the-art solar forecasts interweaved with detailed physics-based models for cloud detection, and cloud optical depth forecasts. HAIMOS will generate accurate irradiance probabilistic forecast to assist in reducing solar generation prediction error. Globally optimized solar forecast models are more likely to impact solar energy stakeholders. The goal of this project was to increase the state-of-the-art forecast skill from their present values of 10 to 35%. At the end of the project, we achieved between 30% and 50% forecast skill across a wide range of horizons for both GHI and DNI.

14 SOLAR ENERGY↗

Mechanical separations of corn stover anatomical fractions in an integrated feedstock preprocessing system: An experimental and data-driven modeling study

High variabilities of material attributes in lignocellulosic biomass present risks for biofuel and biochemical productions and must be mitigated via preprocessing. Since almost no mechanical device is originally designed for processing biomass, how to operate existing apparatuses with efficient performance has not been investigated extensively. This work presents a study on an integrated screening and air classification to separate cobs and stalks from husks and leaves in corn stover. Prototype machine learning models were developed to assess the feasibility of predicting the process outcome based on the measurable parameters. The models trained upon limited experimental data rendered decent predictive accuracy of yield and purity. The experimental data and modeling results collectively suggest decreasing throughput leads to a higher purity. To the contrary, if throughput increases, a lower purity is likely. A possible trade-off between yield and purity of the separated streams indicates the need for optimal combinations of feedstock size, moisture, and throughput to achieve optimized separations. The results of this study also suggest the need to further improve model predictability by developing more accurate formulations for physics governing the integrated unit operations. To accomplish this, additional experimental data needs to be generated for model training.

09 - BIOMASS FUELS↗

Digital Analytics, Causal Knowledge Acquisition and Reasoning for Technical Language Processing

Complex engineering systems such as nuclear power plants (NPPs) generate and collect large amounts of equipment reliability (ER) data elements that contain information on the status of components, assets, and systems. Some of this information is textual in form and can be found in documents such as incident reports (IRs) and work orders (WOs). Analyses of textual data in current NPPs-using natural language processing (NLP) methods-have been expanded over the last decade, and it is only recently that the true potential of such analyses has emerged. So far, applications of NLP methods have mostly been limited to classification and prediction, the goal being to identify the nature of the textual element (e.g., safety or non-safety related). Here, we target a more complex problem: automatically extracting knowledge from a textual element in order to assist system engineers in conducting system health assessments. Knowledge extraction is a very broad concept, and its definition may vary depending on the application context. Our methods are a blend of both rule-based and machine learning (ML) algorithms. For our purposes, knowledge extraction means identifying the systems or assets mentioned in a given textual element, as well as the type of event described (e.g., component failure or maintenance activity). In addition, we want to capture details such as measured quantities and the temporal/cause-effect relations between events. In this tool, we also demonstrate how textual data elements are preprocessed in order to handle typos, acronyms, and abbreviations. One main feature of these methods is that they are not based solely on data, but are in fact model-based. In other words, they also rely on MBSE models that are designed to capture-from a functional point of view-the architecture of the systems/assets under consideration. The main purpose of such models is to digitally emulate system engineers' knowledge of system and asset architecture and to identify dependencies among systems, assets, and components. Provided these models, analyses of textual and numeric ER data can be performed by first identifying the OPM model elements to which the ER data elements are referring. The relationships between ER data elements are then identified by checking for any temporal or logical dependencies.

Mandelli, Diego [Idaho National Laboratory (INL), ↗

Data-driven method for electric vehicle charging demand analysis: Case study in Virginia

Electric vehicle (EV) adoption in the U.S. will be accelerated by the historic $7.5 billion public investments in EV charging infrastructure. Careful analysis of EV charging demands plays a vital role in understanding the energy requirements, power grid impact, and smart charging management opportunities of EVs. To this end, this paper develops a data-driven trip-chaining-based modeling framework including five steps: Trip data acquisition and preprocessing, EV adoption modeling, travel itinerary synthesis, EV charging demand simulation and EV load profile generation. The developed analysis framework was demonstrated using real-world data for one region in Virginia, U.S. The results show that the proposed modeling framework can work effectively. For the study region in 2040, the predicted number of plug-in EVs is 470,114, resulting in a weekly charging demand of 38,078,127 kWh (55% home, 9% work, and 36% public) in September and 45,920,358 kWh (61% home, 9% work, and 30% public) in February.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Sensor Anomaly Detection for Nuclear Reactor Systems Utilizing Linear Regression and K-Means Unsupervised Machine Learning

Nuclear reactors and related systems are becoming increasingly complex due to advancing technologies in next-generation power reactors. This increased complexity necessitates enhanced automation and data management capabilities. To successfully realize autonomous systems, methods must be developed to handle vast volumes of data and effectively distinguish anomalous data from noise and expected data. While impressive models utilizing digital twins and similar approaches are under development, here we propose a simplified model for analyzing fundamental methods and techniques. Initially, we created a general dataset by using initial data from PCTRAN in order to represent ideal steady-state conditions. We then inserted anomalies based on prevalent sensor anomaly types (e.g., point anomalies, linear drift, and downward deviations), along with unusual anomalies such as exponential drift and upward deviations. To detect anomalies, we developed a program that employs data partitioning and linear regression to preprocess and filter the anomalous data. A K-Means machine learning (ML) method was then applied to separate and count the data within the anomalous partition. The results from all datasets—apart from exponential growth—demonstrated positive outcomes, with each returning multiple instances of greaterthan-95% accuracy. We conducted further investigations using Idaho National Laboratory’s RAVEN software to perform a sensitivity analysis on the input variables (R 2 Tolerance, Slope Tolerance, and Window Size) and found that the output variables (Accuracy and Time) were most sensitive to the Window Size. Despite the promising results published, further development is required to effectively apply these methods to nuclear systems. Nevertheless, the strengths of this approach are evident and hold promise for future applications in the field.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Neural MUSE Analysis

Researchers at Oak Ridge National Laboratory (ORNL) created data as part of the MUSE (Multi-Agency Urban Search Experiment Detector and Algorithm Test Bed) project simulating illicit nuclear materials located in various buildings along a road. In the simulation, a truck containing a radiation detector drives down the road gathering listmode data (counting the and energy of incident gamma radiation). Building materials, source shielding, driving speed, truck direction, truck location on the road, source type, and source placement are all varied between runs of the data set. This data was created using deterministic neutron transport and Monte Carlo methods through a combination of SCALE, MAVRIC, MCNP, and GADRAS. As part of a follow-on NA-22 project, two Kaggle competitions were created to determine the best algorithms for finding and identifying gamma sources in this simulated urban environment. The winning algorithm was neural network-based and had a test accuracy of 76.4% accuracy for source identification. This work seeks to build upon this work and improve the results through the application of novel machine learning techniques. As a first step, the data was classified by a simple Convolutional Neural Network (CNN) To accomplish this, the data was first preprocessed into “waterfall plots.” These plots are composed of energy vs count plots that are stacked vertically to show progression in time. The horizontal axis indicating the particle energy incorporated user defined bin spacing with options for in linear-, logarithmic-, square root-, and user-spaced bins. The z or color dimension showed the number of counts corresponding the energy-time combination. This data was then used to generate more data, by generating a local estimate of the mean of the distribution for a bin and then randomly re-sampling that bin from a Poisson distribution. Once all of this data was generated, it was fed into a well-known CNN architecture, ResNet50. The output layer of this model was removed and replaced with layers corresponding to the shape desired isotope outputs. The provided training data was used to train the classifier and the remaining testing data was used to evaluate the model. Results are soon to be forthcoming.

61 RADIATION PROTECTION AND DOSIMETRY↗

The ASHRAE Great Energy Predictor III competition: Overview and results

In late 2019, ASHRAE hosted the Great Energy Predictor III (GEPIII) machine learning competition on the Kaggle platform. This launch marked the third energy prediction competition from ASHRAE and the first since the mid-1990s. In this updated version, the competitors were provided with over 20 million points of training data from 2,380 energy meters collected for 1,448 buildings from 16 sources. This competition’s overall objective was to find the most accurate modeling solutions for the prediction of over 41 million private and public test data points. Furthermore, the competition had 4,370 participants, split across 3,614 teams from 94 countries who submitted 39,403 predictions. In addition to the top five winning workflows, the competitors publicly shared 415 reproducible online machine learning workflow examples (notebooks), including over 40 additional, full solutions. This paper gives a high-level overview of the competition preparation and dataset, competitors and their discussions, machine learning workflows and models generated, winners and their submissions, discussion of lessons learned, and competition outputs and next steps. The most popular and accurate machine learning workflows used large ensembles of mostly gradient boosting tree models, such as LightGBM. Similar to the first predictor competition, preprocessing of the data sets emerged as a key differentiator.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A guideline to document occupant behavior models for advanced building controls

The availability of computational power, and a wealth of data from sensors have boosted the development of model-based predictive control for smart and effective control of advanced buildings in the last decade. More recently occupant-behavior models have been developed for including people in the building control loops. However, while important objectives of scientific research are reproducibility and replicability of results, not all information is available from published documents. Therefore, the aim of this paper is to propose a guideline for a thorough and standardized occupant-behavior model documentation. For that purpose, the literature screening for the existing occupant behavior models in building control was conducted, and the occupant behavior modeling processes were studied to extract practices and gaps for each of the following phases: problem statement, data collection, and preprocessing, model development, model evaluation, and model implementation. Here, the literature screening pointed out that the current state-of-the-art on model documentation shows little unification, which poses a particular burden for the model application and replication in field studies. In addition to the standardized model documentation, this work presented a model-evaluation schema that enabled benchmarking of different models in field settings as well as the recommendations on how OB models are integrated with the building system.

Building control↗

Real-Time Optimization Workflow Status Update

Economically optimal and safe operation of integrated energy systems (IES) requires optimization at many different time scales. A real-time optimization (RTO) workflow will attempt to maximize revenue and minimize operational costs on a time scale of minutes to hours. Such a workflow requires the use of a digital twin (DT), which is a virtual representation of a physical system. The DT is updated using real-time data from the physical system, and serves as a model in an optimization framework. The optimization results are then sent back to the physical system to complete the loop. This report details the progress made in developing building blocks for a DT/RTO framework. The Risk Analysis Virtual Environment (RAVEN) platform within the Framework for Optimization of Resources and Economics (FORCE) tool suite can perform many of the tasks required for building a DT and performing RTO. The first item of this report details RAVEN enhancements that enable RAVEN workflows to be run in various environments. Data communication between the physical system and its DT is essential for successful RTO. This includes preprocessing real-time data, loading data into a data warehouse, and querying the stored data. The second section of this report describes the progress made in implementing an adapter in Python in order for Deep Lynx to handle the data communication. Typical dispatch optimization frameworks are built on linear programming (LP). The prototype RTO workflow developed in this report uses an LP problem as a part of a receding-horizon- or economic model predictive control (EMPC) based optimization. The third section of this report details the framework of an RTO workflow in which the system consists of a simple electrical storage device. A DT can be built from a reduced-order model (ROM). Integrating a ROM into a typical LP optimization framework has been challenging because most optimization packages require the user to write algebraic expressions for the system model. The final section of this report shows how an externally built RAVEN ROM can be integrated in an RTO framework by using the Python package Pyomo. This demonstrates the RTO workflow capability from a software-only perspective and is an important step in demonstrating the capability to implement an RTO workflow for a physical system.

97 MATHEMATICS AND COMPUTING↗