Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Development of Short-Term Forecasting Models Using Plant Asset Data and Feature Selection

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This paper focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 +- 0.014, 0.0026 +- 0, and 0.063 +- 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Using Hyperspectral Imagery to Identify Turfgrass Stresses

The use of a form of remote sensing to aid in the management of large turfgrass fields (e.g. golf courses) has been proposed. A turfgrass field of interest would be surveyed in sunlight by use of an airborne hyperspectral imaging system, then the raw observational data would be preprocessed into hyperspectral reflectance image data. These data would be further processed to identify turfgrass stresses, to determine the spatial distributions of those stresses, and to generate maps showing the spatial distributions. Until now, chemicals and water have often been applied, variously, (1) indiscriminately to an entire turfgrass field without regard to localization of specific stresses or (2) to visible and possibly localized signs of stress for example, browning, damage from traffic, or conspicuous growth of weeds. Indiscriminate application is uneconomical and environmentally unsound; the amounts of water and chemicals consumed could be insufficient in some areas and excessive in most areas, and excess chemicals can leak into the environment. In cases in which developing stresses do not show visible signs at first, it could be more economical and effective to take corrective action before visible signs appear. By enabling early identification of specific stresses and their locations, the proposed method would provide guidance for planning more effective, more economical, and more environmentally sound turfgrass-management practices, including application of chemicals and water, aeration, and mowing. The underlying concept of using hyperspectral imagery to generate stress maps as guides to efficient management of vegetation in large fields is not new; it has been applied in the growth of crops to be harvested. What is new here is the effort to develop an algorithm that processes hyperspectral reflectance data into spectral indices specific to stresses in turfgrass. The development effort has included a study in which small turfgrass plots that were, variously, healthy or subjected to a variety of controlled stresses were observed by use of a hand-held spectroradiometer. The spectroradiometer readings in the wavelength range from 350 to 1,000 nm were processed to extract hyperspectral reflectance data, which, in turn, were analyzed to find correlations with the controlled stresses. Several indices were found to be correlated with drought stress and to be potentially useful for identifying drought stress before visible symptoms appear.

Hutto, Kendall↗

Dynamic Server-Based KML Code Generator Method for Level-of-Detail Traversal of Geospatial Data

Web-based geospatial client applications such as Google Earth and NASA World Wind must listen to data requests, access appropriate stored data, and compile a data response to the requesting client application. This process occurs repeatedly to support multiple client requests and application instances. Newer Web-based geospatial clients also provide user-interactive functionality that is dependent on fast and efficient server responses. With massively large datasets, server-client interaction can become severely impeded because the server must determine the best way to assemble data to meet the client applications request. In client applications such as Google Earth, the user interactively wanders through the data using visually guided panning and zooming actions. With these actions, the client application is continually issuing data requests to the server without knowledge of the server s data structure or extraction/assembly paradigm. A method for efficiently controlling the networked access of a Web-based geospatial browser to server-based datasets in particular, massively sized datasets has been developed. The method specifically uses the Keyhole Markup Language (KML), an Open Geospatial Consortium (OGS) standard used by Google Earth and other KML-compliant geospatial client applications. The innovation is based on establishing a dynamic cascading KML strategy that is initiated by a KML launch file provided by a data server host to a Google Earth or similar KMLcompliant geospatial client application user. Upon execution, the launch KML code issues a request for image data covering an initial geographic region. The server responds with the requested data along with subsequent dynamically generated KML code that directs the client application to make follow-on requests for higher level of detail (LOD) imagery to replace the initial imagery as the user navigates into the dataset. The approach provides an efficient data traversal path and mechanism that can be flexibly established for any dataset regardless of size or other characteristics. The method yields significant improvements in userinteractive geospatial client and data server interaction and associated network bandwidth requirements. The innovation uses a C- or PHP-code-like grammar that provides a high degree of processing flexibility. A set of language lexer and parser elements is provided that offers a complete language grammar for writing and executing language directives. A script is wrapped and passed to the geospatial data server by a client application as a component of a standard KML-compliant statement. The approach provides an efficient means for a geospatial client application to request server preprocessing of data prior to client delivery. Data is structured in a quadtree format. As the user zooms into the dataset, geographic regions are subdivided into four child regions. Conversely, as the user zooms out, four child regions collapse into a single, lower-LOD region. The approach provides an efficient data traversal path and mechanism that can be flexibly established for any dataset regardless of size or other characteristics.

Baxes, Gregory↗

Data Democratization: Challenges and Opportunities

Democratizing Earth data is one of the challenges many organizations around the world face in order to maximize the use of their Earth data for research, applications, education, and societal benefits. For example, at the NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC), over 1600 global and regional datasets in several NASA Earth science focus areas, including atmospheric composition, water and energy cycles, and climate variability, are archived and distributed to the public. Giovanni, the Geospatial Interactive Online Visualization and Analysis Infrastructure, was developed by GES DISC to facilitate data access and exploration, especially for novice users of Earth science. With Giovanni, users can analyze and visualize over 2000 Earth science variables (e.g., precipitation, aerosol, surface wind) without downloading data, software, the expert understanding of data formats and structures, and coding skills, lowering the barrier to data analysis/comparison by preprocessing and accessing to the data. Results of data analysis and visualization can be accessed in several popular formats (e.g., NetCDF, CSV). As a result of Giovanni's efforts, more than 3000 referral papers have been published in various fields. In spite of this, Giovanni is still difficult to use for some users. For instance, if one searches for "precipitation," it will return over 150 related variables. The question is, which one to use? Furthermore, variables from different data providers (e.g., satellites and models) are named differently with different units, further confusing users, especially those outside the communities. Data democratization is complex and multifaceted. Challenges include service and data discovery, user experiences, visualization, data quality, trustworthiness, and more. In this presentation, we will examine Giovanni as an example of challenges and opportunities in developing data democratization services.

data democratization↗

A new approach to telemetry data processing

An approach for a preprocessing system for telemetry data processing was developed. The philosophy of the approach is the development of a preprocessing system to interface with the main processor and relieve it of the burden of stripping information from a telemetry data stream. To accomplish this task, a telemetry preprocessing language was developed. Also, a hardware device for implementing the operation of this language was designed using a cellular logic module concept. In the development of the hardware device and the cellular logic module, a distributed form of control was implemented. This is accomplished by a technique of one-to-one intermodule communications and a set of privileged communication operations. By transferring this control state from module to module, the control function is dispersed through the system. A compiler for translating the preprocessing language statements into an operations table for the hardware device was also developed. Finally, to complete the system design and verify it, a simulator for the collular logic module was written using the APL/360 system.

Broglio, C. J.↗

An Inventory of AI-ready Benchmark Data for US Fires, Heatwaves, and Droughts

Extreme weather events, including fires, heatwaves, and droughts, have significant impacts on earth, environmental, and energy systems. Mechanistic and predictive understanding, as well as probabilistic risk assessment of these extreme weather events, are crucial for detecting, planning for, and responding to these extremes. Records of extreme weather events provide an important data source for understanding present and future extremes, but the existing data needs preprocessing before it can be used for analysis. Moreover, there are many nonstandard metrics defining the levels of severity or impacts of extremes. In this study, we compile a comprehensive benchmark data inventory of extreme weather events, including fires, heatwaves, and droughts. The dataset covers the period from 2001 to 2020 with a daily temporal resolution and a spatial resolution of 0.5°×0.5° (~55km×55km) over the continental United States (CONUS), and a spatial resolution of 1km × 1km over the Pacific Northwest (PNW) region, together with the co-located and relevant meteorological variables. By exploring and summarizing the spatial and temporal patterns of these extremes in various forms of marginal, conditional, and joint probability distributions, we gain a better understanding of the characteristics of climate extremes. The resulting AI/ML-ready data products can be readily applied to ML-based research, fostering and encouraging AI/ML research in the field of extreme weather. This study can contribute significantly to the advancement of extreme weather research, aiding researchers, policymakers, and practitioners in developing improved preparedness and response strategies to protect communities and ecosystems from the adverse impacts of extreme weather events.

54 ENVIRONMENTAL SCIENCES↗

An Inventory of AI-ready Benchmark Data for US Fires, Heatwaves, and Droughts

Extreme weather events, including fires, heatwaves, and droughts, have significant impacts on earth, environmental, and energy systems. Mechanistic and predictive understanding, as well as probabilistic risk assessment of these extreme weather events, are crucial for detecting, planning for, and responding to these extremes. Records of extreme weather events provide an important data source for understanding present and future extremes, but the existing data needs preprocessing before it can be used for analysis. Moreover, there are many nonstandard metrics defining the levels of severity or impacts of extremes. In this study, we compile a comprehensive benchmark data inventory of extreme weather events, including fires, heatwaves, and droughts. The dataset covers the period from 2001 to 2020 with a daily temporal resolution and a spatial resolution of 0.5°×0.5° (~55km×55km) over the continental United States (CONUS), and a spatial resolution of 1km × 1km over the Pacific Northwest (PNW) region, together with the co-located and relevant meteorological variables. By exploring and summarizing the spatial and temporal patterns of these extremes in various forms of marginal, conditional, and joint probability distributions, we gain a better understanding of the characteristics of climate extremes. The resulting AI/ML-ready data products can be readily applied to ML-based research, fostering and encouraging AI/ML research in the field of extreme weather. This study can contribute significantly to the advancement of extreme weather research, aiding researchers, policymakers, and practitioners in developing improved preparedness and response strategies to protect communities and ecosystems from the adverse impacts of extreme weather events. Usage Notes We presented a long term (2001-2020) and comprehensive data inventory of historical extreme events with daily temporal resolution covering the separate spatial extents of CONUS (0.5°×0.5°) and PNW(1km×1km) for various applications and studies. The dataset with 0.5°×0.5° resolution for CONUS can be used to help build more accurate climate models for the entire CONUS, which can help in understanding long-term climate trends, including changes in the frequency and intensity of extreme events, predicting future extreme events as well as understanding the implications of extreme events on society and the environment. The data can also be applied for risk accessment of the extremes. For example, ML/AI models can be developed to predict wildfire risk or forecast HWs by analyzing historical weather data, and past fires or heateave , allowing for early warnings and risk mitigation strategies. Using this dataset, AI-driven risk assessment models can also be built to identify vulnerable energy and utilities infrastructure, imrpove grid resilience and suggest adaptations to withstand extreme weather events. The high-resolution 1km×1km dataset ove PNW are advantageous for real-time, localized and detailed applications. It can enhance the accuracy of early warning systems for extreme weather events, helping authorities and communities prepare for and respond to disasters more effectively. For example, ML models can be developed to provide localized HW predictions for specific neighborhoods or cities, enabling residents and local emergency services to take targeted actions; the assessment of drought severity in specific communities or watersheds within the PNW can help local authorities manage water resources more effectively.

Lin, Xinming↗

Users manual for the US baseline corn and soybean segment classification procedure

A user's manual for the classification component of the FY-81 U.S. Corn and Soybean Pilot Experiment in the Foreign Commodity Production Forecasting Project of AgRISTARS is presented. This experiment is one of several major experiments in AgRISTARS designed to measure and advance the remote sensing technologies for cropland inventory. The classification procedure discussed is designed to produce segment proportion estimates for corn and soybeans in the U.S. Corn Belt (Iowa, Indiana, and Illinois) using LANDSAT data. The estimates are produced by an integrated Analyst/Machine procedure. The Analyst selects acquisitions, participates in stratification, and assigns crop labels to selected samples. In concert with the Analyst, the machine digitally preprocesses LANDSAT data to remove external effects, stratifies the data into field like units and into spectrally similar groups, statistically samples the data for Analyst labeling, and combines the labeled samples into a final estimate.

Horvath, R.↗

Utah FORGE InSAR Data from 2020

Interferometric Synthetic Aperture Radar data from the TerraSAR-X and the TanDEM-X satellite missions operated by the German Space Agency (DLR). Interferometric pairs (interferograms) were created using generic mapping tool GMT-SAR processing software (see link in Resources). Data from January through December 2020.

15 GEOTHERMAL ENERGY↗

Utah FORGE InSAR Data from 2021

Interferometric Synthetic Aperture Radar data from the TerraSAR-X and the TanDEM-X satellite missions operated by the German Space Agency (DLR). Interferometric pairs (interferograms) were created using generic mapping tool GMT-SAR processing software (see link in Resources). Data from January through November 2021.

15 GEOTHERMAL ENERGY↗

Utah FORGE InSAR Data from 2022

Interferometric Synthetic Aperture Radar data from the TerraSAR-X and the TanDEM-X satellite missions operated by the German Space Agency (DLR). Interferometric pairs (interferograms) were created using generic mapping tool GMT-SAR processing software (see link in Resources). Data from January through June 2022.

15 GEOTHERMAL ENERGY↗

Automating Traffic Microsimulation from SYNCHRO UTDF to SUMO

Modern transportation research relies on seamlessly integrating traffic signal data with robust network representation and simulation tools. This study presents utdf2gmns, an open-source Python tool that automates conversion of the Universal Traffic Data Format, including network representation, signalized intersections, and turning volumes into the General Modeling Network Specification (GMNS) Standard. The resulting GMNS-compliant network can be converted for microsimulation in SUMO. By automatically extracting intersection control parameters and aligning them with GMNS conventions, utdf2gmns minimizes manual preprocessing and data loss. utdf2gmns also integrates with the Sigma-X engine to extract and visualize key traffic control metrics, such as phasing diagrams, turning volumes, volume-tocapacity ratios, and control delays. This streamlined workflow enables efficient scenario testing, accurate model building, and consistent data management. Validated through case studies, utdf2gmns reliably models complex urban corridors, promoting reproducibility and standardization. Documentation is available on GitHub and PyPI, supporting easy integration and community engagement.

Luo, Roy [ORNL] (ORCID:0009000312909983)↗

Image data processing of earth resources management

The Earth Resources Technology Satellite (ERTS-1) was launched on July 23, 1972. Since that date, the vehicle has been acquiring and transmitting to earth a vast amount of information relating to the earth's resources. Scientific investigators throughout the world have been analyzing photographic products produced from this data with significant results. Recently, there has been a major shift of emphasis from the visual analysis of photo products to the use of sophisticated digital information extraction techniques. Specialized equipment, such as the GE IMAGE 100 System, has been developed that operates directly from Computer Compatible Tapes, (CCT’s), which contain preprocessed digital data acquired from the ERTS spacecraft. Analysis of CCT data avoids the losses in image resolution and quality which are inherent to all photo processing techniques, thereby greatly increasing the utility of ERTS data. Future ground processing systems will turn out fully corrected (radiometric and geometric) digital tapes which will further increase the utility of ERTS data. This paper is presented in two parts: the first part deals with the subject of present and future generations of image processing and information extraction systems under development at the General Electric Company; the second part describes in more detail the design and operation of GE1s interactive multispectral information extraction systems, IMAGE 100, and discusses results of analyses of ERTS data over a number of U.S. sites.

A W DeSio↗

Study to determine cloud motion from meteorological satellite data

Processing techniques were tested for deducing cloud motion vectors from overlapped portions of pairs of pictures made from meteorological satellites. This was accomplished by programming and testing techniques for estimating pattern motion by means of cross correlation analysis with emphasis placed upon identifying and reducing errors resulting from various factors. Techniques were then selected and incorporated into a cloud motion determination program which included a routine which would select and prepare sample array pairs from the preprocessed test data. The program was then subjected to limited testing with data samples selected from the Nimbus 4 THIR data provided by the 11.5 micron channel.

Clark, B. B.↗

Bhutan Water Resources III: Analyzing Forest Disturbances and Climate Data in Bhutan to Create a Tool for Assisting the Himalayan Environment Rhythm Observation and Evaluation Systems (HEROES) Project

Forest disturbances from bark beetle outbreaks are a major concern in Bhutan, known to cause extensive tree mortality to pine and spruce forests. The NASA DEVELOP team partnered with the Ugyen Wangchuck Institute of Conservation and Environmental Research (UWICER), the Bhutan Foundation, and the Karuna Foundation to assess forest changes for the districts of Bumthang and Haa from 2000 to 2018. The project used preprocessed meteorological data from the Climate Hazards Center Infrared Precipitation with Station (CHIRPS) and Famine Early Warning System Network Land Data Assimilation System (FLDAS), along with Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper plus (ETM+), and Landsat 8 Operational Land Imager (OLI) to assess apparent forest disturbance occurrences and observed climate trends. Shuttle Radar Topography Mission (SRTM) was used to resolve variations in elevation and slope for mountainous regions. Using the Google Earth Engine LandTrendr (LT) code algorithm, along with Landsat data, the team developed an app called Forest Disturbances Detection Toolbox (FDDT) to assess forest changes in Bhutan. The app includes climate variables for the focus districts, along with LT variables, which allows the end users to further examine the cause of disturbances. The team compared geocoordinates for known disturbances with LT disturbance detection products. Although additional work is needed in the future to validate the project end products from the FDDT, the tool will be provided to the project partners to aid forest management efforts in Bhutan.

Tashi Choden↗

Some practical aspects of lossless and nearly-lossless compression of AVHRR imagery

Compression of Advanced Very high Resolution Radiometers (AVHRR) imagery operating in a lossless or nearly-lossless mode is evaluated. Several practical issues are analyzed including: variability of compression over time and among channels, rate-smoothing buffer size, multi-spectral preprocessing of data, day/night handling, and impact on key operational data applications. This analysis is based on a DPCM algorithm employing the Universal Noiseless Coder, which is a candidate for inclusion in many future remote sensing systems. It is shown that compression rates of about 2:1 (daytime) can be achieved with modest buffer sizes (less than or equal to 2.5 Mbytes) and a relatively simple multi-spectral preprocessing step.

Hogan, David B.↗

Data and scripts associated with the manuscript "Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning"

This package contains the data and scripts used in "Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning" (Jiang et al., 2022). The data.zip file contains the flux tower and automated chamber observations used for developing the deep learning model for modeling soil respiration. The scripts.zip file contains the Jupyter notebooks and python scripts for preprocessing the data, training the deep learning models, and postprocessing the results. The src.zip contains the source code for training the deep learning model, performing mutual information analysis, and plotting functions. The trained_models.zip contains multiple folders used for hosting the trained deep-learning models and the associated soil respiration predictions. The whole process is performed using python. We include the REAMD.md to document the python package requirements.Soil respiration in dryland ecosystems is challenging to model due to its complex interactions with environmental drivers. Knowledge-guided deep learning provides a much more effective means of accurately representing these complex interactions than traditional Q10-based models. Mutual information analysis revealed that future soil temperature shares more information with soil respiration than past soil temperature, consistent with their clockwise diel hysteresis. We explicitly encoded diel hysteresis, soil drying, and soil rewetting effects on soil respiration dynamics in a newly designed Long Short Term Memory (LSTM) model. The model takes both past and future environmental drivers as inputs to predict soil respiration. The new LSTM model substantially outperformed three Q10-based models and the Community Land Model when reproducing the observed soil respiration dynamics in a semi-arid ecosystem. The new LSTM model clearly demonstrated its superiority for temporally extrapolating soil respiration dynamics, such that the resulting correlation with observational data is up to 0.7 while the correlations of both Q10-based models and the Community Land Model (CLM) are less than 0.4. Our results underscore the high potential for knowledge-guided deep learning to replace Q10-based soil respiration modules in Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Model Data Archive for Manuscript Titled "Evaluation of a Coupled Surface–Subsurface Hydrologic Model Using Dense Water‑Level Sensors in a Mixed Urban–Rural Watershed"

This archive provides scripts, input files, and datasets used for the implementation and evaluation of a fully coupled surface–subsurface hydrologic model in the Neches River Basin, southeast Texas. The study uses the Advanced Terrestrial Simulator (ATS) to simulate coupled surface–subsurface hydrologic processes over a mixed urban–rural watershed and evaluates model performance using a dense network of 136 in situ water-level sensors, nine U.S. Geological Survey (USGS) stream gauges, and SSEBop-derived evapotranspiration estimates during the period October 2014–June 2024. The workflow is implemented primarily in Python 3 using the Watershed Workflow package. The Jupyter notebooks can be executed using open-source software such as Anaconda JupyterLab or Visual Studio Code. Other data files include TXT, CSV, XML, SHP, TIF, NetCDF, HDF5, and ExodusII files, which can be processed using the provided Python scripts. ATS input files are provided in XML format and can be edited using any commonly used text editor. This archive contains: *Scripts and input files used to generate the ATS model setup, including watershed discretization, mesh generation, parameter mapping, and model configuration. *Jupyter notebooks used for preprocessing observational data, evaluating streamflow, water levels, and evapotranspiration, computing performance metrics, and generating the figures presented in the manuscript. *ATS simulation outputs and processed observational datasets, including OneRain and DD6 water-level sensors, USGS streamflow observations, GIS data, and supporting spatial datasets used throughout the study.

Dense water-level sensor network↗