Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data enhancement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Techniques for storing data to enhance recovery and detection of data corruption errors

Often there are errors when reading data from computer memory. To detect and correct these errors, there are multiple types of error correction codes. Disclosed is an error correction architecture that creates a codeword having a data portion and an error correction code portion. Swizzling rearranges the order of bits and distributes the bits among different codewords. Because the data is redistributed, a potential memory error of up to N contiguous bits, where N for example equals 2 times the number of codewords swizzled together, only affects up to, at most, two bits per swizzled codeword. This keeps the error within the error detecting capabilities of the error correction architecture. Furthermore, this can allow improved error correction and detection without requiring a change to error correcting code generators and checkers.

Mills, Peter↗

Techniques for storing data to enhance recovery and detection of data corruption errors

Often there are errors when reading data from computer memory. To detect and correct these errors, there are multiple types of error correction codes. Disclosed is an error correction architecture that creates a codeword having a data portion and an error correction code portion. Swizzling rearranges the order of bits and distributes the bits among different codewords. Because the data is redistributed, a potential memory error of up to N contiguous bits, where N for example equals 2 times the number of codewords swizzled together, only affects up to, at most, two bits per swizzled codeword. This keeps the error within the error detecting capabilities of the error correction architecture. Furthermore, this can allow improved error correction and detection without requiring a change to error correcting code generators and checkers.

Mills, Peter↗

Digital Image Correlation Data Processing and Analysis Techniques to Enhance Test Data Assessment and Improve Structural Simulations

The NASA Shell Buckling Knockdown Factor Project (SBKF) was established in 2007 by the NASA Engineering and Safety Center (NESC) with the primary goal to develop new analysis-based buckling design factors (a.k.a. knockdown factors) and high-fidelity buckling simulations for selected launch-vehicle-like cylindrical shell structures. A series of tests are being conducted on large-scale metallic and composite cylindrical shells in order to provide validation data for these new factors and simulations. However, the validation of these new factors and simulations is quite demanding and requires test data that is commensurate with their fidelity. Traditional instrumentation, such as linear variable displacement transducers (LVDTs) and electrical-resistance strain gages serve a critical role in providing accurate displacement and strain measurements in these tests, but only allow for data to be recorded at a select number of point locations and are not sufficient to provide all the necessary validation data. Advanced measurement technologies can be used effectively to complement traditional instrumentation and gather additional data required to validate these structural simulations. In particular, three-dimensional digital image correlation (DIC) was implemented during SBKF cylinder testing to characterize the full-field displacement and strain behavior. Commercially available VIC-3DTM software and user-written data processing scripts were used to generate valuable data and insight into the complex buckling response of the cylinders that otherwise would be impossible to gather using traditional instrumentation. In addition, the measured data from DIC was used to verify measured test data obtained from other instrumentation, enhance test and analysis correlation, and help identify the root cause of anomalous test results that may have gone unexplained if only traditional instrumentation was used. Selected test results that demonstrate the use of DIC on the SBKF cylinders are presented and a portion of the data processing methods are described.

Gardner, Nathaniel W.↗

A Data-Fusion Method using Bayesian Approach to Enhance Raw Data Accuracy of Position and Distance Measurements for Connected Vehicles

Accurate positioning of vehicles is a critical element of autonomous and connected vehicle systems. Most of other studies heavily focused on enhancing simultaneous localization and mapping (SLAM) methods, i.e., constructing or updating a map of an unknown environment and tracking an object within the map. This paper provides a method that can, in addition to existing SLAM or relevant methods, enhance the raw measurements of position and distance. The basic idea of this study is to identify and update the error distribution of each data source by combining all available information. A Bayesian approach was incorporated to estimate and update the error distribution of individual data sources or sensors. The proposed method can be conducted in real-time environments, and a self-learning scheme determines whether enough data has been collected to further improve the accuracy of such measurements. The simulated experiments show that the proposed model noticeably improves the accuracy of position and distance measurements. Especially, the estimated biases of position coordinates and distance measures are very close to the biases of true error distributions, with the R-squared over 0.98. A similar approach can also be utilized to enhance accuracy of other sensors or measurements in connected vehicle or relevant systems, where multi-data sources are available.

Lim, Hyeonsup↗

Mathematical analysis study for radar data processing and enhancement. Part 1: Radar data analysis

A study is performed under NASA contract to evaluate data from an AN/FPS-16 radar installed for support of flight programs at Dryden Flight Research Facility of NASA Ames Research Center. The purpose of this study is to provide information necessary for improving post-flight data reduction and knowledge of accuracy of derived radar quantities. Tracking data from six flights are analyzed. Noise and bias errors in raw tracking data are determined for each of the flights. A discussion of an altiude bias error during all of the tracking missions is included. This bias error is defined by utilizing pressure altitude measurements made during survey flights. Four separate filtering methods, representative of the most widely used optimal estimation techniques for enhancement of radar tracking data, are analyzed for suitability in processing both real-time and post-mission data. Additional information regarding the radar and its measurements, including typical noise and bias errors in the range and angle measurements, is also presented. This is in two parts. This is part 1, an analysis of radar data.

James, R.↗

Building Stock Models for Embodied Carbon Emissions—A Review of a Nascent Field

Building stock modeling emerges as a critical tool in the strategic reduction of embodied carbon emissions, which is pivotal in reshaping the evolving construction sector. This review provides an overall view of modern methodologies in building stock modeling, homing in on the nuances of embodied carbon analysis in construction. Examining 23 seminal papers, our study delineates two primary modeling paradigms—top-down and bottom-up—each further compartmentalized into five innovative methods. This study points out the challenges of data scarcity and computational demands, advocating for methodological advancements that promise to refine the precision of building stock models. A groundbreaking trend in recent research is the incorporation of machine learning algorithms, which have demonstrated remarkable capacity, improving stock classification accuracy by 25% and urban material quantification by 40%. Furthermore, the application of remote sensing has revolutionized data acquisition, enhancing data richness by a factor of five. This review offers a critical examination of current practices and charts a course toward an environmentally prudent future. It underscores the transformative impact of building stock modeling in driving ecological stewardship in the construction industry, positioning it as a cornerstone in the quest for sustainability and its significant contribution toward the grand vision of an eco-efficient built environment.

Hu, Ming (ORCID:0000000325831161)↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

Radar range data signal enhancement tracker

The design, fabrication, and performance characteristics are described of two digital data signal enhancement filters which are capable of being inserted between the Space Shuttle Navigation Sensor outputs and the guidance computer. Commonality of interfaces has been stressed so that the filters may be evaluated through operation with simulated sensors or with actual prototype sensor hardware. The filters will provide both a smoothed range and range rate output. Different conceptual approaches are utilized for each filter. The first filter is based on a combination low pass nonrecursive filter and a cascaded simple average smoother for range and range rate, respectively. Filter number two is a tracking filter which is capable of following transient data of the type encountered during burn periods. A test simulator was also designed which generates typical shuttle navigation sensor data.

Source record↗

Optimal Reorganization of NASA Earth Science Data for Enhanced Accessibility and Usability for the Hydrology Community

A long-standing "Digital Divide" in data representation exists between the preferred way of data access by the hydrology community and the common way of data archival by earth science data centers. Typically, in hydrology, earth surface features are expressed as discrete spatial objects (e.g., watersheds), and time-varying data are contained in associated time series. Data in earth science archives, although stored as discrete values (of satellite swath pixels or geographical grids), represent continuous spatial fields, one file per time step. This Divide has been an obstacle, specifically, between the Consortium of Universities for the Advancement of Hydrologic Science, Inc. and NASA earth science data systems. In essence, the way data are archived is conceptually orthogonal to the desired method of access. Our recent work has shown an optimal method of bridging the Divide, by enabling operational access to long-time series (e.g., 36 years of hourly data) of selected NASA datasets. These time series, which we have termed "data rods," are pre-generated or generated on-the-fly. This optimal solution was arrived at after extensive investigations of various approaches, including one based on "data curtains." The on-the-fly generation of data rods uses "data cubes," NASA Giovanni, and parallel processing. The optimal reorganization of NASA earth science data has significantly enhanced the access to and use of the data for the hydrology user community.

data rods↗

Application of SEASAT-1 Synthetic Aperture Radar (SAR) data to enhance and detect geological lineaments and to assist LANDSAT landcover classification mapping

Digital SEASAT-1 synthetic aperture radar (SAR) data were used to enhance linear features to extract geologically significant lineaments in the Appalachian region. Comparison of Lineaments thus mapped with an existing lineament map based on LANDSAT MSS images shows that appropriately processed SEASAT-1 SAR data can significantly improve the detection of lineaments. Merge MSS and SAR data sets were more useful fo lineament detection and landcover classification than LANDSAT or SEASAT data alone. About 20 percent of the lineaments plotted from the SEASAT SAR image did not appear on the LANDSAT image. About 6 percent of minor lineaments or parts of lineaments present in the LANDSAT map were missing from the SEASAT map. Improvement in the landcover classification (acreage and spatial estimation accuracy) was attained by using MSS-SAR merged data. The aerial estimation of residential/built-up and forest categories was improved. Accuracy in estimating the agricultural and water categories was slightly reduced.

Sekhon, R.↗

Guidelines for Publicly Archiving Terrestrial Model Data to Enhance Usability, Intercomparison, and Synthesis

Scientific communities are increasingly publishing data to evaluate, accredit, and build on published research. However, guidelines for curating data for publication are sparse for model-related research, limiting the usability of archived simulation data. In particular, there are no established guidelines for archiving data related to terrestrial models that simulate land processes and their coupled interactions with climate. Terrestrial modelers have a unique set of challenges when publishing data due to the diversity of scientific domains, research questions, and the types and scales of simulations. Researchers in the U.S. Department of Energy’s (DOE) projects use a variety of multiscale models to advance robust predictions of terrestrial and subsurface ecosystem processes. Here, we synthesize archiving needs for data associated with different DOE models, and provide guidelines for publishing terrestrial model data components following FAIR (Findable, Accessible, Interoperable, Reusable) principles. The guidelines recommend archiving model inputs and testing data used in final simulation runs along with associated codes, workflow scripts, and metadata in public repositories. Researchers should consider archiving model outputs if they are within the storage limits of the repository. We also provide considerations for how to bundle files into different data publications with citable digital object identifiers. Finally, we identify repository features and tools that would enable storage and reuse of model data. Given the diversity of DOE terrestrial models, these guidelines are transferable to other model types and will enable efficient reuse of simulation data for purposes such as model intercomparisons, initialization, benchmarking, synthesis, and comparisons with field observations.

58 GEOSCIENCES↗

RHSEG and Subdue: Background and Preliminary Approach for Combining these Technologies for Enhanced Image Data Analysis, Mining and Knowledge Discovery

Under a project recently selected for funding by NASA's Science Mission Directorate under the Applied Information Systems Research (AISR) program, Tilton and Cook will design and implement the integration of the Subdue graph based knowledge discovery system, developed at the University of Texas Arlington and Washington State University, with image segmentation hierarchies produced by the RHSEG software, developed at NASA GSFC, and perform pilot demonstration studies of data analysis, mining and knowledge discovery on NASA data. Subdue represents a method for discovering substructures in structural databases. Subdue is devised for general-purpose automated discovery, concept learning, and hierarchical clustering, with or without domain knowledge. Subdue was developed by Cook and her colleague, Lawrence B. Holder. For Subdue to be effective in finding patterns in imagery data, the data must be abstracted up from the pixel domain. An appropriate abstraction of imagery data is a segmentation hierarchy: a set of several segmentations of the same image at different levels of detail in which the segmentations at coarser levels of detail can be produced from simple merges of regions at finer levels of detail. The RHSEG program, a recursive approximation to a Hierarchical Segmentation approach (HSEG), can produce segmentation hierarchies quickly and effectively for a wide variety of images. RHSEG and HSEG were developed at NASA GSFC by Tilton. In this presentation we provide background on the RHSEG and Subdue technologies and present a preliminary analysis on how RHSEG and Subdue may be combined to enhance image data analysis, mining and knowledge discovery.

Tilton, James C.↗

The VIIRS Ocean Data Simulator Enhancements and Results

The VIIRS Ocean Science Team (VOST) has been developing an Ocean Data Simulator to create realistic VIIRS SDR datasets based on MODIS water-leaving radiances. The simulator is helping to assess instrument performance and scientific processing algorithms. Several changes were made in the last two years to complete the simulator and broaden its usefulness. The simulator is now fully functional and includes all sensor characteristics measured during prelaunch testing, including electronic and optical crosstalk influences, polarization sensitivity, and relative spectral response. Also included is the simulation of cloud and land radiances to make more realistic data sets and to understand their important influence on nearby ocean color data. The atmospheric tables used in the processing, including aerosol and Rayleigh reflectance coefficients, have been modeled using VIIRS relative spectral responses. The capabilities of the simulator were expanded to work in an unaggregated sample mode and to produce scans with additional samples beyond the standard scan. These features improve the capability to realistically add artifacts which act upon individual instrument samples prior to aggregation and which may originate from beyond the actual scan boundaries. The simulator was expanded to simulate all 16 M-bands and the EDR processing was improved to use these bands to make an SST product. The simulator is being used to generate global VIIRS data from and in parallel with the MODIS Aqua data stream. Studies have been conducted using the simulator to investigate the impact of instrument artifacts. This paper discusses the simulator improvements and results from the artifact impact studies.

Robinson, Wayne D.↗

Using NASA Environmental Data to Enhance Public Health Decision Making

The Universities Space Research Association at the NASA Marshall Space Flight Center is collaborating with the University of Alabama at Birmingham (UAB) School of Public Health and the Centers for Disease Control and Prevention (CDC) to address issues of environmental health and enhance public health decision making by utilizing NASA remotely sensed data and products. The objectives of this collaboration are to develop high-quality spatial data sets of environmental variables, and deliver the data sets and associated analyses to local, state and federal end-user groups. These data can be linked spatially and temporally to public health data, such as mortality and disease morbidity, for further analysis and decision making. Three daily environmental data sets have been developed for the conterminous U.S. on different spatial resolutions for the time period 2003-2008: (1) spatial surfaces of estimated fine particulate matter (PM2.5) exposures on a 10-km grid utilizing the US Environmental Protection Agency (EPA) ground observations and NASA s MODerate-resolution Imaging Spectroradiometer (MODIS) data; (2) a 1-km grid of Land Surface Temperature (LST) using MODIS data; and (3) a 12-km grid of daily Solar Insolation (SI) and maximum and minimum air temperature using the North American Land Data Assimilation System (NLDAS) forcing data. These environmental data sets will be linked with public health data from the UAB REasons for Geographic And Racial Differences in Stroke (REGARDS) national cohort study to determine whether exposures to these environmental risk factors are related to cognitive decline and other health outcomes. These environmental datasets and public health linkage analyses will be made available to public health professionals, researchers and the general public through the CDC Wide-ranging Online Data for Epidemiologic Research (WONDER) system and through peer reviewed publications. To date, two of the data sets have been released to the public in CDC WONDER, Daily Air Temperature and Heat Index for years 1979-2010, and Daily Fine Particulate Matter (PM2.5) air quality measures for years 2003-2008. These data in CDC WONDER can be aggregated to the county-level, state-level, or regional-level as per users need and downloaded in tabular, graphical, and map formats. The summary statistical output are available to web and app developers via the WONDER Application Programming Interface (API). The linkage of these data with the CDC WONDER system provides a significant addition to CDC WONDER, allowing public health researchers and policy makers to better include environmental exposure data in the context of other health data available in CDC WONDER online system. It also substantially expands public access to NASA environmental data, making their use by a wide range of decision makers feasible.

Al-Hamdan, Mohammad↗

Connecting People to Data: Enabling Data Connected Communities through Enhancements to the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented a series of new features designed to connect people to data. These features, which are based on feedback from the GDR user community and surveys of the greater geothermal research community, are designed to improve data quality and empower members of all communities to better engage with geothermal data resources by providing universal access to data and by improving the connections between data providers, subject matter experts, and the communities of people using GDR data. This paper will explore some of the recent enhancements made to the GDR to improve data discoverability, reduce submission time, and result in better quality data submissions. These improvements include the ability for users to save a list of their favorite datasets, search for insight into geothermal datasets or data availability, or sign up to receive notifications of future updates to specific datasets. These improvements aim to enhance the overall user experience of the GDR while further connecting communities to the data they need to inform decisions, advance geothermal research, and develop innovative solutions to local energy problems.

access↗

CLAS12 remote data-stream processing using ERSAP framework

Implementing a physics data processing application is relatively straightforward with the use of current containerization technologies and container image runtime services, which are prevalent in most high-performance computing (HPC) environments. However, the process is complicated by the challenges associated with data provisioning and migration, impacting the ease of workflow migration and deployment. Transitioning from traditional file-based batch processing to data-stream processing workflows is suggested as a method to streamline these workflows. This transition not only simplifies file provisioning and migration but also significantly reduces the necessity for extensive disk space. Data-stream processing is particularly effective for real-time processing during data acquisition, thereby enhancing data quality assurance. This paper introduces the integration of the JLAB CLAS12 event reconstruction application within the ERSAP data-stream processing framework that facilitates the execution of streaming event reconstruction at a remote data center and enables the return streaming of reconstructed events to JLAB while circumventing the need for temporary data storage throughout the process.

Gyurjyan, Vardan↗