Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Database of Stress-Strain Properties Auto-generated from the Scientific Literature using ChemDataExtractor

Abstract There has been an ongoing need for information-rich databases in the mechanical-engineering domain to aid in data-driven materials science. To address the lack of suitable property databases, this study employs the latest version of the chemistry-aware natural-language-processing (NLP) toolkit, ChemDataExtractor, to automatically curate a comprehensive materials database of key stress-strain properties. The database contains information about materials and their cognate properties: ultimate tensile strength, yield strength, fracture strength, Young’s modulus, and ductility values. 720,308 data records were extracted from the scientific literature and organized into machine-readable databases formats. The extracted data have an overall precision, recall and F-score of 82.03%, 92.13% and 86.79%, respectively. The resulting database has been made publicly available, aiming to facilitate data-driven research and accelerate advancements within the mechanical-engineering domain.

Kumar, Pankaj

Computational epidemiological tools for pandemic analysis, understanding, and response

This suite of software tools is being developed to enhance and analyze computational epidemiological models that incorporate realistic disease dynamics and human behavior, with the goal of supporting epidemic and pandemic response. Specifically, the tools enable data analysis, feature extraction, data synthesis, machine learning model development, and prediction of key public health outcomes, such as cases, hospitalizations, deaths, and behavioral responses, for airborne infectious diseases like COVID-19 and influenza.

Butts, David

Preliminary efforts toward development of data handling and analysis software for unsteady flow measurements: An application for aeroelastic transonic flow configurations

A few years ago the Structural Dynamics Division at LaRC started ambitious experimental research efforts known as the Benchmark Models Program. The primary objective of this program was to provide experimental data that may serve as a calibration source for computational fluid dynamics (CFD) efforts that deal with aeroelastic unsteady flow configurations. It also focuses on the understanding of complex flow phenomenon associated with unsteady flow developments. The overall plan for the program has been described by Bennett, including a presentation of initial test results of flutter of a rigid wing mounted on flexible supports. An example of a test model employed to measure the dynamic response along with corresponding pressure distributions is shown. This model incorporates eighty pressure transducers distributed along two spanwise stations. In addition, the models are equipped with four accelerometers and two strain gages. The data handling system for the Benchmark Model Program is under development. Several interactive computer routines designed for the user interface, dynamic memory allocation, unsteady flow measurements data extraction, and further data processing were developed. To present a few examples of measured data, the unsteady pressure distributions and the wing model dynamic response were plotted.

Finaish, Fathi

Computer classification of remotely sensed multispectral image data by extraction and classification of homogeneous objects

A method of classification of digitized multispectral images is developed and experimentally evaluated on actual earth resources data collected by aircraft and satellite. The method is designed to exploit the characteristic dependence between adjacent states of nature that is neglected by the more conventional simple-symmetric decision rule. Thus contextual information is incorporated into the classification scheme. The principle reason for doing this is to improve the accuracy of the classification. For general types of dependence this would generally require more computation per resolution element than the simple-symmetric classifier. But when the dependence occurs in the form of redundance, the elements can be classified collectively, in groups, therby reducing the number of classifications required.

Kettig, R. L.

Sensor needs for agricultural applications

The peculiarities of agricultural remotely sensed data requirements evoke special sensor requirements. Vegetative species do not possess significantly different spectral signature at given phases of their development cycle. Hence, the key to their discriminability is the phasing of the phenologic cycle of the subject species. Significant improvements in classification can be obtained by consistently employing multi-temporal observations taken at specific times during the year. The present approach to agricultural data processing results in extracted data equal to approximately .05% of the acquired data. This paper discusses the derivation of agricultural peculiar requirements and the benefits to the end-to-end processing system by judicial utilization and placement of key editing functions such as sample segment extraction, cloudy image removal, sample registration and the elimination of redundant data.

Golden, H.

Scraping Unstructured Data to Explore the Relationship between Rainfall Anomalies and Vector-Borne Disease Outbreaks

According to the World Health Organization (WHO), vector-borne diseases such as malaria and dengue account for 17% of all infectious disease cases and lead to more than 700,000 deaths per year. Tracking and predicting the spread of vector-borne diseases is a vital task that could save hundreds of thousands of lives annually. Oftentimes, the first reports of vector-borne disease outbreaks occur through emails and online reporting systems long before they are officially documented. Tracking and predicting the emergence and spread of vector-borne disease outbreaks requires extracting data from these unstructured sources in combination with historical weather and climate data to understand the underlying background triggers and disease dynamics. In this work, we develop a data extraction pipeline for the online outbreak reporting website ProMED-mail that utilizes a web scraper, transformer neural network summarizer, and named entity recognizer to obtain a dataset of malaria, dengue, zika, and chikungunya outbreaks over the last 30 years. This scraped dataset was further analyzed in association with global rainfall anomalies derived from NASA’s Integrated Multi-satellitE Retrievals for GPM [Global Precipitation Mission] (IMERG) dataset. This preliminary analysis was to understand the effect of global rainfall patterns on the spread of vector-borne diseases. Analysis of the ProMED-mail and GPM data shows that vector-borne disease outbreaks are clustered towards the tropics and outbreaks are often amplified during the rainy seasons. Our scraped dataset can be a valuable tool in creating comprehensive georeferenced disease records for modeling and predicting future outbreaks.

Web scraping

Physical properties, internal structure, and the three‐dimensional petrography of CI chondrites

physical properties and the nature of their breccation, we investigated nine samples of the Ivuna and Orgueil CI chondrites ranging in size from 1 mm to 4 cm in approximate diameter. The combined mass of unique material investigated in this work is 113 g. For our investigations, we use ideal gas pycnometry, 3-D laser scanning, x-ray computed microtomography (μCT), and accompanying digital data extraction techniques. We found that the bulk density of the samples ranged from 1.61 to 2.10 g cm −3 . Larger samples tend to have a lower bulk density. Grain density (ranging from 2.44 to 2.55 g cm −3 ) is significantly less variable than the bulk density in our samples and the quantity of porosity (ranging from 14.6% to 33.8%) is the dominant factor in determining the bulk density of CI chondrite material. Our μCT results show that the visible porosity across all sizes of our CI chondrite samples is in the form of cracks, but these cracks can account for less than two-thirds of the porosity in the CI chondrites. Other porosity is not visible, even at μCT resolutions of 2.7 μm voxel edge −1 and we conclude that it is sub-micron in nature. It is not clear if the cracks seen in our samples are indigenous to the chondrites or are a result of terrestrial processes. We also find that the CI chondrites are excellent examples of the fractal-like nature of brecciation, where clasts can be observed at all scales we imaged. The breccias are composed of sub-equant-shaped and sub-rounded-textured clasts like melt-free impact breccias on other solar system bodies. From our μCT volume and digital data extraction, we determine that the Ivuna CI chondrite breccia is organized: the mostly sub-equant clasts within our ~2 cm chunk of Ivuna have a mean diameter of 1.33 mm and their aligned longest axes define a lineation structure. We speculate that the lineation was imparted after fragmentation of the clasts by slight shear on the parent asteroid which could be the result of seismic-related granular flow or mild non-axial impact-related compaction. These data will help to place returned asteroidal material from asteroids 162173 Ryugu and 101955 Bennu and the CI chondrites into a mutual geological context.

CI chondrite

Earth Science Data Analytics: Preparing for Extracting Knowledge from Information

Data analytics is the process of examining large amounts of data of a variety of types to uncover hidden patterns, unknown correlations and other useful information. Data analytics is a broad term that includes data analysis, as well as an understanding of the cognitive processes an analyst uses to understand problems and explore data in meaningful ways. Analytics also include data extraction, transformation, and reduction, utilizing specific tools, techniques, and methods. Turning to data science, definitions of data science sound very similar to those of data analytics (which leads to a lot of the confusion between the two). But the skills needed for both, co-analyzing large amounts of heterogeneous data, understanding and utilizing relevant tools and techniques, and subject matter expertise, although similar, serve different purposes. Data Analytics takes on a practitioners approach to applying expertise and skills to solve issues and gain subject knowledge. Data Science, is more theoretical (research in itself) in nature, providing strategic actionable insights and new innovative methodologies. Earth Science Data Analytics (ESDA) is the process of examining, preparing, reducing, and analyzing large amounts of spatial (multi-dimensional), temporal, or spectral data using a variety of data types to uncover patterns, correlations and other information, to better understand our Earth. The large variety of datasets (temporal spatial differences, data types, formats, etc.) invite the need for data analytics skills that understand the science domain, and data preparation, reduction, and analysis techniques, from a practitioners point of view. The application of these skills to ESDA is the focus of this presentation. The Earth Science Information Partners (ESIP) Federation Earth Science Data Analytics (ESDA) Cluster was created in recognition of the practical need to facilitate the co-analysis of large amounts of data and information for Earth science. Thus, from a to advance science point of view: On the continuum of ever evolving data management systems, we need to understand and develop ways that allow for the variety of data relationships to be examined, and information to be manipulated, such that knowledge can be enhanced, to facilitate science. Recognizing the importance and potential impacts of the unlimited ways to co-analyze heterogeneous datasets, now and especially in the future, one of the objectives of the ESDA cluster is to facilitate the preparation of individuals to understand and apply needed skills to Earth science data analytics. Pinpointing and communicating the needed skills and expertise is new, and not easy. Information technology is just beginning to provide the tools for advancing the analysis of heterogeneous datasets in a big way, thus, providing opportunity to discover unobvious scientific relationships, previously invisible to the science eye. And it is not easy It takes individuals, or teams of individuals, with just the right combination of skills to understand the data and develop the methods to glean knowledge out of data and information. In addition, whereas definitions of data science and big data are (more or less) available (summarized in Reference 5), Earth science data analytics is virtually ignored in the literature, (barring a few excellent sources).

data analytics

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:

Geological mapping in northwestern Saudi Arabia using LANDSAT multispectral techniques

Various computer enhancement and data extraction systems using LANDSAT data were assessed and used to complement a continuing geologic mapping program. Interactive digital classification techniques using both the parallel-piped and maximum-likelihood statistical approaches achieve very limited success in areas of highly dissected terrain. Computer enhanced imagery developed by color compositing stretched MSS ratio data was constructed for a test site in northwestern Saudi Arabia. Initial results indicate that several igneous and sedimentary rock types can be discriminated.

Blodget, H. W.

Geological mapping in northwestern Saudi Arabia using LANDSAT multispectral techniques

Various computer enhancement and data extraction systems using LANDSAT data were assessed and used to complement a continuing geologic mapping program. Interactive digital classification techniques using both the maximum-likelihood and thresholding statistical approaches achieve very limited success in areas of highly dissected terrain. Computer enhanced imagery developed by color compositing stretched MSS ratio data was constructed for a test site in northwestern Saudi Arabia. Initial results indicate that several igneous and sedimentary rock-types can be discriminated.

Blodget, H. W.

A 'digital' technique for manual extraction of data from aerial photography

The interpretation procedure described uses a grid cell approach. In addition, a random point is located in each cell. The procedure required that the cell/point grid be established on a base map, and identical grids be made to precisely match the scale of the photographic frames. The grid is then positioned on the photography by visual alignment to obvious features. Several alignments on one frame are sometimes required to make a precise match of all points to be interpreted. This system inherently corrects for distortions in the photography. Interpretation is then done cell by cell. In order to meet the time constraints, first order interpretation should be maintained. The data is put onto coding forms, along with other appropriate data, if desired. This 'digital' manual interpretation technique has proven to be efficient, and time and cost effective, while meeting strict requirements for data format and accuracy.

Istvan, L. B.

Analysis of GEOS-3 altimeter data and extraction of ocean wave height and dominant wavelength

When the amplitude and timing biases are removed from the GEOS-3 Sample and Hold (S&H) gates, the mean return waveforms can be excellently fitted with a theoretical template which represents the convolution of: (1) the radar point target response; (2) the range noise (jitter) in the altimeter tracking loop; (3) the sea surface height distribution; and (4) the antenna pattern as a function of the range to mean sea level. Several techniques of varying complexity to remove the effect of the tracking loop jitter in computing the wave height are considered. They include: (1) realigning the S&H gates to their actual positions with respect to mean sea level before averaging; (2) using the observed standard deviation on the altitude measurement to remove the integrated effect of the tracking loop jitter, and (3) using a look-up table to correct for the expected value of range noise. Analysis of skewness in the GEOS return waveform demonstrates the potential of a satellite radar altimeter to determine the dominant wavelength of ocean waves.

Walsh, E. J.

DIOGENES: Expert system for extraction of data system requirements

AA The initial operations concept expresses information about system objectives, and defines the system users, system interfaces, and operational performance constraints. We have developed a prototype expert system which has established the feasibility of automating a scenario-driven methodology for deriving top-level specifications and preliminary designs for user data systems. This scenario-driven methodology uses an initial design, an initial operations concept, and user scenarios as the starting point for system definition. The top-level initial design is a functional description of the system in the form of an annotated data flow diagram. The initial operations concept expresses informationabout system objectives, and defines the system users, system interfaces, and operational performance constraints. The user scenarios are detailed time-lined descriptions of user activities, developed by prospective end users. These scenarios, along with the initial design and operations concept, are analyzed and iterated by the expert system to form a consistent set. The resulting User Scenario-Operation Set plays a key role in the development of requirements and system tests.

Hobbs, Robert W.

Assessment of thermochemical nonequilibrium and slip effects for Orbital Reentry Experiment (OREX)

Results are provided from a viscous shock layer (VSL) analysis of the reentry flowfield around the forebody of the Japanese Orbital Reentry Experiment (OREX) vehicle. This vehicle is a 50 deg. spherically blunted cone with a nose radius of 1.35 m and a base diameter of 3.4 m. Calculations are done for the OREX trajectory from 105 to 48.4 km altitude range. A 7-species chemical model is found adequate for the flowfield analysis. However, for altitudes greater than 84 km, the low density effects (such as thermal nonequilibrium and slip) are to be implemented for good agreement between the predictions and flight inferred heat-transfer rate data. Further, at altitudes lower than 84 km, a finite surface recombination probability is to be employed in place of a non-catalytic surface for better comparison between the calculations and data. VSL results are also compared with the direct simulation Monte Carlo (DSMC) predictions at high altitudes (greater than 80 km) and the electron number density data for three altitudes in the OREX trajectory. Overall, there is a good comparison between the flight data and calculated results. With the ongoing refinements in data extraction procedures, the OREX data should prove valuable for validating theoretical models employed in flowfield codes for calculation of reacting-gas flowfields.

Gupta, Roop N.

Integration Time Required to Extract Accurate Data from Transonic Wind-Tunnel Tests

Because the forces and pressures on wind-tunnel models tested at transonic speeds are not steady, even for static aerodynamic tests, integration time is required to obtain data of acceptable accuracy. The integration time required for both static and dynamic tests is evaluated analytically and confirmed by experimental measurements. It is shown that, for static and dynamic tests, the accuracy obtained is a function of integration time, frequency of the signal, and the ratio of the dynamic amplitude to the full signal of interest. In addition, for the dynamic case, the frequency bandwidth used in analysis is important. Results of this study indicate that, for typical data accuracy desired from models in a large transonic wind tunnel (11- by 11-ft), up to the following integration times are required: static force and moment tests, 0.5 s; static pressure tests, 1 s; flutter tests, 30 to 60 s; and random-dynamic tests, 10 s.

Muhlstein, Lado, Jr.

Chip for CCSDS Compatible Serial Data Streams

A configurable service processor for telemetry ground stations is totally implemented in VLSI/ASIC hardware and finds use in spacecraft systems and other communications systems that operate according to CCSDS and CCSDS-like protocols. The service processor performs the traditional functions of data extraction at very high data and packet rates.

Jason T Dowling