Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Extraction from flight data of longitudinal aerodynamic coefficients in maneuvering flight for F-8C aircraft

Flight-test data were used to extract the longitudinal aerodynamic parameters of the F-8C aircraft. The aircraft was trimmed in a steady turn at angles of attack of approximately 9 deg and 13 deg at Mach numbers of 0.7 and 0.8. The parameters extracted resulted in a good match to the flight data and the values obtained were reasonable. The values were further verified by comparing the period and time to damp to half-amplitude, as calculated by using the extracted parameter values, with the period and time to damp to half-amplitude actually measured from the flight data traces. These results show that for the set of data examined, a mathematical model using linear aerodynamics was adequate to describe the response motions at the test angles of attack.

Suit, W. T.↗

A simulation framework for evaluating electronic order workflows in integrated health records

Electronic health record (EHR) systems are critical to modern healthcare delivery, yet the dynamic workflows that govern electronic order processing remain underexplored. Inefficiencies in these digital pathways can cause delays in care, repetitive workloads, and even patient harm. This study presents a discrete-event simulation framework used to reconstruct and evaluate EHR-based order workflows in a large integrated healthcare system. Using real-world data extracted from the Veterans Health Administration’s Corporate Data Warehouse, the authors mapped order events to standardized state transitions and modeled their progression across different facilities of varying complexity levels. After being calibrated with empirical distributions of transition times and validated against observed time-in-system metrics, the simulation demonstrates close alignment with historical performance. Scenario analyses reveal that resource capacity constraints significantly amplify the impact of electronic order surges, which are reflected in the disproportionate growth in backlogs and processing delays. Adjustments in transition probabilities further increased recirculation and extended workflow paths. Network-based analysis identified Reserved, InProgress, and Completed as structurally critical states that function as hubs within the process network but the transitions in-between also act as major bottlenecks. These results showcased the effectiveness of simulation-based approaches in monitoring EHR order processing performance and evaluating consequences of workflow changes on healthcare network resources planning. The proposed simulation framework provides a scalable data-driven tool to support operational decision-making and improve the efficiency of electronic order management in complex healthcare environments.

Engineering↗

Bringing Analysis Closer to Data: Developing a Visualization Tool for L2 Earth Science Satellite Data

Earth Science satellite missions provide a unique opportunity for scientists to visualize complex and multifaceted observations projected geospatially across maps of the Earth. While visualization tools can help scientists comprehend, analyze, and share data, visualizing Level-2 Earth Sciences data poses its own specific set of challenges. Since the geospatial information in Level-2 data files is stored as independent variables, the plotting process involves matching dimensional information from latitude and longitude with a desired variable. Variables are stored in different ways across various Earth Science data file formats, which complicates the process of extracting data and plotting variables from a given file without requiring extensive user input and prerequisite familiarity with the file type variable structure. In coordination with NASA’s Goddard Earth Sciences Data Information Services Center (GES DISC), the team developed a Level-2 Earth Science data visualization tool that aims to address some of the complexities associated with plotting Level-2 data. This tool offers command-line and user interface support for file and variable selection to accommodate varying use cases and degrees of user familiarity with the structure of a given file. The visualization tool is written in Python 3 and utilizes a modular approach to facilitate continued expansion and reuse. In addressing some common complications involved in plotting Level-2 Earth Sciences data, the tool aims to help to link the process of analysis more directly with data acquisition and visualization, bringing analysis closer to data across levels of processing.

Li, Angela W.↗

Feature extraction and classification algorithms for high dimensional data

Feature extraction and classification algorithms for high dimensional data are investigated. Developments with regard to sensors for Earth observation are moving in the direction of providing much higher dimensional multispectral imagery than is now possible. In analyzing such high dimensional data, processing time becomes an important factor. With large increases in dimensionality and the number of classes, processing time will increase significantly. To address this problem, a multistage classification scheme is proposed which reduces the processing time substantially by eliminating unlikely classes from further consideration at each stage. Several truncation criteria are developed and the relationship between thresholds and the error caused by the truncation is investigated. Next an approach to feature extraction for classification is proposed based directly on the decision boundaries. It is shown that all the features needed for classification can be extracted from decision boundaries. A characteristic of the proposed method arises by noting that only a portion of the decision boundary is effective in discriminating between classes, and the concept of the effective decision boundary is introduced. The proposed feature extraction algorithm has several desirable properties: it predicts the minimum number of features necessary to achieve the same classification accuracy as in the original space for a given pattern recognition problem; and it finds the necessary feature vectors. The proposed algorithm does not deteriorate under the circumstances of equal means or equal covariances as some previous algorithms do. In addition, the decision boundary feature extraction algorithm can be used both for parametric and non-parametric classifiers. Finally, some problems encountered in analyzing high dimensional data are studied and possible solutions are proposed. First, the increased importance of the second order statistics in analyzing high dimensional data is recognized. By investigating the characteristics of high dimensional data, the reason why the second order statistics must be taken into account in high dimensional data is suggested. Recognizing the importance of the second order statistics, there is a need to represent the second order statistics. A method to visualize statistics using a color code is proposed. By representing statistics using color coding, one can easily extract and compare the first and the second statistics.

Lee, Chulhee↗

Feature extraction of multispectral data

A method is presented for feature extraction of multispectral scanner data. Non-training data is used to demonstrate the reduction in processing time that can be obtained by using feature extraction rather than feature selection.

Crane, R. B.↗

Landscape analysis of environmental data sources for linkage with SEER cancer patients database

Abstract One of the challenges associated with understanding environmental impacts on cancer risk and outcomes is estimating potential exposures of individuals diagnosed with cancer to adverse environmental conditions over the life course. Historically, this has been partly due to the lack of reliable measures of cancer patients’ potential environmental exposures before a cancer diagnosis. The emerging sources of cancer-related spatiotemporal environmental data and residential history information, coupled with novel technologies for data extraction and linkage, present an opportunity to integrate these data into the existing cancer surveillance data infrastructure, thereby facilitating more comprehensive assessment of cancer risk and outcomes. In this paper, we performed a landscape analysis of the available environmental data sources that could be linked to historical residential address information of cancer patients’ records collected by the National Cancer Institute’s Surveillance, Epidemiology, and End Results Program. The objective is to enable researchers to use these data to assess potential exposures at the time of cancer initiation through the time of diagnosis and even after diagnosis. The paper addresses the challenges associated with data collection and completeness at various spatial and temporal scales, as well as opportunities and directions for future research.

60 APPLIED LIFE SCIENCES↗

The SeaWiFS Bio-Optical Archive and Storage System (SeaBASS): Current Architecture and Implementation

Satellite ocean color missions require an abundance of high-quality in situ measurements for bio-optical and atmospheric algorithm development and post-launch product validation and sensor calibration. To facilitate the assembly of a global data set, the NASA Sea-viewing Wide Field-of-view (SeaWiFS) Project developed the Seafaring Bio-optical Archive and Storage System (SeaBASS), a local repository for in situ data regularly used in their scientific analyses. The system has since been expanded to contain data sets collected by the NASA Sensor Intercalibration and Merger for Biological and Interdisciplinary Oceanic Studies (SIMBIOS) Project, as part of NASA Research Announcements NRA-96-MTPE-04 and NRA-99-OES-99. SeaBASS is a well moderated and documented hive for bio-optical data with a simple, secure mechanism for locating and extracting data based on user inputs. Its holdings are available to the general public with the exception of the most recently collected data sets. Extensive quality assurance protocols, comprehensive data and system documentation, and the continuation of an archive and relational database management system (RDBMS) suitable for bio-optical data all contribute to the continued success of SeaBASS. This document provides an overview of the current operational SeaBASS system.

Werdell, P. Jeremy↗

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL), ↗

ATS-5 solar cell experiment results after one year in synchronous orbit

The results of the ATS-5 solar cell experiment after one year in synchronous orbit are reported. A partial failure in the experimental electronics package has caused a loss of data from half the 80 experimental solar cells. Procedures for extracting data due to a partial spacecraft failure are described and discussed. Data from the remaining 40 solar cells, including 15 mounted on a thin flexible structure are analyzed. Data are corrected to a solar intensity of 140 mW/sq cm and a temperature of 25 C. It was found that after one year in synchronous orbit: (1) cells with 1.52-mm-thick coverslides did not show a clear-cut advantage over those with 0.15-mm coverslides, (2) cells with solderless grid lines are degrading at the same rate as are cells with solder-dipped grid lines, (3) cells not quite completely covered with coverslides suffered a large power loss in comparison to cells fully covered, (4) no clear-cut advantage of 10-cm cells over 2-cm cells has yet been observed, (5) cells mounted on the flexible panel with relatively little backshielding did not degrade any faster than those with substantial backshielding, and (6) the flight data in large part confirms the adequacy of the ground-based techniques used in our preflight radiation test program.

Anspaugh, B. E.↗

The Space Shuttle Orbiter approach and landing tests - A correlation of flight and predicted performance data

This paper represents the results of a program in which a flight test vehicle was flown for a limited number of flights. The vehicle was the Space Shuttle Orbiter. The total free flight time for the entire flight test program was less than half an hour. The flight regime tested represented only the approach and landing phase of the vehicle's planned flight capability. The program relied heavily on an extensive wind tunnel test program to predict the aerodynamic performance data over the complete flight range, as well as to predict data tolerances. During the flight test program, short maneuvers performed were to provide flight motion data. A data extraction program was developed to produce flight derived aerodynamic performance data in coefficient form from the motion data. The resultant flight test data was correlated with the predicted data and fell within the predicted data tolerances for all phases of the subsonic flight test regime, including ground effects.

Romere, P. O.↗

Security Vulnerability Profiles of Mission Critical Software: Empirical Analysis of Security Related Bug Reports

While some prior research work exists on characteristics of software faults (i.e., bugs) and failures, very little work has been published on analysis of software applications vulnerabilities. This paper aims to contribute towards filling that gap by presenting an empirical investigation of application vulnerabilities. The results are based on data extracted from issue tracking systems of two NASA missions. These data were organized in three datasets: Ground mission IVV issues, Flight mission IVV issues, and Flight mission Developers issues. In each dataset, we identified security related software bugs and classified them in specific vulnerability classes. Then, we created the security vulnerability profiles, i.e., determined where and when the security vulnerabilities were introduced and what were the dominating vulnerabilities classes. Our main findings include: (1) In IVV issues datasets the majority of vulnerabilities were code related and were introduced in the Implementation phase. (2) For all datasets, around 90 of the vulnerabilities were located in two to four subsystems. (3) Out of 21 primary classes, five dominated: Exception Management, Memory Access, Other, Risky Values, and Unused Entities. Together, they contributed from 80 to 90 of vulnerabilities in each dataset.

Goseva-Popstojanova, Katerina↗

Reconstructing Magma Storage Depths for the 2018 Kilauean Eruption from melt inclusion CO2 Contents: The importance of Vapor Bubbles

The 2018 Lower East Rift Zone (LERZ) eruption of Kīlauea Volcano and the accompanying collapse of the summit caldera marked the most destructive phase of activity on Hawai’i in the last 200 years. The integration of petrological data extracted from lava samples collected throughout the eruption with geodetic data examining the caldera collapse event, and estimates of the co-erupted flux of SO2 from the main eruptive fissure (Fissure 8), provides an exceptional opportunity to determine the reservoir geometry and magma transport paths supplying Kīlauea’s LERZ. The forsterite contents of erupted olivines and the degree of disequilibrium with their carrier melts indicate that two distinct olivine populations were erupted from Fissure 8. Melt inclusion entrapment pressures reveal that more evolved olivines (Fo<81.5) crystallized at ~2 km depth within the shallower Halema’uma’u reservoir, while more primitive olivines (Fo>81.5)crystallized within the deeper South Caldera reservoir at ~3–5 km depth. Crucially, primitive olivines experienced extensive post-entrapment crystallization, driving the growth of a vapor bubble. Raman spectroscopy reveals that this bubble contains up to 99% of the total inclusionCO2 budget (median=93%). Measurements of CO2 in only the glass phase would have underestimated entrapment depths by up to 60× (median=11×), and the importance of the SC reservoir as a source of magma to Fissure 8 would have been overlooked. Overall, we demonstrate that Raman measurements of bubbles, along with careful choice of suitably-calibrated H2O-CO2 solubility model, is vital to place accurate constraints on the depths of magma storage regions supplying volcanic eruptions.

SIMS↗

Procedure for extraction of disparate data from maps into computerized data bases

A procedure is presented for extracting disparate sources of data from geographic maps and for the conversion of these data into a suitable format for processing on a computer-oriented information system. Several graphic digitizing considerations are included and related to the NASA Earth Resources Laboratory's Digitizer System. Current operating procedures for the Digitizer System are given in a simplified and logical manner. The report serves as a guide to those organizations interested in converting map-based data by using a comparable map digitizing system.

Junkin, B. G.↗

Integrated Computational System for Aerodynamic Steering and Visualization

In February of 1994, an effort from the Fluid Dynamics and Information Sciences Divisions at NASA Ames Research Center with McDonnel Douglas Aerospace Company and Stanford University was initiated to develop, demonstrate, validate and disseminate automated software for numerical aerodynamic simulation. The goal of the initiative was to develop a tri-discipline approach encompassing CFD, Intelligent Systems, and Automated Flow Feature Recognition to improve the utility of CFD in the design cycle. This approach would then be represented through an intelligent computational system which could accept an engineer's definition of a problem and construct an optimal and reliable CFD solution. Stanford University's role focused on developing technologies that advance visualization capabilities for analysis of CFD data, extract specific flow features useful for the design process, and compare CFD data with experimental data. During the years 1995-1997, Stanford University focused on developing techniques in the area of tensor visualization and flow feature extraction. Software libraries were created enabling feature extraction and exploration of tensor fields. As a proof of concept, a prototype system called the Integrated Computational System (ICS) was developed to demonstrate CFD design cycle. The current research effort focuses on finding a quantitative comparison of general vector fields based on topological features. Since the method relies on topological information, grid matching and vector alignment is not needed in the comparison. This is often a problem with many data comparison techniques. In addition, since only topology based information is stored and compared for each field, there is a significant compression of information that enables large databases to be quickly searched. This report will (1) briefly review the technologies developed during 1995-1997 (2) describe current technologies in the area of comparison techniques, (4) describe the theory of our new method researched during the grant year (5) summarize a few of the results and finally (6) discuss work within the last 6 months that are direct extensions from the grant.

Hesselink, Lambertus↗

Development of VBA Tool for Document Term Search

Employees throughout different agencies such as NASA, have identified that the search of determined terms/words through documents, consume substantial research time of such. These types of searches are substantially limited towards one word in a one document identification; forward one, these usual types of searches lack efficiency & optimization through research aspects of work. Consequently, this reflects in the decrease productivity during work hours etc. The application of VBA (Visual Basic for Applications) is the programming language of Excel, which was conducted for the development of optimized tool for document term search. The project enables the search of single & multiple word/term search through single format documents for paragraph data extraction.

Ssytems Development↗

Holographic interferometric tomography for reconstructing flow fields

Holographic interferometric tomography is a technique for instantaneously capturing and quantitatively reconstructing three-dimensional flow fields. It has a very useful application potential for high-speed aerodynamics. However, three major challenging tasks need to be accomplished before its practical applications. First, fluid flows are mostly unsteady or at least non repeatable. Consequently, a means for Instantaneously recording three-dimensional flow fields, that is, a simple holographic technique for simultaneously recording multi-directional projections, needs to be developed. Second, while holographic interferometry provides enormous data storage capabilities, expeditious data extraction from complicated interferograms is very important for timely near real-time applications. Third, unlike medical applications, flow tomography does not provide complete data sets but instead involves ill-posed reconstruction problems of incomplete projection and limited angular scanning. During this summer research period, new experimental techniques and corresponding hardware were developed and tested to address the above mentioned tasks. The first task was achieved by diffuser illumination. This concept allows instantaneous capture of many projections with a conventional setup for single-projection recording. For the second task, a phase-shifting technique was incorporated. This technique allows one to acquire multiple phase-stepped interferograms for a single projection and thus to extract phase information from intensity data almost at real-time. For the third task, the research that has been extensively conducted previously was utilized. In this research period, a complete experimental setup that provides the above three major capabilities was designed, built, and tested by integrating all the techniques. A simple laboratory experiment for simulating wind-tunnel testing was then conducted. A test flow was produced by employing a relatively simple device that generated a gravity-driven flow. The flow was then experimentally investigated to check the viability of the holographic interferometric tomographic technique before wind-tunnel application.

Cha, Soyoung S.↗

Using MCC Facility Metrics to Size, Inform, and Troubleshoot

The Mission Control Center (MCC) underwent a major architecture update that has been used for Mission Operations since 2016. The MCC Performance team has collected system performance and usage metrics to improve the configuration, troubleshoot incidents, and help size the system to accommodate future programs. The data is collected through MCC custom software and custom scripts to extract data from our Commercial Off The Shelf (COTS) tools. This data has enabled MCC to support more activities concurrently, help our operations and development teams to respond to issues more quickly, and make our directorate informed buyers to meet new requirements when developing project plans for the upcoming Fiscal Year.

Data Science↗